Intern-S2-Preview: Scientific Agentic Foundation Model
一个3970亿参数的科学AI模型能同时处理文本、图像、时间序列并完成需要多步骤操作的长任务
Intern-S2-Preview是一个3970亿参数规模的科学智能体基础模型,设计目标是让模型同时理解科学文献中的文本、图像和数值时间序列,并借助工具完成需要长时间、多步骤推进的任务。训练流程先进行多模态预训练,再经过统一的后训练流程,依次包括监督微调、多任务强化学习、智能体强化学习以及在策略蒸馏。该模型在多个科学、多模态和智能体基准上取得有竞争力或领先的结果,一个独立的记忆扩展模块在不改动被冻结的3970亿参数主干的情况下,把Biology-Instructions平均分从56.92提升到60.32。
METAL MEDIA 解读图
Intern-S2-Preview-397B的结构与训练流程
证据状态已报告实测结果
- 多模态预训练在渲染的科学文档、图文交错数据以及多样化科学语料上学习,同时吸收文本与视觉结构信息
- 统一后训练流程监督微调之后依次进行多任务强化学习、白盒与黑盒智能体强化学习,最后通过在策略蒸馏汇合为一个模型
- 时间序列理解与预测模块在高效的长序列编码器基础上加入专门的数值预测分支来预测未来数值
- Memory Decoder(独立扩展模块)主干3970亿参数模型保持冻结,搭配领域训练的记忆模块,由路由器动态融合以实现针对性专精
- 多基准综合评测在科学、多模态、智能体、通用和时间序列等多个基准上验证,取得有竞争力或领先的结果
他们做了什么
- 科学信息往往分散在文本、图表、公式、页面排版和数值时间序列等多种形式中,该模型被设计成能同时处理这些异质证据,并在长任务过程中反复使用工具,而不是只回答孤立的问题。
- 训练先在渲染的科学文档、图文交错数据以及多样化科学语料上做预训练,再经过统一的后训练流程:监督微调、可扩展的多任务强化学习、白盒与黑盒智能体强化学习,最后用在策略蒸馏把各个专门策略的能力汇聚到一个统一模型中。
- 在架构层面,3970亿参数模型在原有的时间序列理解编码器之上加入了专门的数值预测分支;此外还单独研究了一个名为Memory Decoder(记忆解码器)的模块,它在保持3970亿主干冻结的前提下,通过附加外部记忆来实现对新科学领域的快速专精。
- 在科学、多模态、智能体和通用基准的综合评测中,该模型在多个场景下取得有竞争力或领先的成绩;在SciTS基准上时间序列理解和预测能力都有提升,而专门针对生物学训练的记忆扩展模块Intern-MemDec-4B在不修改被冻结的3970亿主干的情况下,把Biology-Instructions平均分从56.92提升到60.32。

| Provider | Collection | #Tasks | #Environments |
|---|---|---|---|
| SWE-bench | SWE-smith [93] | 59,136 | 222 |
| SWE-Gym | SWE-Gym [59] | 2,438 | 2401 |
| R2E-Gym | R2E-Gym-V1 [38] | 7,480 | 8101 |
| Nebius | SWE-rebench-V2 [6] | 32,100 | 32075 |
| AweAI-Team | Scale-SWE [104] | 20,200 | 19472 |
| NVIDIA | Nemotron-Terminal-Synthetic-Tasks [60] | 80,000 | 8 |
| RUC-AIBOX | ClawGym-Task [7] | 13,500 | 1 |

| SciTS Task ID | ASU01 | ASU03 | BIU01 | BIU03 | EAU01 | MEU01 | NEU06 | PHU01 | PHU04 | RAU01 | RAU02 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Text LLM | GPT-4.1-mini | 67.2 | 15.6 | 0.2 | 12.7 | 67.0 | 44.0 | 16.1 | 24.0 | 52.7 | 24.6 | 10.6 |
| Gemini2.5-Flash | 64.1 | 16.3 | 1.5 | 12.4 | 67.6 | 60.9 | 5.8 | 20.7 | 64.8 | 20.9 | 13.5 | |
| DeepSeek-V3 | 1.1 | 12.3 | 0.0 | 5.8 | 40.2 | 59.3 | 13.6 | 28.9 | 50.7 | 19.4 | 4.2 | |
| VL LLM | GPT-5-mini | 65.7 | 18.9 | 0.8 | 17.9 | 67.6 | 30.4 | 13.3 | 21.4 | 47.8 | 24.3 | 9.1 |
| Gemini2.5-Flash | 61.6 | 15.2 | 0.9 | 8.3 | 72.5 | 64.1 | 11.6 | 22.7 | 59.0 | 31.6 | 11.3 | |
| Intern-S1-Pro | 98.0 | 75.9 | 20.8 | 88.3 | 99.5 | 65.6 | 71.3 | 36.8 | 93.2 | - | - | |
| Intern-S2-Preview-397B | 97.1 | 91.0 | 36.5 | 98.3 | 100.0 | 81.8 | 70.2 | 66.9 | 99.9 | 88.4 | 60.2 |

| SciTS Task ID | ENG02 | ENG03 | MEG03 | NEG03 | PHG02 | URG01 | URG05 | |
|---|---|---|---|---|---|---|---|---|
| Text LLM | GPT-4.1-mini | 125.0 (1.4) | 8.3 (96.0) | 42.1 (49.6) | 95.2 (96.4) | 1.1e3 (94.2) | 320.6 (18.6) | 126.6 (100) |
| Gemini2.5-Flash | 72.5 (5.9) | 9.6 (99.0) | 62.2 (57.9) | 63.5 (99.2) | 110.8 (99.0) | 246.0 (23.3) | 98.6 (100) | |
| DeepSeek-V3 | 117.2 (46.1) | 7.7 (98.0) | 46.4 (30.9) | 4.3 (3.1) | 200.1 (92.2) | 350.0 (18.6) | 296.7 (93.0) | |
| VL LLM | GPT-5-mini | 56.1 (4.5) | 11.2 (76.0) | 37.6 (51.8) | 74.3 (97.2) | 155.3 (97.4) | 182.1 (58.1) | 71.1 (72.9) |
| Gemini2.5-Flash | 103.9 (7.4) | 15.6 (53.0) | 53.1 (37.2) | – | 185.2 (36.9) | 351.9 (16.3) | 114.6 (91.2) | |
| Time Series Models | Moirai-Large [85] | 121.2 (100) | 12.8 (100) | 51.7 (100) | 59.1 (100) | 116.9 (100) | 294.7 (100) | 74.6 (100) |
| TimeMoE-Large [68] | 70.4 (100) | 11.6 (100) | 39.0 (100) | 70.1 (100) | 80.2 (100) | 218.4 (100) | 84.4 (100) | |
| Chronos-bolt-Base [5] | 73.7 (100) | 12.0 (100) | 41.5 (100) | 78.5 (100) | 109.3 (100) | 139.3 (100) | 70.6 (100) | |
| UniTS [31] | 70.1 (100) | 12.8 (100) | 42.0 (100) | 95.2 (46.4) | 135.9 (44.1) | 389.7 (100) | – | |
| TimeOmni [86] | 68.6 (100) | 7.4 (100) | 37.5 (100) | 78.7 (100) | 163.0 (100) | 247.0 (100) | 174.0 (100) | |
| Intern-S2-Preview-397B | 60.2 (100) | 7.1 (100) | 32.8 (100) | 59.2 (100) | 72.2 (100) | 138.9 (100) | 60.6 (100) |

研究结果
- Intern-S2-Preview-397B在Biology-Instructions(56.92)、Mol-Instructions(52.37)和SciReasoner(63.97)上超过了强力的开源与闭源模型,并在内部的MP20和ProteinBinder-9评测集上取得了最好成绩。
- 在MolecularIQ(61.49)、TOMG-Bench(65.66)、XLRS-Bench(51.97)和MicroVQA(68.81)上取得开源模型中的最佳成绩,同时在MMLU-Pro(89.75)、SimpleQA-Verified(69.90)、MMMU-Pro(80.46)和ChartQAPro(69.65)上也是开源模型中的最佳表现。
- 在偏科学的智能体任务上普遍超过DeepSeek-V4-Pro和Qwen3.5-397B,仅次于GLM-5.2;在通用智能体任务上持续超过Qwen3.5-397B,表现与Kimi-K2.7-Code相当。
- 在SciTS时间序列理解任务上持续超过通用文本大模型和视觉语言大模型;尽管参数量不到万亿参数级Intern-S1-Pro的一半,该模型在两者共有的九项任务中的七项上超过了它,其中PHU01任务的F1分数从36.8提升到66.9。
- 专门为生物学训练的记忆扩展模块Intern-MemDec-4B在保持3970亿主干冻结的情况下,把Biology-Instructions平均分从56.92提升到60.32,同时在通用知识、推理和多模态基准上与主干模型表现接近,说明该模块能实现有针对性的专精而不带来明显副作用。
可应用场景
- 需要同时理解科学文献中的图表、公式与文字的文献分析和问答系统
- 希望在不重新训练整个主干模型的情况下,快速为特定领域(如生物学、化学、材料科学)补充专业知识的场景
- 需要理解长时间序列数据并预测未来数值的天文学、地球科学、神经科学、生理信号分析或雷达信号分析等任务
- 需要反复调用工具、经过多步骤才能完成的编程、终端操作和软件工程类智能体任务

局限与待验证事项
- 作者将其称为预览系统,并指出在更长的科学工作流上提升可靠性、扩展领域专用记忆和任务环境、强化验证器以及深化与专业科学工具的整合都是未来工作。
- Memory Decoder的研究只以生物学作为代表领域进行评测,是否在其他科学领域同样有效尚未得到验证。
- 对Memory Decoder的跨领域评测只是检查其在通用基准上的表现是否与被冻结的主干接近,并不能排除在未测试领域中出现更细微副作用的可能。
- 在通用时间序列预测基准GIFT-Eval上,该模型取得0.785的零样本MASE,论文只称其为具有竞争力,并未明确说明是否超过专门的预测模型。
为什么重要
真正的科学研究往往不是回答一道孤立的题目,而是要在混杂的证据之间持续推理、反复使用工具、经历很多步骤才能完成任务,这个模型试图用一个系统覆盖这整个流程。把可插拔的记忆模块接到被冻结的主干模型上以实现领域专精,也为团队提供了一种在不重新训练整个模型、也不损害其通用能力的前提下快速补充专业知识的实用思路。
本文术语
- 智能体强化学习(agentic RL) · 让模型在使用工具、与环境交互并完成多步骤任务的过程中,依据奖励信号进行训练的方法
- Memory Decoder(记忆解码器) · 附加在被冻结的主干模型之上、单独训练的模块,用于提供特定领域知识而不改变主干参数
- 在策略蒸馏 · 把多个专门策略学到的能力汇总合并到一个统一模型中的训练步骤
- 数值预测分支 · 不用文本词元生成数值,而是通过独立的神经网络路径直接预测未来数值,以保持数值精度
- SciTS基准 · 用于评测模型理解和预测科学时间序列信号能力的基准测试
无法转载的图表
- Figure 5: The pipeline of the image retrieval process, including image encoding, vector database construction, and online text-to-image and image-to-image retrieval with post-processing.
论文原文摘要(英文)
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)用AI总结股市新闻发现:简单的摘要方法反而比时髦的检索增强技术更靠谱
METAL MEDIA 最新报道
图片来源: Lei Bai et al., arXiv:2608.13505, CC BY 4.0