Intern-S2-Preview: Scientific Agentic Foundation Model
A 397B-parameter scientific AI model handles text, images, time series, and long multi-step tasks in one system
Intern-S2-Preview is a 397-billion-parameter scientific agentic foundation model built to jointly understand scientific text, images, and time series while carrying out long, tool-using tasks. It goes through multimodal pre-training followed by a unified post-training pipeline combining supervised fine-tuning, multi-task reinforcement learning, agentic RL, and on-policy distillation. It reaches competitive or leading results on multiple scientific, multimodal, and agentic benchmarks, and a separate memory module lifts the Biology-Instructions average score from 56.92 to 60.32 without touching the frozen 397B backbone.
METAL MEDIA explanatory visual
Structure and training flow of Intern-S2-Preview-397B
Evidence statusMeasured results reported
- Multimodal pre-trainingLearns from rendered scientific documents, interleaved image-text data, and diverse scientific corpora to absorb both textual and visual structure
- Unified post-training pipelineSupervised fine-tuning, then multi-task reinforcement learning, then black- and white-box agentic RL, then on-policy distillation into one model
- Time-series understanding and forecasting modulesAn efficient long-sequence encoder is paired with a dedicated numerical forecasting branch for predicting future values
- Memory Decoder (separate extension)Frozen 397B backbone plus a domain-trained memory, dynamically fused by a router for targeted specialization
- Multi-benchmark evaluationTested across scientific, multimodal, agentic, general-purpose, and time-series benchmarks with competitive or leading results
What they did
- Scientific information is scattered across text, figures, tables, equations, layout, and numerical time series, so the model is built to reason over these heterogeneous forms of evidence and to keep working through tools over long task horizons rather than answering isolated questions.
- Training starts with pre-training on rendered scientific documents, interleaved image-text data, and diverse scientific corpora, then moves through a unified post-training pipeline: supervised fine-tuning, scalable multi-task reinforcement learning, black- and white-box agentic RL, and on-policy distillation to merge specialized policies into one model.
- At the architecture level, the 397B model adds a dedicated numerical forecasting branch on top of its time-series understanding encoder, and a separate 'Memory Decoder' module is studied as a way to specialize the model to new scientific domains by attaching an external memory while keeping the 397B backbone frozen.
- Across scientific, multimodal, agentic, and general-purpose benchmarks the model shows competitive or leading results in multiple settings; on the SciTS benchmark it improves time-series understanding and forecasting, and the biology-specific memory extension (Intern-MemDec-4B) raises the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.

| Provider | Collection | #Tasks | #Environments |
|---|---|---|---|
| SWE-bench | SWE-smith [93] | 59,136 | 222 |
| SWE-Gym | SWE-Gym [59] | 2,438 | 2401 |
| R2E-Gym | R2E-Gym-V1 [38] | 7,480 | 8101 |
| Nebius | SWE-rebench-V2 [6] | 32,100 | 32075 |
| AweAI-Team | Scale-SWE [104] | 20,200 | 19472 |
| NVIDIA | Nemotron-Terminal-Synthetic-Tasks [60] | 80,000 | 8 |
| RUC-AIBOX | ClawGym-Task [7] | 13,500 | 1 |

| SciTS Task ID | ASU01 | ASU03 | BIU01 | BIU03 | EAU01 | MEU01 | NEU06 | PHU01 | PHU04 | RAU01 | RAU02 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Text LLM | GPT-4.1-mini | 67.2 | 15.6 | 0.2 | 12.7 | 67.0 | 44.0 | 16.1 | 24.0 | 52.7 | 24.6 | 10.6 |
| Gemini2.5-Flash | 64.1 | 16.3 | 1.5 | 12.4 | 67.6 | 60.9 | 5.8 | 20.7 | 64.8 | 20.9 | 13.5 | |
| DeepSeek-V3 | 1.1 | 12.3 | 0.0 | 5.8 | 40.2 | 59.3 | 13.6 | 28.9 | 50.7 | 19.4 | 4.2 | |
| VL LLM | GPT-5-mini | 65.7 | 18.9 | 0.8 | 17.9 | 67.6 | 30.4 | 13.3 | 21.4 | 47.8 | 24.3 | 9.1 |
| Gemini2.5-Flash | 61.6 | 15.2 | 0.9 | 8.3 | 72.5 | 64.1 | 11.6 | 22.7 | 59.0 | 31.6 | 11.3 | |
| Intern-S1-Pro | 98.0 | 75.9 | 20.8 | 88.3 | 99.5 | 65.6 | 71.3 | 36.8 | 93.2 | - | - | |
| Intern-S2-Preview-397B | 97.1 | 91.0 | 36.5 | 98.3 | 100.0 | 81.8 | 70.2 | 66.9 | 99.9 | 88.4 | 60.2 |

| SciTS Task ID | ENG02 | ENG03 | MEG03 | NEG03 | PHG02 | URG01 | URG05 | |
|---|---|---|---|---|---|---|---|---|
| Text LLM | GPT-4.1-mini | 125.0 (1.4) | 8.3 (96.0) | 42.1 (49.6) | 95.2 (96.4) | 1.1e3 (94.2) | 320.6 (18.6) | 126.6 (100) |
| Gemini2.5-Flash | 72.5 (5.9) | 9.6 (99.0) | 62.2 (57.9) | 63.5 (99.2) | 110.8 (99.0) | 246.0 (23.3) | 98.6 (100) | |
| DeepSeek-V3 | 117.2 (46.1) | 7.7 (98.0) | 46.4 (30.9) | 4.3 (3.1) | 200.1 (92.2) | 350.0 (18.6) | 296.7 (93.0) | |
| VL LLM | GPT-5-mini | 56.1 (4.5) | 11.2 (76.0) | 37.6 (51.8) | 74.3 (97.2) | 155.3 (97.4) | 182.1 (58.1) | 71.1 (72.9) |
| Gemini2.5-Flash | 103.9 (7.4) | 15.6 (53.0) | 53.1 (37.2) | – | 185.2 (36.9) | 351.9 (16.3) | 114.6 (91.2) | |
| Time Series Models | Moirai-Large [85] | 121.2 (100) | 12.8 (100) | 51.7 (100) | 59.1 (100) | 116.9 (100) | 294.7 (100) | 74.6 (100) |
| TimeMoE-Large [68] | 70.4 (100) | 11.6 (100) | 39.0 (100) | 70.1 (100) | 80.2 (100) | 218.4 (100) | 84.4 (100) | |
| Chronos-bolt-Base [5] | 73.7 (100) | 12.0 (100) | 41.5 (100) | 78.5 (100) | 109.3 (100) | 139.3 (100) | 70.6 (100) | |
| UniTS [31] | 70.1 (100) | 12.8 (100) | 42.0 (100) | 95.2 (46.4) | 135.9 (44.1) | 389.7 (100) | – | |
| TimeOmni [86] | 68.6 (100) | 7.4 (100) | 37.5 (100) | 78.7 (100) | 163.0 (100) | 247.0 (100) | 174.0 (100) | |
| Intern-S2-Preview-397B | 60.2 (100) | 7.1 (100) | 32.8 (100) | 59.2 (100) | 72.2 (100) | 138.9 (100) | 60.6 (100) |

Findings
- Intern-S2-Preview-397B outperforms strong open- and closed-source models on Biology-Instructions (56.92), Mol-Instructions (52.37), and SciReasoner (63.97), and achieves state-of-the-art results on the internal MP20 and ProteinBinder-9 evaluation sets.
- It delivers the best results among open-source models on MolecularIQ (61.49), TOMG-Bench (65.66), XLRS-Bench (51.97), and MicroVQA (68.81), and also leads open-source models on MMLU-Pro (89.75), SimpleQA-Verified (69.90), MMMU-Pro (80.46), and ChartQAPro (69.65).
- On science-oriented agentic tasks it generally surpasses DeepSeek-V4-Pro and Qwen3.5-397B, ranking second only to GLM-5.2; on general-purpose agentic tasks it consistently outperforms Qwen3.5-397B and performs comparably to Kimi-K2.7-Code.
- On SciTS time-series understanding it consistently beats general-purpose text and vision-language LLMs, and despite having less than half the parameters of the trillion-scale Intern-S1-Pro, it surpasses that model on seven of nine shared tasks, with the F1 score on PHU01 rising from 36.8 to 66.9.
- The biology-specific Intern-MemDec-4B extension raises the Biology-Instructions average score from 56.92 to 60.32 while keeping performance close to the frozen 397B backbone on general knowledge, reasoning, and multimodal benchmarks, indicating targeted specialization without broad side effects.
Where it can be used
- Systems that need to read scientific literature combining figures, tables, equations, and text alongside plain language for question answering
- Situations where a team wants to quickly add domain-specific knowledge (e.g., biology, chemistry, materials science) without retraining the whole backbone model
- Tasks in astronomy, geoscience, neuroscience, physiological signal analysis, or radar signal analysis that require understanding long numerical time series and predicting future values
- Agentic workflows such as coding, terminal operations, and software engineering that require repeated tool use across multiple steps

Limits and open work
- The authors describe this as a preview system and note that reliability over longer scientific workflows, expansion of domain-specific memories and task environments, stronger verifiers, and deeper integration with scientific tools remain future work.
- The Memory Decoder study was evaluated only for biology as a representative domain; whether the same specialization benefits apply to other scientific fields is not verified here.
- Cross-domain evaluation of the Memory Decoder only checks that performance stays close to the frozen backbone on general benchmarks, which does not rule out subtler side effects in domains not tested.
- On the general-purpose GIFT-Eval time-series benchmark, the model achieves a zero-shot MASE of 0.785 described only as competitive, without a clear claim of surpassing specialized forecasting models there.
Why it matters
Real scientific work rarely ends with a single correct answer to a static question; it requires reading mixed evidence, using tools, and sustaining progress over many steps, and this model is an attempt to build a single system that covers that whole workflow. The idea of attaching a plug-in memory to specialize a frozen backbone also points to a practical way for teams to add domain expertise quickly without retraining or risking damage to a model's general abilities.
Terms in this paper
- agentic RL · reinforcement learning where the model is rewarded for solving multi-step tasks by using tools and interacting with an environment
- Memory Decoder · a separately trained module attached to a frozen backbone model that supplies domain-specific knowledge without changing the backbone's parameters
- on-policy distillation · a training step that merges the abilities learned by several specialized policies into a single unified model
- forecasting branch · a dedicated neural pathway that predicts future numerical values directly instead of generating them as text tokens, preserving numerical precision
- SciTS benchmark · a benchmark for evaluating how well models understand and forecast scientific time-series signals
Figures we cannot republish
- Figure 5: The pipeline of the image retrieval process, including image encoding, vector database construction, and online text-to-image and image-to-image retrieval with post-processing.
Original abstract (English)
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears
Latest from METAL MEDIA
Figures: Lei Bai et al., arXiv:2608.13505, CC BY 4.0