Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

Intern-S2-Preview: Scientific Agentic Foundation Model

arXiv:2608.135052026-08-12

A 397B-parameter scientific AI model handles text, images, time series, and long multi-step tasks in one system

Intern-S2-Preview is a 397-billion-parameter scientific agentic foundation model built to jointly understand scientific text, images, and time series while carrying out long, tool-using tasks. It goes through multimodal pre-training followed by a unified post-training pipeline combining supervised fine-tuning, multi-task reinforcement learning, agentic RL, and on-policy distillation. It reaches competitive or leading results on multiple scientific, multimodal, and agentic benchmarks, and a separate memory module lifts the Biology-Instructions average score from 56.92 to 60.32 without touching the frozen 397B backbone.

METAL MEDIA explanatory visual

Structure and training flow of Intern-S2-Preview-397B

Evidence statusMeasured results reported

  1. Multimodal pre-trainingLearns from rendered scientific documents, interleaved image-text data, and diverse scientific corpora to absorb both textual and visual structure
  2. Unified post-training pipelineSupervised fine-tuning, then multi-task reinforcement learning, then black- and white-box agentic RL, then on-policy distillation into one model
  3. Time-series understanding and forecasting modulesAn efficient long-sequence encoder is paired with a dedicated numerical forecasting branch for predicting future values
  4. Memory Decoder (separate extension)Frozen 397B backbone plus a domain-trained memory, dynamically fused by a router for targeted specialization
  5. Multi-benchmark evaluationTested across scientific, multimodal, agentic, general-purpose, and time-series benchmarks with competitive or leading results
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. Scientific information is scattered across text, figures, tables, equations, layout, and numerical time series, so the model is built to reason over these heterogeneous forms of evidence and to keep working through tools over long task horizons rather than answering isolated questions.
  2. Training starts with pre-training on rendered scientific documents, interleaved image-text data, and diverse scientific corpora, then moves through a unified post-training pipeline: supervised fine-tuning, scalable multi-task reinforcement learning, black- and white-box agentic RL, and on-policy distillation to merge specialized policies into one model.
  3. At the architecture level, the 397B model adds a dedicated numerical forecasting branch on top of its time-series understanding encoder, and a separate 'Memory Decoder' module is studied as a way to specialize the model to new scientific domains by attaching an external memory while keeping the 397B backbone frozen.
  4. Across scientific, multimodal, agentic, and general-purpose benchmarks the model shows competitive or leading results in multiple settings; on the SciTS benchmark it improves time-series understanding and forecasting, and the biology-specific memory extension (Intern-MemDec-4B) raises the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.
Figure 1: Architecture of the separate Memory Decoder extension for Intern-S2-Preview-397B. The frozen Intern-S2-Preview-397B backbone and a domain memory process the same input in parallel and produce separate next-token distributions. A lightweight token-level router uses their hidden states and output-distribution uncertainty features to predict a dynamic fusion weight λ, which controls the contribution of the two distributions to the final prediction.
Figure 1: Architecture of the separate Memory Decoder extension for Intern-S2-Preview-397B. The frozen Intern-S2-Preview-397B backbone and a domain memory process the same input in parallel and produce separate next-token distributions. A lightweight token-level router uses their hidden states and output-distribution uncertainty features to predict a dynamic fusion weight λ, which controls the contribution of the two distributions to the final prediction.
Figure 2: Architecture of the time series modules for long-sequence understanding and numerical forecasting.
Figure 2: Architecture of the time series modules for long-sequence understanding and numerical forecasting.
Table 1: Public sources used to construct executable coding and terminal tasks.
ProviderCollection#Tasks#Environments
SWE-benchSWE-smith [93]59,136222
SWE-GymSWE-Gym [59]2,4382401
R2E-GymR2E-Gym-V1 [38]7,4808101
NebiusSWE-rebench-V2 [6]32,10032075
AweAI-TeamScale-SWE [104]20,20019472
NVIDIANemotron-Terminal-Synthetic-Tasks [60]80,0008
RUC-AIBOXClawGym-Task [7]13,5001
(b) Structure of the time series forecaster.
(b) Structure of the time series forecaster.
Figure 3: Overview of matched text and visual pre-training. The text pathway predicts tokens from parsed PDF content, whereas the visual pathway predicts foreground visual latents from rendered pages, improving alignment between textual and visual document representations.
Figure 3: Overview of matched text and visual pre-training. The text pathway predicts tokens from parsed PDF content, whereas the visual pathway predicts foreground visual latents from rendered pages, improving alignment between textual and visual document representations.
Table 4: Results of time series understanding on SciTS benchmark. F1 scores are reported. Higher F1 indicates better performance.
SciTS Task IDASU01ASU03BIU01BIU03EAU01MEU01NEU06PHU01PHU04RAU01RAU02
Text LLMGPT-4.1-mini67.215.60.212.767.044.016.124.052.724.610.6
Gemini2.5-Flash64.116.31.512.467.660.95.820.764.820.913.5
DeepSeek-V31.112.30.05.840.259.313.628.950.719.44.2
VL LLMGPT-5-mini65.718.90.817.967.630.413.321.447.824.39.1
Gemini2.5-Flash61.615.20.98.372.564.111.622.759.031.611.3
Intern-S1-Pro98.075.920.888.399.565.671.336.893.2--
Intern-S2-Preview-397B97.191.036.598.3100.081.870.266.999.988.460.2
Figure 4: Pipeline for producing the interleaved image-text pair data from PDF documents, including OCR and layout parsing, visual-unit cropping, visual-gain filtering, and document-level sequence assembly.
Figure 4: Pipeline for producing the interleaved image-text pair data from PDF documents, including OCR and layout parsing, visual-unit cropping, visual-gain filtering, and document-level sequence assembly.
Figure 6: Overview of the post-training pipeline for Intern-S2-Preview. The pretrained base model is first enhanced through supervised fine-tuning, followed by multi-task RLVR and black-box agentic RL for general and specialized capability improvement. On-policy distillation then consolidates the resulting scientific reasoning and agentic capabilities into a single unified model.
Figure 6: Overview of the post-training pipeline for Intern-S2-Preview. The pretrained base model is first enhanced through supervised fine-tuning, followed by multi-task RLVR and black-box agentic RL for general and specialized capability improvement. On-policy distillation then consolidates the resulting scientific reasoning and agentic capabilities into a single unified model.
Table 5: Results of time series forecasting on the SciTS benchmark, reported in the format MAPE (success rate %). Lower MAPE indicates better performance, while higher success rate is better.
SciTS Task IDENG02ENG03MEG03NEG03PHG02URG01URG05
Text LLMGPT-4.1-mini125.0 (1.4)8.3 (96.0)42.1 (49.6)95.2 (96.4)1.1e3 (94.2)320.6 (18.6)126.6 (100)
Gemini2.5-Flash72.5 (5.9)9.6 (99.0)62.2 (57.9)63.5 (99.2)110.8 (99.0)246.0 (23.3)98.6 (100)
DeepSeek-V3117.2 (46.1)7.7 (98.0)46.4 (30.9)4.3 (3.1)200.1 (92.2)350.0 (18.6)296.7 (93.0)
VL LLMGPT-5-mini56.1 (4.5)11.2 (76.0)37.6 (51.8)74.3 (97.2)155.3 (97.4)182.1 (58.1)71.1 (72.9)
Gemini2.5-Flash103.9 (7.4)15.6 (53.0)53.1 (37.2)185.2 (36.9)351.9 (16.3)114.6 (91.2)
Time Series ModelsMoirai-Large [85]121.2 (100)12.8 (100)51.7 (100)59.1 (100)116.9 (100)294.7 (100)74.6 (100)
TimeMoE-Large [68]70.4 (100)11.6 (100)39.0 (100)70.1 (100)80.2 (100)218.4 (100)84.4 (100)
Chronos-bolt-Base [5]73.7 (100)12.0 (100)41.5 (100)78.5 (100)109.3 (100)139.3 (100)70.6 (100)
UniTS [31]70.1 (100)12.8 (100)42.0 (100)95.2 (46.4)135.9 (44.1)389.7 (100)
TimeOmni [86]68.6 (100)7.4 (100)37.5 (100)78.7 (100)163.0 (100)247.0 (100)174.0 (100)
Intern-S2-Preview-397B60.2 (100)7.1 (100)32.8 (100)59.2 (100)72.2 (100)138.9 (100)60.6 (100)
Figure 7: Overview of our co-located partial-rollout system based on XTuner and LMDeploy. During rollout, completed requests are continuously replaced to keep the inference engine fully utilized. Once sufficient completed trajectories have been collected for one training batch, the remaining in-flight rollouts are paused and their generated prefixes are retained. The same GPUs then switch to policy training. After the training states are offloaded and the updated model weights are synchronized to the inference engine, the paused requests resume generation from their retained prefixes.
Figure 7: Overview of our co-located partial-rollout system based on XTuner and LMDeploy. During rollout, completed requests are continuously replaced to keep the inference engine fully utilized. Once sufficient completed trajectories have been collected for one training batch, the remaining in-flight rollouts are paused and their generated prefixes are retained. The same GPUs then switch to policy training. After the training states are offloaded and the updated model weights are synchronized to the inference engine, the paused requests resume generation from their retained prefixes.
Figure 8: Comparison of training with and without adaptive length penalty, showing similar reward curves and shorter outputs with regularization.
Figure 8: Comparison of training with and without adaptive length penalty, showing similar reward curves and shorter outputs with regularization.

Findings

  • Intern-S2-Preview-397B outperforms strong open- and closed-source models on Biology-Instructions (56.92), Mol-Instructions (52.37), and SciReasoner (63.97), and achieves state-of-the-art results on the internal MP20 and ProteinBinder-9 evaluation sets.
  • It delivers the best results among open-source models on MolecularIQ (61.49), TOMG-Bench (65.66), XLRS-Bench (51.97), and MicroVQA (68.81), and also leads open-source models on MMLU-Pro (89.75), SimpleQA-Verified (69.90), MMMU-Pro (80.46), and ChartQAPro (69.65).
  • On science-oriented agentic tasks it generally surpasses DeepSeek-V4-Pro and Qwen3.5-397B, ranking second only to GLM-5.2; on general-purpose agentic tasks it consistently outperforms Qwen3.5-397B and performs comparably to Kimi-K2.7-Code.
  • On SciTS time-series understanding it consistently beats general-purpose text and vision-language LLMs, and despite having less than half the parameters of the trillion-scale Intern-S1-Pro, it surpasses that model on seven of nine shared tasks, with the F1 score on PHU01 rising from 36.8 to 66.9.
  • The biology-specific Intern-MemDec-4B extension raises the Biology-Instructions average score from 56.92 to 60.32 while keeping performance close to the frozen 397B backbone on general knowledge, reasoning, and multimodal benchmarks, indicating targeted specialization without broad side effects.
Figure 9: Overview of our agentic RL infrastructure. Heterogeneous white-box and black-box agents are unified by the Agent Gateway and execute against a shared sandbox and model-serving substrate. Semantic trajectories and verifier feedback are retained in the Replay Buffer, while the Rollout Trace Store preserves exact token-level evidence through a per-session incremental PrefixTree. Experience assembly aligns the two views for advantage estimation and policy optimization.
Figure 9: Overview of our agentic RL infrastructure. Heterogeneous white-box and black-box agents are unified by the Agent Gateway and execute against a shared sandbox and model-serving substrate. Semantic trajectories and verifier feedback are retained in the Replay Buffer, while the Rollout Trace Store preserves exact token-level evidence through a per-session incremental PrefixTree. Experience assembly aligns the two views for advantage estimation and policy optimization.
Figure 10: Self-evolving construction of general agentic tasks. Community skills are filtered and composed through compatible skill-state paths before stage-wise synthesis produces validated task bundles. Online and offline rollouts yield curated reusable trajectories, while execution failures update sampling and synthesis components to generate the next task distribution.
Figure 10: Self-evolving construction of general agentic tasks. Community skills are filtered and composed through compatible skill-state paths before stage-wise synthesis produces validated task bundles. Online and offline rollouts yield curated reusable trajectories, while execution failures update sampling and synthesis components to generate the next task distribution.

Where it can be used

  • Systems that need to read scientific literature combining figures, tables, equations, and text alongside plain language for question answering
  • Situations where a team wants to quickly add domain-specific knowledge (e.g., biology, chemistry, materials science) without retraining the whole backbone model
  • Tasks in astronomy, geoscience, neuroscience, physiological signal analysis, or radar signal analysis that require understanding long numerical time series and predicting future values
  • Agentic workflows such as coding, terminal operations, and software engineering that require repeated tool use across multiple steps
Figure 11: Reward trajectories across SWE, general-purpose, and terminal tasks under multiple agent harnesses. Each panel presents a representative example over 160 optimization steps. Curves are locally smoothed to highlight the overall optimization trends.
Figure 11: Reward trajectories across SWE, general-purpose, and terminal tasks under multiple agent harnesses. Each panel presents a representative example over 160 optimization steps. Curves are locally smoothed to highlight the overall optimization trends.

Limits and open work

  • The authors describe this as a preview system and note that reliability over longer scientific workflows, expansion of domain-specific memories and task environments, stronger verifiers, and deeper integration with scientific tools remain future work.
  • The Memory Decoder study was evaluated only for biology as a representative domain; whether the same specialization benefits apply to other scientific fields is not verified here.
  • Cross-domain evaluation of the Memory Decoder only checks that performance stays close to the frozen backbone on general benchmarks, which does not rule out subtler side effects in domains not tested.
  • On the general-purpose GIFT-Eval time-series benchmark, the model achieves a zero-shot MASE of 0.785 described only as competitive, without a clear claim of surpassing specialized forecasting models there.

Why it matters

Real scientific work rarely ends with a single correct answer to a static question; it requires reading mixed evidence, using tools, and sustaining progress over many steps, and this model is an attempt to build a single system that covers that whole workflow. The idea of attaching a plug-in memory to specialize a frozen backbone also points to a practical way for teams to add domain expertise quickly without retraining or risking damage to a model's general abilities.

Terms in this paper

  • agentic RL · reinforcement learning where the model is rewarded for solving multi-step tasks by using tools and interacting with an environment
  • Memory Decoder · a separately trained module attached to a frozen backbone model that supplies domain-specific knowledge without changing the backbone's parameters
  • on-policy distillation · a training step that merges the abilities learned by several specialized policies into a single unified model
  • forecasting branch · a dedicated neural pathway that predicts future numerical values directly instead of generating them as text tokens, preserving numerical precision
  • SciTS benchmark · a benchmark for evaluating how well models understand and forecast scientific time-series signals

Figures we cannot republish

  • Figure 5: The pipeline of the image retrieval process, including image encoding, vector database construction, and online text-to-image and image-to-image retrieval with post-processing.
See the figures in the original paper →

Original abstract (English)

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.

Authors · Lei Bai

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Lei Bai et al., arXiv:2608.13505, CC BY 4.0