Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

arXiv:2608.135582026-08-12

An AI scientist that reads raw lab data instead of pre-made summaries produces measurably better papers

OmniScientist is an AI system that runs full research cycles - forming a hypothesis, running experiments, and writing a paper - while directly inspecting raw scientific evidence such as seismograms, pathology slides, or 3-D models instead of relying on pre-computed feature summaries. Across 36 real-data cases spanning 5 discipline families, it completed every case end-to-end and scored 6.3 out of 10 on average with its main backbone model. When compared against a 'blind' version that only sees precomputed numbers, the version with direct perception won 85% of head-to-head comparisons across all 7 evaluation dimensions.

METAL MEDIA explanatory visual

OmniScientist의 3단계 파이프라인과 지각 레이어

Evidence statusMeasured results reported

  1. 지각 레이어이미지, 신호, 3D 구조, 표, 그래프 등 원본 증거를 4개 계열(지각/기호/정량통계/절차)로 분류해 직접 읽어들인다
  2. 아이디어 구상(Ideation)자료를 관찰하고 문헌을 검색해 검증 가능한 가설을 세우며, 코드 기반 체크가 구조적 완전성과 참신성을 검증한다
  3. 실험(Experiment)코드를 실행해 가설을 검증하고, 실행 출처와 통계적 타당성을 검증하는 엄격성 체크를 통과해야 한다
  4. 작성(Writeup)검증된 실행 기록에서만 주장을 골라 학문 분야별 서식에 맞춰 원고를 작성하고, 주장 검증 체크로 숫자와 진술을 원본 실행 기록과 대조한다
  5. 블라인드 비교 실험원본 데이터 대신 사전 계산된 스칼라 특징만 받는 시스템과 쌍대 비교해 지각의 기여도를 측정한다
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. OmniScientist combines a perception layer with three agents (ideation, experiment, writeup) that operate in a controlled pipeline, letting raw observations shape the research question, the experiment design, and the final claims.
  2. The system organizes evidence into 4 families - perceptual (images, spectra), symbolic (text, formulae, graphs), quantitative-statistical (tables), and procedural (trajectories, simulations) - and applies code-enforced checks for novelty, statistical validity, and traceability of every number reported.
  3. It was tested on a 36-case suite covering 5 discipline families and modalities like images, audio, video, 3-D structures, and knowledge graphs, completing all 36 cases from raw data to a compiled manuscript.
  4. A paired comparison against a 'blind' variant fed only precomputed scalar features showed that direct perception improved every one of 7 review dimensions and won 85% of head-to-head judgments.
  5. Two case studies illustrate the effect: the system found that 21.7% of 'noise'-labelled seismic traces were actually real earthquake signals, and it discovered a texture pattern (local entropy) in chest X-rays that distinguishes pneumonia from healthy lungs with an AUC up to 0.851, beating a raw-pixel baseline of 0.634.
Figure 2: Progression from raw evidence to verified findings across three demonstration cases. The top three rows track workflows in seismology, pathology, and 3-D CAD. From left to right, each row begins with raw evidence (a three-channel seismogram, a stained pathology tile, and a 3-D CAD model), identifies specific structural cues, and outlines the subsequent hypothesis and action sequence. The rightmost column displays the verified findings, such as the discovery that 21.7% of noise labels are real events. The bottom band depicts a precomputed interface where the artifact is reduced to a feature vector, resulting in lost structural relations and a narrower research question space.
Figure 2: Progression from raw evidence to verified findings across three demonstration cases. The top three rows track workflows in seismology, pathology, and 3-D CAD. From left to right, each row begins with raw evidence (a three-channel seismogram, a stained pathology tile, and a 3-D CAD model), identifies specific structural cues, and outlines the subsequent hypothesis and action sequence. The rightmost column displays the verified findings, such as the discovery that 21.7% of noise labels are real events. The bottom band depicts a precomputed interface where the artifact is reduced to a feature vector, resulting in lost structural relations and a narrower research question space.
Table 1: The demonstration suite: 5 categories, 36 cases, one real downloadable dataset each, with the sample count N per dataset. Evidence – perceptual, symbolic, quantitative-statistical, procedural. Modality – image, signal, spectrum, audio, video, 3-D, trajectory, table, formula, sequence, field, graph.
DisciplineRepresentative datasetEvidenceModalityN
Physical sciences5 cases
Condensed matter / nanoNFFA-EUROPE (Aversa et al. 2018)2,655
Vibrational spectroscopyRRUFF (Lafuente et al. 2015)2,000
Materials informaticsUCI superconductor (Hamidieh 2018)21,263
Molecular chemistryPubChem (Kim et al. 2025)30
Symbolic regressionFeynman (Udrescu and Tegmark 2020)12
Earth & space9 cases
Remote sensingEuroSAT (Helber et al. 2019)5,000
Galaxy morphologyGalaxy Zoo (Lintott et al. 2008)1,000
Galaxy cross-surveyGZ DECaLS (Walmsley et al. 2022)210
Gravitational wavesGWOSC (LIGO-Virgo Collaboration 2021)1,500
SeismologySTEAD (Mousavi et al. 2019)1,500
Marine biologyWHOI-Plankton (Orenstein et al. 2015)3,000
Geology / petrophysicsDigital Rocks (Prodanović et al. 2015)375
MeteorologySEVIR (Veillette et al. 2020)384
Cyclone dynamicsIBTrACS (Knapp et al. 2010)400
Life & medical7 cases
PathologyKather CRC (Kather et al. 2016)5,000
RadiologyChest X-ray (Kermany et al. 2018)3,000
Medical imagingMedMNIST CT (Yang et al. 2023a)1,496
CardiologyCinC 2016 (Liu et al. 2016)2,000
Sleep neuroscienceSleep-EDF (Kemp et al. 2000)1,520
Cell biologyCell Tracking Ch. (Ulman et al. 2017)280
GenomicsDNA (H3) (Nguyen et al. 2016)10,000
Agricultural & ecological8 cases
Plant pathologyPlantVillage (Hughes and Salathé 2015)3,002
Precision agricultureIndian Pines (Baumgardner et al. 2015)2,000
Animal behaviorCalMS21 (Sun et al. 2021)2,500
EcoacousticsBird Audio Det. (Stowell et al. 2019)2,000
Marine bioacousticsWatkins MMSD (Sayigh et al. 2016)1,697
Figure 3: Architecture of the OmniScientist framework. At the top, raw evidence from multiple disciplines enters the system, categorized into four evidence families (perceptual, symbolic, quantitative, and procedural) and 12 modalities. The core pipeline consists of three sequential stages. First, the Ideation stage (left) observes materials, searches literature, and formulates falsifiable hypotheses. Next, the Experiment stage (center) designs tests, executes code, and inspects results to generate an execution record containing standard output, figures, data, and configurations. Finally, the Writeup stage (right) selects, grounds, and reports claims supported exclusively by the execution record to compile the final paper. At the bottom, a lifecycle-wide perception layer provides spatial, temporal, cross-channel, statistical, and dynamic analysis capabilities. Dashed arrows indicate that these perception tools are available to all three stages of the pipeline.
Figure 3: Architecture of the OmniScientist framework. At the top, raw evidence from multiple disciplines enters the system, categorized into four evidence families (perceptual, symbolic, quantitative, and procedural) and 12 modalities. The core pipeline consists of three sequential stages. First, the Ideation stage (left) observes materials, searches literature, and formulates falsifiable hypotheses. Next, the Experiment stage (center) designs tests, executes code, and inspects results to generate an execution record containing standard output, figures, data, and configurations. Finally, the Writeup stage (right) selects, grounds, and reports claims supported exclusively by the execution record to compile the final paper. At the bottom, a lifecycle-wide perception layer provides spatial, temporal, cross-channel, statistical, and dynamic analysis capabilities. Dashed arrows indicate that these perception tools are available to all three stages of the pipeline.
Table 2: The 4 families of scientific evidence OmniScientist is designed to perceive.
Evidence familyTypical artifacts
PerceptualImages, video, micrographs, radar, astronomical and remote-sensing imagery, the visual form of scientific plots, audio, and 3-D structure.
SymbolicNatural-language documents, formulae, variables, rules, sequences, knowledge graphs, logical and causal relations, mathematical models.
Quantitative-statisticalTables, measurements, distributions, curves, correlations, significance tests, regression results.
Procedural / dynamicExperimental steps, code execution, agent traces, simulations, dynamic evolution, protocols.
Figure 4: Raw observations and derived discoveries processed by the perception layer across 16 cases, 11 modalities, and all 4 evidence families. The figure presents a four-by-four grid of artifacts, each taken from that case’s own data exactly as the run received it, with the discipline named at the top left of every panel and the modality at the top right. Within each artifact, a red bounding box marks the specific feature flagged by the agent on the raw record. Below the artifact, the saw label reports the direct observation made by the agent, and the found label details the verified experimental result produced by that observation. The first three rows cover perceptual and procedural evidence, spanning images, spectra, signals, audio, video, three-dimensional structures, and trajectories. The bottom row presents quantitative and symbolic evidence, where the layer reads the native numeric structure of a table, a formula, a sequence, or a graph instead of rendering an image.
Figure 4: Raw observations and derived discoveries processed by the perception layer across 16 cases, 11 modalities, and all 4 evidence families. The figure presents a four-by-four grid of artifacts, each taken from that case’s own data exactly as the run received it, with the discipline named at the top left of every panel and the modality at the top right. Within each artifact, a red bounding box marks the specific feature flagged by the agent on the raw record. Below the artifact, the saw label reports the direct observation made by the agent, and the found label details the verified experimental result produced by that observation. The first three rows cover perceptual and procedural evidence, spanning images, spectra, signals, audio, video, three-dimensional structures, and trajectories. The bottom row presents quantitative and symbolic evidence, where the layer reads the native numeric structure of a table, a formula, a sequence, or a graph instead of rendering an image.
Table 3: Detailed review scores across reasoning backbones. Per-dimension means (0–10) are derived from a 2-judge cross-family panel (deepseek-v4-flash and gemini-2.5-flash-lite) over the entire case suite. For these evaluations, the framework and perception models are held fixed, with only the reasoning backbone swapped. Failed runs are excluded; thus, means are computed exclusively over successfully scored papers. The highest value in each column is highlighted, and coverage per backbone is detailed in Table 4. Notably, clarity exhibits the least degradation, whereas factual accuracy and soundness most closely track the underlying backbone strength.
Standard peer-reviewMM-mandatory
BackboneNovelty↑Sound.↑Clarity↑Signif.↑Reprod.↑MM-grnd↑Factual↑Overall↑
Sonnet 5 (Anthropic 2026)6.37.07.06.36.15.17.76.3
GPT-5.6 (OpenAI 2026)5.26.36.35.05.24.27.75.6
GLM-5.2 (Zhipu 2026)6.27.16.86.45.96.67.56.5
Kimi K2.7 (Kimi 2025)6.27.26.76.25.55.88.06.2
Qwen3.5-122B (Qwen Team 2026)4.75.56.24.84.84.86.55.1
Qwen3.5-27B (Qwen Team 2026)5.05.65.94.94.64.96.45.1
Qwen3.5-9B (Qwen Team 2026)4.04.14.83.73.73.94.84.0
Gemma-4-31B (Google 2026)4.75.05.64.54.44.66.54.8
Gemma-4-26B (Google 2026)4.44.45.04.03.73.85.14.2
Figure 5: Two-stage verification pipeline for experimental results and manuscript claims. In the top row from left to right, an unverified experimental result undergoes a rigour check that verifies real execution, accounts for all tests, tests for independence and leakage, and ensures the headline belongs to supported analyses. If a check fails, unsupported analyses are traced and null results trigger re-ideation. Successful validation yields a verified result with certified metrics and attached provenance. Further right, a claim check matches reported numbers (n1​…​nk) to recorded outputs and reported claims (C1​…​Cm) to recorded analyses (E1​…​Em), resulting in a manuscript with fully traced numbers and supported claims. The bottom band displays the execution record, which serves as the source of truth for both checks. This record captures data I/O, standard output, generated figures, and a complete list of all attempted tests including unsupported attempts (Tu).
Figure 5: Two-stage verification pipeline for experimental results and manuscript claims. In the top row from left to right, an unverified experimental result undergoes a rigour check that verifies real execution, accounts for all tests, tests for independence and leakage, and ensures the headline belongs to supported analyses. If a check fails, unsupported analyses are traced and null results trigger re-ideation. Successful validation yields a verified result with certified metrics and attached provenance. Further right, a claim check matches reported numbers (n1​…​nk) to recorded outputs and reported claims (C1​…​Cm) to recorded analyses (E1​…​Em), resulting in a manuscript with fully traced numbers and supported claims. The bottom band displays the execution record, which serves as the source of truth for both checks. This record captures data I/O, standard output, generated figures, and a complete list of all attempted tests including unsupported attempts (Tu).
Table 4: Backbone generality across the 36-case suite. The table reports the number of cases dispatched, the resulting completed papers, and the mean composite score for these successful runs.
BackboneSonnet 5Sonnet5GPT 5.6GPT5.6GLM 5.2GLM5.2Kimi K2.7KimiK2.7Qwen3.5 122BQwen3.5122BQwen3.5 27BQwen3.527BQwen3.5 9B
Sonnet
5
GPT
5.6
GLM
5.2
Kimi
K2.7
Qwen3.5
122B
Qwen3.5
27B
Qwen3.5
9B
Gemma-4
31B
Gemma-4
26B
Cases36101893436323634
Completed↑3691763032183225
Mean↑6.55.76.76.55.45.34.15.04.3
Figure 6: Per-case review profiles across the 7 dimensions. Radar plots are shown for 9 high-coverage cases spanning 4 evidence modalities; each line represents one backbone, scored by a 2-judge panel (on a 0–10 scale). The strong backbones (Sonnet 5, GLM, Kimi) exhibit the largest, most balanced profiles, while the weak open models (Qwen3.5-9B, Gemma-4-26B) collapse inward, particularly in novelty and significance, although clarity varies the least across all models.
Figure 6: Per-case review profiles across the 7 dimensions. Radar plots are shown for 9 high-coverage cases spanning 4 evidence modalities; each line represents one backbone, scored by a 2-judge panel (on a 0–10 scale). The strong backbones (Sonnet 5, GLM, Kimi) exhibit the largest, most balanced profiles, while the weak open models (Qwen3.5-9B, Gemma-4-26B) collapse inward, particularly in novelty and significance, although clarity varies the least across all models.
Table 5: Backbone quality aggregated by evidence modality and discipline family. Results are grouped by modality on the left and by discipline family on the right. The reported metric is the mean composite score evaluated by the 2-judge cross-family panel on a scale of 0 to 10. A dash indicates that a backbone produced no scored papers for that specific category.
By evidence modalityBy discipline
BackboneImageSignalAudioVideo3-DTraj.T&SEarthLifeAgri.Engin.Phys.
Sonnet 56.46.17.16.47.06.36.46.56.76.86.65.8
GPT-5.65.95.55.55.95.66.06.45.2
Qwen3.5-27B5.55.05.95.34.84.95.85.15.55.54.85.5
Qwen3.5-9B3.85.54.14.54.44.33.24.14.54.04.43.2
Qwen3.5-122B4.85.45.95.25.45.85.54.65.76.04.95.4
Gemma-4-31B5.34.95.45.15.94.03.55.15.74.65.04.9
Gemma-4-26B4.14.34.45.04.34.34.84.53.84.34.44.6
GLM-5.26.66.86.66.66.46.96.56.47.5
Kimi K2.76.86.75.96.06.86.26.66.0
Figure 7: Dimension-wise perception gain. For the 5 cases evaluated under both conditions, the chart shows the mean scores with perception removed (pink) and for the full OmniScientist (teal), scored by the same judge across both settings. The 2 panels of this row share one colour key. The largest gain is observed in multimodal grounding, while factual accuracy remains identical since both conditions undergo the same provenance check.
Figure 7: Dimension-wise perception gain. For the 5 cases evaluated under both conditions, the chart shows the mean scores with perception removed (pink) and for the full OmniScientist (teal), scored by the same judge across both settings. The 2 panels of this row share one colour key. The largest gain is observed in multimodal grounding, while factual accuracy remains identical since both conditions undergo the same provenance check.
Table 6: Review rubric performance across the complete evaluation suite using the Sonnet 5 backbone. The 2-judge cross-family panel evaluated all cases on a scale of 0 to 10. Using a single backbone ensures direct comparability across all columns. Factual accuracy reaches 7.0 or higher in 30 of the 36 completed cases.
Standard peer-reviewMM-mandatory
DisciplineNovelty↑Sound.↑Clarity↑Signif.↑Reprod.↑MM-gr.↑Factual↑Overall↑
Physical sciences
Condensed matter2.52.54.02.02.04.51.02.5
Vibrational spectroscopy6.08.07.57.07.06.09.07.0
Materials informatics7.07.07.58.07.04.57.07.0
Molecular chemistry4.55.56.03.56.04.05.04.5
Symbolic regression7.08.07.57.57.53.59.07.0
Earth & space
Remote sensing7.07.06.57.06.05.57.06.5
Galaxy morphology6.56.57.06.06.05.57.56.5
Galaxy cross-survey7.07.06.56.06.05.07.56.5
Gravitational waves5.05.05.55.05.03.55.54.5
Seismology6.58.07.07.56.54.58.57.0
Marine biology6.08.57.57.06.55.59.57.0
Geology / petrophysics6.57.58.07.57.56.09.57.5
Cyclone dynamics6.06.57.06.05.54.57.06.0
Meteorology6.07.08.06.05.04.58.06.5
Life & medical
Pathology7.08.08.06.56.56.59.07.0
Radiology7.07.58.06.56.56.58.57.0
Medical imaging (CT)6.57.56.56.07.04.09.06.0
Cardiology6.58.08.07.56.54.59.07.0
Sleep neuroscience6.54.57.06.05.53.54.04.5
Genomics7.07.07.06.06.05.07.56.5
Cell biology7.06.05.56.55.54.06.05.5
Agricultural & ecological
Plant pathology6.57.58.07.06.56.08.07.0
Precision agriculture6.06.56.56.06.05.08.06.0
Animal behavior6.57.57.05.55.54.58.06.5
Ecoacoustics7.08.08.07.06.55.09.07.5
Figure 8: Breakdown of head-to-head judgments. For each dimension, the 3 bars show the share of all judgments won by OmniScientist, won by the same system with perception removed, and declared a tie. Judges never tie on novelty or significance, the dimensions enhanced by perception, and tie most often on factual accuracy and reproducibility, which both conditions share through the provenance check.
Figure 8: Breakdown of head-to-head judgments. For each dimension, the 3 bars show the share of all judgments won by OmniScientist, won by the same system with perception removed, and declared a tie. Judges never tie on novelty or significance, the dimensions enhanced by perception, and tie most often on factual accuracy and reproducibility, which both conditions share through the provenance check.
Table 7: Comparison of feature utilization between the two systems. In all paired cases, the perceiving system focuses on information inherent in the raw records, whereas the blind system relies solely on the provided scalar features despite sharing the same task and backbone.
CaseEvidence only the raw record carriesQuestion each system asked
Galaxy cross-surveyMorphology read off the imageWith perception: does a vision model’s morphological reading degrade on the shallower survey for the same galaxies? Blind: can the 3 classes be separated along the 8 supplied feature axes?
SeismologyCross-component polarisation of the waveformWith perception: what fraction of noise-labelled traces carry coherent polarised transients? Blind: do frequency-shape features retain a depth imprint after an attenuation correction?
PathologyTexture and nuclear density of the tileWith perception: is the complex class a compositional mixture of the pure tissue prototypes? Blind: does the class confusion matrix follow an a priori similarity ranking?
Mechanical CADPrincipal-axis geometry of the point cloudWith perception: do the function-defined labels correspond to latent geometric morphotypes? Blind: do dimension-standardised part families show tighter descriptor dispersion?
Plant phenotypingPer-point organ labels across repeated scansWith perception: do the two species differ in how many leaves grow at once? Blind: do the species separate on shape descriptors once the size axis is removed?
Figure 9: Component ablation on the seismology case, where each configuration removes a single component with the backbone fixed. Novelty is the judged novelty score and Composite the 7-dimension mean, both evaluated by the DS-V4-Flash judge (0–10). The dashed line marks the composite score of the full system.
Figure 9: Component ablation on the seismology case, where each configuration removes a single component with the backbone fixed. Novelty is the judged novelty score and Composite the 7-dimension mean, both evaluated by the DS-V4-Flash judge (0–10). The dashed line marks the composite score of the full system.
Table 8: Overview of 13 end-to-end runs, grouped by evidence modality and spanning all 4 families. Evaluation metrics are system-specific, extracted verbatim from verified experiment records.
DisciplineEvidenceEvaluation metricHeadline finding
RadiologyimagesupportedPneumonic pediatric lung fields show markedly higher local-entropy heterogeneity (patchiness) than normal.
PathologyimagesupportedThe COMPLEX H&E class is heterogeneous, splitting into compositional sub-clusters.
Galaxy morphologyimagemixedVLM morphology accuracy 83.8% (DECaLS) vs 81.0% (SDSS); the 2.8-pt gap is not significant.
Remote sensingimagemixedColor-only features recover 76.2% of 10-class accuracy vs 83.2% combined, revealing a color shortcut.
Seismologysignalmixed21.7% of noise-labelled STEAD traces carry coherent transient bursts; the instrument-type hypothesis is refuted.
CardiologyaudiosupportedRecording-protocol metadata alone predicts abnormality (AUC 0.60) and collapses out-of-source (0.35), exposing a confound.
EcoacousticsaudiosupportedA mid/high-band bird-presence classifier shows a large, robust drop in discriminability across recording sets.
Mechanical CAD3-DsupportedScale-invariant shape descriptors cluster 1,500 CAD parts into function-agnostic form families without supervision.
Plant phenotyping3-DmixedMaize initiates leaves sequentially where tomato is bursty, separable in 3-D scans.
Materials informaticstablemixedRandom k-fold CV underestimates extrapolation error; leave-one-family-out RMSE is 3.1–7.0× higher.
Symbolic regressionformulasupportedThe Cramér–Rao form Var⁡(a^)=σ2/(N​Var​log⁡x) predicts empirical exponent-estimation variance across all 8 monomial Feynman laws, sampled ranges, and noise levels.
Knowledge engineeringgraphsupportedDisease-associated proteins carry more distinct GO-function annotations than degree-matched controls, and the excess grows with PPI degree; it replicates on withheld test-split edges.
CS/ML methodologytracerefutedRejection-driven repairs are not dominated by omission, contrary to the pre-registered hypothesis.
Figure 10: Review dimension scores across different backbone strengths. Data reflects the 6 backbones with the broadest case coverage, ordered by overall score. The score gap between the strongest and weakest backbones is largest for factual accuracy and smallest for multimodal grounding.
Figure 10: Review dimension scores across different backbone strengths. Data reflects the 6 backbones with the broadest case coverage, ordered by overall score. The score gap between the strongest and weakest backbones is largest for factual accuracy and smallest for multimodal grounding.
Table 9: Numbers the system produced for the STEAD noise audit.
Headline and controlsHeterogeneity and robustness
Full-detector prevalence21.7% (163/750)Channel BH / HH / HN prevalence32.8 / 29.1 / 2.4%
95% confidence interval[18.8, 24.9]%channel χ2p=5.7×10−14
Amplitude-only baseline2.0%Per-network prevalence range0–65%
Ablation, no coincidence term2.0%network χ2p=3.5×10−14
Null false-alarm rate (target 1%)1.07%Station-cluster bootstrap 95% CI[17.9, 25.8]%
Sensitivity (FAR, null, window)stable 19–25%
Figure 11: The seismic audit at a glance. Left: 4 of the three-component traces the agent read, amplitude normalised. The first is a labelled earthquake, shown for reference, and the second a labelled noise trace that really is stationary background. The last two are also labelled noise, yet each carries a coherent onset at the dashed line, with an STA/LTA peak above the upper quartile of the labelled earthquakes and an onset rectilinearity near 1. All 4 were selected by the run’s own stored onset statistics rather than by eye. Right: the share of noise-labelled traces the full detector flags, with its 95% confidence interval, against 3 label-agnostic controls. Removing the cross-channel coincidence term collapses the detector to the amplitude-only rate, which is what identifies timing coincidence across components as the mechanism it is using.
Figure 11: The seismic audit at a glance. Left: 4 of the three-component traces the agent read, amplitude normalised. The first is a labelled earthquake, shown for reference, and the second a labelled noise trace that really is stationary background. The last two are also labelled noise, yet each carries a coherent onset at the dashed line, with an STA/LTA peak above the upper quartile of the labelled earthquakes and an onset rectilinearity near 1. All 4 were selected by the run’s own stored onset statistics rather than by eye. Right: the share of noise-labelled traces the full detector flags, with its 95% confidence interval, against 3 label-agnostic controls. Removing the cross-channel coincidence term collapses the detector to the amplitude-only rate, which is what identifies timing coincidence across components as the mechanism it is using.
Table 10: Numbers the system produced for the paediatric chest radiograph study. As in Table 9, every value is traceable to a real run_python standard output.
Effect and generalisationDiscrimination and robustness
Patchiness effect size, all imagesd=1.25AUC, mean entropy only0.840
development splitd=1.26AUC, mean + patchiness0.851
held-out splitd=1.29AUC, patchiness only0.847
Label difference, Mann-Whitneyp<0.0001AUC, raw-pixel baseline0.634
Independent of the mean levelp=1.7×10−5Window ablation, 8 / 16 / 32 pxd=1.63 / 1.50 / 1.36
Region-of-interest sweepd=1.18 to 1.35
Figure 12: Visualization and evaluation of sliding-window local-entropy features. Left: Normal and pneumonic radiographs paired with their local-entropy maps. The displayed samples correspond to the median patchiness of each class. Notably, the pneumonic entropy map exhibits higher variance (i.e., a mottled appearance) compared to the relatively smooth normal map. Right: Classification performance on the held-out test set. Entropy-derived features achieve score of 0.840–0.851, significantly outperforming the raw-pixel baseline (0.634). This indicates that the discriminative signal relies on local structural patterns rather than raw pixel intensities.
Figure 12: Visualization and evaluation of sliding-window local-entropy features. Left: Normal and pneumonic radiographs paired with their local-entropy maps. The displayed samples correspond to the median patchiness of each class. Notably, the pneumonic entropy map exhibits higher variance (i.e., a mottled appearance) compared to the relatively smooth normal map. Right: Classification performance on the held-out test set. Entropy-derived features achieve score of 0.840–0.851, significantly outperforming the raw-pixel baseline (0.634). This indicates that the discriminative signal relies on local structural patterns rather than raw pixel intensities.
Table 11: Cost per paper by backbone model. The experiment stage dominates the overall cost across all evaluated systems. †Open-weight backbones were run locally.
BackboneTokens in/out↓Cache hit↑$ / paper↓Wall-clock↓
Sonnet 598k / 191k93%$2.6329 min
GPT-5.6122k / 112k89%$4.3412 min
Qwen3.5-27B†1.4M / 75k$0.0630 min
Gemma-4-31B†714k / 41k$0.0314 min
Figure 13: Heatmap visualization of the evaluation scores from Table 3. Darker cells indicate higher panel scores. Model reasoning strength corresponds to row darkness, with the strongest backbones positioned at the top and smaller open-weight models at the bottom. Notably, factual accuracy consistently achieves the highest scores (the darkest column) across all evaluated models.
Figure 13: Heatmap visualization of the evaluation scores from Table 3. Darker cells indicate higher panel scores. Model reasoning strength corresponds to row darkness, with the strongest backbones positioned at the top and smaller open-weight models at the bottom. Notably, factual accuracy consistently achieves the highest scores (the darkest column) across all evaluated models.
Table 12: Perception gain per evaluation dimension. Across 5 blind pairs and 1 vision-off pair, each cell represents the difference in panel scores (original scale: 0–10) between the perception-enabled system and the blind baseline. Positive values indicate that visual perception improves performance. † Cardiology uses the weaker vision-off run.
Standard peer-reviewMM-mandatory
Case (Δ = on − baseline)Novelty↑Sound.↑Clarity↑Signif.↑Reprod.↑MM-gr.↑Factual↑Overall↑
Galaxy cross-survey+2.5+1.2+1.0+2.3+0.7+1.3+2.0+1.7
Seismology+1.0+1.5-1.0+1.5+0.5+1.5+0.5+1.5
Pathology+1.5+1.0+0.5+1.0+0.5+3.0-0.5+1.5
Mechanical CAD+0.5+0.5+3.5+2.5+2.5+1.0+2.0+3.0
Plant phenotyping+1.0+0.0+0.0+0.5+0.7+4.0+0.3+0.3
Cardiology†+0.3+0.3+0.3+0.0+1.3-1.7+0.0+0.3
Macro-average Δ+1.14+0.75+0.72+1.31+1.03+1.53+0.72+1.39
Table 13: Validating the judge before it is trusted; the target column lists the pre-set acceptance thresholds.
Validity checkStatisticTargetMeasured
Inter-judge agreementKrippendorff α>0.60.66
Self-preference biasown − others≈00†
Verbosity biasscore vs. length ρ≈00.16
Table 14: Every condition enforced by the 3 checks. Each row is a separate predicate in the stage’s exit function; failing any predicate returns the agent to the loop with the corresponding demand as the reason. The claim check runs on the drafted manuscript rather than a finalize payload, so it acts on the text itself.
ConditionWhat it demands of the stage output
Idea check (ideation)
SchemaA research question, a hypothesis, an experiment protocol, and a falsification criterion, all non-empty.
BreadthAt least 5 self-screened candidate projects, each rated for novelty risk.
Prior artAt least 3 focused literature searches, one of them aimed at the selected idea specifically.
FeasibilityThe selected idea marked fully computational. A proposal that would need a physical experiment is refused outright.
Minimal claimThe smallest claim worth publishing if the rest of the study fails, stated separately from the hypothesis.
Novelty evidenceWhat the searches returned for and against this particular idea, with citations.
Claim scopeAn explicit statement of what the data cannot establish, separating the measured proxy from any mechanistic or causal reading.
Effective sampleThe decisive-event count for the key test, estimated from the real data counts, and whether it is adequate.
LeakageWhether any step uses ground-truth labels at decision time, and what the label-agnostic counterpart is.
Visual auditRequired once the agent has looked at any raw item: which groups it viewed, and how a disagreement with the given label was resolved.
Novelty languageAbsolute-novelty phrasing (“first”, “unstudied”, “no prior work”) is rejected, because a bounded search cannot support it.
Rigour check (experiment)
VerdictOne of supported, refuted, mixed, null, or infeasible. An honest negative is a valid exit.
Real executionAt least one run_python call that exited 0 and produced real output.
PerceptionA case that carries a look_at_* budget cannot finalise a positive verdict without having looked at the raw evidence at least once.
Key numbersThe decisive numbers the code printed, reported as a structured record.
ProvenanceEvery reported number must appear in the text of a real run_python output.
Real dataSome run must have loaded the actual data rather than hand-coded rows.
ReproducibilityAt least 60% of the reported numbers present in the union of the run’s real outputs.
Multiple testsTwo or more reported p-values require a stated count of every test run and the correction applied to that count.
CircularityWhether the predictor derives from the same representation whose behaviour it predicts, and how that is handled.
Full batteryAt least 4 analyses, each with a saved figure: primary, baseline, ablation, mechanism, breakdown, sensitivity.
Lead selectionThe headline must name an existing analysis, be listed first, and not also appear among the demoted ones.
Lead significanceA lead whose p-values are all ≥0.05 is rejected unless the verdict is itself null or insufficient.
DemotionA non-significant analysis that is not the lead must be demoted, which keeps it in the trace and out of the manuscript.
Correction baseThe stated correction count must cover the demoted analyses too, so demoting cannot shrink the denominator.
Claim check (writeup)
TraceabilityEvery number in the drafted text is matched against the grounded set derived from the experiment record.
Guarded revisionThe prose-polish pass is reverted wholesale if it alters a number, a citation, a claim, or a model name.
Table 15: Check rejections across the 36 primary runs. The two count columns differ because a single run can be rejected multiple times by the same condition. The checks refused 115 finalize attempts in total, and only 4 of the 36 runs reached both stage exits without ever being sent back. The dominant single cause is result selection: in 26 runs, the agent tried to present a non-significant analysis as a finding and was made to demote it.
ConditionWhat firedRejectionsRuns
Idea check (ideation): 28 rejections in 23 of the 36 runs
Schemaa required ideation field left empty1210
Effective samplethe decisive-event estimate left empty44
Novelty evidencesearch evidence for the selected idea left empty33
Visual auditimages inspected but the audit left empty33
Claim scopewhat the data cannot establish left empty33
Novelty languageabsolute-novelty phrasing in the proposal22
Breadthfewer than 5 screened candidates11
Rigour check (experiment): 87 rejections in 32 of the 36 runs
Demotiona non-significant analysis left in the paper5126
Schemaverdict left empty2116
Lead selectionlead unset, not listed first, or a bad demotion name98
Schemano key numbers reported33
Provenancea reported number absent from real standard output11
Perceptionfinalised without looking at the raw evidence11
Multiple teststhe stated test count missing or under-counted11
Total115
Table 16: Run composition across the 36 primary runs. Counts represent issued tool calls recovered from each run’s stored trace. The step budget is 24 for ideation and 50 for experiment, and no run exhausted either.
IdeationExperiment
Per runMeanMedianMaxMeanMedianMax
Agent steps8.891936.03749
Tool calls, all kinds19.916.54937.23851
Code executions (run_python)0.00031.83347
Literature search calls8.79100.000
Perception calls, all channels8.44.5381.007
of which visual (look_at_*)4.53180.907
Exit-check rejections0.8142.427
Table 17: Outcomes of the 36 primary runs. The verdicts are the system’s own, taken from the verified experiment record. Demoted analyses were executed and remain in the trace, but the claim check excludes them from the manuscript.
Outcome over the 36-case suiteCountShare
Self-reported verdict
Supported, the pre-specified hypothesis held1644%
Mixed, part of the hypothesis held1747%
Refuted, the pre-specified hypothesis did not hold26%
No verdict, the experiment stage exhausted its step budget13%
Artifacts produced
Manuscripts drafted in full, with their figures36100%
Experiment stages that exited through the rigour check3597%
Manuscripts scored by the full 2-judge panel36100%
Analyses per run
Analyses carried into the manuscript2657.4 / run
Analyses demoted to the trace671.9 / run
Table 18: The perception layer. A modality automatically unlocks its tools from the specification file, requiring no per-discipline registration. Cases indicates how many of the 36 primary runs had the tool available, and Calls indicates how many times it was invoked.
ToolModalityWhat it returnsCasesCalls
Visual channel: render the artifact, then look at it
look_at_imageimagethe VLM’s reading of specific image files1149
look_at_signalsignalone time-series panel per channel: onsets, bursts, envelopes647
look_at_3d3-Drendered XY, XZ and YZ projections of a cloud or mesh540
look_at_tabletableshape, columns, dtypes, head and summary statistics122
look_at_audioaudiothe rendered waveform and spectrogram417
look_at_videovideoa sample of frames, inspected together313
look_at_trajectorytrajectorythe path coloured by time, plus its speed profile37
Native channel: read the modality in its own terms, no image
analyze_signalsignaltrend, dominant FFT frequencies, peaks, statistics640
analyze_audioaudioduration, rate, RMS, spectral centroid, zero-crossing439
analyze_3d3-Dpoint count, bounding box, centroid, extent, PCA axes521
read_tracetracethe ordered sequence of a track or an agent run log220
analyze_trajectorytrajectorypath length, displacement, straightness, speed, turning angles319
analyze_videovideoframe count, rate, resolution, frame-difference motion33
Table 19: The 5 structural specifications that the writeup stage can resolve to. The abstract column shows the word range for a single-paragraph abstract.
StyleSection orderAbstract
Machine learningIntroduction, Related Work, Method, Experiments, Conclusion, Limitations150–220
BiomedicalIntroduction, Results, Discussion, Methods150–200
Earth & spaceIntroduction, Data, Methods, Results, Discussion, Conclusions150–250
PhysicsIntroduction, Theory and Methods, Results, Discussion, Conclusion150–250
ChemistryIntroduction, Experimental Section, Results and Discussion, Conclusions150–250

Findings

  • 36개 사례 전체에서 원본 데이터부터 완성된 논문까지 전체 경로를 완료했고, 기준 추론 모델(Claude Sonnet 5)로 평균 전체 논문 점수 6.3점(10점 만점)을 받았다.
  • 전처리된 스칼라 특징만 받는 블라인드 시스템과의 쌍대 비교에서, 직접 지각을 사용하는 시스템이 7개 평가 차원 전부에서 더 높은 점수를 받았고 승패 비교의 85%에서 승리했다.
  • 지진학 사례에서 노이즈로 라벨된 파형의 21.7%가 실제 지진 이벤트로 판정되었으며, 이 추정치는 여러 오경보율, 널모델, 윈도우 설정, 417개 관측소에 대한 클러스터 부트스트랩에서도 안정적이었다.
  • 소아 흉부 방사선 사진 사례에서 국소 엔트로피 기반 특징이 held-out 테스트에서 AUC 0.840~0.851을 기록해, 원시 픽셀 기반 baseline(0.634)을 상당히 앞질렀다.
  • 9개 추론 백본 모델 비교에서 강한 폐쇄형 모델(Sonnet 5, GLM, Kimi)이 가장 크고 균형 잡힌 점수 프로파일을 보였고, 약한 오픈모델(Qwen3.5-9B, Gemma-4-26B)은 특히 novelty와 significance에서 점수가 떨어졌다.

Where it can be used

  • 원본 신호·이미지·3D 구조 등 이질적 데이터를 다루는 다분야 연구 워크플로우 자동화 실험에 참고할 수 있다.
  • 가설 생성부터 실험 실행, 논문 작성까지 이어지는 파이프라인에서 코드 기반 검증 절차(참신성·통계적 타당성·주장 추적)를 설계할 때 참고할 수 있다.
  • 의료 영상이나 지구물리 신호처럼 라벨이 불완전하거나 잡음이 섞인 데이터셋에서 이상 패턴을 재검토하는 용도로 응용 아이디어를 얻을 수 있다.

Limits and open work

  • 평가는 저자들이 구성한 36개 사례 모음에 한정되며, 이 사례들이 다른 실제 연구 상황을 얼마나 대표하는지는 별도로 검증되지 않았다.
  • 점수는 2명의 LLM 심사위원(자동화된 채점자)이 매긴 것으로, 인간 전문가 동료 심사와의 일치 여부는 이 요약에 포함된 자료로는 확인되지 않는다.
  • 블라인드 비교는 5개의 쌍(및 1개의 비전-오프 쌍)에 한정되어 있어 전체 36개 사례로 일반화되었는지는 알 수 없다.
  • 지각 도구 목록과 검증 체크는 저자들이 정의한 특정 데이터 형식과 모달리티에 맞춰 설계된 것으로, 다른 종류의 원본 데이터에 대한 일반화는 확인되지 않았다.
  • 비용은 백본 모델에 따라 논문당 0.03달러에서 4.34달러로 보고되었으나, 대규모 실제 연구실 도입 시 비용·시간 효율성은 추가 검증이 필요하다.

Why it matters

Most AI-scientist systems today read data only after it has been turned into text, labels, or a few summary numbers, which can hide the very patterns that matter for a discovery. This work shows that letting an AI agent examine raw signals, images, and structures directly - and building in code-level checks against overclaiming - produces more substantive, better-scoring scientific papers.

Terms in this paper

  • ReAct 루프 · 관찰-추론-행동을 반복하며 스스로 다음 단계를 정하는 에이전트 동작 방식
  • 멀티모달 그라운딩 (multimodal grounding) · 주장이나 결론이 실제 원본 데이터(이미지·신호 등)에 근거하고 있는 정도를 평가하는 항목
  • HARKing · 실험 결과를 본 뒤에 마치 처음부터 그 가설을 세웠던 것처럼 사후에 가설을 짜맞추는 행위
  • 실행 이력(execution record/provenance) · 코드 실행, 표준출력, 생성된 그림 등 실제로 수행된 과정을 남긴 기록으로 주장의 근거를 추적하는 데 쓰임
  • 블라인드 변이(blind variant) · 원본 데이터 대신 미리 계산된 숫자(스칼라 특징)만 받는 비교용 시스템

Original abstract (English)

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.

Authors · Bobo Li

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Bobo Li et al., arXiv:2608.13558, arxiv-nonexclusive