Figure 2: Progression from raw evidence to verified findings across three demonstration cases. The top three rows track workflows in seismology, pathology, and 3-D CAD. From left to right, each row begins with raw evidence (a three-channel seismogram, a stained pathology tile, and a 3-D CAD model), identifies specific structural cues, and outlines the subsequent hypothesis and action sequence. The rightmost column displays the verified findings, such as the discovery that 21.7% of noise labels are real events. The bottom band depicts a precomputed interface where the artifact is reduced to a feature vector, resulting in lost structural relations and a narrower research question space.
Table 1: The demonstration suite: 5 categories, 36 cases, one real downloadable dataset each, with the sample count N per dataset. Evidence – perceptual, symbolic, quantitative-statistical, procedural. Modality – image, signal, spectrum, audio, video, 3-D, trajectory, table, formula, sequence, field, graph.
Discipline
Representative dataset
Evidence
Modality
N
Physical sciences
5 cases
Condensed matter / nano
NFFA-EUROPE (Aversa et al. 2018)
2,655
Vibrational spectroscopy
RRUFF (Lafuente et al. 2015)
2,000
Materials informatics
UCI superconductor (Hamidieh 2018)
21,263
Molecular chemistry
PubChem (Kim et al. 2025)
30
Symbolic regression
Feynman (Udrescu and Tegmark 2020)
12
Earth & space
9 cases
Remote sensing
EuroSAT (Helber et al. 2019)
5,000
Galaxy morphology
Galaxy Zoo (Lintott et al. 2008)
1,000
Galaxy cross-survey
GZ DECaLS (Walmsley et al. 2022)
210
Gravitational waves
GWOSC (LIGO-Virgo Collaboration 2021)
1,500
Seismology
STEAD (Mousavi et al. 2019)
1,500
Marine biology
WHOI-Plankton (Orenstein et al. 2015)
3,000
Geology / petrophysics
Digital Rocks (Prodanović et al. 2015)
375
Meteorology
SEVIR (Veillette et al. 2020)
384
Cyclone dynamics
IBTrACS (Knapp et al. 2010)
400
Life & medical
7 cases
Pathology
Kather CRC (Kather et al. 2016)
5,000
Radiology
Chest X-ray (Kermany et al. 2018)
3,000
Medical imaging
MedMNIST CT (Yang et al. 2023a)
1,496
Cardiology
CinC 2016 (Liu et al. 2016)
2,000
Sleep neuroscience
Sleep-EDF (Kemp et al. 2000)
1,520
Cell biology
Cell Tracking Ch. (Ulman et al. 2017)
280
Genomics
DNA (H3) (Nguyen et al. 2016)
10,000
Agricultural & ecological
8 cases
Plant pathology
PlantVillage (Hughes and Salathé 2015)
3,002
Precision agriculture
Indian Pines (Baumgardner et al. 2015)
2,000
Animal behavior
CalMS21 (Sun et al. 2021)
2,500
Ecoacoustics
Bird Audio Det. (Stowell et al. 2019)
2,000
Marine bioacoustics
Watkins MMSD (Sayigh et al. 2016)
1,697
Figure 3: Architecture of the OmniScientist framework. At the top, raw evidence from multiple disciplines enters the system, categorized into four evidence families (perceptual, symbolic, quantitative, and procedural) and 12 modalities. The core pipeline consists of three sequential stages. First, the Ideation stage (left) observes materials, searches literature, and formulates falsifiable hypotheses. Next, the Experiment stage (center) designs tests, executes code, and inspects results to generate an execution record containing standard output, figures, data, and configurations. Finally, the Writeup stage (right) selects, grounds, and reports claims supported exclusively by the execution record to compile the final paper. At the bottom, a lifecycle-wide perception layer provides spatial, temporal, cross-channel, statistical, and dynamic analysis capabilities. Dashed arrows indicate that these perception tools are available to all three stages of the pipeline.
Table 2: The 4 families of scientific evidence OmniScientist is designed to perceive.
Evidence family
Typical artifacts
Perceptual
Images, video, micrographs, radar, astronomical and remote-sensing imagery, the visual form of scientific plots, audio, and 3-D structure.
Figure 4: Raw observations and derived discoveries processed by the perception layer across 16 cases, 11 modalities, and all 4 evidence families. The figure presents a four-by-four grid of artifacts, each taken from that case’s own data exactly as the run received it, with the discipline named at the top left of every panel and the modality at the top right. Within each artifact, a red bounding box marks the specific feature flagged by the agent on the raw record. Below the artifact, the saw label reports the direct observation made by the agent, and the found label details the verified experimental result produced by that observation. The first three rows cover perceptual and procedural evidence, spanning images, spectra, signals, audio, video, three-dimensional structures, and trajectories. The bottom row presents quantitative and symbolic evidence, where the layer reads the native numeric structure of a table, a formula, a sequence, or a graph instead of rendering an image.
Table 3: Detailed review scores across reasoning backbones. Per-dimension means (0–10) are derived from a 2-judge cross-family panel (deepseek-v4-flash and gemini-2.5-flash-lite) over the entire case suite. For these evaluations, the framework and perception models are held fixed, with only the reasoning backbone swapped. Failed runs are excluded; thus, means are computed exclusively over successfully scored papers. The highest value in each column is highlighted, and coverage per backbone is detailed in Table 4. Notably, clarity exhibits the least degradation, whereas factual accuracy and soundness most closely track the underlying backbone strength.
Standard peer-review
MM-mandatory
Backbone
Novelty↑
Sound.↑
Clarity↑
Signif.↑
Reprod.↑
MM-grnd↑
Factual↑
Overall↑
Sonnet 5 (Anthropic 2026)
6.3
7.0
7.0
6.3
6.1
5.1
7.7
6.3
GPT-5.6 (OpenAI 2026)
5.2
6.3
6.3
5.0
5.2
4.2
7.7
5.6
GLM-5.2 (Zhipu 2026)
6.2
7.1
6.8
6.4
5.9
6.6
7.5
6.5
Kimi K2.7 (Kimi 2025)
6.2
7.2
6.7
6.2
5.5
5.8
8.0
6.2
Qwen3.5-122B (Qwen Team 2026)
4.7
5.5
6.2
4.8
4.8
4.8
6.5
5.1
Qwen3.5-27B (Qwen Team 2026)
5.0
5.6
5.9
4.9
4.6
4.9
6.4
5.1
Qwen3.5-9B (Qwen Team 2026)
4.0
4.1
4.8
3.7
3.7
3.9
4.8
4.0
Gemma-4-31B (Google 2026)
4.7
5.0
5.6
4.5
4.4
4.6
6.5
4.8
Gemma-4-26B (Google 2026)
4.4
4.4
5.0
4.0
3.7
3.8
5.1
4.2
Figure 5: Two-stage verification pipeline for experimental results and manuscript claims. In the top row from left to right, an unverified experimental result undergoes a rigour check that verifies real execution, accounts for all tests, tests for independence and leakage, and ensures the headline belongs to supported analyses. If a check fails, unsupported analyses are traced and null results trigger re-ideation. Successful validation yields a verified result with certified metrics and attached provenance. Further right, a claim check matches reported numbers (n1…nk) to recorded outputs and reported claims (C1…Cm) to recorded analyses (E1…Em), resulting in a manuscript with fully traced numbers and supported claims. The bottom band displays the execution record, which serves as the source of truth for both checks. This record captures data I/O, standard output, generated figures, and a complete list of all attempted tests including unsupported attempts (Tu).
Table 4: Backbone generality across the 36-case suite. The table reports the number of cases dispatched, the resulting completed papers, and the mean composite score for these successful runs.
Backbone
Sonnet 5
Sonnet
5
GPT 5.6
GPT
5.6
GLM 5.2
GLM
5.2
Kimi K2.7
Kimi
K2.7
Qwen3.5 122B
Qwen3.5
122B
Qwen3.5 27B
Qwen3.5
27B
Qwen3.5 9B
Sonnet
5
GPT
5.6
GLM
5.2
Kimi
K2.7
Qwen3.5
122B
Qwen3.5
27B
Qwen3.5
9B
Gemma-4
31B
Gemma-4
26B
Cases
36
10
18
9
34
36
32
36
34
Completed↑
36
9
17
6
30
32
18
32
25
Mean↑
6.5
5.7
6.7
6.5
5.4
5.3
4.1
5.0
4.3
Figure 6: Per-case review profiles across the 7 dimensions. Radar plots are shown for 9 high-coverage cases spanning 4 evidence modalities; each line represents one backbone, scored by a 2-judge panel (on a 0–10 scale). The strong backbones (Sonnet 5, GLM, Kimi) exhibit the largest, most balanced profiles, while the weak open models (Qwen3.5-9B, Gemma-4-26B) collapse inward, particularly in novelty and significance, although clarity varies the least across all models.
Table 5: Backbone quality aggregated by evidence modality and discipline family. Results are grouped by modality on the left and by discipline family on the right. The reported metric is the mean composite score evaluated by the 2-judge cross-family panel on a scale of 0 to 10. A dash indicates that a backbone produced no scored papers for that specific category.
By evidence modality
By discipline
Backbone
Image
Signal
Audio
Video
3-D
Traj.
T&S
Earth
Life
Agri.
Engin.
Phys.
Sonnet 5
6.4
6.1
7.1
6.4
7.0
6.3
6.4
6.5
6.7
6.8
6.6
5.8
GPT-5.6
5.9
5.5
5.5
–
5.9
–
–
5.6
6.0
6.4
5.2
–
Qwen3.5-27B
5.5
5.0
5.9
5.3
4.8
4.9
5.8
5.1
5.5
5.5
4.8
5.5
Qwen3.5-9B
3.8
5.5
4.1
4.5
4.4
4.3
3.2
4.1
4.5
4.0
4.4
3.2
Qwen3.5-122B
4.8
5.4
5.9
5.2
5.4
5.8
5.5
4.6
5.7
6.0
4.9
5.4
Gemma-4-31B
5.3
4.9
5.4
5.1
5.9
4.0
3.5
5.1
5.7
4.6
5.0
4.9
Gemma-4-26B
4.1
4.3
4.4
5.0
4.3
4.3
4.8
4.5
3.8
4.3
4.4
4.6
GLM-5.2
6.6
6.8
6.6
–
6.6
–
–
6.4
6.9
6.5
6.4
7.5
Kimi K2.7
6.8
6.7
5.9
–
6.0
–
–
6.8
6.2
6.6
6.0
–
Figure 7: Dimension-wise perception gain. For the 5 cases evaluated under both conditions, the chart shows the mean scores with perception removed (pink) and for the full OmniScientist (teal), scored by the same judge across both settings. The 2 panels of this row share one colour key. The largest gain is observed in multimodal grounding, while factual accuracy remains identical since both conditions undergo the same provenance check.
Table 6: Review rubric performance across the complete evaluation suite using the Sonnet 5 backbone. The 2-judge cross-family panel evaluated all cases on a scale of 0 to 10. Using a single backbone ensures direct comparability across all columns. Factual accuracy reaches 7.0 or higher in 30 of the 36 completed cases.
Standard peer-review
MM-mandatory
Discipline
Novelty↑
Sound.↑
Clarity↑
Signif.↑
Reprod.↑
MM-gr.↑
Factual↑
Overall↑
Physical sciences
Condensed matter
2.5
2.5
4.0
2.0
2.0
4.5
1.0
2.5
Vibrational spectroscopy
6.0
8.0
7.5
7.0
7.0
6.0
9.0
7.0
Materials informatics
7.0
7.0
7.5
8.0
7.0
4.5
7.0
7.0
Molecular chemistry
4.5
5.5
6.0
3.5
6.0
4.0
5.0
4.5
Symbolic regression
7.0
8.0
7.5
7.5
7.5
3.5
9.0
7.0
Earth & space
Remote sensing
7.0
7.0
6.5
7.0
6.0
5.5
7.0
6.5
Galaxy morphology
6.5
6.5
7.0
6.0
6.0
5.5
7.5
6.5
Galaxy cross-survey
7.0
7.0
6.5
6.0
6.0
5.0
7.5
6.5
Gravitational waves
5.0
5.0
5.5
5.0
5.0
3.5
5.5
4.5
Seismology
6.5
8.0
7.0
7.5
6.5
4.5
8.5
7.0
Marine biology
6.0
8.5
7.5
7.0
6.5
5.5
9.5
7.0
Geology / petrophysics
6.5
7.5
8.0
7.5
7.5
6.0
9.5
7.5
Cyclone dynamics
6.0
6.5
7.0
6.0
5.5
4.5
7.0
6.0
Meteorology
6.0
7.0
8.0
6.0
5.0
4.5
8.0
6.5
Life & medical
Pathology
7.0
8.0
8.0
6.5
6.5
6.5
9.0
7.0
Radiology
7.0
7.5
8.0
6.5
6.5
6.5
8.5
7.0
Medical imaging (CT)
6.5
7.5
6.5
6.0
7.0
4.0
9.0
6.0
Cardiology
6.5
8.0
8.0
7.5
6.5
4.5
9.0
7.0
Sleep neuroscience
6.5
4.5
7.0
6.0
5.5
3.5
4.0
4.5
Genomics
7.0
7.0
7.0
6.0
6.0
5.0
7.5
6.5
Cell biology
7.0
6.0
5.5
6.5
5.5
4.0
6.0
5.5
Agricultural & ecological
Plant pathology
6.5
7.5
8.0
7.0
6.5
6.0
8.0
7.0
Precision agriculture
6.0
6.5
6.5
6.0
6.0
5.0
8.0
6.0
Animal behavior
6.5
7.5
7.0
5.5
5.5
4.5
8.0
6.5
Ecoacoustics
7.0
8.0
8.0
7.0
6.5
5.0
9.0
7.5
Figure 8: Breakdown of head-to-head judgments. For each dimension, the 3 bars show the share of all judgments won by OmniScientist, won by the same system with perception removed, and declared a tie. Judges never tie on novelty or significance, the dimensions enhanced by perception, and tie most often on factual accuracy and reproducibility, which both conditions share through the provenance check.
Table 7: Comparison of feature utilization between the two systems. In all paired cases, the perceiving system focuses on information inherent in the raw records, whereas the blind system relies solely on the provided scalar features despite sharing the same task and backbone.
Case
Evidence only the raw record carries
Question each system asked
Galaxy cross-survey
Morphology read off the image
With perception: does a vision model’s morphological reading degrade on the shallower survey for the same galaxies? Blind: can the 3 classes be separated along the 8 supplied feature axes?
Seismology
Cross-component polarisation of the waveform
With perception: what fraction of noise-labelled traces carry coherent polarised transients? Blind: do frequency-shape features retain a depth imprint after an attenuation correction?
Pathology
Texture and nuclear density of the tile
With perception: is the complex class a compositional mixture of the pure tissue prototypes? Blind: does the class confusion matrix follow an a priori similarity ranking?
Mechanical CAD
Principal-axis geometry of the point cloud
With perception: do the function-defined labels correspond to latent geometric morphotypes? Blind: do dimension-standardised part families show tighter descriptor dispersion?
Plant phenotyping
Per-point organ labels across repeated scans
With perception: do the two species differ in how many leaves grow at once? Blind: do the species separate on shape descriptors once the size axis is removed?
Figure 9: Component ablation on the seismology case, where each configuration removes a single component with the backbone fixed. Novelty is the judged novelty score and Composite the 7-dimension mean, both evaluated by the DS-V4-Flash judge (0–10). The dashed line marks the composite score of the full system.
Table 8: Overview of 13 end-to-end runs, grouped by evidence modality and spanning all 4 families. Evaluation metrics are system-specific, extracted verbatim from verified experiment records.
Discipline
Evidence
Evaluation metric
Headline finding
Radiology
image
supported
Pneumonic pediatric lung fields show markedly higher local-entropy heterogeneity (patchiness) than normal.
Pathology
image
supported
The COMPLEX H&E class is heterogeneous, splitting into compositional sub-clusters.
Galaxy morphology
image
mixed
VLM morphology accuracy 83.8% (DECaLS) vs 81.0% (SDSS); the 2.8-pt gap is not significant.
Remote sensing
image
mixed
Color-only features recover 76.2% of 10-class accuracy vs 83.2% combined, revealing a color shortcut.
Seismology
signal
mixed
21.7% of noise-labelled STEAD traces carry coherent transient bursts; the instrument-type hypothesis is refuted.
Cardiology
audio
supported
Recording-protocol metadata alone predicts abnormality (AUC 0.60) and collapses out-of-source (0.35), exposing a confound.
Ecoacoustics
audio
supported
A mid/high-band bird-presence classifier shows a large, robust drop in discriminability across recording sets.
Mechanical CAD
3-D
supported
Scale-invariant shape descriptors cluster 1,500 CAD parts into function-agnostic form families without supervision.
Plant phenotyping
3-D
mixed
Maize initiates leaves sequentially where tomato is bursty, separable in 3-D scans.
Materials informatics
table
mixed
Random k-fold CV underestimates extrapolation error; leave-one-family-out RMSE is 3.1–7.0× higher.
Symbolic regression
formula
supported
The Cramér–Rao form Var(a^)=σ2/(NVarlogx) predicts empirical exponent-estimation variance across all 8 monomial Feynman laws, sampled ranges, and noise levels.
Knowledge engineering
graph
supported
Disease-associated proteins carry more distinct GO-function annotations than degree-matched controls, and the excess grows with PPI degree; it replicates on withheld test-split edges.
CS/ML methodology
trace
refuted
Rejection-driven repairs are not dominated by omission, contrary to the pre-registered hypothesis.
Figure 10: Review dimension scores across different backbone strengths. Data reflects the 6 backbones with the broadest case coverage, ordered by overall score. The score gap between the strongest and weakest backbones is largest for factual accuracy and smallest for multimodal grounding.
Table 9: Numbers the system produced for the STEAD noise audit.
Headline and controls
Heterogeneity and robustness
Full-detector prevalence
21.7% (163/750)
Channel BH / HH / HN prevalence
32.8 / 29.1 / 2.4%
95% confidence interval
[18.8, 24.9]%
channel χ2
p=5.7×10−14
Amplitude-only baseline
2.0%
Per-network prevalence range
0–65%
Ablation, no coincidence term
2.0%
network χ2
p=3.5×10−14
Null false-alarm rate (target 1%)
1.07%
Station-cluster bootstrap 95% CI
[17.9, 25.8]%
Sensitivity (FAR, null, window)
stable 19–25%
Figure 11: The seismic audit at a glance. Left: 4 of the three-component traces the agent read, amplitude normalised. The first is a labelled earthquake, shown for reference, and the second a labelled noise trace that really is stationary background. The last two are also labelled noise, yet each carries a coherent onset at the dashed line, with an STA/LTA peak above the upper quartile of the labelled earthquakes and an onset rectilinearity near 1. All 4 were selected by the run’s own stored onset statistics rather than by eye. Right: the share of noise-labelled traces the full detector flags, with its 95% confidence interval, against 3 label-agnostic controls. Removing the cross-channel coincidence term collapses the detector to the amplitude-only rate, which is what identifies timing coincidence across components as the mechanism it is using.
Table 10: Numbers the system produced for the paediatric chest radiograph study. As in Table 9, every value is traceable to a real run_python standard output.
Effect and generalisation
Discrimination and robustness
Patchiness effect size, all images
d=1.25
AUC, mean entropy only
0.840
development split
d=1.26
AUC, mean + patchiness
0.851
held-out split
d=1.29
AUC, patchiness only
0.847
Label difference, Mann-Whitney
p<0.0001
AUC, raw-pixel baseline
0.634
Independent of the mean level
p=1.7×10−5
Window ablation, 8 / 16 / 32 px
d=1.63 / 1.50 / 1.36
Region-of-interest sweep
d=1.18 to 1.35
Figure 12: Visualization and evaluation of sliding-window local-entropy features. Left: Normal and pneumonic radiographs paired with their local-entropy maps. The displayed samples correspond to the median patchiness of each class. Notably, the pneumonic entropy map exhibits higher variance (i.e., a mottled appearance) compared to the relatively smooth normal map. Right: Classification performance on the held-out test set. Entropy-derived features achieve score of 0.840–0.851, significantly outperforming the raw-pixel baseline (0.634). This indicates that the discriminative signal relies on local structural patterns rather than raw pixel intensities.
Table 11: Cost per paper by backbone model. The experiment stage dominates the overall cost across all evaluated systems. †Open-weight backbones were run locally.
Backbone
Tokens in/out↓
Cache hit↑
$ / paper↓
Wall-clock↓
Sonnet 5
98k / 191k
93%
$2.63
29 min
GPT-5.6
122k / 112k
89%
$4.34
12 min
Qwen3.5-27B†
1.4M / 75k
–
$0.06
30 min
Gemma-4-31B†
714k / 41k
–
$0.03
14 min
Figure 13: Heatmap visualization of the evaluation scores from Table 3. Darker cells indicate higher panel scores. Model reasoning strength corresponds to row darkness, with the strongest backbones positioned at the top and smaller open-weight models at the bottom. Notably, factual accuracy consistently achieves the highest scores (the darkest column) across all evaluated models.
Table 12: Perception gain per evaluation dimension. Across 5 blind pairs and 1 vision-off pair, each cell represents the difference in panel scores (original scale: 0–10) between the perception-enabled system and the blind baseline. Positive values indicate that visual perception improves performance. † Cardiology uses the weaker vision-off run.
Standard peer-review
MM-mandatory
Case (Δ = on − baseline)
Novelty↑
Sound.↑
Clarity↑
Signif.↑
Reprod.↑
MM-gr.↑
Factual↑
Overall↑
Galaxy cross-survey
+2.5
+1.2
+1.0
+2.3
+0.7
+1.3
+2.0
+1.7
Seismology
+1.0
+1.5
-1.0
+1.5
+0.5
+1.5
+0.5
+1.5
Pathology
+1.5
+1.0
+0.5
+1.0
+0.5
+3.0
-0.5
+1.5
Mechanical CAD
+0.5
+0.5
+3.5
+2.5
+2.5
+1.0
+2.0
+3.0
Plant phenotyping
+1.0
+0.0
+0.0
+0.5
+0.7
+4.0
+0.3
+0.3
Cardiology†
+0.3
+0.3
+0.3
+0.0
+1.3
-1.7
+0.0
+0.3
Macro-average Δ
+1.14
+0.75
+0.72
+1.31
+1.03
+1.53
+0.72
+1.39
Table 13: Validating the judge before it is trusted; the target column lists the pre-set acceptance thresholds.
Validity check
Statistic
Target
Measured
Inter-judge agreement
Krippendorff α
>0.6
0.66
Self-preference bias
own − others
≈0
0†
Verbosity bias
score vs. length ρ
≈0
0.16
Table 14: Every condition enforced by the 3 checks. Each row is a separate predicate in the stage’s exit function; failing any predicate returns the agent to the loop with the corresponding demand as the reason. The claim check runs on the drafted manuscript rather than a finalize payload, so it acts on the text itself.
Condition
What it demands of the stage output
Idea check (ideation)
Schema
A research question, a hypothesis, an experiment protocol, and a falsification criterion, all non-empty.
Breadth
At least 5 self-screened candidate projects, each rated for novelty risk.
Prior art
At least 3 focused literature searches, one of them aimed at the selected idea specifically.
Feasibility
The selected idea marked fully computational. A proposal that would need a physical experiment is refused outright.
Minimal claim
The smallest claim worth publishing if the rest of the study fails, stated separately from the hypothesis.
Novelty evidence
What the searches returned for and against this particular idea, with citations.
Claim scope
An explicit statement of what the data cannot establish, separating the measured proxy from any mechanistic or causal reading.
Effective sample
The decisive-event count for the key test, estimated from the real data counts, and whether it is adequate.
Leakage
Whether any step uses ground-truth labels at decision time, and what the label-agnostic counterpart is.
Visual audit
Required once the agent has looked at any raw item: which groups it viewed, and how a disagreement with the given label was resolved.
Novelty language
Absolute-novelty phrasing (“first”, “unstudied”, “no prior work”) is rejected, because a bounded search cannot support it.
Rigour check (experiment)
Verdict
One of supported, refuted, mixed, null, or infeasible. An honest negative is a valid exit.
Real execution
At least one run_python call that exited 0 and produced real output.
Perception
A case that carries a look_at_* budget cannot finalise a positive verdict without having looked at the raw evidence at least once.
Key numbers
The decisive numbers the code printed, reported as a structured record.
Provenance
Every reported number must appear in the text of a real run_python output.
Real data
Some run must have loaded the actual data rather than hand-coded rows.
Reproducibility
At least 60% of the reported numbers present in the union of the run’s real outputs.
Multiple tests
Two or more reported p-values require a stated count of every test run and the correction applied to that count.
Circularity
Whether the predictor derives from the same representation whose behaviour it predicts, and how that is handled.
Full battery
At least 4 analyses, each with a saved figure: primary, baseline, ablation, mechanism, breakdown, sensitivity.
Lead selection
The headline must name an existing analysis, be listed first, and not also appear among the demoted ones.
Lead significance
A lead whose p-values are all ≥0.05 is rejected unless the verdict is itself null or insufficient.
Demotion
A non-significant analysis that is not the lead must be demoted, which keeps it in the trace and out of the manuscript.
Correction base
The stated correction count must cover the demoted analyses too, so demoting cannot shrink the denominator.
Claim check (writeup)
Traceability
Every number in the drafted text is matched against the grounded set derived from the experiment record.
Guarded revision
The prose-polish pass is reverted wholesale if it alters a number, a citation, a claim, or a model name.
Table 15: Check rejections across the 36 primary runs. The two count columns differ because a single run can be rejected multiple times by the same condition. The checks refused 115 finalize attempts in total, and only 4 of the 36 runs reached both stage exits without ever being sent back. The dominant single cause is result selection: in 26 runs, the agent tried to present a non-significant analysis as a finding and was made to demote it.
Condition
What fired
Rejections
Runs
Idea check (ideation): 28 rejections in 23 of the 36 runs
Schema
a required ideation field left empty
12
10
Effective sample
the decisive-event estimate left empty
4
4
Novelty evidence
search evidence for the selected idea left empty
3
3
Visual audit
images inspected but the audit left empty
3
3
Claim scope
what the data cannot establish left empty
3
3
Novelty language
absolute-novelty phrasing in the proposal
2
2
Breadth
fewer than 5 screened candidates
1
1
Rigour check (experiment): 87 rejections in 32 of the 36 runs
Demotion
a non-significant analysis left in the paper
51
26
Schema
verdict left empty
21
16
Lead selection
lead unset, not listed first, or a bad demotion name
9
8
Schema
no key numbers reported
3
3
Provenance
a reported number absent from real standard output
1
1
Perception
finalised without looking at the raw evidence
1
1
Multiple tests
the stated test count missing or under-counted
1
1
Total
115
Table 16: Run composition across the 36 primary runs. Counts represent issued tool calls recovered from each run’s stored trace. The step budget is 24 for ideation and 50 for experiment, and no run exhausted either.
Ideation
Experiment
Per run
Mean
Median
Max
Mean
Median
Max
Agent steps
8.8
9
19
36.0
37
49
Tool calls, all kinds
19.9
16.5
49
37.2
38
51
Code executions (run_python)
0.0
0
0
31.8
33
47
Literature search calls
8.7
9
10
0.0
0
0
Perception calls, all channels
8.4
4.5
38
1.0
0
7
of which visual (look_at_*)
4.5
3
18
0.9
0
7
Exit-check rejections
0.8
1
4
2.4
2
7
Table 17: Outcomes of the 36 primary runs. The verdicts are the system’s own, taken from the verified experiment record. Demoted analyses were executed and remain in the trace, but the claim check excludes them from the manuscript.
Outcome over the 36-case suite
Count
Share
Self-reported verdict
Supported, the pre-specified hypothesis held
16
44%
Mixed, part of the hypothesis held
17
47%
Refuted, the pre-specified hypothesis did not hold
2
6%
No verdict, the experiment stage exhausted its step budget
1
3%
Artifacts produced
Manuscripts drafted in full, with their figures
36
100%
Experiment stages that exited through the rigour check
35
97%
Manuscripts scored by the full 2-judge panel
36
100%
Analyses per run
Analyses carried into the manuscript
265
7.4 / run
Analyses demoted to the trace
67
1.9 / run
Table 18: The perception layer. A modality automatically unlocks its tools from the specification file, requiring no per-discipline registration. Cases indicates how many of the 36 primary runs had the tool available, and Calls indicates how many times it was invoked.
Tool
Modality
What it returns
Cases
Calls
Visual channel: render the artifact, then look at it
look_at_image
image
the VLM’s reading of specific image files
11
49
look_at_signal
signal
one time-series panel per channel: onsets, bursts, envelopes
6
47
look_at_3d
3-D
rendered XY, XZ and YZ projections of a cloud or mesh
5
40
look_at_table
table
shape, columns, dtypes, head and summary statistics
1
22
look_at_audio
audio
the rendered waveform and spectrogram
4
17
look_at_video
video
a sample of frames, inspected together
3
13
look_at_trajectory
trajectory
the path coloured by time, plus its speed profile
3
7
Native channel: read the modality in its own terms, no image
Table 19: The 5 structural specifications that the writeup stage can resolve to. The abstract column shows the word range for a single-paragraph abstract.
Style
Section order
Abstract
Machine learning
Introduction, Related Work, Method, Experiments, Conclusion, Limitations
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.