컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

AI 과학자가 요약된 숫자 대신 원본 데이터를 직접 읽으면 논문 품질이 눈에 띄게 좋아진다

arXiv:2608.135582026-08-12

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

AI 과학자가 요약된 숫자 대신 원본 데이터를 직접 읽으면 논문 품질이 눈에 띄게 좋아진다

OmniScientist는 지진파, 병리 슬라이드, 3D 모델 같은 원본 과학 데이터를 직접 관찰하면서 가설 세우기, 실험 실행, 논문 작성까지 전체 연구 과정을 자동으로 수행하는 AI 시스템이다. 5개 학문 분야에 걸친 36개 실제 데이터 사례 전부에서 원본 데이터부터 완성된 논문까지의 전체 경로를 완료했고, 주력 모델 기준 평균 논문 점수 10점 만점에 6.3점을 받았다. 원본 대신 미리 계산된 숫자만 보는 블라인드 버전과 비교했을 때, 원본을 직접 보는 버전이 7개 평가 항목 전부에서 더 높은 점수를 받았고 맞대결의 85%에서 승리했다.

METAL MEDIA 해설 도표

OmniScientist의 3단계 파이프라인과 지각 레이어

증거 상태측정 결과가 보고됨

  1. 지각 레이어이미지, 신호, 3D 구조, 표, 그래프 등 원본 증거를 4개 계열(지각/기호/정량통계/절차)로 분류해 직접 읽어들인다
  2. 아이디어 구상(Ideation)자료를 관찰하고 문헌을 검색해 검증 가능한 가설을 세우며, 코드 기반 체크가 구조적 완전성과 참신성을 검증한다
  3. 실험(Experiment)코드를 실행해 가설을 검증하고, 실행 출처와 통계적 타당성을 검증하는 엄격성 체크를 통과해야 한다
  4. 작성(Writeup)검증된 실행 기록에서만 주장을 골라 학문 분야별 서식에 맞춰 원고를 작성하고, 주장 검증 체크로 숫자와 진술을 원본 실행 기록과 대조한다
  5. 블라인드 비교 실험원본 데이터 대신 사전 계산된 스칼라 특징만 받는 시스템과 쌍대 비교해 지각의 기여도를 측정한다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. OmniScientist는 지각(관찰) 레이어와 아이디어 구상·실험·작성을 맡는 3개의 자율 에이전트를 결합해, 통제된 파이프라인 안에서 원본 관찰 내용이 연구 질문, 실험 설계, 최종 주장까지 바꿀 수 있게 한다.
  2. 증거를 지각(이미지, 스펙트럼), 기호(텍스트, 수식, 그래프), 정량-통계(표), 절차(궤적, 시뮬레이션) 4개 계열로 분류하고, 참신성, 통계적 타당성, 보고된 모든 숫자의 추적 가능성을 코드로 강제 검증한다.
  3. 이미지, 오디오, 비디오, 3D 구조, 지식 그래프 등을 포함한 5개 학문 분야, 36개 사례로 구성된 모음에서 시험했으며, 36개 사례 전부에서 원본 데이터부터 완성된 원고까지 완주했다.
  4. 미리 계산된 스칼라 특징만 받는 블라인드 버전과의 쌍대 비교에서, 직접 지각하는 버전이 7개 심사 항목 전부에서 더 나은 점수를 받았고 맞대결의 85%를 이겼다.
  5. 두 사례 연구에서 그 효과를 보여준다. 지진 사례에서는 '노이즈'로 라벨된 파형의 21.7%가 실제 지진 신호였음을 찾아냈고, 흉부 X선 사례에서는 국소 엔트로피(질감) 패턴이 폐렴과 정상을 구분하는 데 AUC 최대 0.851을 기록해 단순 픽셀 기반 기준선(0.634)을 크게 앞질렀다.
Figure 2: Progression from raw evidence to verified findings across three demonstration cases. The top three rows track workflows in seismology, pathology, and 3-D CAD. From left to right, each row begins with raw evidence (a three-channel seismogram, a stained pathology tile, and a 3-D CAD model), identifies specific structural cues, and outlines the subsequent hypothesis and action sequence. The rightmost column displays the verified findings, such as the discovery that 21.7% of noise labels are real events. The bottom band depicts a precomputed interface where the artifact is reduced to a feature vector, resulting in lost structural relations and a narrower research question space.
Figure 2: Progression from raw evidence to verified findings across three demonstration cases. The top three rows track workflows in seismology, pathology, and 3-D CAD. From left to right, each row begins with raw evidence (a three-channel seismogram, a stained pathology tile, and a 3-D CAD model), identifies specific structural cues, and outlines the subsequent hypothesis and action sequence. The rightmost column displays the verified findings, such as the discovery that 21.7% of noise labels are real events. The bottom band depicts a precomputed interface where the artifact is reduced to a feature vector, resulting in lost structural relations and a narrower research question space.
Table 1: The demonstration suite: 5 categories, 36 cases, one real downloadable dataset each, with the sample count N per dataset. Evidence – perceptual, symbolic, quantitative-statistical, procedural. Modality – image, signal, spectrum, audio, video, 3-D, trajectory, table, formula, sequence, field, graph.
DisciplineRepresentative datasetEvidenceModalityN
Physical sciences5 cases
Condensed matter / nanoNFFA-EUROPE (Aversa et al. 2018)2,655
Vibrational spectroscopyRRUFF (Lafuente et al. 2015)2,000
Materials informaticsUCI superconductor (Hamidieh 2018)21,263
Molecular chemistryPubChem (Kim et al. 2025)30
Symbolic regressionFeynman (Udrescu and Tegmark 2020)12
Earth & space9 cases
Remote sensingEuroSAT (Helber et al. 2019)5,000
Galaxy morphologyGalaxy Zoo (Lintott et al. 2008)1,000
Galaxy cross-surveyGZ DECaLS (Walmsley et al. 2022)210
Gravitational wavesGWOSC (LIGO-Virgo Collaboration 2021)1,500
SeismologySTEAD (Mousavi et al. 2019)1,500
Marine biologyWHOI-Plankton (Orenstein et al. 2015)3,000
Geology / petrophysicsDigital Rocks (Prodanović et al. 2015)375
MeteorologySEVIR (Veillette et al. 2020)384
Cyclone dynamicsIBTrACS (Knapp et al. 2010)400
Life & medical7 cases
PathologyKather CRC (Kather et al. 2016)5,000
RadiologyChest X-ray (Kermany et al. 2018)3,000
Medical imagingMedMNIST CT (Yang et al. 2023a)1,496
CardiologyCinC 2016 (Liu et al. 2016)2,000
Sleep neuroscienceSleep-EDF (Kemp et al. 2000)1,520
Cell biologyCell Tracking Ch. (Ulman et al. 2017)280
GenomicsDNA (H3) (Nguyen et al. 2016)10,000
Agricultural & ecological8 cases
Plant pathologyPlantVillage (Hughes and Salathé 2015)3,002
Precision agricultureIndian Pines (Baumgardner et al. 2015)2,000
Animal behaviorCalMS21 (Sun et al. 2021)2,500
EcoacousticsBird Audio Det. (Stowell et al. 2019)2,000
Marine bioacousticsWatkins MMSD (Sayigh et al. 2016)1,697
Figure 3: Architecture of the OmniScientist framework. At the top, raw evidence from multiple disciplines enters the system, categorized into four evidence families (perceptual, symbolic, quantitative, and procedural) and 12 modalities. The core pipeline consists of three sequential stages. First, the Ideation stage (left) observes materials, searches literature, and formulates falsifiable hypotheses. Next, the Experiment stage (center) designs tests, executes code, and inspects results to generate an execution record containing standard output, figures, data, and configurations. Finally, the Writeup stage (right) selects, grounds, and reports claims supported exclusively by the execution record to compile the final paper. At the bottom, a lifecycle-wide perception layer provides spatial, temporal, cross-channel, statistical, and dynamic analysis capabilities. Dashed arrows indicate that these perception tools are available to all three stages of the pipeline.
Figure 3: Architecture of the OmniScientist framework. At the top, raw evidence from multiple disciplines enters the system, categorized into four evidence families (perceptual, symbolic, quantitative, and procedural) and 12 modalities. The core pipeline consists of three sequential stages. First, the Ideation stage (left) observes materials, searches literature, and formulates falsifiable hypotheses. Next, the Experiment stage (center) designs tests, executes code, and inspects results to generate an execution record containing standard output, figures, data, and configurations. Finally, the Writeup stage (right) selects, grounds, and reports claims supported exclusively by the execution record to compile the final paper. At the bottom, a lifecycle-wide perception layer provides spatial, temporal, cross-channel, statistical, and dynamic analysis capabilities. Dashed arrows indicate that these perception tools are available to all three stages of the pipeline.
Table 2: The 4 families of scientific evidence OmniScientist is designed to perceive.
Evidence familyTypical artifacts
PerceptualImages, video, micrographs, radar, astronomical and remote-sensing imagery, the visual form of scientific plots, audio, and 3-D structure.
SymbolicNatural-language documents, formulae, variables, rules, sequences, knowledge graphs, logical and causal relations, mathematical models.
Quantitative-statisticalTables, measurements, distributions, curves, correlations, significance tests, regression results.
Procedural / dynamicExperimental steps, code execution, agent traces, simulations, dynamic evolution, protocols.
Figure 4: Raw observations and derived discoveries processed by the perception layer across 16 cases, 11 modalities, and all 4 evidence families. The figure presents a four-by-four grid of artifacts, each taken from that case’s own data exactly as the run received it, with the discipline named at the top left of every panel and the modality at the top right. Within each artifact, a red bounding box marks the specific feature flagged by the agent on the raw record. Below the artifact, the saw label reports the direct observation made by the agent, and the found label details the verified experimental result produced by that observation. The first three rows cover perceptual and procedural evidence, spanning images, spectra, signals, audio, video, three-dimensional structures, and trajectories. The bottom row presents quantitative and symbolic evidence, where the layer reads the native numeric structure of a table, a formula, a sequence, or a graph instead of rendering an image.
Figure 4: Raw observations and derived discoveries processed by the perception layer across 16 cases, 11 modalities, and all 4 evidence families. The figure presents a four-by-four grid of artifacts, each taken from that case’s own data exactly as the run received it, with the discipline named at the top left of every panel and the modality at the top right. Within each artifact, a red bounding box marks the specific feature flagged by the agent on the raw record. Below the artifact, the saw label reports the direct observation made by the agent, and the found label details the verified experimental result produced by that observation. The first three rows cover perceptual and procedural evidence, spanning images, spectra, signals, audio, video, three-dimensional structures, and trajectories. The bottom row presents quantitative and symbolic evidence, where the layer reads the native numeric structure of a table, a formula, a sequence, or a graph instead of rendering an image.
Table 3: Detailed review scores across reasoning backbones. Per-dimension means (0–10) are derived from a 2-judge cross-family panel (deepseek-v4-flash and gemini-2.5-flash-lite) over the entire case suite. For these evaluations, the framework and perception models are held fixed, with only the reasoning backbone swapped. Failed runs are excluded; thus, means are computed exclusively over successfully scored papers. The highest value in each column is highlighted, and coverage per backbone is detailed in Table 4. Notably, clarity exhibits the least degradation, whereas factual accuracy and soundness most closely track the underlying backbone strength.
Standard peer-reviewMM-mandatory
BackboneNovelty↑Sound.↑Clarity↑Signif.↑Reprod.↑MM-grnd↑Factual↑Overall↑
Sonnet 5 (Anthropic 2026)6.37.07.06.36.15.17.76.3
GPT-5.6 (OpenAI 2026)5.26.36.35.05.24.27.75.6
GLM-5.2 (Zhipu 2026)6.27.16.86.45.96.67.56.5
Kimi K2.7 (Kimi 2025)6.27.26.76.25.55.88.06.2
Qwen3.5-122B (Qwen Team 2026)4.75.56.24.84.84.86.55.1
Qwen3.5-27B (Qwen Team 2026)5.05.65.94.94.64.96.45.1
Qwen3.5-9B (Qwen Team 2026)4.04.14.83.73.73.94.84.0
Gemma-4-31B (Google 2026)4.75.05.64.54.44.66.54.8
Gemma-4-26B (Google 2026)4.44.45.04.03.73.85.14.2
Figure 5: Two-stage verification pipeline for experimental results and manuscript claims. In the top row from left to right, an unverified experimental result undergoes a rigour check that verifies real execution, accounts for all tests, tests for independence and leakage, and ensures the headline belongs to supported analyses. If a check fails, unsupported analyses are traced and null results trigger re-ideation. Successful validation yields a verified result with certified metrics and attached provenance. Further right, a claim check matches reported numbers (n1​…​nk) to recorded outputs and reported claims (C1​…​Cm) to recorded analyses (E1​…​Em), resulting in a manuscript with fully traced numbers and supported claims. The bottom band displays the execution record, which serves as the source of truth for both checks. This record captures data I/O, standard output, generated figures, and a complete list of all attempted tests including unsupported attempts (Tu).
Figure 5: Two-stage verification pipeline for experimental results and manuscript claims. In the top row from left to right, an unverified experimental result undergoes a rigour check that verifies real execution, accounts for all tests, tests for independence and leakage, and ensures the headline belongs to supported analyses. If a check fails, unsupported analyses are traced and null results trigger re-ideation. Successful validation yields a verified result with certified metrics and attached provenance. Further right, a claim check matches reported numbers (n1​…​nk) to recorded outputs and reported claims (C1​…​Cm) to recorded analyses (E1​…​Em), resulting in a manuscript with fully traced numbers and supported claims. The bottom band displays the execution record, which serves as the source of truth for both checks. This record captures data I/O, standard output, generated figures, and a complete list of all attempted tests including unsupported attempts (Tu).
Table 4: Backbone generality across the 36-case suite. The table reports the number of cases dispatched, the resulting completed papers, and the mean composite score for these successful runs.
BackboneSonnet 5Sonnet5GPT 5.6GPT5.6GLM 5.2GLM5.2Kimi K2.7KimiK2.7Qwen3.5 122BQwen3.5122BQwen3.5 27BQwen3.527BQwen3.5 9B
Sonnet
5
GPT
5.6
GLM
5.2
Kimi
K2.7
Qwen3.5
122B
Qwen3.5
27B
Qwen3.5
9B
Gemma-4
31B
Gemma-4
26B
Cases36101893436323634
Completed↑3691763032183225
Mean↑6.55.76.76.55.45.34.15.04.3
Figure 6: Per-case review profiles across the 7 dimensions. Radar plots are shown for 9 high-coverage cases spanning 4 evidence modalities; each line represents one backbone, scored by a 2-judge panel (on a 0–10 scale). The strong backbones (Sonnet 5, GLM, Kimi) exhibit the largest, most balanced profiles, while the weak open models (Qwen3.5-9B, Gemma-4-26B) collapse inward, particularly in novelty and significance, although clarity varies the least across all models.
Figure 6: Per-case review profiles across the 7 dimensions. Radar plots are shown for 9 high-coverage cases spanning 4 evidence modalities; each line represents one backbone, scored by a 2-judge panel (on a 0–10 scale). The strong backbones (Sonnet 5, GLM, Kimi) exhibit the largest, most balanced profiles, while the weak open models (Qwen3.5-9B, Gemma-4-26B) collapse inward, particularly in novelty and significance, although clarity varies the least across all models.
Table 5: Backbone quality aggregated by evidence modality and discipline family. Results are grouped by modality on the left and by discipline family on the right. The reported metric is the mean composite score evaluated by the 2-judge cross-family panel on a scale of 0 to 10. A dash indicates that a backbone produced no scored papers for that specific category.
By evidence modalityBy discipline
BackboneImageSignalAudioVideo3-DTraj.T&SEarthLifeAgri.Engin.Phys.
Sonnet 56.46.17.16.47.06.36.46.56.76.86.65.8
GPT-5.65.95.55.55.95.66.06.45.2
Qwen3.5-27B5.55.05.95.34.84.95.85.15.55.54.85.5
Qwen3.5-9B3.85.54.14.54.44.33.24.14.54.04.43.2
Qwen3.5-122B4.85.45.95.25.45.85.54.65.76.04.95.4
Gemma-4-31B5.34.95.45.15.94.03.55.15.74.65.04.9
Gemma-4-26B4.14.34.45.04.34.34.84.53.84.34.44.6
GLM-5.26.66.86.66.66.46.96.56.47.5
Kimi K2.76.86.75.96.06.86.26.66.0
Figure 7: Dimension-wise perception gain. For the 5 cases evaluated under both conditions, the chart shows the mean scores with perception removed (pink) and for the full OmniScientist (teal), scored by the same judge across both settings. The 2 panels of this row share one colour key. The largest gain is observed in multimodal grounding, while factual accuracy remains identical since both conditions undergo the same provenance check.
Figure 7: Dimension-wise perception gain. For the 5 cases evaluated under both conditions, the chart shows the mean scores with perception removed (pink) and for the full OmniScientist (teal), scored by the same judge across both settings. The 2 panels of this row share one colour key. The largest gain is observed in multimodal grounding, while factual accuracy remains identical since both conditions undergo the same provenance check.
Table 6: Review rubric performance across the complete evaluation suite using the Sonnet 5 backbone. The 2-judge cross-family panel evaluated all cases on a scale of 0 to 10. Using a single backbone ensures direct comparability across all columns. Factual accuracy reaches 7.0 or higher in 30 of the 36 completed cases.
Standard peer-reviewMM-mandatory
DisciplineNovelty↑Sound.↑Clarity↑Signif.↑Reprod.↑MM-gr.↑Factual↑Overall↑
Physical sciences
Condensed matter2.52.54.02.02.04.51.02.5
Vibrational spectroscopy6.08.07.57.07.06.09.07.0
Materials informatics7.07.07.58.07.04.57.07.0
Molecular chemistry4.55.56.03.56.04.05.04.5
Symbolic regression7.08.07.57.57.53.59.07.0
Earth & space
Remote sensing7.07.06.57.06.05.57.06.5
Galaxy morphology6.56.57.06.06.05.57.56.5
Galaxy cross-survey7.07.06.56.06.05.07.56.5
Gravitational waves5.05.05.55.05.03.55.54.5
Seismology6.58.07.07.56.54.58.57.0
Marine biology6.08.57.57.06.55.59.57.0
Geology / petrophysics6.57.58.07.57.56.09.57.5
Cyclone dynamics6.06.57.06.05.54.57.06.0
Meteorology6.07.08.06.05.04.58.06.5
Life & medical
Pathology7.08.08.06.56.56.59.07.0
Radiology7.07.58.06.56.56.58.57.0
Medical imaging (CT)6.57.56.56.07.04.09.06.0
Cardiology6.58.08.07.56.54.59.07.0
Sleep neuroscience6.54.57.06.05.53.54.04.5
Genomics7.07.07.06.06.05.07.56.5
Cell biology7.06.05.56.55.54.06.05.5
Agricultural & ecological
Plant pathology6.57.58.07.06.56.08.07.0
Precision agriculture6.06.56.56.06.05.08.06.0
Animal behavior6.57.57.05.55.54.58.06.5
Ecoacoustics7.08.08.07.06.55.09.07.5
Figure 8: Breakdown of head-to-head judgments. For each dimension, the 3 bars show the share of all judgments won by OmniScientist, won by the same system with perception removed, and declared a tie. Judges never tie on novelty or significance, the dimensions enhanced by perception, and tie most often on factual accuracy and reproducibility, which both conditions share through the provenance check.
Figure 8: Breakdown of head-to-head judgments. For each dimension, the 3 bars show the share of all judgments won by OmniScientist, won by the same system with perception removed, and declared a tie. Judges never tie on novelty or significance, the dimensions enhanced by perception, and tie most often on factual accuracy and reproducibility, which both conditions share through the provenance check.
Table 7: Comparison of feature utilization between the two systems. In all paired cases, the perceiving system focuses on information inherent in the raw records, whereas the blind system relies solely on the provided scalar features despite sharing the same task and backbone.
CaseEvidence only the raw record carriesQuestion each system asked
Galaxy cross-surveyMorphology read off the imageWith perception: does a vision model’s morphological reading degrade on the shallower survey for the same galaxies? Blind: can the 3 classes be separated along the 8 supplied feature axes?
SeismologyCross-component polarisation of the waveformWith perception: what fraction of noise-labelled traces carry coherent polarised transients? Blind: do frequency-shape features retain a depth imprint after an attenuation correction?
PathologyTexture and nuclear density of the tileWith perception: is the complex class a compositional mixture of the pure tissue prototypes? Blind: does the class confusion matrix follow an a priori similarity ranking?
Mechanical CADPrincipal-axis geometry of the point cloudWith perception: do the function-defined labels correspond to latent geometric morphotypes? Blind: do dimension-standardised part families show tighter descriptor dispersion?
Plant phenotypingPer-point organ labels across repeated scansWith perception: do the two species differ in how many leaves grow at once? Blind: do the species separate on shape descriptors once the size axis is removed?
Figure 9: Component ablation on the seismology case, where each configuration removes a single component with the backbone fixed. Novelty is the judged novelty score and Composite the 7-dimension mean, both evaluated by the DS-V4-Flash judge (0–10). The dashed line marks the composite score of the full system.
Figure 9: Component ablation on the seismology case, where each configuration removes a single component with the backbone fixed. Novelty is the judged novelty score and Composite the 7-dimension mean, both evaluated by the DS-V4-Flash judge (0–10). The dashed line marks the composite score of the full system.
Table 8: Overview of 13 end-to-end runs, grouped by evidence modality and spanning all 4 families. Evaluation metrics are system-specific, extracted verbatim from verified experiment records.
DisciplineEvidenceEvaluation metricHeadline finding
RadiologyimagesupportedPneumonic pediatric lung fields show markedly higher local-entropy heterogeneity (patchiness) than normal.
PathologyimagesupportedThe COMPLEX H&E class is heterogeneous, splitting into compositional sub-clusters.
Galaxy morphologyimagemixedVLM morphology accuracy 83.8% (DECaLS) vs 81.0% (SDSS); the 2.8-pt gap is not significant.
Remote sensingimagemixedColor-only features recover 76.2% of 10-class accuracy vs 83.2% combined, revealing a color shortcut.
Seismologysignalmixed21.7% of noise-labelled STEAD traces carry coherent transient bursts; the instrument-type hypothesis is refuted.
CardiologyaudiosupportedRecording-protocol metadata alone predicts abnormality (AUC 0.60) and collapses out-of-source (0.35), exposing a confound.
EcoacousticsaudiosupportedA mid/high-band bird-presence classifier shows a large, robust drop in discriminability across recording sets.
Mechanical CAD3-DsupportedScale-invariant shape descriptors cluster 1,500 CAD parts into function-agnostic form families without supervision.
Plant phenotyping3-DmixedMaize initiates leaves sequentially where tomato is bursty, separable in 3-D scans.
Materials informaticstablemixedRandom k-fold CV underestimates extrapolation error; leave-one-family-out RMSE is 3.1–7.0× higher.
Symbolic regressionformulasupportedThe Cramér–Rao form Var⁡(a^)=σ2/(N​Var​log⁡x) predicts empirical exponent-estimation variance across all 8 monomial Feynman laws, sampled ranges, and noise levels.
Knowledge engineeringgraphsupportedDisease-associated proteins carry more distinct GO-function annotations than degree-matched controls, and the excess grows with PPI degree; it replicates on withheld test-split edges.
CS/ML methodologytracerefutedRejection-driven repairs are not dominated by omission, contrary to the pre-registered hypothesis.
Figure 10: Review dimension scores across different backbone strengths. Data reflects the 6 backbones with the broadest case coverage, ordered by overall score. The score gap between the strongest and weakest backbones is largest for factual accuracy and smallest for multimodal grounding.
Figure 10: Review dimension scores across different backbone strengths. Data reflects the 6 backbones with the broadest case coverage, ordered by overall score. The score gap between the strongest and weakest backbones is largest for factual accuracy and smallest for multimodal grounding.
Table 9: Numbers the system produced for the STEAD noise audit.
Headline and controlsHeterogeneity and robustness
Full-detector prevalence21.7% (163/750)Channel BH / HH / HN prevalence32.8 / 29.1 / 2.4%
95% confidence interval[18.8, 24.9]%channel χ2p=5.7×10−14
Amplitude-only baseline2.0%Per-network prevalence range0–65%
Ablation, no coincidence term2.0%network χ2p=3.5×10−14
Null false-alarm rate (target 1%)1.07%Station-cluster bootstrap 95% CI[17.9, 25.8]%
Sensitivity (FAR, null, window)stable 19–25%
Figure 11: The seismic audit at a glance. Left: 4 of the three-component traces the agent read, amplitude normalised. The first is a labelled earthquake, shown for reference, and the second a labelled noise trace that really is stationary background. The last two are also labelled noise, yet each carries a coherent onset at the dashed line, with an STA/LTA peak above the upper quartile of the labelled earthquakes and an onset rectilinearity near 1. All 4 were selected by the run’s own stored onset statistics rather than by eye. Right: the share of noise-labelled traces the full detector flags, with its 95% confidence interval, against 3 label-agnostic controls. Removing the cross-channel coincidence term collapses the detector to the amplitude-only rate, which is what identifies timing coincidence across components as the mechanism it is using.
Figure 11: The seismic audit at a glance. Left: 4 of the three-component traces the agent read, amplitude normalised. The first is a labelled earthquake, shown for reference, and the second a labelled noise trace that really is stationary background. The last two are also labelled noise, yet each carries a coherent onset at the dashed line, with an STA/LTA peak above the upper quartile of the labelled earthquakes and an onset rectilinearity near 1. All 4 were selected by the run’s own stored onset statistics rather than by eye. Right: the share of noise-labelled traces the full detector flags, with its 95% confidence interval, against 3 label-agnostic controls. Removing the cross-channel coincidence term collapses the detector to the amplitude-only rate, which is what identifies timing coincidence across components as the mechanism it is using.
Table 10: Numbers the system produced for the paediatric chest radiograph study. As in Table 9, every value is traceable to a real run_python standard output.
Effect and generalisationDiscrimination and robustness
Patchiness effect size, all imagesd=1.25AUC, mean entropy only0.840
development splitd=1.26AUC, mean + patchiness0.851
held-out splitd=1.29AUC, patchiness only0.847
Label difference, Mann-Whitneyp<0.0001AUC, raw-pixel baseline0.634
Independent of the mean levelp=1.7×10−5Window ablation, 8 / 16 / 32 pxd=1.63 / 1.50 / 1.36
Region-of-interest sweepd=1.18 to 1.35
Figure 12: Visualization and evaluation of sliding-window local-entropy features. Left: Normal and pneumonic radiographs paired with their local-entropy maps. The displayed samples correspond to the median patchiness of each class. Notably, the pneumonic entropy map exhibits higher variance (i.e., a mottled appearance) compared to the relatively smooth normal map. Right: Classification performance on the held-out test set. Entropy-derived features achieve score of 0.840–0.851, significantly outperforming the raw-pixel baseline (0.634). This indicates that the discriminative signal relies on local structural patterns rather than raw pixel intensities.
Figure 12: Visualization and evaluation of sliding-window local-entropy features. Left: Normal and pneumonic radiographs paired with their local-entropy maps. The displayed samples correspond to the median patchiness of each class. Notably, the pneumonic entropy map exhibits higher variance (i.e., a mottled appearance) compared to the relatively smooth normal map. Right: Classification performance on the held-out test set. Entropy-derived features achieve score of 0.840–0.851, significantly outperforming the raw-pixel baseline (0.634). This indicates that the discriminative signal relies on local structural patterns rather than raw pixel intensities.
Table 11: Cost per paper by backbone model. The experiment stage dominates the overall cost across all evaluated systems. †Open-weight backbones were run locally.
BackboneTokens in/out↓Cache hit↑$ / paper↓Wall-clock↓
Sonnet 598k / 191k93%$2.6329 min
GPT-5.6122k / 112k89%$4.3412 min
Qwen3.5-27B†1.4M / 75k$0.0630 min
Gemma-4-31B†714k / 41k$0.0314 min
Figure 13: Heatmap visualization of the evaluation scores from Table 3. Darker cells indicate higher panel scores. Model reasoning strength corresponds to row darkness, with the strongest backbones positioned at the top and smaller open-weight models at the bottom. Notably, factual accuracy consistently achieves the highest scores (the darkest column) across all evaluated models.
Figure 13: Heatmap visualization of the evaluation scores from Table 3. Darker cells indicate higher panel scores. Model reasoning strength corresponds to row darkness, with the strongest backbones positioned at the top and smaller open-weight models at the bottom. Notably, factual accuracy consistently achieves the highest scores (the darkest column) across all evaluated models.
Table 12: Perception gain per evaluation dimension. Across 5 blind pairs and 1 vision-off pair, each cell represents the difference in panel scores (original scale: 0–10) between the perception-enabled system and the blind baseline. Positive values indicate that visual perception improves performance. † Cardiology uses the weaker vision-off run.
Standard peer-reviewMM-mandatory
Case (Δ = on − baseline)Novelty↑Sound.↑Clarity↑Signif.↑Reprod.↑MM-gr.↑Factual↑Overall↑
Galaxy cross-survey+2.5+1.2+1.0+2.3+0.7+1.3+2.0+1.7
Seismology+1.0+1.5-1.0+1.5+0.5+1.5+0.5+1.5
Pathology+1.5+1.0+0.5+1.0+0.5+3.0-0.5+1.5
Mechanical CAD+0.5+0.5+3.5+2.5+2.5+1.0+2.0+3.0
Plant phenotyping+1.0+0.0+0.0+0.5+0.7+4.0+0.3+0.3
Cardiology†+0.3+0.3+0.3+0.0+1.3-1.7+0.0+0.3
Macro-average Δ+1.14+0.75+0.72+1.31+1.03+1.53+0.72+1.39
Table 13: Validating the judge before it is trusted; the target column lists the pre-set acceptance thresholds.
Validity checkStatisticTargetMeasured
Inter-judge agreementKrippendorff α>0.60.66
Self-preference biasown − others≈00†
Verbosity biasscore vs. length ρ≈00.16
Table 14: Every condition enforced by the 3 checks. Each row is a separate predicate in the stage’s exit function; failing any predicate returns the agent to the loop with the corresponding demand as the reason. The claim check runs on the drafted manuscript rather than a finalize payload, so it acts on the text itself.
ConditionWhat it demands of the stage output
Idea check (ideation)
SchemaA research question, a hypothesis, an experiment protocol, and a falsification criterion, all non-empty.
BreadthAt least 5 self-screened candidate projects, each rated for novelty risk.
Prior artAt least 3 focused literature searches, one of them aimed at the selected idea specifically.
FeasibilityThe selected idea marked fully computational. A proposal that would need a physical experiment is refused outright.
Minimal claimThe smallest claim worth publishing if the rest of the study fails, stated separately from the hypothesis.
Novelty evidenceWhat the searches returned for and against this particular idea, with citations.
Claim scopeAn explicit statement of what the data cannot establish, separating the measured proxy from any mechanistic or causal reading.
Effective sampleThe decisive-event count for the key test, estimated from the real data counts, and whether it is adequate.
LeakageWhether any step uses ground-truth labels at decision time, and what the label-agnostic counterpart is.
Visual auditRequired once the agent has looked at any raw item: which groups it viewed, and how a disagreement with the given label was resolved.
Novelty languageAbsolute-novelty phrasing (“first”, “unstudied”, “no prior work”) is rejected, because a bounded search cannot support it.
Rigour check (experiment)
VerdictOne of supported, refuted, mixed, null, or infeasible. An honest negative is a valid exit.
Real executionAt least one run_python call that exited 0 and produced real output.
PerceptionA case that carries a look_at_* budget cannot finalise a positive verdict without having looked at the raw evidence at least once.
Key numbersThe decisive numbers the code printed, reported as a structured record.
ProvenanceEvery reported number must appear in the text of a real run_python output.
Real dataSome run must have loaded the actual data rather than hand-coded rows.
ReproducibilityAt least 60% of the reported numbers present in the union of the run’s real outputs.
Multiple testsTwo or more reported p-values require a stated count of every test run and the correction applied to that count.
CircularityWhether the predictor derives from the same representation whose behaviour it predicts, and how that is handled.
Full batteryAt least 4 analyses, each with a saved figure: primary, baseline, ablation, mechanism, breakdown, sensitivity.
Lead selectionThe headline must name an existing analysis, be listed first, and not also appear among the demoted ones.
Lead significanceA lead whose p-values are all ≥0.05 is rejected unless the verdict is itself null or insufficient.
DemotionA non-significant analysis that is not the lead must be demoted, which keeps it in the trace and out of the manuscript.
Correction baseThe stated correction count must cover the demoted analyses too, so demoting cannot shrink the denominator.
Claim check (writeup)
TraceabilityEvery number in the drafted text is matched against the grounded set derived from the experiment record.
Guarded revisionThe prose-polish pass is reverted wholesale if it alters a number, a citation, a claim, or a model name.
Table 15: Check rejections across the 36 primary runs. The two count columns differ because a single run can be rejected multiple times by the same condition. The checks refused 115 finalize attempts in total, and only 4 of the 36 runs reached both stage exits without ever being sent back. The dominant single cause is result selection: in 26 runs, the agent tried to present a non-significant analysis as a finding and was made to demote it.
ConditionWhat firedRejectionsRuns
Idea check (ideation): 28 rejections in 23 of the 36 runs
Schemaa required ideation field left empty1210
Effective samplethe decisive-event estimate left empty44
Novelty evidencesearch evidence for the selected idea left empty33
Visual auditimages inspected but the audit left empty33
Claim scopewhat the data cannot establish left empty33
Novelty languageabsolute-novelty phrasing in the proposal22
Breadthfewer than 5 screened candidates11
Rigour check (experiment): 87 rejections in 32 of the 36 runs
Demotiona non-significant analysis left in the paper5126
Schemaverdict left empty2116
Lead selectionlead unset, not listed first, or a bad demotion name98
Schemano key numbers reported33
Provenancea reported number absent from real standard output11
Perceptionfinalised without looking at the raw evidence11
Multiple teststhe stated test count missing or under-counted11
Total115
Table 16: Run composition across the 36 primary runs. Counts represent issued tool calls recovered from each run’s stored trace. The step budget is 24 for ideation and 50 for experiment, and no run exhausted either.
IdeationExperiment
Per runMeanMedianMaxMeanMedianMax
Agent steps8.891936.03749
Tool calls, all kinds19.916.54937.23851
Code executions (run_python)0.00031.83347
Literature search calls8.79100.000
Perception calls, all channels8.44.5381.007
of which visual (look_at_*)4.53180.907
Exit-check rejections0.8142.427
Table 17: Outcomes of the 36 primary runs. The verdicts are the system’s own, taken from the verified experiment record. Demoted analyses were executed and remain in the trace, but the claim check excludes them from the manuscript.
Outcome over the 36-case suiteCountShare
Self-reported verdict
Supported, the pre-specified hypothesis held1644%
Mixed, part of the hypothesis held1747%
Refuted, the pre-specified hypothesis did not hold26%
No verdict, the experiment stage exhausted its step budget13%
Artifacts produced
Manuscripts drafted in full, with their figures36100%
Experiment stages that exited through the rigour check3597%
Manuscripts scored by the full 2-judge panel36100%
Analyses per run
Analyses carried into the manuscript2657.4 / run
Analyses demoted to the trace671.9 / run
Table 18: The perception layer. A modality automatically unlocks its tools from the specification file, requiring no per-discipline registration. Cases indicates how many of the 36 primary runs had the tool available, and Calls indicates how many times it was invoked.
ToolModalityWhat it returnsCasesCalls
Visual channel: render the artifact, then look at it
look_at_imageimagethe VLM’s reading of specific image files1149
look_at_signalsignalone time-series panel per channel: onsets, bursts, envelopes647
look_at_3d3-Drendered XY, XZ and YZ projections of a cloud or mesh540
look_at_tabletableshape, columns, dtypes, head and summary statistics122
look_at_audioaudiothe rendered waveform and spectrogram417
look_at_videovideoa sample of frames, inspected together313
look_at_trajectorytrajectorythe path coloured by time, plus its speed profile37
Native channel: read the modality in its own terms, no image
analyze_signalsignaltrend, dominant FFT frequencies, peaks, statistics640
analyze_audioaudioduration, rate, RMS, spectral centroid, zero-crossing439
analyze_3d3-Dpoint count, bounding box, centroid, extent, PCA axes521
read_tracetracethe ordered sequence of a track or an agent run log220
analyze_trajectorytrajectorypath length, displacement, straightness, speed, turning angles319
analyze_videovideoframe count, rate, resolution, frame-difference motion33
Table 19: The 5 structural specifications that the writeup stage can resolve to. The abstract column shows the word range for a single-paragraph abstract.
StyleSection orderAbstract
Machine learningIntroduction, Related Work, Method, Experiments, Conclusion, Limitations150–220
BiomedicalIntroduction, Results, Discussion, Methods150–200
Earth & spaceIntroduction, Data, Methods, Results, Discussion, Conclusions150–250
PhysicsIntroduction, Theory and Methods, Results, Discussion, Conclusion150–250
ChemistryIntroduction, Experimental Section, Results and Discussion, Conclusions150–250

실제로 확인된 결과

  • 36개 사례 전체에서 원본 데이터부터 완성된 논문까지 전체 경로를 완료했고, 기준 추론 모델(Claude Sonnet 5)로 평균 전체 논문 점수 6.3점(10점 만점)을 받았다.
  • 전처리된 스칼라 특징만 받는 블라인드 시스템과의 쌍대 비교에서, 직접 지각을 사용하는 시스템이 7개 평가 차원 전부에서 더 높은 점수를 받았고 승패 비교의 85%에서 승리했다.
  • 지진학 사례에서 노이즈로 라벨된 파형의 21.7%가 실제 지진 이벤트로 판정되었으며, 이 추정치는 여러 오경보율, 널모델, 윈도우 설정, 417개 관측소에 대한 클러스터 부트스트랩에서도 안정적이었다.
  • 소아 흉부 방사선 사진 사례에서 국소 엔트로피 기반 특징이 held-out 테스트에서 AUC 0.840~0.851을 기록해, 원시 픽셀 기반 baseline(0.634)을 상당히 앞질렀다.
  • 9개 추론 백본 모델 비교에서 강한 폐쇄형 모델(Sonnet 5, GLM, Kimi)이 가장 크고 균형 잡힌 점수 프로파일을 보였고, 약한 오픈모델(Qwen3.5-9B, Gemma-4-26B)은 특히 novelty와 significance에서 점수가 떨어졌다.

어디에 쓸 수 있나

  • 원본 신호·이미지·3D 구조 등 이질적 데이터를 다루는 다분야 연구 워크플로우 자동화 실험에 참고할 수 있다.
  • 가설 생성부터 실험 실행, 논문 작성까지 이어지는 파이프라인에서 코드 기반 검증 절차(참신성·통계적 타당성·주장 추적)를 설계할 때 참고할 수 있다.
  • 의료 영상이나 지구물리 신호처럼 라벨이 불완전하거나 잡음이 섞인 데이터셋에서 이상 패턴을 재검토하는 용도로 응용 아이디어를 얻을 수 있다.

한계와 남은 검증

  • 평가는 저자들이 구성한 36개 사례 모음에 한정되며, 이 사례들이 다른 실제 연구 상황을 얼마나 대표하는지는 별도로 검증되지 않았다.
  • 점수는 2명의 LLM 심사위원(자동화된 채점자)이 매긴 것으로, 인간 전문가 동료 심사와의 일치 여부는 이 요약에 포함된 자료로는 확인되지 않는다.
  • 블라인드 비교는 5개의 쌍(및 1개의 비전-오프 쌍)에 한정되어 있어 전체 36개 사례로 일반화되었는지는 알 수 없다.
  • 지각 도구 목록과 검증 체크는 저자들이 정의한 특정 데이터 형식과 모달리티에 맞춰 설계된 것으로, 다른 종류의 원본 데이터에 대한 일반화는 확인되지 않았다.
  • 비용은 백본 모델에 따라 논문당 0.03달러에서 4.34달러로 보고되었으나, 대규모 실제 연구실 도입 시 비용·시간 효율성은 추가 검증이 필요하다.

왜 중요한가

현재 대부분의 AI 과학자 시스템은 텍스트, 라벨, 미리 계산된 요약 숫자로만 데이터를 접하는데, 이 과정에서 발견에 결정적인 패턴이 사라질 수 있다. 이 연구는 AI 에이전트가 원본 신호·이미지·구조를 직접 살피고, 과도한 주장을 코드 수준에서 걸러내도록 설계하면 더 내실 있고 점수가 높은 과학 논문을 만들 수 있음을 보여준다.

이 논문의 용어

  • ReAct 루프 · 관찰-추론-행동을 반복하며 스스로 다음 단계를 정하는 에이전트 동작 방식
  • 멀티모달 그라운딩(multimodal grounding) · 주장이나 결론이 실제 원본 데이터(이미지·신호 등)에 근거하고 있는 정도를 평가하는 항목
  • HARKing · 실험 결과를 본 뒤 마치 처음부터 그 가설을 세웠던 것처럼 사후에 가설을 짜맞추는 행위
  • 실행 이력(execution record/provenance) · 코드 실행, 표준출력, 생성된 그림 등 실제로 수행된 과정을 남긴 기록으로 주장의 근거를 추적하는 데 쓰인다
  • 블라인드 변이(blind variant) · 원본 데이터 대신 미리 계산된 숫자(스칼라 특징)만 받는 비교용 시스템

저자 · Bobo Li

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Bobo Li et al., arXiv:2608.13558, arxiv-nonexclusive