AI 과학자가 요약된 숫자 대신 원본 데이터를 직접 읽으면 논문 품질이 눈에 띄게 좋아진다
AI 과학자가 요약된 숫자 대신 원본 데이터를 직접 읽으면 논문 품질이 눈에 띄게 좋아진다
OmniScientist는 지진파, 병리 슬라이드, 3D 모델 같은 원본 과학 데이터를 직접 관찰하면서 가설 세우기, 실험 실행, 논문 작성까지 전체 연구 과정을 자동으로 수행하는 AI 시스템이다. 5개 학문 분야에 걸친 36개 실제 데이터 사례 전부에서 원본 데이터부터 완성된 논문까지의 전체 경로를 완료했고, 주력 모델 기준 평균 논문 점수 10점 만점에 6.3점을 받았다. 원본 대신 미리 계산된 숫자만 보는 블라인드 버전과 비교했을 때, 원본을 직접 보는 버전이 7개 평가 항목 전부에서 더 높은 점수를 받았고 맞대결의 85%에서 승리했다.
METAL MEDIA 해설 도표
OmniScientist의 3단계 파이프라인과 지각 레이어
증거 상태측정 결과가 보고됨
- 지각 레이어이미지, 신호, 3D 구조, 표, 그래프 등 원본 증거를 4개 계열(지각/기호/정량통계/절차)로 분류해 직접 읽어들인다
- 아이디어 구상(Ideation)자료를 관찰하고 문헌을 검색해 검증 가능한 가설을 세우며, 코드 기반 체크가 구조적 완전성과 참신성을 검증한다
- 실험(Experiment)코드를 실행해 가설을 검증하고, 실행 출처와 통계적 타당성을 검증하는 엄격성 체크를 통과해야 한다
- 작성(Writeup)검증된 실행 기록에서만 주장을 골라 학문 분야별 서식에 맞춰 원고를 작성하고, 주장 검증 체크로 숫자와 진술을 원본 실행 기록과 대조한다
- 블라인드 비교 실험원본 데이터 대신 사전 계산된 스칼라 특징만 받는 시스템과 쌍대 비교해 지각의 기여도를 측정한다
무엇을 했나
- OmniScientist는 지각(관찰) 레이어와 아이디어 구상·실험·작성을 맡는 3개의 자율 에이전트를 결합해, 통제된 파이프라인 안에서 원본 관찰 내용이 연구 질문, 실험 설계, 최종 주장까지 바꿀 수 있게 한다.
- 증거를 지각(이미지, 스펙트럼), 기호(텍스트, 수식, 그래프), 정량-통계(표), 절차(궤적, 시뮬레이션) 4개 계열로 분류하고, 참신성, 통계적 타당성, 보고된 모든 숫자의 추적 가능성을 코드로 강제 검증한다.
- 이미지, 오디오, 비디오, 3D 구조, 지식 그래프 등을 포함한 5개 학문 분야, 36개 사례로 구성된 모음에서 시험했으며, 36개 사례 전부에서 원본 데이터부터 완성된 원고까지 완주했다.
- 미리 계산된 스칼라 특징만 받는 블라인드 버전과의 쌍대 비교에서, 직접 지각하는 버전이 7개 심사 항목 전부에서 더 나은 점수를 받았고 맞대결의 85%를 이겼다.
- 두 사례 연구에서 그 효과를 보여준다. 지진 사례에서는 '노이즈'로 라벨된 파형의 21.7%가 실제 지진 신호였음을 찾아냈고, 흉부 X선 사례에서는 국소 엔트로피(질감) 패턴이 폐렴과 정상을 구분하는 데 AUC 최대 0.851을 기록해 단순 픽셀 기반 기준선(0.634)을 크게 앞질렀다.

| Discipline | Representative dataset | Evidence | Modality | N |
|---|---|---|---|---|
| Physical sciences | 5 cases | |||
| Condensed matter / nano | NFFA-EUROPE (Aversa et al. 2018) | 2,655 | ||
| Vibrational spectroscopy | RRUFF (Lafuente et al. 2015) | 2,000 | ||
| Materials informatics | UCI superconductor (Hamidieh 2018) | 21,263 | ||
| Molecular chemistry | PubChem (Kim et al. 2025) | 30 | ||
| Symbolic regression | Feynman (Udrescu and Tegmark 2020) | 12 | ||
| Earth & space | 9 cases | |||
| Remote sensing | EuroSAT (Helber et al. 2019) | 5,000 | ||
| Galaxy morphology | Galaxy Zoo (Lintott et al. 2008) | 1,000 | ||
| Galaxy cross-survey | GZ DECaLS (Walmsley et al. 2022) | 210 | ||
| Gravitational waves | GWOSC (LIGO-Virgo Collaboration 2021) | 1,500 | ||
| Seismology | STEAD (Mousavi et al. 2019) | 1,500 | ||
| Marine biology | WHOI-Plankton (Orenstein et al. 2015) | 3,000 | ||
| Geology / petrophysics | Digital Rocks (Prodanović et al. 2015) | 375 | ||
| Meteorology | SEVIR (Veillette et al. 2020) | 384 | ||
| Cyclone dynamics | IBTrACS (Knapp et al. 2010) | 400 | ||
| Life & medical | 7 cases | |||
| Pathology | Kather CRC (Kather et al. 2016) | 5,000 | ||
| Radiology | Chest X-ray (Kermany et al. 2018) | 3,000 | ||
| Medical imaging | MedMNIST CT (Yang et al. 2023a) | 1,496 | ||
| Cardiology | CinC 2016 (Liu et al. 2016) | 2,000 | ||
| Sleep neuroscience | Sleep-EDF (Kemp et al. 2000) | 1,520 | ||
| Cell biology | Cell Tracking Ch. (Ulman et al. 2017) | 280 | ||
| Genomics | DNA (H3) (Nguyen et al. 2016) | 10,000 | ||
| Agricultural & ecological | 8 cases | |||
| Plant pathology | PlantVillage (Hughes and Salathé 2015) | 3,002 | ||
| Precision agriculture | Indian Pines (Baumgardner et al. 2015) | 2,000 | ||
| Animal behavior | CalMS21 (Sun et al. 2021) | 2,500 | ||
| Ecoacoustics | Bird Audio Det. (Stowell et al. 2019) | 2,000 | ||
| Marine bioacoustics | Watkins MMSD (Sayigh et al. 2016) | 1,697 |

| Evidence family | Typical artifacts |
|---|---|
| Perceptual | Images, video, micrographs, radar, astronomical and remote-sensing imagery, the visual form of scientific plots, audio, and 3-D structure. |
| Symbolic | Natural-language documents, formulae, variables, rules, sequences, knowledge graphs, logical and causal relations, mathematical models. |
| Quantitative-statistical | Tables, measurements, distributions, curves, correlations, significance tests, regression results. |
| Procedural / dynamic | Experimental steps, code execution, agent traces, simulations, dynamic evolution, protocols. |

| Standard peer-review | MM-mandatory | |||||||
|---|---|---|---|---|---|---|---|---|
| Backbone | Novelty↑ | Sound.↑ | Clarity↑ | Signif.↑ | Reprod.↑ | MM-grnd↑ | Factual↑ | Overall↑ |
| Sonnet 5 (Anthropic 2026) | 6.3 | 7.0 | 7.0 | 6.3 | 6.1 | 5.1 | 7.7 | 6.3 |
| GPT-5.6 (OpenAI 2026) | 5.2 | 6.3 | 6.3 | 5.0 | 5.2 | 4.2 | 7.7 | 5.6 |
| GLM-5.2 (Zhipu 2026) | 6.2 | 7.1 | 6.8 | 6.4 | 5.9 | 6.6 | 7.5 | 6.5 |
| Kimi K2.7 (Kimi 2025) | 6.2 | 7.2 | 6.7 | 6.2 | 5.5 | 5.8 | 8.0 | 6.2 |
| Qwen3.5-122B (Qwen Team 2026) | 4.7 | 5.5 | 6.2 | 4.8 | 4.8 | 4.8 | 6.5 | 5.1 |
| Qwen3.5-27B (Qwen Team 2026) | 5.0 | 5.6 | 5.9 | 4.9 | 4.6 | 4.9 | 6.4 | 5.1 |
| Qwen3.5-9B (Qwen Team 2026) | 4.0 | 4.1 | 4.8 | 3.7 | 3.7 | 3.9 | 4.8 | 4.0 |
| Gemma-4-31B (Google 2026) | 4.7 | 5.0 | 5.6 | 4.5 | 4.4 | 4.6 | 6.5 | 4.8 |
| Gemma-4-26B (Google 2026) | 4.4 | 4.4 | 5.0 | 4.0 | 3.7 | 3.8 | 5.1 | 4.2 |
| Backbone | Sonnet 5 | Sonnet | 5 | GPT 5.6 | GPT | 5.6 | GLM 5.2 | GLM | 5.2 | Kimi K2.7 | Kimi | K2.7 | Qwen3.5 122B | Qwen3.5 | 122B | Qwen3.5 27B | Qwen3.5 | 27B | Qwen3.5 9B |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Sonnet | |||||||||||||||||||
| 5 | |||||||||||||||||||
| GPT | |||||||||||||||||||
| 5.6 | |||||||||||||||||||
| GLM | |||||||||||||||||||
| 5.2 | |||||||||||||||||||
| Kimi | |||||||||||||||||||
| K2.7 | |||||||||||||||||||
| Qwen3.5 | |||||||||||||||||||
| 122B | |||||||||||||||||||
| Qwen3.5 | |||||||||||||||||||
| 27B | |||||||||||||||||||
| Qwen3.5 | |||||||||||||||||||
| 9B | |||||||||||||||||||
| Gemma-4 | |||||||||||||||||||
| 31B | |||||||||||||||||||
| Gemma-4 | |||||||||||||||||||
| 26B | |||||||||||||||||||
| Cases | 36 | 10 | 18 | 9 | 34 | 36 | 32 | 36 | 34 | ||||||||||
| Completed↑ | 36 | 9 | 17 | 6 | 30 | 32 | 18 | 32 | 25 | ||||||||||
| Mean↑ | 6.5 | 5.7 | 6.7 | 6.5 | 5.4 | 5.3 | 4.1 | 5.0 | 4.3 |
| By evidence modality | By discipline | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Backbone | Image | Signal | Audio | Video | 3-D | Traj. | T&S | Earth | Life | Agri. | Engin. | Phys. |
| Sonnet 5 | 6.4 | 6.1 | 7.1 | 6.4 | 7.0 | 6.3 | 6.4 | 6.5 | 6.7 | 6.8 | 6.6 | 5.8 |
| GPT-5.6 | 5.9 | 5.5 | 5.5 | – | 5.9 | – | – | 5.6 | 6.0 | 6.4 | 5.2 | – |
| Qwen3.5-27B | 5.5 | 5.0 | 5.9 | 5.3 | 4.8 | 4.9 | 5.8 | 5.1 | 5.5 | 5.5 | 4.8 | 5.5 |
| Qwen3.5-9B | 3.8 | 5.5 | 4.1 | 4.5 | 4.4 | 4.3 | 3.2 | 4.1 | 4.5 | 4.0 | 4.4 | 3.2 |
| Qwen3.5-122B | 4.8 | 5.4 | 5.9 | 5.2 | 5.4 | 5.8 | 5.5 | 4.6 | 5.7 | 6.0 | 4.9 | 5.4 |
| Gemma-4-31B | 5.3 | 4.9 | 5.4 | 5.1 | 5.9 | 4.0 | 3.5 | 5.1 | 5.7 | 4.6 | 5.0 | 4.9 |
| Gemma-4-26B | 4.1 | 4.3 | 4.4 | 5.0 | 4.3 | 4.3 | 4.8 | 4.5 | 3.8 | 4.3 | 4.4 | 4.6 |
| GLM-5.2 | 6.6 | 6.8 | 6.6 | – | 6.6 | – | – | 6.4 | 6.9 | 6.5 | 6.4 | 7.5 |
| Kimi K2.7 | 6.8 | 6.7 | 5.9 | – | 6.0 | – | – | 6.8 | 6.2 | 6.6 | 6.0 | – |
| Standard peer-review | MM-mandatory | |||||||
|---|---|---|---|---|---|---|---|---|
| Discipline | Novelty↑ | Sound.↑ | Clarity↑ | Signif.↑ | Reprod.↑ | MM-gr.↑ | Factual↑ | Overall↑ |
| Physical sciences | ||||||||
| Condensed matter | 2.5 | 2.5 | 4.0 | 2.0 | 2.0 | 4.5 | 1.0 | 2.5 |
| Vibrational spectroscopy | 6.0 | 8.0 | 7.5 | 7.0 | 7.0 | 6.0 | 9.0 | 7.0 |
| Materials informatics | 7.0 | 7.0 | 7.5 | 8.0 | 7.0 | 4.5 | 7.0 | 7.0 |
| Molecular chemistry | 4.5 | 5.5 | 6.0 | 3.5 | 6.0 | 4.0 | 5.0 | 4.5 |
| Symbolic regression | 7.0 | 8.0 | 7.5 | 7.5 | 7.5 | 3.5 | 9.0 | 7.0 |
| Earth & space | ||||||||
| Remote sensing | 7.0 | 7.0 | 6.5 | 7.0 | 6.0 | 5.5 | 7.0 | 6.5 |
| Galaxy morphology | 6.5 | 6.5 | 7.0 | 6.0 | 6.0 | 5.5 | 7.5 | 6.5 |
| Galaxy cross-survey | 7.0 | 7.0 | 6.5 | 6.0 | 6.0 | 5.0 | 7.5 | 6.5 |
| Gravitational waves | 5.0 | 5.0 | 5.5 | 5.0 | 5.0 | 3.5 | 5.5 | 4.5 |
| Seismology | 6.5 | 8.0 | 7.0 | 7.5 | 6.5 | 4.5 | 8.5 | 7.0 |
| Marine biology | 6.0 | 8.5 | 7.5 | 7.0 | 6.5 | 5.5 | 9.5 | 7.0 |
| Geology / petrophysics | 6.5 | 7.5 | 8.0 | 7.5 | 7.5 | 6.0 | 9.5 | 7.5 |
| Cyclone dynamics | 6.0 | 6.5 | 7.0 | 6.0 | 5.5 | 4.5 | 7.0 | 6.0 |
| Meteorology | 6.0 | 7.0 | 8.0 | 6.0 | 5.0 | 4.5 | 8.0 | 6.5 |
| Life & medical | ||||||||
| Pathology | 7.0 | 8.0 | 8.0 | 6.5 | 6.5 | 6.5 | 9.0 | 7.0 |
| Radiology | 7.0 | 7.5 | 8.0 | 6.5 | 6.5 | 6.5 | 8.5 | 7.0 |
| Medical imaging (CT) | 6.5 | 7.5 | 6.5 | 6.0 | 7.0 | 4.0 | 9.0 | 6.0 |
| Cardiology | 6.5 | 8.0 | 8.0 | 7.5 | 6.5 | 4.5 | 9.0 | 7.0 |
| Sleep neuroscience | 6.5 | 4.5 | 7.0 | 6.0 | 5.5 | 3.5 | 4.0 | 4.5 |
| Genomics | 7.0 | 7.0 | 7.0 | 6.0 | 6.0 | 5.0 | 7.5 | 6.5 |
| Cell biology | 7.0 | 6.0 | 5.5 | 6.5 | 5.5 | 4.0 | 6.0 | 5.5 |
| Agricultural & ecological | ||||||||
| Plant pathology | 6.5 | 7.5 | 8.0 | 7.0 | 6.5 | 6.0 | 8.0 | 7.0 |
| Precision agriculture | 6.0 | 6.5 | 6.5 | 6.0 | 6.0 | 5.0 | 8.0 | 6.0 |
| Animal behavior | 6.5 | 7.5 | 7.0 | 5.5 | 5.5 | 4.5 | 8.0 | 6.5 |
| Ecoacoustics | 7.0 | 8.0 | 8.0 | 7.0 | 6.5 | 5.0 | 9.0 | 7.5 |
| Case | Evidence only the raw record carries | Question each system asked |
|---|---|---|
| Galaxy cross-survey | Morphology read off the image | With perception: does a vision model’s morphological reading degrade on the shallower survey for the same galaxies? Blind: can the 3 classes be separated along the 8 supplied feature axes? |
| Seismology | Cross-component polarisation of the waveform | With perception: what fraction of noise-labelled traces carry coherent polarised transients? Blind: do frequency-shape features retain a depth imprint after an attenuation correction? |
| Pathology | Texture and nuclear density of the tile | With perception: is the complex class a compositional mixture of the pure tissue prototypes? Blind: does the class confusion matrix follow an a priori similarity ranking? |
| Mechanical CAD | Principal-axis geometry of the point cloud | With perception: do the function-defined labels correspond to latent geometric morphotypes? Blind: do dimension-standardised part families show tighter descriptor dispersion? |
| Plant phenotyping | Per-point organ labels across repeated scans | With perception: do the two species differ in how many leaves grow at once? Blind: do the species separate on shape descriptors once the size axis is removed? |
| Discipline | Evidence | Evaluation metric | Headline finding |
|---|---|---|---|
| Radiology | image | supported | Pneumonic pediatric lung fields show markedly higher local-entropy heterogeneity (patchiness) than normal. |
| Pathology | image | supported | The COMPLEX H&E class is heterogeneous, splitting into compositional sub-clusters. |
| Galaxy morphology | image | mixed | VLM morphology accuracy 83.8% (DECaLS) vs 81.0% (SDSS); the 2.8-pt gap is not significant. |
| Remote sensing | image | mixed | Color-only features recover 76.2% of 10-class accuracy vs 83.2% combined, revealing a color shortcut. |
| Seismology | signal | mixed | 21.7% of noise-labelled STEAD traces carry coherent transient bursts; the instrument-type hypothesis is refuted. |
| Cardiology | audio | supported | Recording-protocol metadata alone predicts abnormality (AUC 0.60) and collapses out-of-source (0.35), exposing a confound. |
| Ecoacoustics | audio | supported | A mid/high-band bird-presence classifier shows a large, robust drop in discriminability across recording sets. |
| Mechanical CAD | 3-D | supported | Scale-invariant shape descriptors cluster 1,500 CAD parts into function-agnostic form families without supervision. |
| Plant phenotyping | 3-D | mixed | Maize initiates leaves sequentially where tomato is bursty, separable in 3-D scans. |
| Materials informatics | table | mixed | Random k-fold CV underestimates extrapolation error; leave-one-family-out RMSE is 3.1–7.0× higher. |
| Symbolic regression | formula | supported | The Cramér–Rao form Var(a^)=σ2/(NVarlogx) predicts empirical exponent-estimation variance across all 8 monomial Feynman laws, sampled ranges, and noise levels. |
| Knowledge engineering | graph | supported | Disease-associated proteins carry more distinct GO-function annotations than degree-matched controls, and the excess grows with PPI degree; it replicates on withheld test-split edges. |
| CS/ML methodology | trace | refuted | Rejection-driven repairs are not dominated by omission, contrary to the pre-registered hypothesis. |
| Headline and controls | Heterogeneity and robustness | ||
|---|---|---|---|
| Full-detector prevalence | 21.7% (163/750) | Channel BH / HH / HN prevalence | 32.8 / 29.1 / 2.4% |
| 95% confidence interval | [18.8, 24.9]% | channel χ2 | p=5.7×10−14 |
| Amplitude-only baseline | 2.0% | Per-network prevalence range | 0–65% |
| Ablation, no coincidence term | 2.0% | network χ2 | p=3.5×10−14 |
| Null false-alarm rate (target 1%) | 1.07% | Station-cluster bootstrap 95% CI | [17.9, 25.8]% |
| Sensitivity (FAR, null, window) | stable 19–25% |
| Effect and generalisation | Discrimination and robustness | ||
|---|---|---|---|
| Patchiness effect size, all images | d=1.25 | AUC, mean entropy only | 0.840 |
| development split | d=1.26 | AUC, mean + patchiness | 0.851 |
| held-out split | d=1.29 | AUC, patchiness only | 0.847 |
| Label difference, Mann-Whitney | p<0.0001 | AUC, raw-pixel baseline | 0.634 |
| Independent of the mean level | p=1.7×10−5 | Window ablation, 8 / 16 / 32 px | d=1.63 / 1.50 / 1.36 |
| Region-of-interest sweep | d=1.18 to 1.35 |

| Backbone | Tokens in/out↓ | Cache hit↑ | $ / paper↓ | Wall-clock↓ |
|---|---|---|---|---|
| Sonnet 5 | 98k / 191k | 93% | $2.63 | 29 min |
| GPT-5.6 | 122k / 112k | 89% | $4.34 | 12 min |
| Qwen3.5-27B† | 1.4M / 75k | – | $0.06 | 30 min |
| Gemma-4-31B† | 714k / 41k | – | $0.03 | 14 min |

| Standard peer-review | MM-mandatory | |||||||
|---|---|---|---|---|---|---|---|---|
| Case (Δ = on − baseline) | Novelty↑ | Sound.↑ | Clarity↑ | Signif.↑ | Reprod.↑ | MM-gr.↑ | Factual↑ | Overall↑ |
| Galaxy cross-survey | +2.5 | +1.2 | +1.0 | +2.3 | +0.7 | +1.3 | +2.0 | +1.7 |
| Seismology | +1.0 | +1.5 | -1.0 | +1.5 | +0.5 | +1.5 | +0.5 | +1.5 |
| Pathology | +1.5 | +1.0 | +0.5 | +1.0 | +0.5 | +3.0 | -0.5 | +1.5 |
| Mechanical CAD | +0.5 | +0.5 | +3.5 | +2.5 | +2.5 | +1.0 | +2.0 | +3.0 |
| Plant phenotyping | +1.0 | +0.0 | +0.0 | +0.5 | +0.7 | +4.0 | +0.3 | +0.3 |
| Cardiology† | +0.3 | +0.3 | +0.3 | +0.0 | +1.3 | -1.7 | +0.0 | +0.3 |
| Macro-average Δ | +1.14 | +0.75 | +0.72 | +1.31 | +1.03 | +1.53 | +0.72 | +1.39 |
| Validity check | Statistic | Target | Measured |
|---|---|---|---|
| Inter-judge agreement | Krippendorff α | >0.6 | 0.66 |
| Self-preference bias | own − others | ≈0 | 0† |
| Verbosity bias | score vs. length ρ | ≈0 | 0.16 |
| Condition | What it demands of the stage output |
|---|---|
| Idea check (ideation) | |
| Schema | A research question, a hypothesis, an experiment protocol, and a falsification criterion, all non-empty. |
| Breadth | At least 5 self-screened candidate projects, each rated for novelty risk. |
| Prior art | At least 3 focused literature searches, one of them aimed at the selected idea specifically. |
| Feasibility | The selected idea marked fully computational. A proposal that would need a physical experiment is refused outright. |
| Minimal claim | The smallest claim worth publishing if the rest of the study fails, stated separately from the hypothesis. |
| Novelty evidence | What the searches returned for and against this particular idea, with citations. |
| Claim scope | An explicit statement of what the data cannot establish, separating the measured proxy from any mechanistic or causal reading. |
| Effective sample | The decisive-event count for the key test, estimated from the real data counts, and whether it is adequate. |
| Leakage | Whether any step uses ground-truth labels at decision time, and what the label-agnostic counterpart is. |
| Visual audit | Required once the agent has looked at any raw item: which groups it viewed, and how a disagreement with the given label was resolved. |
| Novelty language | Absolute-novelty phrasing (“first”, “unstudied”, “no prior work”) is rejected, because a bounded search cannot support it. |
| Rigour check (experiment) | |
| Verdict | One of supported, refuted, mixed, null, or infeasible. An honest negative is a valid exit. |
| Real execution | At least one run_python call that exited 0 and produced real output. |
| Perception | A case that carries a look_at_* budget cannot finalise a positive verdict without having looked at the raw evidence at least once. |
| Key numbers | The decisive numbers the code printed, reported as a structured record. |
| Provenance | Every reported number must appear in the text of a real run_python output. |
| Real data | Some run must have loaded the actual data rather than hand-coded rows. |
| Reproducibility | At least 60% of the reported numbers present in the union of the run’s real outputs. |
| Multiple tests | Two or more reported p-values require a stated count of every test run and the correction applied to that count. |
| Circularity | Whether the predictor derives from the same representation whose behaviour it predicts, and how that is handled. |
| Full battery | At least 4 analyses, each with a saved figure: primary, baseline, ablation, mechanism, breakdown, sensitivity. |
| Lead selection | The headline must name an existing analysis, be listed first, and not also appear among the demoted ones. |
| Lead significance | A lead whose p-values are all ≥0.05 is rejected unless the verdict is itself null or insufficient. |
| Demotion | A non-significant analysis that is not the lead must be demoted, which keeps it in the trace and out of the manuscript. |
| Correction base | The stated correction count must cover the demoted analyses too, so demoting cannot shrink the denominator. |
| Claim check (writeup) | |
| Traceability | Every number in the drafted text is matched against the grounded set derived from the experiment record. |
| Guarded revision | The prose-polish pass is reverted wholesale if it alters a number, a citation, a claim, or a model name. |
| Condition | What fired | Rejections | Runs |
|---|---|---|---|
| Idea check (ideation): 28 rejections in 23 of the 36 runs | |||
| Schema | a required ideation field left empty | 12 | 10 |
| Effective sample | the decisive-event estimate left empty | 4 | 4 |
| Novelty evidence | search evidence for the selected idea left empty | 3 | 3 |
| Visual audit | images inspected but the audit left empty | 3 | 3 |
| Claim scope | what the data cannot establish left empty | 3 | 3 |
| Novelty language | absolute-novelty phrasing in the proposal | 2 | 2 |
| Breadth | fewer than 5 screened candidates | 1 | 1 |
| Rigour check (experiment): 87 rejections in 32 of the 36 runs | |||
| Demotion | a non-significant analysis left in the paper | 51 | 26 |
| Schema | verdict left empty | 21 | 16 |
| Lead selection | lead unset, not listed first, or a bad demotion name | 9 | 8 |
| Schema | no key numbers reported | 3 | 3 |
| Provenance | a reported number absent from real standard output | 1 | 1 |
| Perception | finalised without looking at the raw evidence | 1 | 1 |
| Multiple tests | the stated test count missing or under-counted | 1 | 1 |
| Total | 115 |
| Ideation | Experiment | |||||
|---|---|---|---|---|---|---|
| Per run | Mean | Median | Max | Mean | Median | Max |
| Agent steps | 8.8 | 9 | 19 | 36.0 | 37 | 49 |
| Tool calls, all kinds | 19.9 | 16.5 | 49 | 37.2 | 38 | 51 |
| Code executions (run_python) | 0.0 | 0 | 0 | 31.8 | 33 | 47 |
| Literature search calls | 8.7 | 9 | 10 | 0.0 | 0 | 0 |
| Perception calls, all channels | 8.4 | 4.5 | 38 | 1.0 | 0 | 7 |
| of which visual (look_at_*) | 4.5 | 3 | 18 | 0.9 | 0 | 7 |
| Exit-check rejections | 0.8 | 1 | 4 | 2.4 | 2 | 7 |
| Outcome over the 36-case suite | Count | Share |
|---|---|---|
| Self-reported verdict | ||
| Supported, the pre-specified hypothesis held | 16 | 44% |
| Mixed, part of the hypothesis held | 17 | 47% |
| Refuted, the pre-specified hypothesis did not hold | 2 | 6% |
| No verdict, the experiment stage exhausted its step budget | 1 | 3% |
| Artifacts produced | ||
| Manuscripts drafted in full, with their figures | 36 | 100% |
| Experiment stages that exited through the rigour check | 35 | 97% |
| Manuscripts scored by the full 2-judge panel | 36 | 100% |
| Analyses per run | ||
| Analyses carried into the manuscript | 265 | 7.4 / run |
| Analyses demoted to the trace | 67 | 1.9 / run |
| Tool | Modality | What it returns | Cases | Calls |
|---|---|---|---|---|
| Visual channel: render the artifact, then look at it | ||||
| look_at_image | image | the VLM’s reading of specific image files | 11 | 49 |
| look_at_signal | signal | one time-series panel per channel: onsets, bursts, envelopes | 6 | 47 |
| look_at_3d | 3-D | rendered XY, XZ and YZ projections of a cloud or mesh | 5 | 40 |
| look_at_table | table | shape, columns, dtypes, head and summary statistics | 1 | 22 |
| look_at_audio | audio | the rendered waveform and spectrogram | 4 | 17 |
| look_at_video | video | a sample of frames, inspected together | 3 | 13 |
| look_at_trajectory | trajectory | the path coloured by time, plus its speed profile | 3 | 7 |
| Native channel: read the modality in its own terms, no image | ||||
| analyze_signal | signal | trend, dominant FFT frequencies, peaks, statistics | 6 | 40 |
| analyze_audio | audio | duration, rate, RMS, spectral centroid, zero-crossing | 4 | 39 |
| analyze_3d | 3-D | point count, bounding box, centroid, extent, PCA axes | 5 | 21 |
| read_trace | trace | the ordered sequence of a track or an agent run log | 2 | 20 |
| analyze_trajectory | trajectory | path length, displacement, straightness, speed, turning angles | 3 | 19 |
| analyze_video | video | frame count, rate, resolution, frame-difference motion | 3 | 3 |
| Style | Section order | Abstract |
|---|---|---|
| Machine learning | Introduction, Related Work, Method, Experiments, Conclusion, Limitations | 150–220 |
| Biomedical | Introduction, Results, Discussion, Methods | 150–200 |
| Earth & space | Introduction, Data, Methods, Results, Discussion, Conclusions | 150–250 |
| Physics | Introduction, Theory and Methods, Results, Discussion, Conclusion | 150–250 |
| Chemistry | Introduction, Experimental Section, Results and Discussion, Conclusions | 150–250 |
실제로 확인된 결과
- 36개 사례 전체에서 원본 데이터부터 완성된 논문까지 전체 경로를 완료했고, 기준 추론 모델(Claude Sonnet 5)로 평균 전체 논문 점수 6.3점(10점 만점)을 받았다.
- 전처리된 스칼라 특징만 받는 블라인드 시스템과의 쌍대 비교에서, 직접 지각을 사용하는 시스템이 7개 평가 차원 전부에서 더 높은 점수를 받았고 승패 비교의 85%에서 승리했다.
- 지진학 사례에서 노이즈로 라벨된 파형의 21.7%가 실제 지진 이벤트로 판정되었으며, 이 추정치는 여러 오경보율, 널모델, 윈도우 설정, 417개 관측소에 대한 클러스터 부트스트랩에서도 안정적이었다.
- 소아 흉부 방사선 사진 사례에서 국소 엔트로피 기반 특징이 held-out 테스트에서 AUC 0.840~0.851을 기록해, 원시 픽셀 기반 baseline(0.634)을 상당히 앞질렀다.
- 9개 추론 백본 모델 비교에서 강한 폐쇄형 모델(Sonnet 5, GLM, Kimi)이 가장 크고 균형 잡힌 점수 프로파일을 보였고, 약한 오픈모델(Qwen3.5-9B, Gemma-4-26B)은 특히 novelty와 significance에서 점수가 떨어졌다.
어디에 쓸 수 있나
- 원본 신호·이미지·3D 구조 등 이질적 데이터를 다루는 다분야 연구 워크플로우 자동화 실험에 참고할 수 있다.
- 가설 생성부터 실험 실행, 논문 작성까지 이어지는 파이프라인에서 코드 기반 검증 절차(참신성·통계적 타당성·주장 추적)를 설계할 때 참고할 수 있다.
- 의료 영상이나 지구물리 신호처럼 라벨이 불완전하거나 잡음이 섞인 데이터셋에서 이상 패턴을 재검토하는 용도로 응용 아이디어를 얻을 수 있다.
한계와 남은 검증
- 평가는 저자들이 구성한 36개 사례 모음에 한정되며, 이 사례들이 다른 실제 연구 상황을 얼마나 대표하는지는 별도로 검증되지 않았다.
- 점수는 2명의 LLM 심사위원(자동화된 채점자)이 매긴 것으로, 인간 전문가 동료 심사와의 일치 여부는 이 요약에 포함된 자료로는 확인되지 않는다.
- 블라인드 비교는 5개의 쌍(및 1개의 비전-오프 쌍)에 한정되어 있어 전체 36개 사례로 일반화되었는지는 알 수 없다.
- 지각 도구 목록과 검증 체크는 저자들이 정의한 특정 데이터 형식과 모달리티에 맞춰 설계된 것으로, 다른 종류의 원본 데이터에 대한 일반화는 확인되지 않았다.
- 비용은 백본 모델에 따라 논문당 0.03달러에서 4.34달러로 보고되었으나, 대규모 실제 연구실 도입 시 비용·시간 효율성은 추가 검증이 필요하다.
왜 중요한가
현재 대부분의 AI 과학자 시스템은 텍스트, 라벨, 미리 계산된 요약 숫자로만 데이터를 접하는데, 이 과정에서 발견에 결정적인 패턴이 사라질 수 있다. 이 연구는 AI 에이전트가 원본 신호·이미지·구조를 직접 살피고, 과도한 주장을 코드 수준에서 걸러내도록 설계하면 더 내실 있고 점수가 높은 과학 논문을 만들 수 있음을 보여준다.
이 논문의 용어
- ReAct 루프 · 관찰-추론-행동을 반복하며 스스로 다음 단계를 정하는 에이전트 동작 방식
- 멀티모달 그라운딩(multimodal grounding) · 주장이나 결론이 실제 원본 데이터(이미지·신호 등)에 근거하고 있는 정도를 평가하는 항목
- HARKing · 실험 결과를 본 뒤 마치 처음부터 그 가설을 세웠던 것처럼 사후에 가설을 짜맞추는 행위
- 실행 이력(execution record/provenance) · 코드 실행, 표준출력, 생성된 그림 등 실제로 수행된 과정을 남긴 기록으로 주장의 근거를 추적하는 데 쓰인다
- 블라인드 변이(blind variant) · 원본 데이터 대신 미리 계산된 숫자(스칼라 특징)만 받는 비교용 시스템
최신 논문
- AI 코딩 에이전트에게 과학 소프트웨어 수리를 시켜보니, 절반도 제대로 못 고쳤다AI 코딩 에이전트에게 과학 소프트웨어 수리를 시켜보니, 절반도 제대로 못 고쳤다
- 논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
- 고객상담 AI 상담원이 규정을 '한 번의 행동'이 아니라 '전체 절차'로 지키게 만드는 방법고객상담 AI 상담원이 규정을 '한 번의 행동'이 아니라 '전체 절차'로 지키게 만드는 방법
- 로봇 팔에게 사람의 시연 없이 새 일 시키기, 말 잘하는 AI가 대신 가르친다로봇 팔에게 사람의 시연 없이 새 일 시키기, 말 잘하는 AI가 대신 가르친다
- 에이전트 학습용 환경을 새로 만드는 대신, 기존 환경에 '패치 부품'을 씌워 그 에이전트의 약점에 맞게 바꾸는 방법에이전트 학습용 환경을 새로 만드는 대신, 기존 환경에 '패치 부품'을 씌워 그 에이전트의 약점에 맞게 바꾸는 방법
- AI 모델을 '소유'하지 못한 조직은 안전 통제도 절반밖에 못 한다AI 모델을 '소유'하지 못한 조직은 안전 통제도 절반밖에 못 한다
- AI가 선생님 모델을 따라 배우다가, 정답에 다가가는 '좋은 생각'까지 억누르는 문제를 잡아낸다AI가 선생님 모델을 따라 배우다가, 정답에 다가가는 '좋은 생각'까지 억누르는 문제를 잡아낸다
- AI가 특정 사람 말투를 흉내내도록 시켜봤더니, 결국 AI 자신의 말투에서 못 벗어난다AI가 특정 사람 말투를 흉내내도록 시켜봤더니, 결국 AI 자신의 말투에서 못 벗어난다
METAL MEDIA 최신 기사
그림 출처: Bobo Li et al., arXiv:2608.13558, arxiv-nonexclusive