The Problem Is the Problem: Towards Scalable Mathematical Discovery
让AI从海量论文里自动寻找未解决的数学问题、尝试求解并层层筛选,把人力集中在最后少数值得评审的成果上
这篇论文指出,当前AI辅助数学研究的流程里,人的精力集中在两头——挑选要研究的问题,以及审阅AI给出的结果——这两处正成为瓶颈。作者提出让数学家只需指定一个感兴趣的研究方向,系统FAR(Find, Attempt, and Recommend)便会在文献语料库中自动寻找候选问题、尝试求解,并逐层过滤出少数结果供专家评审。在组合数学领域的试点中,流程从5,245篇论文出发,最终筛出77项交由作者团队评审,其中作者亲自核查的15项全部数学上正确。
METAL MEDIA 解读图
FAR流水线:从文献语料库到专家评审结果
证据状态已报告实测结果
- Find 寻找按研究方向给论文打标签,抽取未解决的陈述,并核查其是否仍未解决,构成含4,717个猜想的可尝试问题池
- Attempt 尝试流水线中最强的模型对问题池中每个猜想各尝试一次,产生1,050个被标记为KNOWN/NEW/FIX/NONE之一的结果
- Recommend 推荐多位评判者核查NEW结果的正确性(598个通过),再由分级智能体将通过项分为已知、次要或可独立发表,最终保留77项
- 专家评审作者团队人工核查其中15项感兴趣的成果,全部确认正确,涉及多个知名猜想与问题
他们做了什么
- 在当前多数AI辅助数学研究的工作流中,人的精力集中在流程的开头(挑选合适的问题)和结尾(审阅AI产出的结果)这两个阶段,作者认为这正在成为高水平数学研究的瓶颈。
- 作者提出一种新范式:数学家不再事先选定单一问题,而只需给出一个自己感兴趣且擅长的研究方向,系统随后在广泛的文献语料库中搜索该方向下的候选问题,这一思路借鉴了搜索与推荐系统。
- 他们构建了FAR(Find, Attempt, and Recommend)流水线,依次完成:给论文打方向标签、从论文中抽取未解决的陈述、核查这些陈述是否仍然悬而未决(构成可尝试问题池),再用较强模型逐一尝试求解,最后对结果进行评判和推荐,交给专家评审。
- 在组合数学试点中,流程从51,110篇数学论文出发,标签阶段保留5,245篇组合数学论文,抽取阶段得到6,453个候选猜想或未解决问题,核查后剩4,717个看似表述良好且仍未解决的猜想构成问题池;后续尝试产生1,050个声称已解决的结果,评判阶段接受598个,分级阶段最终选出77项交由作者团队评审。
- 作者依据自身兴趣人工核查了其中15项,全部在数学上正确,涉及Davies-Jenssen-Perkins-Roberts、Erdős-Straus、Ikenmeyer-Pak-Panova以及Lund-Saraf-Wolf等人提出的猜想或问题。

| Field | Value |
|---|---|
| Source paper | Ikenmeyer, Pak, and Panova, Positivity of the Symmetric Group Characters is as Hard as the Polynomial Time Hierarchy |
| Extracted label | Conjecture 5.3.2 |
| Extracted statement | The problem ComputeCharBinary is 𝖦𝖺𝗉𝖯-complete under many-one reductions. |
| Status | Open. The check found no credible resolution, and records that the completeness question is still unsettled. |
| strategy | B=10 | B=25 | B=50 | B=75 | B=100 | B=200 | B=300 |
|---|---|---|---|---|---|---|---|
| uniform random | 0.17 | 0.43 | 0.86 | 1.29 | 1.72 | 3.44 | 5.17 |
| rank on p^ | 0.40 | 0.97 | 1.64 | 2.29 | 2.95 | 5.42 | 7.92 |
| rank on ip^ | 0.15 | 0.47 | 1.02 | 1.59 | 2.24 | 4.78 | 7.25 |
| rank on p^ inside the top 1/10 | 0.30 | 0.70 | 0.99 | 1.01 | 1.06 | 1.99 | 2.00 |
| rank on p^ inside the top 1/5 | 0.17 | 0.42 | 0.86 | 1.25 | 1.68 | 2.99 | 3.72 |
| strategy | B=10 | B=25 | B=50 | B=75 | B=100 | B=200 | B=300 |
|---|---|---|---|---|---|---|---|
| uniform random | 0.08 | 0.21 | 0.42 | 0.63 | 0.84 | 1.68 | 2.51 |
| rank on p^ | 0.16 | 0.40 | 0.73 | 1.05 | 1.38 | 2.62 | 3.87 |
| rank on ip^ | 0.09 | 0.28 | 0.60 | 0.89 | 1.22 | 2.49 | 3.72 |
| rank on p^ inside the top 1/10 | 0.26 | 0.60 | 0.84 | 0.86 | 0.90 | 1.69 | 1.70 |
| rank on p^ inside the top 1/5 | 0.13 | 0.34 | 0.69 | 1.00 | 1.34 | 2.44 | 3.06 |
| strategy | B=10 | B=25 | B=50 | B=75 | B=100 | B=200 | B=300 |
|---|---|---|---|---|---|---|---|
| uniform random | 0.078 | 0.176 | 0.299 | 0.386 | 0.448 | 0.574 | 0.629 |
| rank on p^ | 0.157 | 0.354 | 0.437 | 0.467 | 0.485 | 0.499 | 0.500 |
| rank on ip^ | 0.090 | 0.280 | 0.546 | 0.597 | 0.600 | 0.600 | 0.600 |
| rank on p^ inside the top 1/10 | 0.259 | 0.596 | 0.841 | 0.850 | 0.850 | 0.850 | 0.850 |
| rank on p^ inside the top 1/5 | 0.126 | 0.294 | 0.540 | 0.690 | 0.779 | 0.850 | 0.850 |

| units digit | X-block | X′-block | selector digit | |
|---|---|---|---|---|
| Q0 | Q1,…,Qk | Qk+1,…,Qk+k′ | Qk+k′+1 | |
| aS (S∈C) | 0 | 𝟏S | 𝟎 | 0 |
| bS′ (S′∈C′) | 0 | 𝟎 | 𝟏S′ | 0 |
| cA | 1 | 𝟎 | 𝟏 | 1 |
| cB | 0 | 𝟏 | 𝟎 | 1 |
| H | 2 | 𝟏 | 𝟏 | 1 |
| s | 1 | 𝟏 | 𝟏 | 1 |
| s−1 | 0 | 𝟏 | 𝟏 | 1 |
研究结果
- 在组合数学试点中,51,110篇数学论文经标签筛选后剩5,245篇,抽取阶段得到6,453个候选猜想或未解决问题,核查后剩4,717个仍未解决、构成可尝试问题池。
- 对问题池中每个猜想各尝试一次,共产生1,050个声称已解决的结果,评判阶段接受了598个,分级阶段最终留下77项作为交由作者团队评审的成果。
- 作者依兴趣人工核查的15项成果全部在数学上正确,包括对Davies-Jenssen-Perkins-Roberts猜想和Lund-Saraf-Wolf猜想的反例、对Ikenmeyer-Pak-Panova猜想的证明,以及对Erdős-Straus提出问题的回答。
- 模型给出的难度分数对'该问题最终未获接受结果'的预测AUC为0.69,重要性分数对'被接受结果是否评为可发表'的预测AUC为0.60,二者的相关性在统计上均显著(p分别小于10^-40和为0.008)。
- 在比较预算分配策略时,按估计的成功概率(或成功概率乘以重要性)排序分配的策略,相比均匀分配基线能产生更多可发表成果,且最优策略会随目标(成果数量、总重要性还是单个最高重要性)不同而改变。
可应用场景
- 在某一特定数学领域(如组合数学)中,自动扫描大量文献以发现仍未解决、值得尝试的猜想或问题
- 在专家评审资源有限的场景下,用多阶段自动化审核流程对AI生成的大量候选证明或反例结果进行预筛选,减轻专家负担
- 为如何在多个候选问题之间分配有限的模型推理预算提供参考,依据目标是追求更多成果数量、更高总重要性还是单个高影响力结果来选择分配策略
局限与待验证事项
- 试点仅在组合数学这一作者本身具备核查能力的领域进行,尚未验证该方法在其他数学领域的适用性。
- 实际运行中每个猜想只分配了一次尝试(相当于老虎机算法的初始化阶段),更复杂、允许重复尝试的分配算法留待未来研究。
- 难度分数和重要性分数均由同一模型给出,二者的斯皮尔曼等级相关系数高达0.83,并非完全独立的信号。
- 流水线自身对'是否已被解决'的核查并不完美,试点中就有一项结果其实在运行前几个月已由其他途径被解决,而级联中的各阶段搜索都未发现这一点。
- 文中给出的最优预算分配策略是在假定预先知道每个问题成功概率和重要性的理论前提下推导的,论文并未报告实际估计误差对分配效果的具体影响。
为什么重要
前沿模型的推理能力和专家的数学评审能力都是稀缺资源,如何合理分配这两种资源直接决定了AI辅助数学研究能否真正提高效率。这项工作实际演示了一种把人力瓶颈从逐一挑选和审阅问题,转移到只需审阅少数经过多层筛选的高质量候选结果的具体做法。
本文术语
- FAR(Find, Attempt, and Recommend) · 一个从文献中寻找问题、尝试求解、再推荐评审结果的三阶段流水线
- 可尝试问题池(𝒫) · 从文献中抽取并核实为仍未解决的、有据可查的候选猜想集合
- 推荐级联(recommendation cascade) · 搜索或推荐系统中常见的做法:用一系列逐渐更严格、更昂贵的过滤步骤,把庞大的候选集缩减到很小的一批
- 多臂老虎机(bandit)视角 · 把有限的求解尝试次数如何分配给多个候选问题,看作多臂老虎机式的资源分配问题
- AUC(ROC曲线下面积) · 衡量一个打分能否把结果更好的项目排在前面的指标,即随机抽一好一坏两项、好的排名更高的概率
论文原文摘要(英文)
AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math workflows, human effort is concentrated at the beginning and end, in selecting suitable research problems and later reviewing the resulting artifacts. These two stages are becoming bottlenecks for research-level mathematics. We address them by proposing a new human-AI discovery paradigm. The human input is no longer a single problem selected in advance, but a research direction in which the experts have interest and expertise. The system then searches a broad literature corpus for candidate problems in that direction. Inspired by search and recommender systems, we build Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering. In a combinatorics pilot, the pipeline starts from 5,245 combinatorics papers, recovers 6,453 candidate conjectures or open problems, and filters them to 4,717 apparently well-posed and still-open conjectures. Subsequent reasoning and automated triage stages surface 598 potential resolutions and select 77 items for author-team review. Among them, we identify many interesting discoveries, including results on conjectures and questions of Davies--Jenssen--Perkins--Roberts, Erdős--Straus, Ikenmeyer--Pak--Panova, and Lund--Saraf--Wolf. These results demonstrate the effectiveness of this new mode of human-AI collaboration for mathematical discovery.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调
METAL MEDIA 最新报道
图片来源: Zeyu Zheng et al., arXiv:2608.16977, cc-by-nc-sa-4.0