The Problem Is the Problem: Towards Scalable Mathematical Discovery
A pipeline lets AI search literature for open problems, attempt them, and filter the results so human effort in mathematical research goes only to the small set worth reviewing
The paper argues that in AI-assisted math research, human effort gets stuck at two bottlenecks: picking which problem to work on, and reviewing whatever comes out. Instead of choosing one problem in advance, the authors let a mathematician specify only a research direction, and their FAR (Find, Attempt, and Recommend) pipeline searches a literature corpus, tries candidate problems, and filters the outputs down to a handful for expert review. In a combinatorics pilot starting from 5,245 papers, the pipeline surfaced 77 items for author review, and 15 that the authors manually checked were all mathematically correct.
METAL MEDIA explanatory visual
The FAR pipeline: from a literature corpus to expert-reviewed results
Evidence statusMeasured results reported
- FindLabels papers matching the research direction, extracts unresolved statements, and checks which are still open, forming the attemptable pool of 4,717 conjectures
- AttemptThe strongest model in the pipeline makes one attempt per conjecture in the pool, producing 1,050 claimed resolutions labeled KNOWN, NEW, FIX, or NONE
- RecommendMultiple judges verify correctness of NEW outcomes (598 pass), then a grading agent sorts passed results into already-known, minor, or publishable, keeping 77 artifacts
- Expert reviewThe author team manually checks 15 artifacts of interest; all are found mathematically correct, including results tied to several named conjectures
What they did
- In most current AI-for-math workflows, human effort concentrates at the start (choosing a problem) and the end (reviewing the AI's output), and the authors argue these two stages are becoming bottlenecks for research-level mathematics.
- Instead of a single pre-selected problem, the authors propose that experts specify a research direction, and the system automatically searches a broad literature corpus for candidate problems within it, inspired by search and recommender systems.
- They build FAR (Find, Attempt, and Recommend), a cascade with stages that label relevant papers, extract unresolved statements, check whether those statements are still open (forming an attemptable pool), attempt resolutions with a strong model, and judge/recommend the results for expert review.
- In a combinatorics pilot, starting from 51,110 papers, labeling kept 5,245 combinatorics papers, extraction recovered 6,453 candidate conjectures/open problems, checking left 4,717 still-open conjectures, attempts produced 1,050 claimed resolutions, judging accepted 598, and grading selected 77 items for author-team review.
- The authors manually reviewed 15 of these artifacts chosen by their own interest and found every one mathematically correct, including results touching conjectures/questions of Davies-Jenssen-Perkins-Roberts, Erdős-Straus, Ikenmeyer-Pak-Panova, and Lund-Saraf-Wolf.

| Field | Value |
|---|---|
| Source paper | Ikenmeyer, Pak, and Panova, Positivity of the Symmetric Group Characters is as Hard as the Polynomial Time Hierarchy |
| Extracted label | Conjecture 5.3.2 |
| Extracted statement | The problem ComputeCharBinary is 𝖦𝖺𝗉𝖯-complete under many-one reductions. |
| Status | Open. The check found no credible resolution, and records that the completeness question is still unsettled. |
| strategy | B=10 | B=25 | B=50 | B=75 | B=100 | B=200 | B=300 |
|---|---|---|---|---|---|---|---|
| uniform random | 0.17 | 0.43 | 0.86 | 1.29 | 1.72 | 3.44 | 5.17 |
| rank on p^ | 0.40 | 0.97 | 1.64 | 2.29 | 2.95 | 5.42 | 7.92 |
| rank on ip^ | 0.15 | 0.47 | 1.02 | 1.59 | 2.24 | 4.78 | 7.25 |
| rank on p^ inside the top 1/10 | 0.30 | 0.70 | 0.99 | 1.01 | 1.06 | 1.99 | 2.00 |
| rank on p^ inside the top 1/5 | 0.17 | 0.42 | 0.86 | 1.25 | 1.68 | 2.99 | 3.72 |
| strategy | B=10 | B=25 | B=50 | B=75 | B=100 | B=200 | B=300 |
|---|---|---|---|---|---|---|---|
| uniform random | 0.08 | 0.21 | 0.42 | 0.63 | 0.84 | 1.68 | 2.51 |
| rank on p^ | 0.16 | 0.40 | 0.73 | 1.05 | 1.38 | 2.62 | 3.87 |
| rank on ip^ | 0.09 | 0.28 | 0.60 | 0.89 | 1.22 | 2.49 | 3.72 |
| rank on p^ inside the top 1/10 | 0.26 | 0.60 | 0.84 | 0.86 | 0.90 | 1.69 | 1.70 |
| rank on p^ inside the top 1/5 | 0.13 | 0.34 | 0.69 | 1.00 | 1.34 | 2.44 | 3.06 |
| strategy | B=10 | B=25 | B=50 | B=75 | B=100 | B=200 | B=300 |
|---|---|---|---|---|---|---|---|
| uniform random | 0.078 | 0.176 | 0.299 | 0.386 | 0.448 | 0.574 | 0.629 |
| rank on p^ | 0.157 | 0.354 | 0.437 | 0.467 | 0.485 | 0.499 | 0.500 |
| rank on ip^ | 0.090 | 0.280 | 0.546 | 0.597 | 0.600 | 0.600 | 0.600 |
| rank on p^ inside the top 1/10 | 0.259 | 0.596 | 0.841 | 0.850 | 0.850 | 0.850 | 0.850 |
| rank on p^ inside the top 1/5 | 0.126 | 0.294 | 0.540 | 0.690 | 0.779 | 0.850 | 0.850 |

| units digit | X-block | X′-block | selector digit | |
|---|---|---|---|---|
| Q0 | Q1,…,Qk | Qk+1,…,Qk+k′ | Qk+k′+1 | |
| aS (S∈C) | 0 | 𝟏S | 𝟎 | 0 |
| bS′ (S′∈C′) | 0 | 𝟎 | 𝟏S′ | 0 |
| cA | 1 | 𝟎 | 𝟏 | 1 |
| cB | 0 | 𝟏 | 𝟎 | 1 |
| H | 2 | 𝟏 | 𝟏 | 1 |
| s | 1 | 𝟏 | 𝟏 | 1 |
| s−1 | 0 | 𝟏 | 𝟏 | 1 |
Findings
- In the combinatorics pilot, out of 51,110 math papers, labeling kept 5,245 papers, extraction recovered 6,453 candidate conjectures/open problems, and checking left 4,717 apparently well-posed, still-open conjectures forming the attemptable pool.
- One attempt per conjecture produced 1,050 claimed resolutions; judging accepted 598 of them, and grading selected 77 as artifacts substantial enough for author-team review.
- Of 15 artifacts the authors manually checked by their own interest, all were mathematically correct, including counterexamples to Davies-Jenssen-Perkins-Roberts and Lund-Saraf-Wolf conjectures, a proof of the Ikenmeyer-Pak-Panova conjecture, and an answer to an Erdős-Straus question.
- The model's difficulty score achieved an AUC of 0.69 for predicting whether an attempt would fail to yield an accepted resolution, and the importance score achieved an AUC of 0.60 for predicting whether an accepted resolution would be graded publishable, both statistically significant (p<10^-40 and p=0.008 respectively).
- Comparing effort-allocation strategies, sorting conjectures by estimated success probability (or by probability times importance) outperformed a uniform-allocation baseline, with the best strategy depending on which objective (number of artifacts, total importance, or max importance) was being maximized.
Where it can be used
- Automatically scanning a large body of literature in a specific mathematical field to surface still-open conjectures or questions worth attempting
- Using a multi-stage automated triage process to pre-filter AI-generated candidate proofs or counterexamples before scarce expert reviewers spend time on them
- Informing how to allocate a limited budget of model attempts across many candidate problems depending on whether the goal is more solved problems, higher total importance, or one high-impact result
Limits and open work
- The pilot is limited to combinatorics, a field the authors themselves have expertise to verify, so generalization to other mathematical domains has not been demonstrated.
- Only a single-attempt-per-conjecture strategy (the bandit initialization step) was actually run; investigating more sophisticated bandit algorithms with repeated attempts is left for future work.
- The difficulty and importance scores come from the same model and are highly correlated (Spearman rank correlation 0.83), so they may not provide fully independent signals.
- The pipeline's own novelty checks are imperfect: one reviewed result had already been solved a few months earlier by a route none of the search stages found.
- The optimal allocation strategies derived assume the success and importance probabilities of each conjecture are known in advance; the paper does not report how estimation error in real deployments would affect performance.
Why it matters
Frontier-model reasoning and expert mathematical review are both scarce resources, so how they get allocated is central to whether AI can meaningfully speed up mathematical discovery. This work demonstrates a concrete way to move the human bottleneck from picking and reviewing individual problems to reviewing a small, pre-filtered set of promising results surfaced automatically from the literature.
Terms in this paper
- FAR (Find, Attempt, and Recommend) · A three-stage pipeline that finds candidate problems in literature, attempts to solve them, and recommends the best results for human review
- attemptable pool (𝒫) · The set of source-grounded conjectures extracted from papers and verified to still be open, ready to be attempted
- recommendation cascade · A search/recommender-system design where successive stages filter a large candidate set down to a small final set using progressively more expensive checks
- bandit interpretation · Framing the choice of which problems to attempt as a multi-armed bandit problem, where each conjecture is an arm and an attempt is a pull
- AUC (area under the ROC curve) · The probability that a score ranks a randomly chosen positive example above a randomly chosen negative one, used here to measure how well difficulty/importance scores predict outcomes
Original abstract (English)
AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math workflows, human effort is concentrated at the beginning and end, in selecting suitable research problems and later reviewing the resulting artifacts. These two stages are becoming bottlenecks for research-level mathematics. We address them by proposing a new human-AI discovery paradigm. The human input is no longer a single problem selected in advance, but a research direction in which the experts have interest and expertise. The system then searches a broad literature corpus for candidate problems in that direction. Inspired by search and recommender systems, we build Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering. In a combinatorics pilot, the pipeline starts from 5,245 combinatorics papers, recovers 6,453 candidate conjectures or open problems, and filters them to 4,717 apparently well-posed and still-open conjectures. Subsequent reasoning and automated triage stages surface 598 potential resolutions and select 77 items for author-team review. Among them, we identify many interesting discoveries, including results on conjectures and questions of Davies--Jenssen--Perkins--Roberts, Erdős--Straus, Ikenmeyer--Pak--Panova, and Lund--Saraf--Wolf. These results demonstrate the effectiveness of this new mode of human-AI collaboration for mathematical discovery.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)Testing AI summaries of stock news, the simple approach beat the trendy retrieval-based one
Latest from METAL MEDIA
Figures: Zeyu Zheng et al., arXiv:2608.16977, cc-by-nc-sa-4.0