Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox

METAL MEDIA

The Problem Is the Problem: Towards Scalable Mathematical Discovery

arXiv:2608.169772026-08-16

A pipeline lets AI search literature for open problems, attempt them, and filter the results so human effort in mathematical research goes only to the small set worth reviewing

The paper argues that in AI-assisted math research, human effort gets stuck at two bottlenecks: picking which problem to work on, and reviewing whatever comes out. Instead of choosing one problem in advance, the authors let a mathematician specify only a research direction, and their FAR (Find, Attempt, and Recommend) pipeline searches a literature corpus, tries candidate problems, and filters the outputs down to a handful for expert review. In a combinatorics pilot starting from 5,245 papers, the pipeline surfaced 77 items for author review, and 15 that the authors manually checked were all mathematically correct.

METAL MEDIA explanatory visual

The FAR pipeline: from a literature corpus to expert-reviewed results

Evidence statusMeasured results reported

  1. FindLabels papers matching the research direction, extracts unresolved statements, and checks which are still open, forming the attemptable pool of 4,717 conjectures
  2. AttemptThe strongest model in the pipeline makes one attempt per conjecture in the pool, producing 1,050 claimed resolutions labeled KNOWN, NEW, FIX, or NONE
  3. RecommendMultiple judges verify correctness of NEW outcomes (598 pass), then a grading agent sorts passed results into already-known, minor, or publishable, keeping 77 artifacts
  4. Expert reviewThe author team manually checks 15 artifacts of interest; all are found mathematically correct, including results tied to several named conjectures
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.

What they did

  1. In most current AI-for-math workflows, human effort concentrates at the start (choosing a problem) and the end (reviewing the AI's output), and the authors argue these two stages are becoming bottlenecks for research-level mathematics.
  2. Instead of a single pre-selected problem, the authors propose that experts specify a research direction, and the system automatically searches a broad literature corpus for candidate problems within it, inspired by search and recommender systems.
  3. They build FAR (Find, Attempt, and Recommend), a cascade with stages that label relevant papers, extract unresolved statements, check whether those statements are still open (forming an attemptable pool), attempt resolutions with a strong model, and judge/recommend the results for expert review.
  4. In a combinatorics pilot, starting from 51,110 papers, labeling kept 5,245 combinatorics papers, extraction recovered 6,453 candidate conjectures/open problems, checking left 4,717 still-open conjectures, attempts produced 1,050 claimed resolutions, judging accepted 598, and grading selected 77 items for author-team review.
  5. The authors manually reviewed 15 of these artifacts chosen by their own interest and found every one mathematically correct, including results touching conjectures/questions of Davies-Jenssen-Perkins-Roberts, Erdős-Straus, Ikenmeyer-Pak-Panova, and Lund-Saraf-Wolf.
Figure 1: From choosing a problem to choosing a direction. The upper part shows the problem-level interface; the lower part shows our approach. Figure 3 gives the details of FAR.
Figure 1: From choosing a problem to choosing a direction. The upper part shows the problem-level interface; the lower part shows our approach. Figure 3 gives the details of FAR.
Table 1: A recovered candidate, from source text to the pool.
FieldValue
Source paperIkenmeyer, Pak, and Panova, Positivity of the Symmetric Group Characters is as Hard as the Polynomial Time Hierarchy
Extracted labelConjecture 5.3.2
Extracted statementThe problem ComputeCharBinary is 𝖦𝖺𝗉𝖯-complete under many-one reductions.
StatusOpen. The check found no credible resolution, and records that the completeness question is still unsettled.
Figure 2: Schematic view of the current reachable region. The figure is illustrative, and both axes should be read qualitatively. A single pass over the pool probes which conjectures can produce artifacts worth review under the current model. As model capability improves, more conjectures may become reachable.
Figure 2: Schematic view of the current reachable region. The figure is illustrative, and both axes should be read qualitatively. A single pass over the pool probes which conjectures can produce artifacts worth review under the current model. As model capability improves, more conjectures may become reachable.
Table 2: Expected number of artifacts, the objective f1.
strategyB=10B=25B=50B=75B=100B=200B=300
uniform random0.170.430.861.291.723.445.17
rank on p^0.400.971.642.292.955.427.92
rank on i​p^0.150.471.021.592.244.787.25
rank on p^ inside the top 1/100.300.700.991.011.061.992.00
rank on p^ inside the top 1/50.170.420.861.251.682.993.72
Figure 3: From papers to recommendations for expert review. The upper row shows a search or recommender pipeline that recalls and filters candidates from a large corpus. Its numbers indicate typical orders of magnitude. The lower row shows the analogous FAR pipeline. Numbers in the lower row are counts from our pilot run detailed in Section 4.
Figure 3: From papers to recommendations for expert review. The upper row shows a search or recommender pipeline that recalls and filters candidates from a large corpus. Its numbers indicate typical orders of magnitude. The lower row shows the analogous FAR pipeline. Numbers in the lower row are counts from our pilot run detailed in Section 4.
Table 3: Expected total importance of the artifacts returned, the objective f2.
strategyB=10B=25B=50B=75B=100B=200B=300
uniform random0.080.210.420.630.841.682.51
rank on p^0.160.400.731.051.382.623.87
rank on i​p^0.090.280.600.891.222.493.72
rank on p^ inside the top 1/100.260.600.840.860.901.691.70
rank on p^ inside the top 1/50.130.340.691.001.342.443.06
Figure 4: Each score against the quantity it judges. Panel (a) plots δ(d−1[a,b)) on an axis starting at 50%, panel (b) plots ι(i−1[a,b)). n counts attempts in (a) and accepted resolutions in (b). Candidates that a later status recheck reclassified as solved or invalid are excluded.
Figure 4: Each score against the quantity it judges. Panel (a) plots δ(d−1[a,b)) on an axis starting at 50%, panel (b) plots ι(i−1[a,b)). n counts attempts in (a) and accepted resolutions in (b). Candidates that a later status recheck reclassified as solved or invalid are excluded.
Table 4: Expected maximum importance among the artifacts returned, the objective f3.
strategyB=10B=25B=50B=75B=100B=200B=300
uniform random0.0780.1760.2990.3860.4480.5740.629
rank on p^0.1570.3540.4370.4670.4850.4990.500
rank on i​p^0.0900.2800.5460.5970.6000.6000.600
rank on p^ inside the top 1/100.2590.5960.8410.8500.8500.8500.850
rank on p^ inside the top 1/50.1260.2940.5400.6900.7790.8500.850
The Problem Is the Problem: Towards Scalable Mathematical Discovery figure 4
Table 5: The base-Q digits of the parts of μ and of the two targets s and s−1. Here 𝟏 and 𝟎 denote the all-ones and all-zeros patterns on the block indicated, and 𝟏S denotes the indicator pattern of S on the X-block, with 𝟏S′ defined analogously on the X′-block. No column can reach Q, so subset sums may be compared digit by digit: a subset of parts summing to s must take cA, hence miss the X′-block entirely and pick out an exact cover of (X,C), while a subset summing to s−1 must take cB and pick out an exact cover of (X′,C′).
units digitX-blockX′-blockselector digit
Q0Q1,…,QkQk+1,…,Qk+k′Qk+k′+1
aS (S∈C)0𝟏S𝟎0
bS′ (S′∈C′)0𝟎𝟏S′0
cA1𝟎𝟏1
cB0𝟏𝟎1
H2𝟏𝟏1
s1𝟏𝟏1
s−10𝟏𝟏1
Figure 5: Allocation strategies against the budget. Each fit is made on four fifths of 𝒫 and applied to the remaining fifth, from which B/5 conjectures are drawn. The five selections together make one set of B conjectures, and each point averages what that set returns over 1000 random partitions. Ties are broken at random, and the uniform baseline is computed exactly from its closed form. Here, B only counts the allocated attempts.
Figure 5: Allocation strategies against the budget. Each fit is made on four fifths of 𝒫 and applied to the remaining fifth, from which B/5 conjectures are drawn. The five selections together make one set of B conjectures, and each point averages what that set returns over 1000 random partitions. Ties are broken at random, and the uniform baseline is computed exactly from its closed form. Here, B only counts the allocated attempts.

Findings

  • In the combinatorics pilot, out of 51,110 math papers, labeling kept 5,245 papers, extraction recovered 6,453 candidate conjectures/open problems, and checking left 4,717 apparently well-posed, still-open conjectures forming the attemptable pool.
  • One attempt per conjecture produced 1,050 claimed resolutions; judging accepted 598 of them, and grading selected 77 as artifacts substantial enough for author-team review.
  • Of 15 artifacts the authors manually checked by their own interest, all were mathematically correct, including counterexamples to Davies-Jenssen-Perkins-Roberts and Lund-Saraf-Wolf conjectures, a proof of the Ikenmeyer-Pak-Panova conjecture, and an answer to an Erdős-Straus question.
  • The model's difficulty score achieved an AUC of 0.69 for predicting whether an attempt would fail to yield an accepted resolution, and the importance score achieved an AUC of 0.60 for predicting whether an accepted resolution would be graded publishable, both statistically significant (p<10^-40 and p=0.008 respectively).
  • Comparing effort-allocation strategies, sorting conjectures by estimated success probability (or by probability times importance) outperformed a uniform-allocation baseline, with the best strategy depending on which objective (number of artifacts, total importance, or max importance) was being maximized.
Figure 6: Each score against the quantity it judges, with Wilson intervals. Panel (a) gives δ against the difficulty score, panel (b) gives ι against the importance score.
Figure 6: Each score against the quantity it judges, with Wilson intervals. Panel (a) gives δ against the difficulty score, panel (b) gives ι against the importance score.

Where it can be used

  • Automatically scanning a large body of literature in a specific mathematical field to surface still-open conjectures or questions worth attempting
  • Using a multi-stage automated triage process to pre-filter AI-generated candidate proofs or counterexamples before scarce expert reviewers spend time on them
  • Informing how to allocate a limited budget of model attempts across many candidate problems depending on whether the goal is more solved problems, higher total importance, or one high-impact result

Limits and open work

  • The pilot is limited to combinatorics, a field the authors themselves have expertise to verify, so generalization to other mathematical domains has not been demonstrated.
  • Only a single-attempt-per-conjecture strategy (the bandit initialization step) was actually run; investigating more sophisticated bandit algorithms with repeated attempts is left for future work.
  • The difficulty and importance scores come from the same model and are highly correlated (Spearman rank correlation 0.83), so they may not provide fully independent signals.
  • The pipeline's own novelty checks are imperfect: one reviewed result had already been solved a few months earlier by a route none of the search stages found.
  • The optimal allocation strategies derived assume the success and importance probabilities of each conjecture are known in advance; the paper does not report how estimation error in real deployments would affect performance.

Why it matters

Frontier-model reasoning and expert mathematical review are both scarce resources, so how they get allocated is central to whether AI can meaningfully speed up mathematical discovery. This work demonstrates a concrete way to move the human bottleneck from picking and reviewing individual problems to reviewing a small, pre-filtered set of promising results surfaced automatically from the literature.

Terms in this paper

  • FAR (Find, Attempt, and Recommend) · A three-stage pipeline that finds candidate problems in literature, attempts to solve them, and recommends the best results for human review
  • attemptable pool (𝒫) · The set of source-grounded conjectures extracted from papers and verified to still be open, ready to be attempted
  • recommendation cascade · A search/recommender-system design where successive stages filter a large candidate set down to a small final set using progressively more expensive checks
  • bandit interpretation · Framing the choice of which problems to attempt as a multi-armed bandit problem, where each conjecture is an arm and an attempt is a pull
  • AUC (area under the ROC curve) · The probability that a score ranks a randomly chosen positive example above a randomly chosen negative one, used here to measure how well difficulty/importance scores predict outcomes

Original abstract (English)

AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math workflows, human effort is concentrated at the beginning and end, in selecting suitable research problems and later reviewing the resulting artifacts. These two stages are becoming bottlenecks for research-level mathematics. We address them by proposing a new human-AI discovery paradigm. The human input is no longer a single problem selected in advance, but a research direction in which the experts have interest and expertise. The system then searches a broad literature corpus for candidate problems in that direction. Inspired by search and recommender systems, we build Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering. In a combinatorics pilot, the pipeline starts from 5,245 combinatorics papers, recovers 6,453 candidate conjectures or open problems, and filters them to 4,717 apparently well-posed and still-open conjectures. Subsequent reasoning and automated triage stages surface 598 potential resolutions and select 77 items for author-team review. Among them, we identify many interesting discoveries, including results on conjectures and questions of Davies--Jenssen--Perkins--Roberts, Erdős--Straus, Ikenmeyer--Pak--Panova, and Lund--Saraf--Wolf. These results demonstrate the effectiveness of this new mode of human-AI collaboration for mathematical discovery.

Authors · Zeyu Zheng

Read on arXiv

Latest papers

All papers →

Latest from METAL MEDIA

Figures: Zeyu Zheng et al., arXiv:2608.16977, cc-by-nc-sa-4.0