Mathematics in the age of AI
If AI can eventually solve research-level math proofs, what should the math community actually protect?
This essay, based on a lecture at the 2026 International Congress of Mathematicians, sidesteps the debate over whether AI can already do research-level math and instead assumes it eventually will. It then asks what mathematicians actually value and optimize for, using problem solving as a case study. The author argues that simply producing a correct answer is not enough — understanding, exposition, and community acceptance matter just as much.
METAL MEDIA explanatory visual
If AI can eventually solve research-level math proofs, what should the math community actually protect?
- 01Instead of arguing about current AI math capability, the author adopts a 'Working Hypothesis' that AI will eventually handle a reasonable share of research-level math tasks, and asks what mathematical values would need re-examining as a result.
- 02Using problem solving as a case study, the goal is refined step by step: from 'solve as many problems as possible' to adding correctness verification, clear communication, community digestion/acceptance, and finally incorporation into a field's definitive, canonical theory.
- 03The essay warns that an AI-generated proof can be formally verified yet understood by nobody, or so smoothly polished that the genuinely difficult steps become indistinguishable from routine ones, erasing cues that normally guide readers.
- 04As concrete evidence, the second batch of the First Proof project tested four AI systems on ten genuinely new research-level problems never posted online; seven of the ten received at least one passing, expert-refereed grade, at compute costs of roughly tens to hundreds of dollars per problem.
- 05The piece quotes several recommendations from the June 2026 Leiden Declaration on Artificial Intelligence and Mathematics — endorsed by the International Mathematical Union — on disclosing AI tool use, supporting reviewers, affirming human authorship, and proper attribution.
What they did
- Instead of arguing about current AI math capability, the author adopts a 'Working Hypothesis' that AI will eventually handle a reasonable share of research-level math tasks, and asks what mathematical values would need re-examining as a result.
- Using problem solving as a case study, the goal is refined step by step: from 'solve as many problems as possible' to adding correctness verification, clear communication, community digestion/acceptance, and finally incorporation into a field's definitive, canonical theory.
- The essay warns that an AI-generated proof can be formally verified yet understood by nobody, or so smoothly polished that the genuinely difficult steps become indistinguishable from routine ones, erasing cues that normally guide readers.
- As concrete evidence, the second batch of the First Proof project tested four AI systems on ten genuinely new research-level problems never posted online; seven of the ten received at least one passing, expert-refereed grade, at compute costs of roughly tens to hundreds of dollars per problem.
- The piece quotes several recommendations from the June 2026 Leiden Declaration on Artificial Intelligence and Mathematics — endorsed by the International Mathematical Union — on disclosing AI tool use, supporting reviewers, affirming human authorship, and proper attribution.
![Figure 3. A page from a 1991 paper of Bourgain [3], annotated by my much younger (and very frustrated) self. But by fighting my way through these texts, I came to understand Bourgain’s way of thinking, and in time I actively sought out his papers to read. See also [19], [17].](https://media.metallab.ai/papers/2608.16753/f0.jpg)
Why it matters
As AI starts generating proofs faster than the community can verify, write up, referee, or canonicalize them, existing institutions built for scarcity of proofs may break down under abundance. The essay is a reminder — relevant to anyone thinking about knowledge production, not just mathematicians — that a fast, correct answer is not the same as understood, trusted knowledge.
Terms in this paper
- Goodhart's law · When a measure becomes a target, it stops being a reliable measure of what it was meant to track
- autoformalization · automatically converting human-written math proofs into machine-checkable proof assistant languages like Lean
- canonicalization · the slow process of restating a proven result in its natural, general form and absorbing it into standard textbooks and toolkits
- First Proof project · an independent evaluation testing AI systems on brand-new, never-published research-level math problems
- Leiden Declaration · a set of 23 recommendations on AI and mathematics, published June 2026 and endorsed by the International Mathematical Union
Original abstract (English)
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears
Latest from METAL MEDIA
Figures: jonbaer et al., arXiv:2608.16753, arxiv-nonexclusive