MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement
先检索数学库知识、再让编译器和语义判定反复纠错,能让一个8B小模型在把数学题转成可验证代码这件事上打败32B专用模型
自动形式化是把自然语言写的数学命题转换成Lean 4这类机器可验证代码的任务,但它不只是翻译,还需要把数学概念准确映射到Mathlib这个庞大数学库的定义和类型体系中。作者提出MathForm框架,在生成前先从Mathlib检索相关知识,生成后再用编译器诊断和语义一致性反馈进行多轮修正,由此构建了约36.7万条经过验证的数据集FormalVerse。用该数据训练的MathForm-8B在六个基准测试上平均语法通过率达88.06%,语义一致通过率达72.37%,超过多个32B级专用自动形式化模型。
METAL MEDIA 解读图
先检索数学库知识、再让编译器和语义判定反复纠错,能让一个8B小模型在把数学题转成可验证代码这件事上打败32B专用模型
- 01以往方法过度依赖模型记忆的知识,容易引用Mathlib中不存在的定理或违反库的命名习惯,而常见的数据构建流程只是对大批量单次生成结果做事后筛选,无法指出错误出在哪里
- 02MathForm在生成前用检索规划器从Mathlib取回相关定义和已有形式化案例,生成后经过最多三轮修正,依据编译诊断和语义一致性判定(数据构建阶段用QwQ-32B作裁判,最终评测用gpt-oss-120b)不断改进候选结果
- 03基于约36.7万条验证数据FormalVerse,先对Qwen3-8B做监督微调,再用强化学习方法DAPO进一步训练,得到MathForm-8B
- 04在最具挑战性的FATE-H和FATE-X测试集上,MathForm-8B的语义一致通过率分别达到63%和37%,比最强专用基线高出10和12个百分点,且在抽象代数等更依赖库知识的领域优势更明显
- 05在相同10万条数据规模的对照实验中,用FormalVerse训练的模型比用其他公开Lean 4数据集训练的模型语义一致通过率最高高出18.83个百分点
他们做了什么
- 以往方法过度依赖模型记忆的知识,容易引用Mathlib中不存在的定理或违反库的命名习惯,而常见的数据构建流程只是对大批量单次生成结果做事后筛选,无法指出错误出在哪里
- MathForm在生成前用检索规划器从Mathlib取回相关定义和已有形式化案例,生成后经过最多三轮修正,依据编译诊断和语义一致性判定(数据构建阶段用QwQ-32B作裁判,最终评测用gpt-oss-120b)不断改进候选结果
- 基于约36.7万条验证数据FormalVerse,先对Qwen3-8B做监督微调,再用强化学习方法DAPO进一步训练,得到MathForm-8B
- 在最具挑战性的FATE-H和FATE-X测试集上,MathForm-8B的语义一致通过率分别达到63%和37%,比最强专用基线高出10和12个百分点,且在抽象代数等更依赖库知识的领域优势更明显
- 在相同10万条数据规模的对照实验中,用FormalVerse训练的模型比用其他公开Lean 4数据集训练的模型语义一致通过率最高高出18.83个百分点
| AVG | FormalMATH | ProverBench | CombiBench | FATE-M | FATE-H | FATE-X | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | SC | CC | SC | CC | SC | CC | SC | CC | SC | CC | SC | CC | SC | CC |
| Specialized Autoformalizers | ||||||||||||||
| Herald Translator-7B | 64.12 | 27.63 | 95.29 | 47.76 | 78.74 | 37.36 | 77.00 | 5.00 | 70.67 | 54.67 | 42.00 | 15.00 | 21.00 | 6.00 |
| Kimina-Autoformalizer-7B | 73.20 | 34.37 | 99.29 | 76.24 | 96.55 | 56.32 | 95.00 | 16.00 | 77.33 | 44.67 | 43.00 | 8.00 | 28.00 | 5.00 |
| Mathesis-HPO-7B | 76.20 | 34.96 | 99.06 | 79.29 | 97.13 | 59.77 | 96.00 | 15.00 | 84.00 | 48.67 | 50.00 | 4.00 | 31.00 | 3.00 |
| StepFun-Formalizer-7B | 58.12 | 39.55 | 97.41 | 81.41 | 89.66 | 59.20 | 79.00 | 28.00 | 60.67 | 52.67 | 17.00 | 12.00 | 5.00 | 4.00 |
| StepFun-Formalizer-32B | 63.65 | 44.47 | 99.06 | 85.88 | 92.53 | 64.94 | 86.00 | 32.00 | 71.33 | 60.00 | 23.00 | 17.00 | 10.00 | 7.00 |
| Goedel-Formalizer-V2-8B | 78.24 | 60.08 | 98.82 | 94.12 | 98.28 | 89.66 | 89.00 | 42.00 | 87.33 | 82.67 | 62.00 | 44.00 | 34.00 | 8.00 |
| Goedel-Formalizer-V2-32B | 78.28 | 63.74 | 99.06 | 94.59 | 98.28 | 92.53 | 91.00 | 49.00 | 89.33 | 85.33 | 63.00 | 48.00 | 29.00 | 13.00 |
| ReForm-8B | 81.76 | 66.21 | 99.06 | 94.12 | 98.85 | 90.80 | 86.00 | 47.00 | 94.67 | 91.33 | 67.00 | 53.00 | 45.00 | 21.00 |
| ReForm-32B | 81.61 | 68.41 | 99.06 | 95.53 | 98.28 | 94.25 | 93.00 | 55.00 | 91.33 | 88.67 | 69.00 | 52.00 | 39.00 | 25.00 |
| Ours | ||||||||||||||
| MathForm-8B-SFT | 84.38 | 66.53 | 99.29 | 91.06 | 100.00 | 90.80 | 83.00 | 43.00 | 98.00 | 91.33 | 80.00 | 58.00 | 46.00 | 25.00 |
| MathForm-8B | 88.06 | 72.37 | 100.00 | 95.06 | 100.00 | 94.83 | 93.00 | 47.00 | 99.33 | 97.33 | 82.00 | 63.00 | 54.00 | 37.00 |

| AVG | FATE-M | FATE-H | FATE-X | |||||
|---|---|---|---|---|---|---|---|---|
| Method | SC | CC | SC | CC | SC | CC | SC | CC |
| gpt-oss-120b | ||||||||
| Single | 27.33 | 26.43 | 52.00 | 51.30 | 21.00 | 19.00 | 9.00 | 9.00 |
| BoN | 42.67 | 41.10 | 70.00 | 69.30 | 40.00 | 37.00 | 18.00 | 17.00 |
| Feedback | 42.57 | 40.77 | 68.70 | 67.30 | 41.00 | 38.00 | 18.00 | 17.00 |
| Retrieval | 32.23 | 29.00 | 56.70 | 52.00 | 27.00 | 26.00 | 13.00 | 9.00 |
| MathForm | 49.67 | 48.00 | 76.00 | 74.00 | 46.00 | 44.00 | 27.00 | 26.00 |
| Qwen3-235B-A22B-Thinking-2507 | ||||||||
| Single | 7.43 | 7.43 | 13.30 | 13.30 | 5.00 | 5.00 | 4.00 | 4.00 |
| BoN | 18.77 | 18.77 | 33.30 | 33.30 | 14.00 | 14.00 | 9.00 | 9.00 |
| Feedback | 28.77 | 28.77 | 51.30 | 51.30 | 22.00 | 22.00 | 13.00 | 13.00 |
| Retrieval | 9.90 | 8.57 | 16.70 | 16.70 | 8.00 | 6.00 | 5.00 | 3.00 |
| MathForm | 37.57 | 36.23 | 60.70 | 58.70 | 32.00 | 31.00 | 20.00 | 19.00 |
| AVG | FormalMATH | ProverBench | CombiBench | FATE-M | FATE-H | FATE-X | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Training Dataset | SC | CC | SC | CC | SC | CC | SC | CC | SC | CC | SC | CC | SC | CC |
| NuminaMath-LEAN | 66.24 | 41.49 | 99.53 | 85.18 | 96.55 | 72.41 | 87.00 | 25.00 | 69.33 | 49.33 | 35.00 | 16.00 | 10.00 | 1.00 |
| FineLeanCorpus | 78.25 | 46.53 | 100.00 | 84.47 | 98.85 | 74.71 | 96.00 | 29.00 | 90.67 | 68.00 | 53.00 | 17.00 | 31.00 | 6.00 |
| FormalVerse | 77.17 | 60.32 | 98.82 | 90.59 | 98.85 | 89.66 | 78.00 | 36.00 | 95.33 | 84.67 | 62.00 | 46.00 | 30.00 | 15.00 |
| Judge Model | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|
| gpt-oss-120b | 0.8917 | 0.8755 | 0.9133 | 0.8940 |
| QwQ-32B | 0.8567 | 0.8367 | 0.8867 | 0.8609 |
| gpt-oss-20b | 0.8500 | 0.8142 | 0.9067 | 0.8579 |
| Hyperparameter | Value |
|---|---|
| Maximum sequence length | 16,384 |
| Global batch size | 128 |
| Learning rate | 2.0×10−5 |
| Epochs | 3 |
| LR scheduler | Cosine |
| Warmup ratio | 0.1 |
| Precision | bf16 |
| AVG | FormalMATH | ProverBench | CombiBench | FATE-M | FATE-H | FATE-X | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | SC | CC | SC | CC | SC | CC | SC | CC | SC | CC | SC | CC | SC | CC |
| General-Purpose LLMs | ||||||||||||||
| DeepSeek-V4-Pro | 78.34 | 76.54 | 97.88 | 96.24 | 94.83 | 93.68 | 85.00 | 81.00 | 87.33 | 87.33 | 70.00 | 68.00 | 35.00 | 33.00 |
| Qwen3.7-Plus | 86.33 | 83.38 | 99.29 | 98.59 | 100.00 | 97.70 | 92.00 | 87.00 | 94.67 | 92.00 | 78.00 | 77.00 | 54.00 | 48.00 |
| Qwen3-235B-A22B-Thinking-2507 | 58.94 | 55.37 | 91.29 | 88.24 | 81.03 | 75.29 | 59.00 | 50.00 | 71.33 | 70.67 | 32.00 | 30.00 | 19.00 | 18.00 |
| Qwen3-32B | 43.07 | 36.55 | 80.94 | 72.94 | 61.49 | 50.00 | 43.00 | 27.00 | 46.00 | 43.33 | 20.00 | 19.00 | 7.00 | 7.00 |
| DeepSeek-R1-0528-Qwen3-8B | 47.28 | 37.65 | 72.24 | 61.18 | 52.30 | 40.23 | 36.00 | 15.00 | 36.00 | 30.00 | 4.00 | 4.00 | 4.00 | 1.00 |
| Qwen3-8B | 27.25 | 17.06 | 61.88 | 42.12 | 37.93 | 27.59 | 29.00 | 7.00 | 26.67 | 22.67 | 4.00 | 3.00 | 4.00 | 0.00 |
| Ours | ||||||||||||||
| MathForm-8B | 88.06 | 72.37 | 100.00 | 95.06 | 100.00 | 94.83 | 93.00 | 47.00 | 99.33 | 97.33 | 82.00 | 63.00 | 54.00 | 37.00 |
| FATE-M | FATE-H | |||
|---|---|---|---|---|
| Model | SC | CC | SC | CC |
| StepFun-Formalizer-32B | 42.67 | 36.00 | 16.00 | 12.00 |
| Goedel-Formalizer-V2-32B | 75.33 | 68.00 | 38.00 | 27.00 |
| ReForm-32B | 78.67 | 74.00 | 50.00 | 41.00 |
| MathForm-8B | 86.67 | 76.67 | 54.00 | 42.00 |
为什么重要
训练能证明数学定理的AI系统需要大量人工难以生产的、经过验证的形式化数据,这项工作展示了一种更准确的自动生产方式。同时它证明了搭配优质数据和验证流程的小模型可以超越大得多的专用系统,这对追求性价比的AI开发很有参考价值。
本文术语
- 自动形式化(Autoformalization) · 把自然语言写的数学命题自动转换成Lean等机器可验证形式语言的过程
- Lean 4 / Mathlib · 一种用于机器验证数学证明的证明助理语言(Lean)及其内置的大型定义与定理库(Mathlib)
- 语法检查(SC)/一致性检查(CC) · SC检查生成代码是否能正确编译,CC检查生成代码是否真正保留了原始命题的含义
- Pass@8 · 衡量模型在每题8次尝试中成功比例的评测指标
- DAPO/强化学习(RL) · 一种通过比较模型生成的多个候选结果并给予奖励信号来调整模型、使其偏好更优结果的训练方法
论文原文摘要(英文)
Autoformalization is commonly framed as translating natural-language mathematical statements into machine-verifiable formal languages such as Lean 4. However, faithful formalization requires more than translation. Models must map mathematical concepts to the complex hierarchy of types and definitions in formal libraries such as Mathlib, while ensuring that generated statements preserve the meaning of the source propositions. Existing approaches struggle because they rely heavily on the model's parametric memory for library-specific knowledge, while common data construction pipelines often resort to filtering single-pass outputs and lack mechanisms for feedback-driven revision. To address these challenges, we introduce MathForm, an autoformalization framework for constructing verified training data through Mathlib knowledge retrieval and verification-guided iterative refinement. Before generation, a retrieval planner gathers relevant definitions and existing formalizations from Mathlib to guide the formalization generator. Generated statements are then revised using compiler diagnostics and semantic-consistency feedback. Using this framework, we construct FormalVerse, a Lean 4 dataset containing approximately 367K verified examples across diverse mathematical domains and sources. We then train MathForm-8B through supervised fine-tuning followed by reinforcement learning. Across six benchmarks, MathForm-8B achieves average Pass@8 rates of 88.06% under Syntax Check (SC) and 72.37% under Consistency Check (CC), outperforming multiple specialized 32B autoformalizers. On the challenging FATE-H and FATE-X subsets, it attains CC pass rates of 63% and 37%, exceeding the strongest specialized baselines in both cases.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调
METAL MEDIA 最新报道
图片来源: Lushi Pu et al., arXiv:2608.14221, CC BY 4.0