Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Mask-Proof: An LLM-based Automated Data Curation Pipeline on Mathematical Proofs

About

Large language models (LLMs) are increasingly capable of mathematical problem solving and can even assist with research-level proofs, yet we still lack a scalable and reproducible way to measure step-level reasoning in long proofs across diverse sources. This evaluation gap limits trustworthy AI assistance in proof-certified scientific progress. Existing evaluations often emphasize final answers or rely on costly expert grading, while end-to-end proof generation remains open-ended and hard to verify automatically. We introduce Mask-Proof, a pipeline that turns real proofs into automatically checkable masked-step tasks. It masks key formula steps, provides the necessary surrounding context, and evaluates model reconstructions with an LLM-based equivalence judge using repeated votes for stability. The resulting Mask-ProofBench contains 292 curated problems across diverse research areas. Experiments with 17 models show that reasoning-enhanced models outperform standard models by 12% to 27%. Our evaluator achieves 96.8% agreement with expert annotators, enabling faithful, reproducible, and comparable measurement of step-level mathematical reasoning. Benchmark, annotations, and code are available at https://github.com/weating/Mask-Proof.

Jierui Zhang, Siyuan Tan, Xinhang Li, Longzhuangzhi Lin, Dailin Li, Chengfeng Gu, Xinping Li, Yaxian Hao, Shengjia Liang, Yuxiang Ren, Wenhao Liu• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical Proof ReasoningIMO-ProofBench Advanced (full set)--
14
Mathematical Proof ReasoningIMO-ProofBench-Advanced Novel (22 problems)--
7
Mathematical Proof ReasoningIMO-ProofBench-Advanced USAMO 2025 (5 problems)--
7
Showing 3 of 3 rows

Other info

Follow for update