DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
About
We introduce DeepSeek-Prover-V1.5, an open-source language model designed for theorem proving in Lean 4, which enhances DeepSeek-Prover-V1 by optimizing both training and inference processes. Pre-trained on DeepSeekMath-Base with specialization in formal mathematical languages, the model undergoes supervised fine-tuning using an enhanced formal theorem proving dataset derived from DeepSeek-Prover-V1. Further refinement is achieved through reinforcement learning from proof assistant feedback (RLPAF). Beyond the single-pass whole-proof generation approach of DeepSeek-Prover-V1, we propose RMaxTS, a variant of Monte-Carlo tree search that employs an intrinsic-reward-driven exploration strategy to generate diverse proof paths. DeepSeek-Prover-V1.5 demonstrates significant improvements over DeepSeek-Prover-V1, achieving new state-of-the-art results on the test set of the high school level miniF2F benchmark ($63.5\%$) and the undergraduate level ProofNet benchmark ($25.3\%$).
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Mathematical Reasoning | AIME 2024 | Accuracy9.3 | 394 | |
| Mathematical Reasoning | AIME 2025 | Accuracy7.3 | 378 | |
| Formal Theorem Proving | MiniF2F (test) | Pass@163.5 | 157 | |
| Mathematical Reasoning | MATH500 | Accuracy85.1 | 124 | |
| Automated Theorem Proving | MiniF2F (test) | Success Rate63.5 | 100 | |
| Formal Theorem Proving | PutnamBench | Solved Count23 | 56 | |
| Mathematical Reasoning | AMC23 | Mean Accuracy66.3 | 42 | |
| Mathematical Reasoning | OlympiadBench | Accuracy54.4 | 38 | |
| Math Reasoning | Math Reasoning Evaluation Suite (Math-500, AMC23, GSM8k, Minerva, Olympiad) | Math-5008.1 | 34 | |
| Logical reasoning | Countdown | Accuracy52.6 | 31 |