Trust Region Policy Distillation
About
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Mathematical Reasoning | AIME 2024 | Pass@32 Accuracy20.31 | 19 | |
| Mathematical Reasoning | MATH 500 | Accuracy (Avg@8)77.33 | 15 | |
| Mathematical Reasoning | AIME 2025 | Pass@3217.71 | 14 | |
| Mathematical Reasoning | AIME26 | AIME26 Accuracy44.06 | 5 | |
| Mathematical Reasoning | AMC 2023 | Avg@3254.38 | 5 | |
| Mathematical Reasoning | Olympiad | Average Score @843.06 | 5 |