Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Trust Region Policy Distillation

About

Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.

Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningAIME 2024
Pass@32 Accuracy20.31
19
Mathematical ReasoningMATH 500
Accuracy (Avg@8)77.33
15
Mathematical ReasoningAIME 2025
Pass@3217.71
14
Mathematical ReasoningAIME26
AIME26 Accuracy44.06
5
Mathematical ReasoningAMC 2023
Avg@3254.38
5
Mathematical ReasoningOlympiad
Average Score @843.06
5
Showing 6 of 6 rows

Other info

Follow for update