Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision

About

On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence. Incoherent rollouts yield noisy gradients; already-mastered tokens yield redundant ones. This creates waste at three scales (tokens, training phases, and prompts) yet existing methods supervise uniformly. We introduce SEAD, which uses entropy as a unified probe of this competence-dependent degradation at three scales: (1) joint teacher-student entropy partitions tokens into zones receiving tailored divergences or zero gradient (approx. 50% skipped); (2) a cosine schedule anneals from forward to reverse KL as competence grows; (3) a competence-gated curriculum introduces prompts easy-to-hard. These components are symbiotically necessary: token selection requires coherent rollouts (curriculum), annealing requires monotonic improvement (also curriculum). On OLMo-3 (7B to 32B), SEAD achieves +4.8 avg accuracy over vanilla OPD across six math benchmarks, with ablations confirming super-additive interactions.

Chia-Hsuan Lee, Zelei Cheng, Yu Wang, Renkun Ni, Sambit Sahu, Shi-Xiong Zhang, William Campbell• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningMathematical Reasoning Aggregate
Average Score64
46
Mathematical ReasoningOlympiadBench
Pass@1 Accuracy62.5
38
Mathematical ReasoningAMC 2023
Accuracy (AMC 2023)94.7
32
Mathematical ReasoningMATH 500
Pass@1 Accuracy91.2
23
Mathematical ReasoningMinerva Math
Accuracy56.6
16
Mathematical ReasoningAMC 2023
Pass@3289.8
9
Showing 6 of 6 rows

Other info

Follow for update