SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs

About

Knowledge distillation (KD) is a standard route to compress Large Language Models (LLMs) into compact students, yet most pipelines uniformly apply token-wise loss regardless of teacher confidence. This indiscriminate supervision amplifies noisy, high-entropy signals and is especially harmful under large teacher-student capacity gaps. We introduce SelecTKD, a plug-and-play Selective Token-Weighted distillation framework that shifts the focus from "how to measure divergence" to "where to apply learning". At each step, the student proposes tokens that are verified by the teacher through a robust propose-and-verify procedure with two variants: greedy Top-k and non-greedy Spec-k. Accepted tokens receive full loss, while rejected tokens are masked or down-weighted. This objective-agnostic design works with on- and off-policy data, induces an implicit curriculum quantified by Token Acceptance Rate (TAR), and stabilizes optimization. Across instruction following, mathematical reasoning, code generation, and a VLM setting, SelecTKD consistently improves strong baselines and achieves state-of-the-art results for small models without architectural changes or extra reference models.

Haiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju Ren• 2025

Related benchmarks

Task	Dataset	Result
Logical reasoning	ZebraLogic	Accuracy72.1	86
Mathematical Reasoning	HMMT 25	Accuracy (HMMT 25)33.54	50
Knowledge Reasoning	GPQA Diamond	Accuracy58.46	48
Instruction Following	IF-Eval	Accuracy49.17	14
Coding	LCB v6	Pass@129.71	6
Coding	LCB v5	Pass@154.48	6
Preference-based Generation	Arena CW	Score35.6	6
General Evaluation	LiveBench 1125	Score41.8	6
Medical Reasoning QA	MedQA USMLE	Accuracy85.78	6
Medical Reasoning QA	MedXpertQA text	Accuracy23.51	6

Showing 10 of 11 rows

Other info

Follow for update

@wizwand_team Discord