Sink-Aware Pruning for Diffusion Language Models

About

Diffusion Language Models (DLMs) incur high inference cost due to iterative denoising, motivating efficient pruning. Existing pruning heuristics largely inherited from autoregressive (AR) LLMs, typically preserve attention sink tokens because AR sinks serve as stable global anchors. We show that this assumption does not hold for DLMs: the attention-sink position exhibits substantially higher variance over the full generation trajectory (measured by how the dominant sink locations shift across timesteps), indicating that sinks are often transient and less structurally essential than in AR models. Based on this observation, we propose ${\bf \texttt{Sink-Aware Pruning}}$, which automatically identifies and prunes unstable sinks in DLMs (prior studies usually keep sinks for AR LLMs). Without retraining, our method achieves a better quality-efficiency trade-off and outperforms strong prior pruning baselines under matched compute. Our code is available at https://github.com/VILA-Lab/Sink-Aware-Pruning.

Aidar Myrzakhan, Tianyi Li, Bowei Guo, Shengkun Tang, Zhiqiang Shen• 2026

Related benchmarks

Task	Dataset	Result
Commonsense Reasoning	WinoGrande	Accuracy50.67	1442
Language Understanding	MMLU	Accuracy33.3	844
Physical Commonsense Reasoning	PIQA	Accuracy60.34	696
Science Question Answering	ARC Challenge	Accuracy38.2	354
Question Answering	GPQA	Accuracy25	258
Science Question Answering	ARC Easy	Accuracy71.75	162
commonsense inference	HellaSwag	Accuracy52.1	123
Reading Comprehension	RACE	Accuracy28.2	75
Language Understanding	LLM Benchmark Suite (MMLU, ARC-C, PIQA, WinoG, GSM8K, HellaSwag, GPQA, RACE) (test)	Overall Accuracy57.68	13
Zero-shot Language Understanding and Reasoning	LLM Evaluation Suite (MMLU, ARC-C, PIQA, WinoG, GSM8K, HellaSwag, GPQA, RACE) zero-shot LLaDA1.5	Average Score58.47	13

Showing 10 of 13 rows

Other info

GitHub

Follow for update

@wizwand_team Discord