Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models

About

Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. Existing methods typically rely on Gumbel-Softmax to approximate discrete selection during training. However, the optimization is driven by surrogate gradients rather than the true selection process, leading to unreliable learning of token importance. In this paper, we propose DiffPrune, which reformulates pruning as continuous control of token information instead of discrete selection learning. Specifically, we introduce an Information Throttler that modulates each token using variance-preserving noise conditioned on importance scores, where higher scores induce less information suppression during training. This design directly operates on token representations, naturally providing a fully differentiable optimization path for learning token importance. At inference, tokens are removed via hard thresholding on the learned scores. Across ten VLM benchmarks, DiffPrune retains 96.5% of full-model accuracy while accelerating LLM prefill by 2.85x, with only 0.69 ms of inference overhead.

Landi He, Mingde Yao, Shawn Young, Lijian Xu• 2026

Related benchmarks

Task	Dataset	Result
Object Hallucination Evaluation	POPE	--	2056
Multimodal Understanding	MMBench CN	--	302
Multimodal Evaluation	MME	Total Score1.72e+3	30
Multimodal Understanding	MMBench	MMB Score64.7	26
Multimodal Understanding and Question Answering	LLaVA 7B Evaluation Suite (GQA, MMBench, MMBench-CN, MME, POPE, ScienceQA, VQAv2, TextVQA, SEED-Bench, VizWiz) 1.5	GQA Accuracy57.8	22
Visual Question Answering	TextVQA	VQAText Score56.7	21
Science Question Answering	ScienceQA	SQA Score72.7	19
Visual Question Answering	GQA	GQA Score62.3	14
Vision-Language Multi-task Evaluation	Qwen2.5-VL Evaluation Suite MMB, MME, POPE, SQA, VQAText (test)	MMB Score81.7	10
Visual Question Answering	VQAv2	VQAv2 Score78.6	7

Showing 10 of 10 rows

Other info

Follow for update

@wizwand_team Discord