Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation

About

Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low token budget. Our analysis shows that as the pruning budget decreases, accuracy degradation is often accompanied by larger feature distribution shifts. Critically, the degree of this distribution shift strongly correlates with performance degradation. To better characterize this phenomenon, we introduce a lightweight distribution consistency metric to estimate the distribution shift between retained and full tokens. Motivated by these observations, we propose a two-stage pruning framework consisting of Anchor-Context Graph Recovery (ACGR) and Text-Aware Token Cluster Selection (TATCS). Specifically, ACGR transfers contextual information before token removal, while TATCS dynamically re-selects representative tokens when severe distribution shift is detected. Extensive experiments demonstrate that our method achieves superior and more stable performance under ultra-low token budget. Notably, it retains 92.1% of the upper-bound average performance on LLaVA-1.5-7B with only 16 visual tokens.

Xifeng Xue, Xiaokang Wang, Zirui Li, Ming-Ming Cheng, Guolei Sun• 2026

Related benchmarks

TaskDatasetResultRank
Multimodal EvaluationMME
Score2.13e+3
902
Object Hallucination EvaluationPOPE
Accuracy86.5
259
Video Question AnsweringMSVD
Accuracy52.76
169
Video Question AnsweringMSRVTT
Accuracy47.7
117
Text-based Visual Question AnsweringTextVQA
Accuracy64.6
106
Visual Question AnsweringSQA
Accuracy69.3
86
Video Question AnsweringTGIF
Top-1 Acc33.6
75
Science Question AnsweringSQA
Accuracy (SQA)79.1
52
Visual Question AnsweringGQA
Exact Match58.5
32
Perception and Cognition EvaluationMME
Score1.76e+3
25
Showing 10 of 13 rows

Other info

Follow for update