Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MI-Pruner: Crossmodal Mutual Information-guided Token Pruner for Efficient MLLMs

About

For multimodal large language models (MLLMs), visual information is relatively sparse compared with text. As a result, research on visual pruning emerges for efficient inference. Current approaches typically measure token importance based on the attention scores in the visual encoder or in the LLM decoder, then select visual tokens with high attention scores while pruning others. In this paper, we pursue a different and more surgical approach. Instead of relying on mechanism-specific signals, we directly compute Mutual Information (MI) between visual and textual features themselves, prior to their interaction. This allows us to explicitly measure crossmodal dependency at the feature levels. Our MI-Pruner is simple, efficient and non-intrusive, requiring no access to internal attention maps or architectural modifications. Experimental results demonstrate that our approach outperforms previous attention-based pruning methods with minimal latency.

Jiameng Li, Aleksei Tiulpin, Matthew B. Blaschko• 2026

Related benchmarks

TaskDatasetResultRank
Object Hallucination EvaluationPOPE--
2056
Video UnderstandingVideoMME
Score (Overall)67.1
369
Science Question AnsweringScienceQA (SQA)
Accuracy69.81
338
Long Video UnderstandingLongVideoBench--
290
Video Question AnsweringVideoMME--
254
Multimodal EvaluationMM-Vet--
249
Visual Question AnsweringTextVQA
TextVQA Accuracy55.9
210
Video Question AnsweringMSVD
Accuracy70.6
169
Visual Question AnsweringGQA
GQA Score57.01
152
Multimodal EvaluationMME
MME-P Score1.43e+3
139
Showing 10 of 21 rows

Other info

Follow for update