Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference

About

Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Existing acceleration methods usually remove visual tokens or skip visual-token updates in entire layers, but these coarse strategies may discard fine-grained evidence or suppress useful operators together with redundant ones. In this paper, we study visual-token computation from an answer-observable perspective and find that late visual-token updates can remain large while having little effect on answer-token representations. Motivated by this answer-silent redundancy, we decompose each Transformer layer into attention and FFN operators and show that useful visual computation is often operator-dominant and layer-dependent. We propose an operator-level visual-token skipping framework that preserves the full visual-token sequence while selectively bypassing redundant attention, FFN, or both. Experiments across three MLLM architectures and 10 VQA benchmarks show that our method achieves strong efficiency-accuracy trade-offs, reducing \textbf{33.7\%} TFLOPs on Qwen3-VL while retaining \textbf{99.5\%} of the vanilla model performance.

Zhaoyang Luo, Runmin Dong, Miao Yang, Fan Wei, Yushan Lai, Bin Luo, Haohuan Fu• 2026

Related benchmarks

TaskDatasetResultRank
Diagram UnderstandingAI2D
Accuracy82.55
377
Multimodal Perception and CognitionMME
Overall Score2.34e+3
344
Visual Question AnsweringVizWiz
Accuracy71.7
193
Multimodal UnderstandingMMBench
Accuracy83.08
137
Multimodal UnderstandingMMMU
Accuracy36.78
107
Scientific Question AnsweringScienceQA
Accuracy94.35
83
Multimodal ReasoningMMMU
Accuracy52.89
77
OCR & Document UnderstandingOCRBench
Score80.4
77
Object Hallucination EvaluationPOPE
Average Accuracy98.8
53
General Visual Question AnsweringGQA
Accuracy60.84
35
Showing 10 of 19 rows

Other info

Follow for update