Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation

About

While knowledge distillation (KD) is widely adopted for training lightweight models by leveraging supervision from larger teacher models, relying solely on output token distributions has proven insufficient for compressing Multimodal Large Language Models (MLLMs). Since output tokens are a byproduct of the model attending to visual inputs, prior works have explored explicitly distilling attention to provide a direct supervisory signal. While promising, the precise utility of which attention signals to distill remains under-explored. In this work, we challenge the conventional reliance on prompt-to-vision attention by revealing that downstream performance correlates strongly with response-to-vision attention similarity to the teacher, but negligibly with that of prompt-conditioned attention. Furthermore, we observe that attention distributions exhibit significant variance across individual tokens, indicating that a uniform distillation objective is suboptimal. To this end, we introduce Token-level Response-visual Attention Guidance (TRAG), a distillation objective that 1) shifts the focus to response-to-vision signals and 2) employs token-specific objectives by adaptively weighting the Kullback-Leibler divergence based on attention entropy, effectively guiding the student to mirror the teacher's precise visual focus. Extensive experimental results on multiple benchmarks demonstrate that TRAG significantly outperforms prior distillation baselines.

Jaehyun Jang, Eunseop Yoon, Hee Suk Yoon, SooHwan Eom, Mark A. Hasegawa-Johnson, Chang D. Yoo• 2026

Related benchmarks

TaskDatasetResultRank
Visual Question AnsweringScienceQA
Accuracy73.2
525
Visual Question AnsweringMMBench (MMB)
Accuracy70.8
169
Multimodal UnderstandingMMMU
Accuracy38.7
107
Visual Question AnsweringMMBench CN
Accuracy69.3
99
Compositional ReasoningSugarCrepe
Overall Accuracy87.3
95
Compositional ReasoningWinoground--
33
Multimodal Perception and ReasoningMME
MME Score71.2
31
Compositional ReasoningBIVLC
Accuracy93.8
16
Showing 8 of 8 rows

Other info

Follow for update