FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

About

Recent advances on Multi-modal Large Language Models have demonstrated that high-resolution image input is crucial for model capabilities, especially for fine-grained tasks. However, high-resolution images lead to a quadratic increase in the number of visual tokens input into LLMs, resulting in significant computational costs. Current work develop visual token compression methods to achieve efficiency improvements, often at the expense of performance. We argue that removing visual redundancy can simultaneously improve both efficiency and performance. We build a coarse-to-fine visual token compression method, with a vision-guided sampler for compressing redundant regions with low information density, and a text-guided sampler for selecting visual tokens that are strongly correlated with the user instructions.With these two modules, the proposed FocusLLaVA achieves improvements in both efficiency and performance. We validate the effectiveness of our approach on a wide range of evaluation datasets.

Yuke Zhu, Chi Xie, Shuang Liang, Bo Zheng, Sheng Guo• 2024

Related benchmarks

Task	Dataset	Result
Object Hallucination Evaluation	POPE	Accuracy87.7	2019
Visual Question Answering	TextVQA	Accuracy70	1453
Visual Question Answering	GQA	Accuracy66	1425
Multimodal Evaluation	MME	--	727
Multimodal Capability Evaluation	MM-Vet	Score41.3	393
Multimodal Model Evaluation	MMBench	Accuracy74.7	204
Multimodal Evaluation	MMBench CN	Accuracy70.3	120
Question Answering	ScienceQA	Accuracy79	96
Multimodal Evaluation	LLaVA-Bench In-the-Wild	Score65.6	73

Showing 9 of 9 rows

Other info

Follow for update

@wizwand_team Discord