LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

About

How to efficiently transform large language models (LLMs) into instruction followers is recently a popular research direction, while training LLM for multi-modal reasoning remains less explored. Although the recent LLaMA-Adapter demonstrates the potential to handle visual inputs with LLMs, it still cannot generalize well to open-ended visual instructions and lags behind GPT-4. In this paper, we present LLaMA-Adapter V2, a parameter-efficient visual instruction model. Specifically, we first augment LLaMA-Adapter by unlocking more learnable parameters (e.g., norm, bias and scale), which distribute the instruction-following ability across the entire LLaMA model besides adapters. Secondly, we propose an early fusion strategy to feed visual tokens only into the early LLM layers, contributing to better visual knowledge incorporation. Thirdly, a joint training paradigm of image-text pairs and instruction-following data is introduced by optimizing disjoint groups of learnable parameters. This strategy effectively alleviates the interference between the two tasks of image-text alignment and instruction following and achieves strong multi-modal reasoning with only a small-scale image-text and instruction dataset. During inference, we incorporate additional expert models (e.g. captioning/OCR systems) into LLaMA-Adapter to further enhance its image understanding capability without incurring training costs. Compared to the original LLaMA-Adapter, our LLaMA-Adapter V2 can perform open-ended multi-modal instructions by merely introducing 14M parameters over LLaMA. The newly designed framework also exhibits stronger language-only instruction-following capabilities and even excels in chat interactions. Our code and models are available at https://github.com/ZrrSkywalker/LLaMA-Adapter.

Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, Hongsheng Li, Yu Qiao• 2023

Related benchmarks

Task	Dataset	Result
Visual Question Answering	VizWiz	Accuracy42.7	1820
Visual Question Answering	TextVQA	Accuracy43.8	1453
Visual Question Answering	VQA v2	Accuracy70.7	1429
Visual Question Answering	GQA	Accuracy45.1	1425
Multimodal Understanding	MMBench	Accuracy39.5	847
Multimodal Evaluation	MME	Score1.33e+3	727
Image Captioning	MS COCO Karpathy (test)	CIDEr1.222	706
Multimodal Understanding	MM-Vet	MM-Vet Score31.4	631
Multimodal Reasoning	MM-Vet	MM-Vet Score31.4	517
Multimodal Understanding	SEED-Bench	Accuracy32.7	516

Showing 10 of 65 rows

Other info

Code

Follow for update

@wizwand_team Discord