MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

About

Multimodal large language models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, a generalist MLLM typically underperforms compared with a specialist MLLM on most VL tasks, which can be attributed to task interference. In this paper, we propose a mixture of multimodal experts (MoME) to mitigate task interference and obtain a generalist MLLM. Our MoME is composed of two key components, a mixture of vision experts (MoVE) and a mixture of language experts (MoLE). MoVE can adaptively modulate the features transformed from various vision encoders, and has a strong compatibility in transformation architecture. MoLE incorporates sparsely gated experts into LLMs to achieve painless improvements with roughly unchanged inference costs. In response to task interference, our MoME specializes in both vision and language modality to adapt to task discrepancies. Extensive experiments show that MoME significantly improves the performance of generalist MLLMs across various VL tasks. The source code is released at https://github.com/JiuTian-VL/MoME

Leyang Shen, Gongwei Chen, Rui Shao, Weili Guan, Liqiang Nie• 2024

Related benchmarks

Task	Dataset	Result
Visual Question Answering	VizWiz	Accuracy81	1863
Visual Question Answering	TextVQA	Accuracy53.2	1455
Visual Question Answering	GQA	Accuracy81.2	1445
Visual Question Answering	ChartQA	Accuracy57.2	620
Visual Question Answering	ScienceQA	Accuracy80.4	525
Visual Question Answering	VQA v2	Accuracy75.7	347
Multimodal Evaluation	MM-Vet	--	249
Multimodal Evaluation	MMStar	Accuracy48.1	177
Visual Question Answering	DocVQA	ANLS50.8	75
Visual Question Answering	MRAG-Bench	Overall Accuracy66.78	14

Showing 10 of 10 rows

Other info

Follow for update

@wizwand_team Discord