Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

About

High-performance Multimodal Large Language Models (MLLMs) are heavily dependent on data quality. To advance fine-grained image recognition within MLLMs, we introduce a novel data synthesis method inspired by contrastive learning and image difference captioning. Our key idea involves challenging the model to discern both matching and distinct elements by scrutinizing object differences in detailed regions across similar images. We begin by generating pairs of similar images that emphasize object variations. Following this, we employ a Difference Area Generator to pinpoint object differences, and subsequently, a Difference Captions Generator to articulate these differences. This process results in a high-quality dataset of "object replacement" samples, termed Img-Diff, which can be scaled as needed due to its automated nature. We leverage this generated dataset to fine-tune state-of-the-art (SOTA) MLLMs, such as InternVL2, achieving substantial improvements across various image difference and Visual Question Answering tasks. Notably, the trained models significantly outperform existing SOTA models like GPT-4V and Gemini on the MMVP benchmark. Additionally, we conduct comprehensive evaluations to validate the dataset's diversity, quality, and robustness, offering several insights into the synthesis of such contrastive datasets. We release our codes and dataset to encourage further research on multimodal data synthesis and MLLMs' fundamental capabilities for image understanding.

Qirui Jiao, Daoyuan Chen, Yilun Huang, Bolin Ding, Yaliang Li, Ying Shen• 2024

Related benchmarks

Task	Dataset	Result
Object Hallucination Evaluation	POPE	--	2056
Visual Question Answering	GQA	Accuracy62.8	1445
Visual Question Answering	VQA v2	Accuracy81.8	1429
Multimodal Capability Evaluation	MM-Vet	Score52.6	429
Multimodal Reasoning	MMBench	Accuracy82.7	180
Science Question Answering	ScienceQA SQA-I	Accuracy96.6	149
Multimodal Reasoning	MMBench CN	Accuracy81.4	119
Multimodal Reasoning	SEED-Bench	Accuracy69.9	59
Image Difference Captioning	Image-Edit-Request (test)	BLEU16.6	22
Image Difference Captioning	Spot-the-Diff (test)	METEOR13.1	22

Showing 10 of 10 rows

Other info

Code

Follow for update

@wizwand_team Discord