Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs

About

Large Language Models (LLMs) suffer from hallucinations due to their static parametric knowledge. Retrieval-Augmented Generation (RAG) and GraphRAG mitigate this issue by incorporating external knowledge and structured reasoning over knowledge graphs (KGs). However, existing approaches remain largely text-centric, as constructing fine-grained multimodal knowledge graphs (MMKGs) with explicit cross-modal semantics remains challenging. In this paper, we propose MMGraphRAG, a framework for building interpretable MMKGs that unify textual and visual knowledge. Our approach represents visual content as structured scene graphs and integrates them with textual KGs through a novel cross-modal entity linking method, SpecLink, which leverages spectral clustering to jointly model semantic similarity and graph structure. This design preserves explicit entities, relations, and reasoning paths across modalities, enabling structure-aware retrieval and generation. To support evaluation, we introduce the CMEL dataset, a benchmark for fine-grained cross-modal entity alignment. Experimental results on CMEL demonstrate improved entity linking accuracy, while evaluations on DocBench and MMLongBench show that MMGraphRAG achieves superior performance and stronger robustness, particularly in complex multimodal reasoning scenarios.

Xueyao Wan, Hang Yu• 2025

Related benchmarks

TaskDatasetResultRank
Multimodal ReasoningScienceQA
Average Accuracy78.21
45
Multi-Modal Long Document Question AnsweringVisDoMBench (Full)
SPIQA Score69.91
25
Knowledge-based VQAInfoSeek
Unseen-Q Performance0.69
18
Multimodal Document Question AnsweringDocBench
Accuracy (Academia)60.7
17
Multimodal ClassificationCrisisMMD
BC Accuracy68.4
16
Knowledge-based VQAE-VQA
Single-Hop Accuracy19.12
16
Multimodal Document Question AnsweringMMLongBench (test)
Chart Acc.34.7
12
Multimodal Document QAVisDoMBench FetaTab (full)
Accuracy72.4
11
Multimodal Document QAVisDoMBench PaperTab (full)
Accuracy56.36
11
Multimodal Document QAVisDoMBench SciGraphQA (full)
Accuracy64.11
11
Showing 10 of 13 rows

Other info

Follow for update