Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models

About

Multimodal Large Language Models (MLLMs) have demonstrated strong perception and reasoning capabilities. However, most existing models focus on isolated objects and neglect structured relationships for efficient target navigation, limiting their performance on visually intensive tasks. To address this challenge, we introduce Scene Graph Thinking (SaGe), a novel paradigm that enables fine-grained and structured visual reasoning through explicit scene-graph representations. Specifically, we first introduce an automated data engine that converts flat image-text corpora into structured scene graphs, where hierarchical entities constitute the nodes and diverse visual relations define the edges. Building upon this, we construct 120K high-quality training data by sampling reasoning traces from scene graphs. Then, two-stage graph-aligned post-training paradigms are introduced, where supervised fine-tuning internalizes MLLMs with structured reasoning, and subsequent reinforcement fine-tuning proposes node-as-proxy graph rewards to consolidate efficient graph exploration. With curated data and graph-aligned training, our approach achieves significant improvements across eight multimodal benchmarks, demonstrating strong effectiveness on fine-grained perception and reasoning tasks. Code is available at https://github.com/zwyang6/SaGe.

Zhiwei Yang, Yuanchen Wu, Nan Zhang, Yucong Meng, Ke Yan, Shouhong Ding• 2026

Related benchmarks

TaskDatasetResultRank
General image understandingMMStar
Accuracy62.3
67
Fine-Grained PerceptionHR-Bench-4K
Overall Score76.5
43
GroundingRefCOCO (val)--
23
Fine-Grained PerceptionVStarBench
Overall Accuracy89
19
Spatial UnderstandingCVBench 2D
Overall Score79.4
9
Spatial UnderstandingCVBench 3D
Overall Score80.5
9
Chart AnalysisChartQA
ChartQA Score87.2
7
Showing 7 of 7 rows

Other info

Follow for update