Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

About

Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial reasoning. Existing spatial MLLMs rely on large-scale datasets, explicit 3D inputs, architecture-specific modifications, or sparse Reinforcement Learning (RL) methods that provide insufficient guidance for spatially-grounded reasoning. We introduce SpatialThinker. To our knowledge, it is the first MLLM unifying Scene Graph Generation (SGG) and visual reasoning in a single pass via online RL. The model simulates human-like spatial perception by constructing a mental scene graph of task-relevant objects and relations, and reasoning toward an answer via dense spatial rewards. Our contributions are threefold: (1) SGG-grounded reasoning: integrating SGG directly within the reasoning chain rather than as a disjoint preprocessing step; (2) STVQA-7K: a high-quality spatial VQA training dataset via a scalable synthesis pipeline; and (3) a dense spatial reward design that enforces structured grounding during RL and generalizes to improve broad visual perception. SpatialThinker-7B achieves 3.6$\times$ larger gains over SFT and $1.7\times$ better in- and out-of-distribution generalization than sparse RL. Trained on only 7K samples, SpatialThinker-7B matches GPT-5 and outperforms GPT-4o, while SpatialThinker-30B surpasses both GPT-5 and Claude 4 Sonnet on average across 14 spatial and real-world benchmarks, demonstrating that structured spatial grounding with reward-aligned reasoning enables robust spatial understanding with limited data.

Hunar Batra, Haoqin Tu, Hardy Chen, Yuanze Lin, Cihang Xie, Ronald Clark• 2025

Related benchmarks

TaskDatasetResultRank
Visual Question AnsweringRealworldQA
Accuracy74.9
327
Visual Question AnsweringMMStar
Accuracy66.9
151
Visual Question AnsweringV*Bench
Accuracy85.9
111
Spatial ReasoningMindCube (tiny)
Accuracy45.4
84
3D Spatial Reasoning3DSRBench
Accuracy62.1
64
Real-world Visual Question AnsweringMME-RealWorld-Lite (MMERW)
Accuracy49.2
31
Spatial ReasoningMMVP
Accuracy79.7
29
Spatial UnderstandingCV-Bench--
29
Egocentric Spatial Reasoning3DSRBench Egocentric (test)
Orientation Accuracy (Cam.V)0.3052
24
Spatial UnderstandingBLINK (val)
Spatial Relation88.1
23
Showing 10 of 21 rows

Other info

Follow for update