Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

About

Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations beyond basic object recognition. Existing methods typically build such representations through textual reasoning or 3D reconstruction. However, they often falter during multi-step reasoning, particularly when required to dynamically re-anchor evidence to the specific camera-, object-, or direction-centric reference frames demanded by complex queries. To address this, we propose OmniView-Space, a framework designed to maintain spatial consistency through multimodal egocentric evidence. Our approach consists of three core components: (1) Multi-Perspective Spatial Mapping (MPSM), which re-anchors reconstructed geometry into a query-aligned visual cognitive map and a textual spatial graph; (2) Tool-Guided Egocentric Reasoning, an interleaved policy trained to actively select the ego anchor required by the query and request the corresponding MPSM evidence; and (3) Cognitive-Map Distillation, which uses MPSM-generated trajectories and ego-frame rewards to train the model to reason with self-generated cognitive maps. Experiments on single- and multi-image spatial reasoning benchmarks show that OmniView-Space achieves state-of-the-art performance. Furthermore, the distilled model maintains this performance while reducing reliance on external geometry pipelines.

Xudong Li, Mengdan Zhang, Peixian Chen, Jiaxi Tan, Zihao Huang, Jingyuan Zheng, Yan Zhang, Xiawu Zheng, Xing Sun, Rongrong Ji• 2026

Related benchmarks

TaskDatasetResultRank
Multi-view spatial reasoningMindCube (tiny)
Overall Accuracy71.5
84
Spatial Relationship ReasoningSPAR-Bench
Accuracy (Avg)44.8
57
Multi-image Spatial ReasoningSPAR-Bench-MV + MindCube-Tiny + MMSI-Bench (test)
Overall Score50.6
32
Multi-view spatial reasoningMMSI-Bench
Overall Accuracy35.5
17
Single-image spatial reasoningOmniSpatial Perspective
Average Score54.53
7
Single-image spatial reasoning3DSRBench Allocentric
Front Accuracy67.15
7
Showing 6 of 6 rows

Other info

Follow for update