Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

About

Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making it a challenging testbed for spatial reasoning in multimodal large language models (MLLMs) deployed in physical environments. However, current MLLMs lack systematic evaluation and training frameworks for these capabilities. We introduce ReasonMatch-Bench, a benchmark stratified by viewpoint displacement and matching granularity across indoor, outdoor, and object-centric scenarios, and show that current MLLMs still struggle with fine-grained wide-baseline correspondence: on a difficult 90-sample subset, human annotators achieve 84.0 F1, while the best existing baseline reaches 37.2. To bridge this gap, we build a scalable data-generation pipeline that automatically extracts wide-baseline view pairs from large-scale video-3D corpora, including RGB-D videos and SfM reconstructions, yielding diverse and verifiable supervision. We further propose Dynamic Correspondence Reinforcement Learning (DCRL), which combines Image-Level Viewpoint Progression and Point-Level Correspondence Curriculum to improve WBM training through verifiable rewards without explicit CoT supervision. Extensive experiments show that DCRL substantially improves ReasonMatch-Bench and transfers to related spatial benchmarks, while maintaining general visual understanding performance with modest gains on several benchmarks.

Hao Zhong, Muzhi Zhu, Shenyan Zeng, Anzhou Li, Cong Chen, Hua Geng, Duochao Shi, Wentao Ye, Tao Lin, Hao Chen, Chunhua Shen• 2026

Related benchmarks

TaskDatasetResultRank
General Visual UnderstandingRealworldQA
Accuracy70.5
64
Image UnderstandingMMStar
Score62.5
56
Spatial ReasoningSAT (real)
Accuracy (Pass@1)75.3
25
Wide-Baseline Correspondence ReasoningReasonMatch-Bench Overall 90-sample high-divergence subset
F1 Score70.6
18
Wide Baseline MatchingReasonMatch-Bench Overall (test)
F1 Score70.5
10
Wide Baseline MatchingReasonMatch-Bench Outdoor (test)
L1 Score90.9
10
Wide Baseline MatchingReasonMatch-Bench Object (test)
L1 Metric45.6
10
Wide Baseline MatchingReasonMatch-Bench Indoor (test)
L1 Error84.6
10
Spatial ReasoningMindCube
Overall Accuracy43.52
9
Cross-view CorrespondenceDL3DV ReasonMatch-Bench high-divergence samples
Precision53.3
6
Showing 10 of 16 rows

Other info

GitHub

Follow for update