Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning

About

While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and multi-step visual interaction. This limitation suggests that relying solely on implicit visual representations from vision encoders is insufficient for recovering fine-grained spatial evidence. We introduce PERception-Interaction-reason Agent (PERIA), a tool-augmented visual agent for spatial reasoning tasks across map reasoning, visual probing, and vision reconstruction. PERIA uses two lightweight tool families: vision perception tools for exposing textual, symbolic, and spatial evidence, and vision interaction tools for manipulating visual context, tracing paths, and verifying spatial relations. To train PERIA, we develop a unified recipe that combines supervised tool-use trajectory synthesis, composite rewards, and Observation-Relaxed Group-in-Group Policy Optimization (OR-GIGPO) for effective multi-tool behavior. Experiments on 13 benchmarks from 8 datasets show that PERIA-8B improves over the Qwen3-8B backbone by 10.0% on in-distribution benchmarks and 4.4% on out-of-distribution benchmarks, while outperforming previous state-of-the-art baselines of similar size by 7.0%-14.8%. It also achieves performance comparable to much larger models such as Qwen3-VL-235B-A22B-Thinking and GPT-5, demonstrating the effectiveness of PERIA in enhancing spatial reasoning capabilities.

Changye Li, Meng Lu, Yi Wu, Ligeng Zhu• 2026

Related benchmarks

TaskDatasetResultRank
Route tracingMapTrace
Accuracy82.4
16
Visual ProbingVisualProbe Hard
Accuracy44.3
16
Paper FoldingPaper Folding
Accuracy15.7
16
Visual ProbingVisualProbe Medium
Accuracy41.2
16
Spatial ReasoningReasonMap
Accuracy45.1
16
Visual ProbingVisualProbe Easy
Accuracy61.7
16
Spatial ReasoningREASONMAP-PLUS
Accuracy72.6
16
Spatial ReasoningV*
Accuracy82.2
16
Ball TrackingBall Tracking
Accuracy19.6
16
Spatial ReasoningBabyVision
Accuracy15.5
16
Showing 10 of 13 rows

Other info

Follow for update