Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought

About

Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning. However, current vision-language models rely on text-only or coordinate-augmented chain-of-thought, where entity references remain implicit and ambiguous. This may cause the reasoning process to decouple from visual evidence, entity references to drift across steps, and a causal disconnection between the reasoning trajectory and the final answer, with these problems further amplified in multi-view scenarios due to cross-view appearance changes. To address these issues, we propose Pinned Chain-of-Thought (PinCoT), a structured reasoning paradigm that pins every reasoning step to visual evidence. PinCoT introduces the concept of reasoning anchor, which binds each task-relevant entity to a structured visual anchor with entity name, unique identity, view index, and spatial grounding, enabling consistent entity tracking across reasoning steps and views. We build a fully automated data generation pipeline to construct PIN-170K, a high-quality PinCoT-formatted reasoning dataset. We then train RoboPIN through three-stage post-training that progressively injects embodied knowledge, structured reasoning ability, and process-supervised alignment, with rewards that directly constrain both anchor localization and identity consistency during reasoning. On 14 benchmarks covering embodied spatial reasoning, multi-view reasoning, and pointing, RoboPIN with only 4B parameters consistently outperforms 7B level open-source embodied models, achieving a 12% average improvement over the strongest 7B baseline, Mimo-Embodied. Further analysis shows that PinCoT improves grounding accuracy and cross-step identity consistency, validating the effectiveness of process supervision.

Yaoting Huang, Yifu Yuan, Linqi Han, Chengwen Li, Shuoheng Zhang, Xianze Yao, Hongyao Tang, Yan Zheng, Jianye Hao• 2026

Related benchmarks

TaskDatasetResultRank
Spatial ReasoningEmbSpatial
Overall Accuracy83.1
131
Spatial ReasoningCV-Bench
Accuracy85.5
89
Spatial ReasoningBLINK
Spa. Score87.4
57
Spatial ReasoningROBOSPATIAL
Accuracy66.6
48
Spatial ReasoningSAT
Overall Acc78
21
Spatial ReasoningERQA
Accuracy44
17
Multi-view spatial reasoningCrossPoint
Accuracy72.2
10
Visual GroundingRef-Spatial
Pointing Performance50.5
10
Visual GroundingRobo-Refit
Pointing Performance89
10
Visual GroundingRobo Afford
Pointing Performance72.8
10
Showing 10 of 16 rows

Other info

Follow for update