Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

About

Vision-language-action (VLA) models typically inject proprioception only as a late conditioning signal, preventing robot state from grounding instruction understanding or directing visual attention. We introduce ThinkProprio, which discretizes proprioception into VLM-vocabulary tokens and uses them jointly with the instruction to gate visual patches before VLM computation, steering the model toward action-relevant evidence while discarding redundant tokens early. We find that proprioception added as a passive conditioning signal leaves performance essentially unchanged; its value emerges when token-form state acts as an active query that, with the instruction, selects which visual patches the VLM processes. Systematic ablations show that VLM-vocabulary tokens outperform learned projectors as the state encoding, and that retaining only about \SI{12}{\percent} of the visual tokens surpasses on CALVIN ABC$\to$D. Across CALVIN, LIBERO, and real-world manipulation, ThinkProprio reduces end-to-end inference latency while improving the matched full-token baseline.

Fangyuan Wang, Peng Zhou, Jiaming Qi, Shipeng Lyu, Chengyang He, David Navarro-Alarcon, Guodong Guo• 2026

Related benchmarks

TaskDatasetResultRank
Robot ManipulationLIBERO
Object Achievement98.4
1025
Long-horizon robot manipulationCalvin ABCD→D
Task 1 Completion Rate97.7
140
Robot Manipulation8 held-out robot manipulation tasks (test)
Success Rate91.3
12
Long-horizon task successCALVIN D→D long-horizon
Success Rate (LH-1)99.5
11
Language-conditioned imitation learningLIBERO (test)
Spatial Score97.6
8
Showing 5 of 5 rows

Other info

Follow for update