Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

About

Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments. We argue that this stems from the lack of a unified geometry-aware manipulation representation, leaving existing VLAs vulnerable to low-level trajectory supervision, misaligned 3D features, and embodiment differences. To address this, we propose GEAR-VLA, a VLA framework for learning unified geometry-aware action representations for generalizable robotic manipulation. GEAR-VLA adopts coarse-to-fine action learning, where multi-source embodied pretraining equips the VLM with embodied reasoning and discrete action understanding before latent action tokens connect action semantics to a gradient-decoupled DiT continuous action expert. It further performs semantic-aligned 3D integration by aligning a trainable 3D spatial backbone with the VLA representation while freezing the original VLM-aligned visual pathway. To share this representation across robots, GEAR-VLA uses embodiment canonicalization, where embodiment-aware states and embodiment-invariant actions confine robot differences to the low-level interface. Extensive simulation and real-world experiments demonstrate strong generalization: GEAR-VLA achieves state-of-the-art performance on LIBERO, zero-shot LIBERO-Plus, and RoboTwin 2.0, reaches 85.9% success on AgileX and 81.0% on the pretraining-unseen LDT-01 embodiment, and obtains 90.1% success on a 6,360-trial universal grasping benchmark with 212 unseen objects. Code and models will be released at https://github.com/babynabeauty/GEAR-VLA.

Yuan Zhang, Shiqi Zhang, Yedong Shen, Shuai Dong, Jiajun Deng, Xin Zhang, Yuxuan Gao, Jiajia Wu, Xin Nie, Zhiyuan Cheng, Jianmin Ji, Yanyong Zhang, Xingyi Zhang, Jia Pan• 2026

Related benchmarks

TaskDatasetResultRank
Spatial ReasoningEmbSpatial--
131
Robot ManipulationRoboTwin Randomized 2.0
Overall Success Rate89.9
100
Robot ManipulationRoboTwin Clean 2.0
Success Rate91.1
74
Visual Perception and ReasoningBLINK
Accuracy85.69
70
Robot ManipulationLIBERO Average
Success Rate98.7
41
Embodied Visual Question AnsweringERQA
Accuracy43.5
39
Egocentric PlanningEgo-Plan 2
Score47
35
Pick-&-PlaceReal-world Pick-and-place Sparse v1
Success Rate93.5
18
Pick-&-PlaceReal-world Pick-and-place Dense v1
Success Rate91.7
18
Pick-&-PlaceReal-world Pick-and-place BG Light v1
Success Rate92
18
Showing 10 of 18 rows

Other info

Follow for update