G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models
About
Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrained vision-language backbones, but their visual tokens remain grounded in 2D image coordinates rather than the calibrated geometry of the robot's cameras -- a mismatch especially pronounced in multi-camera setups, where views are coupled by known intrinsics and extrinsics yet processed as independent images. We propose G$^3$VLA, a camera-aware geometric module that injects calibrated structure into the visual-token stream of a pretrained VLA without altering its action space or imitation objective, combining intrinsic-conditioned ray embeddings, projective positional encoding (PRoPE), and bidirectional cross-view fusion. Geometric supervision is provided either from ground-truth point maps when available, or from confidence-gated $\pi^3$X teacher predictions, requiring no depth sensors or manual annotations. Instantiated on $\pi_0$, G$^3$VLA yields consistent gains across the LIBERO suites, RoboCasa24, RoboTwin2.0, and real-robot settings, with the largest improvements on spatially and object-sensitive tasks. We further validate on $\pi_{0.5}$ and GR00T 1.5, with results suggesting that geometric transfer is most effective when geometry-aware tokens have direct access to the action generation pathway. Our project page is at https://sites.google.com/view/g3vla
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Robot Manipulation | LIBERO | Spatial Success Rate96.6 | 223 | |
| Robotic Manipulation | Pour manipulation task | Success Rate (SR)100 | 16 | |
| Pick-and-Place Test Tube | Pick-and-Place Test Tube ID | Success Rate75 | 12 | |
| Pick-and-Place Test Tube | Pick-and-Place Test Tube OOD | Success Rate58.3 | 12 | |
| Pick-and-Place Test Tube | Pick-and-Place Test Tube Overall | Success Rate65 | 12 | |
| Robot Manipulation | Pouring Nut Task (OOD) | Success Rate87.5 | 8 | |
| Robot Manipulation | Pouring Nut Task (Overall) | Success Rate92.5 | 8 | |
| Robot Manipulation | RoboCasa24 total average | Success Rate37.1 | 3 | |
| Robot Manipulation | RoboTwin handover_block 2.0 | Success Rate49 | 3 | |
| Robot Manipulation | LIBERO | Goal Success Rate88.4 | 3 |