Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Pelican-VLA 0.5: Attending Before Acting Benefits Generalization

About

In this report, we present Pelican-VLA 0.5, a unified VLA model that integrates vision-language understanding, future-frame generation, and action prediction within a single architecture. Pelican-VLA 0.5 achieves attention-level generalization: without object annotations, segmentation masks, attention supervision, or task-specific fine-tuning, its action pathway already focuses on the manipulation-relevant object and contact region. This behavior persists across unseen scenes and unseen robot embodiments, and is substantially stronger than in other open-source VLA baselines. We verify that this ability originates from the learnable Bottleneck Token inserted between perception and action: by routing task-relevant visual information through a compact bottleneck, the tokens interface induces manipulation-centric attention during pre-training and remains effective across different policy structures, including a MoT-style architecture.

Zeyuan Ding, Wenhai Liu, Yang Xu, Jiayu Hu, Yinda Chen, Yi Zhang, Yong Dai, Jian Tang, Xiaozhu Ju• 2026

Related benchmarks

TaskDatasetResultRank
Robotic ManipulationRoboTwin 50-task (Seen Tasks)
Average Success Rate91.2
27
Showing 1 of 1 rows

Other info

Follow for update