Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

About

In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pre-training and supervised fine-tuning, RL post-training, and real-world deployment. Each component serves a distinct role in this stack.

He Zhang, Lingzhu Xiang, Haitao Lin, Zeyu Huang, Minghui Wang, Dingyan Zhong, Yubo Dong, Yihao Wu, Yongming Rao, Dongsheng Zhang, Wanjia He, Ling Chen, Kai Huang, Jiahao Chen, Sichang Su, Xumin Yu, Ziyi Wang, Chengwei Zhu, Xiao Teng, Yuchun Guo, Yufeng Zhang, Yuandong Liu, Rui Wang, Zisheng Lu, Han Hu, Zhengyou Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Robot ManipulationRoboTwin Randomized 2.0
Overall Success Rate90.1
100
Robotic ManipulationRoboTwin 50-task (Seen Tasks)
Average Success Rate90.5
27
Robotics Task ExecutionRoboTwin 2.0 (Clean)
Success Rate90.9
20
Bimanual Robot ManipulationRoboTwin Easy 2.0
Average Success Rate90.9
12
Bimanual Robotic ManipulationRoboTwin Hard 2.0
Success Rate (Overall)90.1
12
Showing 5 of 5 rows

Other info

Follow for update