Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

About

Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether this scaling recipe can be applied to robotic manipulation to achieve genuine generalization. This is challenging because, unlike text, manipulation data is heterogeneous by nature, expensive to collect, and narrow in diversity, making alignment and scale simultaneously difficult. We present Qwen-RobotManip, a generalizable Vision-Language-Action foundation model built on Qwen-VL. Qwen-RobotManip introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. This alignment capability in turn enables Qwen-RobotManip to absorb manipulation data at a scale that prior training regimes could not sustain. A human-to-robot synthesis pipeline converts egocentric hand demonstrations into robot trajectories across 15 platforms, and a rigorous curation pipeline harmonizes heterogeneous datasets. Using only open-source datasets and human videos without proprietary data collection, Qwen-RobotManip constructs a ~38,100-hour pretraining corpus and exhibits emergent generalization capabilities, including zero-shot instruction following, robustness to perturbations, reactive error recovery, and cross-embodiment transfer. We find that standard benchmarks fail to capture pretraining quality and instead adopt OOD settings including RoboCasa365, LIBERO-Plus, EBench, RoboTwin-Clean2Rand, RoboTwin-IF, and RoboTwin-XE. Qwen-RobotManip substantially outperforms prior state-of-the-art models, including $\pi$0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.

Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wu, Xiong-Hui Chen• 2026

Related benchmarks

TaskDatasetResultRank
Robotic ManipulationLIBERO-Plus
Language Understanding Score86.5
414
Robotic ManipulationLIBERO v1 (test)
Average Success Rate99.2
118
Robotic ManipulationLIBERO-Plus (test)
Lighting Robustness Score98.6
52
Robotic ManipulationRoboTwin Easy
Average Success Rate (AVG)93.7
18
Robotic ManipulationRoboCasa365
Average Success Rate35.9
18
Robotic ManipulationRoboTwin Hard
Average Success Rate94
12
Robot ManipulationRoboTwin 2.0
Clean (Easy) Success Rate93.7
11
Robot ManipulationRoboCasa365 pretraining (test)
Average Success Rate35.9
11
Robotic ManipulationRoboTwin-Clean2Rand (test)
Success Rate (Easy)85
9
Robot ManipulationLIBERO
Success Rate99.2
8
Showing 10 of 30 rows

Other info

GitHub

Follow for update