Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

WMPO: World Model-based Policy Optimization for Vision-Language-Action Models

About

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation, but their reliance on expert demonstrations limits their ability to learn from failures and perform self-corrections. Reinforcement learning (RL) addresses these through self-improving interactions with the physical environment, but suffers from high sample complexity on real robots. We introduce World-Model-based Policy Optimization (WMPO), a principled framework for on-policy VLA RL without interacting with the real environment. In contrast to widely used latent world models, WMPO focuses on pixel-based predictions that align the "imagined" trajectories with the VLA features pretrained with web-scale images. Crucially, WMPO enables the policy to perform on-policy GRPO that provides stronger performance than the often-used off-policy methods. Extensive experiments in both simulation and real-robot settings demonstrate that WMPO (i) substantially improves sample efficiency, (ii) achieves stronger overall performance, (iii) exhibits emergent behaviors such as self-correction, and (iv) demonstrates robust generalization and lifelong learning capabilities.

Fangqi Zhu, Zhengyang Yan, Zicong Hong, Quanxin Shou, Xiao Ma, Song Guo• 2025

Related benchmarks

TaskDatasetResultRank
Robot ManipulationLIBERO
Object Achievement48
1025
InsertionReal-world
Success Rate82
16
Robot Manipulation AggregateFranka Manipulation Real-World (Evaluation)
Mean Success Rate52
16
World Model GenerationLIBERO
FPS7
12
Bread-to-toasterFranka manipulation Real-world
Task Success Rate (TSR)60
7
Robot ManipulationSafeLIBERO Level I
TSR (Spatial)54
7
Apple selectionReal-world Franka manipulation
Task Success Rate (TSR)70
7
Bowl-to-plateFranka manipulation Real-world
Task Success Rate70
7
Dual-arm stackingReal-world Franka manipulation
Task Success Rate (TSR)30
7
Block-stackingFranka manipulation Real-world
Task Success Rate (TSR)30
7
Showing 10 of 14 rows

Other info

GitHub

Follow for update