$\pi^{*}_{0.6}$: a VLA That Learns From Experience
About
We study how vision-language-action (VLA) models can improve through real-world deployments via reinforcement learning (RL). We present a general-purpose method, RL with Experience and Corrections via Advantage-conditioned Policies (RECAP), that provides for RL training of VLAs via advantage conditioning. Our method incorporates heterogeneous data into the self-improvement process, including demonstrations, data from on-policy collection, and expert teleoperated interventions provided during autonomous execution. RECAP starts by pre-training a generalist VLA with offline RL, which we call $\pi^{*}_{0.6}$, that can then be specialized to attain high performance on downstream tasks through on-robot data collection. We show that the $\pi^{*}_{0.6}$ model trained with the full RECAP method can fold laundry in real homes, reliably assemble boxes, and make espresso drinks using a professional espresso machine. On some of the hardest tasks, RECAP more than doubles task throughput and roughly halves the task failure rate.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| USB Insertion | USB Insertion | Success Rate93 | 19 | |
| open drawer | Open Drawer | Success Rate40 | 13 | |
| Pencil-Case Packing | CASE | Success Rate (SR)90 | 10 | |
| Cosmetic Packaging | PACK | Success Rate94 | 10 | |
| Pen-Cap Assembly | CAP | Success Rate94 | 10 | |
| Robotic Assembly | AirPods assembly | Grasp Case Success Rate100 | 10 | |
| Robot Manipulation | Real-world robot manipulation tasks | Success Rate (Socks)60 | 7 | |
| Box Closing | Box Closing Real-world 1.0 (test) | Success Rate60 | 6 | |
| Dynamic Brick Sorting | Dynamic Brick Sorting Real-world 1.0 (test) | Success Rate50 | 6 | |
| Backpack Packing | Backpack Packing Real-world 1.0 (test) | Success Rate40 | 6 |