Phantom: Training Robots Without Robots Using Only Human Videos
About
Training general-purpose robots requires learning from large and diverse data sources. Current approaches rely heavily on teleoperated demonstrations which are difficult to scale. We present a scalable framework for training manipulation policies directly from human video demonstrations, requiring no robot data. Our method converts human demonstrations into robot-compatible observation-action pairs using hand pose estimation and visual data editing. We inpaint the human arm and overlay a rendered robot to align the visual domains. This enables zero-shot deployment on real hardware without any fine-tuning. We demonstrate strong success rates-up to 92%-on a range of tasks including deformable object manipulation, multi-object sweeping, and insertion. Our approach generalizes to novel environments and supports closed-loop execution. By demonstrating that effective policies can be trained using only human videos, our method broadens the path to scalable robot learning.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Visual Fidelity Evaluation | TACO and Aria Dataset | FD470.6 | 15 | |
| Robot Manipulation | RoboTwin novel egocentric viewpoint 2.0 | Adjust Success Rate52 | 6 | |
| Robot Manipulation | RoboTwin standard egocentric viewpoint 2.0 | Adjust96 | 6 | |
| Video Generation | RoboTwin Simulation novel egocentric viewpoint 2.0 | PSNR16.97 | 5 | |
| Drawer | Aria (Real Robot) (test) | Success Rate5 | 4 | |
| Flower | Aria (Real Robot) (test) | Success Rate (SR)0.00e+0 | 4 | |
| Hammer | Aria (Real Robot) (test) | Success Rate (SR)0.00e+0 | 4 | |
| Mustard | Aria (Real Robot) (test) | SR0.00e+0 | 4 | |
| Cross-embodiment video editing | Cross-embodiment video editing dataset | FVD1.95e+3 | 3 |