PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
About
Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action models and world models struggle with accurate 3D geometry and physically meaningful forecasting. We propose PhysMani, a framework that couples a physics-principled 3D Gaussian world model with a future-aware action policy model. The world model learns a divergence-free Gaussian velocity field via online optimization for fast and physically grounded future dynamics prediction. The policy model integrates the predicted 3D scene future dynamics through a learnable token based cross-attention module. We introduce PhysMani-Bench, a dynamic manipulation benchmark with 16 tasks, and demonstrate a superior success rate over strong baselines in both simulation and real-world robot experiments.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Robotic Manipulation | PhysMani-Bench | Mean Success Rate45.9 | 7 | |
| Robotic Manipulation | PhysMani-Bench High Speed | Beat Buzz Success Rate (H)49.2 | 7 | |
| Dynamic Object Manipulation | PhysMani-Bench | Training Time (hours)53 | 7 | |
| Future Frame Prediction | PhysMani-Bench next 1st frame | PSNR26.9 | 3 | |
| Future Frame Prediction | PhysMani-Bench next 5th frame | PSNR22.42 | 2 | |
| Future Frame Prediction | PhysMani-Bench next 10th frame | PSNR20 | 2 | |
| Pick from Belt | Real-world Physical Robot | SR81.3 | 2 | |
| Place on Belt | Real-world Physical Robot | Success Rate62.5 | 2 | |
| Place on Rack | Real-world Physical Robot | SR25 | 2 | |
| Real-world Dynamic Tasks (Aggregate) | Real-world Physical Robot | Mean SR (%)62.5 | 2 |