Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera
About
Recent visual imitation learning systems have widely adopted multi-camera setups with wrist-mounted cameras as the de facto standard. However, manipulation from a single global view remains challenging, as the policy should capture fine-grained interaction details and identify task-relevant regions without local wrist views. To address this challenge, we present Spatially Conditioned Diffusion Policy (SCDP), a diffusion-based visuomotor policy that achieves precise and robust manipulation in a single-camera setting. Our key idea is that end-effector trajectories can serve as visual attention anchors that reflect task-relevant regions. Building on this idea, SCDP consists of two key components: (i) a visual encoder that produces multi-scale feature maps to capture both broader context and fine-grained visual features, and (ii) a spatial conditioning module that samples point-wise features along intermediate end-effector trajectories in the diffusion loop. Extensive simulation experiments show that SCDP consistently outperforms strong single-view baselines and achieves performance comparable to multi-camera baselines. Real-world experiments further demonstrate precise manipulation and robustness to visual distractors, highlighting the potential of single-camera imitation learning.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Robotic Manipulation | DexArt | -- | 29 | |
| Robot Manipulation | Meta-World | Success Rate (Easy)91.5 | 11 | |
| Battery Insertion | Real-World Battery Insertion | Grasp Success Rate100 | 5 | |
| Cup Handle Grasping | Real-World Cup Handle Grasping Original | Success Rate95 | 5 | |
| Cup Handle Grasping | Real-World Cup Handle Grasping Distractor | Success Rate85 | 5 | |
| Cup Wall Grasping | Real-World Cup Wall Grasping Original | Success Rate95 | 5 | |
| Cup Wall Grasping | Real-World Cup Wall Grasping Distractor | Success Rate80 | 5 | |
| General Robot Manipulation | Real-World Manipulation Aggregate | Average Success Rate73.8 | 5 | |
| Long-horizon manipulation | Real-World Long Horizon Task | Open Drawer Success Rate90 | 5 | |
| Push Cube | Real-World Push Cube Distractor | Success Rate75 | 5 |