Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

About

This paper introduces Diffusion Policy, a new way of generating robot behavior by representing a robot's visuomotor policy as a conditional denoising diffusion process. We benchmark Diffusion Policy across 12 different tasks from 4 different robot manipulation benchmarks and find that it consistently outperforms existing state-of-the-art robot learning methods with an average improvement of 46.9%. Diffusion Policy learns the gradient of the action-distribution score function and iteratively optimizes with respect to this gradient field during inference via a series of stochastic Langevin dynamics steps. We find that the diffusion formulation yields powerful advantages when used for robot policies, including gracefully handling multimodal action distributions, being suitable for high-dimensional action spaces, and exhibiting impressive training stability. To fully unlock the potential of diffusion models for visuomotor policy learning on physical robots, this paper presents a set of key technical contributions including the incorporation of receding horizon control, visual conditioning, and the time-series diffusion transformer. We hope this work will help motivate a new generation of policy learning techniques that are able to leverage the powerful generative modeling capabilities of diffusion models. Code, data, and training details is publicly available diffusion-policy.cs.columbia.edu

Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, Shuran Song• 2023

Related benchmarks

TaskDatasetResultRank
Robot ManipulationLIBERO
Object Achievement92.5
1025
Vision-Language NavigationR2R-CE (val-unseen)
Success Rate (SR)30
779
Robotic ManipulationLIBERO
Spatial Success Rate91.4
570
Vision-Language NavigationRxR-CE (val-unseen)
SR23
512
Robotic ManipulationLIBERO-Plus
Language Understanding Score77
414
Robot ManipulationLIBERO (test)
Average Success Rate76.1
237
Robot ManipulationLIBERO
Spatial Success Rate78.5
223
Robotic ManipulationLIBERO
Long-horizon Success Rate68.3
165
Long-horizon robot manipulationCalvin ABCD→D
Task 1 Completion Rate86.3
140
Robotic ManipulationCalvin ABCD→D
Avg Length0.56
139
Showing 10 of 1423 rows
...

Other info

Follow for update