Diffusion Policy
扩散策略DPEssentialAn imitation-learning policy that generates robot action sequences by gradually denoising from random noise, using a diffusion model.
Diffusion Policy was introduced by Cheng Chi and colleagues in Shuran Song's group at Columbia University, working with Toyota Research Institute and MIT; it appeared on arXiv in March 2023, was published at RSS 2023, and an extended version appeared in IJRR in 2024. It formulates a robot policy as a conditional denoising diffusion process: conditioned on observations such as camera images, it starts from random noise and denoises step by step into a chunk of future actions. The advantage is that it can represent multimodal actions — when a scene has several equally valid ways to act, it doesn't average them into one wrong action the way direct regression would. The paper also combines this with receding-horizon control (predict a chunk, execute part of it, then re-predict), and beat the best prior methods by an average of 46.9% across 12 tasks in 4 benchmarks. It became an important source for the diffusion and flow-matching action heads later used in models like 3D Diffusion Policy and π0.
ExamplePush-T is its signature task: the robot uses the end of a cylindrical rod to push a T-shaped block on a table into a target position and orientation. Since going around from the left or the right both work, and both appear in the demonstrations, Diffusion Policy learns both, but commits to just one of them on each run.
- Also called
- DP, Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Related
- Diffusion Model · Action Multimodality · Denoising Diffusion Probabilistic Model · Action Chunking · 3D Diffusion Policy · Push-T
- Sources
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)
Diffusion Policy 项目主页 (Chinese) - As of
- 2024-03