Diffusion Policy Policy Optimization
DPPO(扩散策略策略优化)DPPOAdvancedFine-tunes a diffusion policy directly with PPO policy-gradient reinforcement learning by treating each denoising step as a decision.
DPPO was proposed in September 2024 by Allen Z. Ren and colleagues at Princeton, together with researchers at MIT, Toyota Research Institute, Carnegie Mellon University, and Harvard, published at ICLR 2025. A diffusion policy is trained with imitation learning, so its performance is capped by the quality of the demonstrations; policy-gradient reinforcement learning methods like PPO need to compute action probabilities, and multi-step diffusion denoising doesn't lend itself to that directly, so this kind of fine-tuning was widely assumed to be inefficient. DPPO splits the process into two nested Markov decision processes: the outer one is interacting with the environment, and the inner one treats each denoising step as a Gaussian-sampling “action,” whose probability can be computed, making it possible to fine-tune end to end with PPO. Experiments show it explores close to the demonstration data's distribution, trains stably, and produces more robust policies after fine-tuning, making it a common baseline for reinforcement-learning fine-tuning of diffusion and flow-matching policies.
ExampleOn the Furniture-Bench furniture-assembly simulation tasks, DPPO raised a pretrained diffusion policy's success rate from 57% to 97% on One-Leg and from 12% to 87% on Lamp, and the assembly policy also transferred zero-shot to a real robot.
- Also called
- DPPO
- Related
- Diffusion Policy · Proximal Policy Optimization · Reinforcement Fine-Tuning (RL Fine-Tuning) · Policy Gradient · Sim-to-Real Transfer · ReinFlow
- Sources
- Diffusion Policy Policy Optimization (arXiv:2409.00588)
DPPO 项目主页 (Chinese)
Allen Z. Ren 个人主页(论文列表,标注 ICLR 2025) (Chinese) - As of
- 2025