Diffusion Steering via Reinforcement Learning
DSRL(扩散策略噪声空间强化学习)DSRLAdvancedLeaves a diffusion policy's weights untouched and uses reinforcement learning only to pick its input noise, steering it toward better actions.
DSRL was proposed in June 2025 by Andrew Wagenmaker and colleagues in Sergey Levine's group at UC Berkeley, together with researchers at the University of Washington and Amazon, published at CoRL 2025. A diffusion policy generates an action by first sampling random noise and then denoising it, and for the same observation, different initial noise leads to different actions. DSRL treats this initial noise as a new “action space,” training a small reinforcement-learning policy (an MLP) that outputs noise given the observation, which is then handed to the frozen diffusion policy to denoise. The base policy's weights are never touched — it's only called as a black box — and because the noise always maps to some reasonable action from the demonstration data, exploration is better directed and needs less real-robot interaction. It's well suited to quickly and autonomously improving a behavior-cloned policy in a new environment, and it's a representative example of steering a policy through its noise space.
ExampleThe authors treated a public π0 checkpoint (trained on DROID) as a black box and ran reinforcement learning only in its noise space, improving performance on real-robot manipulation tasks without fine-tuning π0 itself.
- Also called
- DSRL, Steering Your Diffusion Policy with Latent Space RL
- Related
- Noise-Space Policy Steering · Diffusion Policy · π0 · Real-World Reinforcement Learning · Residual Reinforcement Learning · Diffusion Policy Policy Optimization
- Sources
- Steering Your Diffusion Policy with Latent Space Reinforcement Learning (arXiv:2506.15799)
DSRL 项目主页 (Chinese) - As of
- 2025-06