Embodied AI Glossary中文

Noise-Space Policy Steering

噪声空间策略引导Advanced

Keeping a diffusion policy's weights fixed and using reinforcement learning to pick its input noise instead, changing its output.

Diffusion and flow-matching policies generate actions by sampling a random noise vector and gradually denoising it into an action; feed the same model a different starting noise and it produces a different action. Noise-space policy steering treats that noise as a controllable “action”: it freezes the original policy and trains a small, separate reinforcement-learning policy that, given the current observation, outputs which noise to use so the original policy produces a better action. The representative method is DSRL, proposed by a UC Berkeley-led team in 2025. It only needs black-box calls to the original policy, with no backpropagation through the multi-step denoising process, and it never touches the large model's weights, so it is sample-efficient and well suited to online improvement on a real robot, including on-the-fly adaptation of general-purpose VLAs such as π0. Correspondingly, how much it can improve things is bounded by the range of actions the original policy is capable of producing in the first place.

ExampleDSRL steers a public π0 checkpoint trained on DROID data to open a toaster with a Franka arm; after about 80 online episodes, success rate rises from 5/20 to 18/20.

Also called
Diffusion Noise-Space RL, Latent Noise-Space RL, Diffusion Steering
Related
Diffusion Steering via Reinforcement Learning · Diffusion Policy · Flow Matching · Reinforcement Fine-Tuning (RL Fine-Tuning) · Residual Reinforcement Learning · π0
Sources
Wagenmaker et al. 2025: Steering Your Diffusion Policy with Latent Space Reinforcement Learning (DSRL)
As of
2025-06

See it in the full glossary →