Embodied AI Glossary中文

ReinFlow

Advanced

A method that injects learnable noise into a flow-matching policy so it can be fine-tuned with online reinforcement learning.

ReinFlow was released by Tonghe Zhang, Chao Yu, Yu Wang, and colleagues at Carnegie Mellon University, Tsinghua University, and others in May 2025, published at NeurIPS 2025. Flow-matching policies (generative policies that turn noise into actions step by step using a velocity field) hit an obstacle when researchers try to improve them further with reinforcement learning: the generation process is deterministic, so it has no well-defined action probability and little exploration. ReinFlow adds a learnable Gaussian noise term at each denoising step, turning generation into a discrete-time Markov process, which makes the likelihood exactly computable and lets the policy be fine-tuned with policy gradients; the noise network is used only during training and dropped at inference time. On legged-locomotion tasks, this raises the episode reward of rectified-flow policies by about 135% on average, with roughly 83% less compute than DPPO, a diffusion-policy fine-tuning method. ReinFlow belongs to the same line of work as DPPO and πRL: using reinforcement learning to fine-tune generative policies.

ExampleA Shortcut flow policy that needs very few denoising steps is first trained on demonstrations, then fine-tuned with ReinFlow using online reinforcement learning on robomimic simulated manipulation tasks, producing a clear jump in success rate.

Also called
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
Related
Flow Matching · Diffusion Policy Policy Optimization · πRL · Reinforcement Fine-Tuning (RL Fine-Tuning) · Rectified Flow · Policy Gradient
Sources
ReinFlow (arXiv 2505.22094)
ReinFlow project page
As of
2025-05

See it in the full glossary →