Embodied AI Glossary中文

πRL

Advanced

An open-source framework for online reinforcement-learning fine-tuning of flow-matching VLAs like π0 and π0.5.

πRL was released in October 2025 by teams from Tsinghua University, Peking University, the Institute of Automation at the Chinese Academy of Sciences, and others, built on RLinf, an open-source reinforcement learning framework. The difficulty is that a flow-matching VLA generates actions through multi-step denoising, which has no tractable log-probability for an action, making it hard to directly apply policy-gradient algorithms like PPO. πRL offers two solutions: Flow-Noise models the denoising process as a discrete-time MDP with a learnable noise network that makes the log-likelihood exactly computable, and Flow-SDE turns the deterministic ODE sampling into stochastic SDE sampling, forming a two-level MDP out of denoising and environment interaction that makes exploration easier. Both substantially improve a policy already fine-tuned with a small amount of supervised data, on both LIBERO and ManiSkill, and it has also been validated on GR00T N1.5.

ExampleA π0 policy fine-tuned with supervised learning on a small number of demonstrations reaches only 57.6% success on LIBERO; after reinforcement learning with πRL, this rises to 97.6%, and π0.5 goes from 77.1% to 98.3%.

Also called
piRL, pi_RL, πRL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
Related
π0 · π0.5 · RLinf · Reinforcement Fine-Tuning (RL Fine-Tuning) · Flow Matching · Proximal Policy Optimization
Sources
πRL (arXiv:2510.25889)
πRL 论文 HTML 版 (Chinese)
As of
2026-01

See it in the full glossary →