RLinf
AdvancedOpen-source, large-scale reinforcement learning training framework for embodied and agentic foundation models.
RLinf is a reinforcement learning infrastructure open-sourced in 2025, jointly developed by institutions including Tsinghua University, aimed at making large-model RL training both fast and flexible. It schedules simulation, inference (rollout generation), and training onto the same pool of GPUs, switching between them or running them in a pipeline as needed to keep GPU utilization high. On the embodied side, it supports online reinforcement-learning fine-tuning of vision-language-action models such as OpenVLA, OpenVLA-OFT, and π0 inside simulators like ManiSkill and LIBERO; it can also be used for reasoning-oriented RL on large language models. It suits researchers who want to apply RL post-training to a VLA without building their own distributed system from scratch.
ExampleUse RLinf's provided configuration to fine-tune OpenVLA-OFT with PPO or GRPO reinforcement learning inside the ManiSkill simulator, improving pick-and-place success rate.
- Related
- Reinforcement Fine-Tuning (RL Fine-Tuning) · Vision-Language-Action Model · veRL (Volcano Engine Reinforcement Learning) · SimpleVLA-RL · ManiSkill · LIBERO Benchmark
- Sources
- RLinf on GitHub
- As of
- 2025-09