VLA-RL
AdvancedAn early framework that further improves an autoregressive VLA, such as OpenVLA, using online reinforcement learning.
VLA-RL was proposed in May 2025 by researchers at Tsinghua Shenzhen International Graduate School and Nanyang Technological University, with code open-sourced. A VLA trained purely by imitating demonstrations has only seen a limited set of states, and tends to fail as soon as it drifts out of distribution; VLA-RL instead lets a pretrained autoregressive VLA keep improving online by trying things itself in the environment. It treats one robot manipulation trajectory as a multi-turn multimodal conversation and applies trajectory-level reinforcement learning with PPO (Proximal Policy Optimization); to ease the sparse-reward problem, it fine-tunes a vision-language model into a robotic process reward model, with training labels coming from automatically segmented task stages (using keyframes such as moments when the gripper becomes stable). On the engineering side, it also uses curriculum-based task selection, GPU-load-balanced parallel environments, batched decoding, and value-network warm-up. VLA-RL is one of the earlier works to systematically demonstrate that 'VLA plus online RL' is workable, and the authors also observed that increasing test-time optimization keeps improving results further.
ExampleAcross 40 manipulation tasks in LIBERO, VLA-RL raised OpenVLA-7B's average success rate from 76.5% to 81.0%, on par with π0-FAST.
- Also called
- VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
- Related
- Reinforcement Fine-Tuning (RL Fine-Tuning) · Proximal Policy Optimization · Reward Model · OpenVLA · LIBERO Benchmark · SimpleVLA-RL
- Sources
- VLA-RL (arXiv 2505.18719)
GuanxingLu/vlarl (GitHub) - As of
- 2025-05