Online Reinforcement Learning
在线强化学习Online RLCommonLearning while interacting with the environment, continuously collecting new data with the latest policy during training.
This is reinforcement learning's most classic setup: the agent interacts with the environment using its current policy, updates the policy once new data comes in, then keeps collecting with the updated policy, repeating the cycle. It covers both on-policy methods that use only the freshest data, like PPO, and off-policy methods that store history in a replay buffer, like SAC; what defines it is that training keeps getting new data throughout, which is exactly the line separating it from offline reinforcement learning. Online learning can explore actively and improve on its own mistakes, but it demands a lot of interaction: simulation speeds this up through parallelism, while real robots are limited by safety, time, and how often a scene can be reset. A common recipe now is to start from demonstrations or offline data and then fine-tune with online reinforcement learning — for example, using RL to further improve a VLA.
ExampleHIL-SERL runs online reinforcement learning on a real robot with a human correcting it as needed, bringing an arm to near-perfect success on tasks like precision assembly and bimanual coordination within 1 to 2.5 hours.
- Also called
- Online RL
- Related
- Offline Reinforcement Learning · Offline-to-Online Reinforcement Learning · Real-World Reinforcement Learning · Reinforcement Fine-Tuning (RL Fine-Tuning) · On-Policy · Off-Policy
- Sources
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems (Fig. 1)
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)