Embodied AI Glossary中文

Reinforcement Learning with Prior Data

利用先验数据的强化学习RLPDAdvanced

A simple, efficient way to do online RL by mixing offline data half-and-half with freshly collected data in every batch.

RLPD was proposed by Ball, Smith, Kostrikov, and Levine at ICML 2023, addressing how to make the best use of existing offline data — expert demonstrations or a large amount of suboptimal trajectories — once online interaction begins. The authors found that no elaborate offline pretraining was needed: just a few small changes to the off-policy algorithm SAC were enough. Every training batch draws half from the offline data and half from the online replay buffer (symmetric sampling); the critic network gets layer normalization to stop it from overestimating the value of unseen actions; and the critic is an ensemble of 10 networks trained with a higher update-to-data ratio. The paper reports roughly a 2.5x improvement over prior methods across several benchmarks. The real-robot RL framework SERL uses RLPD as its core algorithm.

ExampleSERL trains on a real robot using RLPD: human demonstrations go into the offline buffer, and data the robot generates through its own interaction goes into the online buffer, with every training batch drawing half from each.

Also called
RLPD, Efficient Online Reinforcement Learning with Offline Data
Related
Offline-to-Online Reinforcement Learning · Soft Actor-Critic · Update-to-Data Ratio · Experience Replay · SERL · HIL-SERL
Sources
Ball et al. 2023: Efficient Online Reinforcement Learning with Offline Data (RLPD)
GitHub: ikostrikov/rlpd
Luo et al. 2024: SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning

See it in the full glossary →