Embodied AI Glossary中文

Experience Replay

经验回放Common

Storing an agent's past interactions in a buffer and sampling them randomly during training, instead of using each one only once.

Experience replay is a data-reuse mechanism in reinforcement learning, proposed by Long-Ji Lin in the early 1990s and popularized when DeepMind's DQN (Deep Q-Network) combined it with deep networks in 2013. At every step, the agent stores the transition — state, action, reward, next state — into a buffer of limited capacity, overwriting the oldest entries once it's full; training then samples random mini-batches from this buffer to update the network. This has two benefits: each piece of data can be reused many times, improving sample efficiency; and random sampling breaks the strong correlation between consecutive steps, making training more stable. Because the sampled data comes from an older policy, any algorithm using it must be off-policy, such as DQN or SAC. Real-robot reinforcement learning, where samples are expensive, relies on this especially heavily.

ExampleDQN playing Atari keeps the most recent 1 million frames in its replay buffer. HIL-SERL uses two buffers — one for human demonstrations and intervention data, one for the policy's own interactions — and samples half of each training batch from each.

Also called
Replay Buffer
Related
Off-Policy · Deep Q-Network · Soft Actor-Critic · Hindsight Experience Replay · Sample Efficiency · HIL-SERL
Sources
Playing Atari with Deep Reinforcement Learning (DQN, 2013)
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL)

See it in the full glossary →