Hindsight Experience Replay
后见之明经验回放HERAdvancedRelabeling a failed attempt as having succeeded at whatever goal it actually reached, so even sparse rewards can be learned from.
Introduced by OpenAI's Andrychowicz and colleagues in 2017 (NIPS 2017). In goal-conditioned tasks with only a binary success/failure reward, a robot almost never succeeds early on, so the replay buffer fills up with nothing but failures and there's nothing to learn from. HER's fix is to store a second copy of every failed trajectory, with its goal swapped out for whatever state that trajectory actually reached (such as its final state, or a state at some later point) — which turns it into a successful example. It can be attached after any off-policy algorithm (one that can learn from old data, such as DDPG), and in effect works like an automatically generated curriculum. The original paper completed pushing, sliding, and pick-and-place tasks on a Fetch arm and deployed a policy trained in simulation to the real robot. This hindsight-relabeling idea was later widely borrowed by methods like goal-conditioned behavior cloning.
ExampleA robot arm tries to push a block to a red dot but ends up pushing it 10 cm to the side instead; HER rewrites that trajectory's goal to be that 10-cm-off position, turning it into a successful experience.
- Also called
- HER
- Related
- Sparse Reward · Goal-Conditioned Reinforcement Learning · Experience Replay · Hindsight Relabeling · Deep Deterministic Policy Gradient · Off-Policy
- Sources
- Hindsight Experience Replay (arXiv 1707.01495)