Hindsight Relabeling
事后重标注AdvancedAfter data is collected, relabeling a trajectory's goal or instruction to match whatever it actually achieved.
The idea behind hindsight relabeling is: a trajectory that failed to reach its original goal still ended up somewhere, so that actual outcome can be relabeled as the goal after the fact — turning a failed sample into a success at “achieving a different goal.” The technique became popular through OpenAI's 2017 Hindsight Experience Replay (HER), which addresses how little a policy can learn under sparse rewards (a reward given only on success), and was validated on robot-arm tasks like pushing, sliding, and pick-and-place. It was later extended to imitation learning and language-conditioned policies: a large amount of manipulation data is collected with no task label at all, and afterward it's labeled with an image goal or a language instruction based on what the footage shows. Google's DIAL, for instance, uses a vision-language model like CLIP to automatically add language instructions to 80,000 demonstrations, 96.5% of which had no human annotation to begin with. The precondition for this technique is that the policy is conditioned on a goal or instruction — otherwise there's nothing to relabel.
ExampleA robot arm meant to push a block to point A instead pushes it to point B; after relabeling, this trajectory is stored as a successful demonstration with “goal = point B,” and used to train a goal-conditioned policy.
- Also called
- Hindsight Goal Relabeling
- Related
- Hindsight Experience Replay · Goal-Conditioned Reinforcement Learning · Sparse Reward · Language Annotation · Auto-labeling · Instruction Augmentation
- Sources
- Hindsight Experience Replay (arXiv)
Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models (DIAL, arXiv)