Recovery and Correction Data
纠偏数据CommonDemonstration data specifically recorded to show how a robot gets back on track after starting to go wrong.
Most teleoperated demonstrations are “perfect trajectories” done right the first time, so a policy never sees what a drifted-off state looks like and has no idea how to save itself once it starts going wrong, and the error compounds. Recovery and correction data fills exactly this gap: a policy is left to run on its own, and just as it's about to fail, a human takes over, steers the robot back on track, and finishes the task, with that takeover segment added to the training set. The idea traces back to DAgger (Dataset Aggregation) from 2011, with later variants such as Human-Gated DAgger. A 2025 method called RaC turns it into a dedicated stage after imitation learning: the operator first rewinds the robot to a familiar state, then demonstrates a correction, and on long-horizon tasks such as hanging a shirt or sealing a lunchbox it beats prior methods using roughly one-tenth the collection time.
ExampleIn RaC's shirt-hanging task, just as the policy is about to hang the shirt crookedly, the operator takes over, first pulls the arm back to a normal prior pose, then demonstrates hanging it straight on the hanger — that takeover becomes one piece of recovery and correction data.
- Also called
- Recovery Data, Correction Data
- Related
- Human Intervention Data · Failure Recovery · DAgger · Human-Gated DAgger · Compounding Error · RaC
- Sources
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) - As of
- 2025-09