Embodied AI Glossary中文

Failure Data

失败数据Common

Trajectories that didn't complete the task, useful for learning to recover from mistakes and for training reward models.

Failure data refers to trajectories that failed to complete the task, coming from teleoperation mistakes, a policy's own autonomous failures, or segments recorded before a human stepped in to correct things. Pure imitation learning usually throws these away, since copying them would teach the policy to fail the same way, but they record exactly what conditions lead to errors, which has real uses: training success detectors, reward models, and value functions, serving as negative feedback in reinforcement learning, and, paired with correction data, teaching a model to recover from mistakes. The π0 paper points out that a model trained only on high-quality data never learns to recover, because this kind of data rarely contains mistakes at all. Physical Intelligence's RECAP method trains π*0.6 on a mix of failure trajectories, autonomous-run data, and human corrections, reporting more than double the throughput and roughly half the failure rate on harder tasks.

ExampleBesides its 76,000 successful trajectories, the DROID dataset also separately releases about 16,000 trajectories the collectors labeled “unsuccessful,” usable for training a success detector or for offline reinforcement learning.

Also called
Failed Trajectories, Failed Demonstrations
Related
Recovery and Correction Data · Human Intervention Data · Success Detector · Reward Model · RECAP · Offline Reinforcement Learning
Sources
π*0.6: a VLA That Learns From Experience (arXiv 2511.14759)
DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset (arXiv 2403.12945)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)
As of
2025-11

See it in the full glossary →