Embodied AI Glossary中文

Imitation from Observation

从观测中模仿学习IfOAdvanced

Learning to imitate from only the demonstrator's states or video, with no recorded action labels at all.

Ordinary imitation learning needs observation-action pairs; imitation from observation gets only a sequence of the demonstrator's states — video of a person doing something, say — with no idea what control command produced each step. This opens the door to using internet video and other huge collections of human video, but it also has to deal with differing viewpoints and body structures. Notable examples: UC Berkeley's Liu and colleagues (2017) used context translation to convert human video into a robot's viewpoint before running reinforcement learning; UT Austin's Torabi and colleagues (2018) proposed BCO, which first lets the agent explore on its own to learn an inverse dynamics model (inferring the action from a pair of consecutive frames), then uses it to fill in action labels for expert video so behavior cloning can be applied; adversarial approaches exist too. Today's embodied-AI work that pretrains on human video, or uses latent actions or pseudo-action labels, is solving exactly this same problem.

ExampleIn Liu and colleagues' 2017 paper, a robot watches only video of a person sweeping, scooping almonds, and pushing objects — no joint recordings at all — and learns to perform the same actions with tools.

Also called
IfO, Imitation Learning from Observation, Learning from Video
Related
Action-free Video · Inverse Dynamics Model · Human Video Data · Latent Action · Pseudo Action Labels · LAPA
Sources
Recent Advances in Imitation Learning from Observation (IJCAI 2019 survey)
Behavioral Cloning from Observation (IJCAI 2018)
Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation (ICRA 2018)

See it in the full glossary →