Imitation Learning
模仿学习ILEssentialTeaching a robot a skill by having it imitate an expert's — usually a human's — demonstrated actions.
Imitation learning studies how to get an agent to learn behavior from expert demonstrations, and it's well suited to situations where “just show it once” is easier than writing rules or designing a reward. Demonstrations are usually recorded from a human through teleoperation, kinesthetic teaching (physically guiding the robot's arm), or motion capture. There are three main approaches: behavioral cloning, which directly uses supervised learning to fit the demonstrated actions; inverse reinforcement learning, which first infers a reward function from the demonstrations and then optimizes a policy against it with reinforcement learning; and adversarial imitation methods such as GAIL, which push the policy's behavior distribution to match the expert's. The core difficulty is distribution shift: during execution, the policy ends up in states the demonstrations never covered, and methods like DAgger address this by asking the expert to label additional actions in exactly those states. The core training of virtually every mainstream VLA model today is, at heart, large-scale imitation learning.
ExampleMobile ALOHA collected just 50 human teleoperated demonstrations per task, co-trained with existing static-ALOHA data, and learned mobile manipulation tasks like sautéing shrimp and plating it, calling an elevator and riding it, and opening a cabinet to put away a pot.
- Also called
- Learning from Demonstration (LfD), Programming by Demonstration, IL
- Related
- Behavior Cloning · Inverse Reinforcement Learning · Generative Adversarial Imitation Learning · DAgger · Demonstration Data · Teleoperation
- Sources
- Osa et al. 2018: An Algorithmic Perspective on Imitation Learning
Wikipedia: Imitation learning
Fu et al. 2024: Mobile ALOHA