Policy Distillation
策略蒸馏AdvancedHaving a student policy imitate one or more already-trained teacher policies, compressing or merging their skills.
Policy distillation is knowledge distillation applied to decision-making: instead of learning directly from reward, the student policy fits the actions or action distributions the teacher policy outputs at each state. DeepMind's Rusu and colleagues systematically proposed this on Atari games in 2015, for two purposes: compressing a large network into a small one, and merging several single-game experts into one multi-task policy whose combined performance beats each expert trained alone. A common robotics pattern is “specialist to generalist”: train separate specialist policies for different objects or tasks, sometimes using privileged information (ground-truth state only available in simulation), and then distill them into one generalist policy, often paired with DAgger so the student queries the teacher for labels at the states it actually visits.
ExamplePeking University's He Wang and colleagues' UniDexGrasp++ first groups thousands of objects by geometric features, trains a specialist grasping policy for each group, then iteratively distills them into one generalist policy, ultimately reaching grasp success rates of 85.4% on training objects and 78.2% on held-out ones.
- Also called
- Specialist-to-Generalist Distillation
- Related
- Knowledge Distillation · Teacher-Student Distillation · Specialist Policy · Generalist Policy · DAgger · On-Policy Distillation
- Sources
- Rusu et al. 2015: Policy Distillation
Wan et al. 2023: UniDexGrasp++
He et al. 2024: HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots