Embodied AI Glossary中文

Knowledge Distillation

知识蒸馏KDCommon

Training a small student model to mimic a large teacher model's outputs, compressing the teacher's ability into a smaller model.

Knowledge distillation was systematically formalized by Hinton, Vinyals, and Dean in a 2015 paper: first train a large, strong teacher model, then train a small student model to match the teacher's output probability distribution (called soft labels) rather than only the ground-truth labels. The paper raises the softmax “temperature” to make the soft labels smoother, which exposes similarity relationships between classes so the student learns more from each example. This addresses the problem that large models are expensive to deploy and slow to run inference on. In embodied AI, distillation often takes a teacher-student form: a teacher trained in simulation with privileged information unavailable on a real robot (such as exact terrain shape or contact state), and a student that only sees real-robot sensors and learns to imitate the teacher's actions — this is one route to sim-to-real transfer. It's also used to distill a multi-step denoising diffusion policy into a model that needs far fewer steps, cutting inference latency.

ExampleIn Lee and colleagues' 2020 Science Robotics work on the ANYmal quadruped, the teacher policy sees ground-truth terrain and foot-contact information in simulation, while the student policy takes only a history of onboard proprioception and learns by imitating the teacher, eventually deploying zero-shot on mud, snow, and rubble in the wild.

Also called
KD, Model Distillation
Related
Teacher-Student Distillation · Privileged Information · Policy Distillation · On-Policy Distillation · Kullback-Leibler Divergence · Diffusion Step Distillation
Sources
Distilling the Knowledge in a Neural Network (arXiv 1503.02531)
Learning Quadrupedal Locomotion over Challenging Terrain (Science Robotics 2020, arXiv 2010.11251)

See it in the full glossary →