Privileged Information
特权信息CommonExtra information available only during training, not at deployment, like a simulator's exact terrain shape or friction values.
This term comes from Vapnik and Vashist's 2009 “learning using privileged information” (LUPI): extra information helps during training but is unavailable at test time. In robotics, it usually refers to quantities a simulator can read directly but a real robot's sensors cannot measure — terrain height, foot contact forces, friction coefficients, external disturbances. Training with reinforcement learning using only what a real robot can actually observe often fails to learn at all, so a common approach is to first train a teacher policy that can see the privileged information, then distill it into a student policy that uses only proprioception or camera input (teacher-student distillation); alternatively, only the critic (the network estimating value) is given privileged information, called an asymmetric actor-critic.
ExampleETH's ANYmal quadruped (Lee et al., Science Robotics 2020): the teacher policy sees terrain shape, foot contact state and contact forces, friction coefficients, and external disturbances in simulation; the student policy imitates the teacher using only a history of proprioception like joint states and IMU readings, and deploys directly to the real robot, walking across mud, snow, rubble, and dense vegetation.
- Also called
- Privileged Observation, Learning Using Privileged Information, LUPI
- Related
- Teacher-Student Distillation · Asymmetric Actor-Critic · Sim-to-Real Transfer · Rapid Motor Adaptation · Proprioception · Knowledge Distillation
- Sources
- Learning Quadrupedal Locomotion over Challenging Terrain (Lee et al., Science Robotics 2020)
Learning by Cheating (Chen et al., CoRL 2019)