Asymmetric Actor-Critic
非对称演员-评论家AdvancedLetting the critic see the full simulated state while the actor sees only what a real robot could actually observe.
Asymmetric actor-critic was introduced by Lerrel Pinto, Marcin Andrychowicz, Pieter Abbeel, and colleagues in 2017. In an actor-critic algorithm, the critic (the value network) is used only during training and discarded at deployment, so while training in simulation it can be given the full true state — exact object poses, velocities, and other privileged information — while the actor (the policy) receives only inputs a real robot can actually get, such as camera images. The critic's value estimates become more accurate, training is faster and more stable, and the policy that's left over can go straight onto a real robot. The original paper combined this with domain randomization to complete grasping and pushing tasks on a 7-DOF Fetch arm using no real-robot data at all. OpenAI's Dactyl dexterous-hand project used the same approach, and it's also common today to let the critic read privileged observations like terrain and friction when training legged locomotion with PPO.
ExampleWhen training Dactyl, the value network could read extra information unavailable on the real robot, while the policy used only observations available on the real robot; once trained, the policy deployed directly to a Shadow dexterous hand to rotate a block.
- Also called
- Asymmetric AC
- Related
- Privileged Information · Teacher-Student Distillation · Value Function · Partially Observable Markov Decision Process · Sim-to-Real Transfer · Domain Randomization
- Sources
- Asymmetric Actor Critic for Image-Based Robot Learning (arXiv 1710.06542)
Learning Dexterous In-Hand Manipulation (OpenAI Dactyl, arXiv 1808.00177)