Unsupervised Skill Discovery
无监督技能发现AdvancedWith no task reward, letting an agent practice into a set of distinguishable skills it can call on later.
Unsupervised skill discovery lets an agent spontaneously learn a variety of different behaviors with no external reward at all. The best known method, DIAYN (“Diversity is All You Need”), was proposed by Eysenbach, Levine, and colleagues in 2018: the policy is given a skill number as input, and a discriminator is trained at the same time to guess which skill was used just from the states reached, so the intrinsic reward rises the more accurately it can be guessed, plus a max-entropy term to encourage varied actions; underneath, this maximizes the mutual information between skill and state. Simulated robots have spontaneously learned to walk, hop, and more this way. Skills can serve as a pretraining initialization, or be composed by a higher-level policy to solve sparse-reward tasks. METRA (ICLR 2024) instead learns skills in a latent space that preserves temporal distance, easing the weak-exploration problem that mutual-information methods tend to have.
ExampleSharma and colleagues (2020) turned the skill-discovery algorithm DADS into an off-policy version and, with no reward and no demonstrations at all, had a real quadruped robot learn distinct gaits and headings, then chained them together with model-predictive control to perform navigation.
- Also called
- DIAYN, Reward-Free Skill Learning
- Related
- Intrinsic Motivation · Hierarchical Reinforcement Learning · Exploration vs. Exploitation · Behavior Foundation Model · BFM-Zero · Entropy Regularization
- Sources
- Eysenbach et al. 2018: Diversity is All You Need: Learning Skills without a Reward Function
Sharma et al. 2020: Emergent Real-World Robotic Skills via Unsupervised Off-Policy Reinforcement Learning
Park, Rybkin, Levine 2023: METRA: Scalable Unsupervised RL with Metric-Aware Abstraction (ICLR 2024)