Embodied AI Glossary中文

Self-Play

自博弈Advanced

Having an agent compete against itself, or past versions of itself, to keep improving through the outcomes.

Self-play is a training approach within multi-agent reinforcement learning where the opponent is not a human or a fixed script but the agent's own current or past self. As the agent gets stronger, its opponent gets stronger in lockstep, which amounts to an automatically generated curriculum of rising difficulty, and it needs no human match data at all. The most famous example is DeepMind's 2017 AlphaGo Zero: starting from random moves and using only self-play plus Monte Carlo tree search, after three days of training it beat the earlier AlphaGo version that had defeated the human champion, 100 games to 0. In robotics, adversarial tasks commonly use it: DeepMind trained the small humanoid robot OP3 to play one-on-one soccer by first training standing-up and scoring skills separately, distilling both into one policy, and then having it compete against snapshots of its own past self. Playing only against the very latest version of itself tends to cause cycling, so opponents are usually drawn from a pool of historical versions instead.

ExampleIn OP3's soccer training, opponents were drawn from a pool of the agent's own periodically saved historical snapshots; an ablation study showed that an agent trained without self-play performed worse even against a fixed opponent.

Also called
Self-Competition
Related
Multi-Agent Reinforcement Learning · Reinforcement Learning · Monte Carlo Tree Search · Curriculum Learning · Population-Based Training · OP3 Soccer (DeepMind)
Sources
Google DeepMind Blog: AlphaGo Zero: Starting from scratch
Haarnoja et al. 2023: Learning Agile Soccer Skills for a Bipedal Robot with Deep Reinforcement Learning

See it in the full glossary →