MuZero
AdvancedA DeepMind algorithm that masters Go and Atari through tree search inside a learned internal model, with no rules given.
MuZero was posted as a preprint by Julian Schrittwieser, David Silver, and colleagues at DeepMind in November 2019, and published in Nature in December 2020, the successor to AlphaGo and AlphaZero. AlphaZero needs to know the rules of a game to simulate it out “in its head”; MuZero no longer needs the rules at all: it learns a latent-space model that predicts only the three things most useful for decision-making — reward, policy, and value — without reconstructing the full image, then uses Monte Carlo tree search to plan ahead and choose an action inside that learned model. It reached AlphaZero's level at Go, chess, and shogi, and beat every algorithm at the time across 57 Atari games. It is a representative example of model-based reinforcement learning, embodying the idea that a world model only needs to serve planning, and is often discussed alongside the Dreamer series.
ExamplePlaying an Atari game, MuZero is never told the rules; seeing only the screen and the score, it simulates the consequences of several different moves inside its own learned internal model and picks whichever has the highest expected return.
- Also called
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
- Related
- Model-Based Reinforcement Learning · World Model · Monte Carlo Tree Search · Latent World Model · Value Function · DreamerV3
- Sources
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (arXiv 1911.08265)
MuZero: Mastering Go, chess, shogi and Atari without rules (Google DeepMind blog) - As of
- 2020-12