Embodied AI Glossary中文

PlaNet

Advanced

A model-based reinforcement learning agent that learns a latent-space world model from pixels and plans actions by imagining outcomes.

PlaNet was released by Danijar Hafner and colleagues at Google (a Google Brain and DeepMind collaboration) in November 2018 and published at ICML 2019. It works from images alone: it first learns a world model that compresses each frame into a latent state and predicts how that latent state changes under different actions, and how much reward it would earn. The model's core is the Recurrent State-Space Model (RSSM), which combines a deterministic path and a stochastic path so multi-step predictions stay stable; the paper also introduces a multi-step training objective called latent overshooting. At decision time, PlaNet does not train a separate policy network — instead, it imagines many candidate action sequences inside the latent space and picks the one with the highest predicted return, executing only its first step before replanning (online planning). On the DeepMind Control continuous-control benchmark, it needs far fewer interaction episodes than model-free methods; Google's blog reported roughly 50 times better data efficiency on average. The later Dreamer series reuses the RSSM but learns a policy inside imagination instead.

ExampleOn the 'cheetah run' task, PlaNet works from camera images alone, comparing tens of thousands of imagined action sequences in latent space at every step and executing just the first action of the sequence with the highest predicted reward.

Also called
Deep Planning Network, PlaNet: Learning Latent Dynamics for Planning from Pixels
Related
Recurrent State-Space Model · DreamerV3 · World Model · Model-Based Reinforcement Learning · Model Predictive Control · Cross-Entropy Method
Sources
arXiv 1811.04551: Learning Latent Dynamics for Planning from Pixels
ICML 2019 论文页(PMLR v97) (Chinese)
Google Research Blog: Introducing PlaNet
As of
2019-06

See it in the full glossary →