Embodied AI Glossary中文

Model-Based Reinforcement Learning

基于模型的强化学习MBRLCommon

Learning a model that predicts how the environment will respond, then using it to plan or “imagine” training data for a policy.

This is a major branch of reinforcement learning. Beyond learning a policy, the agent also learns a dedicated environment model (also called a dynamics model or world model): given the current state and action, it predicts the next state and reward. With a model in hand, an agent can plan by rolling the model forward before acting (as in model predictive control), or generate large numbers of “imagined” trajectories from the model to train a policy — Sutton's Dyna architecture is an early example of the latter. The benefit is saving real interaction and improving sample efficiency, which matters a lot when trial and error on a real robot is costly; the risk is that once the model is inaccurate, the policy learns to exploit the model's blind spots, called model bias. The Dreamer series and TD-MPC2 both belong to this family, and it's also where world-model research and robot reinforcement learning meet.

ExampleDayDreamer (2022) ran the Dreamer algorithm directly on real hardware: a quadruped robot learned to roll over, stand up, and walk from scratch in just 1 hour, with no simulator involved.

Also called
MBRL, Model-Based RL
Related
Model-Free Reinforcement Learning · World Model · Learning in Imagination · Sample Efficiency · DreamerV3 · Model Predictive Control
Sources
OpenAI Spinning Up: Kinds of RL Algorithms
Model-based Reinforcement Learning: A Survey (Moerland et al.)
DayDreamer: World Models for Physical Robot Learning

See it in the full glossary →