World Model
世界模型WMEssentialA model that predicts what the world will look like after a given action is taken.
A world model is one that has learned how an environment changes: given the current state or frame and an action, it predicts the next state or frame. Ha and Schmidhuber's 2018 paper “World Models” popularized the term: it compresses game frames into a low-dimensional vector with a variational autoencoder, then uses a recurrent network to predict the next step, letting an agent learn a policy inside the model's generated “dream” and transfer it back to the real game. Today the term covers models that predict purely in latent space, meaning compressed features, such as DreamerV3, which serves reinforcement learning, and Meta's V-JEPA 2, used for robot planning, as well as large models that directly generate video frames, such as Google DeepMind's Genie 3 and NVIDIA's Cosmos. In embodied AI, world models are used to generate synthetic training data, evaluate policies “in imagination,” and plan, all to cut down on real-robot trial and error.
ExampleGenie 3, released in August 2025, generates a 720p, 24-frames-per-second scene from a text description; as the user presses direction keys to move around, it generates what comes next in real time, and can remember what it saw roughly a minute earlier.
- Also called
- WM, World Simulator
- Related
- Latent World Model · World Foundation Model · World Action Model · NVIDIA Cosmos · Genie 3 · Model-Based Reinforcement Learning
- Sources
- World Models (Ha & Schmidhuber, arXiv:1803.10122)
Genie 3: A new frontier for world models (Google DeepMind)
What Are World Models? (NVIDIA Glossary)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv:2506.09985) - As of
- 2025-08