Embodied AI Glossary中文

DIAMOND

DIAMOND(扩散世界模型)Advanced

Generates frames one at a time with a diffusion model as a world model, letting an RL agent train inside a generated game.

DIAMOND was released in May 2024 by researchers at the University of Geneva, the University of Edinburgh, and Microsoft Research, a NeurIPS 2024 Spotlight paper. Many earlier world models (such as IRIS and DreamerV3) compress frames into discrete tokens or a latent variable before predicting, which easily loses small but important visual details. DIAMOND instead uses a diffusion model directly to generate the next frame from recent frames and the action, and works out the key design choices needed to keep this stable over long rollouts. An agent is trained with reinforcement learning entirely inside this generated environment, then tested in the real game: on the Atari 100k benchmark it reaches an average human-normalized score of 1.46, the best result at the time for an agent trained purely with a world model. It also shows that a diffusion world model can serve as a playable, interactive game engine, in a similar spirit to GameNGen and Genie.

ExampleThe authors trained a 381-million-parameter diffusion world model on 87 hours of human Counter-Strike: Global Offensive (CS:GO) match data; it runs at about 10 frames per second on an RTX 3090, and a person can “play” inside it in real time with a keyboard and mouse.

Also called
Diffusion for World Modeling, DIAMOND: Visual Details Matter in Atari
Related
World Model · Diffusion Model · Learning in Imagination · GameNGen · DreamerV3 · Arcade Learning Environment (ALE) / Atari 100k
Sources
Diffusion for World Modeling: Visual Details Matter in Atari (arXiv:2405.12399)
DIAMOND 项目主页 (Chinese)
As of
2024-10

See it in the full glossary →