Embodied AI Glossary中文

Genie 2

Advanced

Google DeepMind's large-scale world model that generates a controllable 3D world from a single image.

Genie 2 is a foundation world model released by Google DeepMind on December 4, 2024, the successor to the original Genie. Given a single image — either a real photo or output from the Imagen 3 text-to-image model — it generates a 3D environment that can be controlled with a keyboard and mouse, staying visually consistent for up to about a minute, though most examples hold for 10–20 seconds. Its architecture is an autoregressive latent diffusion model trained on large-scale video data: it denoises and generates frame by frame inside a compressed latent space. It can simulate effects like gravity, water, and smoke, along with object interactions, and remembers content that has moved out of view. Its main use is providing diverse training and evaluation environments for agents such as SIMA; its successor is Genie 3.

ExampleGiven a real-world photo, Genie 2 can turn it into an interactive 3D scene that you can walk through in first person using the keyboard.

Related
Genie (Original) · Genie 3 · World Model · Interactive World Model · Latent Diffusion Model · SIMA 2
Sources
Genie 2: A large-scale foundation world model (Google DeepMind Blog)
As of
2024-12

See it in the full glossary →