Embodied AI Glossary中文

Sora

Sora(视频生成即世界模拟器)Common

OpenAI's text-to-video model, introduced with a technical report arguing that video generation models can double as world simulators.

Sora is OpenAI's video generation model, first shown in February 2024 and opened to paying users that December, alongside a technical report titled “Video generation models as world simulators.” It's a diffusion Transformer: video is compressed into a latent space, cut into spacetime patches used as tokens, and then generated by progressive denoising. The report argues that scaling up video generation models naturally produces abilities like 3D consistency and object permanence, and may be a path toward a simulator of the physical world. This connected the video-generation and world-model lines of research, and also sparked debate over whether it truly understands physics — the report itself acknowledges that interactions like glass shattering aren't simulated accurately. Sora 2, with a companion social app, launched on September 30, 2025; OpenAI announced it was discontinuing Sora in March 2026, the app was taken down on April 26, and the API was shut off on September 24.

ExampleIn the technical report, Sora generated Minecraft footage while simultaneously controlling the player character and rendering the surrounding world — cited as an example of it “simulating a digital world.”

Also called
Sora 2, Video Generation Models as World Simulators
Related
Video Generation Model · World Model · Diffusion Transformer · Spacetime Patches · World Foundation Model · Intuitive Physics
Sources
Sora (text-to-video model) - Wikipedia
Sora (人工智能模型) - 维基百科 (Chinese)
As of
2026-09

See it in the full glossary →