Embodied AI Glossary中文

Video Generation Model

视频生成模型Common

A generative model that produces a new, coherent video from text, an image, or an existing clip.

A video generation model can produce new video from text, an image, or an existing clip. Google's 2022 paper Video Diffusion Models extended image diffusion models to video; since then, the dominant approach has been to compress video into latent space with a video VAE first, then denoise multiple frames together step by step with a diffusion Transformer or U-Net, so consecutive frames stay coherent. Notable examples include OpenAI's Sora, Google's Veo, and Alibaba's open-source Wan. Because these models learn how objects move and get pushed around from huge amounts of video, they're often treated as a starting point for world models: NVIDIA's Cosmos and other world foundation models, 'generate video then infer actions' policies like UniPi, and world action models like DreamZero are all built on top of video generation models.

ExampleWan is released open-source in two sizes, 1.3B and 14B; the 1.3B version needs only about 8 GB of GPU memory, so text-to-video runs on a consumer graphics card.

Also called
Video Diffusion Model
Related
Text-to-Video / Image-to-Video · Diffusion Model · Diffusion Transformer · Video Prediction Model · World Model · Wan (Alibaba Video Generation Model)
Sources
Video Diffusion Models (arXiv 2204.03458)
Wan: Open and Advanced Large-Scale Video Generative Models (arXiv 2503.20314)
As of
2025-04

See it in the full glossary →