Embodied AI Glossary中文

Stable Video Diffusion

SVDAdvanced

Stability AI's open-source image-to-video diffusion model, commonly used in robotics research as a backbone for video prediction.

Stable Video Diffusion is an open-source video generation model released by Stability AI on November 21, 2023, a latent diffusion model (which compresses images into a low-dimensional latent space and denoises step by step within it). It adds temporal layers on top of the Stable Diffusion image model and summarizes training into three stages — text-to-image pretraining, large-scale video pretraining, and high-quality video fine-tuning — emphasizing how much video-data curation matters. The initial release included two image-to-video models generating 14 and 25 frames respectively (the latter called SVD-XT), with a settable frame rate of 3 to 30 fps; it launched as a research preview, not for commercial use. Thanks to its open weights and modest size of about 1.5 billion parameters, SVD became a commonly used video backbone in embodied AI — for example, Video Prediction Policy (VPP) adds language conditioning on top of SVD, fine-tunes it on robot video, and extracts predictive representations of the future from it to output actions.

ExampleVideo Prediction Policy (VPP) fine-tunes SVD into a manipulation-video prediction model: given the current frame and the instruction 'open the drawer,' it first predicts features of the upcoming frames and then uses them to output the robot arm's actions.

Also called
SVD-XT, Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Related
Video Generation Model · Latent Diffusion Model · Video Prediction Policy · Text-to-Video / Image-to-Video · World Model · Diffusion Model
Sources
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets (arXiv 2311.15127)
Introducing Stable Video Diffusion (Stability AI)
Video Prediction Policy (arXiv 2412.14803)
As of
2023-11

See it in the full glossary →