Wan (Alibaba Video Generation Model)
通义万相AdvancedAlibaba's Tongyi-team video generation model; Wan2.1 and Wan2.2 are open-weight and widely used as a base for world models.
Wan (Tongyi Wanxiang in Chinese) is the video generation model family from Alibaba's Tongyi team. Wan2.1, open-sourced in late February 2025, comes in 1.3B and 14B sizes under an Apache 2.0 license; it's a diffusion Transformer trained with flow matching, ships with its own 3D causal video VAE (Wan-VAE) and a umT5 text encoder, and supports text-to-video, image-to-video, and video editing. Wan2.2, released in July 2025, switched to a mixture-of-experts design with two roughly 14B experts — a high-noise expert that sets overall layout and a low-noise expert that fills in detail, with only one active per step — plus a 5B text-and-image-to-video model; audio-driven (S2V) and character-animation (Animate) variants followed. Because Wan has learned a great deal about how objects move from massive video data, and its weights are open, researchers often use it as a starting point for world models or world-action models. As of September 2026, Wan2.2 remains the main open-weight release.
ExampleNVIDIA's DreamZero builds on the Wan2.1-I2V-14B-480P image-to-video model as its backbone, continuing training on robot data so the model generates both future images and the corresponding actions at once — functioning as a world-action model that directly controls the robot.
- Also called
- Tongyi Wanxiang, Wan2.1, Wan2.2
- Related
- Video Generation Model · Diffusion Transformer · Video Tokenizer · Flow Matching · Mixture of Experts · DreamZero
- Sources
- Wan: Open and Advanced Large-Scale Video Generative Models (arXiv:2503.20314)
Wan-Video/Wan2.2 (GitHub)
World Action Models are Zero-shot Policies (DreamZero, arXiv:2602.15922) - As of
- 2026-09