Unified Multimodal Model
统一多模态模型UMMAdvancedA single model that can both understand images and video by answering questions about them, and generate or edit images and video.
A unified multimodal model puts multimodal understanding (answering questions about an image) and visual generation (drawing or editing an image, or generating video, from text) into a single model. These used to be two separate lines of work: understanding mostly used autoregressive multimodal large language models, and generation mostly used diffusion models. GPT-4o's native image generation brought wide attention to this direction, and a leading open-source example is ByteDance Seed's BAGEL, released in May 2025 (a Mixture-of-Transformers design, with 7 billion active parameters and 14 billion total). Roughly speaking, these fall into three categories by generation method: autoregressive, diffusion, and hybrid autoregressive-diffusion. Embodied AI cares about this direction because a single model that can both understand a scene and 'imagine what it will look like after this action' can double as a policy and a world model at once.
ExampleAlibaba DAMO Academy's WorldVLA puts a VLA and a world model into the same autoregressive framework: it outputs an action given the current view, and also predicts the next frame given the view and an action; the paper reports the two reinforce each other, outperforming a standalone action model or world model.
- Also called
- UMM, Unified Understanding and Generation Model
- Related
- Multimodal Large Language Model · Native Multimodal · BAGEL · Hybrid Autoregressive-Diffusion Architecture · World Model · WorldVLA
- Sources
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities (arXiv:2505.02567)
BAGEL (ByteDance-Seed GitHub)
WorldVLA: Towards Autoregressive Action World Model (arXiv:2506.21539) - As of
- 2025-06