Embodied AI Glossary中文

BAGEL

Advanced

ByteDance Seed's open-source unified multimodal model that can both understand images and generate or edit them.

BAGEL is a unified multimodal foundation model ByteDance's Seed team open-sourced in May 2025, with code and weights released under the Apache 2.0 license; a single model handles image-text understanding, text-to-image generation, and image editing all at once. It's built on the Qwen2.5 large language model and uses a Mixture-of-Transformers (MoT) design: understanding and generation each get their own set of parameters, but every token shares the same self-attention at each layer, giving it 7B active parameters and 14B total. The understanding side uses a SigLIP2 vision encoder; the generation side uses FLUX's VAE to compress images into latent space. The model was pretrained on trillions of tokens of interleaved image-text, video, and web data, and the paper reports that capabilities such as free-form image editing, future-frame prediction, viewpoint rotation, and 'world navigation' emerge as scale increases. It demonstrates putting understanding and generation into one model, which is relevant to the unified-multimodal-model and world-model directions in embodied AI.

ExampleGiven BAGEL an image and an editing instruction in words, it can output the edited image directly; given navigation commands like 'move forward' or 'turn,' it can generate the view after that viewpoint change.

Also called
BAGEL-7B-MoT, Emerging Properties in Unified Multimodal Pretraining
Related
Unified Multimodal Model · Mixture-of-Transformers · SigLIP · Variational Autoencoder · World Model · ByteDance Seed
Sources
BAGEL: Emerging Properties in Unified Multimodal Pretraining (arXiv 2505.14683)
ByteDance-Seed/Bagel GitHub
BAGEL 论文 HTML 版 (Chinese)
As of
2025-05

See it in the full glossary →