Embodied AI Glossary中文

Hybrid Autoregressive-Diffusion Architecture

自回归-扩散混合架构Advanced

An architecture where discrete content is generated autoregressively and continuous content is generated by diffusion, in the same model.

Autoregressive means predicting the next token step by step in sequence, as large language models do, which suits discrete content like text and reasoning; a diffusion model starts from noise and denoises step by step, which suits continuous, high-dimensional data like images and motion. A hybrid architecture puts both inside the same network. Transfusion, proposed by Meta and others in 2024, uses a single Transformer to process a sequence that interleaves text and images, computing a next-token-prediction loss for text and a diffusion loss for images; Kaiming He and colleagues' MAR instead models each continuous-valued token with a diffusion loss inside an autoregressive framework, eliminating the need for vector quantization. In robotics, 2025's HybridVLA has a single large language model do both diffusion denoising and autoregressive action prediction, then adaptively fuses the two resulting actions. The motivation is that forcing continuous actions into discrete tokens loses precision, while a purely diffusion-based action head has more trouble directly using a language model's reasoning ability.

ExampleInside the same language model, HybridVLA both generates a continuous action through diffusion denoising and predicts discretized action tokens autoregressively, then fuses the two results into the final action at execution time.

Also called
AR + Diffusion, Hybrid AR-Diffusion
Related
HybridVLA · Diffusion Model · Autoregressive Decoding · Unified Multimodal Model · Diffusion Action Head · Action Tokenizer
Sources
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model (arXiv:2408.11039)
Autoregressive Image Generation without Vector Quantization (MAR, arXiv:2406.11838)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified VLA Model (arXiv:2503.10631)
As of
2025-06

See it in the full glossary →