Embodied AI Glossary中文

Diffusion Forcing

扩散强制Advanced

A causal diffusion sequence model trained by adding an independently random noise level to each token in the sequence.

Diffusion Forcing was proposed in 2024 by MIT's Boyuan Chen, Vincent Sitzmann, Russ Tedrake, and colleagues, published at NeurIPS 2024. Sequence generation traditionally comes in two flavors: next-token prediction generates one token at a time with flexible length, but tends to drift off course over long rollouts; full-sequence diffusion denoises an entire span at once, which allows guiding the whole trajectory but fixes its length. Diffusion Forcing trains a causal model in which each token in the sequence gets its own independently random noise level; at inference, it can then generate frame by frame like an autoregressive model, extending past the training length, while also using guidance like a diffusion model to steer the whole trajectory toward a desired goal. It's been used for long video generation, planning, and robot control, and is a training scheme that autoregressive video generation and world-model work often reuses or compares against.

ExampleIn the paper's real-robot experiment, an arm has to swap two randomly placed pieces of fruit using a third slot as a buffer, which requires remembering the initial positions; Diffusion Forcing completed the task, while an imitation-learning baseline without memory failed.

Related
Diffusion Model · Autoregressive Video Generation · Next-Token Prediction · Teacher Forcing · World Model · Self Forcing
Sources
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion (arXiv:2407.01392)
Diffusion Forcing project page

See it in the full glossary →