Diffusion Forcing
扩散强制AdvancedA causal diffusion sequence model trained by adding an independently random noise level to each token in the sequence.
Diffusion Forcing was proposed in 2024 by MIT's Boyuan Chen, Vincent Sitzmann, Russ Tedrake, and colleagues, published at NeurIPS 2024. Sequence generation traditionally comes in two flavors: next-token prediction generates one token at a time with flexible length, but tends to drift off course over long rollouts; full-sequence diffusion denoises an entire span at once, which allows guiding the whole trajectory but fixes its length. Diffusion Forcing trains a causal model in which each token in the sequence gets its own independently random noise level; at inference, it can then generate frame by frame like an autoregressive model, extending past the training length, while also using guidance like a diffusion model to steer the whole trajectory toward a desired goal. It's been used for long video generation, planning, and robot control, and is a training scheme that autoregressive video generation and world-model work often reuses or compares against.
ExampleIn the paper's real-robot experiment, an arm has to swap two randomly placed pieces of fruit using a third slot as a buffer, which requires remembering the initial positions; Diffusion Forcing completed the task, while an imitation-learning baseline without memory failed.
- Related
- Diffusion Model · Autoregressive Video Generation · Next-Token Prediction · Teacher Forcing · World Model · Self Forcing
- Sources
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion (arXiv:2407.01392)
Diffusion Forcing project page