Teacher Forcing
教师强制AdvancedTraining a sequence model by feeding it the real previous step at every step, instead of its own prediction.
Teacher forcing is the standard way to train autoregressive sequence models, named by Williams and Zipser in 1989: when predicting step t, the input uses the real first t−1 steps from the data, not the model's own generated output. This lets every step's loss be computed in parallel, so training is fast and stable, and it's how large language models and VLAs that discretize actions into tokens, such as RT-2 and OpenVLA, are trained. The cost is a mismatch between training and inference: at inference the model can only keep conditioning on its own output, so small early errors accumulate, a problem called exposure bias. Scheduled sampling (2015) gradually mixes in the model's own predictions during training; Self Forcing (2025), in video generation, instead unrolls the model autoregressively during training itself.
ExampleWhen training OpenVLA, each action is discretized into 7 tokens, and the k-th token is predicted conditioned on the true first k−1 tokens; at deployment, though, it can only keep decoding after the tokens it just generated itself.
- Related
- Next-Token Prediction · Autoregressive Decoding · Exposure Bias · Self Forcing · Diffusion Forcing · Compounding Error
- Sources
- Wikipedia: Teacher forcing
Bengio et al. 2015: Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
Huang et al. 2025: Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion