Embodied AI Glossary中文

Next-Token Prediction

下一个 token 预测NTPCommon

The training objective of predicting the next token in a sequence given everything that came before it.

This is the core training objective behind the GPT family of large language models. Text is cut into tokens; the model reads all the preceding tokens and outputs a probability distribution over the next one, and cross-entropy loss pushes up the probability assigned to the actual next token. It needs only raw text, no human labeling, so it scales to enormous pretraining datasets — OpenAI's GPT-1 (2018) already used it for generative pretraining. At inference time, the model generates tokens one after another, called autoregressive decoding. In embodied AI, RT-2 and OpenVLA discretize continuous actions into tokens and place them in the same sequence as text, reusing this same objective to train a VLA; a UC Berkeley team in 2024 even modeled humanoid-robot walking as next-token prediction. The cost is that discretization loses precision and generating tokens one at a time is slow, which is why some models switch to diffusion or flow-matching action heads instead.

ExampleOpenVLA divides each action dimension into 256 bins based on the training-data distribution, represents them using the 256 least-used tokens in the Llama vocabulary, and trains with the standard next-token-prediction objective, computing cross-entropy only on the action tokens.

Also called
NTP, Autoregressive Pretraining, Causal Language Modeling
Related
Autoregressive Decoding · Action Tokenizer · Action Binning · Teacher Forcing · Cross-Entropy · Vision-Language-Action Model
Sources
Improving Language Understanding by Generative Pre-Training (GPT-1, OpenAI 2018)
OpenVLA: An Open-Source Vision-Language-Action Model
Humanoid Locomotion as Next Token Prediction

See it in the full glossary →