Embodied AI Glossary中文

Token

token(词元)Essential

The basic unit a model processes data in; text, image patches, and actions can all be broken into tokens.

A token is the basic unit that Transformer-style models read and produce. Text is first cut by a tokenizer into words or sub-word pieces; each piece has an index in a vocabulary, which is then turned into a vector, an embedding, and fed into the model — a large language model's training objective is exactly to predict the next token. In English, a token is often a word or part of a word; in Chinese, a single character may be its own token, or may get split up or merged with neighboring characters, depending on the tokenizer. The idea has since spread to other modalities: ViT cuts an image into 16×16-pixel patches, each one becoming a visual token; RT-2 and OpenVLA discretize each action dimension into 256 bins, each bin becoming a token, so a language model can output actions the same way it writes text. Context length is measured in tokens, and inference time grows with token count too.

ExampleOpenVLA repurposes the 256 least-used tokens in the Llama tokenizer's vocabulary as action bins, so a 7-dimensional action — position, orientation, and gripper open/close — comes out as 7 tokens.

Also called
Marker
Related
Tokenizer · Embedding · Visual Token · Action Tokenizer · Next-Token Prediction · Context Length
Sources
Tokenizers (Hugging Face LLM Course, Chapter 2)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT, arXiv:2010.11929)

See it in the full glossary →