Embodied AI Glossary中文

Tokenizer

分词器Common

The preprocessing step that splits text, images, or actions into tokens and converts them into integer IDs.

A tokenizer is the preprocessing module in front of a model: it splits the input into tokens (the smallest units a model works with), then converts them into integer IDs using a fixed vocabulary; it's also used in reverse, to turn the model's output IDs back into text. Large language models mostly use subword algorithms such as byte-pair encoding (BPE), which repeatedly merges the most frequent adjacent character pairs — common words stay as a single token, rare words get split into several — and GPT-2's vocabulary has 50,257 tokens. In embodied AI the term has been extended: a video tokenizer compresses video into tokens, and an action tokenizer turns continuous joint motion into discrete tokens, so a language model can output actions the same way it outputs words. How something is tokenized determines how long the sequence is and what the model can represent, making it a key design choice for a VLA.

ExampleOpenVLA divides each action dimension into 256 bins and directly repurposes the 256 least-used tokens in the Llama tokenizer's vocabulary to represent them, so one step of a 7-dimensional action becomes 7 tokens.

Related
Token · Byte-Pair Encoding · Action Tokenizer · Video Tokenizer · Visual Token · Action Binning
Sources
Hugging Face Transformers Docs: Tokenization algorithms
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)

See it in the full glossary →