Embodied AI Glossary中文

Byte-Pair Encoding

字节对编码BPEAdvanced

A tokenization algorithm that builds a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair.

Byte-pair encoding was first proposed by Philip Gage in 1994 as a text-compression algorithm; in 2016, Sennrich and colleagues applied it to tokenization for neural machine translation (ACL 2016), and it has since become the dominant scheme for tokenizers in large language models like GPT. The procedure starts from individual characters or bytes, counts the most frequent adjacent symbol pair in the corpus, merges it into a new symbol added to the vocabulary, and repeats until the vocabulary reaches a target size. This way, common words become a single token while rare words split into a few subwords, so the model never hits an out-of-vocabulary word and sequences don't get as long as character-by-character splitting would make them. Byte-level BPE first converts text into UTF-8 bytes before merging, so it can encode any text. In embodied AI, Physical Intelligence's FAST action tokenizer also uses BPE to compress action sequences.

ExampleFAST first applies a discrete cosine transform to each action dimension, scales and rounds the result, flattens it, and trains BPE on top (the paper defaults to a scale of 10 and a vocabulary of 1,024), compressing a stretch of action into a short token sequence for a VLA to predict autoregressively.

Also called
BPE, Byte-level BPE
Related
Tokenizer · Token · Action Tokenizer · Discrete Cosine Transform · π0-FAST · Large Language Model
Sources
Neural Machine Translation of Rare Words with Subword Units (arXiv 1508.07909)
FAST: Efficient Action Tokenization for Vision-Language-Action Models (arXiv 2501.09747)
Byte-pair encoding - Wikipedia

See it in the full glossary →