Embodied AI Glossary中文

Visual Token

视觉 tokenCommon

One of the vectors produced by cutting an image into patches and encoding them — the basic unit a model uses for images.

A visual token is the basic unit an image is broken into before entering a Transformer. The common approach traces back to ViT: the image is cut into fixed-size patches, and each patch is turned into a vector by a vision encoder — that vector is one visual token; a vision-language model then uses a projector to map these into the same space as text tokens, concatenating them with the instruction into a single sequence for joint processing. The token count grows with the square of resolution: PaliGemma produces 256 tokens for a 224-pixel image, 1,024 for 448 pixels, and 4,096 for 896 pixels. Robots typically have multiple cameras and need to run inference many times per second, so visual tokens often make up the bulk of the sequence and directly drive up inference latency — which is why compression methods like visual token pruning and resampling exist. The equivalent unit in video models is the spacetime patch.

Exampleπ0 is built on PaliGemma: each camera's image is first encoded into visual tokens, which are then concatenated with the language instruction's tokens into the same sequence fed into the model.

Also called
Image Token, Patch Token
Related
Vision Transformer · Token · Projector / Connector · Visual Token Pruning · Spacetime Patches · Vision Encoder
Sources
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (arXiv 2010.11929)
PaliGemma: A versatile 3B VLM for transfer (arXiv 2407.07726)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)

See it in the full glossary →