Embodied AI Glossary中文

Context Length

上下文长度Advanced

The maximum number of tokens a model can process at once, which sets how much input it can 'see' at a time.

Context length, also called the context window, is the upper limit on the total number of tokens a Transformer-style model can process at once — measured in tokens, not characters or words. Anything beyond that limit either gets truncated or has to be summarized before being fed in. The compute cost of self-attention grows with the square of sequence length, and the GPU memory used by the KV cache grows linearly with it, so a longer context means slower inference and more memory used. According to an IBM roundup from October 2024, GPT-4o and Llama 3.1 support 128K tokens, while Gemini 1.5 Pro supports up to 2 million. For a VLA, images eat up context especially fast: at 224×224 resolution, PaliGemma turns a single image into 256 visual tokens, and multiple cameras plus a history of several frames can make the sequence grow very quickly — one reason many VLAs only look at the current frame, or need a separate memory module or visual-token pruning instead.

ExampleEncoding 3 camera views at 224×224 with PaliGemma uses 768 tokens for images alone; adding 4 frames of history per camera brings that to 12 images and 3,072 tokens.

Also called
Context Window
Related
Token · Key-Value Cache · Visual Token · Embodied Memory · Memory-Augmented VLA · Visual Token Pruning
Sources
What is a context window? (IBM Think)
PaliGemma – Google's Cutting-Edge Open Vision Language Model (Hugging Face Blog)
As of
2024-10

See it in the full glossary →