Visual Token Pruning
视觉 token 剪枝AdvancedDropping or merging unimportant image tokens at inference time so vision-language and VLA models run faster.
VLMs and VLAs cut each image into hundreds of visual tokens — vectors representing small image patches — before feeding them into the language model, and the count multiplies with multiple cameras or multiple frames. Since attention computation scales roughly quadratically with the number of tokens, this is a major source of inference latency. Visual token pruning keeps only a small set of important tokens, ranked by some importance score (commonly how much attention text or action tokens pay to them), and drops or merges the rest; many methods need no retraining and can be applied directly. The landmark method FastV (ECCV 2024) found that a model's attention to image tokens becomes very sparse after the second layer, so it prunes half of them after the shallow layers, cutting LLaVA-1.5-13B's computation by about 45% with almost no drop in performance. In robotics this directly affects control frequency: EfficientVLA combines visual token selection, layer skipping, and caching to speed up CogACT by 1.93×, and 2025–2026 saw a wave of pruning methods designed specifically for VLAs.
ExampleIn FastV, the image first passes through the language model's first two layers as usual; after that, only the half of image tokens with the highest attention scores continue into later layers. No weights change and no retraining is needed, yet LLaVA-1.5's inference computation drops noticeably.
- Also called
- Token Pruning, Visual Token Compression
- Related
- Visual Token · Inference Latency · Pruning · Attention Mechanism · Key-Value Cache · On-Device / Edge Deployment
- Sources
- An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models (FastV, arXiv:2403.06764)
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models (arXiv:2506.10100) - As of
- 2026-09