Key-Value Cache
KV 缓存KV CacheCommonStoring already-computed attention keys and values so later tokens can reuse them instead of recomputing.
The KV cache is the standard trick for speeding up Transformer inference. In self-attention, each token produces a query (Q), key (K), and value (V) vector; during autoregressive generation, only one new token is added at each step, and the K and V vectors for all earlier tokens never change. Without caching, every step would have to recompute the entire preceding sequence; with caching, each step only computes the new token and then attends to the stored K and V vectors, which is much faster. The cost is GPU memory: the cache grows linearly with context length and often becomes a bottleneck for long contexts, which is why techniques like cache quantization and sliding windows exist. The KV cache matters just as much for VLA models: π0 caches the K and V vectors for the prefix formed by the image and language inputs, so its 10-step flow-matching denoising process only has to recompute the action part, making one full inference pass take about 73 milliseconds.
ExampleOne π0 inference pass: image encoding takes about 14 ms and the prefix forward pass about 32 ms, computed only once; the action expert's 10 denoising steps take about 27 ms total, each step reusing the same cached prefix KV.
- Also called
- KV Cache, KV Caching
- Related
- Self-Attention · Causal Attention · Autoregressive Decoding · Inference Latency · Context Length · π0
- Sources
- Hugging Face Transformers: Cache strategies
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)