Speculative Decoding
投机解码AdvancedHaving a small model quickly guess a few tokens, then having the large model verify them all in one parallel pass, with no change in output.
Speculative decoding is a large-model inference speedup technique, proposed by Google's Leviathan and colleagues in late 2022 (ICML 2023), with DeepMind's Chen and colleagues independently proposing the same idea as 'speculative sampling' around the same time. An autoregressive model has to run the full large model for every single token it generates, which is slow. Speculative decoding instead has a cheap draft model guess several tokens in a row, then has the large model score all of them in one forward pass; a specific acceptance rule keeps the prefix of correct guesses and resamples from the first wrong position onward — so the output distribution exactly matches what the large model would have produced generating alone, with no retraining needed. Google measured a 2-3x speedup on T5-XXL. In embodied AI, VLAs that output discrete action tokens, such as OpenVLA, use it to cut inference latency too, and 2026 saw improved methods that incorporate kinematic information.
ExampleSpec-VLA (EMNLP 2025) found that applying standard speculative decoding directly to a VLA gave limited speedup, so it relaxes the acceptance condition using the relative distance between action tokens, raising accepted length by 44% on OpenVLA for a 1.42x speedup with no drop in success rate.
- Also called
- Speculative Sampling
- Related
- Autoregressive Decoding · Inference Latency · Parallel Decoding · Key-Value Cache · Vision-Language-Action Model · OpenVLA
- Sources
- Fast Inference from Transformers via Speculative Decoding (arXiv:2211.17192)
Accelerating Large Language Model Decoding with Speculative Sampling (arXiv:2302.01318)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance (arXiv:2507.22424) - As of
- 2026-03