Speculative Decoding
Speculative decoding uses a small draft model to guess several tokens ahead, then verifies them all in one parallel pass with the full model, speeding up generation.
Speculative decoding speeds up LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. by exploiting spare GPU compute during the memory-bandwidth-bound decode phase: a small, fast "draft" model generates several candidate next tokens ahead of time, and the full model verifies all of them in a single parallel forward pass. Matching guesses are accepted together — several tokens for roughly the cost of one decode step; a wrong guess falls back to the full model's actual choice. Output is mathematically identical to normal decoding, just faster.
How it works
Each round has two halves. The draft model generates k candidate
tokens autoregressively — cheap, because it is small. The target model
then scores all k positions in one forward pass, exactly as it would
score a prompt during prefill, producing a distribution for every
candidate at once.
The candidates are accepted left to right by a rejection-sampling rule that compares the target and draft probabilities for each token. At the first rejection, that token is resampled from a corrected distribution and the remaining drafts are discarded. Because the acceptance rule is constructed to preserve the target model's distribution, the output matches ordinary sampling. Variants drop the separate draft model entirely and predict candidates with extra heads on the target model itself.
When it breaks
- Acceptance rate is everything. If the draft disagrees often, you pay for drafting plus a near-normal decode step and end up slower than baseline. Draft and target must share a tokenizerTokenizationTokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE). and ideally training data.
- It only helps when compute is spare. The gain comes from unused FLOPs during memory-bound decode; under heavy batchingBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish. the server is already compute-bound and speculation competes with real work.
- Extra memory. Draft weights and their own KV cache sit alongside the target model's, reducing the cache budget available for concurrency.
- Tuning
kis workload-specific. Long drafts win on predictable text and waste work on high-entropy output, and high sampling temperature lowers acceptance.
See also: InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering., KV CacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable.
Learn more: Inference & Serving
Mentioned in
Lessons where this comes up in context.
KV Cache
The KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable.
Tensor Parallelism
Tensor parallelism splits individual weight matrices across multiple GPUs so a model too large for one GPU's memory can still be trained or served.