KV Cache
The KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable.
The KV cache stores each token's Key and Value vectors (from attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.) the first time they're computed during inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering., reusing them for every later generation step instead of recomputing the whole sequence from scratch each time. This turns a quadratic-cost operation into a linear one, and is essential to making LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. generation fast enough to serve. Its main cost is GPU memory, which grows with sequence length and batch size.
How it works
In a decoder-only transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs., the query for the newest token attends over the Keys and Values of every preceding token. Because causal maskingCausal MaskingCausal masking restricts self-attention so each token can only attend to earlier positions, not future ones — required for training and running autoregressive generation. means earlier positions never see later ones, their K and V vectors are fixed once computed. The cache stores them per layer and per head, so each decode step computes K and V for exactly one new position, appends them, and attends over the stored tensors.
Size is predictable: roughly 2 * layers * kv_heads * head_dim * seq_len * batch * bytes_per_element, where the leading 2 covers Keys
and Values. Grouped-query attention shrinks kv_heads below the number
of query heads specifically to cut this term, and the cache is often
kept in fp8 or int8 for the same reason.
When it breaks
- It is the real memory ceiling. Weights are fixed, but cache grows with every concurrent request and every token generated, so long contexts starve batchingBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish. capacity.
- Fragmentation wastes capacity. Reserving a contiguous block for each request's maximum length leaves most of it unused; paged allocation schemes exist to reclaim that waste.
- Any prefix edit invalidates it. Inserting a system prompt, trimming old turns, or reordering retrieved context changes every downstream Key and Value, forcing a full prefill again.
- Cache quantization is not free. Dropping to very low precision degrades long-range recall well before it degrades short prompts, so it often passes quick tests and fails in production.
See also: InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering., AttentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.
Learn more: Inference & Serving
Mentioned in
Lessons where this comes up in context.
- Attention & TransformersThe core architecture behind modern AI
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- Inference & ServingBatching, KV caching, quantization, latency/throughput tradeoffs
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
Knowledge Distillation
Knowledge distillation trains a small "student" model to mimic a larger "teacher" model's output distribution, transferring most of its capability at a fraction of the serving cost.
Speculative Decoding
Speculative decoding uses a small draft model to guess several tokens ahead, then verifies them all in one parallel pass with the full model, speeding up generation.