Inference & Serving
Batching, KV caching, quantization, latency/throughput tradeoffs
Everything so far has been about training — producing a model's weights. InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. is using a trained model to generate output, and it turns out to be its own deep engineering problem: the naive way to run a transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. is far too slow and expensive to serve to real users at scale. This lesson covers the main techniques that make LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. serving practical.
Autoregressive generation is inherently sequential
An LLM generates text one token at a time: predict a distribution over the next token, sample one, append it to the sequence, and repeat — using its own previous output as input to produce the next token. This is called autoregressive generation, and it has an unavoidable consequence: you cannot generate token 50 before you've generated token 49, because token 50's prediction depends on it. Unlike training (where the whole target sequence is known upfront and attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. can be computed in parallel across all positions), inference is fundamentally a sequential loop.
This sequential nature is the root cause of most of the engineering problems in this lesson.
Prefill vs. decode: two different phases
Serving a request actually involves two phases with very different performance characteristics:
- Prefill: process the entire input prompt at once. Every prompt token's Key and Value can be computed in parallel — no sequential dependency yet, since none of these tokens depend on a model-generated token. This phase is compute-bound: it's a large, parallel matrix operation, so it's mostly limited by raw GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. compute throughput.
- Decode: generate the output, one token at a time, autoregressively. Each step does comparatively little computation (one new token) but must read the entire growing KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable. from GPU memory at every step. This phase is memory-bandwidth-bound: the bottleneck is how fast data can move, not how much math the GPU can do.
Put end to end, a single request moves through a fixed pipeline — one prefill pass, then a loop that runs once per generated token:
Why is generating a response slower than reading the prompt?
This distinction is why time to first token and tokens per second (below) behave so differently and are optimized somewhat separately: a long prompt mostly costs prefill time; a long generated response mostly costs decode time, dominated by repeatedly reading an ever-growing KV cache.
| Prefill | Decode | |
|---|---|---|
| Tokens per pass | The whole prompt | Exactly one |
| Parallelism | Across all prompt positions | None across time |
| Bottleneck | GPU compute | Memory bandwidth |
| Metric it governs | Time to first token | Tokens per second |
The useful quantity is arithmetic intensity: how many floating-point operations a kernel performs per byte it has to read from memory.
A decode step with batch size over a model with parameters does roughly FLOPs — two operations (a multiply and an add) per parameter per sequence. But it must read every weight once regardless of , so at bytes per weight it moves about bytes. The intensity is therefore:
The model size cancels. With 2-byte weights, intensity is just — around 1 FLOP per byte for a single request. Current accelerators need several hundred FLOPs per byte before compute becomes the limiting factor, so decode at small batch sizes leaves the arithmetic units almost entirely idle while the memory system runs flat out.
Prefill is the opposite case. A prompt of tokens reuses each weight times in the same pass, giving intensity — comfortably compute-bound for any prompt of realistic length.
Two consequences follow directly. Raising the batch size is the main lever for making decode efficient, because it raises intensity without reading any more weights. And speculative decodingSpeculative DecodingSpeculative decoding uses a small draft model to guess several tokens ahead, then verifies them all in one parallel pass with the full model, speeding up generation. (below) works precisely because that idle compute is free to spend on verifying guesses.
KV caching: don't recompute what hasn't changed
Recall from Attention & Transformers that every token produces a Key and Value vector used by attention. Naively, generating each new token would mean re-running the entire sequence so far back through every layer, recomputing Keys and Values for every previous token all over again — wasteful, since those earlier tokens haven't changed.
The KV cache fixes this: store each token's Key and Value vectors the first time they're computed, and reuse them for every subsequent generation step. Generating a new token then only requires computing that one new token's Query, Key, and Value, and attending against the cached Keys/Values of everything before it — turning an operation whose cost grows with the square of the sequence length (recompute everything, every step) into one that grows linearly with it.
The tradeoff: the KV cache consumes GPU memory proportional to sequence length and batch size, and at long context lengths or high concurrent request counts, KV cache memory — not raw compute — is often the binding constraint on how many requests a server can handle at once.
Every layer stores one Key and one Value vector per attention head per token. For a model with layers, key/value heads, head dimension , and bytes per number, a single sequence of length in a batch of occupies:
The leading is the Key and the Value. Note what is absent: the feed-forward weights, which dominate parameter count, contribute nothing. Cache size is driven by depth and attention width, not by total parameters.
Everything except and is fixed once the model is trained, so the formula is really "a constant number of bytes per token." That constant is exactly what grouped-query attention attacks: sharing one Key/Value head across several Query heads shrinks — and therefore the cache — by the sharing ratio, with no change to or . It is an inference-time memory optimization baked into the architecture at training time.
Take a large model with 80 layers, 8 key/value heads, head dimension 128, running in 16-bit precision. Per token:
That sounds negligible. Now fill a 32,000-token context:
- One request: 10.7 GB
- Eight concurrent requests: ~86 GB
Eight users have exhausted an 80 GB accelerator on cache alone — before loading a single model weight. This is why long context is expensive to serve, not just to train, and why serving systems page, evict, and compress KV cache as aggressively as operating systems manage RAM.
Generate some tokens both ways below and watch the gap widen — these are simulated costs illustrating the shape of the difference, not measured hardware timing, but the quadratic-vs-linear growth is exactly the real effect:
Without KV cache (recompute every step)0 ms
With KV cache0 ms
Simulated costs, not measured hardware timing — the point is the shape: without caching, each step redoes work proportional to everything generated so far (quadratic total cost); with caching, each step is roughly constant work (linear total cost). The speedup grows the longer the generation runs.
Batching: trading latency for throughput
GPUs are dramatically more efficient processing many requests simultaneously (as one batched matrix operation) than one at a time — most of the cost of a GPU operation is fixed overhead, so batchingBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish. amortizes it across many requests. A server therefore usually batches multiple users' requests together into one forward pass rather than running each in full isolation.
This creates a real tradeoff:
- Larger batches → higher throughput (more total tokens generated per second across all users) → but individual requests may wait longer before being included in a batch, increasing latency (time to first token / time per token, for one specific user).
- Smaller batches (or no batching) → lower latency per request, but the GPU spends more time on fixed overhead relative to useful work, lowering total throughput.
Production LLM servers use techniques like continuous batching (dynamically adding new requests into an in-flight batch as earlier ones finish, rather than waiting for a whole batch to complete together) to get much of the throughput benefit of large batches without forcing every request to wait for the slowest one in a fixed batch.
Under static batching, a batch of requests is admitted together and returns together. Every slot stays occupied until the longest generation finishes, so if request needs tokens, the useful work is while the batch is held for token-slots. Utilisation is therefore:
Output lengths in real traffic vary by orders of magnitude — a one-word answer shares a batch with a long essay — so and most slots spend most of their time computing padding.
Continuous batching schedules at the granularity of a single decode step instead of a whole request. When a sequence emits its stop token, its slot is returned and a queued request takes it on the very next step. approaches 1 regardless of how uneven the length distribution is, and queueing delay drops from "wait for the current batch to finish" to "wait for any one slot."
The cost is bookkeeping: slots now hold caches of differing lengths, so the attention kernel must handle ragged sequences and the memory allocator must hand out cache space in small blocks rather than one contiguous slab per request.
Drag the batch size below and watch that split play out: the step time stays flat, then breaks upward once the batch is large enough to be compute-bound — while latency climbs the whole way, mostly from queueing:
Time per step
20 ms
Latency (queue + step)
26 ms
Throughput (all users)
200 tok/s
Still bandwidth-bound: the step barely slows down as batch size grows — almost all of the added latency is queueing, not compute.
Quantization: trading precision for speed and memory
Model weights are normally stored as 16-bit or 32-bit floating point numbers. QuantizationQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy. reduces that precision — commonly to 8-bit or even 4-bit integers — which shrinks the model's memory footprint (fewer bits per weight) and can speed up computation (many GPUs execute low-bit integer math faster than full floating point).
The obvious question is why this doesn't just break the model: neural networks turn out to be fairly robust to reduced numerical precision — small rounding errors on individual weights tend not to meaningfully change the model's overall output, especially with quantization schemes designed to preserve the range and distribution of the original weights rather than naively rounding. There's a real accuracy cost as bit width drops (4-bit is noticeably riskier than 8-bit), which is why quantization level is itself a tuned tradeoff — lower precision saves memory and cost, but can measurably degrade output quality if pushed too far for a given model and use case.
Speculative decoding: guessing ahead to go faster
Because decode is memory-bandwidth-bound rather than compute-bound (see above), the GPU often has spare compute capacity during generation even though it's the bottleneck on speed. Speculative decoding exploits this: use a small, fast "draft" model to quickly generate several candidate next tokens (say, 4–5) ahead of time, then have the full model verify all of them in a single parallel forward pass — which the GPU has the spare compute for, since checking several tokens at once costs barely more than checking one, thanks to that same prefill-style parallelism.
If the draft model's guesses match what the full model would have generated (checked exactly, not approximately — the output is mathematically identical to normal decoding), all of them are accepted at once, several tokens for roughly the cost of one decode step; wherever a guess is wrong, generation falls back to the full model's actual choice from that point on. The speedup depends entirely on how often the small model's guesses match the large model's — most effective when the two models are closely related (e.g. distilled from each other) or the text is fairly predictable.
Serving models too large for one GPU
Some models don't fit in a single GPU's memory even quantized, requiring the same kind of model parallelism used in large-scale training (see Tooling & The Dev Stack) — most commonly tensor parallelismTensor ParallelismTensor parallelism splits individual weight matrices across multiple GPUs so a model too large for one GPU's memory can still be trained or served., which splits individual weight matrices across multiple GPUs, each computing a portion of every matrix multiplication and exchanging partial results. This adds inter-GPU communication overhead on the critical path of every single forward pass, which is why serving very large models efficiently also depends heavily on fast GPU-to-GPU interconnects, not just raw GPU count.
Putting it together: the tradeoffs that define an LLM API
When you send a request to an LLM API, several of these mechanisms are working simultaneously: your request is likely batched with others (continuous batching), each generation step reuses a KV cache built up token by token, and the model itself may be running at reduced precision (quantized) to fit more capacity per GPU. The metrics that fall out of this stack, and that serving systems are explicitly tuned against:
- Time to first token (TTFT): latency before generation starts — dominated by how long a request waits to be batched and the cost of processing the input prompt.
- Tokens per second (throughput): how fast tokens stream out once generation starts — dominated by batch size, KV cache efficiency, and quantization.
- Cost per token: a direct function of how efficiently GPU compute and memory are being used — which is exactly what batching, KV caching, and quantization are all optimizing.
There is no free lunch here: better throughput/cost usually costs some latency or some accuracy, and serving systems are built around choosing a deliberate point on that tradeoff curve for a given product's needs (a chat assistant cares more about per-token latency; a batch summarization job cares more about total throughput and cost).
Recap and what's next
Inference is a different engineering problem from training: it's an inherently sequential loop, and the techniques covered here — KV caching, batching, quantization — exist specifically to make that loop fast and cheap enough to serve at scale. Now that a model is trained and served, the next question is how you actually know if it's any good — the next lesson covers evaluation and benchmarking, before the final lesson looks at what gets built on top of a served LLM: retrieval, tool use, and multi-step agentic systems.
Reinforcement Learning
MDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
Evaluation & Benchmarks
How models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance