Systems

Inference & Serving

Batching, KV caching, quantization, latency/throughput tradeoffs

Everything so far has been about training — producing a model's weights. Inference is using a trained model to generate output, and it turns out to be its own deep engineering problem: the naive way to run a transformer is far too slow and expensive to serve to real users at scale. This lesson covers the main techniques that make LLM serving practical.

Autoregressive generation is inherently sequential

An LLM generates text one token at a time: predict a distribution over the next token, sample one, append it to the sequence, and repeat — using its own previous output as input to produce the next token. This is called autoregressive generation, and it has an unavoidable consequence: you cannot generate token 50 before you've generated token 49, because token 50's prediction depends on it. Unlike training (where the whole target sequence is known upfront and attention can be computed in parallel across all positions), inference is fundamentally a sequential loop.

This sequential nature is the root cause of most of the engineering problems in this lesson.

Prefill vs. decode: two different phases

Serving a request actually involves two phases with very different performance characteristics:

  • Prefill: process the entire input prompt at once. Every prompt token's Key and Value can be computed in parallel — no sequential dependency yet, since none of these tokens depend on a model-generated token. This phase is compute-bound: it's a large, parallel matrix operation, so it's mostly limited by raw GPU compute throughput.
  • Decode: generate the output, one token at a time, autoregressively. Each step does comparatively little computation (one new token) but must read the entire growing KV cache from GPU memory at every step. This phase is memory-bandwidth-bound: the bottleneck is how fast data can move, not how much math the GPU can do.

Put end to end, a single request moves through a fixed pipeline — one prefill pass, then a loop that runs once per generated token:

Why is generating a response slower than reading the prompt?

This distinction is why time to first token and tokens per second (below) behave so differently and are optimized somewhat separately: a long prompt mostly costs prefill time; a long generated response mostly costs decode time, dominated by repeatedly reading an ever-growing KV cache.

PrefillDecode
Tokens per passThe whole promptExactly one
ParallelismAcross all prompt positionsNone across time
BottleneckGPU computeMemory bandwidth
Metric it governsTime to first tokenTokens per second

KV caching: don't recompute what hasn't changed

Recall from Attention & Transformers that every token produces a Key and Value vector used by attention. Naively, generating each new token would mean re-running the entire sequence so far back through every layer, recomputing Keys and Values for every previous token all over again — wasteful, since those earlier tokens haven't changed.

The KV cache fixes this: store each token's Key and Value vectors the first time they're computed, and reuse them for every subsequent generation step. Generating a new token then only requires computing that one new token's Query, Key, and Value, and attending against the cached Keys/Values of everything before it — turning an operation whose cost grows with the square of the sequence length (recompute everything, every step) into one that grows linearly with it.

The tradeoff: the KV cache consumes GPU memory proportional to sequence length and batch size, and at long context lengths or high concurrent request counts, KV cache memory — not raw compute — is often the binding constraint on how many requests a server can handle at once.

When the cache outgrows the GPU

Take a large model with 80 layers, 8 key/value heads, head dimension 128, running in 16-bit precision. Per token:

2×80×8×128×2 B≈0.33 MB2 \times 80 \times 8 \times 128 \times 2\,\text{B} \approx 0.33\ \text{MB}

That sounds negligible. Now fill a 32,000-token context:

  • One request: 0.33 MB×32,000≈0.33\,\text{MB} \times 32{,}000 \approx 10.7 GB
  • Eight concurrent requests: ~86 GB

Eight users have exhausted an 80 GB accelerator on cache alone — before loading a single model weight. This is why long context is expensive to serve, not just to train, and why serving systems page, evict, and compress KV cache as aggressively as operating systems manage RAM.

Generate some tokens both ways below and watch the gap widen — these are simulated costs illustrating the shape of the difference, not measured hardware timing, but the quadratic-vs-linear growth is exactly the real effect:

Without KV cache (recompute every step)0 ms

With KV cache0 ms

0 / 30 tokens

Simulated costs, not measured hardware timing — the point is the shape: without caching, each step redoes work proportional to everything generated so far (quadratic total cost); with caching, each step is roughly constant work (linear total cost). The speedup grows the longer the generation runs.

Batching: trading latency for throughput

GPUs are dramatically more efficient processing many requests simultaneously (as one batched matrix operation) than one at a time — most of the cost of a GPU operation is fixed overhead, so batching amortizes it across many requests. A server therefore usually batches multiple users' requests together into one forward pass rather than running each in full isolation.

This creates a real tradeoff:

  • Larger batches → higher throughput (more total tokens generated per second across all users) → but individual requests may wait longer before being included in a batch, increasing latency (time to first token / time per token, for one specific user).
  • Smaller batches (or no batching) → lower latency per request, but the GPU spends more time on fixed overhead relative to useful work, lowering total throughput.

Production LLM servers use techniques like continuous batching (dynamically adding new requests into an in-flight batch as earlier ones finish, rather than waiting for a whole batch to complete together) to get much of the throughput benefit of large batches without forcing every request to wait for the slowest one in a fixed batch.

Drag the batch size below and watch that split play out: the step time stays flat, then breaks upward once the batch is large enough to be compute-bound — while latency climbs the whole way, mostly from queueing:

Time per step

20 ms

Latency (queue + step)

26 ms

Throughput (all users)

200 tok/s

Still bandwidth-bound: the step barely slows down as batch size grows — almost all of the added latency is queueing, not compute.

Quantization: trading precision for speed and memory

Model weights are normally stored as 16-bit or 32-bit floating point numbers. Quantization reduces that precision — commonly to 8-bit or even 4-bit integers — which shrinks the model's memory footprint (fewer bits per weight) and can speed up computation (many GPUs execute low-bit integer math faster than full floating point).

The obvious question is why this doesn't just break the model: neural networks turn out to be fairly robust to reduced numerical precision — small rounding errors on individual weights tend not to meaningfully change the model's overall output, especially with quantization schemes designed to preserve the range and distribution of the original weights rather than naively rounding. There's a real accuracy cost as bit width drops (4-bit is noticeably riskier than 8-bit), which is why quantization level is itself a tuned tradeoff — lower precision saves memory and cost, but can measurably degrade output quality if pushed too far for a given model and use case.

Speculative decoding: guessing ahead to go faster

Because decode is memory-bandwidth-bound rather than compute-bound (see above), the GPU often has spare compute capacity during generation even though it's the bottleneck on speed. Speculative decoding exploits this: use a small, fast "draft" model to quickly generate several candidate next tokens (say, 4–5) ahead of time, then have the full model verify all of them in a single parallel forward pass — which the GPU has the spare compute for, since checking several tokens at once costs barely more than checking one, thanks to that same prefill-style parallelism.

If the draft model's guesses match what the full model would have generated (checked exactly, not approximately — the output is mathematically identical to normal decoding), all of them are accepted at once, several tokens for roughly the cost of one decode step; wherever a guess is wrong, generation falls back to the full model's actual choice from that point on. The speedup depends entirely on how often the small model's guesses match the large model's — most effective when the two models are closely related (e.g. distilled from each other) or the text is fairly predictable.

Serving models too large for one GPU

Some models don't fit in a single GPU's memory even quantized, requiring the same kind of model parallelism used in large-scale training (see Tooling & The Dev Stack) — most commonly tensor parallelism, which splits individual weight matrices across multiple GPUs, each computing a portion of every matrix multiplication and exchanging partial results. This adds inter-GPU communication overhead on the critical path of every single forward pass, which is why serving very large models efficiently also depends heavily on fast GPU-to-GPU interconnects, not just raw GPU count.

Putting it together: the tradeoffs that define an LLM API

When you send a request to an LLM API, several of these mechanisms are working simultaneously: your request is likely batched with others (continuous batching), each generation step reuses a KV cache built up token by token, and the model itself may be running at reduced precision (quantized) to fit more capacity per GPU. The metrics that fall out of this stack, and that serving systems are explicitly tuned against:

  • Time to first token (TTFT): latency before generation starts — dominated by how long a request waits to be batched and the cost of processing the input prompt.
  • Tokens per second (throughput): how fast tokens stream out once generation starts — dominated by batch size, KV cache efficiency, and quantization.
  • Cost per token: a direct function of how efficiently GPU compute and memory are being used — which is exactly what batching, KV caching, and quantization are all optimizing.

There is no free lunch here: better throughput/cost usually costs some latency or some accuracy, and serving systems are built around choosing a deliberate point on that tradeoff curve for a given product's needs (a chat assistant cares more about per-token latency; a batch summarization job cares more about total throughput and cost).

Recap and what's next

Inference is a different engineering problem from training: it's an inherently sequential loop, and the techniques covered here — KV caching, batching, quantization — exist specifically to make that loop fast and cheap enough to serve at scale. Now that a model is trained and served, the next question is how you actually know if it's any good — the next lesson covers evaluation and benchmarking, before the final lesson looks at what gets built on top of a served LLM: retrieval, tool use, and multi-step agentic systems.

On this page