Inference & Applied Systems

Inference

Inference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering.

Inference is using a trained model to produce output, as distinct from training (producing the weights in the first place). For LLMs, inference is autoregressive: generate one token at a time, feeding each output back in as input for the next step — which is inherently sequential and drives most of the engineering in this area, including KV caching, quantization, and request batching.

How it works

Generation splits into two phases with opposite performance profiles. Prefill runs the whole prompt through the model in one pass; every position is processed in parallel, so it is compute-bound and its cost scales with prompt length. Decode then emits one token per forward pass, each pass reading the entire weight set to produce a single token, which makes it memory-bandwidth-bound. Between the two, the KV cache carries the per-token Key/Value vectors forward so decode never re-processes earlier positions.

That split explains the standard metrics: time-to-first-token measures prefill, inter-token latency measures decode, and throughput is measured in tokens per second across a batch. It also explains why batching helps decode so much and prefill so little.

When it breaks

  • Decode cannot be parallelised away. Token n+1 depends on token n, so adding GPUs raises throughput but not single-stream speed. Only tricks like speculative decoding shorten the chain.
  • Memory, not FLOPs, is usually the wall. Weights plus KV cache must fit; exceeding capacity causes eviction, recompute, or outright request failure rather than graceful slowdown.
  • Benchmarks mislead. Single-request latency measured on an idle server has little to do with latency at production concurrency, where queueing and batch scheduling dominate.
  • Numerics drift. Batch size, kernel choice, and GPU model change floating-point reduction order, so identical inputs and a fixed seed can still produce different outputs.

See also: KV Cache, Quantization

Learn more: Inference & Serving · Wikipedia: Inference engine

Mentioned in

Lessons where this comes up in context.

On this page