Inference
Inference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering.
Inference is using a trained model to produce output, as distinct from training (producing the weights in the first place). For LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama., inference is autoregressive: generate one token at a time, feeding each output back in as input for the next step — which is inherently sequential and drives most of the engineering in this area, including KV cachingKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable., quantizationQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy., and request batching.
How it works
Generation splits into two phases with opposite performance profiles. Prefill runs the whole prompt through the model in one pass; every position is processed in parallel, so it is compute-bound and its cost scales with prompt length. Decode then emits one token per forward pass, each pass reading the entire weight set to produce a single token, which makes it memory-bandwidth-bound. Between the two, the KV cache carries the per-token Key/Value vectors forward so decode never re-processes earlier positions.
That split explains the standard metrics: time-to-first-token measures prefill, inter-token latency measures decode, and throughput is measured in tokens per second across a batch. It also explains why batchingBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish. helps decode so much and prefill so little.
When it breaks
- Decode cannot be parallelised away. Token n+1 depends on token n, so adding GPUs raises throughput but not single-stream speed. Only tricks like speculative decodingSpeculative DecodingSpeculative decoding uses a small draft model to guess several tokens ahead, then verifies them all in one parallel pass with the full model, speeding up generation. shorten the chain.
- Memory, not FLOPs, is usually the wall. Weights plus KV cache must fit; exceeding capacity causes eviction, recompute, or outright request failure rather than graceful slowdown.
- Benchmarks mislead. Single-request latency measured on an idle server has little to do with latency at production concurrency, where queueing and batch scheduling dominate.
- Numerics drift. Batch size, kernel choice, and GPU model change floating-point reduction order, so identical inputs and a fixed seed can still produce different outputs.
See also: KV CacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable., QuantizationQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy.
Learn more: Inference & Serving · Wikipedia: Inference engine
Mentioned in
Lessons where this comes up in context.
- Attention & TransformersThe core architecture behind modern AI
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- Inference & ServingBatching, KV caching, quantization, latency/throughput tradeoffs
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- RAG & Vector DatabasesRetrieval-augmented generation grounds an LLM's answers in retrieved text at request time — and the vector databases and similarity search that make it work
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
PPO (Proximal Policy Optimization)
PPO is a policy-gradient reinforcement-learning algorithm that takes conservative, clipped update steps, and is the algorithm most commonly used to optimize LLMs during RLHF.
Quantization
Quantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy.