Batching (Continuous Batching)
Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish.
Batching groups multiple users' requests into one forward pass, since GPUs are far more efficient processing many requests at once than one at a time — most of the cost of a GPU operation is fixed overhead that batching amortizes. Larger batches raise throughput but can raise latency per request. Continuous batching dynamically adds new requests into an in-flight batch as earlier ones finish, capturing most of the throughput benefit of large batches without forcing every request to wait for a whole batch to complete.
How it works
Decoding one token for one request multiplies a weight matrix by a
single row of activations. The GPU spends most of that
time reading weights from memory, not doing arithmetic, so stacking B
requests into a matrix of shape [B, hidden] reuses each weight read
B times and costs barely more wall-clock time. That is the whole
trick: decode is memory-bandwidth-bound, and batching converts it
toward compute-bound.
Static batching stalls because requests finish at different lengths — the batch waits for its slowest member. Continuous batching (also called in-flight batching) schedules per decode step instead: when a sequence emits its stop token, its slot is freed and a queued request takes it on the next step, with its own KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable. entry.
When it breaks
- KV cache, not compute, is the ceiling. Each concurrent sequence holds its own cache, so memory per request grows with context length. Long-context traffic collapses the feasible batch size.
- Prefill blocks decode. A newly admitted request's prompt pass is compute-heavy and stalls every decoding sequence in the batch, showing up as inter-token latency spikes for users already streaming. Chunked prefill exists to limit this.
- Throughput and latency trade off. Bigger batches raise tokens per second for the fleet while raising time-to-first-token for each user. Tuning to peak throughput will violate a p99 latency SLO.
- Preemption thrash. Under memory pressure servers evict and later recompute a sequence's cache, wasting the work already done.
See also: InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering., GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires.
Learn more: Inference & Serving
Mentioned in
Lessons where this comes up in context.
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- Inference & ServingBatching, KV caching, quantization, latency/throughput tradeoffs
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
Tensor Parallelism
Tensor parallelism splits individual weight matrices across multiple GPUs so a model too large for one GPU's memory can still be trained or served.
RAG (Retrieval-Augmented Generation)
RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining.