Quantization
Quantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy.
Quantization reduces a model's weight precision — commonly from 16/32-bit floating point down to 8-bit or 4-bit integers — shrinking memory footprint and often speeding up computation. Neural networks are fairly robust to this reduced precision, though accuracy risk increases as bit width drops. It's central to running LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. locally (tools like llama.cpp and Ollama) and to fitting larger models into limited GPU memory in production.
How it works
A group of floating-point weights is mapped onto a small integer grid by
storing a scale (and often a zero-point) alongside the integers, so the
original value is approximated as scale * (q - zero_point). Scales are
kept per channel or per small block rather than per tensor, because one
global scale lets a single outlier weight crush the resolution of
everything else.
Two families matter in practice. Post-training quantization converts an already-trained model, sometimes using a small calibration set to choose scales — this is what GPTQ and AWQ do. Quantization-aware training simulates the rounding during training so the weights adapt to it. For decode in a transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs., the win is mostly bandwidth: fewer bytes read per forward pass means faster generation, even when the math runs in higher precision.
When it breaks
- Activation outliers, not weights, cause the damage. Transformers produce a few extreme activation channels; naive low-bit activation quantization destroys them and tanks quality, which is why many schemes quantize weights only.
- Perplexity hides the loss. A 4-bit model often matches its
fp16parent on perplexity while getting visibly worse at long-context recall, code, and multi-step reasoning. Evaluate on the actual task. - Smaller does not always mean faster. If a kernel dequantizes back
to
fp16to do the matmul, you save memory but may gain little speed — and at small batchBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish. sizes the overhead can dominate. - Stacking with fine-tuning compounds error. Quantizing a model that was already adapted with low-rank adaptersLoRA (Low-Rank Adaptation)LoRA is a parameter-efficient fine-tuning method that trains a small number of additional low-rank parameters instead of updating an entire model's weights. can undo much of the adaptation.
See also: InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering., GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires.
Learn more: Inference & Serving · Tooling & The Dev Stack
Mentioned in
Lessons where this comes up in context.
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- Inference & ServingBatching, KV caching, quantization, latency/throughput tradeoffs
- RAG & Vector DatabasesRetrieval-augmented generation grounds an LLM's answers in retrieved text at request time — and the vector databases and similarity search that make it work
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
Inference
Inference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering.
Knowledge Distillation
Knowledge distillation trains a small "student" model to mimic a larger "teacher" model's output distribution, transferring most of its capability at a fraction of the serving cost.