Decoding Strategies (Temperature, Top-k, Top-p)
Decoding strategies turn an LLM's next-token probability distribution into an actual chosen token, using temperature, top-k, or top-p (nucleus) sampling.
Decoding turns an LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.'s output probability distribution over the vocabulary into an actual chosen token at each generation step. Greedy decoding always picks the highest-probability token (deterministic, often repetitive). Temperature rescales the distribution before sampling — lower is more focused, higher is more random. Top-k sampling restricts sampling to the k most likely tokens; top-p (nucleus) sampling keeps the smallest set of tokens whose cumulative probability exceeds p, adapting to model confidence at each step.
How it works
Each step produces a logit vector of size vocab_size. Temperature
divides the logits before the softmax: softmax(logits / T). T < 1
sharpens the distribution toward the argmax, T > 1 flattens it, and
T → 0 is greedy decoding. Top-k keeps the k highest logits and
renormalizes; top-p sorts descending, takes the shortest prefix whose
cumulative probability exceeds p, and renormalizes that. The filters
stack — most APIs expose temperature, top_k, and top_p together —
and repetition or frequency penalties subtract directly from the logits
of already-emitted tokens. Beam search is different in kind: it
keeps several partial sequences alive and ranks them by total
log-probability rather than sampling one token at a time.
When it breaks
- Degenerate repetition. Greedy and low-temperature decoding fall into loops, because a repeated phrase keeps raising its own probability. Top-p or a repetition penalty is the usual fix.
- Temperature is not confidence. Raising
Tadds noise, not knowledge; it widens the tail that produces hallucinationsHallucinationA hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts. rather than improving coverage of correct answers. - Beam search hurts open-ended text. Maximizing total log-probability favors short, bland continuations — useful for translation, bad for chat.
- "Deterministic" isn't. Greedy decoding still varies across batch sizes and hardware, because floating-point reduction order changes with batchingBatching (Continuous Batching)Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish. and kernel selection.
See also: LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama., InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering.
Learn more: LLMs
Mixture of Experts
A mixture-of-experts layer routes each input to only a handful of specialized subnetworks, growing a model's total parameter count without a proportional rise in compute per input.
Hallucination
A hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts.