Hallucination
A hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts.
A hallucination is a confidently stated, fluent LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. output that's factually wrong. It follows directly from what the model is: a next-token predictor trained to produce plausible-sounding text, not a database lookup with a built-in sense of "I don't know." When the model lacks reliable knowledge, the training objective still rewards some fluent continuation. RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining. reduces hallucination by grounding generation in retrieved, verifiable text instead of relying purely on what the model absorbed during pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability..
How it works
Nothing in the training objective separates "true" from "likely." Next-token prediction rewards the most plausible continuation given the context, and a fabricated citation in the right format is highly plausible. Several mechanisms compound this:
- No abstention signal. Base models are never trained to say "I don't know" about facts they lack, so the distribution puts little mass on refusal.
- Alignment pressure. RLHFRLHF (Reinforcement Learning from Human Feedback)RLHF trains an LLM to match human preferences by learning a reward model from ranked response comparisons, then optimizing the LLM against that reward. raters tend to prefer confident, complete answers, which pushes the model away from hedging.
- Sampling. Anything with nonzero probability can be emitted, and higher-temperature decodingDecoding Strategies (Temperature, Top-k, Top-p)Decoding strategies turn an LLM's next-token probability distribution into an actual chosen token, using temperature, top-k, or top-p (nucleus) sampling. reaches further into the tail.
- Lossy storage. Facts seen once during training are compressed into weights and recalled as something that merely rhymes with the original.
When it breaks
- Fluency masks it. Wrong output is stylistically identical to right output, so reviewers skim past it; token probabilities are only a weak correlate of correctness.
- Retrieval doesn't guarantee grounding. If retrieved passages are irrelevant or contradictory, the model still answers — now with a citation that doesn't support the claim.
- Self-checking is circular. Asking the same model whether it was correct produces another confident answer generated the same way.
- Compounding in loops. An agentAgentAn agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal. that acts on a fabricated API name or file path carries the error into every subsequent step.
See also: LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama., RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining.
Learn more: LLMs · RAG & Vector Databases
Mentioned in
Lessons where this comes up in context.
- Evaluation & BenchmarksHow models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- Multimodal ModelsHow a single model handles more than one kind of input — CLIP's shared embedding space, how vision gets fed into a language model as tokens, and what actually breaks
Decoding Strategies (Temperature, Top-k, Top-p)
Decoding strategies turn an LLM's next-token probability distribution into an actual chosen token, using temperature, top-k, or top-p (nucleus) sampling.
MDP (Markov Decision Process)
An MDP formalizes the reinforcement-learning setup — an agent observing a state, choosing an action, and receiving a reward, with no dataset of correct actions to imitate.