Inference & Applied Systems

RAG (Retrieval-Augmented Generation)

RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining.

RAG (Retrieval-Augmented Generation) gives an LLM relevant information at request time instead of relying only on what it learned during pretraining. Documents are split into chunks, converted to embeddings, and stored in a vector database; at query time, the most similar chunks are retrieved and inserted into the prompt so the model can generate an answer grounded in that text. This lets a model answer questions about private or current data — and reduces hallucination — without any retraining.

How it works

RAG has two pipelines. The indexing pipeline runs offline: parse documents, split them into chunks sized to survive retrieval intact, embed each chunk with an encoder model, and write the vectors plus the original text and metadata to the index. The query pipeline runs per request: embed the user's question with the same encoder, retrieve the top k nearest chunks, and paste them into the prompt above the question with an instruction to answer only from the provided context.

Production systems usually add two stages. Hybrid search combines dense vector similarity with keyword scoring such as BM25, since embeddings are weak at exact identifiers. Reranking then scores a wider candidate set with a cross-encoder and keeps only the best few, which costs more per query but sharply improves what reaches the model.

When it breaks

  • Chunking decides the ceiling. Split too small and a fact loses the context that makes it interpretable; too large and the real answer is diluted by irrelevant text. No prompt fixes a bad split.
  • Retrieval failures are silent. If the right chunk is not in the top k, the model answers confidently from parametric memory anyway — grounding reduces hallucination, it does not prevent it.
  • Embedding mismatch. Questions and documents are worded differently, so semantic similarity misses paraphrases and, worse, exact codes and part numbers.
  • Stale or conflicting sources. The index happily returns two contradictory documents, and the model picks one with no signal about which is current.
  • Retrieved text is untrusted input. A poisoned document can carry instructions the model follows; see prompt injection.

See also: Embedding, Vector Database, Agent, Hallucination, Prompt Injection

Learn more: RAG & Vector Databases

Mentioned in

Lessons where this comes up in context.

On this page