RAG (Retrieval-Augmented Generation)
RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining.
RAG (Retrieval-Augmented Generation) gives an LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. relevant information at request time instead of relying only on what it learned during pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability.. Documents are split into chunks, converted to embeddingsEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space., and stored in a vector databaseVector DatabaseA vector database stores embedding vectors and answers nearest-neighbor queries fast, using approximate nearest neighbor (ANN) indexes like HNSW — the backbone of RAG retrieval.; at query time, the most similar chunks are retrieved and inserted into the prompt so the model can generate an answer grounded in that text. This lets a model answer questions about private or current data — and reduces hallucinationHallucinationA hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts. — without any retraining.
How it works
RAG has two pipelines. The indexing pipeline runs offline: parse
documents, split them into chunks sized to survive retrieval intact,
embed each chunk with an encoder model, and write the vectors plus the
original text and metadata to the index. The query pipeline runs per
request: embed the user's question with the same encoder, retrieve the
top k nearest chunks, and paste them into the prompt above the
question with an instruction to answer only from the provided context.
Production systems usually add two stages. Hybrid search combines dense vector similarity with keyword scoring such as BM25, since embeddings are weak at exact identifiers. Reranking then scores a wider candidate set with a cross-encoder and keeps only the best few, which costs more per query but sharply improves what reaches the model.
When it breaks
- Chunking decides the ceiling. Split too small and a fact loses the context that makes it interpretable; too large and the real answer is diluted by irrelevant text. No prompt fixes a bad split.
- Retrieval failures are silent. If the right chunk is not in the
top
k, the model answers confidently from parametric memory anyway — grounding reduces hallucination, it does not prevent it. - Embedding mismatch. Questions and documents are worded differently, so semantic similarity misses paraphrases and, worse, exact codes and part numbers.
- Stale or conflicting sources. The index happily returns two contradictory documents, and the model picks one with no signal about which is current.
- Retrieved text is untrusted input. A poisoned document can carry instructions the model follows; see prompt injectionPrompt InjectionPrompt injection is an attack where text an LLM processes — user input, a retrieved document, or a tool's output — contains instructions that override the application's intended behavior..
See also: EmbeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space., Vector DatabaseVector DatabaseA vector database stores embedding vectors and answers nearest-neighbor queries fast, using approximate nearest neighbor (ANN) indexes like HNSW — the backbone of RAG retrieval., AgentAgentAn agent puts an LLM in a loop with tool access, letting it decide autonomously which tools to call and in what order to accomplish a multi-step goal., HallucinationHallucinationA hallucination is a confidently stated, fluent LLM output that is factually wrong, a direct consequence of models being trained to produce plausible text rather than verified facts., Prompt InjectionPrompt InjectionPrompt injection is an attack where text an LLM processes — user input, a retrieved document, or a tool's output — contains instructions that override the application's intended behavior.
Learn more: RAG & Vector Databases
Mentioned in
Lessons where this comes up in context.
- Agents & Tool UseGiving an LLM the ability to take actions and chain multiple steps together — tool use, MCP, and the agent loop
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- Prompt EngineeringShaping an LLM's output by changing the input alone — few-shot examples, chain-of-thought, and system prompts — with no training involved
- RAG & Vector DatabasesRetrieval-augmented generation grounds an LLM's answers in retrieved text at request time — and the vector databases and similarity search that make it work
Batching (Continuous Batching)
Batching groups multiple requests into one GPU forward pass for efficiency; continuous batching dynamically adds new requests into an in-flight batch as others finish.
Vector Database
A vector database stores embedding vectors and answers nearest-neighbor queries fast, using approximate nearest neighbor (ANN) indexes like HNSW — the backbone of RAG retrieval.