Applied

RAG & Vector Databases

Retrieval-augmented generation grounds an LLM's answers in retrieved text at request time — and the vector databases and similarity search that make it work

Prompt Engineering covered shaping behavior for free, but it can't fix the deeper problem: a deployed LLM's knowledge is frozen at training time, and it knows nothing about private data it was never trained on. RAG (retrieval-augmented generation) addresses that by giving the model relevant information at request time, rather than trying to bake everything into its weights via training.

The RAG pattern

Index: split your documents (docs, PDFs, support tickets, whatever the knowledge source is) into chunks, and convert each chunk into an embedding vector — conceptually the same kind of embedding from LLMs, but here used to represent a whole passage of text rather than a single token. Store these in a vector database.

Retrieve: when a user asks a question, embed the question the same way, and find the stored chunks whose embeddings are most similar (nearest neighbors in vector space) — the intuition being that semantically related text ends up with similar embedding vectors, so this retrieves passages relevant to the question even if they don't share exact keywords with it.

Generate: insert the retrieved chunks into the prompt, along with the user's question, and let the LLM generate an answer grounded in that retrieved text.

RAG's appeal: it lets a model answer questions about private, current, or otherwise-untrained-on data without any retraining — you update the answers by updating the indexed documents, not by fine-tuning the model. Its main failure modes are retrieval quality (if the wrong chunks are retrieved, the model confidently answers from irrelevant context) and context limits (a model can only be given so much retrieved text before a prompt gets too long or too noisy to reason over well).

A vector database is a database built specifically to store embedding vectors and answer one kind of query fast: "which stored vectors are closest to this query vector?" That's a different problem from what a normal database indexes for — exact matches and ranges on structured fields — which is why RAG systems reach for a purpose-built one rather than bolting vector search onto a regular table.

"Closest" is measured with a similarity metric, most commonly cosine similarity (the angle between two vectors, ignoring their magnitude) or plain dot product. Two embeddings from LLMs that represent semantically similar text end up pointing in similar directions in that high-dimensional space, which is exactly what makes nearest-neighbor search double as "find text with a similar meaning."

Click any point below to make it the query and see its real cosine-similarity nearest neighbors. The positions are a hand-placed 2D illustration, not output from an actual embedding model — real embeddings have hundreds to thousands of dimensions — but the similarity scores themselves are computed live, the same formula shown above. One pair is a deliberate trap: dense embeddings compress out negation and direction far more than they compress out topic — a named blind spot, not a quirk of this toy example — so try querying it to see the effect directly.

elephantcatdoglionpythonjavascriptrustapplebananamangohappysadangryto Parisfrom Paris
AnimalsCodeFruitEmotionDirection

Click any word to make it the query and rank every other word by cosine similarity to it — including "to Paris" and "from Paris", which aren't as different as they sound.

Naively, finding the true nearest neighbors means comparing the query against every stored vector — fine for thousands of chunks, far too slow once you're indexing millions. Vector databases solve this with approximate nearest neighbor (ANN) indexes — HNSW (Hierarchical Navigable Small World) is the most common — that organize vectors into a searchable graph structure ahead of time, trading a small amount of recall accuracy for orders-of-magnitude faster lookups. This tradeoff (exact vs. approximate, accuracy vs. speed) is a recurring theme in systems that need to operate at scale, the same way quantization trades numerical precision for speed at inference time.

In practice, "vector database" covers a range of tools: dedicated managed services (Pinecone), open-source databases built vector-first (Weaviate, Milvus, Qdrant), a lightweight option for local/small-scale use (Chroma), and vector-search extensions added to databases you might already run (pgvector for Postgres). Which one to reach for mostly comes down to scale, whether you want a separate system to operate versus extending an existing database, and how much of the ANN tuning you want to manage yourself.

ToolTypeBest for
PineconeManaged serviceProduction, without operating your own infrastructure
Weaviate / Milvus / QdrantOpen-source, vector-firstSelf-hosting at scale with full control
ChromaLightweight, embeddedLocal development, small-scale projects
pgvectorExtension for an existing databaseAdding vector search to a Postgres database you already run

Recap and what's next

RAG grounds generation in retrieved text, and vector databases are the infrastructure that makes retrieval fast at scale — but a model that can only read still can't act. The next lesson, Agents & Tool Use, covers giving a model the ability to call functions, take actions, and chain multiple steps together — including using retrieval itself as one tool among several.

On this page