Vector Database
A vector database stores embedding vectors and answers nearest-neighbor queries fast, using approximate nearest neighbor (ANN) indexes like HNSW — the backbone of RAG retrieval.
A vector database is built to answer one kind of query fast: "which stored embeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. vectors are closest to this query vector?" — a different problem from the exact-match and range queries a normal database indexes for. Closeness is measured with cosine similarity or dot product; since semantically similar text produces similar embeddings, nearest-neighbor search doubles as "find text with a similar meaning." Comparing a query against every stored vector doesn't scale, so vector databases use approximate nearest neighbor (ANN) indexes — most commonly HNSW — trading a little accuracy for large speedups. Common tools: Pinecone, Weaviate, Milvus, Qdrant, Chroma, and the pgvector extension for Postgres.
How it works
HNSW builds a layered proximity graph. The top layers are sparse and
act as long-range highways; the bottom layer contains every vector with
short-range links. A search enters at the top, greedily walks toward the
query, and drops a layer when it can no longer improve — turning a
linear scan into roughly logarithmic hops. Two knobs govern the trade:
M, the number of links per node, fixes index size and build time;
efSearch, the size of the candidate list kept during traversal, trades
latency for recall at query time.
A usable database wraps that index with the payload the application
needs: the original text a transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.
embedded, plus metadata for filtering by tenant, date, or permission —
which must be applied during traversal, not after, or the top k comes
back empty.
When it breaks
- Recall is invisible. ANN search returns something for every
query, so a badly tuned
efSearchsilently drops the right answer. Measure recall against exact search on a sample. - Filtering fights the graph. A highly selective filter can disconnect the region the walk needs to traverse, so results degrade exactly where precision matters most.
- Updates are awkward. HNSW handles inserts far better than deletes, which are usually tombstoned and only reclaimed by a rebuild, so churn-heavy corpora bloat and slow down.
- Re-embedding is a migration. Changing the encoder model invalidates every stored vector; old and new embeddings are not comparable and mixing them corrupts results for the whole LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. pipeline downstream.
See also: EmbeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space., RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining.
Learn more: RAG & Vector Databases
Mentioned in
Lessons where this comes up in context.
- Agents & Tool UseGiving an LLM the ability to take actions and chain multiple steps together — tool use, MCP, and the agent loop
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- Prompt EngineeringShaping an LLM's output by changing the input alone — few-shot examples, chain-of-thought, and system prompts — with no training involved
- RAG & Vector DatabasesRetrieval-augmented generation grounds an LLM's answers in retrieved text at request time — and the vector databases and similarity search that make it work
RAG (Retrieval-Augmented Generation)
RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining.
Prompt Engineering
Prompt engineering shapes an LLM's output by changing the input phrasing alone — via few-shot examples, chain-of-thought, or system prompts — with no training involved.