Inference & Applied Systems

Vector Database

A vector database stores embedding vectors and answers nearest-neighbor queries fast, using approximate nearest neighbor (ANN) indexes like HNSW — the backbone of RAG retrieval.

A vector database is built to answer one kind of query fast: "which stored embedding vectors are closest to this query vector?" — a different problem from the exact-match and range queries a normal database indexes for. Closeness is measured with cosine similarity or dot product; since semantically similar text produces similar embeddings, nearest-neighbor search doubles as "find text with a similar meaning." Comparing a query against every stored vector doesn't scale, so vector databases use approximate nearest neighbor (ANN) indexes — most commonly HNSW — trading a little accuracy for large speedups. Common tools: Pinecone, Weaviate, Milvus, Qdrant, Chroma, and the pgvector extension for Postgres.

How it works

HNSW builds a layered proximity graph. The top layers are sparse and act as long-range highways; the bottom layer contains every vector with short-range links. A search enters at the top, greedily walks toward the query, and drops a layer when it can no longer improve — turning a linear scan into roughly logarithmic hops. Two knobs govern the trade: M, the number of links per node, fixes index size and build time; efSearch, the size of the candidate list kept during traversal, trades latency for recall at query time.

A usable database wraps that index with the payload the application needs: the original text a transformer embedded, plus metadata for filtering by tenant, date, or permission — which must be applied during traversal, not after, or the top k comes back empty.

When it breaks

  • Recall is invisible. ANN search returns something for every query, so a badly tuned efSearch silently drops the right answer. Measure recall against exact search on a sample.
  • Filtering fights the graph. A highly selective filter can disconnect the region the walk needs to traverse, so results degrade exactly where precision matters most.
  • Updates are awkward. HNSW handles inserts far better than deletes, which are usually tombstoned and only reclaimed by a rebuild, so churn-heavy corpora bloat and slow down.
  • Re-embedding is a migration. Changing the encoder model invalidates every stored vector; old and new embeddings are not comparable and mixing them corrupts results for the whole LLM pipeline downstream.

See also: Embedding, RAG

Learn more: RAG & Vector Databases

Mentioned in

Lessons where this comes up in context.

On this page