Transformers & LLMs

Embedding

An embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space.

An embedding is a learned vector of numbers (typically hundreds to thousands of dimensions) representing a token, word, or passage of text in a way a model can compute with. Items used in similar contexts during training end up with similar embedding vectors — learned entirely from data. Token embeddings are the raw input to a transformer; document/passage embeddings power similarity search in RAG systems via vector databases.

How it works

An embedding table is a learned matrix of shape (vocab_size, d_model); looking up token id i returns row i. The rows start random and are updated by the same gradients that train the rest of the network, so the geometry emerges from the training objective rather than being designed. Static methods such as word2vec give one fixed vector per word; inside a transformer the vector at each layer is contextual, because attention mixes in surrounding tokens, so the same word gets different representations in different sentences.

Passage embeddings for search are produced by pooling a model's final hidden states into a single vector, then comparing vectors by cosine similarity or dot product. Such models are typically trained with a contrastive objective on paired text.

When it breaks

  • Dimensions mean nothing individually. Only relative distances carry signal; inspecting component 47 is not interpretation.
  • Similarity is not relevance. Contrastive training optimizes for the pairs it saw, so negations and antonyms often land close together — "flight to Paris" and "flight from Paris" are nearly identical vectors.
  • Normalization mismatches. Cosine similarity assumes unit-norm vectors; mixing normalized and unnormalized vectors, or using different pooling at index time and query time, quietly destroys recall.
  • Version lock-in. Vectors from two different models are not comparable, so changing the embedding model means reindexing the entire corpus.

See also: Tokenization, RAG, Vector Database

Learn more: LLMs · RAG & Vector Databases · Wikipedia: Word embedding

Mentioned in

Lessons where this comes up in context.

On this page