Embedding
An embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space.
An embedding is a learned vector of numbers (typically hundreds to thousands of dimensions) representing a token, word, or passage of text in a way a model can compute with. Items used in similar contexts during training end up with similar embedding vectors — learned entirely from data. Token embeddings are the raw input to a transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.; document/passage embeddings power similarity search in RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining. systems via vector databasesVector DatabaseA vector database stores embedding vectors and answers nearest-neighbor queries fast, using approximate nearest neighbor (ANN) indexes like HNSW — the backbone of RAG retrieval..
How it works
An embedding table is a learned matrix of shape
(vocab_size, d_model); looking up token id i returns row i. The
rows start random and are updated by the same gradients that train the
rest of the network, so the geometry emerges from the training
objective rather than being designed. Static methods such as
word2vecword2vecword2vec (2013) was a technique for learning dense vector representations of words from raw text, a direct precursor to modern token embeddings. give one fixed vector per word; inside
a transformer the vector at each layer is contextual, because
attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. mixes in surrounding tokens, so the
same word gets different representations in different sentences.
Passage embeddings for search are produced by pooling a model's final hidden states into a single vector, then comparing vectors by cosine similarity or dot product. Such models are typically trained with a contrastive objective on paired text.
When it breaks
- Dimensions mean nothing individually. Only relative distances carry signal; inspecting component 47 is not interpretation.
- Similarity is not relevance. Contrastive training optimizes for the pairs it saw, so negations and antonyms often land close together — "flight to Paris" and "flight from Paris" are nearly identical vectors.
- Normalization mismatches. Cosine similarity assumes unit-norm vectors; mixing normalized and unnormalized vectors, or using different pooling at index time and query time, quietly destroys recall.
- Version lock-in. Vectors from two different models are not comparable, so changing the embedding model means reindexing the entire corpus.
See also: TokenizationTokenizationTokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE)., RAGRAG (Retrieval-Augmented Generation)RAG grounds an LLM's answers in retrieved documents at request time, letting it answer questions about private or current data without retraining., Vector DatabaseVector DatabaseA vector database stores embedding vectors and answers nearest-neighbor queries fast, using approximate nearest neighbor (ANN) indexes like HNSW — the backbone of RAG retrieval.
Learn more: LLMs · RAG & Vector Databases · Wikipedia: Word embedding
Mentioned in
Lessons where this comes up in context.
- Attention & TransformersThe core architecture behind modern AI
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- Multimodal ModelsHow a single model handles more than one kind of input — CLIP's shared embedding space, how vision gets fed into a language model as tokens, and what actually breaks
- RAG & Vector DatabasesRetrieval-augmented generation grounds an LLM's answers in retrieved text at request time — and the vector databases and similarity search that make it work
Tokenization
Tokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE).
Pretraining
Pretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability.