word2vec
word2vec (2013) was a technique for learning dense vector representations of words from raw text, a direct precursor to modern token embeddings.
word2vec (2013) was a technique for learning dense vector representations of words directly from raw text — words used in similar contexts end up with similar vectors, learned without manual labeling. It was a major step in bringing deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. into language processing, and a direct precursor to the token embeddingsEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. used throughout modern LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama..
How it works
word2vec trains a shallow two-matrix model over a sliding context
window. The skip-gram variant takes a center word and learns to
predict the words around it; CBOW reverses that, predicting the
center word from its context. Scoring the full vocabulary on every step
is far too expensive, so training uses negative sampling: raise the
dot product of a true (center, context) pair and lower it for a handful
of randomly drawn pairs. Very frequent words are subsampled so the
and of don't dominate the gradient. When training finishes the model
is discarded and its input matrix is kept as a lookup table — one dense
vector per word. Similarity is read off as cosine distance, and the
famous analogy arithmetic (king - man + woman landing near queen)
falls out of that geometry.
When it breaks
- One vector per word, regardless of context. bank gets a single representation averaging the river and financial senses — the limitation that contextual models built on transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. removed by producing a different vector per occurrence.
- Out-of-vocabulary words have no vector at all, a gap subword tokenizationTokenizationTokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE). later closed.
- Cosine proximity conflates relations. good and bad sit close together because they occur in near-identical contexts, which breaks naive sentiment applications.
- It still earns its place where cheap static vectors are enough — retrieval candidate generation, and embedding non-text items like products or songs from co-occurrence sequences.
See also: EmbeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space., LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.
Learn more: History & Landscape · LLMs · Paper: Efficient Estimation of Word Representations in Vector Space
Mentioned in
Lessons where this comes up in context.
Perceptron
The Perceptron (1958) was the first learning system built from an artificial neuron, and the direct ancestor of the neural network training loop.
AI (Artificial Intelligence)
AI is the field of building systems that perform tasks normally requiring human intelligence — reasoning, perception, language, and decision-making.