Attention & Transformers
The core architecture behind modern AI
CNNsCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. (previous lesson) encode an assumption that works great for images: nearby things are related. Language breaks that assumption constantly — in "the trophy didn't fit in the suitcase because it was too big," what "it" refers to depends on a word far earlier in the sentence, and that distance varies sentence to sentence. Before 2017, sequence models (RNNsRNN (Recurrent Neural Network)An RNN is a neural network for sequences that processes input one step at a time, carrying a hidden state forward — the main predecessor to transformers., LSTMs) processed text one token at a time, left to right, which made long-range relationships hard to learn and impossible to parallelize during training. The transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. architecture solved both problems with one mechanism: self-attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer..
Here is the shape of the thing before we take it apart. A transformer is the same block repeated dozens of times, and one block looks like this:
The two arrows that skip past attention and past the MLP are the residual connectionsResidual ConnectionA residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable. — the block adds to its input rather than replacing it. Everything below fills in those boxes.
The core idea: attention as weighted lookup
Self-attention lets every position in a sequence directly look at every other position and decide how much to "pay attention" to it, regardless of distance. For the sentence above, when processing the word "it," attention lets the model directly weigh "trophy" and "suitcase" against each other and learn — from data, via training — that "it" should attend strongly to "trophy" in this context.
Mechanically, every token produces three vectors, each obtained by multiplying the token's embeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. by a learned weight matrix:
- Query (Q): "what am I looking for?" — from this token's perspective
- Key (K): "what do I contain?" — advertised by every other token
- Value (V): "what information do I actually offer if attended to?"
Attention for a given token compares its Query against every token's Key (via a dot product — a similarity score), turns those scores into weights that sum to 1 (via softmax), and produces an output as the weighted sum of every token's Value:
Attention(Q, K, V) = softmax(Q · Kᵀ / √d_k) · VThe √d_k term just rescales the dot products to keep training stable;
the important part is softmax(Q · Kᵀ): a matrix of "how much should
token i attend to token j" for every pair of tokens, learned entirely
from data, not hand-specified.
For a sequence of tokens, stack the per-token queries, keys and values into matrices and . Then:
The product is an matrix whose entry is the raw score . The softmax is applied per row, so each row becomes a probability distributionProbability DistributionA probability distribution assigns a likelihood to each possible value of a random variable — the object every loss function is secretly built to measure the fit of. over all :
and the output for token is — a weighted average of value vectors, with weights summing to 1.
The is not cosmetic. If the entries of and are roughly independent with unit variance, their dot product is a sum of such terms and so has variance — meaning typical scores grow like . With that is already an eightfold stretch, and softmax on widely separated inputs saturates: one weight goes to ~1, the rest to ~0, and the gradient through the softmax goes to nearly zero. Dividing by puts the scores back at roughly unit variance so the softmax stays in a region where it still has useful gradients.
Click through a few example sentences below to build intuition for what that attention matrix actually looks like. These weights are hand-set to illustrate the pattern a trained model learns, not live model output — but they show the same kind of long-range dependency real attention heads pick up on:
Click "it" — coreference resolution, exactly the lesson's own example.
Click any token above to see what it attends to.
Why "self"-attention, and why "multi-head"
It's called self-attention because Q, K, and V all come from the same sequence — every token attends to (potentially) every other token in the same input, including itself.
Why does attention need multiple heads?
In practice, transformers use multi-head attention: run several attention computations in parallel, each with its own learned Q/K/V projection matrices, then combine the results. Each "head" can specialize in a different kind of relationship — one head might learn to track subject-verb agreement, another might track coreference (like "it" → "trophy"), another might track adjacent-word syntax. No one designs these specializations; they emerge from training, similar to how CNN layers learn their own feature hierarchy.
Multi-head attention does not make the model wider. With model dimension and heads, each head works in a smaller subspace, typically . Head has its own learned projections :
The outputs — each — are concatenated back to width and passed through one more learned matrix :
Two things follow. First, the parameter count is essentially unchanged versus a single head of full width, so multiple heads are close to free — they buy specialization rather than capacity. Second, the reason you want them is that a single softmax row must spend its total weight of 1 on one blend of positions; it cannot simultaneously attend to the subject of the sentence and the previous token. Splitting into independent softmaxes lets the block attend to different things at once, and learns how to mix those results back together.
The rest of the transformer block
Self-attention alone is just weighted averaging — it needs to be paired with a few other components to form a working architecture:
- Positional information: attention itself has no notion of word order — swap two tokens and, by itself, attention treats it identically. Transformers inject order via positional encodingsPositional EncodingPositional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set. added to each token's embedding before attention, so the model can tell "dog bites man" from "man bites dog." The original transformer used a fixed sinusoidal encoding (a deterministic pattern of sine/cosine waves, one per position); many models instead learn positional embeddings directly from data. Most current LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. use a newer scheme called RoPERoPE (Rotary Positional Embedding)RoPE encodes token position by rotating Query/Key attention vectors by an amount depending on position, and is used in most current LLMs. (rotary positional embedding), which encodes position by rotating the Query/Key vectors by an amount depending on their position — it tends to generalize better to sequence lengths longer than what the model was trained on, which matters a lot for long-context models.
- Feedforward layer: after attention mixes information across tokens, a regular feedforward network (same idea as Neural Networks & Backprop) is applied independently to each token, adding representational capacity.
- Residual connections and normalization: each sub-layer's output is added back to its input (a "skip connection") and normalized. This is the same fix mentioned in the previous lesson for keeping gradients well-behaved across many stacked layers — transformers are typically dozens of layers deep, and residuals are what make that trainable at all.
A full transformer stacks many of these blocks (attention → feedforward, repeated), with the final layer producing a prediction — in a language model, a probability distribution over the next token.
Causal masking: attention for generation
Plain self-attention, as described above, lets every token look at every other token — including tokens later in the sequence. That's fine for understanding a complete sentence, but breaks generation: predicting the next word can't be allowed to "see" the actual next word, or the task becomes trivial and the model learns nothing useful.
Causal (masked) self-attention fixes this by zeroing out (masking to
-∞ before the softmax) every attention score where a token would attend
to a position after itself. Token 5 can attend to tokens 1–5, never to
token 6 onward. This single constraint is what makes autoregressive
generation — the token-by-token process covered in
Inference & Serving — trainable at all: the
model is trained on the exact task it will perform at generation time,
predicting each next token from only what came before it.
Masking is also where attention's main cost shows up. Because every token compares itself against every other token, the score matrix has one entry per pair of positions — so the work grows with the square of the sequence length, not linearly with it.
One attention head builds an score matrix, one float per pair of tokens. At 2 bytes per entry:
- 1K tokens → 1M entries → ~2 MB per head
- 8K tokens → 64M entries → ~128 MB per head
- 128K tokens → 16.4B entries → ~33 GB per head
Eight times the context, 64 times the score matrix. And that is one head — a model with 32 heads across 80 layers would need that many times over if it ever held them all at once. This single quadratic is why long context was a research problem rather than a configuration flag, and why real implementations never materialize the full matrix.
For sequence length and head dimension , forming costs operations and, done naively, memory. Compare that to an RNN's — linear in . Attention wins whenever and loses badly once grows large, which is exactly the regime long-context models live in.
Causal maskingCausal MaskingCausal masking restricts self-attention so each token can only attend to earlier positions, not future ones — required for training and running autoregressive generation. is implemented inside the softmax rather than after it. Define the masked scores
Since , future positions contribute nothing to the normalizing sum, and each row still sums to 1 over the positions it is allowed to see. Zeroing the weights after the softmax would not work — the remaining weights would no longer sum to 1.
This triangular structure has a second consequence: token 's output never depends on anything after , so a single forward pass over a sequence of length yields a valid next-token prediction at every position at once. That is what makes training parallel over the sequence even though generation is inherently sequential. It is also what makes the KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable. described in Inference & Serving correct: past keys and values can never change, so they can be computed once and reused.
Encoder, decoder, and encoder-decoder transformers
The original transformer paper actually describes two halves working together, and different modern model families use different halves:
- Encoder-only (e.g. BERT): uses full (non-causal) self-attention — every token sees every other token, including future ones. Good at understanding a fixed piece of text (classification, extracting answers from a passage), not at generating new text token by token.
- Decoder-only (e.g. GPT, and most current LLMs): uses causal self-attention exclusively, exactly as described above. This is the architecture behind essentially every general-purpose chat/completion LLM today, because it directly matches the pretrainingPretrainingPretraining is self-supervised training of a base LLM on massive amounts of text to predict the next token, the primary source of its knowledge and language ability. objective (predict the next token) covered in LLMs.
- Encoder-decoder (e.g. T5, the original translation-focused transformer): an encoder processes the full input (say, an English sentence) with unmasked attention, and a decoder generates the output (its French translation) with causal attention, additionally attending back to the encoder's output at each step. Common for tasks with a clear "input document → output document" shape, like translation or summarization.
Decoder-only won out as the default for general-purpose LLMs largely because a single decoder-only model, trained purely on next-token prediction, turned out to generalize to both understanding and generation tasks reasonably well via prompting — removing the need to pick an architecture per task.
Why this architecture won
Three properties made transformers displace RNNs almost entirely for sequence modeling:
- Parallelizable training: an RNN must process token 1, then token 2, then token 3, sequentially — attention computes relationships between all token pairs simultaneously, which maps extremely well onto GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires. hardware built for large parallel matrix operations.
- Long-range dependencies: any two tokens are directly connected through one attention computation, regardless of how far apart they are — no information has to be relayed step-by-step across dozens of intermediate tokens and risk degrading along the way.
- Scaling behavior: transformers turned out to keep improving predictably as you make them bigger and train them on more data (the "scaling lawsScaling LawsScaling laws are empirical relationships between a model's loss and its parameter count, dataset size, and compute budget, used to plan how large to train a new model." mentioned in History & Landscape) — a property that directly justified building the huge models covered later in LLMs.
| RNN / LSTM | Self-attention | |
|---|---|---|
| Path between two tokens | Up to steps | 1 step |
| Training over a sequence | Sequential | Parallel |
| Cost in sequence length | Linear | Quadratic |
Note the trade that was actually made: transformers bought a shorter path and parallel training by accepting a worse asymptotic cost in sequence length. That was the right call on hardware built for big parallel matrix multiplies — and it is also why nearly every long-context technique since has been an attempt to claw that quadratic back.
Recap
Self-attention replaces "process the sequence step by step" with "let every position directly weigh every other position, with weights learned from data." Stack that mechanism into blocks with feedforward layers, positional encodings, and residual connections, and you get the transformer — the architecture underneath essentially every modern language model, and increasingly vision and audio models too — including the image-generating diffusion modelsDiffusion ModelA diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step. covered next, before returning to language and scaling this architecture up into large language models.