Architectures

Attention & Transformers

The core architecture behind modern AI

CNNs (previous lesson) encode an assumption that works great for images: nearby things are related. Language breaks that assumption constantly — in "the trophy didn't fit in the suitcase because it was too big," what "it" refers to depends on a word far earlier in the sentence, and that distance varies sentence to sentence. Before 2017, sequence models (RNNs, LSTMs) processed text one token at a time, left to right, which made long-range relationships hard to learn and impossible to parallelize during training. The transformer architecture solved both problems with one mechanism: self-attention.

Here is the shape of the thing before we take it apart. A transformer is the same block repeated dozens of times, and one block looks like this:

The two arrows that skip past attention and past the MLP are the residual connections — the block adds to its input rather than replacing it. Everything below fills in those boxes.

The core idea: attention as weighted lookup

Self-attention lets every position in a sequence directly look at every other position and decide how much to "pay attention" to it, regardless of distance. For the sentence above, when processing the word "it," attention lets the model directly weigh "trophy" and "suitcase" against each other and learn — from data, via training — that "it" should attend strongly to "trophy" in this context.

Mechanically, every token produces three vectors, each obtained by multiplying the token's embedding by a learned weight matrix:

  • Query (Q): "what am I looking for?" — from this token's perspective
  • Key (K): "what do I contain?" — advertised by every other token
  • Value (V): "what information do I actually offer if attended to?"

Attention for a given token compares its Query against every token's Key (via a dot product — a similarity score), turns those scores into weights that sum to 1 (via softmax), and produces an output as the weighted sum of every token's Value:

Attention(Q, K, V) = softmax(Q · Kᵀ / √d_k) · V

The √d_k term just rescales the dot products to keep training stable; the important part is softmax(Q · Kᵀ): a matrix of "how much should token i attend to token j" for every pair of tokens, learned entirely from data, not hand-specified.

Click through a few example sentences below to build intuition for what that attention matrix actually looks like. These weights are hand-set to illustrate the pattern a trained model learns, not live model output — but they show the same kind of long-range dependency real attention heads pick up on:

Click "it" — coreference resolution, exactly the lesson's own example.

Click any token above to see what it attends to.

Why "self"-attention, and why "multi-head"

It's called self-attention because Q, K, and V all come from the same sequence — every token attends to (potentially) every other token in the same input, including itself.

Why does attention need multiple heads?

In practice, transformers use multi-head attention: run several attention computations in parallel, each with its own learned Q/K/V projection matrices, then combine the results. Each "head" can specialize in a different kind of relationship — one head might learn to track subject-verb agreement, another might track coreference (like "it" → "trophy"), another might track adjacent-word syntax. No one designs these specializations; they emerge from training, similar to how CNN layers learn their own feature hierarchy.

The rest of the transformer block

Self-attention alone is just weighted averaging — it needs to be paired with a few other components to form a working architecture:

  • Positional information: attention itself has no notion of word order — swap two tokens and, by itself, attention treats it identically. Transformers inject order via positional encodings added to each token's embedding before attention, so the model can tell "dog bites man" from "man bites dog." The original transformer used a fixed sinusoidal encoding (a deterministic pattern of sine/cosine waves, one per position); many models instead learn positional embeddings directly from data. Most current LLMs use a newer scheme called RoPE (rotary positional embedding), which encodes position by rotating the Query/Key vectors by an amount depending on their position — it tends to generalize better to sequence lengths longer than what the model was trained on, which matters a lot for long-context models.
  • Feedforward layer: after attention mixes information across tokens, a regular feedforward network (same idea as Neural Networks & Backprop) is applied independently to each token, adding representational capacity.
  • Residual connections and normalization: each sub-layer's output is added back to its input (a "skip connection") and normalized. This is the same fix mentioned in the previous lesson for keeping gradients well-behaved across many stacked layers — transformers are typically dozens of layers deep, and residuals are what make that trainable at all.

A full transformer stacks many of these blocks (attention → feedforward, repeated), with the final layer producing a prediction — in a language model, a probability distribution over the next token.

Causal masking: attention for generation

Plain self-attention, as described above, lets every token look at every other token — including tokens later in the sequence. That's fine for understanding a complete sentence, but breaks generation: predicting the next word can't be allowed to "see" the actual next word, or the task becomes trivial and the model learns nothing useful.

Causal (masked) self-attention fixes this by zeroing out (masking to -∞ before the softmax) every attention score where a token would attend to a position after itself. Token 5 can attend to tokens 1–5, never to token 6 onward. This single constraint is what makes autoregressive generation — the token-by-token process covered in Inference & Serving — trainable at all: the model is trained on the exact task it will perform at generation time, predicting each next token from only what came before it.

Masking is also where attention's main cost shows up. Because every token compares itself against every other token, the score matrix has one entry per pair of positions — so the work grows with the square of the sequence length, not linearly with it.

Why long context is expensive

One attention head builds an n×nn \times n score matrix, one float per pair of tokens. At 2 bytes per entry:

  • 1K tokens → 1M entries → ~2 MB per head
  • 8K tokens → 64M entries → ~128 MB per head
  • 128K tokens → 16.4B entries → ~33 GB per head

Eight times the context, 64 times the score matrix. And that is one head — a model with 32 heads across 80 layers would need that many times over if it ever held them all at once. This single quadratic is why long context was a research problem rather than a configuration flag, and why real implementations never materialize the full matrix.

Encoder, decoder, and encoder-decoder transformers

The original transformer paper actually describes two halves working together, and different modern model families use different halves:

  • Encoder-only (e.g. BERT): uses full (non-causal) self-attention — every token sees every other token, including future ones. Good at understanding a fixed piece of text (classification, extracting answers from a passage), not at generating new text token by token.
  • Decoder-only (e.g. GPT, and most current LLMs): uses causal self-attention exclusively, exactly as described above. This is the architecture behind essentially every general-purpose chat/completion LLM today, because it directly matches the pretraining objective (predict the next token) covered in LLMs.
  • Encoder-decoder (e.g. T5, the original translation-focused transformer): an encoder processes the full input (say, an English sentence) with unmasked attention, and a decoder generates the output (its French translation) with causal attention, additionally attending back to the encoder's output at each step. Common for tasks with a clear "input document → output document" shape, like translation or summarization.

Decoder-only won out as the default for general-purpose LLMs largely because a single decoder-only model, trained purely on next-token prediction, turned out to generalize to both understanding and generation tasks reasonably well via prompting — removing the need to pick an architecture per task.

Why this architecture won

Three properties made transformers displace RNNs almost entirely for sequence modeling:

  1. Parallelizable training: an RNN must process token 1, then token 2, then token 3, sequentially — attention computes relationships between all token pairs simultaneously, which maps extremely well onto GPU hardware built for large parallel matrix operations.
  2. Long-range dependencies: any two tokens are directly connected through one attention computation, regardless of how far apart they are — no information has to be relayed step-by-step across dozens of intermediate tokens and risk degrading along the way.
  3. Scaling behavior: transformers turned out to keep improving predictably as you make them bigger and train them on more data (the "scaling laws" mentioned in History & Landscape) — a property that directly justified building the huge models covered later in LLMs.
RNN / LSTMSelf-attention
Path between two tokensUp to nn steps1 step
Training over a sequenceSequentialParallel
Cost in sequence lengthLinearQuadratic

Note the trade that was actually made: transformers bought a shorter path and parallel training by accepting a worse asymptotic cost in sequence length. That was the right call on hardware built for big parallel matrix multiplies — and it is also why nearly every long-context technique since has been an attempt to claw that quadratic back.

Recap

Self-attention replaces "process the sequence step by step" with "let every position directly weigh every other position, with weights learned from data." Stack that mechanism into blocks with feedforward layers, positional encodings, and residual connections, and you get the transformer — the architecture underneath essentially every modern language model, and increasingly vision and audio models too — including the image-generating diffusion models covered next, before returning to language and scaling this architecture up into large language models.

On this page