Attention (Self-Attention)
Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.
Self-attention is the mechanism at the core of the transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.: every token produces a Query, Key, and Value vector; attention compares each token's Query against every other token's Key to compute weights (via softmax), then produces an output as the weighted sum of every token's Value. This lets a model directly relate any two tokens regardless of distance, learned entirely from data. Multi-head attention runs several such computations in parallel, each able to specialize in a different kind of relationship.
How it works
Each input vector is projected three ways: Q = xW_q, K = xW_k,
V = xW_v. Raw scores are Q @ K.T / sqrt(d_k) — the division keeps
dot products from growing with dimension and saturating the softmax.
Softmax over the last axis turns scores into weights summing to 1 per
query, and the output is weights @ V. With batch and heads the
tensors are (batch, heads, seq, d_head) and the score matrix is
(batch, heads, seq, seq).
Multi-head attention splits d_model across h heads, runs the above
in parallel, concatenates the per-head outputs, and applies one final
projection W_o. Each head can specialize — one on local syntax,
another on long-range coreference — without competing for the same
subspace.
When it breaks
- Quadratic cost. The score matrix is
seq × seqper head, so memory and compute grow with the square of context length. Naive implementations exhaust VRAM long before the model runs out of capability; fused kernels like FlashAttention avoid materializing it. - Order blindness. Attention is permutation-equivariant on its own, so without positional encodingPositional EncodingPositional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set. the model sees an unordered bag of tokens.
- Mask bugs. Padding and causal masksCausal MaskingCausal masking restricts self-attention so each token can only attend to earlier positions, not future ones — required for training and running autoregressive generation. must be applied before the softmax as large negative values; applying them afterward leaks future or padding tokens silently.
- Weights aren't explanations. High attention on a token does not reliably mean the model's output depended on it.
See also: TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs., LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.
Learn more: Attention & Transformers · Paper: Attention Is All You Need
Mentioned in
Lessons where this comes up in context.
- Agents & Tool UseGiving an LLM the ability to take actions and chain multiple steps together — tool use, MCP, and the agent loop
- Attention & TransformersThe core architecture behind modern AI
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- How ChatGPT Was Actually BuiltA worked narrative tying pretraining, alignment, inference, and security together as one pipeline, instead of as separate topics — illustrative, synthesized from public research, not an insider account
- Inference & ServingBatching, KV caching, quantization, latency/throughput tradeoffs
- InterpretabilityReverse-engineering what a trained network's weights actually compute — probing, superposition, sparse autoencoders, and circuits, instead of judging a model by its outputs alone
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Multimodal ModelsHow a single model handles more than one kind of input — CLIP's shared embedding space, how vision gets fed into a language model as tokens, and what actually breaks
- Neural Networks & BackpropFrom Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
Transformer
The transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.
Positional Encoding
Positional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set.