Transformer
The transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.
The transformer is the neural network architecture introduced in the 2017 paper "Attention Is All You Need," built entirely around self-attention — letting every position in a sequence directly weigh every other position, regardless of distance. It replaced older sequence models (RNNs) because it's parallelizable to train, handles long-range dependencies directly, and scales predictably with size and data. It underlies essentially every modern LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama., and increasingly vision and audio models too.
How it works
A decoder-only transformer stacks N identical blocks. Each block has
two sublayers: multi-head self-attention and a position-wise
feed-forward network that typically expands d_model to 4 * d_model
and back. Both sublayers sit inside a
residual connectionResidual ConnectionA residual (skip) connection adds a layer's input back to its output, giving gradients a direct path backward and making very deep networks trainable. with layer
normalization, which is what makes depths of dozens of layers
trainable.
The path through the model is: token ids →
embeddingEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. lookup →
positional encodingPositional EncodingPositional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set. → blocks → a
final linear projection to vocabulary-sized logits. Shapes stay
(batch, seq, d_model) from the embedding to the last block. Because
every position is processed at once rather than step by step, an entire
training sequence fits in a single parallel forward pass.
When it breaks
- Context is a hard wall. Attention cost grows quadratically with sequence length, so the context window is a deliberate budget — exceeding it truncates rather than degrading gracefully.
- Pre-norm vs post-norm matters. The original post-norm layout needs learning-rate warmup to train stably at depth, which is why most modern implementations moved normalization inside the residual branch.
- Training parallelism doesn't carry to generation. Decoding is still one token at a time, so a KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable. is mandatory for usable inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. latency.
- Data hunger. With no convolutional or recurrent inductive bias, transformers underperform on small datasets unless pretrained.
See also: AttentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer., LLMLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama., RNNRNN (Recurrent Neural Network)An RNN is a neural network for sequences that processes input one step at a time, carrying a hidden state forward — the main predecessor to transformers.
Learn more: Attention & Transformers · Paper: Attention Is All You Need
Mentioned in
Lessons where this comes up in context.
- Applied & Agentic SystemsHow prompting, RAG, and agents combine to turn a single trained LLM into a real, capable application
- Attention & TransformersThe core architecture behind modern AI
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- Inference & ServingBatching, KV caching, quantization, latency/throughput tradeoffs
- InterpretabilityReverse-engineering what a trained network's weights actually compute — probing, superposition, sparse autoencoders, and circuits, instead of judging a model by its outputs alone
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Multimodal ModelsHow a single model handles more than one kind of input — CLIP's shared embedding space, how vision gets fed into a language model as tokens, and what actually breaks
- Neural Networks & BackpropFrom Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
- Optimization & Training DynamicsWhy plain gradient descent isn't what actually trains modern models — Adam's mechanics, why warmup exists, how learning-rate schedules are shaped, and what batch size actually trades off at scale
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
Diffusion Model
A diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step.
Attention (Self-Attention)
Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.