Transformers & LLMs

Transformer

The transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.

The transformer is the neural network architecture introduced in the 2017 paper "Attention Is All You Need," built entirely around self-attention — letting every position in a sequence directly weigh every other position, regardless of distance. It replaced older sequence models (RNNs) because it's parallelizable to train, handles long-range dependencies directly, and scales predictably with size and data. It underlies essentially every modern LLM, and increasingly vision and audio models too.

How it works

A decoder-only transformer stacks N identical blocks. Each block has two sublayers: multi-head self-attention and a position-wise feed-forward network that typically expands d_model to 4 * d_model and back. Both sublayers sit inside a residual connection with layer normalization, which is what makes depths of dozens of layers trainable.

The path through the model is: token ids → embedding lookup → positional encoding → blocks → a final linear projection to vocabulary-sized logits. Shapes stay (batch, seq, d_model) from the embedding to the last block. Because every position is processed at once rather than step by step, an entire training sequence fits in a single parallel forward pass.

When it breaks

  • Context is a hard wall. Attention cost grows quadratically with sequence length, so the context window is a deliberate budget — exceeding it truncates rather than degrading gracefully.
  • Pre-norm vs post-norm matters. The original post-norm layout needs learning-rate warmup to train stably at depth, which is why most modern implementations moved normalization inside the residual branch.
  • Training parallelism doesn't carry to generation. Decoding is still one token at a time, so a KV cache is mandatory for usable inference latency.
  • Data hunger. With no convolutional or recurrent inductive bias, transformers underperform on small datasets unless pretrained.

See also: Attention, LLM, RNN

Learn more: Attention & Transformers · Paper: Attention Is All You Need

Mentioned in

Lessons where this comes up in context.

On this page