Causal Masking
Causal masking restricts self-attention so each token can only attend to earlier positions, not future ones — required for training and running autoregressive generation.
Causal (masked) self-attention restricts attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. so a token can only attend to positions at or before itself, never later ones — by zeroing out attention scores to future positions before the softmax. This is what makes autoregressive inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. trainable: the model is trained on exactly the task it performs at generation time, predicting each next token from only what came before it. Decoder-only LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. (GPT-style) use causal masking throughout.
How it works
An additive mask is applied to the raw score matrix before the softmax:
entry (i, j) is set to -inf — in practice a large negative constant
— whenever j > i. Softmax maps those entries to zero probability, so
position i's output is a weighted sum over positions 0..i only. The
mask is a fixed upper-triangular pattern, built once with something
like torch.triu(..., diagonal=1), and has no learned parameters.
This is what lets one training forward pass supply seq supervised
examples at once: every position predicts its own next token without
ever seeing it. The same constraint holds by construction at generation
time, which is also what makes a KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable. valid
— past keys and values can never change.
When it breaks
- Off-by-one in the mask. Masking the diagonal stops a token from
attending to itself and wrecks training; using
>=instead of>leaks the target. Both surface as an implausibly low or flat loss rather than a crash. - Padding interacts with it. Batched inputs need the causal mask and a padding mask combined; forgetting the second lets real tokens attend to pad embeddings.
- Bidirectional tasks suffer. Classification and retrieval usually want full context, which is why encoder-decoderEncoder-Decoder ArchitectureEncoder-only, decoder-only, and encoder-decoder describe which half of the original transformer a model uses, determining whether it's built for understanding, generation, or both. and encoder-only models drop the mask on the input side.
See also: AttentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer., InferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering.
Learn more: Attention & Transformers
Mentioned in
Lessons where this comes up in context.
RoPE (Rotary Positional Embedding)
RoPE encodes token position by rotating Query/Key attention vectors by an amount depending on position, and is used in most current LLMs.
Encoder-Decoder Architecture
Encoder-only, decoder-only, and encoder-decoder describe which half of the original transformer a model uses, determining whether it's built for understanding, generation, or both.