Positional Encoding
Positional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set.
Positional encoding injects order information into a transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs., since self-attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. alone has no notion of sequence order — swap two tokens and attention treats it identically. The original transformer used a fixed sinusoidal encoding; many models learn positional embeddings directly from data instead. Most current LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama. use RoPERoPE (Rotary Positional Embedding)RoPE encodes token position by rotating Query/Key attention vectors by an amount depending on position, and is used in most current LLMs., a newer scheme that tends to generalize better to sequence lengths longer than training.
How it works
Sinusoidal encoding builds a fixed matrix of shape
(max_seq, d_model) in which each dimension is a sine or cosine of
position at a different frequency, then adds it to the token
embeddingsEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space.. The geometric spread of frequencies
means nearby positions get similar vectors while distant ones stay
distinguishable. Learned absolute encodings replace that matrix with a
trainable table of the same shape.
Both are absolute: they tag each position with an identity. Later
schemes encode relative distance instead, biasing the attention score
between positions i and j by a function of i - j. That is usually
what you want, since language depends on how far apart two tokens are
rather than where they sit in the buffer.
When it breaks
- Extrapolation fails. Learned tables have no entry beyond
max_seq, and sinusoidal encodings, though defined everywhere, produce out-of-distribution inputs past the trained length — quality collapses rather than degrading gently. - Added, not concatenated. Absolute encodings share dimensions with the token signal, so the two compete for the same capacity.
- Off-by-one on the offset. With a KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable., each new token must be encoded at its index in the full sequence, not at index 0 of the current step. Getting this wrong yields fluent text that ignores the prompt.
- Padding shifts positions. Left-padding a batch offsets every real token unless positions are derived from the attention mask.
See also: RoPERoPE (Rotary Positional Embedding)RoPE encodes token position by rotating Query/Key attention vectors by an amount depending on position, and is used in most current LLMs., TransformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.
Learn more: Attention & Transformers
Mentioned in
Lessons where this comes up in context.
Attention (Self-Attention)
Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.
RoPE (Rotary Positional Embedding)
RoPE encodes token position by rotating Query/Key attention vectors by an amount depending on position, and is used in most current LLMs.