Transformers & LLMs

Positional Encoding

Positional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set.

Positional encoding injects order information into a transformer, since self-attention alone has no notion of sequence order — swap two tokens and attention treats it identically. The original transformer used a fixed sinusoidal encoding; many models learn positional embeddings directly from data instead. Most current LLMs use RoPE, a newer scheme that tends to generalize better to sequence lengths longer than training.

How it works

Sinusoidal encoding builds a fixed matrix of shape (max_seq, d_model) in which each dimension is a sine or cosine of position at a different frequency, then adds it to the token embeddings. The geometric spread of frequencies means nearby positions get similar vectors while distant ones stay distinguishable. Learned absolute encodings replace that matrix with a trainable table of the same shape.

Both are absolute: they tag each position with an identity. Later schemes encode relative distance instead, biasing the attention score between positions i and j by a function of i - j. That is usually what you want, since language depends on how far apart two tokens are rather than where they sit in the buffer.

When it breaks

  • Extrapolation fails. Learned tables have no entry beyond max_seq, and sinusoidal encodings, though defined everywhere, produce out-of-distribution inputs past the trained length — quality collapses rather than degrading gently.
  • Added, not concatenated. Absolute encodings share dimensions with the token signal, so the two compete for the same capacity.
  • Off-by-one on the offset. With a KV cache, each new token must be encoded at its index in the full sequence, not at index 0 of the current step. Getting this wrong yields fluent text that ignores the prompt.
  • Padding shifts positions. Left-padding a batch offsets every real token unless positions are derived from the attention mask.

See also: RoPE, Transformer

Learn more: Attention & Transformers

Mentioned in

Lessons where this comes up in context.

On this page