RoPE (Rotary Positional Embedding)
RoPE encodes token position by rotating Query/Key attention vectors by an amount depending on position, and is used in most current LLMs.
RoPE (Rotary Positional Embedding) is a positional encodingPositional EncodingPositional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set. scheme used in most current LLMsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama.. Instead of adding a positional vector to each token's embedding, it encodes position by rotating the Query and Key vectors used in attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. by an amount depending on their position in the sequence. It tends to generalize better to sequence lengths longer than what the model was trained on, which matters for long-context models.
How it works
RoPE treats each Query and Key vector as d_head / 2 two-dimensional
pairs and rotates pair m by an angle pos * theta_m, where the
frequencies theta_m decrease geometrically across pairs — fast
rotation in early dimensions, slow in later ones. Because the dot
product of two rotated vectors depends only on the difference of
their angles, the resulting attention score is a function of i - j:
absolute positions go in, relative distance comes out.
The rotation is applied to Q and K after their projections but before scores are computed, and never to V. It adds no parameters and no KV cacheKV CacheThe KV cache stores each token's Key and Value attention vectors so they don't need to be recomputed at every generation step, making LLM inference tractable. overhead, since keys are stored already rotated at their absolute index.
When it breaks
- Still bounded by training length. RoPE degrades more gracefully than a learned table but does not extrapolate for free; going past the trained window needs frequency rescaling such as position interpolation or YaRN, usually with a short fine-tuneFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions..
- The base frequency is part of the contract. Changing the base
(commonly
10000) without retraining invalidates every cached key and scrambles relative positions. - The cache index must be absolute. Rotating a resumed token by its offset within the current chunk rather than the full sequence produces coherent output that ignores the prompt.
- Long-range decay. Attention logits attenuate with distance, which helps locality but weakens genuinely long dependencies.
See also: Positional EncodingPositional EncodingPositional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set., AttentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer.
Learn more: Attention & Transformers
Mentioned in
Lessons where this comes up in context.
Positional Encoding
Positional encoding injects word-order information into a transformer, since self-attention alone treats a sequence's tokens as an unordered set.
Causal Masking
Causal masking restricts self-attention so each token can only attend to earlier positions, not future ones — required for training and running autoregressive generation.