“Dog bites man” and “man bites dog” use the same three words. Only the order differs, and the order is the whole story. Yet attention, as you computed it two lessons ago, cannot tell them apart.
Attention sees a set, not a sequence
Nothing in mentions positions. Each token’s query, key, and value come from its embedding alone, and every query is compared with every key in the same way. If you shuffle the input tokens, the output vectors come out shuffled in the same way and are otherwise identical.
So the model has to be told where each token sits, using some form of Position encodingInformation about where each token sits in the sequence, given to a transformer because attention on its own ignores word order.Open in glossary. There are several ways to do it:
- Learned absolute embeddings. A trainable vector for each position, added to the token embedding. GPT-2 does this for positions 0 through 1,023.
- Sinusoidal encodings. A fixed vector for each position, built from sine and cosine waves, added to the token embedding. This is what the original transformer used.
- Rotary embeddings (RoPE). Rotate each query and key by an angle that depends on its position, inside every attention layer. Used by Llama, Mistral, Qwen, and many other recent models.
Other schemes exist, such as ALiBi, which subtracts a penalty proportional to distance from each attention score. This lesson covers the two most influential fixed schemes.
Sinusoidal encodings
Vaswani et al. (2017) gave each position a vector of length , defined in pairs of dimensions:
and added it to the token’s embedding before the first layer: .
Read the formula one pair at a time. Pair is a sine and cosine wave whose wavelength is positions. The first pair repeats every positions; each later pair is slower, and the last ones take tens of thousands of positions to complete a cycle. Together they work like the hands of a clock moving at different speeds: any single hand is ambiguous, but all of them together pin down the time.
Sinusoidal position encodings
Each row is the vector added to the token at that position. Each column is a wave; columns to the right oscillate more slowly.
How similar is position 12 to every other position? (dot product, scaled so a position with itself scores 1)
Try this
- Drag the dimension slider from 0 to the right. The wavelength grows from 6.3 positions to tens of thousands, and the columns on the right barely change across all 48 positions.
- Look at the similarity chart. It is highest at the position itself and generally lower further away, with some ripples. Move the position slider: the curve slides along without changing shape.
- Switch d_model to 16. With fewer waves to average, the curve becomes bumpier, so similarity is a less reliable guide to distance: some far positions score higher than some near ones.
The sliding curve is not a coincidence. Using on each pair,
which depends only on the offset . The paper chose this design partly because is a fixed linear transformation of for any offset , which might make relative positions easy to learn.
There is a catch. Once the encoding has been added to the embedding and multiplied by and , attention scores mix content and position together, and the clean “offset only” property no longer holds for the scores themselves. That observation led to methods that put relative position directly into attention.
Rotary position embeddings
Su et al. (2021) proposed Rotary position embeddingA way to encode position by rotating each pair of query and key dimensions by an angle proportional to the token's position, so attention scores depend on position only through the offset between tokens.Open in glossary, or RoPE. Instead of adding anything to the embeddings, RoPE rotates the query and key vectors just before their dot product, in every attention layer.
Split a query or key into pairs of dimensions . At position , rotate pair by the angle , with , where is the length of the query and key vectors:
Rotations preserve lengths, and rotating one vector by and another by changes the angle between them by . So the dot product of a rotated query at position with a rotated key at position depends on the positions only through .
Rotary position embedding (RoPE)
RoPE rotates each pair of query and key dimensions by an angle proportional to the token's position.
Dashed arrows are this pair of dimensions before rotation. The query turns by m × θ = 4.00 radians and the key by n × θ = 1.00. The angle between them changes by (m − n) × θ, so the dot product depends only on how far apart the two tokens are.
Press Both +1 or Both +5 to move both tokens later in the text. Both arrows turn, but neither dot product changes. Pair 0 spins fastest (one radian per position); higher pairs turn so slowly that they encode long distances.
Try this
- Press Both +5 a few times. Both arrows turn, but neither dot product changes, not for this pair and not for the full 64-dimensional vectors.
- Move only the key position. The angle between the arrows changes, and so does the dot product.
- Slide the dimension pair up toward 31. Here is tiny, and the arrows barely move even across 40 positions. These slow pairs keep distant positions distinguishable.
RoPE has become a common default for decoder language models, partly because it needs no extra parameters and works with the KV cache: a token’s rotated key never changes once computed. The base of 10,000 is a choice, not a law. Llama 3, for example, raised it to 500,000 to suit longer contexts, and techniques for extending a trained model’s context window often work by rescaling these angles.
Why rotation gives relative positionOptional
Treat each pair as a complex number, . Rotating by multiplies it by . The dot product of two pairs is the real part of one times the conjugate of the other, so
The absolute positions and disappear, leaving only . Summing over all pairs gives the full attention score, which therefore also depends only on the offset and the content of and .
Key ideas
- Self-attention is order blind: shuffling the tokens just shuffles the outputs. Position must be supplied.
- Sinusoidal encodings add a vector of sine and cosine waves to each token’s embedding, with wavelengths from up to about positions.
- The dot product of two sinusoidal encodings depends only on the distance between positions.
- RoPE rotates query and key pairs by angles proportional to position, in every attention layer, so attention scores depend on position only through the offset between tokens.
- Fast-rotating pairs resolve nearby positions; slow ones distinguish distant positions.