Lesson 4 of 8

Where words are

Attention ignores word order unless the model is told where each token sits. Sinusoidal encodings add position as waves; rotary embeddings rotate queries and keys so attention sees relative distance.

Advanced18 min

In this lesson you will

  • Explain why self-attention needs to be given position information
  • Describe sinusoidal position encodings and what each dimension's wave represents
  • Explain how rotary position embeddings make attention scores depend on relative position

“Dog bites man” and “man bites dog” use the same three words. Only the order differs, and the order is the whole story. Yet attention, as you computed it two lessons ago, cannot tell them apart.

Attention sees a set, not a sequence

Nothing in softmax(QK⊤/dk)V\mathrm{softmax}(QK^\top/\sqrt{d_k})V mentions positions. Each token’s query, key, and value come from its embedding alone, and every query is compared with every key in the same way. If you shuffle the input tokens, the output vectors come out shuffled in the same way and are otherwise identical.

So the model has to be told where each token sits, using some form of Position encodingInformation about where each token sits in the sequence, given to a transformer because attention on its own ignores word order.Open in glossary. There are several ways to do it:

  • Learned absolute embeddings. A trainable vector for each position, added to the token embedding. GPT-2 does this for positions 0 through 1,023.
  • Sinusoidal encodings. A fixed vector for each position, built from sine and cosine waves, added to the token embedding. This is what the original transformer used.
  • Rotary embeddings (RoPE). Rotate each query and key by an angle that depends on its position, inside every attention layer. Used by Llama, Mistral, Qwen, and many other recent models.

Other schemes exist, such as ALiBi, which subtracts a penalty proportional to distance from each attention score. This lesson covers the two most influential fixed schemes.

Sinusoidal encodings

Vaswani et al. (2017) gave each position pp a vector PE(p)\mathrm{PE}(p) of length dmodeld_{\text{model}}, defined in pairs of dimensions:

PE(p,2i)=sin⁡ ⁣(p100002i/dmodel),PE(p,2i+1)=cos⁡ ⁣(p100002i/dmodel)\mathrm{PE}(p, 2i) = \sin\!\left(\frac{p}{10000^{2i/d_{\text{model}}}}\right), \qquad \mathrm{PE}(p, 2i+1) = \cos\!\left(\frac{p}{10000^{2i/d_{\text{model}}}}\right)

and added it to the token’s embedding before the first layer: xp=embedding(tokenp)+PE(p)x_p = \mathrm{embedding}(\text{token}_p) + \mathrm{PE}(p).

Read the formula one pair at a time. Pair ii is a sine and cosine wave whose wavelength is 2π⋅100002i/dmodel2\pi \cdot 10000^{2i/d_{\text{model}}} positions. The first pair repeats every 2π≈6.32\pi \approx 6.3 positions; each later pair is slower, and the last ones take tens of thousands of positions to complete a cycle. Together they work like the hands of a clock moving at different speeds: any single hand is ambiguous, but all of them together pin down the time.

Sinusoidal position encodings

Each row is the vector added to the token at that position. Each column is a wave; columns to the right oscillate more slowly.

How similar is position 12 to every other position? (dot product, scaled so a position with itself scores 1)

10.80.60.40.2012243647
d_model
12
0
Wavelength of dimension 06.3 positions
Similarity of 12 to 130.966

Try this

  • Drag the dimension slider from 0 to the right. The wavelength grows from 6.3 positions to tens of thousands, and the columns on the right barely change across all 48 positions.
  • Look at the similarity chart. It is highest at the position itself and generally lower further away, with some ripples. Move the position slider: the curve slides along without changing shape.
  • Switch d_model to 16. With fewer waves to average, the curve becomes bumpier, so similarity is a less reliable guide to distance: some far positions score higher than some near ones.

The sliding curve is not a coincidence. Using sin⁡asin⁡b+cos⁡acos⁡b=cos⁡(a−b)\sin a \sin b + \cos a \cos b = \cos(a - b) on each pair,

PE(p)⋅PE(q)=∑icos⁡(ωi(p−q)),ωi=10000−2i/dmodel,\mathrm{PE}(p) \cdot \mathrm{PE}(q) = \sum_i \cos\big(\omega_i (p - q)\big), \qquad \omega_i = 10000^{-2i/d_{\text{model}}},

which depends only on the offset p−qp - q. The paper chose this design partly because PE(p+k)\mathrm{PE}(p + k) is a fixed linear transformation of PE(p)\mathrm{PE}(p) for any offset kk, which might make relative positions easy to learn.

There is a catch. Once the encoding has been added to the embedding and multiplied by WQW_Q and WKW_K, attention scores mix content and position together, and the clean “offset only” property no longer holds for the scores themselves. That observation led to methods that put relative position directly into attention.

Rotary position embeddings

Su et al. (2021) proposed Rotary position embeddingA way to encode position by rotating each pair of query and key dimensions by an angle proportional to the token's position, so attention scores depend on position only through the offset between tokens.Open in glossary, or RoPE. Instead of adding anything to the embeddings, RoPE rotates the query and key vectors just before their dot product, in every attention layer.

Split a query or key into pairs of dimensions (x2i,x2i+1)(x_{2i}, x_{2i+1}). At position mm, rotate pair ii by the angle mθim\theta_i, with θi=10000−2i/d\theta_i = 10000^{-2i/d}, where dd is the length of the query and key vectors:

(x2i′x2i+1′)=(cos⁡mθi−sin⁡mθisin⁡mθicos⁡mθi)(x2ix2i+1)\begin{pmatrix} x'_{2i} \\ x'_{2i+1} \end{pmatrix} = \begin{pmatrix} \cos m\theta_i & -\sin m\theta_i \\ \sin m\theta_i & \cos m\theta_i \end{pmatrix} \begin{pmatrix} x_{2i} \\ x_{2i+1} \end{pmatrix}

Rotations preserve lengths, and rotating one vector by mθm\theta and another by nθn\theta changes the angle between them by (m−n)θ(m - n)\theta. So the dot product of a rotated query at position mm with a rotated key at position nn depends on the positions only through m−nm - n.

Rotary position embedding (RoPE)

RoPE rotates each pair of query and key dimensions by an angle proportional to the token's position.

k at 1q at 4

Dashed arrows are this pair of dimensions before rotation. The query turns by m × θ = 4.00 radians and the key by n × θ = 1.00. The angle between them changes by (m − n) × θ, so the dot product depends only on how far apart the two tokens are.

Press Both +1 or Both +5 to move both tokens later in the text. Both arrows turn, but neither dot product changes. Pair 0 spins fastest (one radian per position); higher pairs turn so slowly that they encode long distances.

4
1
0 (θ = 1.0)
Relative position, m − n3
Pair 0 dot product-0.898
Full 64-dim q · k-6.853

Try this

  • Press Both +5 a few times. Both arrows turn, but neither dot product changes, not for this pair and not for the full 64-dimensional vectors.
  • Move only the key position. The angle between the arrows changes, and so does the dot product.
  • Slide the dimension pair up toward 31. Here θ\theta is tiny, and the arrows barely move even across 40 positions. These slow pairs keep distant positions distinguishable.

RoPE has become a common default for decoder language models, partly because it needs no extra parameters and works with the KV cache: a token’s rotated key never changes once computed. The base of 10,000 is a choice, not a law. Llama 3, for example, raised it to 500,000 to suit longer contexts, and techniques for extending a trained model’s context window often work by rescaling these angles.

Why rotation gives relative positionOptional

Treat each pair as a complex number, z=x2i+i x2i+1z = x_{2i} + \mathrm{i}\,x_{2i+1}. Rotating by mθm\theta multiplies it by eimθe^{\mathrm{i} m\theta}. The dot product of two pairs is the real part of one times the conjugate of the other, so

Re ⁣[(q eimθ) (k einθ)‾]=Re ⁣[q kˉ ei(m−n)θ].\mathrm{Re}\!\left[(q\, e^{\mathrm{i} m\theta})\,\overline{(k\, e^{\mathrm{i} n\theta})}\right] = \mathrm{Re}\!\left[q\, \bar{k}\, e^{\mathrm{i}(m - n)\theta}\right].

The absolute positions mm and nn disappear, leaving only m−nm - n. Summing over all pairs gives the full attention score, which therefore also depends only on the offset and the content of qq and kk.

Key ideas

  • Self-attention is order blind: shuffling the tokens just shuffles the outputs. Position must be supplied.
  • Sinusoidal encodings add a vector of sine and cosine waves to each token’s embedding, with wavelengths from 2π2\pi up to about 10000⋅2π10000 \cdot 2\pi positions.
  • The dot product of two sinusoidal encodings depends only on the distance between positions.
  • RoPE rotates query and key pairs by angles proportional to position, in every attention layer, so attention scores depend on position only through the offset between tokens.
  • Fast-rotating pairs resolve nearby positions; slow ones distinguish distant positions.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1Why does a transformer need position information at all?
2In the sinusoidal encoding, which dimensions change fastest as the position increases?
3With RoPE, a query at position 10 meets a key at position 7. Then the same two vectors appear at positions 110 and 107. How do the attention scores compare?

Progress is saved in this browser only.

Up nextThe transformer block
Next
Transformers and LLMs
  1. 1Predicting the next token
  2. 2Attention, step by step
  3. 3Many heads and the causal mask
  4. 4Where words are
  5. 5The transformer block
  6. 6Inside a real language model
  7. 7How LLMs are trained
  8. 8Generating fast: the KV cache

Try "embedding", "softmax", "overfitting", or "backpropagation".