Position encodings

Explore the sine and cosine waves of the original transformer's position encodings, then rotate queries and keys with RoPE and watch their dot product depend on position only through the offset.

AdvancedExplained in Where words are

Sinusoidal position encodings

Each row is the vector added to the token at that position. Each column is a wave; columns to the right oscillate more slowly.

How similar is position 12 to every other position? (dot product, scaled so a position with itself scores 1)

10.80.60.40.2012243647
d_model
12
0
Wavelength of dimension 06.3 positions
Similarity of 12 to 130.966

Try this

  • Tap anywhere on the heatmap to pick a position and a dimension at once.
  • Compare the similarity curve at d_model 16 and 128.

Rotary position embedding (RoPE)

RoPE rotates each pair of query and key dimensions by an angle proportional to the token's position.

k at 1q at 4

Dashed arrows are this pair of dimensions before rotation. The query turns by m × θ = 4.00 radians and the key by n × θ = 1.00. The angle between them changes by (m − n) × θ, so the dot product depends only on how far apart the two tokens are.

Press Both +1 or Both +5 to move both tokens later in the text. Both arrows turn, but neither dot product changes. Pair 0 spins fastest (one radian per position); higher pairs turn so slowly that they encode long distances.

4
1
0 (θ = 1.0)
Relative position, m − n3
Pair 0 dot product-0.898
Full 64-dim q · k-6.853

Try this

  • Set the query and key to the same position. The dot product is the same as with no rotation at all.
  • Find a pair where shifting both positions by 5 visibly turns the arrows, and confirm the dot product stays put.

Try "embedding", "softmax", "overfitting", or "backpropagation".