Attention calculator

Drag queries, keys, and values on a plane and watch scaled dot-product attention compute scores, weights, and outputs for every token.

AdvancedExplained in Attention, step by step

Attention, one query at a time

Pick a token. Its query is compared with every key, the scores become weights, and the weights mix the values.

Choose the query token
k Thek catk satk onk thek matq sat

Drag the arrow tips. A key that points the same way as the query gets a high score.

Scores and weights for the query "sat" against every key
Keyq · k÷ √2softmax weight
-2.52-1.780%
5.884.1676%
2.281.616%
-0.96-0.681%
-3.48-2.460%
3.782.6717%

qsat · kcat = 2.40 × 2.20 + 0.60 × 1.00 = 5.88

output = Σ weight × v = (1.54, 0.48)

All six queries at once: the attention weight matrix, softmax(QKᵀ/√d_k)

Attention weights. Rows are queries, columns are keys. Select a row to make it the current query.
query ↓ key →Thecatsatonthemat
0.010.620.110.020.010.23
0.040.160.580.190.010.01
0.000.760.060.010.000.17
0.010.220.010.010.010.74
0.010.310.030.010.020.62
0.090.050.500.330.030.00

Shapes: Q, K, V are 6 × 2. QKᵀ is 6 × 6, one row per query. The output is 6 × 2, one mixed vector per token.

The plane shows
Query tokensat
Attends most tocat, 76%
Output vector(1.54, 0.48)

Try this

  • Pick mat as the query. Which key does it favor, and what happens to the output when you drag that key’s value somewhere new?
  • Drag two keys to exactly the same point. Their weights become equal for every query. (Same direction is not enough: a longer key gets a larger score.)
  • Turn off scaling and compare the matrix with scaling on. Every row gets sharper.

Try "embedding", "softmax", "overfitting", or "backpropagation".