Attention, one query at a time
Pick a token. Its query is compared with every key, the scores become weights, and the weights mix the values.
Drag the arrow tips. A key that points the same way as the query gets a high score.
| Key | q · k | ÷ √2 | softmax weight |
|---|---|---|---|
| -2.52 | -1.78 | 0% | |
| 5.88 | 4.16 | 76% | |
| 2.28 | 1.61 | 6% | |
| -0.96 | -0.68 | 1% | |
| -3.48 | -2.46 | 0% | |
| 3.78 | 2.67 | 17% |
qsat · kcat = 2.40 × 2.20 + 0.60 × 1.00 = 5.88
output = Σ weight × v = (1.54, 0.48)
All six queries at once: the attention weight matrix, softmax(QKᵀ/√d_k)
| query ↓ key → | The | cat | sat | on | the | mat |
|---|---|---|---|---|---|---|
| 0.01 | 0.62 | 0.11 | 0.02 | 0.01 | 0.23 | |
| 0.04 | 0.16 | 0.58 | 0.19 | 0.01 | 0.01 | |
| 0.00 | 0.76 | 0.06 | 0.01 | 0.00 | 0.17 | |
| 0.01 | 0.22 | 0.01 | 0.01 | 0.01 | 0.74 | |
| 0.01 | 0.31 | 0.03 | 0.01 | 0.02 | 0.62 | |
| 0.09 | 0.05 | 0.50 | 0.33 | 0.03 | 0.00 |
Shapes: Q, K, V are 6 × 2. QKᵀ is 6 × 6, one row per query. The output is 6 × 2, one mixed vector per token.
Query tokensat
Attends most tocat, 76%
Output vector(1.54, 0.48)
Try this
- Pick mat as the query. Which key does it favor, and what happens to the output when you drag that key’s value somewhere new?
- Drag two keys to exactly the same point. Their weights become equal for every query. (Same direction is not enough: a longer key gets a larger score.)
- Turn off scaling and compare the matrix with scaling on. Every row gets sharper.