Five heads, one sentence
Each head applies its own scoring rule to the same tokens. Pick a head, then pick a query token to see where it looks.
Each token finds an earlier copy of itself and looks at the token that came next, which helps predict a repeat.
Induction heads are described by Elhage et al. (2021) and Olsson et al. (2022), who tie them to in-context learning.
| Query | the | cat | chased | the | dog | . | the | cat | chased | the |
|---|---|---|---|---|---|---|---|---|---|---|
| 1.00 | masked | masked | masked | masked | masked | masked | masked | masked | masked | |
| 0.88 | 0.12 | masked | masked | masked | masked | masked | masked | masked | masked | |
| 0.79 | 0.11 | 0.11 | masked | masked | masked | masked | masked | masked | masked | |
| 0.05 | 0.94 | 0.01 | 0.01 | masked | masked | masked | masked | masked | masked | |
| 0.65 | 0.09 | 0.09 | 0.09 | 0.09 | masked | masked | masked | masked | masked | |
| 0.60 | 0.08 | 0.08 | 0.08 | 0.08 | 0.08 | masked | masked | masked | masked | |
| 0.02 | 0.48 | 0.00 | 0.00 | 0.48 | 0.00 | 0.00 | masked | masked | masked | |
| 0.05 | 0.01 | 0.92 | 0.01 | 0.01 | 0.01 | 0.01 | 0.01 | masked | masked | |
| 0.05 | 0.01 | 0.01 | 0.91 | 0.01 | 0.01 | 0.01 | 0.01 | 0.01 | masked | |
| 0.02 | 0.32 | 0.00 | 0.00 | 0.32 | 0.00 | 0.00 | 0.32 | 0.00 | 0.00 |
Try this
- For each head, predict where position 9 will look before you tap it.
- Turn off the causal mask and look at which heads change. The previous-token and first-token heads barely change: their high scores are never on later tokens, so the future gets only a sliver of weight.
Splitting the model width into heads
More heads means narrower heads. The total size of the attention layer does not change.
One token's 512-dimensional vector, cut into 8 heads of 64 dimensions. Each head has its own query, key, and value projections of shape 512 × 64, runs attention independently, and the 8 outputs are concatenated back to width 512 and multiplied by WO (512 × 512).