Anatomy of a transformer
Select any part to see what it does, its tensor shapes, and how many parameters it holds.
Multi-head self-attention
The only place where tokens exchange information. Each of 12 heads projects every token to a query, key, and value of size 64, scores each query against the keys of the tokens at or before it (the causal mask), and takes a softmax-weighted sum of their values. The heads' outputs are joined and projected back to 768 numbers by W_O, then added to the stream.
| W_Q | 768 x 768 |
|---|---|
| W_K | 768 x 768 |
| W_V | 768 x 768 |
| W_O | 768 x 768 |
| Scores, per head | 8 x 8 |
| Out | 8 x 768 |
2,362,368 parameters per block, 28,348,416 across all 12, 23% of the model.
Where GPT-2 small's 124,439,808 parameters live
- Token and position embeddings39.4M, 32%
- Attention, all blocks28.3M, 23%
- MLP, all blocks56.7M, 46%
- Norms38.4K, under 0.1%
Try this
- Select every part of GPT-2 small and add up the parameters yourself. The total matches the released checkpoint exactly.
- Raise the sequence length and watch which shapes change. Weights never do.
- Compare where the parameters live in GPT-2 small and in Llama 3 70B.
Parameter counts are computed from each model’s published configuration and checked against the released checkpoints in this site’s tests.