Transformer anatomy

Select any part of a decoder-only transformer to see its job, its tensor shapes, and its parameter count, for models from GPT-2 to Llama 3.

AdvancedExplained in The transformer block

Anatomy of a transformer

Select any part to see what it does, its tensor shapes, and how many parameters it holds.

Parts of GPT-2 small

Residual stream, 8 x 768

Block 1 of 12

11 more blocks with the same shape and their own weights

Multi-head self-attention

The only place where tokens exchange information. Each of 12 heads projects every token to a query, key, and value of size 64, scores each query against the keys of the tokens at or before it (the causal mask), and takes a softmax-weighted sum of their values. The heads' outputs are joined and projected back to 768 numbers by W_O, then added to the stream.

Shapes for Multi-head self-attention
W_Q768 x 768
W_K768 x 768
W_V768 x 768
W_O768 x 768
Scores, per head8 x 8
Out8 x 768

2,362,368 parameters per block, 28,348,416 across all 12, 23% of the model.

Where GPT-2 small's 124,439,808 parameters live

  • Token and position embeddings39.4M, 32%
  • Attention, all blocks28.3M, 23%
  • MLP, all blocks56.7M, 46%
  • Norms38.4K, under 0.1%
Model
8 tokens

Up to this model's context of 1,024.

Parameters124M
Without embeddings85.1M
Forward FLOPs per token at T = 80.25 G

Try this

  • Select every part of GPT-2 small and add up the parameters yourself. The total matches the released checkpoint exactly.
  • Raise the sequence length and watch which shapes change. Weights never do.
  • Compare where the parameters live in GPT-2 small and in Llama 3 70B.

Parameter counts are computed from each model’s published configuration and checked against the released checkpoints in this site’s tests.

Try "embedding", "softmax", "overfitting", or "backpropagation".