You have now met the two ideas that make a TransformerA neural network architecture built from stacked layers of attention and feed-forward networks, introduced by Vaswani et al. in 2017. Most modern language models are transformers.Open in glossary work: AttentionA mechanism that lets each token build a new vector as a weighted mix of other tokens' value vectors, with weights set by how well its query matches their keys.Open in glossary, which lets tokens look at each other, and position encodings, which tell them where they are. A language model wraps those ideas in a block and repeats the block, 12 times in GPT-2 small and 80 times in Llama 3 70B. This lesson takes one block apart and then puts the whole model back together.
From token IDs to logits
Every decoder-only language model, from GPT-2 to Llama 3, has the same outline:
- Embed. Each token ID selects a row of a learned table, giving a vector of numbers per token. For tokens that is a matrix.
- Repeat the block times. Each block takes a matrix and returns one of the same shape.
- Unembed. A final normalization, then a matrix multiplication turns each position’s numbers into scores, one per vocabulary token. These are the logits from the first lesson of this track.
The second step is what makes depth possible. Because a block’s input and output have the same shape, block 7 does not care whether it follows block 6 or block 1, and adding more blocks is a one-line change to the configuration.
Anatomy of a transformer
Select any part to see what it does, its tensor shapes, and how many parameters it holds.
Multi-head self-attention
The only place where tokens exchange information. Each of 12 heads projects every token to a query, key, and value of size 64, scores each query against the keys of the tokens at or before it (the causal mask), and takes a softmax-weighted sum of their values. The heads' outputs are joined and projected back to 768 numbers by W_O, then added to the stream.
| W_Q | 768 x 768 |
|---|---|
| W_K | 768 x 768 |
| W_V | 768 x 768 |
| W_O | 768 x 768 |
| Scores, per head | 8 x 8 |
| Out | 8 x 768 |
2,362,368 parameters per block, 28,348,416 across all 12, 23% of the model.
Where GPT-2 small's 124,439,808 parameters live
- Token and position embeddings39.4M, 32%
- Attention, all blocks28.3M, 23%
- MLP, all blocks56.7M, 46%
- Norms38.4K, under 0.1%
Try this
- With GPT-2 small, select each part from top to bottom and read its shapes. Which ones mention , and which do not?
- Drag the sequence length from 8 to 1,024. The parameter count does not move, because weights never depend on how many tokens you feed in. Select attention: its score matrix grows as .
- Switch to Llama 3 8B and select attention. Compare the shape of with , then read why they differ.
- Compare the breakdown bar for GPT-2 small with Llama 3 70B. The embedding tables are almost a third of the small model and about 3% of the large one, counting the separate unembedding. Where did the rest go?
The residual stream
Look at the diagram again: the blocks hang off a vertical line instead of sitting in a chain. That line is the Residual streamThe running vector for each token that flows through a transformer from the embedding to the output. Every attention and MLP sublayer reads from it and adds its result back.Open in glossary, the matrix that flows from the embedding to the unembedding. Each sublayer reads a normalized copy of the stream, computes something, and adds its result back:
Here is the stream entering the block and is the stream leaving it. Nothing is ever overwritten. Unrolled over all blocks, the final stream at a position is the token’s embedding plus the sum of every sublayer’s contribution at that position. Elhage et al. (2021) popularized the name “residual stream” and suggested picturing it as a shared channel: each layer reads what earlier layers wrote and writes its own additions for later layers to read.
This design solves a training problem. In a plain chain of 80 layers, the gradient has to pass through 80 transformations on its way back to the first layer and tends to shrink or explode along the way. The identity path in gives it a direct route: the derivative of with respect to is the identity plus the derivative of . The same trick, residual connections, first made very deep image networks trainable (He et al., 2016).
Normalization
Each sublayer starts by normalizing its input, so it sees vectors on a consistent scale no matter how large the stream has grown. GPT-2 uses Layer normalizationRescaling each token's vector to a standard mean and spread, then applying a learned scale (and, in LayerNorm, a shift). Keeps every sublayer's input on a stable scale.Open in glossary (Ba et al., 2016), which works on one token’s vector at a time:
where and are the mean and variance of the entries of , is a small constant for numerical safety, and and are learned vectors of length . Llama models use RMSNorm (Zhang and Sennrich, 2019), which drops the mean subtraction and the shift:
The MLP
The second sublayer is a small neural network, the Feed-forward networkThe MLP inside each transformer block. It processes every token's vector on its own, usually widening it about four times, applying a nonlinearity, and projecting back. It holds most of each block's parameters.Open in glossary or MLP, applied to every position separately. In GPT-2 it widens each vector by a factor of four, applies the GELU nonlinearity, and projects back:
Llama models use a gated variant called SwiGLU (Shazeer, 2020), with three matrices instead of two: , where multiplies elementwise.
The two sublayers divide the work. Attention moves information between positions; the MLP transforms the information at each position. Some interpretability research suggests that MLP layers hold much of a model’s stored knowledge, acting partly like lookup tables from patterns to associated outputs (Geva et al., 2021; Meng et al., 2022), though how knowledge is spread across a network is still actively studied.
The MLP also holds most of each block’s parameters. In a GPT-2 block, attention has about weights (the query, key, value, and output projections) and the MLP has about , so the MLP is two thirds of every block. As models grow wider, these terms dominate and the embedding tables shrink to a small share, which is what you saw in the breakdown bar.
Counting the parameters
With the pieces named, you can count GPT-2 small’s parameters from its configuration alone: vocabulary , context , width , and blocks.
The full count for GPT-2 smallOptional
| Part | Formula | Parameters |
|---|---|---|
| Token embedding | 38,597,376 | |
| Position embedding | 786,432 | |
| Attention, per block | 2,362,368 | |
| MLP, per block | 4,722,432 | |
| Two LayerNorms, per block | 3,072 | |
| One block | 7,087,872 | |
| Twelve blocks | 85,054,464 | |
| Final LayerNorm | 1,536 | |
| Unembedding | tied to the token embedding | 0 |
| Total | 124,439,808 |
The attention line counts four matrices with a bias vector each. The MLP line counts the and matrices with biases of length and . The total matches the released checkpoint exactly, which is a good check that nothing is missing. The GPT-2 paper itself listed this model as 117M parameters; the released checkpoint has the 124M counted here.
A useful shortcut from Kaplan et al. (2020): ignoring embeddings and the small bias and normalization terms, a transformer has about parameters. For GPT-2 small that gives 84.9 million, against an exact 85.1 million.
Key ideas
- A decoder-only transformer embeds tokens, applies the same kind of block times, and unembeds to logits. Every block maps to .
- The residual stream runs through the whole model. Each sublayer reads a normalized copy and adds its output back, which keeps gradients flowing in deep stacks.
- Attention is the only part that mixes positions. The MLP processes each position on its own and holds most of each block’s parameters.
- Normalization keeps each sublayer’s input on a stable scale. Modern models normalize before each sublayer (pre-norm).
- Parameter counts follow from the configuration: about plus the embedding tables.