Lesson 5 of 8

The transformer block

A transformer is one block repeated many times. Each block reads from a shared residual stream, mixes information across tokens with attention, transforms each token with an MLP, and adds the results back.

Advanced18 min

In this lesson you will

  • Trace a token from its ID to the logits, with the tensor shape at every step
  • Explain the residual stream and why each sublayer adds to it instead of replacing it
  • Say what normalization and the MLP contribute, and where a model's parameters live
  • Count the parameters of GPT-2 small from its configuration

You have now met the two ideas that make a TransformerA neural network architecture built from stacked layers of attention and feed-forward networks, introduced by Vaswani et al. in 2017. Most modern language models are transformers.Open in glossary work: AttentionA mechanism that lets each token build a new vector as a weighted mix of other tokens' value vectors, with weights set by how well its query matches their keys.Open in glossary, which lets tokens look at each other, and position encodings, which tell them where they are. A language model wraps those ideas in a block and repeats the block, 12 times in GPT-2 small and 80 times in Llama 3 70B. This lesson takes one block apart and then puts the whole model back together.

From token IDs to logits

Every decoder-only language model, from GPT-2 to Llama 3, has the same outline:

  1. Embed. Each token ID selects a row of a learned table, giving a vector of dd numbers per token. For TT tokens that is a T×dT \times d matrix.
  2. Repeat the block LL times. Each block takes a T×dT \times d matrix and returns one of the same shape.
  3. Unembed. A final normalization, then a matrix multiplication turns each position’s dd numbers into VV scores, one per vocabulary token. These are the logits from the first lesson of this track.

The second step is what makes depth possible. Because a block’s input and output have the same shape, block 7 does not care whether it follows block 6 or block 1, and adding more blocks is a one-line change to the configuration.

Anatomy of a transformer

Select any part to see what it does, its tensor shapes, and how many parameters it holds.

Parts of GPT-2 small

Residual stream, 8 x 768

Block 1 of 12

11 more blocks with the same shape and their own weights

Multi-head self-attention

The only place where tokens exchange information. Each of 12 heads projects every token to a query, key, and value of size 64, scores each query against the keys of the tokens at or before it (the causal mask), and takes a softmax-weighted sum of their values. The heads' outputs are joined and projected back to 768 numbers by W_O, then added to the stream.

Shapes for Multi-head self-attention
W_Q768 x 768
W_K768 x 768
W_V768 x 768
W_O768 x 768
Scores, per head8 x 8
Out8 x 768

2,362,368 parameters per block, 28,348,416 across all 12, 23% of the model.

Where GPT-2 small's 124,439,808 parameters live

  • Token and position embeddings39.4M, 32%
  • Attention, all blocks28.3M, 23%
  • MLP, all blocks56.7M, 46%
  • Norms38.4K, under 0.1%
Model
8 tokens

Up to this model's context of 1,024.

Parameters124M
Without embeddings85.1M
Forward FLOPs per token at T = 80.25 G

Try this

  • With GPT-2 small, select each part from top to bottom and read its shapes. Which ones mention TT, and which do not?
  • Drag the sequence length from 8 to 1,024. The parameter count does not move, because weights never depend on how many tokens you feed in. Select attention: its score matrix grows as T×TT \times T.
  • Switch to Llama 3 8B and select attention. Compare the shape of WKW_K with WQW_Q, then read why they differ.
  • Compare the breakdown bar for GPT-2 small with Llama 3 70B. The embedding tables are almost a third of the small model and about 3% of the large one, counting the separate unembedding. Where did the rest go?

The residual stream

Look at the diagram again: the blocks hang off a vertical line instead of sitting in a chain. That line is the Residual streamThe running vector for each token that flows through a transformer from the embedding to the output. Every attention and MLP sublayer reads from it and adds its result back.Open in glossary, the T×dT \times d matrix that flows from the embedding to the unembedding. Each sublayer reads a normalized copy of the stream, computes something, and adds its result back:

h=x+Attn(Norm(x))x′=h+MLP(Norm(h))\begin{aligned} h &= x + \mathrm{Attn}\big(\mathrm{Norm}(x)\big) \\ x' &= h + \mathrm{MLP}\big(\mathrm{Norm}(h)\big) \end{aligned}

Here xx is the stream entering the block and x′x' is the stream leaving it. Nothing is ever overwritten. Unrolled over all LL blocks, the final stream at a position is the token’s embedding plus the sum of every sublayer’s contribution at that position. Elhage et al. (2021) popularized the name “residual stream” and suggested picturing it as a shared channel: each layer reads what earlier layers wrote and writes its own additions for later layers to read.

This design solves a training problem. In a plain chain of 80 layers, the gradient has to pass through 80 transformations on its way back to the first layer and tends to shrink or explode along the way. The identity path in x+f(x)x + f(x) gives it a direct route: the derivative of x+f(x)x + f(x) with respect to xx is the identity plus the derivative of ff. The same trick, residual connections, first made very deep image networks trainable (He et al., 2016).

Normalization

Each sublayer starts by normalizing its input, so it sees vectors on a consistent scale no matter how large the stream has grown. GPT-2 uses Layer normalizationRescaling each token's vector to a standard mean and spread, then applying a learned scale (and, in LayerNorm, a shift). Keeps every sublayer's input on a stable scale.Open in glossary (Ba et al., 2016), which works on one token’s vector at a time:

LayerNorm(x)=γ⊙x−μσ2+ϵ+β\mathrm{LayerNorm}(x) = \gamma \odot \frac{x - \mu}{\sqrt{\sigma^2 + \epsilon}} + \beta

where μ\mu and σ2\sigma^2 are the mean and variance of the dd entries of xx, ϵ\epsilon is a small constant for numerical safety, and γ\gamma and β\beta are learned vectors of length dd. Llama models use RMSNorm (Zhang and Sennrich, 2019), which drops the mean subtraction and the shift:

RMSNorm(x)=γ⊙x1d∑ixi2+ϵ\mathrm{RMSNorm}(x) = \gamma \odot \frac{x}{\sqrt{\tfrac{1}{d}\sum_i x_i^2 + \epsilon}}

The MLP

The second sublayer is a small neural network, the Feed-forward networkThe MLP inside each transformer block. It processes every token's vector on its own, usually widening it about four times, applying a nonlinearity, and projecting back. It holds most of each block's parameters.Open in glossary or MLP, applied to every position separately. In GPT-2 it widens each vector by a factor of four, applies the GELU nonlinearity, and projects back:

MLP(x)=W2 GELU(W1x+b1)+b2,W1∈R4d×d,  W2∈Rd×4d\mathrm{MLP}(x) = W_2\,\mathrm{GELU}(W_1 x + b_1) + b_2, \qquad W_1 \in \mathbb{R}^{4d \times d},\; W_2 \in \mathbb{R}^{d \times 4d}

Llama models use a gated variant called SwiGLU (Shazeer, 2020), with three matrices instead of two: W2(SiLU(Wgx)⊙W1x)W_2\big(\mathrm{SiLU}(W_g x) \odot W_1 x\big), where ⊙\odot multiplies elementwise.

The two sublayers divide the work. Attention moves information between positions; the MLP transforms the information at each position. Some interpretability research suggests that MLP layers hold much of a model’s stored knowledge, acting partly like lookup tables from patterns to associated outputs (Geva et al., 2021; Meng et al., 2022), though how knowledge is spread across a network is still actively studied.

The MLP also holds most of each block’s parameters. In a GPT-2 block, attention has about 4d24d^2 weights (the query, key, value, and output projections) and the MLP has about 8d28d^2, so the MLP is two thirds of every block. As models grow wider, these d2d^2 terms dominate and the embedding tables shrink to a small share, which is what you saw in the breakdown bar.

Counting the parameters

With the pieces named, you can count GPT-2 small’s parameters from its configuration alone: vocabulary V=50,257V = 50{,}257, context 1,0241{,}024, width d=768d = 768, and L=12L = 12 blocks.

The full count for GPT-2 smallOptional
PartFormulaParameters
Token embeddingV×dV \times d38,597,376
Position embedding1024×d1024 \times d786,432
Attention, per block4d2+4d4d^2 + 4d2,362,368
MLP, per block8d2+5d8d^2 + 5d4,722,432
Two LayerNorms, per block4d4d3,072
One block12d2+13d12d^2 + 13d7,087,872
Twelve blocks85,054,464
Final LayerNorm2d2d1,536
Unembeddingtied to the token embedding0
Total124,439,808

The attention line counts four d×dd \times d matrices with a bias vector each. The MLP line counts the d×4dd \times 4d and 4d×d4d \times d matrices with biases of length 4d4d and dd. The total matches the released checkpoint exactly, which is a good check that nothing is missing. The GPT-2 paper itself listed this model as 117M parameters; the released checkpoint has the 124M counted here.

A useful shortcut from Kaplan et al. (2020): ignoring embeddings and the small bias and normalization terms, a transformer has about 12Ld212 L d^2 parameters. For GPT-2 small that gives 84.9 million, against an exact 85.1 million.

Key ideas

  • A decoder-only transformer embeds tokens, applies the same kind of block LL times, and unembeds to logits. Every block maps T×dT \times d to T×dT \times d.
  • The residual stream runs through the whole model. Each sublayer reads a normalized copy and adds its output back, which keeps gradients flowing in deep stacks.
  • Attention is the only part that mixes positions. The MLP processes each position on its own and holds most of each block’s parameters.
  • Normalization keeps each sublayer’s input on a stable scale. Modern models normalize before each sublayer (pre-norm).
  • Parameter counts follow from the configuration: about 12Ld212 L d^2 plus the embedding tables.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1Which update describes the attention sublayer of a pre-norm transformer block, as used in GPT-2 and Llama?
2Which part of a transformer block lets one token use information from another?
3GPT-2 small has a model width of 768. About how many parameters does one of its blocks hold?

Progress is saved in this browser only.

Up nextInside a real language model
Next
Transformers and LLMs
  1. 1Predicting the next token
  2. 2Attention, step by step
  3. 3Many heads and the causal mask
  4. 4Where words are
  5. 5The transformer block
  6. 6Inside a real language model
  7. 7How LLMs are trained
  8. 8Generating fast: the KV cache

Try "embedding", "softmax", "overfitting", or "backpropagation".