Transformers and LLMs
Take a transformer apart. Compute attention by hand, see how position is encoded, trace a token through every layer, and run a real language model in your browser to inspect its predictions.
- Level
- Advanced
- Length
- 8 lessons, 2 hr 21 min
- Assumes
- Neural networks and embeddings, or equivalent background.
- Builds on
- Neural networks and Language as vectors
- 1Predicting the next tokenAt every step a language model scores every token in its vocabulary. Softmax turns the scores into probabilities, and a sampling rule picks one. Temperature, top-k, and top-p shape that choice.16 min
- 2Attention, step by stepAttention lets every token gather information from the other tokens in its context. Work through it one query at a time, from scores to weights to a mixed output.20 min
- 3Many heads and the causal maskTransformers run several attention heads side by side, each free to learn its own pattern, and language models hide the future from every token with a causal mask.18 min
- 4Where words areAttention ignores word order unless the model is told where each token sits. Sinusoidal encodings add position as waves; rotary embeddings rotate queries and keys so attention sees relative distance.18 min
- 5The transformer blockA transformer is one block repeated many times. Each block reads from a shared residual stream, mixes information across tokens with attention, transforms each token with an MLP, and adds the results back.18 min
- 6Inside a real language modelRun SmolLM2-135M in your browser. Watch it split your text into tokens, score every token in its vocabulary, and write one token at a time.15 min
- 7How LLMs are trainedPretraining on trillions of tokens builds a base model; fine-tuning and preference tuning turn it into an assistant. Scaling laws tell you how to split a compute budget between model size and data.20 min
- 8Generating fast: the KV cacheEach new token needs the keys and values of every token before it. Storing them instead of recomputing them makes generation fast, and the store grows with every token.16 min
Labs in this track
- A live language modelRun SmolLM2-135M in your browser, see how it tokenizes your text, read its real next-token probabilities, and generate one token at a time.
- Attention calculatorDrag queries, keys, and values on a plane and watch scaled dot-product attention compute scores, weights, and outputs for every token.
- Attention heads and the causal maskCompare hand-built versions of five attention patterns seen in trained transformers, switch the causal mask on and off, and see how heads split the model width.
- Compute and scaling lawsChoose a model size and a number of training tokens, see the compute it costs, and compare it with the compute-optimal split under two fitted scaling laws.
- Position encodingsExplore the sine and cosine waves of the original transformer's position encodings, then rotate queries and keys with RoPE and watch their dot product depend on position only through the offset.
- Sampling the next tokenReshape a next-token distribution with temperature, top-k, and top-p, then draw from it and compare the counts.
- The KV cacheStep through generation with and without a key-value cache, then size the cache for real models, context lengths, and batches.
- Transformer anatomySelect any part of a decoder-only transformer to see its job, its tensor shapes, and its parameter count, for models from GPT-2 to Llama 3.