Transformers and LLMs

Take a transformer apart. Compute attention by hand, see how position is encoded, trace a token through every layer, and run a real language model in your browser to inspect its predictions.

Level
Advanced
Length
8 lessons, 2 hr 21 min
Assumes
Neural networks and embeddings, or equivalent background.
Builds on
Neural networks and Language as vectors
  1. 1Predicting the next tokenAt every step a language model scores every token in its vocabulary. Softmax turns the scores into probabilities, and a sampling rule picks one. Temperature, top-k, and top-p shape that choice.16 min
  2. 2Attention, step by stepAttention lets every token gather information from the other tokens in its context. Work through it one query at a time, from scores to weights to a mixed output.20 min
  3. 3Many heads and the causal maskTransformers run several attention heads side by side, each free to learn its own pattern, and language models hide the future from every token with a causal mask.18 min
  4. 4Where words areAttention ignores word order unless the model is told where each token sits. Sinusoidal encodings add position as waves; rotary embeddings rotate queries and keys so attention sees relative distance.18 min
  5. 5The transformer blockA transformer is one block repeated many times. Each block reads from a shared residual stream, mixes information across tokens with attention, transforms each token with an MLP, and adds the results back.18 min
  6. 6Inside a real language modelRun SmolLM2-135M in your browser. Watch it split your text into tokens, score every token in its vocabulary, and write one token at a time.15 min
  7. 7How LLMs are trainedPretraining on trillions of tokens builds a base model; fine-tuning and preference tuning turn it into an assistant. Scaling laws tell you how to split a compute budget between model size and data.20 min
  8. 8Generating fast: the KV cacheEach new token needs the keys and values of every token before it. Storing them instead of recomputing them makes generation fast, and the store grows with every token.16 min

Labs in this track

Try "embedding", "softmax", "overfitting", or "backpropagation".