Frontiers

How image generators turn noise into pictures, how one model learns both images and text, and how language models use tools and think longer to solve harder problems.

Level
Advanced
Length
5 lessons, 1 hr 29 min
Assumes
The transformers track.
Builds on
Transformers and LLMs
  1. 1Diffusion: from noise to dataImage generators start from pure noise and remove it a little at a time. Watch the exact reverse process turn random dots into a dataset, and see what a real model has to learn.22 min
  2. 2One space for images and textCLIP learned to place a picture and a sentence that describes it at nearly the same point in one embedding space. Train a tiny version and see why that enables zero-shot classification.18 min
  3. 3Mixture of expertsSome of the largest language models run only a fraction of their weights for each token. See how a router picks experts, and why total and active parameter counts differ so much.16 min
  4. 4Thinking longerLanguage models often answer better when they write out intermediate steps, sample several attempts, or are trained to reason at length. See exactly when voting over attempts helps and when it hurts.18 min
  5. 5Models that use toolsA language model only produces text, yet it can check the weather, search, or run code. Step through the loop that connects a model to tools, and see where it fails and where safety controls belong.15 min

Labs in this track

Try "embedding", "softmax", "overfitting", or "backpropagation".