Frontiers
How image generators turn noise into pictures, how one model learns both images and text, and how language models use tools and think longer to solve harder problems.
- Level
- Advanced
- Length
- 5 lessons, 1 hr 29 min
- Assumes
- The transformers track.
- Builds on
- Transformers and LLMs
- 1Diffusion: from noise to dataImage generators start from pure noise and remove it a little at a time. Watch the exact reverse process turn random dots into a dataset, and see what a real model has to learn.22 min
- 2One space for images and textCLIP learned to place a picture and a sentence that describes it at nearly the same point in one embedding space. Train a tiny version and see why that enables zero-shot classification.18 min
- 3Mixture of expertsSome of the largest language models run only a fraction of their weights for each token. See how a router picks experts, and why total and active parameter counts differ so much.16 min
- 4Thinking longerLanguage models often answer better when they write out intermediate steps, sample several attempts, or are trained to reason at length. See exactly when voting over attempts helps and when it hurts.18 min
- 5Models that use toolsA language model only produces text, yet it can check the weather, search, or run code. Step through the loop that connects a model to tools, and see where it fails and where safety controls belong.15 min
Labs in this track
- Contrastive learning, CLIP styleTrain an image encoder and a text encoder together until each picture lands next to its caption in one shared space.
- Diffusion from noise to dataWatch noise turn into a 2D dataset using the exact reverse diffusion process, with the score field drawn as arrows.
- Mixture of expertsSee a router send each token to its top experts, then compare total and active parameter counts for a Mixtral-shaped model.
- The tool-use loopStep through how an application and a model trade messages to use tools, and switch on failures to see how the loop copes.
- Voting over sampled answersCompute exactly how majority voting over n sampled answers changes accuracy, and see when it helps and when it hurts.