In the first track you saw a computer learn from examples. This track opens up how that learning actually happens. The surprising part is how little machinery is involved: the same short loop trains a straight line through a few points and a language model with hundreds of billions of numbers inside it.
Learning from answers
The most common kind of machine learning is Supervised learningLearning from examples that come with the right answer attached, such as photos labeled "cat" or houses with their sale prices.Open in glossary. You have examples of inputs, each with the right answer attached, and you want a rule that produces the right answer for new inputs too.
When the answer is a number, such as a price or a temperature, the task is called RegressionA prediction task where the answer is a number, such as a price, a temperature, or a time.Open in glossary. When it is a category, such as spam or not spam, it is called ClassificationA prediction task where the answer is one of a fixed set of categories, such as spam or not spam, or which digit a drawing shows.Open in glossary. This lesson uses regression because it is the easiest to see.
Three ingredients
Every supervised learning setup has the same three parts.
A model. A formula with some adjustable numbers in it, called ParameterA number inside a model that training adjusts, such as the slope of a line or a weight in a neural network. Large language models have billions of them.Open in glossary. Our model is a straight line:
Here is the input, (read “y-hat”) is the model’s guess, is the slope, and is where the line crosses the vertical axis. Learning means finding good values for and .
A loss. A Loss functionA formula that turns a model's mistakes on the training data into a single number. Lower is better, and training tries to make it as low as possible.Open in glossary turns all of the model’s mistakes into a single number, so that “better” has a precise meaning. We will use the Mean squared errorThe average of the squared differences between predictions and true values. Squaring makes big misses count much more than small ones.Open in glossary: for each training point, take the miss , square it, and average over all points.
An optimizer. An OptimizerThe rule that decides how to change a model's parameters at each training step, given the gradient. Gradient descent, momentum, and Adam are optimizers.Open in glossary is the rule for changing the parameters to lower the loss. Ours is Gradient descentAn optimization method that repeatedly moves the parameters a small step in the direction that lowers the loss fastest, the opposite of the gradient.Open in glossary: work out which way each parameter should move to reduce the loss, then move it a small step that way. The size of that step is the Learning rateThe step size in gradient descent. Too small and training crawls; too large and each step overshoots, so the loss can grow instead of shrink.Open in glossary.
Run the loop
Training is these ingredients in a cycle: predict, measure, find the direction, update, and again. The demo below starts with a line that is badly wrong. Step through a round slowly first, then let it run.
The learning loop
A model (the line) improves by repeating four steps. Step through one round at a time, or let it run.
- 1PredictUse the current line to guess y for every training x.
- 2Measure the lossAverage the squared misses into one number.
- 3Find the directionWork out which way to nudge w and b to lower the loss.
- 4UpdateMove w and b a small step that way. Repeat.
- Training points (used to learn)
- Test points (held back)
- Current line
The line starts out wrong. Press Next step to begin the first round.
Try this
- Press Next step four times and read the note under the plot after each press. You have just done one round of training by hand.
- Press Run rounds. In the chart, the training loss drops steeply at first and then levels off as the line settles.
- Reset, then drag the step size down to 0.02 and run. The line still learns, but it crawls.
- Reset, drag the step size to the top (about 1.6) and run. Each update overshoots further than the last, and the loss explodes.
The line never sees the rule that generated the points. It only ever sees its own misses on the training data, and it keeps nudging and in whichever direction makes those misses smaller. That is all “learning” means here.
The step size matters more than it looks. Too small and you waste rounds. Too large and every update jumps past the best values, which is why the loss can grow instead of shrink. You will see exactly why in the gradient descent lesson.
Where the update comes fromOptional
To lower , we need to know how changes when or change: its slopes with respect to each parameter. For mean squared error they are
If is positive, increasing would increase the loss, so we decrease it. The update with learning rate (eta) is
These two slopes together are the gradient of the loss. The demo computes exactly these formulas each round.
Training data and test data
Look at the two lines in the chart. The solid one is the loss on the training points, the circles the model learns from. The dashed one is the loss on the triangles, which the model never learns from. They form a Test setExamples held back from training and used only to measure how well a model does on data it has never seen.Open in glossary.
Why hold data back? Because doing well on the training data is not the goal. A model can do very well on examples it has already seen and still fail on new ones. The test loss is your honest estimate of how the model will do on data it has never encountered.
In this demo the two curves stay close, because a straight line is too simple to memorize twelve points. In the overfitting lesson you will give the model enough flexibility to memorize, and watch the two curves split apart.
The same loop, at every scale
A large language model is trained with this same loop. The model is a neural network with billions of parameters instead of two. The loss measures how well it predicts the next word piece in real text. The optimizer is a refined version of gradient descent. The main practical difference is that computing the loss on all the training data every round would be far too slow, so each round uses a small random batch of examples. That variant is called Stochastic gradient descentGradient descent where each step uses a small random batch of training examples instead of all of them. Steps are noisier but far cheaper, which is how large models are trained.Open in glossary, and it is how almost every modern model is trained.
Key ideas
- Supervised learning starts from examples with known answers. Regression predicts numbers; classification predicts categories.
- Every setup has a model with parameters, a loss that scores its mistakes, and an optimizer that adjusts the parameters.
- Training is a loop: predict, measure the loss, find the direction that lowers it, take a small step, repeat.
- The step size controls the pace. Too large and the loss can explode instead of shrink.
- Test data is held back to measure performance on examples the model has never seen.