Lesson 1 of 8

The learning loop

Every trained model, from a straight line to a chatbot, improves by repeating the same four steps. Run the loop yourself and watch a line learn.

Beginner14 min

In this lesson you will

  • Name the three ingredients of supervised learning, a model, a loss, and an optimizer
  • Step through one round of training and say what each step does
  • Explain why some data is held back for testing and what test loss tells you

In the first track you saw a computer learn from examples. This track opens up how that learning actually happens. The surprising part is how little machinery is involved: the same short loop trains a straight line through a few points and a language model with hundreds of billions of numbers inside it.

Learning from answers

The most common kind of machine learning is Supervised learningLearning from examples that come with the right answer attached, such as photos labeled "cat" or houses with their sale prices.Open in glossary. You have examples of inputs, each with the right answer attached, and you want a rule that produces the right answer for new inputs too.

When the answer is a number, such as a price or a temperature, the task is called RegressionA prediction task where the answer is a number, such as a price, a temperature, or a time.Open in glossary. When it is a category, such as spam or not spam, it is called ClassificationA prediction task where the answer is one of a fixed set of categories, such as spam or not spam, or which digit a drawing shows.Open in glossary. This lesson uses regression because it is the easiest to see.

Three ingredients

Every supervised learning setup has the same three parts.

A model. A formula with some adjustable numbers in it, called ParameterA number inside a model that training adjusts, such as the slope of a line or a weight in a neural network. Large language models have billions of them.Open in glossary. Our model is a straight line:

y^=w x+b\hat{y} = w\,x + b

Here xx is the input, y^\hat{y} (read “y-hat”) is the model’s guess, ww is the slope, and bb is where the line crosses the vertical axis. Learning means finding good values for ww and bb.

A loss. A Loss functionA formula that turns a model's mistakes on the training data into a single number. Lower is better, and training tries to make it as low as possible.Open in glossary turns all of the model’s mistakes into a single number, so that “better” has a precise meaning. We will use the Mean squared errorThe average of the squared differences between predictions and true values. Squaring makes big misses count much more than small ones.Open in glossary: for each training point, take the miss y^i−yi\hat{y}_i - y_i, square it, and average over all nn points.

L(w,b)=1n∑i=1n(y^i−yi)2L(w, b) = \frac{1}{n} \sum_{i=1}^{n} \left(\hat{y}_i - y_i\right)^2

An optimizer. An OptimizerThe rule that decides how to change a model's parameters at each training step, given the gradient. Gradient descent, momentum, and Adam are optimizers.Open in glossary is the rule for changing the parameters to lower the loss. Ours is Gradient descentAn optimization method that repeatedly moves the parameters a small step in the direction that lowers the loss fastest, the opposite of the gradient.Open in glossary: work out which way each parameter should move to reduce the loss, then move it a small step that way. The size of that step is the Learning rateThe step size in gradient descent. Too small and training crawls; too large and each step overshoots, so the loss can grow instead of shrink.Open in glossary.

Run the loop

Training is these ingredients in a cycle: predict, measure, find the direction, update, and again. The demo below starts with a line that is badly wrong. Step through a round slowly first, then let it run.

The learning loop

A model (the line) improves by repeating four steps. Step through one round at a time, or let it run.

  1. 1PredictUse the current line to guess y for every training x.
  2. 2Measure the lossAverage the squared misses into one number.
  3. 3Find the directionWork out which way to nudge w and b to lower the loss.
  4. 4UpdateMove w and b a small step that way. Repeat.
  • Training points (used to learn)
  • Test points (held back)
  • Current line

The line starts out wrong. Press Next step to begin the first round.

0.50

How far each update moves w and b.

Round0
Training loss0.1007
Test loss0.0858
Slope w-0.5
Intercept b0.85

Try this

  • Press Next step four times and read the note under the plot after each press. You have just done one round of training by hand.
  • Press Run rounds. In the chart, the training loss drops steeply at first and then levels off as the line settles.
  • Reset, then drag the step size down to 0.02 and run. The line still learns, but it crawls.
  • Reset, drag the step size to the top (about 1.6) and run. Each update overshoots further than the last, and the loss explodes.

The line never sees the rule that generated the points. It only ever sees its own misses on the training data, and it keeps nudging ww and bb in whichever direction makes those misses smaller. That is all “learning” means here.

The step size matters more than it looks. Too small and you waste rounds. Too large and every update jumps past the best values, which is why the loss can grow instead of shrink. You will see exactly why in the gradient descent lesson.

Where the update comes fromOptional

To lower LL, we need to know how LL changes when ww or bb change: its slopes with respect to each parameter. For mean squared error they are

∂L∂w=2n∑i=1n(y^i−yi)xi,∂L∂b=2n∑i=1n(y^i−yi).\frac{\partial L}{\partial w} = \frac{2}{n}\sum_{i=1}^{n}\left(\hat{y}_i - y_i\right)x_i, \qquad \frac{\partial L}{\partial b} = \frac{2}{n}\sum_{i=1}^{n}\left(\hat{y}_i - y_i\right).

If ∂L/∂w\partial L / \partial w is positive, increasing ww would increase the loss, so we decrease it. The update with learning rate η\eta (eta) is

w←w−η ∂L∂w,b←b−η ∂L∂b.w \leftarrow w - \eta\,\frac{\partial L}{\partial w}, \qquad b \leftarrow b - \eta\,\frac{\partial L}{\partial b}.

These two slopes together are the gradient of the loss. The demo computes exactly these formulas each round.

Training data and test data

Look at the two lines in the chart. The solid one is the loss on the training points, the circles the model learns from. The dashed one is the loss on the triangles, which the model never learns from. They form a Test setExamples held back from training and used only to measure how well a model does on data it has never seen.Open in glossary.

Why hold data back? Because doing well on the training data is not the goal. A model can do very well on examples it has already seen and still fail on new ones. The test loss is your honest estimate of how the model will do on data it has never encountered.

In this demo the two curves stay close, because a straight line is too simple to memorize twelve points. In the overfitting lesson you will give the model enough flexibility to memorize, and watch the two curves split apart.

The same loop, at every scale

A large language model is trained with this same loop. The model is a neural network with billions of parameters instead of two. The loss measures how well it predicts the next word piece in real text. The optimizer is a refined version of gradient descent. The main practical difference is that computing the loss on all the training data every round would be far too slow, so each round uses a small random batch of examples. That variant is called Stochastic gradient descentGradient descent where each step uses a small random batch of training examples instead of all of them. Steps are noisier but far cheaper, which is how large models are trained.Open in glossary, and it is how almost every modern model is trained.

Key ideas

  • Supervised learning starts from examples with known answers. Regression predicts numbers; classification predicts categories.
  • Every setup has a model with parameters, a loss that scores its mistakes, and an optimizer that adjusts the parameters.
  • Training is a loop: predict, measure the loss, find the direction that lowers it, take a small step, repeat.
  • The step size controls the pace. Too large and the loss can explode instead of shrink.
  • Test data is held back to measure performance on examples the model has never seen.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1Which of these is the loss in the demo?
2Why are the test points never used to update the line?
3You raise the step size and the training loss starts growing every round. What is happening?

Progress is saved in this browser only.

Up nextLinear regression and loss
Next
How machines learn
  1. 1The learning loop
  2. 2Linear regression and loss
  3. 3Gradient descent
  4. 4Better optimizers
  5. 5Classification and probability
  6. 6Overfitting and generalization
  7. 7Measuring a classifier
  8. 8Finding groups: k-means

Try "embedding", "softmax", "overfitting", or "backpropagation".