Imagine you run an ice cream stand. On hot days you sell more, on cool days fewer, and you would like to predict tomorrow’s sales from the weather forecast so you buy the right amount of cream. You have notes from sixteen past days: the temperature and how many cones you sold.
The simplest prediction is a straight line: the hotter it is, the more you sell. But which line? There are infinitely many. To pick one, you first need a way to say how good a line is.
Measuring a miss
For any line, look at each past day. The line predicts some number of cones for that day’s temperature, and you know how many you really sold. The difference is the miss for that day.
One miss per day is a lot to keep track of. To compare lines, we want a single number. Adding up the misses does not work: a line that predicts 20 too many on one day and 20 too few on another would add up to zero, as if it were perfect.
The standard fix is to square each miss before averaging. Squaring makes every miss positive, so they cannot cancel. It also makes big misses count extra: a miss of 20 counts four times as much as a miss of 10. The result is called the Mean squared errorThe average of the squared differences between predictions and true values. Squaring makes big misses count much more than small ones.Open in glossary. Take its square root and you are back in cones, which gives you a “typical miss” you can read directly.
A number that measures how wrong a model is, like this one, is called a Loss functionA formula that turns a model's mistakes on the training data into a single number. Lower is better, and training tries to make it as low as possible.Open in glossary. Lower is better.
Fit a line to the ice cream sales
Each dot is one day: the temperature and how many cones sold. Drag the two handles to place a line that predicts sales from temperature.
Try this
- With Show squared errors on, each miss is drawn as a square. Drag the handles and watch the squares shrink and grow. The loss is proportional to the average area of those squares.
- Try to get your line’s typical miss within 2 cones of the best possible. It is harder than it looks.
- Notice how one far-away day makes a huge square. Squaring means the line works hardest to avoid big misses.
- Press Let the computer fit it and watch it make small adjustments until it reaches the best line.
Learning is getting less wrong, one step at a time
When you dragged the handles, you were doing what the computer does, only by feel. The computer has no eyes, so it does something methodical. At each step it asks: if I tilted the line slightly, or moved it up or down slightly, would the error go up or down? Then it moves the line a small amount in whichever direction lowers the error, and asks again.
This is called Gradient descentAn optimization method that repeatedly moves the parameters a small step in the direction that lowers the loss fastest, the opposite of the gradient.Open in glossary, and it is how almost every modern AI model learns, including the ones behind chatbots. The model there has billions of adjustable numbers instead of two, but the loop is the same:
- Make predictions with the current settings.
- Measure the loss.
- Nudge every setting a little in the direction that lowers the loss.
- Repeat, often thousands or even millions of times.
The machine learning track looks at this loop in detail, including what can go wrong when the steps are too big or too small.
The best line still misses
Even the best line misses by about 18 cones on a typical day. That is not a failure. Sales depend on more than temperature: weekends, rain, a school trip passing by. A straight line through temperature alone cannot know about any of that.
You could draw a wiggly curve that passes through every single point, with zero error on these sixteen days. It would be useless for tomorrow, because it would be following the accidents of those particular days rather than the real pattern. A good model captures the trend and accepts some error on the examples it learned from.
The formula for mean squared errorOptional
Call the temperatures and the cones sold . A line predicts , where is the slope and is where the line crosses zero degrees. The mean squared error is
and the “typical miss” in the demo is , called the root mean squared error.
For a straight line there is a formula that gives the best and directly (ordinary least squares). The demo uses step-by-step gradient descent instead because that is the method that works for every kind of model, including neural networks where no such formula exists. To keep the animation short, its steps are scaled to the spread of the temperatures, which lets it converge in a couple of seconds instead of thousands of tiny steps.
Key ideas
- To improve a model you first need a single number that says how wrong it is: the loss.
- Mean squared error averages the squared misses, so misses cannot cancel and big misses count extra.
- Gradient descent improves a model by repeatedly nudging its settings in the direction that lowers the loss.
- The best model still makes errors on its examples. Chasing zero error usually means memorizing noise.