Suppose you want to predict an apartment’s rent from its size, or a runner’s finish time from their training miles. You have a pile of past examples, each a pair of numbers. Plotted, they lean in a direction. The simplest useful model is a straight line through them, and fitting that line is called Linear regressionPredicting a number with a weighted sum of the inputs plus a constant. With one input, it is fitting a straight line.Open in glossary.
What makes a line good?
Any line makes a guess for every . For each example, the miss is the guess minus the true value, . Some misses are positive (the line is too high) and some negative (too low).
You cannot just add the misses up, because they cancel: a line far above half the points and far below the other half can total zero. So every miss is made positive first, in one of two common ways.
- Square it. Averaging the squares gives the Mean squared errorThe average of the squared differences between predictions and true values. Squaring makes big misses count much more than small ones.Open in glossary (MSE).
- Take its size, ignoring the sign. Averaging those gives the Mean absolute errorThe average size of the differences between predictions and true values, ignoring sign. Less sensitive to a few wild outliers than mean squared error.Open in glossary (MAE).
The squared error has a picture: each squared miss is the area of a square whose side is the miss. The MSE is the average area of those squares.
Fit it yourself
Use the sliders to move the line. Try to make the orange squares as small as possible overall.
Fitting a line
Drag points to move them, tap empty space to add one, tap a point to remove it.
- Data points
- Squared misses (area)
- Your line
Try this
- Adjust the slope and intercept until the mean squared error stops going down. The readout tells you when you have reached the best possible value.
- Notice how one big square can outweigh several small ones. Squaring makes large misses dominate.
- Drag a point far from the line and watch its square grow much faster than the distance.
- Switch to Best fit to see the exact answer.
You were doing what an optimizer does: change the parameters, watch the loss, keep the changes that lower it. For a straight line under squared error there is also a shortcut. Calculus gives a formula for the best slope and intercept directly, which is what Best fit uses.
The least-squares formulaOptional
Write and for the averages of the and values. The line that minimizes the mean squared error has
You get these by setting the derivatives of the MSE with respect to and to zero and solving. The best line always passes through the point .
Absolute error has no such simple formula, because has a sharp corner at zero. This demo finds the best absolute-error line exactly by checking every line through two of the points: some best line always passes through at least two of them.
Which loss? Outliers decide
The two losses agree when the data is well behaved. They disagree sharply about OutlierA data point far from the pattern of the rest, from a rare event or a recording error. Some loss functions are pulled hard by outliers.Open in glossary, points far from the general pattern, such as a typo in the data or a genuinely rare event.
Fitting a line
Drag points to move them, tap empty space to add one, tap a point to remove it.
- Data points
- Squared misses (area)
- Best line for squared error
- Best line for absolute error
Try this
- Press Add an outlier. The solid line (best for squared error) tilts noticeably toward the new point. The dashed line (best for absolute error) barely moves.
- Switch the loss to Absolute error. Now the solid line is the robust one, and the squared-error line is dashed.
- Drag the outlier further away and compare how each line responds.
A miss of 6 costs 36 under squared error but only 6 under absolute error, so under squared error a single distant point can be worth more than all the others combined. The line bends toward it to shrink that one enormous square.
Neither loss is simply better. Squared error is smooth everywhere, which makes the optimization easy, and it is the right choice when large errors really are much worse than small ones. Absolute error is more robust when the data may contain a few bad values. Choosing a loss is choosing what “wrong” means for your problem.
More than one input
Real predictions use many inputs: size, location, age, and number of bedrooms, not just size. Linear regression extends directly with one weight per input:
With two inputs the model is a flat plane instead of a line, and with more it is a flat surface in higher dimensions that you cannot draw. The loss and the fitting work exactly the same way.
Key ideas
- Linear regression fits a straight line (or flat plane) by choosing the slope and intercept.
- Misses are squared or made positive before averaging, so they cannot cancel.
- Squared error lets big misses dominate, so outliers pull the fit hard. Absolute error is more robust to them.
- For squared error, the best line has an exact formula. In general, models are fit by searching for parameters that lower the loss.