Lesson 2 of 8

Linear regression and loss

What makes one line fit better than another? Measure the misses, find the best line by hand, and see how the choice of loss decides which line wins.

Beginner15 min

In this lesson you will

  • Explain why misses are squared or made positive before they are averaged
  • Fit a line by hand by watching the loss, then compare it with the exact best fit
  • Predict how squared error and absolute error react to an outlier

Suppose you want to predict an apartment’s rent from its size, or a runner’s finish time from their training miles. You have a pile of past examples, each a pair of numbers. Plotted, they lean in a direction. The simplest useful model is a straight line through them, and fitting that line is called Linear regressionPredicting a number with a weighted sum of the inputs plus a constant. With one input, it is fitting a straight line.Open in glossary.

What makes a line good?

Any line makes a guess y^\hat{y} for every xx. For each example, the miss is the guess minus the true value, y^i−yi\hat{y}_i - y_i. Some misses are positive (the line is too high) and some negative (too low).

You cannot just add the misses up, because they cancel: a line far above half the points and far below the other half can total zero. So every miss is made positive first, in one of two common ways.

  • Square it. Averaging the squares gives the Mean squared errorThe average of the squared differences between predictions and true values. Squaring makes big misses count much more than small ones.Open in glossary (MSE).
  • Take its size, ignoring the sign. Averaging those gives the Mean absolute errorThe average size of the differences between predictions and true values, ignoring sign. Less sensitive to a few wild outliers than mean squared error.Open in glossary (MAE).
MSE=1n∑i=1n(y^i−yi)2MAE=1n∑i=1n∣y^i−yi∣\text{MSE} = \frac{1}{n}\sum_{i=1}^{n}\left(\hat{y}_i - y_i\right)^2 \qquad \text{MAE} = \frac{1}{n}\sum_{i=1}^{n}\left|\hat{y}_i - y_i\right|

The squared error has a picture: each squared miss is the area of a square whose side is the miss. The MSE is the average area of those squares.

Fit it yourself

Use the sliders to move the line. Try to make the orange squares as small as possible overall.

Fitting a line

Drag points to move them, tap empty space to add one, tap a point to remove it.

  • Data points
  • Squared misses (area)
  • Your line
Line
Loss
0.20
4.50
Mean squared error0.998
Mean absolute error0.781
Best possible MSE0.41
Liney = 0.2x + 4.5

Try this

  • Adjust the slope and intercept until the mean squared error stops going down. The readout tells you when you have reached the best possible value.
  • Notice how one big square can outweigh several small ones. Squaring makes large misses dominate.
  • Drag a point far from the line and watch its square grow much faster than the distance.
  • Switch to Best fit to see the exact answer.

You were doing what an optimizer does: change the parameters, watch the loss, keep the changes that lower it. For a straight line under squared error there is also a shortcut. Calculus gives a formula for the best slope and intercept directly, which is what Best fit uses.

The least-squares formulaOptional

Write xˉ\bar{x} and yˉ\bar{y} for the averages of the xx and yy values. The line that minimizes the mean squared error has

w=∑i(xi−xˉ)(yi−yˉ)∑i(xi−xˉ)2,b=yˉ−w xˉ.w = \frac{\sum_i (x_i - \bar{x})(y_i - \bar{y})}{\sum_i (x_i - \bar{x})^2}, \qquad b = \bar{y} - w\,\bar{x}.

You get these by setting the derivatives of the MSE with respect to ww and bb to zero and solving. The best line always passes through the point (xˉ,yˉ)(\bar{x}, \bar{y}).

Absolute error has no such simple formula, because ∣⋅∣|\cdot| has a sharp corner at zero. This demo finds the best absolute-error line exactly by checking every line through two of the points: some best line always passes through at least two of them.

Which loss? Outliers decide

The two losses agree when the data is well behaved. They disagree sharply about OutlierA data point far from the pattern of the rest, from a rare event or a recording error. Some loss functions are pulled hard by outliers.Open in glossary, points far from the general pattern, such as a typo in the data or a genuinely rare event.

Fitting a line

Drag points to move them, tap empty space to add one, tap a point to remove it.

  • Data points
  • Squared misses (area)
  • Best line for squared error
  • Best line for absolute error
Line
Loss
Mean squared error0.41
Mean absolute error0.472
Liney = 0.41x + 2.66

Try this

  • Press Add an outlier. The solid line (best for squared error) tilts noticeably toward the new point. The dashed line (best for absolute error) barely moves.
  • Switch the loss to Absolute error. Now the solid line is the robust one, and the squared-error line is dashed.
  • Drag the outlier further away and compare how each line responds.

A miss of 6 costs 36 under squared error but only 6 under absolute error, so under squared error a single distant point can be worth more than all the others combined. The line bends toward it to shrink that one enormous square.

Neither loss is simply better. Squared error is smooth everywhere, which makes the optimization easy, and it is the right choice when large errors really are much worse than small ones. Absolute error is more robust when the data may contain a few bad values. Choosing a loss is choosing what “wrong” means for your problem.

More than one input

Real predictions use many inputs: size, location, age, and number of bedrooms, not just size. Linear regression extends directly with one weight per input:

y^=w1x1+w2x2+⋯+wdxd+b\hat{y} = w_1 x_1 + w_2 x_2 + \cdots + w_d x_d + b

With two inputs the model is a flat plane instead of a line, and with more it is a flat surface in higher dimensions that you cannot draw. The loss and the fitting work exactly the same way.

Key ideas

  • Linear regression fits a straight line (or flat plane) by choosing the slope and intercept.
  • Misses are squared or made positive before averaging, so they cannot cancel.
  • Squared error lets big misses dominate, so outliers pull the fit hard. Absolute error is more robust to them.
  • For squared error, the best line has an exact formula. In general, models are fit by searching for parameters that lower the loss.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1Why not measure a line's fit by simply adding up the misses (prediction minus true value)?
2One point is far from all the others. Which loss lets that single point pull the fitted line the most?
3In the best-fit mode, the squared-error line has a mean squared error of 0.41. Can any other straight line have a lower mean squared error on these points?

Progress is saved in this browser only.

Up nextGradient descent
Next
How machines learn
  1. 1The learning loop
  2. 2Linear regression and loss
  3. 3Gradient descent
  4. 4Better optimizers
  5. 5Classification and probability
  6. 6Overfitting and generalization
  7. 7Measuring a classifier
  8. 8Finding groups: k-means

Try "embedding", "softmax", "overfitting", or "backpropagation".