Lesson 5 of 8

Classification and probability

To sort things into categories, a model turns a score into a probability. Meet the sigmoid, the cross-entropy loss, and the decision boundary, and train a classifier live.

Beginner17 min

In this lesson you will

  • Explain how the sigmoid turns any score into a probability between 0 and 1
  • Read cross-entropy as "how surprised the model was by the right answer"
  • Train a logistic regression classifier and interpret its decision boundary

Is this email spam? Is this tumor benign? Which digit is this? These are ClassificationA prediction task where the answer is one of a fixed set of categories, such as spam or not spam, or which digit a drawing shows.Open in glossary problems: the answer is a category, not a number. A good classifier does more than name a category. It says how sure it is, because “90% likely to be spam” and “51% likely to be spam” call for different actions.

From a score to a probability

Start with something you already know: a weighted sum of the inputs, like linear regression, which gives a score z=w x+bz = w\,x + b. The score can be any number, from very negative to very positive. A probability must stay between 0 and 1. The SigmoidThe S-shaped function 1 / (1 + e^-z), which squashes any number into the range 0 to 1 so it can be read as a probability.Open in glossary function bridges the two:

σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

Large positive scores map to probabilities near 1, large negative scores to probabilities near 0, and a score of exactly 0 maps to 0.5. Putting a weighted sum through a sigmoid gives Logistic regressionA classifier that computes a weighted sum of the inputs and passes it through the sigmoid to get the probability of the positive class. Despite the name, it is used for classification.Open in glossary. Despite the name, it is a classifier.

From score to probability

Made-up data: hours of practice before a driving test, and whether the test was passed. The curve is the model's probability of passing.

  • Failed (y = 0)
  • Passed (y = 1)
  • Gap to the right answer (thickest: the costliest point)
0.40
-1
Mean cross-entropy0.378
Accuracy at 50%81%
50% point2.5 hours

Try this

  • Drag the weight up. The curve gets steeper: the model becomes more decisive about the change from fail to pass.
  • Drag the bias. The whole curve slides left or right, moving the 50% point.
  • Find the setting with the lowest mean cross-entropy you can, then press Train (Adam) and see how close you got.
  • Make the curve very steep in the wrong place. The thick orange gap marks the costliest point, and the loss jumps.

Measuring a probabilistic mistake

Squared error is a poor fit for probabilities. What we care about is how much probability the model gave to the answer that turned out to be right. The standard loss is Cross-entropyA loss for predicted probabilities: minus the log of the probability the model gave to the right answer. Confident wrong answers are punished very heavily.Open in glossary: minus the logarithm of that probability.

loss=−ln⁡pcorrect\text{loss} = -\ln p_{\text{correct}}
Probability given to the right answerLoss
0.990.01
0.90.11
0.50.69
0.12.30
0.014.61

A confident right answer costs almost nothing, a coin flip costs about 0.69, and a confident wrong answer costs a lot. The loss keeps growing without limit as the probability for the right answer approaches zero, so a model is strongly discouraged from being sure and wrong. One way to read it: cross-entropy measures how surprised the model was by the truth.

Why the gradient is so simpleOptional

For a label y∈{0,1}y \in \{0, 1\} and predicted probability p=σ(z)p = \sigma(z), the loss on one example is

ℓ=−[ yln⁡p+(1−y)ln⁡(1−p) ].\ell = -\left[\,y \ln p + (1 - y)\ln(1 - p)\,\right].

Using the fact that σ′(z)=σ(z) (1−σ(z))\sigma'(z) = \sigma(z)\,(1 - \sigma(z)), the derivative with respect to the score simplifies to

∂ℓ∂z=p−y.\frac{\partial \ell}{\partial z} = p - y.

So the gradient for each weight is just (predicted probability minus label) times that weight’s input. A point the model gets right with confidence contributes almost nothing; a point it gets badly wrong pushes hard. The demos compute exactly this, averaged over the data.

Two inputs: the decision boundary

With two inputs, the score is z=w1x1+w2x2+bz = w_1 x_1 + w_2 x_2 + b. Each point on the plane gets a probability, and the classifier predicts class B wherever that probability is above 50%. The line where it is exactly 50%, where z=0z = 0, is the Decision boundaryThe line or surface where a classifier switches from predicting one class to another.Open in glossary.

Logistic regression

Shading shows the model's probability for each spot: blue for class A, orange for class B, gray where it is unsure. Tap to add points.

  • Class A
  • Class B
  • Decision boundary (50%)
  • 10% and 90% lines
0.50
Data
Tap adds
Steps0
Mean cross-entropy0.6931
Accuracy50%
Weights w1, w2, b0, 0, 0

Try this

  • Press Train. The boundary swings into place between the two clouds, and the gray band of uncertainty narrows.
  • Keep training on the Separated data. The boundary barely moves, but the dashed 10% and 90% lines keep squeezing toward it: the model grows ever more confident.
  • Switch to Overlapping. Training settles at a loss well above zero, and the uncertain band stays wide, because some points really are ambiguous.
  • Add a few class B points deep inside the blue region and train again. The boundary tilts to accommodate them, but a straight line can only do so much.
  • Raise the step size to the maximum. On this problem training still converges, just in bigger jumps; try it on the overlapping data too.

The boundary of logistic regression is always a straight line, or a flat plane with more inputs, because the probability depends only on a weighted sum. When the classes are arranged in a ring or a spiral, no straight line separates them. You can fix that by inventing extra input features, such as x12x_1^2, or by letting a neural network learn its own features, which is what the neural networks track is about.

Key ideas

  • A classifier outputs probabilities, not just labels, so you know how sure it is.
  • The sigmoid squashes any score into a probability between 0 and 1. Logistic regression is the sigmoid of a weighted sum.
  • Cross-entropy is minus the log of the probability given to the right answer: cheap when confident and right, very expensive when confident and wrong.
  • The decision boundary of logistic regression is a straight line where the probability is 50%.
  • On perfectly separable data, the weights grow without limit unless regularization stops them.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1A model gives the right answer a probability of 0.5. A second model gives the right answer 0.01. How do their cross-entropy losses compare?
2For logistic regression with two inputs, what shape is the decision boundary?
3On data where the two classes are perfectly separated, you keep training logistic regression. What happens to the weights?

Progress is saved in this browser only.

Up nextOverfitting and generalization
Next
How machines learn
  1. 1The learning loop
  2. 2Linear regression and loss
  3. 3Gradient descent
  4. 4Better optimizers
  5. 5Classification and probability
  6. 6Overfitting and generalization
  7. 7Measuring a classifier
  8. 8Finding groups: k-means

Try "embedding", "softmax", "overfitting", or "backpropagation".