A spam filter does not really decide anything. Like the logistic regression from the classification lesson, it produces a score for each email, a number from 0 to 1. Something else turns that score into an action: a threshold. Emails scoring above it go to the spam folder. Where you put that threshold, and how you judge the result, matters as much as the model itself.
Four kinds of outcome
Every flagged-or-not decision lands in one of four boxes, laid out in a Confusion matrixA table counting a classifier's results by true class and predicted class: true positives, false positives, false negatives, and true negatives.Open in glossary.
| Flagged as spam | Delivered | |
|---|---|---|
| Really spam | True positive: spam caught | False negative: spam missed |
| Really good | False positive: good mail flagged | True negative: good mail delivered |
“Positive” just means the class you are looking for, here spam. A false positive is a false alarm; a false negative is a miss.
Why accuracy is not enough
Accuracy is the share of all decisions that were right. It sounds like the obvious score, but it hides the kind of mistakes being made. Suppose only 2 in 100 emails are spam. A filter that never flags anything is right 98 times out of 100, a 98% accuracy, and it is completely useless.
Two questions get closer to what you care about:
- PrecisionOf the items a classifier flagged as positive, the fraction that really are positive. High precision means few false alarms.Open in glossary: of the emails I flagged, what fraction really were spam? This is how much you can trust a flag.
- RecallOf all the items that really are positive, the fraction the classifier flagged. High recall means few misses.Open in glossary: of all the spam that arrived, what fraction did I flag? This is how much you catch.
Here , , and are the counts of true positives, false positives, and false negatives.
Choosing a threshold
A simulated spam filter gives each of 1000 emails a score from 0 to 1. Emails scoring at or above the threshold are flagged as spam.
- Spam (above the line)
- Good email (below)
- Threshold (drag it)
| Flagged | Delivered | |
|---|---|---|
| Really spam | 275spam caught (true positives) | 25spam missed (false negatives) |
| Really good | 70good mail flagged (false positives) | 630good mail delivered (true negatives) |
ROC curve. The dot is the current threshold; the dashed diagonal is random guessing.
Try this
- Drag the threshold in the histogram all the way left. Every email is flagged: recall is 100%, but precision falls to the share of spam.
- Drag it all the way right. Nothing is flagged: precision is undefined (there are no flags to judge) and recall is 0%, yet accuracy looks respectable.
- Set the share of spam to 2% and a high threshold. Compare the accuracy with the “flags nothing” line under the demo.
- Keep the threshold fixed and lower the share of spam from 30% to 5%. Recall and the false positive rate change only a little (the wobble comes from having few spam examples), but precision drops sharply: when spam is rare, even a small false alarm rate swamps the real catches.
- Drag model quality to 0. The two humps sit on top of each other and the ROC curve falls onto the diagonal: the scores carry no information.
Every threshold is a trade-off
Moving the threshold trades one kind of mistake for the other. Lower it and you catch more spam but flag more good mail. Raise it and your flags become trustworthy but more spam slips through. No threshold removes both kinds of error unless the model’s scores separate the classes perfectly.
So the right threshold depends on what each mistake costs. For a spam filter, a lost job offer in the spam folder is worse than a little extra spam, so filters favor precision. For a medical screening test followed by a careful exam, missing a real case is far worse than an extra appointment, so screening favors recall.
Judging the model, not the threshold
To compare models before choosing a threshold, sweep the threshold from strict to lenient and plot, at every setting, recall against the false positive rate, the share of good mail wrongly flagged. That plot is the ROC curveA plot of the true positive rate against the false positive rate as a classifier's threshold sweeps from strict to lenient. The area under it (AUC) summarizes how well the scores rank positives above negatives.Open in glossary. A model whose scores carry no information lies on the diagonal; a perfect model hugs the top-left corner.
The area under the ROC curve, AUC, summarizes it in one number with a neat meaning: it is the probability that a randomly chosen spam email gets a higher score than a randomly chosen good one. An AUC of 0.5 is coin flipping and 1.0 is perfect ranking. Because it looks only at how scores rank the two classes, it does not change with the threshold or with how common spam is.
F1 and other single-number summariesOptional
When you need one number that balances precision and recall at a chosen threshold, a common choice is the F1 score, their harmonic mean:
The harmonic mean is pulled toward the smaller of the two, so a classifier cannot score well on F1 by maxing out one and ignoring the other. Like precision, F1 depends on how common the positive class is. For heavily imbalanced problems, the precision-recall curve, which plots precision against recall across thresholds, is often more informative than the ROC curve.
Key ideas
- A classifier produces scores; a threshold turns them into decisions.
- The confusion matrix counts true positives, false positives, false negatives, and true negatives.
- Accuracy can be high for a useless model when one class is rare.
- Precision asks how trustworthy a flag is; recall asks how much is caught. The threshold trades one for the other.
- The ROC curve and its area judge how well the scores rank positives above negatives, independent of the threshold.