Lesson 7 of 7

What AI gets wrong

Models learn only from their data, and can be confidently wrong outside it. See a model fail far from its training data, and learn why fluent chatbots can state falsehoods.

Novice14 min

In this lesson you will

  • Explain why a model can fail on data unlike its training data
  • Recognize that a model's confidence is not a measure of its accuracy
  • Describe how gaps in training data become unfair errors
  • Explain why a fluent chatbot answer can still be false

Over the last six lessons you have seen models learn from examples: shading regions by vote, fitting lines, setting weights, writing sentences. Every one of them learned only from its examples. That single fact explains most of the ways AI goes wrong, and knowing it makes you a much sharper user of AI.

A model only knows its data

Imagine a model trained on data collected in one place, then used somewhere else. The demo below makes that concrete. The model is a single neuron, like the one from the artificial neuron lesson. It learned from the examples in the dashed box on the left. The true rule that decides each point’s class is a curve, but in that box the curve is almost straight, so a straight line fits it nicely.

Now slide the solid test box to try the model on new data from further right.

Confident, and wrong

A single neuron learned from data in the dashed box on the left. Slide the solid box to test it on new data from elsewhere.

Inside the training area
Accuracy where it learned99%
Accuracy on the new data100%
How sure it sounds on the new data97%

Try this

  • Start with the test box over the training area. The model gets nearly everything right.
  • Slide it slowly to the right. Watch Accuracy on the new data fall to around 70%, with more and more ringed mistakes.
  • Now watch How sure it sounds. It barely moves, staying around 95%.
  • Turn on Show the true rule. The truth bends upward, but the model never saw that part, so it simply continued its straight line.

When the data a model meets differs from the data it learned from, it is called Distribution shiftWhen the data a model meets in use differs from the data it was trained on. Models often get worse under shift, sometimes without becoming any less confident.Open in glossary. The model in the demo is not broken. Inside its experience it is excellent. It was never shown the right side of the plane, so it guessed by extending what it knew, and it had no way to tell that its guess was poor.

The most important thing to notice is the confidence number. It barely changes as the model goes from nearly perfect to wrong on almost a third of the examples. This model’s confidence reflects how far a point is from its boundary, not how familiar the situation is. Nothing in this model measures “I have never seen anything like this.”

Gaps in data become unfair errors

Distribution shift is not only about places. It is also about people. If some group of people is rare or missing in a model’s training data, the model is effectively out of its experience whenever it meets them. This is one common form of Data biasWhen training data misrepresents the situations or people a model will be used on, for example by leaving some groups out, so the model works worse for them.Open in glossary.

Real systems have shown exactly this. A 2018 study called Gender Shades tested commercial face analysis systems and found they misclassified the gender of darker-skinned women far more often than that of lighter-skinned men: error rates up to about 35% for the first group, under 1% for the second. A 2020 study of major speech recognition systems found they made roughly twice as many word errors for Black speakers as for white speakers. In both cases, a likely contributor was training data that underrepresented the people the systems later failed.

Data also carries patterns people would rather a model not learn. A model trained on past hiring decisions learns whatever preferences shaped those decisions, fair or not. The model cannot tell a real pattern from an unfair habit; both are just patterns in the numbers.

Fluent is not the same as true

Chatbots have their own version of this problem. In the last lesson you wrote with a model that picks likely next words. Its sentences were often fluent, and they regularly stitched together events that never happened in any fable. The counting model had no idea which sentences were true. It only knew which words tend to follow which.

Large language models are vastly better at producing accurate text, because accurate text is a large part of what they learned from. But their writing process is the same kind of prediction, and it contains no built-in step that checks a statement against reality. When the training data does not settle a question, the model can produce text that has the shape of a correct answer, such as a plausible quote, a realistic-looking citation, or a confident statistic, that is simply invented. This is called HallucinationWhen a language model states something false or invented, such as a fake quote or citation, in the same fluent, confident style as true statements.Open in glossary.

Developers reduce it in several ways: further training that rewards accurate and honest answers, letting the model search the web or a document collection and quote what it finds, and training models to say when they do not know. These help a lot. They do not make the problem disappear.

What models are good at

None of this means AI is unreliable everywhere. Inside the range of their training data, well-built models can be extremely accurate, often more consistent than people at narrow tasks. The skill is knowing where that range ends: what the model was trained on, whether your situation looks like that data, and whether anyone has checked how it performs for cases like yours.

Key ideas

  • A model learns only from its training data. Far from that data, it guesses by extending patterns that may no longer hold.
  • When data in use differs from training data, accuracy can collapse while confidence stays high.
  • Groups that are rare or missing in training data get worse results, one common source of unfair AI errors.
  • Chatbots write by predicting likely text, with no built-in fact check, so fluent answers can be false.
  • Check AI answers in proportion to how much they matter and how unusual the question is.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1In the demo, the model's accuracy falls far from its training data. What happens to its confidence?
2A speech recognizer was trained mostly on recordings of a few accents. What is the most likely result?
3Why can a chatbot state a false fact in a confident, fluent sentence?

Progress is saved in this browser only.

Up next in How machines learnThe learning loop
Next
AI from zero
  1. 1What is AI, really?
  2. 2Turning things into numbers
  3. 3Getting less wrong
  4. 4The artificial neuron
  5. 5Networks of neurons
  6. 6How a chatbot writes
  7. 7What AI gets wrong

Try "embedding", "softmax", "overfitting", or "backpropagation".