Over the last six lessons you have seen models learn from examples: shading regions by vote, fitting lines, setting weights, writing sentences. Every one of them learned only from its examples. That single fact explains most of the ways AI goes wrong, and knowing it makes you a much sharper user of AI.
A model only knows its data
Imagine a model trained on data collected in one place, then used somewhere else. The demo below makes that concrete. The model is a single neuron, like the one from the artificial neuron lesson. It learned from the examples in the dashed box on the left. The true rule that decides each point’s class is a curve, but in that box the curve is almost straight, so a straight line fits it nicely.
Now slide the solid test box to try the model on new data from further right.
Confident, and wrong
A single neuron learned from data in the dashed box on the left. Slide the solid box to test it on new data from elsewhere.
Try this
- Start with the test box over the training area. The model gets nearly everything right.
- Slide it slowly to the right. Watch Accuracy on the new data fall to around 70%, with more and more ringed mistakes.
- Now watch How sure it sounds. It barely moves, staying around 95%.
- Turn on Show the true rule. The truth bends upward, but the model never saw that part, so it simply continued its straight line.
When the data a model meets differs from the data it learned from, it is called Distribution shiftWhen the data a model meets in use differs from the data it was trained on. Models often get worse under shift, sometimes without becoming any less confident.Open in glossary. The model in the demo is not broken. Inside its experience it is excellent. It was never shown the right side of the plane, so it guessed by extending what it knew, and it had no way to tell that its guess was poor.
The most important thing to notice is the confidence number. It barely changes as the model goes from nearly perfect to wrong on almost a third of the examples. This model’s confidence reflects how far a point is from its boundary, not how familiar the situation is. Nothing in this model measures “I have never seen anything like this.”
Gaps in data become unfair errors
Distribution shift is not only about places. It is also about people. If some group of people is rare or missing in a model’s training data, the model is effectively out of its experience whenever it meets them. This is one common form of Data biasWhen training data misrepresents the situations or people a model will be used on, for example by leaving some groups out, so the model works worse for them.Open in glossary.
Real systems have shown exactly this. A 2018 study called Gender Shades tested commercial face analysis systems and found they misclassified the gender of darker-skinned women far more often than that of lighter-skinned men: error rates up to about 35% for the first group, under 1% for the second. A 2020 study of major speech recognition systems found they made roughly twice as many word errors for Black speakers as for white speakers. In both cases, a likely contributor was training data that underrepresented the people the systems later failed.
Data also carries patterns people would rather a model not learn. A model trained on past hiring decisions learns whatever preferences shaped those decisions, fair or not. The model cannot tell a real pattern from an unfair habit; both are just patterns in the numbers.
Fluent is not the same as true
Chatbots have their own version of this problem. In the last lesson you wrote with a model that picks likely next words. Its sentences were often fluent, and they regularly stitched together events that never happened in any fable. The counting model had no idea which sentences were true. It only knew which words tend to follow which.
Large language models are vastly better at producing accurate text, because accurate text is a large part of what they learned from. But their writing process is the same kind of prediction, and it contains no built-in step that checks a statement against reality. When the training data does not settle a question, the model can produce text that has the shape of a correct answer, such as a plausible quote, a realistic-looking citation, or a confident statistic, that is simply invented. This is called HallucinationWhen a language model states something false or invented, such as a fake quote or citation, in the same fluent, confident style as true statements.Open in glossary.
Developers reduce it in several ways: further training that rewards accurate and honest answers, letting the model search the web or a document collection and quote what it finds, and training models to say when they do not know. These help a lot. They do not make the problem disappear.
What models are good at
None of this means AI is unreliable everywhere. Inside the range of their training data, well-built models can be extremely accurate, often more consistent than people at narrow tasks. The skill is knowing where that range ends: what the model was trained on, whether your situation looks like that data, and whether anyone has checked how it performs for cases like yours.
Key ideas
- A model learns only from its training data. Far from that data, it guesses by extending patterns that may no longer hold.
- When data in use differs from training data, accuracy can collapse while confidence stays high.
- Groups that are rare or missing in training data get worse results, one common source of unfair AI errors.
- Chatbots write by predicting likely text, with no built-in fact check, so fluent answers can be false.
- Check AI answers in proportion to how much they matter and how unusual the question is.