Lesson 2 of 7

Turning things into numbers

Models only see numbers. Learn how pictures and objects become numbers, and why choosing the right measurements decides what a model can learn.

Novice12 min

In this lesson you will

  • See how a picture becomes a grid of numbers
  • Describe an example with features and a label
  • Explain why some features help a model and others do not

In the last lesson the computer only ever saw each example as two numbers: how far across and how far up. That was not a shortcut for the demo. It is how every model works. A model takes numbers in and gives numbers out, so before a computer can learn anything about photos, fruit, songs, or sentences, those things have to become numbers.

This lesson is about that step, and about a surprising fact: which numbers you choose matters as much as how clever the model is.

A picture is a grid of numbers

A digital picture is a grid of tiny squares called PixelOne tiny square of a digital image. A grayscale pixel is stored as one brightness number; a color pixel as three (red, green, and blue).Open in glossary. In a black-and-white picture, each pixel is stored as one number for its brightness. Below, 0 is black and 255 is white.

A picture is a grid of numbers

Each square is one pixel, stored as a brightness from 0 (black) to 255 (white). Paint to change the numbers.

235235235202020202352352352352020235235235235202023523520235235235235235235202352023523520235235202352352020235235235235235235235235202023520235235235235202352020235235202020202352352023520235235235235235235202352352020235235235235202023523523523520202020235235235

A color photo stores three numbers per pixel (red, green, and blue), so a 12-megapixel phone photo is about 36 million numbers.

Paint with
Numbers in this picture100
Dark pixels36

Try this

  • Paint a few squares white and watch their numbers change to 255. The picture and the numbers are the same thing, shown two ways.
  • Turn off Show the numbers. You see a face. The computer never does: it only has the hundred numbers.
  • Paint with gray. Brightness is not just on or off; every value from 0 to 255 is allowed.

To you, the picture is obviously a smiling face. To the computer, it is a list of 100 numbers, mostly 235 with some 20s. Recognizing the face means finding a pattern in those numbers, and that is exactly the kind of thing models learn.

Sound works the same way: a recording is a long list of numbers measuring air pressure, typically 44,100 of them for every second of audio. Text needs a few more steps, which you will see in the language track, but it also ends up as numbers.

Describing things with features

Not everything starts as a picture. Often you describe each example with a handful of measurements. Each measurement is called a FeatureOne measurable property of an example, given to a model as a number, such as a fruit's weight or a pixel's brightness.Open in glossary, and the answer you want the model to learn is called the LabelThe answer a model should learn to give for an example, such as "apple" or "orange". Labeled examples are what supervised learning learns from.Open in glossary.

Here are a few fruit described that way. Each row is one example. The first four columns are features; the last is the label.

Weight (g)Width (cm)Skin colorSkin bumpiness (0 to 10)Label
1677.2red0.4apple
2348.0green0.4apple
1507.4orange7.4orange
1727.6orange8.2orange

Skin color is written as a word here so you can read it, but the model gets a number: the color’s position on a color wheel, where red is near 0, orange near 30, yellow near 60, and the yellowish green of an unripe apple near 90.

Some features are better than others

Below are sixty fruit, thirty apples and thirty oranges, each measured four ways. The fruit are made up for this demo, but their measurements follow realistic ranges. Pick two features to plot. Each dot is one fruit, placed by its two numbers.

Which measurements tell apples from oranges?

Sixty fruit, measured four ways. Choose two measurements to plot and see how well a model can tell the fruit apart using only those.

Across
Up
Fruit a model gets right27 of 60 (45%)
VerdictAbout as good as guessing

Try this

  • Start with Weight across and Width up. Apples and oranges are mixed together, and the model gets only about half of them right.
  • Change Up to Bumpiness. The two kinds of fruit pull apart into separate groups, and the score jumps.
  • Try Color with Weight. Color helps, but notice that apples form two groups, red and green, with oranges in between.
  • Before switching, predict the score for Width and Color. Then check.

The model never changed. It is the same nearest-neighbors method from the last lesson every time. What changed was the information it was given. With weight and width, apples and oranges look the same, so the model is guessing. With skin bumpiness, the difference is right there in the numbers.

That is the first lesson of practical machine learning: a model can only find patterns that are present in the numbers it receives. No amount of cleverness recovers information you left out.

For decades, people chose features by hand, a job called feature engineering. A big change brought by Deep learningMachine learning with neural networks that have many layers. It powers modern image recognition, speech recognition, translation, and chatbots.Open in glossary is that models now often start from raw numbers, like every pixel of a photo, and learn useful features on their own. You will see how in the neural network lessons.

Why the demo rescales the numbersOptional

Weight is measured in grams and runs from about 110 to 250. Bumpiness runs from 0 to 10. If you measured distance between fruit using the raw numbers, a 20-gram difference in weight would swamp the biggest possible difference in bumpiness, and the model would effectively ignore bumpiness.

So before measuring distances, the demo rescales each feature: it subtracts that feature’s average and divides by how spread out it is (its standard deviation). After that, every feature has a similar range, and each one gets a fair say. This step, called standardization, is routine in machine learning.

The score in the demo is measured fairly too: each fruit is guessed from its five nearest other fruit, so no fruit can vote for itself.

Key ideas

  • Models only see numbers. Pictures become grids of brightness values; sounds become lists of measurements.
  • An example is described by features (the measurements) and a label (the answer to learn).
  • A model can only learn patterns that are present in its features. Good features make learning easy; useless ones make it impossible.
  • Modern deep learning often learns its own features from raw data instead of relying on hand-picked ones.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1A model is learning to tell apples from oranges. Which is the label?
2In the demo, weight and width together got about half the fruit right. Why?
3A black-and-white picture is 28 pixels wide and 28 pixels tall. How many numbers does it take to store it?

Progress is saved in this browser only.

Up nextGetting less wrong
Next
AI from zero
  1. 1What is AI, really?
  2. 2Turning things into numbers
  3. 3Getting less wrong
  4. 4The artificial neuron
  5. 5Networks of neurons
  6. 6How a chatbot writes
  7. 7What AI gets wrong

Try "embedding", "softmax", "overfitting", or "backpropagation".