Lesson 6 of 6

Reading handwriting

A real network, trained on 60,000 handwritten digits, running in your browser. Draw a digit and watch 784 pixels flow through two hidden layers into ten probabilities.

Intermediate15 min

In this lesson you will

  • Follow a drawing through preprocessing, two hidden layers, and a softmax output
  • Interpret test accuracy and explain why the network still makes confident mistakes
  • Describe what individual hidden neurons respond to, and why they are hard to interpret

Every idea in this track meets here. The network below is the one whose shapes you counted in the first lesson: 784 inputs, a hidden layer of 64 ReLU neurons, a hidden layer of 32, and 10 outputs, 52,650 parameters in all. It was trained with backpropagation and Adam, as in the playground, but on real data.

The data: MNIST

MNISTA classic dataset of 70,000 handwritten digits, each a 28 x 28 grayscale image: 60,000 for training and 10,000 for testing.Open in glossary is a set of 70,000 handwritten digits assembled by Yann LeCun, Corinna Cortes, and Christopher Burges from forms filled in by American Census Bureau employees and high school students. Each digit is a 28 × 28 grayscale image. It is split into 60,000 training images and 10,000 test images, written by different people, so the test measures how well a model reads handwriting it has never seen.

Before any training, every digit was prepared the same way: scaled to fit inside a 20 × 20 box while keeping its proportions, then placed in the 28 × 28 frame so that its center of mass sits in the middle. The demo does the same to your drawing. That is not cosmetic. A network learns the regularities of its training data, including how that data was prepared.

How this network was trained

The network trained for 14 epochs on the 60,000 training digits, with each digit randomly rotated by up to 12 degrees, rescaled by up to 12%, and shifted by up to about 2.5 pixels every time it was shown. That is Data augmentationMaking extra training examples by applying small, label-preserving changes to existing ones, such as shifting, rotating, or rescaling an image.Open in glossary: hand-drawn digits on a screen vary more than the scanned forms in MNIST, and showing the network varied versions helps it cope. On the 10,000 test digits it reaches 97.8% accuracy. The training script is in the project’s scripts/digit-recognizer/ folder, and the same weights run here in plain TypeScript, matching the training framework’s outputs to within rounding.

Read handwriting with a trained network

A network trained on 60,000 handwritten digits. Draw one and watch every layer respond.

What the network sees: 28 × 28 = 784 inputs

Hidden layer 1: 64 neurons (tap one to see its weights)

Hidden layer 2: 32 neurons

Output: probability of each digit

  1. 0
  2. 1
  3. 2
  4. 3
  5. 4
  6. 5
  7. 6
  8. 7
  9. 8
  10. 9
Network's answerDraw something

Try this

  • Draw a clear 3. Then draw it small in a corner. In the 28 × 28 preview it keeps the same size and position, because preprocessing rescales and recenters it. Its strokes do get thicker, because the pen stays the same width.
  • Draw an ambiguous digit, such as a 4 with a closed top or a 1 with a long flag, and watch the probability move between two classes as you add strokes.
  • Draw something that is not a digit, like a letter or a smiley face. The network still names a digit, often with high confidence.
  • Press Show a real test digit and step through a few. These are digits the network never trained on.
  • Tap neurons in the first hidden layer to see their weights as a 28 × 28 picture: red pixels push the neuron to fire, blue pixels hold it back.

Reading the layers

Watch the hidden layers as you draw. Each square’s color shows one neuron’s activation, scaled so the most active neuron in its layer has the strongest color; a plain square means zero. Many neurons stay silent for any given digit, because ReLU sets every negative weighted sum to zero. Different digits light up different patterns.

The output layer turns its ten numbers into probabilities with a SoftmaxA function that turns a list of real-valued scores into probabilities that are positive and sum to 1, by exponentiating each score and dividing by the total.Open in glossary, so they are positive and add up to 1. The network’s answer is the largest one.

The weight pictures show what a first-layer neuron responds to: its weighted sum is large when ink falls on its red pixels and avoids its blue ones. Some look like fragments of strokes. Most look like nothing in particular. That is typical. The network spreads what it knows across many neurons, and each digit is recognized from a combination of partial clues, not by one neuron per digit.

How good is good?

A 97.8% test accuracy means about 218 of the 10,000 test digits are misread. Simple fully connected networks of about this size usually score around 97 to 98%. Convolutional networks, which build in the locality and weight sharing from the convolutions lesson, do much better: the best published results misread fewer than 20 of the 10,000. MNIST is now considered easy, but it was a standard benchmark for years, and it remains a good place to see a whole network at once.

What the network does not knowOptional

The network was never told what a digit is. It found weights that make its outputs match the labels on 60,000 examples, and those weights generalize to new examples that resemble them. Nothing in it represents loops, strokes, or numbers as concepts. This is why it confidently names a digit for a doodle: its output layer has no way to express “this is not like my training data”. Detecting such out-of-distribution inputs is a research area of its own, and What AI gets wrong explores the same problem from the beginning of the course.

Key ideas

  • This network is 784 → 64 → 32 → 10 with ReLU and softmax: 52,650 parameters, trained on 60,000 MNIST digits.
  • Inputs must be prepared the way the training data was: scaled into 20 × 20 and centered by mass.
  • Test accuracy (97.8%) is measured on unseen MNIST digits, not on arbitrary drawings.
  • Softmax forces every input into one of the ten classes, so unfamiliar inputs still get confident answers.
  • Knowledge is spread across many neurons; individual neurons are rarely easy to interpret.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1You draw the letter A and the network answers "4" with 70% confidence. Why does it not say "none of these"?
2The network is 97.8% accurate on the 10,000 MNIST test digits. Roughly how many of those does it get wrong?
3Why does the demo shrink and center your drawing before the network sees it?

Progress is saved in this browser only.

Up next in Language as vectorsTokens
Next
Neural networks
  1. 1Layers are matrix multiplications
  2. 2Why nonlinearity matters
  3. 3Backpropagation
  4. 4Training a network
  5. 5Convolutions
  6. 6Reading handwriting

Try "embedding", "softmax", "overfitting", or "backpropagation".