Lesson 4 of 6

Training a network

Put layers, activations, backpropagation, and gradient descent together and train real networks in your browser. See what each setting changes, and what each neuron learns.

Intermediate20 min

In this lesson you will

  • Describe the training loop in terms of epochs, minibatches, and gradient steps
  • Predict how depth, width, activation, and learning rate change what a network can learn and how fast
  • Recognize overfitting from training and test loss curves

You now have every piece. A network is layers of matrix multiplications with activations in between. A loss measures how wrong it is. Backpropagation computes the gradient of that loss for every weight, and gradient descent nudges every weight downhill. Repeat that thousands of times and the network learns. This lesson puts it all in your hands.

The training loop

Training a network on a dataset looks like this:

  1. Start every weight at a small random value.
  2. Shuffle the training examples and split them into small groups called minibatches.
  3. For each minibatch: run the forward pass, compute the loss, backpropagate, and update every weight.
  4. When every minibatch has been used once, one EpochOne full pass through the training data. Training usually runs for many epochs, shuffling the examples each time.Open in glossary is done. Go back to step 2.

Using a minibatch instead of the whole dataset for each step is stochastic gradient descent: each step’s gradient is a noisy estimate of the true one, but the steps are cheap and frequent. The Batch sizeHow many training examples are used to compute each gradient step. Small batches give noisy but frequent updates; large batches give smoother but fewer updates per epoch.Open in glossary sets that trade.

The network below trains exactly this way, in your browser. Each point is an example with two inputs, its coordinates. The big square shows the network’s prediction for every possible input. The diagram shows every neuron as a small square of its own: what that neuron outputs for every point on the plane.

Neural network playground

Train a real network in your browser. Shading shows its prediction everywhere; small squares show what each neuron computes.

Blue regions are predicted blue, orange regions orange. Paler means less sure.

Tap a neuron to show its output full size. Lines are weights: red positive, blue negative, thicker is larger.

Data
0.20
Hidden layers
4
Activation
0.03
Optimizer
Batch size
Epoch0
Training loss0.000
Test loss0.000
Test accuracy0%
Parameters0

Try this

  • Press Train with the default settings. Within a few dozen epochs the circle is separated and the training loss drops toward zero.
  • Set Hidden layers to 0 and train again. The network is now just the output layer, which can only draw a straight line. The loss stays near 0.69, the loss of guessing 50/50, because no line can enclose a circle.
  • Go back to 2 hidden layers and watch the neuron squares while it trains. Each first-layer neuron is a soft line across the plane. Later neurons combine those lines into curves. Tap any neuron to see it full size.
  • Set the learning rate to 1. The loss jumps around and the network never settles. Then try 0.001 with SGD: it barely moves. Switch to Adam at 0.001 and it learns, slowly.
  • Choose Spiral, ReLU, 8 neurons per hidden layer, a learning rate near 0.01, and train for a few hundred epochs. A harder shape needs more capacity and more patience.

What each setting does

Depth and width set the network’s capacity: how complicated a boundary it can form. Zero hidden layers gives a straight line. One hidden layer of a few neurons can enclose a region. Spirals need more neurons and usually more layers, because each layer can build on the shapes the previous one made.

The activation function shapes the pieces. ReLU networks build boundaries out of straight segments, so their predictions have corners. Tanh and sigmoid networks build smooth curves. With no activation, any number of layers is still a single straight line, as you saw in the previous lesson on nonlinearity.

The Learning rateThe step size in gradient descent. Too small and training crawls; too large and each step overshoots, so the loss can grow instead of shrink.Open in glossary is the size of each step. Too large and steps overshoot, so the loss bounces or blows up. Too small and training crawls. Adam adapts the step size for each weight separately based on the history of its gradients, which is why it copes with a wider range of learning rates than plain SGD. The optimizers lesson shows why.

Random starting weights matter more than you might expect. Press New starting weights a few times on the spiral: some starts find a good solution quickly, others take longer or get stuck. Starting at random is also essential. If every neuron in a layer began identical, every one would get the same gradient and they would stay identical forever.

Watching for overfitting

The dashed line in the loss chart is the test loss, computed on points the network never trains on. Turn on Show test data to see them as hollow marks. As long as the two curves fall together, the network is learning patterns that hold for new data.

Try this: make it overfit

  • Choose Circle, set Noise to 1, use 2 hidden layers of 8 neurons, and train for a few hundred epochs.
  • Watch the two loss curves. Training loss keeps falling, but after a few dozen epochs the test loss turns around and climbs.
  • Turn on Show test data and look at the boundary. It has grown small pockets around individual noisy training points, pockets that the test points do not follow.

That is OverfittingWhen a model fits the quirks and noise of its training data so closely that it does worse on new data. Training error keeps falling while test error rises.Open in glossary: with enough capacity, the network starts fitting the noise in its particular training points. The usual remedies are more data, a smaller network, regularization, or simply stopping when test loss stops improving. The overfitting lesson in the machine learning track covers them in depth.

Key ideas

  • Training repeats one loop: shuffle, split into minibatches, forward, loss, backward, update. One pass over the data is an epoch.
  • Depth and width set capacity. With no hidden layer, the boundary is a straight line.
  • Hidden neurons in the first layer each draw a soft line; later layers combine lines into curves.
  • The learning rate trades speed against stability. Adam tolerates a wider range than plain SGD.
  • Training loss falling while test loss rises is the signature of overfitting.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1What is one epoch?
2Why are a network's weights started at random values instead of all zeros?
3Training loss keeps falling, but test loss has started to rise. What is happening?

Progress is saved in this browser only.

Up nextConvolutions
Next
Neural networks
  1. 1Layers are matrix multiplications
  2. 2Why nonlinearity matters
  3. 3Backpropagation
  4. 4Training a network
  5. 5Convolutions
  6. 6Reading handwriting

Try "embedding", "softmax", "overfitting", or "backpropagation".