The straight line in the last lesson had a formula for its best fit. Almost nothing else does. A neural network can have millions of parameters, and there is no equation you can solve to get their best values. What you can do is improve them a little at a time. The method that does this, in one form or another, trains nearly every model in use today. It is called Gradient descentAn optimization method that repeatedly moves the parameters a small step in the direction that lowers the loss fastest, the opposite of the gradient.Open in glossary.
Downhill in fog
Imagine you are on a hillside in thick fog and want to reach the bottom of the valley. You cannot see the valley, but you can feel the slope under your feet. A reasonable plan: take a step in the steepest downhill direction, feel the slope again, and repeat.
That is the whole algorithm. The hillside is the loss, your position is the model’s parameters, and “feeling the slope” is computing the DerivativeThe slope of a function at a point: how much the output changes for a tiny change in the input.Open in glossary of the loss.
One parameter
With a single parameter , the slope at a point tells you two things. Its sign tells you which way is uphill: a positive slope means the loss rises to the right, so you should move left. Its size tells you how steep it is. Gradient descent uses both:
Here is the slope of the loss at the current , and (eta) is the Learning rateThe step size in gradient descent. Too small and training crawls; too large and each step overshoots, so the loss can grow instead of shrink.Open in glossary, also called the step size. The minus sign is what makes it go downhill.
Downhill in one dimension
The ball shows the current parameter value. Drag it to choose a starting point, then step.
The slope is negative (uphill to the left), so the step goes right. The arrow along the bottom shows the next step: step size times slope, 0.3 times -2.5.
Try this
- With the step size at 0.3, press Keep stepping. The steps start large and get smaller as the ball nears the bottom, even though the step size never changes.
- Reset and set the step size to exactly 1. On this bowl one step lands precisely at the bottom.
- Try 1.5. The ball overshoots, bounces back and forth across the bottom, and still settles.
- Try 2.2. Each bounce is bigger than the last, and the ball flies off the chart.
- Switch to Two dips. Starting from the left, the ball settles in the shallow dip and never finds the deeper one on the right. Drag the ball to the right side and step again.
Two lessons are hiding in that demo. First, the step size matters enormously: too small wastes time, a bit too large makes the ball bounce, and much too large makes training blow up. Second, gradient descent only knows the slope where it stands. It stops at the bottom of whatever dip it is in, which may be a Local minimumA point lower than everything around it but not necessarily the lowest point overall. Gradient descent can settle in one and stop improving.Open in glossary rather than the lowest point overall.
Where the limit of 2 comes fromOptional
The simple bowl is , with slope . One step gives
Every step multiplies by . If that factor is between 0 and 1, so shrinks steadily toward 0. If it is 0, and you land on the minimum in one step. If the factor is between and , so flips sign each step while shrinking: the bouncing you saw. If the factor is below and grows without limit.
For a bowl that curves more sharply, , the same argument gives the condition . Steeper curvature demands a smaller step.
Two parameters: the gradient
With two parameters the loss is a surface, like a landscape. We draw it from above as a contour map, the way hiking maps show terrain: each line joins points of equal loss, and lines bunched close together mean steep ground.
Now there is a slope in each direction. Collected together, the slopes form the GradientThe list of slopes of a function, one per parameter. It points in the direction in which the function increases fastest, so its opposite points downhill.Open in glossary, written . It points in the direction of steepest ascent, straight across the contour lines. Gradient descent steps the opposite way:
Each is a partial derivative: the slope when you change only and hold everything else fixed.
Gradient descent on a loss surface
Contour lines join points of equal loss; the stronger the tint, the lower the loss. Tap anywhere to start there.
One minimum, but steeper in one direction. Too large a step makes the path zig-zag.
Tap the map to choose a starting point, or press Step. The arrow points downhill: the direction of the negative gradient.
Try this
- On the Stretched bowl, press Run with the default step size 0.1. The path curves toward the minimum (marked with a cross) instead of heading straight for it.
- Reset, set the step size near 0.2, and run. The path zig-zags across the narrow direction of the bowl.
- Now try 0.26. The zig-zag grows instead of shrinking, and the run diverges.
- Choose Many valleys and tap different starting points. Each start rolls into its own valley.
- Choose Banana valley and turn on Fast forward. The path drops into the curved valley quickly, then crawls along its floor. The loss chart flattens into a long, slow slide.
- Choose Saddle point. The start sits almost exactly on a ridge, so the path heads into the flat middle and lingers there before it slides off to one side.
The stretched bowl shows the central difficulty. It curves 8 times more steeply in the vertical direction than the horizontal one. The vertical direction needs a step below to stay stable, but with such a small step the horizontal direction, which is gentle, takes many steps to cross. One step size has to serve every direction at once, and on a stretched surface no single value is good for all of them. The next lesson is about optimizers that fix exactly this.
From two parameters to millions
You cannot draw a landscape with a million dimensions, but the algorithm does not care. The gradient becomes a list of a million slopes, one per parameter, and each parameter takes its own step in the downhill direction. Computing those slopes efficiently for a neural network is the job of backpropagation, covered in the neural networks track.
High dimensions also change the picture of getting stuck. A point where the gradient is zero is a minimum only if the loss curves upward in every direction. With millions of directions, it is much more likely that at least one curves down, which makes the point a Saddle pointA flat point that curves up in some directions and down in others, like the middle of a horse saddle. The gradient is zero there, but it is not a minimum.Open in glossary that a little noise or momentum can escape. Researchers have argued that in large networks saddle points and long flat regions are a bigger obstacle than poor local minima, which is part of why gradient descent works better on big models than small pictures suggest.
Key ideas
- Gradient descent improves parameters step by step: compute the slope of the loss, then move a small amount in the downhill direction.
- Each step is the learning rate times the slope, so steps shrink naturally as the surface flattens near a minimum.
- Too small a learning rate is slow. Too large overshoots, bounces, and can diverge; along a direction with curvature c, it must stay below 2 / c.
- The gradient points across contour lines, not at the minimum, so paths curve and zig-zag on stretched surfaces.
- Gradient descent can settle in a local minimum or linger near a saddle point. It only knows the slope where it stands.