Lesson 4 of 8

Better optimizers

Plain gradient descent struggles in narrow valleys and on flat ground. Momentum, RMSProp, and Adam fix this with two simple ideas. Race them on the same surfaces.

Beginner16 min

In this lesson you will

  • Explain the two weaknesses of plain gradient descent, narrow ravines and flat regions
  • Describe what momentum adds and what per-parameter step sizes add
  • Compare optimizers on the same surface and explain why none wins everywhere

In the last lesson, plain gradient descent had two clear weaknesses. On the stretched bowl and the banana-shaped valley, it zig-zagged across the narrow direction while creeping along the long one. Near the saddle point, where the ground is almost flat, its steps became tiny and it lingered. Training a real network meets both problems constantly, so practitioners almost never use plain gradient descent. They use an OptimizerThe rule that decides how to change a model's parameters at each training step, given the gradient. Gradient descent, momentum, and Adam are optimizers.Open in glossary built on one or both of two ideas.

Idea one: momentum

Picture a heavy ball rolling down the landscape instead of a hiker taking careful steps. The ball has speed. If the slope keeps pointing the same way, it accelerates. If the slope flips back and forth, as it does when you zig-zag across a ravine, the pushes cancel and the ball settles into the middle.

MomentumAn optimizer that keeps a running sum of past gradients, like a ball gathering speed downhill, so it moves faster along consistent slopes and damps zig-zags.Open in glossary does exactly this. It keeps a velocity vv, a running sum of recent gradients, and steps along the velocity instead of the raw gradient:

v←β v+∇L,θ←θ−η vv \leftarrow \beta\,v + \nabla L, \qquad \theta \leftarrow \theta - \eta\,v

Here θ\theta (theta) stands for all the parameters, ∇L\nabla L is the gradient, η\eta is the learning rate, and β\beta is how much of the old velocity to keep, typically 0.90.9. On a long, steady slope the velocity grows to about 1/(1−β)=101/(1-\beta) = 10 times a single gradient, so momentum moves up to ten times faster there.

Idea two: a step size for each parameter

The stretched bowl is steep in one direction and gentle in the other. One learning rate cannot suit both. What if each parameter had its own?

RMSProp keeps, for every parameter, a running average ss of its squared gradients, which measures how big that parameter’s gradients usually are. It then divides each gradient by the square root of that average:

s←α s+(1−α) g2,θ←θ−η gs+ϵs \leftarrow \alpha\,s + (1-\alpha)\,g^2, \qquad \theta \leftarrow \theta - \eta\,\frac{g}{\sqrt{s} + \epsilon}

Here gg is the gradient for one parameter, α\alpha is typically 0.990.99, and ϵ\epsilon is a tiny number that prevents division by zero. The effect is that every parameter moves roughly η\eta per step whatever the size of its gradient. Steep directions get reined in, and gentle or nearly flat directions get boosted. One catch: ss starts at zero, so early on it underestimates the typical gradient and the steps come out larger than η\eta. With α=0.99\alpha = 0.99 the very first step is 10η10\eta.

Both at once: Adam

AdamA widely used optimizer that combines momentum with a separate, automatically scaled step size for each parameter. Most large neural networks are trained with Adam or a close variant.Open in glossary combines a momentum-style running average of gradients with RMSProp’s per-parameter scaling, plus a small correction for the first few steps, when both running averages are still biased toward zero. With its usual defaults it works reasonably well on a wide range of problems, which is why it, and a variant called AdamW, are the standard choice for training large neural networks, including language models.

The Adam updateOptional

For each parameter, with gradient gtg_t at step tt:

m←β1m+(1−β1) gts←β2s+(1−β2) gt2m^=m1−β1 t,s^=s1−β2 tθ←θ−η m^s^+ϵ\begin{aligned} m &\leftarrow \beta_1 m + (1-\beta_1)\,g_t \\ s &\leftarrow \beta_2 s + (1-\beta_2)\,g_t^2 \\ \hat{m} &= \frac{m}{1-\beta_1^{\,t}}, \qquad \hat{s} = \frac{s}{1-\beta_2^{\,t}} \\ \theta &\leftarrow \theta - \eta\,\frac{\hat{m}}{\sqrt{\hat{s}} + \epsilon} \end{aligned}

Common defaults are β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, and ϵ=10−8\epsilon = 10^{-8}. Because mm and ss start at zero, they are too small early on; dividing by 1−βt1-\beta^t corrects for that. On the very first step m^=g\hat{m} = g and s^=g2\hat{s} = g^2, so every parameter moves by almost exactly η\eta. The demo below implements these update rules exactly as written, matching PyTorch’s conventions.

Race them

Each optimizer gets a base step size suited to it, shown next to its name, and the slider scales all of them together.

Optimizer race

Four optimizers start from the same point on the same surface. Tap the map to move the start.

The Rosenbrock function: easy to find the valley, slow to follow its gently curving floor to the minimum.

Tap to move the start
  • Gradient descent
  • Momentum
  • RMSProp
  • Adam
Surface
1.00x
Racers
Steps0
Gradient descent7.016
Momentum7.016
RMSProp7.016
Adam7.016

Try this

  • On the Banana valley, press Race and turn on Fast forward. All four reach the valley floor, but momentum and Adam reach the minimum far sooner. The loss chart shows them dropping while plain gradient descent slides slowly.
  • Scale the step sizes to 4x and race again. Plain gradient descent diverges within a few steps. RMSProp survives but keeps jittering, because its steps never shrink below about the learning rate. Momentum and Adam still converge.
  • Switch to Saddle point. RMSProp and Adam slide off the ridge within a few steps, before they even reach the flat middle; plain gradient descent and momentum drift into the middle and linger there for dozens.
  • Switch to Many valleys. Here momentum carries enough speed to roll past the first shallow valley into the deeper one, while the others stop in the first valley they meet.

Each result traces back to one of the two ideas. Momentum’s velocity is what carries it along the banana’s curved floor and over the small rise between valleys. Per-parameter scaling is what turns the saddle’s tiny gradient into a full-sized step.

What practitioners actually do

Large models are trained with Adam or AdamW on small random batches of data, which is the stochastic gradient descent you met in the first lesson with Adam’s update rule on top. The learning rate is still the most important setting. It is also rarely fixed: typical runs start with a short warmup from a tiny learning rate, then gradually decay it, so training can move quickly early and settle precisely at the end.

Key ideas

  • Plain gradient descent zig-zags in narrow ravines and crawls on flat ground.
  • Momentum steps along a running sum of gradients. It speeds up along consistent slopes and cancels back-and-forth oscillation.
  • RMSProp gives each parameter its own step size by dividing by the typical size of its gradients, which helps on flat regions and saddle points.
  • Adam combines both ideas and is the standard optimizer for large networks.
  • No optimizer is best on every surface. The learning rate still matters most.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1How does momentum change the update compared with plain gradient descent?
2Near a saddle point the gradient is tiny. Why do RMSProp and Adam escape it much faster than plain gradient descent?
3Which optimizer is most commonly used to train large neural networks today?

Progress is saved in this browser only.

Up nextClassification and probability
Next
How machines learn
  1. 1The learning loop
  2. 2Linear regression and loss
  3. 3Gradient descent
  4. 4Better optimizers
  5. 5Classification and probability
  6. 6Overfitting and generalization
  7. 7Measuring a classifier
  8. 8Finding groups: k-means

Try "embedding", "softmax", "overfitting", or "backpropagation".