In the last lesson, plain gradient descent had two clear weaknesses. On the stretched bowl and the banana-shaped valley, it zig-zagged across the narrow direction while creeping along the long one. Near the saddle point, where the ground is almost flat, its steps became tiny and it lingered. Training a real network meets both problems constantly, so practitioners almost never use plain gradient descent. They use an OptimizerThe rule that decides how to change a model's parameters at each training step, given the gradient. Gradient descent, momentum, and Adam are optimizers.Open in glossary built on one or both of two ideas.
Idea one: momentum
Picture a heavy ball rolling down the landscape instead of a hiker taking careful steps. The ball has speed. If the slope keeps pointing the same way, it accelerates. If the slope flips back and forth, as it does when you zig-zag across a ravine, the pushes cancel and the ball settles into the middle.
MomentumAn optimizer that keeps a running sum of past gradients, like a ball gathering speed downhill, so it moves faster along consistent slopes and damps zig-zags.Open in glossary does exactly this. It keeps a velocity , a running sum of recent gradients, and steps along the velocity instead of the raw gradient:
Here (theta) stands for all the parameters, is the gradient, is the learning rate, and is how much of the old velocity to keep, typically . On a long, steady slope the velocity grows to about times a single gradient, so momentum moves up to ten times faster there.
Idea two: a step size for each parameter
The stretched bowl is steep in one direction and gentle in the other. One learning rate cannot suit both. What if each parameter had its own?
RMSProp keeps, for every parameter, a running average of its squared gradients, which measures how big that parameter’s gradients usually are. It then divides each gradient by the square root of that average:
Here is the gradient for one parameter, is typically , and is a tiny number that prevents division by zero. The effect is that every parameter moves roughly per step whatever the size of its gradient. Steep directions get reined in, and gentle or nearly flat directions get boosted. One catch: starts at zero, so early on it underestimates the typical gradient and the steps come out larger than . With the very first step is .
Both at once: Adam
AdamA widely used optimizer that combines momentum with a separate, automatically scaled step size for each parameter. Most large neural networks are trained with Adam or a close variant.Open in glossary combines a momentum-style running average of gradients with RMSProp’s per-parameter scaling, plus a small correction for the first few steps, when both running averages are still biased toward zero. With its usual defaults it works reasonably well on a wide range of problems, which is why it, and a variant called AdamW, are the standard choice for training large neural networks, including language models.
The Adam updateOptional
For each parameter, with gradient at step :
Common defaults are , , and . Because and start at zero, they are too small early on; dividing by corrects for that. On the very first step and , so every parameter moves by almost exactly . The demo below implements these update rules exactly as written, matching PyTorch’s conventions.
Race them
Each optimizer gets a base step size suited to it, shown next to its name, and the slider scales all of them together.
Optimizer race
Four optimizers start from the same point on the same surface. Tap the map to move the start.
The Rosenbrock function: easy to find the valley, slow to follow its gently curving floor to the minimum.
- Gradient descent
- Momentum
- RMSProp
- Adam
Try this
- On the Banana valley, press Race and turn on Fast forward. All four reach the valley floor, but momentum and Adam reach the minimum far sooner. The loss chart shows them dropping while plain gradient descent slides slowly.
- Scale the step sizes to 4x and race again. Plain gradient descent diverges within a few steps. RMSProp survives but keeps jittering, because its steps never shrink below about the learning rate. Momentum and Adam still converge.
- Switch to Saddle point. RMSProp and Adam slide off the ridge within a few steps, before they even reach the flat middle; plain gradient descent and momentum drift into the middle and linger there for dozens.
- Switch to Many valleys. Here momentum carries enough speed to roll past the first shallow valley into the deeper one, while the others stop in the first valley they meet.
Each result traces back to one of the two ideas. Momentum’s velocity is what carries it along the banana’s curved floor and over the small rise between valleys. Per-parameter scaling is what turns the saddle’s tiny gradient into a full-sized step.
What practitioners actually do
Large models are trained with Adam or AdamW on small random batches of data, which is the stochastic gradient descent you met in the first lesson with Adam’s update rule on top. The learning rate is still the most important setting. It is also rarely fixed: typical runs start with a short warmup from a tiny learning rate, then gradually decay it, so training can move quickly early and settle precisely at the end.
Key ideas
- Plain gradient descent zig-zags in narrow ravines and crawls on flat ground.
- Momentum steps along a running sum of gradients. It speeds up along consistent slopes and cancels back-and-forth oscillation.
- RMSProp gives each parameter its own step size by dividing by the typical size of its gradients, which helps on flat regions and saddle points.
- Adam combines both ideas and is the standard optimizer for large networks.
- No optimizer is best on every surface. The learning rate still matters most.