Um momento
0x70Lesson 8 of 17

Gradient descent and optimizers

Turn gradients into weight updates with SGD, momentum and Adam, and tune the learning rate.

28 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Explain gradient descent and the role of the learning rate
  • Compare full-batch, stochastic and mini-batch gradient descent
  • Implement momentum and Adam, and choose a learning-rate schedule

Gradients point uphill - the direction in which the loss grows fastest. So to learn, step the other way:

θ←θ−η ∇θL\theta \leftarrow \theta - \eta \, \nabla_\theta L

where θ\theta stands for all the parameters and η\eta (eta) is the learning rate. It is the single most important setting in deep learning: too small and training crawls; too large and the loss bounces around or explodes.

Try it

Roll the ball downhill

Try several learning rates. Find one that converges smoothly, one that crawls, and one that overshoots back and forth - or flies off.

x = -1.50
loss = 11.13
gradient = -4.50
steps taken: 0

How much data per step?

  • Full-batch gradient descent uses the whole dataset for every step: accurate but slow.
  • Stochastic gradient descent (SGD) uses one example: fast, very noisy.
  • Mini-batch - 32 to a few thousand examples - is what everyone actually does. Its noise even helps escape poor regions. (Confusingly, people call this “SGD” too.)

One pass over the whole dataset is an epoch.

Momentum and Adam

Plain gradient descent zigzags in long, narrow valleys: the steep direction makes it bounce, the shallow direction makes it crawl. Momentum keeps a running velocity, so consistent directions build up speed and the zigzags cancel out:

v←βv+∇L,θ←θ−ηvv \leftarrow \beta v + \nabla L, \qquad \theta \leftarrow \theta - \eta v (typically β=0.9\beta = 0.9)

Adam (2014), the most popular default, also gives each parameter its own step size, scaled down where gradients have been large:

m←β1m+(1−β1)g,s←β2s+(1−β2)g2m \leftarrow \beta_1 m + (1 - \beta_1) g, \qquad s \leftarrow \beta_2 s + (1 - \beta_2) g^2

m^=m1−β1t,s^=s1−β2t,θ←θ−ηm^s^+ϵ\hat{m} = \frac{m}{1 - \beta_1^t}, \quad \hat{s} = \frac{s}{1 - \beta_2^t}, \qquad \theta \leftarrow \theta - \eta \frac{\hat{m}}{\sqrt{\hat{s}} + \epsilon}

with β1=0.9,β2=0.999,ϵ=10−8\beta_1 = 0.9, \beta_2 = 0.999, \epsilon = 10^{-8}. The hats are bias corrections: m and s start at zero, which would make the first steps too small. AdamW adds decoupled weight decay and is the default for training transformers.

race.py
1import numpy as np
2
3def gradient(p):                      # a long, narrow valley: f = 0.5 (x² + 25 y²)
4    return np.array([p[0], 25 * p[1]])
5
6def loss(p):
7    return 0.5 * (p[0] ** 2 + 25 * p[1] ** 2)
8
9def run(lr, beta=0.0):                # beta = 0 is plain gradient descent
10    p, v = np.array([10.0, 1.0]), np.zeros(2)
11    for step in range(1, 1001):
12        v = beta * v + gradient(p)
13        p = p - lr * v
14        if loss(p) > 1e6:
15            return "diverged"
16        if loss(p) < 1e-6:
17            return step
18    return "too slow"
19
20print("SGD at its best, lr 0.077:   ", run(0.077))
21print("SGD, lr 0.09:                ", run(0.09))
22print("momentum 0.5, lr 0.1:        ", run(0.1, beta=0.5))
23print("momentum 0.9, lr 0.1:        ", run(0.1, beta=0.9))
Output
SGD at its best, lr 0.077:    113
SGD, lr 0.09:                 diverged
momentum 0.5, lr 0.1:         28
momentum 0.9, lr 0.1:         160

Plain gradient descent can’t raise its learning rate past 0.08 - the steep y-direction would blow up - so the gentle x-direction crawls: 113 steps at best. With momentum 0.5 it takes 28. But momentum is a setting too: at 0.9, velocity builds up so much that the ball overshoots and circles the minimum for 160 steps. The popular 0.9 shines on big, noisy problems; on a tiny, clean one like this, less is more.

Learning-rate schedules

The best learning rate changes during training: big steps early to make progress, small steps late to settle into a minimum. Common schedules:

  • Step decay: divide by 10 every N epochs.
  • Cosine decay: glide smoothly from the starting rate down to near zero.
  • Warmup: start tiny and ramp up over the first few hundred steps, then decay. Transformers almost always use warmup, because early updates on random weights are unstable.

Key takeaways

  • Gradient descent steps against the gradient: θ ← θ − η∇L; the learning rate is the key setting.

  • Mini-batch gradient descent balances speed and noise; an epoch is one pass over the data.

  • Momentum accumulates velocity; Adam adds per-parameter step sizes with bias correction.

  • Schedules - warmup, then decay - adapt the learning rate during training.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

The momentum race

+25 XP

Each input line is sgd lr or momentum lr beta. Starting from (10, 1) on f(x,y)=0.5(x2+25y2)f(x, y) = 0.5(x^2 + 25y^2), count the steps until the loss drops below 1e-6 (at most 1000 steps). Print sgd lr=0.077: 113 steps, or momentum lr=0.1 beta=0.5: 28 steps, or diverged instead of the steps if the loss ever exceeds 1e6, or too slow if 1000 steps aren’t enough.

  • Five racers
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Implement Adam

+25 XP

Minimize f(θ)=(θ−3)2f(\theta) = (\theta - 3)^2 with Adam (β1 = 0.9, β2 = 0.999, ε = 1e-8). The input has the starting θ and the learning rate. Print θ to 4 decimals after steps 1, 2, 3 and after step 200, like step 1: theta=0.1000.

  • From 0 with lr 0.1
  • From 10 with lr 0.5
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: