Um momento
0x70Lesson 8 of 12

How networks learn: loss and gradient descent

Roll downhill on the loss, pick a learning rate that neither crawls nor explodes, and see how backpropagation trains every weight.

22 min 6-question quiz 2 code exercises
By the end of this lesson you can
  • Explain the gradient descent update rule
  • Predict what too small and too large a learning rate do
  • Describe backpropagation, epochs and batches in plain words

Training a network means finding weights that make the loss small. Picture the loss as a hilly landscape: every possible setting of the weights is a spot on the ground, and the height is how wrong the model is there. We want the bottom of the valley.

You’re standing on the hillside in thick fog. You can’t see the valley, but you can feel which way the ground slopes under your feet. So you take a step downhill, feel again, and repeat. That is gradient descent.

The gradient is the slope: which direction makes the loss go up, and how steeply. So we step the opposite way:

wnew=wold−η⋅gradientw_{\text{new}} = w_{\text{old}} - \eta \cdot \text{gradient}

The Greek letter η\eta (“eta”) is the learning rate: how big a step to take.

Try it

Roll the ball downhill

Press Step to take one gradient step and watch the ball head for the bottom of the curve. Then experiment with the learning rate: at 0.05 the ball crawls; around 0.5-1.0 it gets there fast; push it above 1.0 and watch it overshoot back and forth. Above 2 it would fly off entirely.

x = -1.50
loss = 11.13
gradient = -4.50
steps taken: 0
gradient_descent.py
1# Loss: (w - 3)² has its minimum at w = 3. Its slope (gradient) is 2 * (w - 3).
2weight = -1.0
3learning_rate = 0.2
4for step in range(1, 9):
5    gradient = 2 * (weight - 3)
6    weight -= learning_rate * gradient
7    print(f"step {step}: weight = {weight:.3f}, loss = {(weight - 3) ** 2:.4f}")
Output
step 1: weight = 0.600, loss = 5.7600
step 2: weight = 1.560, loss = 2.0736
step 3: weight = 2.136, loss = 0.7465
step 4: weight = 2.482, loss = 0.2687
step 5: weight = 2.689, loss = 0.0967
step 6: weight = 2.813, loss = 0.0348
step 7: weight = 2.888, loss = 0.0125
step 8: weight = 2.933, loss = 0.0045

Fitting a line by gradient descent

The same idea fits a whole model. Below, a line y=wx+by = wx + b starts flat at w=b=0w = b = 0 and gradient descent adjusts both parameters at once to shrink the mean squared error. The right panel shows the loss falling step by step.

Try it

Train a line

Pick a learning rate and press “10 steps” a few times. With 0.01 the line creeps; with 0.05 it gets there several times faster. Then try 0.12: each step overshoots further than the last, the loss explodes, and the widget stops you.

Learning rate:
y = 0.00x + 0.00
Loss (log scale) over 0 steps · now 48.7

Backpropagation, epochs and batches

A real network has millions of weights, so it needs the gradient for every single one. Backpropagation computes them all efficiently. It starts at the output, where the error is known, and works backward layer by layer, using the chain rule from calculus to share out the blame: how much did each weight contribute to the mistake?

Training then repeats:

  • A batch is a small group of examples (say 32). The gradient is averaged over the batch and the weights take one step.
  • An epoch is one full pass through the training data. With 10,000 examples and batches of 100, one epoch is 100 steps.
  • Training runs for many epochs, while you watch the loss on held-out data to stop before overfitting.

Key takeaways

  • Gradient descent repeats w←w−η⋅gradientw \leftarrow w - \eta \cdot \text{gradient}: a small step downhill on the loss.

  • The learning rate η\eta sets the step size: too small crawls, too large overshoots and diverges.

  • Backpropagation computes the gradient for every weight, working backward from the output.

  • A batch gives one update; an epoch is one full pass through the data.

Lesson quiz

6 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Take gradient steps

+25 XP

The loss is L(w)=(w−3)2L(w) = (w - 3)^2 with gradient 2(w−3)2(w - 3). The input is one line: start learning_rate steps. Run that many gradient descent steps and print the weight after each one with 4 decimal places. Then print converged if the final weight is within 0.01 of 3, diverged if it is farther from 3 than the start was, and still going otherwise.

  • Good rate
  • Too small
  • Too big
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Plan a training run

+25 XP

The input is examples batch_size epochs. The last batch of an epoch may be smaller than the others. Print batches per epoch: B, last batch size: S and total updates: U.

  • Even split
  • Leftover
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: