Backpropagation
Compute every gradient in a network with one backward sweep of the chain rule.
- Apply the chain rule along a computation graph
- Compute gradients backward as “upstream gradient × local derivative”, summing over branches
- Derive the gradients of a linear layer and check them numerically
To improve a weight, training needs its gradient: how much the loss changes when that weight changes a little, . A network can have billions of weights. Backpropagation gets every one of those gradients in a single backward pass that costs about the same as the forward pass. It’s the algorithm that makes deep learning possible.
The idea is the chain rule. If depends on , and depends on , then:
Break any computation into simple steps - a computation graph - and every node only needs its local derivative. Going backward from the loss, each node’s gradient is the gradient flowing in from above (upstream) times its local derivative. When a value feeds into several places, the gradients from each place add up.
The local derivatives you need are few:
| Node | Local derivative | So the gradient... |
|---|---|---|
| c = a + b | ∂c/∂a = 1, ∂c/∂b = 1 | is copied to both inputs |
| c = a − b | 1 and −1 | is copied, negated for b |
| c = a × b | ∂c/∂a = b, ∂c/∂b = a | is multiplied by the other input |
| c = a² | 2a | is scaled by 2a |
| c = relu(a) | 1 if a > 0 else 0 | passes through or is blocked |
| c = σ(a) | σ(a)(1 − σ(a)) | shrinks (at most ×0.25) |
Try it
Backprop stepper
Run each graph forward one node at a time, then backward. Before some gradients are revealed, predict them: take the gradient of the node above and multiply by the local derivative.
In the last graph, watch what a ReLU that’s switched off does to everything below it.
A one-weight model predicts and is scored with , for and target .
| Node | Value | Gradient dL/d· | How |
|---|---|---|---|
| w (input) | ? | ||
| x (input) | ? | ||
| b (input) | ? | ||
| y (input) | ? | ||
| z = w × x | ? | ||
| p = z + b | ? | ||
| d = p − y | ? | ||
| L = d² | ? |
Backprop through a whole layer
Real layers work on matrices, but the rules are the same. For a linear layer with input batch (N × in), and the upstream gradient (N × out):
A good sanity check: every gradient has the same shape as the thing it’s the gradient of. And dX is what you pass on, upstream, to the layer before.
1import numpy as np
2
3rng = np.random.default_rng(1)
4X, W, b = rng.normal(size=(4, 3)), rng.normal(size=(3, 2)), rng.normal(size=2)
5target = rng.normal(size=(4, 2))
6
7def loss(W):
8 return np.mean((X @ W + b - target) ** 2)
9
10# Backprop: L = mean(D²) with D = XW + b - target, so dL/dD = 2D / D.size
11G = 2 * (X @ W + b - target) / target.size
12dW = X.T @ G
13
14# Numerical check: nudge each weight by h and watch the loss
15h = 1e-6
16numeric = np.zeros_like(W)
17for i in range(W.shape[0]):
18 for j in range(W.shape[1]):
19 nudged = W.copy()
20 nudged[i, j] += h
21 numeric[i, j] = (loss(nudged) - loss(W)) / h
22print(dW.shape, np.abs(dW - numeric).max() < 1e-5)(3, 2) True
Key takeaways
Backprop applies the chain rule backward through a computation graph, getting every gradient in one sweep.
Each node’s gradient = upstream gradient × local derivative; branches add up.
For Y = XW + b: dW = Xᵀ G, db = sum of G’s rows, dX = G Wᵀ.
Check gradients numerically when you write anything custom.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Gradients of a sigmoid neuron
A neuron computes and is scored with . Each input line is w x b y. Backpropagate by hand (chain rule!) and print the loss and the gradients dL/dw and dL/db to 4 decimals:
L=0.2500 dw=-0.5000 db=-0.2500The starter checks your answer numerically, so you can see whether you got it right.
- The stepper’s neuron
- Two more
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Backward pass of a linear layer
The starter builds a batch X (from the input lines), weights W, bias b, and an upstream gradient G. Compute dW, db and dX for and print their shapes, then each one’s values to 2 decimals (row by row, space-separated). Finally print whether a numerical check of dW agrees (check: True).
- Two examples
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…