Um momento
0x80Lesson 9 of 17

Train a network end to end

Put forward pass, loss, backprop and updates together and train a network from scratch.

34 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Write the complete training loop: forward, loss, backward, update
  • Backpropagate through a two-layer network by hand in numpy
  • Shuffle data into mini-batches and track loss and accuracy over epochs

Time to put it all together and train Neuro’s first real network. Every deep learning training run, from XOR to Claude, is this loop:

1for each epoch:
2    shuffle the training data
3    for each mini-batch:
4        1. forward:  predictions = model(batch)
5        2. loss:     how wrong are they?
6        3. backward: gradient of the loss for every parameter
7        4. update:   parameters -= learning_rate * gradients
8    check progress (loss, accuracy)

XOR, learned this time

In the first lesson you hand-wired XOR. Now the network finds weights itself. It’s 2 → 8 (tanh) → 1 (sigmoid), trained with binary cross-entropy. The backward pass uses a famous shortcut: for sigmoid followed by binary cross-entropy (and likewise softmax with cross-entropy), the gradient of the loss with respect to the logits simplifies to just predictions − labels.

train_xor.py
1import numpy as np
2
3X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]], dtype=float)
4y = np.array([[0], [1], [1], [0]], dtype=float)
5
6rng = np.random.default_rng(3)
7W1, b1 = rng.normal(0, 1, size=(2, 8)), np.zeros(8)
8W2, b2 = rng.normal(0, 1, size=(8, 1)), np.zeros(1)
9lr = 0.5
10
11for epoch in range(2001):
12    # 1. forward
13    h = np.tanh(X @ W1 + b1)
14    p = 1 / (1 + np.exp(-(h @ W2 + b2)))
15    # 2. loss (binary cross-entropy)
16    loss = -np.mean(y * np.log(p) + (1 - y) * np.log(1 - p))
17    # 3. backward
18    d_logits = (p - y) / len(X)             # sigmoid + BCE shortcut
19    dW2, db2 = h.T @ d_logits, d_logits.sum(axis=0)
20    d_h = d_logits @ W2.T * (1 - h ** 2)    # back through tanh
21    dW1, db1 = X.T @ d_h, d_h.sum(axis=0)
22    # 4. update
23    W1 -= lr * dW1; b1 -= lr * db1
24    W2 -= lr * dW2; b2 -= lr * db2
25    if epoch % 500 == 0:
26        print(f"epoch {epoch:4}: loss {loss:.4f}")
27
28print("predictions:", np.round(p.ravel(), 2))
Output
epoch    0: loss 0.8362
epoch  500: loss 0.0086
epoch 1000: loss 0.0038
epoch 1500: loss 0.0023
epoch 2000: loss 0.0016
predictions: [0. 1. 1. 0.]

Trace the shapes in the backward pass: d_logits is (4, 1) like the logits; dW2 = h.T @ d_logits is (8, 1) like W2; d_h is (4, 8) like h, multiplied element-wise by tanh’s derivative 1 − h²; and dW1 is (2, 8) like W1. Every gradient matches its parameter’s shape.

Mini-batches and epochs

XOR has 4 examples, so every step used all of them. With real data you shuffle every epoch and walk through mini-batches. Shuffling matters: if the data were sorted by class, each batch would push the network toward one class at a time.

1for epoch in range(epochs):
2    order = rng.permutation(len(X))
3    for start in range(0, len(X), batch_size):
4        batch = order[start:start + batch_size]
5        loss, grads = forward_backward(X[batch], y[batch])
6        update(grads)

Watch the loss as you go - and hold back a validation set the network never trains on, to see whether it learns patterns or memorizes examples. That’s Unit 3’s topic.

Key takeaways

  • Training is a loop: forward, loss, backward, update - over shuffled mini-batches, epoch after epoch.

  • Softmax/sigmoid with cross-entropy gives the gradient predictions − labels at the logits.

  • Every gradient has the same shape as its parameter - a free bug check.

  • Track the loss and keep a validation set to see whether learning generalizes.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Write the backward pass

+25 XP

The starter trains a 2 → 8 → 1 network on a logic gate given by the input (each line: a b label), but the backward pass is missing, so it never learns. Fill in the gradients for W1, b1, W2, b2 (tanh hidden layer, sigmoid output, binary cross-entropy).

It prints the loss at epochs 0 and 2000, then the rounded predictions and the accuracy.

  • XOR
  • XNOR
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Shuffle into mini-batches

+25 XP

The input is examples batch_size epochs seed. Using rng = np.random.default_rng(seed), print the mini-batches of example indexes for every epoch: draw a fresh rng.permutation(examples) each epoch and slice it into batches (the last batch may be smaller). Print epoch 0: [3 0 2] [1 4] style lines, then steps per epoch: 2.

  • Five examples
  • Even split
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: