Layers and the forward pass
Stack layers into a multilayer perceptron, run a batch through it, and count its parameters.
- Describe a multilayer perceptron as a chain of x @ W + b and activations
- Track tensor shapes through every layer of a forward pass
- Count a network’s parameters and estimate its memory
A multilayer perceptron (MLP) - also called a fully connected or dense network - chains layers. For Neuro’s doodles (5×5 = 25 pixels) and 4 kinds of doodle, a small MLP looks like:
input (batch, 25)
→ Linear 25→16 → ReLU hidden (batch, 16)
→ Linear 16→4 → softmax output (batch, 4) probabilitiesRunning inputs through all the layers to get predictions is the forward pass. In code it’s strikingly short:
1import numpy as np
2
3rng = np.random.default_rng(42)
4W1, b1 = rng.normal(0, 0.3, size=(25, 16)), np.zeros(16)
5W2, b2 = rng.normal(0, 0.3, size=(16, 4)), np.zeros(4)
6
7def softmax(z):
8 exps = np.exp(z - z.max(axis=1, keepdims=True))
9 return exps / exps.sum(axis=1, keepdims=True)
10
11X = rng.integers(0, 2, size=(3, 25)).astype(float) # 3 random doodles
12hidden = np.maximum(0, X @ W1 + b1) # (3, 16)
13probabilities = softmax(hidden @ W2 + b2) # (3, 4)
14print(X.shape, hidden.shape, probabilities.shape)
15print(np.round(probabilities.sum(axis=1), 6))
16print(probabilities.argmax(axis=1))(3, 25) (3, 16) (3, 4) [1. 1. 1.] [3 2 3]
Notice the softmax works row by row (axis=1, keepdims=True), so every example gets its own distribution. argmax picks the most likely class. With random weights, Neuro’s guesses are, of course, nonsense - training is how they get good.
Counting parameters
A linear layer from to features has weights plus biases. Activations have no parameters. For the doodle MLP: 25·16 + 16 + 16·4 + 4 = 484 parameters.
Real networks are bigger: a classic MNIST MLP, 784→128→64→10, has about 109,000 parameters; GPT-3 has 175 billion. Each 32-bit float takes 4 bytes, so parameters × 4 is the memory just to store the weights - training needs several times more for gradients and optimizer state.
Key takeaways
An MLP chains linear layers and activations: h = relu(x @ W1 + b1), out = softmax(h @ W2 + b2).
The batch dimension rides along untouched; each layer changes only the feature dimension.
A linear layer has in × out weights plus out biases; activations have none.
Parameters × 4 bytes is the storage for float32 weights; training needs more.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Run Neuro’s forward pass
The starter has a tiny trained network (3 inputs → 4 hidden → 2 outputs, ReLU then softmax). Each input line is one example. Run the whole batch through it in one go, and for each example print the two probabilities to 3 decimals and the predicted class (argmax):
0.905 0.095 -> 0- Three examples
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Parameter budget
Each input line describes an MLP as layer sizes, like 784 128 64 10. Print the number of weights, biases, the total, and the memory for float32 parameters in megabytes (1 MB = 1,000,000 bytes) to 2 decimals:
784 128 64 10: weights=109184 biases=202 total=109386 memory=0.44 MB- Two networks
- A wide one
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…