Tensors, shapes and broadcasting
Store data as multi-dimensional arrays, reason about their shapes, and let broadcasting do the loops.
- Describe tensors by rank and shape, including the batch dimension
- Predict the result shape of matrix multiplication, broadcasting and reductions
- Reshape, transpose and flatten tensors without losing track of the data
Neural networks speak one language: tensors - arrays of numbers with any number of dimensions. In numpy (and PyTorch, which copies numpy’s style) a tensor has a shape, a tuple giving the size of each dimension:
| Rank | Example | Shape |
|---|---|---|
| 0 (scalar) | a loss value | () |
| 1 (vector) | one doodle’s 25 pixels | (25,) |
| 2 (matrix) | a batch of 32 doodles | (32, 25) |
| 3 | a batch of 32 grayscale 5×5 images | (32, 5, 5) |
| 4 | 32 color images, 3 channels, 64×64 | (32, 3, 64, 64) in PyTorch |
The first dimension is almost always the batch: networks process many examples at once, because one big matrix multiplication is much faster than many small ones.
1import numpy as np
2
3doodles = np.arange(2 * 3 * 4).reshape(2, 3, 4) # 2 doodles, 3 rows, 4 columns
4print(doodles.shape, doodles.ndim, doodles.size)
5flat = doodles.reshape(2, -1) # -1: "work this one out"
6print(flat.shape)
7print(flat[1, :5])
8print(doodles.transpose(0, 2, 1).shape)
9print(doodles.sum(axis=0).shape, doodles.sum(axis=(1, 2)))(2, 3, 4) 3 24 (2, 12) [12 13 14 15 16] (2, 4, 3) (3, 4) [ 66 210]
Reductions like sum, mean and max take an axis: that dimension disappears from the shape. Pass keepdims=True to keep it as size 1 - handy when you want to divide by the result.
reshape never moves data; it reinterprets the same numbers in a new shape, filling the last dimension first. transpose swaps dimensions around.
Matrix multiplication
The workhorse of every layer is @, matrix multiplication. Shapes must line up on the inside, and the outside dimensions survive:
For a layer: inputs (batch, in_features) @ (in_features, out_features) gives (batch, out_features). Each output is a weighted sum - one dot product per row and column.
Broadcasting
Element-wise operations between tensors of different shapes broadcast: numpy lines the shapes up from the right, and a dimension of size 1 (or a missing one) stretches to match. No data is copied.
| Shapes | Result |
|---|---|
(32, 10) + (10,) | (32, 10) - one bias per column, added to every row |
(32, 10) − (32, 1) | (32, 10) - one value per row |
(3, 1) × (1, 4) | (3, 4) - an outer product |
(32, 10) + (32,) | error: 10 and 32 don’t match |
1import numpy as np
2
3X = np.array([[1.0, 2.0, 3.0],
4 [4.0, 5.0, 6.0]]) # 2 examples, 3 features
5bias = np.array([10.0, 20.0, 30.0])
6print(X + bias)
7row_means = X.mean(axis=1, keepdims=True) # shape (2, 1)
8print(X - row_means)
9W = np.ones((3, 4))
10print((X @ W).shape)
11try:
12 X + np.array([1.0, 2.0])
13except ValueError as error:
14 print("Error:", str(error).split(" with ")[0])[[11. 22. 33.] [14. 25. 36.]] [[-1. 0. 1.] [-1. 0. 1.]] (2, 4) Error: operands could not be broadcast together
Try it
Will it broadcast?
Line the shapes up from the right: each pair of dimensions must be equal, or one of them must be 1 (or missing). For @, the inner dimensions must match.
“(64, 128) + (128,)”
“(64, 128) + (64,)”
“(64, 1) * (1, 10)”
“(8, 3) @ (3, 5)”
“(8, 3) @ (8, 3)”
“(16, 3, 32, 32) + (3, 1, 1)”
“(5, 4) - (5, 4).mean(axis=1)”
Key takeaways
Tensors are n-dimensional arrays; their shape lists each dimension’s size, batch first.
@multiplies (N, K) by (K, M) into (N, M) - the core of every layer.Broadcasting aligns shapes from the right and stretches size-1 dimensions.
Reductions remove the axis you name;
keepdims=Truekeeps it as size 1.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Standardize the doodle features
Each input line is one doodle’s features (same count on every line). Standardize each column: subtract the column mean and divide by the column standard deviation (np.std, the default). Print each row with values to 2 decimals separated by spaces, then the shape of the result.
- Three doodles
- Four doodles
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
A layer for a whole batch
Each input line is one example with 3 features. Stack them into a batch X, then compute a layer’s outputs X @ W + b with the given W (3 inputs → 2 outputs) and b, for all examples at once (no loop over rows). Print the shapes of X and of the output, then each output row to 1 decimal.
- Two examples
- Three examples
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…