Convolutional networks
Slide small learned filters over images, pool the results, and build the CNNs that taught computers to see.
- Explain why convolutions suit images: local patterns, shared weights
- Compute convolution outputs and output sizes with padding and stride
- Describe how conv, activation and pooling layers stack into a CNN, and count their parameters
Feed a 224×224 color photo to a fully connected layer with 1,000 neurons and you need 150 million weights - for one layer. Worse, a cat in the top-left corner and a cat in the bottom-right would be learned separately.
A convolutional layer fixes both problems with a small kernel (filter), say 3×3, that slides across the image. At each position it computes a weighted sum of the pixels under it. The same kernel - the same 9 weights - is used everywhere (weight sharing), and each output only looks at a small neighborhood (locality). The output is a feature map showing where the kernel’s pattern appears.
The kernels aren’t designed by hand: the network learns them by backprop, just like any other weights. Early layers end up with edge and color detectors much like the classic ones below.
Try it
Slide a kernel
Choose a kernel and click any output pixel to see the multiplication behind it. Edge kernels light up where brightness changes; notice that a vertical-edge detector ignores horizontal edges completely.
Output at row 6, column 6 = sum of (pixel × weight) over the highlighted 3×3 window (pixels outside the image count as 0):
220×-1 + 220×0 + 220×1 + 20×-2 + 20×0 + 20×2 + 20×-1 + 20×0 + 220×1 = 200
1import numpy as np
2
3def conv2d(image, kernel, stride=1, padding=0):
4 image = np.pad(image, padding)
5 k = kernel.shape[0]
6 out_size = (image.shape[0] - k) // stride + 1
7 out = np.zeros((out_size, out_size))
8 for i in range(out_size):
9 for j in range(out_size):
10 patch = image[i * stride:i * stride + k, j * stride:j * stride + k]
11 out[i, j] = np.sum(patch * kernel)
12 return out
13
14doodle = np.array([[0, 0, 1, 1, 0],
15 [0, 0, 1, 1, 0],
16 [0, 0, 1, 1, 0],
17 [0, 0, 1, 1, 0],
18 [0, 0, 1, 1, 0]], dtype=float)
19vertical_edges = np.array([[-1, 0, 1], [-2, 0, 2], [-1, 0, 1]])
20print(conv2d(doodle, vertical_edges))
21print(conv2d(doodle, vertical_edges, padding=1).shape, conv2d(doodle, vertical_edges, stride=2).shape)[[ 4. 4. -4.] [ 4. 4. -4.] [ 4. 4. -4.]] (5, 5) (2, 2)
Positive values mark where the doodle turns from dark to bright (left to right), negative where it turns back. (Strictly, deep learning “convolution” doesn’t flip the kernel - mathematicians call it cross-correlation - but since kernels are learned, it doesn’t matter.)
Output size. For input width , kernel , padding and stride :
- Padding adds a border of zeros; with (“same” padding) the size is preserved.
- Stride skips positions: stride 2 roughly halves each dimension.
- Channels: real images have 3 channels (RGB), and layers produce many feature maps. A kernel spans all input channels, so a layer with inputs, outputs and K×K kernels has parameters - tiny compared with a dense layer.
Pooling and the CNN recipe
Pooling shrinks feature maps by summarizing small windows - usually taking the maximum of each 2×2 block (max pooling). It makes the representation smaller and a little tolerant to small shifts: the feature was somewhere in that window.
Try it
Max pooling
Step the 2×2 window across the feature map with stride 2. Each output keeps only the strongest response in its window.
max(1, 3, 4, 2) = 4
A classic CNN repeats conv → ReLU → pool a few times, so feature maps get smaller but more numerous (more channels), and each neuron sees a larger region of the original image - its receptive field grows layer by layer. Then it flattens the result and finishes with dense layers to classify. LeNet (1998) read handwritten digits this way; AlexNet, VGG and ResNet scaled the same recipe up.
Key takeaways
Convolutions slide small, shared, learned kernels over the input; each output sees only a local patch.
Output width = ⌊(W − K + 2P)/S⌋ + 1; “same” padding preserves size, stride shrinks it.
A conv layer has K·K·C_in·C_out + C_out parameters - far fewer than a dense layer.
CNNs stack conv → ReLU → pool, growing channels and receptive fields, then classify.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Write a convolution
The first input line is stride padding; the remaining lines are a square image. Implement conv2d(image, kernel, stride, padding) (no kernel flip, zero padding) and apply the horizontal-edge kernel [[-1, -2, -1], [0, 0, 0], [1, 2, 1]]. Print the output shape, then its rows with values as integers separated by spaces.
- A horizontal bar
- Padded and strided
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Trace a CNN’s shapes
The first input line is the input shape channels height width. Each following line is a layer: conv out_channels kernel stride padding, pool size (max pooling with stride = size), or flatten. Print the shape after each layer and that layer’s parameter count, then the total:
1conv 16 3 1 1 -> (16, 32, 32) params 448
2pool 2 -> (16, 16, 16) params 0
3flatten -> (4096,) params 0
4total params 448- A small CNN
- LeNet-style
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…