Deep learning with PyTorch
Map everything you built by hand onto PyTorch: tensors, autograd, modules, optimizers, data loaders and fine-tuning.
- Translate numpy tensors, your autograd engine and your training loop into PyTorch
- Define models with nn.Module, train them with an optimizer, and switch between train and eval modes
- Use GPUs, save and load weights, and fine-tune a pretrained model
You’ve built deep learning from scratch. In real projects you use a framework - most often PyTorch - that provides the same pieces, fast and on GPUs. The good news: you already understand every one of them.
| You built | PyTorch |
|---|---|
| numpy arrays | torch.tensor (same broadcasting, @, .shape) |
the Value autograd engine | requires_grad=True and loss.backward() |
x @ W + b layers | nn.Linear, nn.Conv2d, nn.LSTM, nn.MultiheadAttention |
| ReLU, softmax + cross-entropy | nn.ReLU, nn.CrossEntropyLoss (takes logits!) |
| SGD, momentum, Adam updates | torch.optim.SGD, torch.optim.AdamW |
| dropout, layer norm | nn.Dropout, nn.LayerNorm |
| shuffled mini-batches | DataLoader(dataset, batch_size=32, shuffle=True) |
PyTorch can’t run in this browser, so this lesson’s examples aren’t runnable here - install it with pip install torch and try them. Every output shown was checked against PyTorch 2.9.
1import torch
2
3w = torch.tensor(0.5, requires_grad=True)
4x, b, y = torch.tensor(2.0), torch.tensor(-1.0, requires_grad=True), torch.tensor(1.0)
5p = torch.sigmoid(w * x + b)
6loss = (p - y) ** 2
7loss.backward() # your Value.backward(), on tensors
8print(f"loss={loss.item():.4f} dw={w.grad.item():.4f} db={b.grad.item():.4f}")loss=0.2500 dw=-0.5000 db=-0.2500
Those are exactly the numbers from the backprop stepper’s sigmoid neuron. Now the training loop - the same four steps as yours, with zero_grad() doing the gradient reset you had to remember yourself:
1import torch
2from torch import nn
3
4torch.manual_seed(0)
5X = torch.tensor([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
6y = torch.tensor([0, 1, 1, 0])
7
8model = nn.Sequential(nn.Linear(2, 8), nn.Tanh(), nn.Linear(8, 2))
9loss_fn = nn.CrossEntropyLoss() # takes logits
10optimizer = torch.optim.Adam(model.parameters(), lr=0.05)
11
12model.train()
13for epoch in range(300):
14 optimizer.zero_grad() # reset gradients
15 loss = loss_fn(model(X), y) # forward + loss
16 loss.backward() # backward
17 optimizer.step() # update
18
19model.eval()
20with torch.no_grad(): # no graph needed for predictions
21 predictions = model(X).argmax(dim=1)
22print("predictions:", predictions.tolist(), f"loss: {loss.item():.3f}")
23print("parameters:", sum(p.numel() for p in model.parameters()))predictions: [0, 1, 1, 0] loss: 0.000 parameters: 42
For anything bigger than nn.Sequential can express, write a class:
1class DoodleNet(nn.Module):
2 def __init__(self):
3 super().__init__()
4 self.features = nn.Sequential(
5 nn.Conv2d(1, 16, kernel_size=3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
6 nn.Conv2d(16, 32, kernel_size=3, padding=1), nn.ReLU(), nn.MaxPool2d(2))
7 self.classifier = nn.Sequential(nn.Flatten(), nn.Dropout(0.3), nn.Linear(32 * 7 * 7, 10))
8
9 def forward(self, x): # x: (batch, 1, 28, 28)
10 return self.classifier(self.features(x))And the practical essentials:
- GPU:
device = "cuda" if torch.cuda.is_available() else "cpu", thenmodel.to(device)andbatch.to(device). - Modes:
model.train()turns dropout and batch-norm training behavior on;model.eval()turns them off. Wrap evaluation intorch.no_grad()to save memory. - Saving:
torch.save(model.state_dict(), "doodle.pt"), and latermodel.load_state_dict(torch.load("doodle.pt")).
1import torch
2from torch import nn
3
4features = nn.Sequential(
5 nn.Conv2d(1, 16, kernel_size=3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
6 nn.Conv2d(16, 32, kernel_size=3, padding=1), nn.ReLU(), nn.MaxPool2d(2))
7batch = torch.zeros(8, 1, 28, 28)
8print(tuple(features(batch).shape))
9print(sum(p.numel() for p in features.parameters()))(8, 32, 7, 7) 4800
Transfer learning
Training a big network from scratch takes huge data and compute. Usually you don’t: you start from a model pretrained on a large dataset (ImageNet for vision, web text for language) and fine-tune it on your task. The early layers’ edge, texture and grammar detectors transfer remarkably well.
1from torchvision import models
2model = models.resnet18(weights="IMAGENET1K_V1") # pretrained
3for parameter in model.parameters():
4 parameter.requires_grad = False # freeze everything...
5model.fc = nn.Linear(model.fc.in_features, 4) # ...then a new head for 4 doodle classes
6optimizer = torch.optim.AdamW(model.fc.parameters(), lr=1e-3)With a little more data, unfreeze the later layers too and fine-tune them with a smaller learning rate. Hugging Face hosts thousands of pretrained models for this.
Key takeaways
PyTorch tensors are numpy arrays with autograd and GPU support;
loss.backward()is your Value.backward().Models are nn.Modules; the loop is zero_grad → forward → loss → backward → step.
Use model.train()/model.eval(), torch.no_grad() for inference, and .to(device) for GPUs.
Fine-tune pretrained models: freeze early layers, replace the head, train with a small learning rate.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Count a model’s parameters
Each input line is a layer in PyTorch-style shorthand: Linear in out, Conv2d in out kernel, LayerNorm features, Embedding vocabulary dim, or a parameter-free layer like ReLU, Dropout or MaxPool2d. Print each layer’s parameter count and then the total, matching what sum(p.numel() for p in model.parameters()) would report:
Linear 784 128: 100480
ReLU: 0
total: 100480(LayerNorm has a learned scale and shift per feature; Embedding has one vector per vocabulary entry and no bias.)
- An MLP
- A tiny transformer piece
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Fine-tuning budget
The first input line lists the layer names to train (the rest are frozen). Each following line is layer_name parameter_count for a pretrained model. Print the trainable and frozen counts and the trainable percentage to 2 decimals: trainable 2052 of 11178564 (0.02%), frozen 11176512.
- New head only
- Head and last block
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…