Um momento
0xF0Lesson 16 of 17

Deep learning with PyTorch

Map everything you built by hand onto PyTorch: tensors, autograd, modules, optimizers, data loaders and fine-tuning.

30 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Translate numpy tensors, your autograd engine and your training loop into PyTorch
  • Define models with nn.Module, train them with an optimizer, and switch between train and eval modes
  • Use GPUs, save and load weights, and fine-tune a pretrained model

You’ve built deep learning from scratch. In real projects you use a framework - most often PyTorch - that provides the same pieces, fast and on GPUs. The good news: you already understand every one of them.

You builtPyTorch
numpy arraystorch.tensor (same broadcasting, @, .shape)
the Value autograd enginerequires_grad=True and loss.backward()
x @ W + b layersnn.Linear, nn.Conv2d, nn.LSTM, nn.MultiheadAttention
ReLU, softmax + cross-entropynn.ReLU, nn.CrossEntropyLoss (takes logits!)
SGD, momentum, Adam updatestorch.optim.SGD, torch.optim.AdamW
dropout, layer normnn.Dropout, nn.LayerNorm
shuffled mini-batchesDataLoader(dataset, batch_size=32, shuffle=True)

PyTorch can’t run in this browser, so this lesson’s examples aren’t runnable here - install it with pip install torch and try them. Every output shown was checked against PyTorch 2.9.

autograd_in_pytorch.py
1import torch
2
3w = torch.tensor(0.5, requires_grad=True)
4x, b, y = torch.tensor(2.0), torch.tensor(-1.0, requires_grad=True), torch.tensor(1.0)
5p = torch.sigmoid(w * x + b)
6loss = (p - y) ** 2
7loss.backward()                      # your Value.backward(), on tensors
8print(f"loss={loss.item():.4f} dw={w.grad.item():.4f} db={b.grad.item():.4f}")
Output
loss=0.2500 dw=-0.5000 db=-0.2500

Those are exactly the numbers from the backprop stepper’s sigmoid neuron. Now the training loop - the same four steps as yours, with zero_grad() doing the gradient reset you had to remember yourself:

train_xor_pytorch.py
1import torch
2from torch import nn
3
4torch.manual_seed(0)
5X = torch.tensor([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
6y = torch.tensor([0, 1, 1, 0])
7
8model = nn.Sequential(nn.Linear(2, 8), nn.Tanh(), nn.Linear(8, 2))
9loss_fn = nn.CrossEntropyLoss()                      # takes logits
10optimizer = torch.optim.Adam(model.parameters(), lr=0.05)
11
12model.train()
13for epoch in range(300):
14    optimizer.zero_grad()            # reset gradients
15    loss = loss_fn(model(X), y)      # forward + loss
16    loss.backward()                  # backward
17    optimizer.step()                 # update
18
19model.eval()
20with torch.no_grad():                # no graph needed for predictions
21    predictions = model(X).argmax(dim=1)
22print("predictions:", predictions.tolist(), f"loss: {loss.item():.3f}")
23print("parameters:", sum(p.numel() for p in model.parameters()))
Output
predictions: [0, 1, 1, 0] loss: 0.000
parameters: 42

For anything bigger than nn.Sequential can express, write a class:

1class DoodleNet(nn.Module):
2    def __init__(self):
3        super().__init__()
4        self.features = nn.Sequential(
5            nn.Conv2d(1, 16, kernel_size=3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
6            nn.Conv2d(16, 32, kernel_size=3, padding=1), nn.ReLU(), nn.MaxPool2d(2))
7        self.classifier = nn.Sequential(nn.Flatten(), nn.Dropout(0.3), nn.Linear(32 * 7 * 7, 10))
8
9    def forward(self, x):            # x: (batch, 1, 28, 28)
10        return self.classifier(self.features(x))

And the practical essentials:

  • GPU: device = "cuda" if torch.cuda.is_available() else "cpu", then model.to(device) and batch.to(device).
  • Modes: model.train() turns dropout and batch-norm training behavior on; model.eval() turns them off. Wrap evaluation in torch.no_grad() to save memory.
  • Saving: torch.save(model.state_dict(), "doodle.pt"), and later model.load_state_dict(torch.load("doodle.pt")).
shapes_in_pytorch.py
1import torch
2from torch import nn
3
4features = nn.Sequential(
5    nn.Conv2d(1, 16, kernel_size=3, padding=1), nn.ReLU(), nn.MaxPool2d(2),
6    nn.Conv2d(16, 32, kernel_size=3, padding=1), nn.ReLU(), nn.MaxPool2d(2))
7batch = torch.zeros(8, 1, 28, 28)
8print(tuple(features(batch).shape))
9print(sum(p.numel() for p in features.parameters()))
Output
(8, 32, 7, 7)
4800

Transfer learning

Training a big network from scratch takes huge data and compute. Usually you don’t: you start from a model pretrained on a large dataset (ImageNet for vision, web text for language) and fine-tune it on your task. The early layers’ edge, texture and grammar detectors transfer remarkably well.

1from torchvision import models
2model = models.resnet18(weights="IMAGENET1K_V1")      # pretrained
3for parameter in model.parameters():
4    parameter.requires_grad = False                    # freeze everything...
5model.fc = nn.Linear(model.fc.in_features, 4)         # ...then a new head for 4 doodle classes
6optimizer = torch.optim.AdamW(model.fc.parameters(), lr=1e-3)

With a little more data, unfreeze the later layers too and fine-tune them with a smaller learning rate. Hugging Face hosts thousands of pretrained models for this.

Key takeaways

  • PyTorch tensors are numpy arrays with autograd and GPU support; loss.backward() is your Value.backward().

  • Models are nn.Modules; the loop is zero_grad → forward → loss → backward → step.

  • Use model.train()/model.eval(), torch.no_grad() for inference, and .to(device) for GPUs.

  • Fine-tune pretrained models: freeze early layers, replace the head, train with a small learning rate.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Count a model’s parameters

+25 XP

Each input line is a layer in PyTorch-style shorthand: Linear in out, Conv2d in out kernel, LayerNorm features, Embedding vocabulary dim, or a parameter-free layer like ReLU, Dropout or MaxPool2d. Print each layer’s parameter count and then the total, matching what sum(p.numel() for p in model.parameters()) would report:

Linear 784 128: 100480
ReLU: 0
total: 100480

(LayerNorm has a learned scale and shift per feature; Embedding has one vector per vocabulary entry and no bias.)

  • An MLP
  • A tiny transformer piece
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Fine-tuning budget

+25 XP

The first input line lists the layer names to train (the rest are frozen). Each following line is layer_name parameter_count for a pretrained model. Print the trainable and frozen counts and the trainable percentage to 2 decimals: trainable 2052 of 11178564 (0.02%), frozen 11176512.

  • New head only
  • Head and last block
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: