Um momento
0xC0Lesson 13 of 18

Generative AI: how machines create

From telling things apart to making new ones: autoregressive text, GANs, latent spaces, and diffusion models that paint by removing noise.

24 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Tell discriminative models from generative ones
  • Explain how diffusion models generate images by learning to remove noise
  • Describe what a latent space is and what interpolating in it does

Most of the models in this track so far are discriminative: they look at something and decide - spam or not, cat or dog, which digit. Generative models do the reverse: they produce new examples that look like they came from the training data - a sentence, a face, a song, a molecule.

Underneath, a generative model learns the probability distribution of its data: which pixel patterns, words or notes are likely together. Generating means sampling from that distribution. There are a few big families:

  • Autoregressive models build the output one piece at a time, each piece conditioned on the ones before. Chatbots work this way, one token after another.
  • GANs (generative adversarial networks, 2014) pit two networks against each other: a generator makes fakes and a discriminator tries to spot them. Each makes the other better - like a forger and a detective.
  • VAEs (variational autoencoders) squeeze data into a small latent space of numbers and learn to rebuild it; picking a new point in that space and decoding it creates something new.
  • Diffusion models (Stable Diffusion, DALL·E, Midjourney, many video models) learn to remove noise. That idea is surprisingly simple - let’s take it apart.

Try it

Discriminative or generative?

Sort each AI task: is the model deciding something about an input, or creating something new?

0 of 7 sortedScore 0/0
  • “Flagging an email as spam”

  • “Writing a poem about your cat”

  • “Detecting pedestrians in a camera frame”

  • “Turning “a fox in a spacesuit” into a picture”

  • “Predicting whether a tumour on a scan is benign”

  • “Creating a realistic voice reading your text”

  • “Proposing new protein shapes for a medicine”

Diffusion: painting by removing noise

Take a photo and add a sprinkle of random noise. Then a bit more. Keep going for, say, 1,000 steps, and the photo becomes pure static. This forward process needs no learning at all - it’s a fixed recipe:

xt=αˉt x0+1−αˉt εx_t = \sqrt{\bar\alpha_t}\, x_0 + \sqrt{1 - \bar\alpha_t}\, \varepsilon

where x0x_0 is the original, ε\varepsilon is random noise, and the schedule αˉt\bar\alpha_t slides from 1 (all image) to 0 (all noise).

Now the clever part. Train a neural network on millions of noisy images to answer one question: “what noise was added here?” If it can predict the noise, you can subtract a little of it and get a slightly cleaner image. Repeat from pure static, step by step, and an image emerges that never existed before. Add a text prompt as an extra input, and the network removes noise in the direction of “a fox in a spacesuit”.

forward: add a little noise each step (fixed recipe)x₀ cleanstep 1step 2step 3x_T staticreverse: a neural network predicts and removes noisestart from fresh static (+ a text prompt) → a brand-new image
Diffusion in two directions. Training: add noise step by step (easy, no learning needed). Generating: a neural network removes a little noise at a time, starting from pure static.

Try it

Drown an image in noise - and bring it back

Drag the slider to add noise step by step, and watch the chart: the blue signal fades as the red noise takes over. Around which step can you no longer recognise the picture? Then press Denoise to play the process backwards - the job a diffusion model learns to do without knowing the original.

x₀ - the original
x₀ - after 0 noising steps

Crisp and clear - nothing has been added yet.

10step t →signal √ᾱnoise √(1−ᾱ)

The trick: here “denoise” cheats by replaying noise we already know. A real diffusion model never sees the original - a neural network looks at the noisy image and predicts the noise, one small step at a time. Train it on millions of images and you can start from fresh static and “denoise” your way to a brand-new picture.

noise_schedule.py
1import math
2
3def signal_level(t, T, s=0.008):
4    # The "cosine schedule": how much of the original image survives after t of T steps.
5    f = lambda step: math.cos((step / T + s) / (1 + s) * math.pi / 2) ** 2
6    return f(t) / f(0)
7
8T = 1000
9for t in [0, 100, 250, 500, 750, 900, 1000]:
10    a = max(0.0, min(1.0, signal_level(t, T)))
11    signal, noise = math.sqrt(a), math.sqrt(1 - a)
12    bar = "#" * round(a * 20) + "." * (20 - round(a * 20))
13    print(f"t={t:4}  signal {signal:.2f}  noise {noise:.2f}  {bar}")
Output
t=   0  signal 1.00  noise 0.00  ####################
t= 100  signal 0.99  noise 0.17  ###################.
t= 250  signal 0.92  noise 0.39  #################...
t= 500  signal 0.70  noise 0.71  ##########..........
t= 750  signal 0.38  noise 0.93  ###.................
t= 900  signal 0.16  noise 0.99  ....................
t=1000  signal 0.00  noise 1.00  ....................

Latent space: a map of possibilities

Generative models usually work in a compressed latent space: a handful of numbers that describe the important features of an output. Nearby points give similar outputs, so walking in a straight line from one point to another morphs smoothly between them - the trick behind those videos where one face melts into another.

latent_walk.py
1# Pretend a generator maps 3 latent numbers to a face: (smile, hair length, glasses).
2happy_short = [0.9, 0.1, 0.0]
3grumpy_long = [-0.8, 0.9, 1.0]
4
5def mix(a, b, amount):
6    return [round(x + (y - x) * amount, 2) for x, y in zip(a, b)]
7
8for step in range(5):
9    amount = step / 4
10    print(f"{amount:.2f} -> {mix(happy_short, grumpy_long, amount)}")
Output
0.00 -> [0.9, 0.1, 0.0]
0.25 -> [0.47, 0.3, 0.25]
0.50 -> [0.05, 0.5, 0.5]
0.75 -> [-0.38, 0.7, 0.75]
1.00 -> [-0.8, 0.9, 1.0]

Key takeaways

  • Discriminative models decide about an input; generative models create new samples from a learned distribution.

  • Families: autoregressive (one piece at a time), GANs (generator vs discriminator), VAEs (latent space), diffusion (denoising).

  • Diffusion: noising is a fixed recipe; a network learns to predict the noise, so generation runs the process backwards from static.

  • A latent space is a compact map of outputs; interpolating in it morphs smoothly between them.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Read a noise schedule

+25 XP

Line 1: the total number of steps T. Line 2: some step numbers t.

Use the cosine schedule with s = 0.008: let f(t)=cos⁡2 ⁣(t/T+s1+s⋅π2)f(t) = \cos^2\!\left(\frac{t/T + s}{1 + s} \cdot \frac{\pi}{2}\right) and αˉt=f(t)/f(0)\bar\alpha_t = f(t) / f(0), clamped to the range 0-1. For each t, print t=T_VALUE signal S noise N where S = αˉt\sqrt{\bar\alpha_t} and N = 1−αˉt\sqrt{1 - \bar\alpha_t}, both to 3 decimal places.

  • A thousand steps
  • Ten steps
  • Fifty steps
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Walk through latent space

+25 XP

Line 1 and line 2 are two latent vectors of the same length. Line 3 is a number of steps n. Print the n + 1 evenly spaced points from the first vector to the second (both included), one per line, each value formatted to 2 decimal places and separated by spaces.

  • Four steps
  • Three dimensions
  • Thirds
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: