Generative AI: how machines create
From telling things apart to making new ones: autoregressive text, GANs, latent spaces, and diffusion models that paint by removing noise.
- Tell discriminative models from generative ones
- Explain how diffusion models generate images by learning to remove noise
- Describe what a latent space is and what interpolating in it does
Most of the models in this track so far are discriminative: they look at something and decide - spam or not, cat or dog, which digit. Generative models do the reverse: they produce new examples that look like they came from the training data - a sentence, a face, a song, a molecule.
Underneath, a generative model learns the probability distribution of its data: which pixel patterns, words or notes are likely together. Generating means sampling from that distribution. There are a few big families:
- Autoregressive models build the output one piece at a time, each piece conditioned on the ones before. Chatbots work this way, one token after another.
- GANs (generative adversarial networks, 2014) pit two networks against each other: a generator makes fakes and a discriminator tries to spot them. Each makes the other better - like a forger and a detective.
- VAEs (variational autoencoders) squeeze data into a small latent space of numbers and learn to rebuild it; picking a new point in that space and decoding it creates something new.
- Diffusion models (Stable Diffusion, DALL·E, Midjourney, many video models) learn to remove noise. That idea is surprisingly simple - let’s take it apart.
Try it
Discriminative or generative?
Sort each AI task: is the model deciding something about an input, or creating something new?
“Flagging an email as spam”
“Writing a poem about your cat”
“Detecting pedestrians in a camera frame”
“Turning “a fox in a spacesuit” into a picture”
“Predicting whether a tumour on a scan is benign”
“Creating a realistic voice reading your text”
“Proposing new protein shapes for a medicine”
Diffusion: painting by removing noise
Take a photo and add a sprinkle of random noise. Then a bit more. Keep going for, say, 1,000 steps, and the photo becomes pure static. This forward process needs no learning at all - it’s a fixed recipe:
where is the original, is random noise, and the schedule slides from 1 (all image) to 0 (all noise).
Now the clever part. Train a neural network on millions of noisy images to answer one question: “what noise was added here?” If it can predict the noise, you can subtract a little of it and get a slightly cleaner image. Repeat from pure static, step by step, and an image emerges that never existed before. Add a text prompt as an extra input, and the network removes noise in the direction of “a fox in a spacesuit”.
Try it
Drown an image in noise - and bring it back
Drag the slider to add noise step by step, and watch the chart: the blue signal fades as the red noise takes over. Around which step can you no longer recognise the picture? Then press Denoise to play the process backwards - the job a diffusion model learns to do without knowing the original.
Crisp and clear - nothing has been added yet.
The trick: here “denoise” cheats by replaying noise we already know. A real diffusion model never sees the original - a neural network looks at the noisy image and predicts the noise, one small step at a time. Train it on millions of images and you can start from fresh static and “denoise” your way to a brand-new picture.
1import math
2
3def signal_level(t, T, s=0.008):
4 # The "cosine schedule": how much of the original image survives after t of T steps.
5 f = lambda step: math.cos((step / T + s) / (1 + s) * math.pi / 2) ** 2
6 return f(t) / f(0)
7
8T = 1000
9for t in [0, 100, 250, 500, 750, 900, 1000]:
10 a = max(0.0, min(1.0, signal_level(t, T)))
11 signal, noise = math.sqrt(a), math.sqrt(1 - a)
12 bar = "#" * round(a * 20) + "." * (20 - round(a * 20))
13 print(f"t={t:4} signal {signal:.2f} noise {noise:.2f} {bar}")t= 0 signal 1.00 noise 0.00 #################### t= 100 signal 0.99 noise 0.17 ###################. t= 250 signal 0.92 noise 0.39 #################... t= 500 signal 0.70 noise 0.71 ##########.......... t= 750 signal 0.38 noise 0.93 ###................. t= 900 signal 0.16 noise 0.99 .................... t=1000 signal 0.00 noise 1.00 ....................
Latent space: a map of possibilities
Generative models usually work in a compressed latent space: a handful of numbers that describe the important features of an output. Nearby points give similar outputs, so walking in a straight line from one point to another morphs smoothly between them - the trick behind those videos where one face melts into another.
1# Pretend a generator maps 3 latent numbers to a face: (smile, hair length, glasses).
2happy_short = [0.9, 0.1, 0.0]
3grumpy_long = [-0.8, 0.9, 1.0]
4
5def mix(a, b, amount):
6 return [round(x + (y - x) * amount, 2) for x, y in zip(a, b)]
7
8for step in range(5):
9 amount = step / 4
10 print(f"{amount:.2f} -> {mix(happy_short, grumpy_long, amount)}")0.00 -> [0.9, 0.1, 0.0] 0.25 -> [0.47, 0.3, 0.25] 0.50 -> [0.05, 0.5, 0.5] 0.75 -> [-0.38, 0.7, 0.75] 1.00 -> [-0.8, 0.9, 1.0]
Key takeaways
Discriminative models decide about an input; generative models create new samples from a learned distribution.
Families: autoregressive (one piece at a time), GANs (generator vs discriminator), VAEs (latent space), diffusion (denoising).
Diffusion: noising is a fixed recipe; a network learns to predict the noise, so generation runs the process backwards from static.
A latent space is a compact map of outputs; interpolating in it morphs smoothly between them.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Read a noise schedule
Line 1: the total number of steps T. Line 2: some step numbers t.
Use the cosine schedule with s = 0.008: let and , clamped to the range 0-1. For each t, print t=T_VALUE signal S noise N where S = and N = , both to 3 decimal places.
- A thousand steps
- Ten steps
- Fifty steps
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Walk through latent space
Line 1 and line 2 are two latent vectors of the same length. Line 3 is a number of steps n. Print the n + 1 evenly spaced points from the first vector to the second (both included), one per line, each value formatted to 2 decimal places and separated by spaces.
- Four steps
- Three dimensions
- Thirds
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…