Generative Models
GANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
Every model so far in this course understands or predicts: a CNNCNN (Convolutional Neural Network)A CNN is a neural network built around the convolution operation, which encodes locality and translation invariance for processing images efficiently. classifies an image, a transformer predicts the next token. This lesson is about models built to generate new content — new images, audio, video — that didn't exist in the training data, which turns out to need a genuinely different training setup than the supervised loss functionsLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. covered so far.
The core problem: there's no single "correct" output
Supervised learning (see ML Fundamentals) works by comparing a prediction to one correct answer. Generation breaks that assumption: there is no single "correct" image of a cat — millions of valid ones exist, and a good generative model should be capable of producing any of them, with the right relative likelihoods. This is why generative models need training objectives fundamentally different from cross-entropy against a known label — each approach below solves this "no single right answer" problem differently.
GANs: two networks competing
A GANGAN (Generative Adversarial Network)A GAN trains a generator and a discriminator against each other, the generator learning to fool the discriminator into mistaking its output for real data. (Generative Adversarial Network) trains two networks against each other:
- The generator takes random noise as input and tries to produce realistic-looking outputs (e.g. images).
- The discriminator is a classifier trained to distinguish the generator's fake outputs from real training examples.
The generator is trained to fool the discriminator; the discriminator is trained to catch the generator. As training progresses, both improve together — the discriminator gets better at spotting fakes, forcing the generator to produce increasingly realistic ones. This is a genuinely different training loop from anything else in this course: two networks with opposing objectives, updated in alternation, rather than one network minimizing a single fixed loss.
Why are GANs so hard to train?
GANs are notoriously hard to train
The adversarial setup that makes GANs powerful also makes them unstable: if the discriminator gets too good too fast, the generator stops receiving useful gradient signal and training stalls (a failure mode called mode collapse, where the generator produces only a narrow range of outputs). Diffusion modelsDiffusion ModelA diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step., below, have largely displaced GANs for state-of-the-art image generation partly because their training is far more stable.
A GAN has no single loss to minimise. The generator and discriminator play against each other over one shared objective:
wants to push toward 1 on real data and toward 0 on fakes; wants the opposite. The target is not a minimum but a Nash equilibrium — a point where neither network can improve unilaterally.
That difference is the source of the instability. Ordinary training descends a fixed landscape; here the landscape moves every time the opponent updates, so the pair can oscillate or drift without converging. It also explains mode collapse directly: the objective only asks to produce samples cannot distinguish from real. Nothing in it rewardsRewardA reward is the scalar feedback signal a reinforcement-learning agent receives after taking an action — the only learning signal it gets, and often a delayed one. covering the data distribution, so emitting one very convincing cat forever is a perfectly good solution to the equation — just not to the problem.
VAEs: learning a compressed, generative latent space
A VAEVAE (Variational Autoencoder)A VAE encodes inputs into a distribution over a compact latent space rather than a fixed point, so sampling from that space and decoding produces plausible new outputs. (Variational Autoencoder) takes a different approach: an encoder compresses an input into a compact latent representation (a vector of numbers, similar in spirit to the embeddingsEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space. from LLMs), and a decoder reconstructs the original input from that latent vector. The "variational" part is what makes it generative rather than just compression: instead of encoding an input to a single fixed point, the encoder produces a distribution (see Probability & Statistics Foundations) over the latent space, and the model is trained so that sampling randomly from that latent space and decoding produces plausible new outputs — not just reconstructions of inputs it has already seen.
VAEs train more stably than GANs (a single well-defined loss function, no adversarial dynamics) but historically produced blurrier, less sharp outputs — one of the reasons diffusion models became the dominant approach for high-fidelity image generation.
Diffusion models: destroy, then learn to reverse it
Diffusion models — the approach behind most current state-of-the-art image generators — take a strikingly different strategy: train a model to reverse a process of gradually adding random noise to an image.
The forward direction is fixed arithmetic that needs no learning at all. Only the reverse direction is a network, and it is trained to answer one narrow question at every noise level: what noise is in this image?
Forward process (fixed, not learned): take a real training image and progressively add small amounts of random noise over many steps, until it becomes indistinguishable from pure noise.
Train a model to reverse one step: at each noise level, train a network to predict the noise that was added, so it can be subtracted back out — a supervised prediction target, despite the end goal being generation.
Generate by reversing from pure noise: start from random noise and repeatedly apply the trained denoising step, gradually turning noise into a coherent image over many iterations.
The forward process adds Gaussian noise on a fixed schedule , each step slightly shrinking the image and mixing in noise:
Taken literally, producing would mean running sequential steps — unworkable when is in the hundreds or thousands and you need a fresh sample for every training example. But because the sum of independent Gaussians is Gaussian, the whole chain collapses. Writing and :
One draw of and one interpolation gives you the image at any noise level directly from the clean image. This is what makes training cheap and fully parallel: sample a random , jump straight there, and train on that one noise level. The schedule is chosen so , which is exactly the statement that is indistinguishable from pure noise — the distribution you can sample from for free at generation time.
The closed form above hands you a labelled example for free. You know because you built it, and you know because you drew it — so the noise itself is the target, and the network is trained with mean squared error:
That is it — the same MSE from ML Fundamentals, with sampled uniformly so every noise level gets trained. No adversary, no likelihood bound, no sampling during training.
The subtlety is that a noise prediction is equivalent to an image prediction: rearranging the forward equation recovers . Predicting is the better-conditioned choice because the target has unit variance at every , so no noise level dominates the loss.
Drag the slider below to run that exact closed-form equation yourself, at any noise level:
Forward is the real closed-form equation from the lesson — x_t = √ᾱ_t · x0 + √(1-ᾱ_t) · ε — computed fresh at every position of the slider. Reverse is simulated: there's no trained network here to actually predict and remove noise, so it just replays the same known frames backward. A real diffusion model never gets to see the clean image — it has to guess the noise at each step from x_t alone.
This reframes "generate a realistic image" as a long sequence of much simpler "remove a little noise" prediction problems — each individually easy to train with an ordinary supervised loss, with no adversarial instability. The tradeoff is inference cost: generating one image requires running the denoising network many times in sequence (often dozens to hundreds of steps), making diffusion models considerably slower to sample from than a GAN's single forward pass — a direct analog of the autoregressive generation cost covered in Inference & Serving, just for images instead of tokens.
Modern diffusion models (e.g. Stable Diffusion, Midjourney's underlying approach) commonly run this denoising process in a compressed VAE latent space rather than on raw pixels directly — combining both techniques covered in this lesson — and increasingly use transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.-based denoising networks instead of pure CNNs, another instance of attentionAttention (Self-Attention)Attention is a mechanism letting each position in a sequence weigh every other position via learned Query/Key/Value vectors, forming the core of the transformer. proving to be a general-purpose mechanism rather than a language-specific one.
A 512×512 RGB image is 512 × 512 × 3 ≈ 786,000 values. Stable Diffusion's VAE downsamples by 8 per side into a 64 × 64 × 4 latent: 64 × 64 × 4 ≈ 16,400 values — roughly 48× smaller.
The spatial grid shrinks harder still: 512² = 262,144 positions become 64² = 4,096, a 64× reduction. Since attention cost grows with the square of the number of positions, an attention layer over the latent grid is ~64² ≈ 4,000× cheaper than the same layer over pixels.
And that saving is paid on every denoising step. At 50 steps, the difference between running a heavy network 50 times on 786,000 values versus 16,400 is the difference between a research cluster and a consumer GPUGPU (Graphics Processing Unit)GPUs, originally built for rendering graphics, turned out to be extremely well-suited to the parallel matrix multiplications deep learning requires..
The four families side by side
| Approach | Sample quality | Diversity | Training stability | Sampling speed |
|---|---|---|---|---|
| GAN | Sharp | Prone to mode collapse (narrow output variety) | Fragile, adversarial | One pass — fastest |
| VAE | Historically blurry | Good coverage | Stable, single loss | One pass — fast |
| Diffusion | State of the art | Good coverage | Stable, plain regression | Many passes — slow |
| Autoregressive | Strong | Good coverage | Stable, next-token loss | One pass per element |
The pattern worth noticing: the two stable, well-covering approaches pay for it in sampling cost, and the fastest approach pays for it in training fragility. Latent diffusion and reduced-step samplers are both attempts to keep diffusion's training properties while clawing back its sampling bill.
Recap and what's next
Generative models solve a fundamentally different problem than the classifiers and predictors covered earlier: producing new, plausible content rather than a single correct label. GANs pit two networks against each other; VAEs learn a compressed, sampleable latent space; diffusion models reframe generation as iteratively reversing noise, and currently dominate state-of-the-art image generation. The next lesson returns to text, applying the transformer architecture from the previous lesson at scale: large language modelsLLM (Large Language Model)An LLM is a large transformer trained to predict the next token on massive text corpora, then fine-tuned to follow instructions — the architecture behind GPT, Claude, Gemini, and Llama..