AI Foundationspredict · compress · act
act VIII

The Attention Revolution

Diffusion models learn to subtract noise, one step at a time — and in doing so learn to create.

33

Noise into Images

before this →Teaching Taste
2021Diffusion models — Learn to undo noise, and you can make an image from scratch.
2022Stable Diffusion — Image generation runs on a home GPU and spreads everywhere.

Generation has a second great recipe besides next-token prediction. Destroy an image with noise, step by step, until nothing remains. Then train a network to run the film backwards. That is a diffusion model.

The recipe before this one

Diffusion gets called the way to generate images, but it is the third recipe, and the first one deserves its paragraph. A generative adversarial network (GAN, 2014) trains two networks against each other: a generator that makes fake images, and a discriminator that tries to tell fake from real. Nobody ever tells the generator what a cat looks like — it only ever hears “not convincing enough”, and it learns to fool the critic. For most of a decade this made the sharpest images anyone had produced.

It had a characteristic failure. A mode collapse happens when the generator finds a single output the discriminator cannot reject — one convincing cat — and then produces that cat forever. Training a two-player game is also unstable: the two networks have to improve in step, and if one gets too far ahead the signal for the other disappears. Diffusion replaced GANs not because it made better pictures at first, but because its training is a regression — predict the noise that was added — and a regression has a stable answer to converge on. Stability beat sharpness, and then eventually it beat it on sharpness too. The trade-off is real and still shows: a GAN produces an image in one pass, where diffusion needs many.

forward: add noisereverse: the model predicts and removes the noise — learned once, run at sampling time
forward: add noise. reverse: learn to remove it. generation: start from pure noise.

Two directions

The forward process is fixed, not learned: mix in a little Gaussian noise at each step until the image is pure static. The reverse process is the model: given a noisy image and the step number, predict the noise that was added. Train it on millions of images. Then, at generation time, start from random noise and iteratively denoise. An image appears.

Why it works so well

Denoising is an easier learning problem than generating from scratch. Each step is a small, local correction. The model only needs to be good at noticing and removing noise. The result is an extraordinary range of outputs and precise control through conditioning.

Two refinements matter in practice. Latent diffusion runs the whole process in a compressed latent space rather than on pixels, which makes it fast enough to be practical. Classifier-free guidance trades diversity for prompt fidelity: it overshoots toward the conditioning, so a prompt is followed more literally, at the cost of variety.

Two roads to generation

autoregressive vs. diffusion

Language models write left to right, one token at a time. Diffusion models refine the whole output at once, over many passes. Both learn a distribution and both sample from it; they take different paths.

the unificationRecent systems combine them: diffusion for images and video, autoregressive transformers for text, and increasingly one architecture learning both. The boundary is dissolving.
NOISEstart from pure static
DENOISEpredict and subtract, step by step
OUTPUTan image, guided by your prompt
introduces →generative adversarial networkmode collapsediffusion modeldenoisinglatent spaceclassifier-free guidanceIan Goodfellow
← previousTeaching Tastenext →Opening the Black Box