AI Foundationspredict · compress · act
act VII

The Attention Revolution

Diffusion models learn to subtract noise, one step at a time — and in doing so learn to create.

29

Noise into Images

before this →Teaching Taste
2021Diffusion models — Learn to undo noise, and you can make an image from scratch.
2022Stable Diffusion — Image generation runs on a home GPU and spreads everywhere.

Generation has a second great recipe besides next-token prediction. Destroy an image with noise, step by step, until nothing remains. Then train a network to run the film backwards. That is a diffusion model.

forward: add noisereverse: the model predicts and removes the noise — learned once, run at sampling time
forward: add noise. reverse: learn to remove it. generation: start from pure noise.

Two directions

The forward process is fixed, not learned: mix in a little Gaussian noise at each step until the image is pure static. The reverse process is the model: given a noisy image and the step number, predict the noise that was added. Train it on millions of images. Then, at generation time, start from random noise and iteratively denoise — and an image appears.

Why it works so well

Denoising is an easier learning problem than generating from scratch. Each step is a small, local correction, and the model only needs to be good at noticing and removing noise. The result is an extraordinary range of outputs and precise control through conditioning.

Two refinements matter in practice. Latent diffusion runs the whole process in a compressed latent space rather than on pixels, which makes it fast enough to be practical. Classifier-free guidance trades diversity for prompt fidelity: it overshoots toward the conditioning, so a prompt is followed more literally, at the cost of variety.

Two roads to generation

autoregressive vs. diffusion

Language models write left to right, one token at a time. Diffusion models refine the whole output at once, over many passes. Both learn a distribution and both sample from it; they just take different paths.

the unificationRecent systems combine them: diffusion for images and video, autoregressive transformers for text, and increasingly one architecture learning both. The boundary is dissolving.
NOISEstart from pure static
DENOISEpredict and subtract, step by step
OUTPUTan image, guided by your prompt
introduces →diffusion modeldenoisinglatent spaceclassifier-free guidance
← previousTeaching Tastenext →Opening the Black Box