Noise into Images
Generation has a second great recipe besides next-token prediction. Destroy an image with noise, step by step, until nothing remains. Then train a network to run the film backwards. That is a diffusion model.
Two directions
The forward process is fixed, not learned: mix in a little Gaussian noise at each step until the image is pure static. The reverse process is the model: given a noisy image and the step number, predict the noise that was added. Train it on millions of images. Then, at generation time, start from random noise and iteratively denoise — and an image appears.
Why it works so well
Denoising is an easier learning problem than generating from scratch. Each step is a small, local correction, and the model only needs to be good at noticing and removing noise. The result is an extraordinary range of outputs and precise control through conditioning.
Two refinements matter in practice. Latent diffusion runs the whole process in a compressed latent space rather than on pixels, which makes it fast enough to be practical. Classifier-free guidance trades diversity for prompt fidelity: it overshoots toward the conditioning, so a prompt is followed more literally, at the cost of variety.
Two roads to generation
Language models write left to right, one token at a time. Diffusion models refine the whole output at once, over many passes. Both learn a distribution and both sample from it; they just take different paths.