Image Models
The diffusion chapter gave the recipe: destroy an image with noise, then train a network to undo it. That idea has not changed. What changed is where the network works, what shape it has, and who is willing to release one.
Two shifts
Latent diffusion. Instead of denoising full-resolution pixels, the model works in a small compressed latent space and a separate decoder turns the result back into an image. This is what made image generation run on ordinary hardware — Stable Diffusion’s whole reputation rested on it.
The transformer takes over. The original denoiser was a U-Net of convolutions. Newer models replace it with a diffusion transformer (DiT) — the same attention machinery as a language model, applied to patches. Many now train with flow matching, a cleaner objective that produces a straighter path from noise to image. The architecture from the attention chapter turned out to be general.
Making it obey
A prompt alone gives loose control. Three additions fixed that, and each maps to a different question:
- Text encoders (CLIP, T5) turn your prompt into the conditioning signal — the reason wording matters so much.
- ControlNet takes a structural input — a pose, an edge map, a depth image — and forces the output to follow it.
- LoRA adapts a frozen base model to a style, a face, or a character with a small add-on file. The same technique from the open-weights act, applied to images.
The ecosystem, and why it opened
Text models mostly stayed behind APIs; image models went open fast. The reason is structural: image models are smaller, the community had already built a culture of checkpoints and fine-tunes, and a LoRA is easy to share. The result is a dense field.
| Open weights | Closed |
|---|---|
| FLUX (Black Forest Labs) — Apache 2.0 for Schnell, custom for Dev | Midjourney — product, no weights |
| Stable Diffusion / SDXL / SD 3.5 (Stability AI) | Imagen (Google) — via API |
| Qwen-Image, Hunyuan-Image, HiDream, Z-Image | DALL·E — via API |
What people actually do with them
Editing now matters more than generating from nothing: remove an object, change the light, extend the frame, or rewrite the image from an instruction. Identity consistency — the same character across many images — is the problem the whole field is chasing, because it is what making a film or a comic actually requires. And the workflow is rarely one model: a base model, a set of LoRAs, a ControlNet, and an upscaler, chained together.
Images are only the first non-text modality. The next chapter is about models that drop the boundaries between modalities altogether — reading and speaking, seeing and drawing, in a single system.