AI Foundationspredict · compress · act
act XII

Beyond Text

Diffusion explained how an image is made. This is what changed since: a new architecture, a control layer, and an ecosystem that went open almost overnight.

53

Image Models

frontieras of 2026-09
before this →Noise into Images

The diffusion chapter gave the recipe: destroy an image with noise, then train a network to undo it. That idea has not changed. What changed is where the network works, what shape it has, and who is willing to release one.

CONDITIONING — what you tell ittext promptstructure (ControlNet)style (LoRA)DENOISERin latent spacelatent →decoder → imageTHE SHAPE OF THE DENOISERU-Net — convolutions, 2022–2023Stable Diffusion 1.5, SDXLTransformer (DiT) — 2024 onwardSD3, FLUX, Qwen-Image · often with flow matchingclassifier-free guidance (from the diffusion chapter) still steers how hard the prompt is followed.
two shifts in image models: work in a compressed latent instead of pixels, and swap the U-Net for a transformer. Text, structure and style enter as separate controls.

Two shifts

Latent diffusion. Instead of denoising full-resolution pixels, the model works in a small compressed latent space and a separate decoder turns the result back into an image. This is what made image generation run on ordinary hardware — Stable Diffusion’s whole reputation rested on it.

The transformer takes over. The original denoiser was a U-Net of convolutions. Newer models replace it with a diffusion transformer (DiT) — the same attention machinery as a language model, applied to patches. Many now train with flow matching, a cleaner objective that produces a straighter path from noise to image. The architecture from the attention chapter turned out to be general.

Making it obey

A prompt alone gives loose control. Three additions fixed that, and each maps to a different question:

The ecosystem, and why it opened

Text models mostly stayed behind APIs; image models went open fast. The reason is structural: image models are smaller, the community had already built a culture of checkpoints and fine-tunes, and a LoRA is easy to share. The result is a dense field.

Open weights Closed
FLUX (Black Forest Labs) — Apache 2.0 for Schnell, custom for Dev Midjourney — product, no weights
Stable Diffusion / SDXL / SD 3.5 (Stability AI) Imagen (Google) — via API
Qwen-Image, Hunyuan-Image, HiDream, Z-Image DALL·E — via API

What people actually do with them

generation is the easy part

Editing now matters more than generating from nothing: remove an object, change the light, extend the frame, or rewrite the image from an instruction. Identity consistency — the same character across many images — is the problem the whole field is chasing, because it is what making a film or a comic actually requires. And the workflow is rarely one model: a base model, a set of LoRAs, a ControlNet, and an upscaler, chained together.

the honest noteModel names above change monthly, and this table will be stale within a quarter. Treat it as a map of the kind of thing that exists, not a current ranking. The architecture — latent diffusion, transformer denoiser, conditioning — is the part that lasts.
GENERATEtext or structure → an image
EDITchange part of it, keep the rest
CONSISTkeep a character identical across many

Images are only the first non-text modality. The next chapter is about models that drop the boundaries between modalities altogether — reading and speaking, seeing and drawing, in a single system.

introduces →latent diffusiondiffusion transformerflow matchingControlNetopen image model
← previousThe Harnessnext →Omni-Models