Video Models
The image chapter ended with a recipe: destroy a picture with noise, learn to undo it, work in a compressed latent. Point that recipe at video and you might expect it to just work. It does not — because a video is not a pile of pictures. Every frame has to agree with the frame before it, and all of them together have to agree with the way the world moves.
Time is a new axis, not a new model
The move that made video work was almost boring: treat time as a third dimension and patch over it. A block of pixels measured across space and a handful of frames becomes a single token, and the same attention machinery from chapter 28 runs over those tokens. Nothing about the transformer had to change. A video model is a diffusion model that denoises a spacetime volume instead of a picture — which is why image models and video models appeared in the same two years, from the same architecture.
The hard part is not the architecture. It is temporal consistency: frame 400 has to agree with frame 399 and with frame 3. Two things break this. Error accumulates, so a small mistake at frame 50 is a visible one at frame 300. And memory is finite — a model generating a long clip cannot keep every earlier frame in view, so once a frame leaves the window it is forgotten. A face changes shape; a shirt changes colour; a room quietly rearranges itself. That is the failure mode the whole field is measured on.
Sound stopped being a separate step
Early pipelines added audio afterwards, as a second model. Newer ones generate it natively — the same model produces picture and sound together, so they can agree. When OpenAI shipped Sora 2 in September 2025 it listed synchronized dialogue and sound effects alongside the visuals, and DeepMind describes Veo 3 as adding “sound effects, ambient noise, and even dialogue” natively. That is a real change: lip movement and speech are one problem, and a model that does not know that will always look dubbed.
The frontier, late 2026
Closed. Sora 2 (OpenAI, 30 September 2025) — realistic physics, synchronised audio, control across multiple shots. Veo 3.1 (Google DeepMind) — eight-second clips at up to 4K with natively generated audio. Kling 2.5 Turbo (Kuaishou, September 2025), the strongest of the Chinese services.
Open. Alibaba’s Wan 2.2, Tencent’s HunyuanVideo, plus LTX-Video, CogVideoX, Open-Sora and Mochi — a video model you can download and fine-tune is ordinary now, exactly as it became ordinary for images.
What still fails
Be careful with any demo you see. Everything above produces clips measured in seconds. Physics still breaks — objects merge, hands multiply, cups pass through tables — and the longer the clip, the more it drifts. And the table above is the least durable part of this chapter: model names in video turn over in months, while the three ideas — time as an axis, native audio, bounded memory — are the parts that will still be true.
Which raises a smaller question first. These models already generate sound natively — but that sound is attached, not understood. The next chapter is about the modality AI handled first and modelled last: the one that talks back.