AI Foundationspredict · compress · act
act XIII

Beyond Text

An image is one frame. Video is thousands of frames that must agree with each other, and with physics — which turns out to be a different problem.

59

Video Models

frontieras of 2026-09
before this →Image ModelsOmni-Models
2025Sora 2 — Video generation gets synchronised dialogue and sound effects, made by the same model as the picture.
2026Sora is shut down — OpenAI closes the Sora app and API and points the research team at world simulation and robotics.

The image chapter ended with a recipe: destroy a picture with noise, learn to undo it, work in a compressed latent. Point that recipe at video and you might expect it to just work. It does not — because a video is not a pile of pictures. Every frame has to agree with the frame before it, and all of them together have to agree with the way the world moves.

THE FRAMES — each one is an imagett+1t+2t+3time →frame after frame, this isan image model run repeatedly.ONE TOKEN SPANS SPACE AND TIMEone patch =one tokenAttention runs over patches in spaceand time at once — the transformer doesnot know or care which axis is which.BUT MEMORY IS BOUNDEDthe window it can still seeeverything older is gone — so colour, facesand layout slowly drift. This is the failure.temporal consistency: the frames must agree with each other, not just look good one at a time.
video adds one axis to the problem. The model denoises a block of space and time together, and can only look back at a limited window of what it already made — which is why a scene drifts the longer it runs.

Time is a new axis, not a new model

The move that made video work was almost boring: treat time as a third dimension and patch over it. A block of pixels measured across space and a handful of frames becomes a single token, and the same attention machinery from chapter 28 runs over those tokens. Nothing about the transformer had to change. A video model is a diffusion model that denoises a spacetime volume instead of a picture — which is why image models and video models appeared in the same two years, from the same architecture.

The hard part is not the architecture. It is temporal consistency: frame 400 has to agree with frame 399 and with frame 3. Two things break this. Error accumulates, so a small mistake at frame 50 is a visible one at frame 300. And memory is finite — a model generating a long clip cannot keep every earlier frame in view, so once a frame leaves the window it is forgotten. A face changes shape; a shirt changes colour; a room quietly rearranges itself. That is the failure mode the whole field is measured on.

Sound stopped being a separate step

Early pipelines added audio afterwards, as a second model. Newer ones generate it natively — the same model produces picture and sound together, so they can agree. When OpenAI shipped Sora 2 in September 2025 it listed synchronized dialogue and sound effects alongside the visuals, and DeepMind describes Veo 3 as adding “sound effects, ambient noise, and even dialogue” natively. That is a real change: lip movement and speech are one problem, and a model that does not know that will always look dubbed.

The frontier, late 2026

seconds, not minutes — and names that rot

Closed. Sora 2 (OpenAI, 30 September 2025) — realistic physics, synchronised audio, control across multiple shots. Veo 3.1 (Google DeepMind) — eight-second clips at up to 4K with natively generated audio. Kling 2.5 Turbo (Kuaishou, September 2025), the strongest of the Chinese services.

Open. Alibaba’s Wan 2.2, Tencent’s HunyuanVideo, plus LTX-Video, CogVideoX, Open-Sora and Mochi — a video model you can download and fine-tune is ordinary now, exactly as it became ordinary for images.

what happened to SoraIn March 2026 OpenAI discontinued Sora — the app and the API — and said the research team would move to world simulation and robotics. The most famous video product in the world was shut down to go and build the subject of the chapter after next. That is the clearest signal available about where the frontier thinks the value is.
DAYSone frame at a time
SECONDSspacetime patches, native audio
MINUTES?bounded memory is the wall

What still fails

Be careful with any demo you see. Everything above produces clips measured in seconds. Physics still breaks — objects merge, hands multiply, cups pass through tables — and the longer the clip, the more it drifts. And the table above is the least durable part of this chapter: model names in video turn over in months, while the three ideas — time as an axis, native audio, bounded memory — are the parts that will still be true.

Which raises a smaller question first. These models already generate sound natively — but that sound is attached, not understood. The next chapter is about the modality AI handled first and modelled last: the one that talks back.

introduces →video modeltemporal consistencytemporal driftnative audioSora 2VeoWan
← previousOmni-Modelsnext →Audio Models