World Models
Go back to chapter 46 for the idea: instead of reacting to the world, a mind can hold a model of it and imagine what happens next. That chapter was about why — sample efficiency, planning in imagination, predicting in representation space. It was written as theory, with JEPA and model-based reinforcement learning as the evidence.
Then it became an industry. In 2025 three very different companies shipped things they all called world models, and they were not the same kind of object at all. Sorting them out is the work of this chapter.
The word is doing three jobs
World Labs — the lab Fei-Fei Li co-founded — wrote a short taxonomy in 2026 that is the clearest thing anyone has published on this, and it is worth taking at face value. It splits world models by output:
- A renderer “outputs observations in the form of pixels meant for human eyes, and the quality that matters most is visual fidelity.”
- A simulator “outputs state: a geometrically, physically or dynamically faithful representation of the world that humans and computer programs can both compute on and interact with.”
- A planner “outputs actions” — given an observation and a goal, it answers what the agent should do next.
The vocabulary earns its place because those three systems are judged by completely different things, and lumping them together under “world model” makes the field look more confused than it is. The same article adds the caveat that keeps this honest: the categories “are not fundamentally separate”, but “three projections of a single underlying understanding.”
A world you can walk around in
The renderer side became real in August 2025. Genie 3 (Google DeepMind) takes a text prompt and generates a photorealistic environment you can move through in real time — 720p at 24 frames per second, holding together “for a few minutes”. It was the first real-time interactive world model at that fidelity, and the difference from a video model is the whole point: a video model answers what does this look like, a Genie answers what happens if I turn left. DeepMind opened it to Google AI Ultra subscribers as Project Genie in January 2026.
The simulator side is where Marble (World Labs, generally available November 2025) sits: it takes a prompt, a photo, a video, or a panorama and produces a persistent 3D world — geometry you can re-enter, not frames you re-generate. That persistence is what separates a world from a clip. It is why simulators are the ones people can build games and training environments on.
The planner side is V-JEPA 2 (Meta, 2025). Its approach is the JEPA idea from chapter 46 at scale: pretrain on over a million hours of video without actions, then post-train a small action-conditioned version on a little real robot data, and it plans — including zero-shot control of a robot in a place it has never been. Note what it did not need: a rendering of the future, or even a picture of it.
And then there is NVIDIA Cosmos (January 2025), which is best understood as the industrial version of the argument: a family of open world foundation models built for generating physics-aware video and world states, aimed squarely at robots and autonomous vehicles rather than at viewers.
Why anyone wants a simulated world
You cannot crash a car ten million times to teach one to drive, and you cannot put a robot in ten thousand kitchens. A world model is a place to practice — unlimited, parallel, and resetable. DeepMind’s own framing is blunt about the stakes: world models make it possible “to train AI agents in an unlimited curriculum of rich simulation environments”, and they call it “a key stepping stone on the path to AGI”.
Nobody can score them yet
Interactive world models have an evaluation problem, and it is worse than the one in the omni chapter. A benchmark called WBench (2026) tried to fix it with 289 multi-turn cases across five dimensions, scoring 42 models on 22 metrics checked against human judgements. Its headline finding is the honest one: no model dominates all dimensions. Which means a leaderboard position tells you almost nothing — a system can be the best in the world at staying consistent and the worst at responding to what you did, and both facts are true at once.
What is still missing
Hold the claims at arm’s length. “Consistency for a few minutes” is the actual number, not hours. A renderer, by the taxonomy’s own admission, “carries no explicit understanding of three-dimensional structure” — it produces what a camera would see, not what the room is. Physics that looks right in a demo fails on contact and deformation. And every one of these systems is expensive to run.
Two things are solid, though. The output split is real and durable. It points somewhere specific: the planner is the one that ends in a body. The next chapter is about what happens when a model’s output is a movement.