AI Foundationspredict · compress · act
act X

The Frontier

Predicting the next token is one way to learn how the world works. Building an internal simulator is another — and maybe the deeper one.

42

Models of the World

before this →Thinking Before Answering

Everything so far predicts tokens. But a mind that can imagine what happens next — without acting — is doing something different. It has a world model: an internal simulator of how situations evolve.

model-freepolicyworldact, observe, repeat — no imaginationmodel-basedpolicyworld modelrolloutplan inside the model before acting
model-free: react. model-based: simulate, then act.

Model-based reinforcement learning

A model-based RL agent learns two things: a predictor of what the world does, and a policy that plans within it. This is far more sample-efficient — you can “practice” in imagination, generating hypothetical rollouts without risking the real world. The catch is that model errors compound: a small inaccuracy becomes a wildly wrong prediction many steps out.

Learning physics from video

Self-supervised video models train on the same trick as language models — predict what comes next — but on frames instead of tokens. To predict the next moment of a video, you must implicitly learn objects, motion, gravity, and cause and effect. There is evidence that capable video predictors develop genuinely useful internal representations of physical structure.

JEPA (Joint Embedding Predictive Architecture) takes a further step: predict in representation space rather than in pixels. Instead of hallucinating exact future pixels — which are often unpredictable and unimportant — the model predicts an abstract embedding of the future. The argument is that this is closer to how abstraction works, and avoids wasting capacity on irrelevant detail.

Two roadmaps to general intelligence

predict tokens vs. model the world

One camp says scale next-token prediction and the world model emerges for free. The other says a dedicated simulator is necessary, because pixels and tokens are the wrong level of abstraction. The truth is unresolved — and likely involves both.

the through-lineA world model is compression at the level of dynamics. It is the same thesis as chapter 1, applied to time.
OBSERVEwatch how the world moves
COMPRESSlearn a simulator, not a recording
PLANimagine, evaluate, then act
introduces →world modelmodel-based RLself-supervised videoJEPA
← previousThinking Before Answeringnext →The Open Question