Models of the World
Everything so far predicts tokens. But a mind that can imagine what happens next — without acting — is doing something different. It has a world model: an internal simulator of how situations evolve.
Model-based reinforcement learning
A model-based RL agent learns two things: a predictor of what the world does, and a policy that plans within it. This is far more sample-efficient — you can “practice” in imagination, generating hypothetical rollouts without risking the real world. The catch is that model errors compound: a small inaccuracy becomes a wildly wrong prediction many steps out.
Learning physics from video
Self-supervised video models train on the same trick as language models — predict what comes next — but on frames instead of tokens. To predict the next moment of a video, you must implicitly learn objects, motion, gravity, and cause and effect. There is evidence that capable video predictors develop genuinely useful internal representations of physical structure.
JEPA (Joint Embedding Predictive Architecture) takes a further step: predict in representation space rather than in pixels. Instead of hallucinating exact future pixels — which are often unpredictable and unimportant — the model predicts an abstract embedding of the future. The argument is that this is closer to how abstraction works, and avoids wasting capacity on irrelevant detail.
Two roadmaps to general intelligence
One camp says scale next-token prediction and the world model emerges for free. The other says a dedicated simulator is necessary, because pixels and tokens are the wrong level of abstraction. The truth is unresolved — and likely involves both.