AI Foundationspredict · compress · act
act XIV

AI for Science

The world-model chapter built simulators so agents could practise. This one builds simulators for systems where the ground truth is a physical measurement — the weather, and the space of possible materials.

64

Simulating the World

frontieras of 2026-09
before this →Seeing StructureWorld Models
2023GraphCast — A graph network forecasts the weather better than the operational physics system it learned from.
2023GNoME — Millions of candidate crystals are screened by a network — candidates, still waiting for a lab.
2024GenCast — A forecast becomes a spread of possible futures rather than a single answer — because weather is uncertain.

Chapter 61 asked for a model you could step: give it a state, take an action, get the next state. This chapter uses the same shape for a different purpose. Here the simulator is not a place to practise — it stands in for a physical calculation that is correct but far too slow to run as often as we would like.

The trick is old and unglamorous. It works precisely because there is something to check the answer against.

THE EXPENSIVE WAY — solve the physicsstate at time tintegrate the equationsfluid dynamics, thermodynamics…state at time t+1accuracy: yes · cost: a supercomputer, per stepTHE LEARNED WAY — imitate itstate at time ta neural networktrained to match the simulatorstate at time t+1accuracy: approximate · cost: a GPU, thousands of times fasterWHAT MAKES THIS SAFE TO TRUSTA surrogate model is only useful because the slow version exists and is right. The network istrained on the simulator’s own output, and judged against the simulator or the instrument.So the physics is not being replaced. It is being distilled — the expensive thing runs once toproduce the data, and the cheap approximation is what you then use a thousand times a day.this is also why the approach fails when there is no ground truth to distil from.
a surrogate model. The physics is solved step by step on a supercomputer — accurate, and enormously expensive. A learned model imitates state → next state directly: thousands of times faster and only approximately right, which is the trade being made.

Weather, where the stakes are daily

Numerical weather prediction — NWP — is one of the largest computations performed on Earth. Supercomputers integrate the equations of the atmosphere on a three-dimensional grid, advancing the whole planet’s state in small time steps, and the major weather services have been doing this for decades. It is accurate, it is physics, and it costs a great deal of money every day.

GraphCast (Google DeepMind, Science 2023) replaced the integrator with a learned model. It works on a 0.25° latitude–longitude grid, uses a graph neural network — a network whose units are points in space, whose edges are neighbourhoods, so information propagates locally along the mesh instead of everything attending to everything — and predicts the state ten days ahead autoregressively, over 227 variables. It outperformed “the most accurate operational medium-range weather forecasting system in the world”.

GenCast (DeepMind, Nature 2024) addressed the thing a single forecast cannot express: uncertainty. A weather forecast is inherently probabilistic — the question is not “will it rain” but “how likely is rain, and how much”. GenCast generates an ensemble of plausible futures using a diffusion model, the same family from the image chapters, so it produces a distribution over outcomes rather than one answer.

Two honest notes, because they are the load-bearing ones. First, both models are trained on reanalysis — the historical record of atmospheric states produced by the physics models. The learned model is distilled from decades of NWP output; it does not remove the physics, it inherits it. Second, “faster” is the entire selling point. Nothing here discovered anything about the atmosphere.

Materials, where the search space is the problem

Finding a new stable crystal is a search problem with an astronomically large space. The honest test of a candidate is density functional theory (DFT) — a quantum-mechanical calculation of whether the structure holds together — and it is expensive enough that screening millions of candidates is not feasible.

GNoME (Graph Networks for Materials Exploration, DeepMind, Nature 2023) used a graph network over crystal structures to do the screening, and reported the discovery of 2.2 million new crystals, of which roughly 380,000 were predicted to be stable — DeepMind’s own framing being that this is “equivalent to nearly 800 years’ worth of knowledge”. The candidate set was added to the Materials Project so that other researchers could use it.

The catch is in the word “predicted”. A stable-in-the-model crystal is a candidate. The pipeline from candidate to a material you can hold runs through synthesis, which is slow, fiddly and happens in a laboratory. The result is a vastly better map of where to look; it is not a warehouse.

Why these worked

the same reason proteins did

The ground truth is a physical calculation or a measurement, and the model’s job is to approximate it quickly. That means every claim is checkable: a forecast is verified against what the weather did, a crystal against whether it can be made. There is no question of whether the model “really” understands — only whether it is close enough, often enough.

Which is also the limitation. A surrogate cannot tell you anything the slow version could not. It accelerates a science; it does not redirect it. Every genuinely new result in this chapter is a new place to look, not a new law.

the through-lineNote the convergence with the world-model chapter. Both learn dynamics instead of deriving them. The difference is not the architecture — it is that here there is an instrument to be wrong against.
EXPENSIVEphysics, or a lab
DISTILlearn to imitate it
SCREENask it a million times

So the pattern holds a second time. Structure, forecast, crystal — in each case the target is a well-defined object and something external decides whether the answer is right.

That leaves the part of science that has no such object: deciding what to look at. Which is a different kind of problem, and the one where the hype outruns the evidence by the widest margin.

introduces →surrogate modelneural weather modelgraph neural networkGraphCastGenCastmaterials discoveryGNoME
← previousSeeing Structurenext →Doing Science