Simulating the World
Chapter 61 asked for a model you could step: give it a state, take an action, get the next state. This chapter uses the same shape for a different purpose. Here the simulator is not a place to practise — it stands in for a physical calculation that is correct but far too slow to run as often as we would like.
The trick is old and unglamorous. It works precisely because there is something to check the answer against.
Weather, where the stakes are daily
Numerical weather prediction — NWP — is one of the largest computations performed on Earth. Supercomputers integrate the equations of the atmosphere on a three-dimensional grid, advancing the whole planet’s state in small time steps, and the major weather services have been doing this for decades. It is accurate, it is physics, and it costs a great deal of money every day.
GraphCast (Google DeepMind, Science 2023) replaced the integrator with a learned model. It works on a 0.25° latitude–longitude grid, uses a graph neural network — a network whose units are points in space, whose edges are neighbourhoods, so information propagates locally along the mesh instead of everything attending to everything — and predicts the state ten days ahead autoregressively, over 227 variables. It outperformed “the most accurate operational medium-range weather forecasting system in the world”.
GenCast (DeepMind, Nature 2024) addressed the thing a single forecast cannot express: uncertainty. A weather forecast is inherently probabilistic — the question is not “will it rain” but “how likely is rain, and how much”. GenCast generates an ensemble of plausible futures using a diffusion model, the same family from the image chapters, so it produces a distribution over outcomes rather than one answer.
Two honest notes, because they are the load-bearing ones. First, both models are trained on reanalysis — the historical record of atmospheric states produced by the physics models. The learned model is distilled from decades of NWP output; it does not remove the physics, it inherits it. Second, “faster” is the entire selling point. Nothing here discovered anything about the atmosphere.
Materials, where the search space is the problem
Finding a new stable crystal is a search problem with an astronomically large space. The honest test of a candidate is density functional theory (DFT) — a quantum-mechanical calculation of whether the structure holds together — and it is expensive enough that screening millions of candidates is not feasible.
GNoME (Graph Networks for Materials Exploration, DeepMind, Nature 2023) used a graph network over crystal structures to do the screening, and reported the discovery of 2.2 million new crystals, of which roughly 380,000 were predicted to be stable — DeepMind’s own framing being that this is “equivalent to nearly 800 years’ worth of knowledge”. The candidate set was added to the Materials Project so that other researchers could use it.
The catch is in the word “predicted”. A stable-in-the-model crystal is a candidate. The pipeline from candidate to a material you can hold runs through synthesis, which is slow, fiddly and happens in a laboratory. The result is a vastly better map of where to look; it is not a warehouse.
Why these worked
The ground truth is a physical calculation or a measurement, and the model’s job is to approximate it quickly. That means every claim is checkable: a forecast is verified against what the weather did, a crystal against whether it can be made. There is no question of whether the model “really” understands — only whether it is close enough, often enough.
Which is also the limitation. A surrogate cannot tell you anything the slow version could not. It accelerates a science; it does not redirect it. Every genuinely new result in this chapter is a new place to look, not a new law.
So the pattern holds a second time. Structure, forecast, crystal — in each case the target is a well-defined object and something external decides whether the answer is right.
That leaves the part of science that has no such object: deciding what to look at. Which is a different kind of problem, and the one where the hype outruns the evidence by the widest margin.