Why Bigger Works
Three chapters have given us a floor: data has a price, descriptions have a length, and representations have a channel. None of them explained the most striking fact in the field, which the scaling chapter states as an observation — loss falls as a power law in compute, model size and data. Plot it on a log scale and you get a straight line, holding across more than seven orders of magnitude.
A straight line that stable is asking to be explained. This chapter is the attempt, and it starts by admitting that the classical theory predicted the opposite.
First, the old curve broke
Chapter 11 showed the classical picture: bias–variance, and the familiar U-shaped curve. Small models underfit, large models overfit, and the best model is somewhere in the middle. That curve was taught for decades.
In 2019 it stopped describing reality. Mikhail Belkin and colleagues, and separately Preetum Nakkiran and colleagues, found the same thing from different directions: error falls, then rises as the model approaches the point where it has just enough capacity to fit the training data perfectly, then falls again and keeps falling as the model grows past that point. Two descents, not one. The dip is not a curiosity — it sits exactly where the interpolation threshold is, the boundary between a model that must generalise and a model that can memorise.
Nakkiran’s experiments found it in convolutional networks, residual networks and transformers. It was triggered by model size, by dataset size and training time. And the authors were explicit that this is a problem for theory rather than a triumph for it: the effect appears to be near-universal, and “we don’t yet fully understand why it happens.”
So the honest position after 2019 is that every practitioner intuition — bigger works, train longer — is empirically correct, and the classical theory that says otherwise is describing a regime that modern models left behind.
Then, a straight line
The scaling exponent is the reason the scaling laws matter. If loss falls as a power law, then loss ≈ k · compute^(−α), and the only thing that matters is α. A small change in the exponent changes everything about how much compute you should buy.
And α is a small number — on the order of a few hundredths. That is worth sitting with. It means a hundred-fold increase in compute buys a modest, predictable improvement, not a transformation, and that the improvement arrives at a known rate you can plan against. It is the most valuable empirical constant in the field and it was found by fitting curves, not by deriving them.
Why that exponent?
Two answers exist, and they are not the same kind of answer.
It is a property of the data. The strongest version of this is Bahri and colleagues’ explanation, which identifies four scaling regimes — variance-limited and resolution-limited, for both dataset size and model size — and derives them from the geometry of the problem rather than from optimisation. The argument: a model with infinite data has a well-defined best achievable loss, and finite data adds a variance term that decays predictably; separately, a model of finite size can only resolve structure down to some scale, and the data has structure at every scale. If the data’s structure is scale-free — as natural language and images experimentally are — the loss has no characteristic scale to stop at. You get a power law instead of a curve that flattens quickly.
In 2026 that line of argument got its sharpest result yet: for the data-limited case, exponents can be predicted from two statistical properties of the language alone — chiefly the way pairwise token correlations decay with distance. If that holds up, the exponent is not a property of transformers at all. It is a property of English.
It is a property of the model. The competing reading is that capacity and optimisation determine where the resolution limit sits. The exponent reflects how efficiently a given architecture uses the data it is given. This is the camp that would say a better architecture buys a better α.
Nothing in the guide can settle this, because nothing in the field has.
Two things to hold on to
Solid: the entropy floor from chapter 4 is real and unavoidable. Power laws cannot continue forever, because a model cannot do better than the intrinsic unpredictability of its data. Every straight line on that log plot is a straight line for now, and the end of it is the thing the open question at the end of the guide is actually about.
Not solid: “emergent abilities” — the claim that a capability appears suddenly at some scale. Schaeffer and colleagues argued in 2023 that this is often an artefact of the metric: a benchmark that scores an answer as entirely right or entirely wrong will turn a smooth improvement into an apparent jump. The capability may have been growing all along.
That closes the floor. Four chapters, one shape: learning has a price in data, a price in description length, a channel it must squeeze through, and a curve whose exponent comes from the data’s own statistics. Every one of those is the compression thesis from chapter 2, applied at a different level — and the one thing they share is that the theory is always one step behind the practice it is trying to explain.
The guide now returns to the story, at a point where machines had just learned to see.