AI Foundationspredict · compress · act
act IV

The Learning Floor

The last chapter of the floor. Loss falls on a straight line when you plot it against compute on a log scale — and the honest answer to 'why' is that the curve is a property of the data, and we only half understand it.

16

Why Bigger Works

frontieras of 2026-09
before this →The BottleneckMore Is Different

Three chapters have given us a floor: data has a price, descriptions have a length, and representations have a channel. None of them explained the most striking fact in the field, which the scaling chapter states as an observation — loss falls as a power law in compute, model size and data. Plot it on a log scale and you get a straight line, holding across more than seven orders of magnitude.

A straight line that stable is asking to be explained. This chapter is the attempt, and it starts by admitting that the classical theory predicted the opposite.

THE CURVE THE OLD THEORY PREDICTEDthe diperrorcapacity →WHAT IS ACTUALLY MEASURED — DOUBLE DESCENTcan now fit thetraining data exactlymodel size →rises, then fallsTHE POWER LAW (log–log both axes)compute, log scale →loss(log)loss ∝ compute^(−α)a straight line, across 7+ ordersof magnitude (Kaplan et al., 2020)the entropy floor of the datathe line cannot go below it — so it must bend, eventually
two curves that should not both be true. Small: the classical picture said error falls then rises — the U-curve. Large: what is actually measured is a straight line on a log scale, so the fitted exponent is the whole story.

First, the old curve broke

Chapter 11 showed the classical picture: bias–variance, and the familiar U-shaped curve. Small models underfit, large models overfit, and the best model is somewhere in the middle. That curve was taught for decades.

In 2019 it stopped describing reality. Mikhail Belkin and colleagues, and separately Preetum Nakkiran and colleagues, found the same thing from different directions: error falls, then rises as the model approaches the point where it has just enough capacity to fit the training data perfectly, then falls again and keeps falling as the model grows past that point. Two descents, not one. The dip is not a curiosity — it sits exactly where the interpolation threshold is, the boundary between a model that must generalise and a model that can memorise.

Nakkiran’s experiments found it in convolutional networks, residual networks and transformers. It was triggered by model size, by dataset size and training time. And the authors were explicit that this is a problem for theory rather than a triumph for it: the effect appears to be near-universal, and “we don’t yet fully understand why it happens.”

So the honest position after 2019 is that every practitioner intuition — bigger works, train longer — is empirically correct, and the classical theory that says otherwise is describing a regime that modern models left behind.

Then, a straight line

The scaling exponent is the reason the scaling laws matter. If loss falls as a power law, then loss ≈ k · compute^(−α), and the only thing that matters is α. A small change in the exponent changes everything about how much compute you should buy.

And α is a small number — on the order of a few hundredths. That is worth sitting with. It means a hundred-fold increase in compute buys a modest, predictable improvement, not a transformation, and that the improvement arrives at a known rate you can plan against. It is the most valuable empirical constant in the field and it was found by fitting curves, not by deriving them.

Why that exponent?

Two answers exist, and they are not the same kind of answer.

It is a property of the data. The strongest version of this is Bahri and colleagues’ explanation, which identifies four scaling regimes — variance-limited and resolution-limited, for both dataset size and model size — and derives them from the geometry of the problem rather than from optimisation. The argument: a model with infinite data has a well-defined best achievable loss, and finite data adds a variance term that decays predictably; separately, a model of finite size can only resolve structure down to some scale, and the data has structure at every scale. If the data’s structure is scale-free — as natural language and images experimentally are — the loss has no characteristic scale to stop at. You get a power law instead of a curve that flattens quickly.

In 2026 that line of argument got its sharpest result yet: for the data-limited case, exponents can be predicted from two statistical properties of the language alone — chiefly the way pairwise token correlations decay with distance. If that holds up, the exponent is not a property of transformers at all. It is a property of English.

It is a property of the model. The competing reading is that capacity and optimisation determine where the resolution limit sits. The exponent reflects how efficiently a given architecture uses the data it is given. This is the camp that would say a better architecture buys a better α.

Nothing in the guide can settle this, because nothing in the field has.

Two things to hold on to

what is solid, and what is not

Solid: the entropy floor from chapter 4 is real and unavoidable. Power laws cannot continue forever, because a model cannot do better than the intrinsic unpredictability of its data. Every straight line on that log plot is a straight line for now, and the end of it is the thing the open question at the end of the guide is actually about.

Not solid: “emergent abilities” — the claim that a capability appears suddenly at some scale. Schaeffer and colleagues argued in 2023 that this is often an artefact of the metric: a benchmark that scores an answer as entirely right or entirely wrong will turn a smooth improvement into an apparent jump. The capability may have been growing all along.

the honest noteBoth of the popular narratives about scale — that it will keep working, and that it is about to stop — are predictions beyond the evidence. The curve is real, its exponent is measurable, and its cause is contested. That is what a frontier looks like from the inside.
MEASUREloss falls as a power law
ASK WHYthe data's statistics, or the model's
THE FLOORit ends at the data's entropy

That closes the floor. Four chapters, one shape: learning has a price in data, a price in description length, a channel it must squeeze through, and a curve whose exponent comes from the data’s own statistics. Every one of those is the compression thesis from chapter 2, applied at a different level — and the one thing they share is that the theory is always one step behind the practice it is trying to explain.

The guide now returns to the story, at a point where machines had just learned to see.

introduces →double descentscaling exponentJared KaplanMikhail BelkinPreetum Nakkiran
← previousThe Bottlenecknext →Can Machines Think?