AI Foundationspredict · compress · act
act IV

The Learning Floor

Description length says which explanation to keep. This chapter asks the same question from inside the model: what should a representation throw away, and how much of the pretty theory survives contact with experiment?

15

The Bottleneck

decade
before this →The Shortest ExplanationEntropy: The Price of Surprise

Chapter 7 introduced mutual information — how much one thing tells you about another. This chapter spends that currency on the question a representation has to answer: of everything in the input, what is worth keeping?

It is the same question as MDL, asked from the other side. MDL compares whole hypotheses from outside. The bottleneck asks what an internal layer should preserve. And the theory that answers it is one of the most elegant ideas in the guide — followed by one of the most instructive failures.

I(X;Z) — how much of the input it still remembers →I(Z;Y) — how muchit knows about the answerthe bottleneck frontier1 · fitting2 · compression?depends on theactivation functiona ReLU netclimbs — noturning backabove the frontier is impossible · below it is merely wasteful
the information plane. A representation should sit high — knowing a lot about the answer — and far left, keeping little about the input. The curve is the best possible trade; everything under it is achievable, nothing above it is.

Keep what matters, drop the rest

In 1999 Naftali Tishby, Fernando Pereira and William Bialek proposed the information bottleneck: a good representation Z of an input X, for the purpose of predicting a target Y, should maximise the information it carries about the target while minimising the information it carries about the input.

Written out, the objective is a trade between two mutual informations:

maximise   I(Z; Y)      — keep what tells you the answer
minimise   I(X; Z)      — throw away everything else

That second term is the bottleneck, and it is where the theory earns its name: the representation is squeezed through a channel deliberately too narrow to carry the input whole, so it is forced to keep only what is relevant. This is not a vague appeal to simplicity like Occam’s razor. It is a quantity you can compute.

It is also Shannon’s other theorem wearing new clothes. Rate–distortion theory — the lossy-compression dual of the channel-coding chapter — asks how little distortion you can achieve at a given number of bits. The bottleneck is rate–distortion where “distortion” is measured against the label instead of the pixels. Same shape, different definition of what counts as damage.

The story that was too good

The obvious next move is to claim that this is what deep networks do. In 2015 Tishby and Zaslavsky argued exactly that. The picture was gorgeous: training has two phases. First a fitting phase, where the network pulls in information about the label and I(Z;Y) rises. Then a compression phase, where it starts discarding input detail and I(X;Z) falls. That compression is what generalisation is.

If true, it would have been the missing theory. It would explain why generalisation happens, why it takes longer for bigger models, why the training curve looks the way it does. On the information plane, training traces a curve that rises and then hooks back toward the left. The hook is the explanation.

Where it broke

In 2018 Andrew Saxe and colleagues tested all three of the theory’s claims directly, and did not find them. The summary of what they report:

So the beautiful picture was, to a large extent, a property of the activation function rather than of learning. That is the instructive part: it is a good theory, mathematically sound and worth knowing, that did not survive being checked against what real networks do.

Two different verdicts

the principle survives, the story does not

The principle is intact. The information bottleneck is still a working method: if you build a representation by explicitly optimising the trade, you get useful compressed features. It is a design objective, and a good one.

The story is not. “Networks generalise because they compress” does not follow from the evidence, and the pretty training trajectory does not generalise past a particular activation function. Anyone who read the 2015 argument and took it as settled was reading a hypothesis, not a result.

the through-lineNotice the pattern across these three chapters: a bound that is sound but not explanatory, a criterion that is right but not an algorithm, and a theory that is elegant but not true of the thing it was about. Theory in this field is usually a compass, rarely a map.
KEEPwhat predicts the answer
DISCARDeverything else
SQUEEZEthe narrow channel does it

Three chapters in, the score is: we can bound how much data we need, we can say which explanation is best, and we can describe what a representation ought to keep — but none of it explains the single most striking empirical fact in the field, which is that loss falls on a straight line when you plot it against compute on a log scale. That is the last question, and it is the one with the least settled answer.

introduces →information bottleneckrate-distortion theoryinformation planeClaude ShannonNaftali Tishby
← previousThe Shortest Explanationnext →Why Bigger Works