The Bottleneck
Chapter 7 introduced mutual information — how much one thing tells you about another. This chapter spends that currency on the question a representation has to answer: of everything in the input, what is worth keeping?
It is the same question as MDL, asked from the other side. MDL compares whole hypotheses from outside. The bottleneck asks what an internal layer should preserve. And the theory that answers it is one of the most elegant ideas in the guide — followed by one of the most instructive failures.
Keep what matters, drop the rest
In 1999 Naftali Tishby, Fernando Pereira and William Bialek proposed the information bottleneck: a good representation Z of an input X, for the purpose of predicting a target Y, should maximise the information it carries about the target while minimising the information it carries about the input.
Written out, the objective is a trade between two mutual informations:
maximise I(Z; Y) — keep what tells you the answer
minimise I(X; Z) — throw away everything else
That second term is the bottleneck, and it is where the theory earns its name: the representation is squeezed through a channel deliberately too narrow to carry the input whole, so it is forced to keep only what is relevant. This is not a vague appeal to simplicity like Occam’s razor. It is a quantity you can compute.
It is also Shannon’s other theorem wearing new clothes. Rate–distortion theory — the lossy-compression dual of the channel-coding chapter — asks how little distortion you can achieve at a given number of bits. The bottleneck is rate–distortion where “distortion” is measured against the label instead of the pixels. Same shape, different definition of what counts as damage.
The story that was too good
The obvious next move is to claim that this is what deep networks do. In 2015 Tishby and Zaslavsky argued exactly that. The picture was gorgeous: training has two phases. First a fitting phase, where the network pulls in information about the label and I(Z;Y) rises. Then a compression phase, where it starts discarding input detail and I(X;Z) falls. That compression is what generalisation is.
If true, it would have been the missing theory. It would explain why generalisation happens, why it takes longer for bigger models, why the training curve looks the way it does. On the information plane, training traces a curve that rises and then hooks back toward the left. The hook is the explanation.
Where it broke
In 2018 Andrew Saxe and colleagues tested all three of the theory’s claims directly, and did not find them. The summary of what they report:
- The compression phase does not always happen. Whether a network compresses depends on the activation function. With saturating nonlinearities like tanh it appears; with ReLU — the function modern networks actually use — it does not.
- Compression is not what causes generalisation. The two can be decoupled: networks can generalise without compressing, and can compress while memorising.
- The phase transition is an artefact of how mutual information is estimated in a deterministic network, not a property of learning.
So the beautiful picture was, to a large extent, a property of the activation function rather than of learning. That is the instructive part: it is a good theory, mathematically sound and worth knowing, that did not survive being checked against what real networks do.
Two different verdicts
The principle is intact. The information bottleneck is still a working method: if you build a representation by explicitly optimising the trade, you get useful compressed features. It is a design objective, and a good one.
The story is not. “Networks generalise because they compress” does not follow from the evidence, and the pretty training trajectory does not generalise past a particular activation function. Anyone who read the 2015 argument and took it as settled was reading a hypothesis, not a result.
Three chapters in, the score is: we can bound how much data we need, we can say which explanation is best, and we can describe what a representation ought to keep — but none of it explains the single most striking empirical fact in the field, which is that loss falls on a straight line when you plot it against compute on a log scale. That is the last question, and it is the one with the least settled answer.