Intelligence Is Compression
Here is the bold claim that organizes this entire guide:
To understand something is to be able to compress it.
If you can predict the next word, you can store the text in fewer bits. If you can predict tomorrow’s weather, you can store the weather in fewer bits. A model that predicts well has, in a precise mathematical sense, found the pattern — and the pattern is exactly what is left after the randomness is squeezed out.
Why this is not just a slogan
Suppose I ask you to guess the next letter in a sentence. If you are good at it, then a file that stores only your surprises can rebuild the sentence exactly. Every time you guessed right, the encoder needed to send nothing at all. Your accuracy is the compression rate.
A model that memorizes every sentence it has ever seen compresses its training data perfectly and nothing else. That is not intelligence; it is a very large photocopier. The interesting thing is generalization: compressing the rules so tightly that they also describe data you have never seen.
Memorize vs. generalize
A lookup table copies. A model that has found the structure can throw the table away and still answer. The second kind is smaller than the data it explains — and that smallness is why we trust it.