AI Foundationspredict · compress · act
act II

The Substrate

If entropy is the price, coding is the shopping. Match code lengths to how often things happen.

05

Codes: Saying More with Less

before this →Entropy: The Price of Surprise

Entropy told us the minimum average length of a message. Coding is the craft of getting close to it. The rule is simple and beautiful:

Common things get short codes. Rare things get long codes.

1.0.55.45e · 1t · 01a · 000q · 0010 = left · 1 = rightno code is a prefix of another
a Huffman tree for letters by frequency — decoding is a walk from the root

The prefix rule

A prefix code makes decoding unambiguous without separators: no code word is the beginning of another. 0 and 01 cannot both be codes, or 01 could mean one symbol or two. That single constraint is what forces rare symbols to be long.

Huffman coding builds the optimal prefix code for any known frequencies: repeatedly merge the two least likely symbols into a new node. The result is provably the shortest possible average code. Source coding — the theorem behind it — says you can never beat entropy, but Huffman gets you within one bit of it on average.

The rate is the whole game

why compression keeps showing up

The rate of a code is its average length in bits per symbol. Every storage format, every video codec, and (as we will see) every good model is judged by the rate it achieves on data it has not seen.

the linkA language model assigning probabilities to tokens is a code: the better its probabilities, the shorter the code. Predicting and compressing are literally the same arithmetic.
FREQUENCYhow often each symbol occurs
CODEshort for common, long for rare
RATEaverage bits per symbol — approach entropy
introduces →source codingprefix codeHuffman codingrate
← previousEntropy: The Price of Surprisenext →Channels: Pushing Through Noise