AI Foundationspredict · compress · act
act II

The Substrate

KL divergence measures how wrong a guess is. Mutual information measures how much two things know about each other.

07

Distance Between Beliefs

before this →Channels: Pushing Through Noise

We often need to compare two probability distributions: what a model believes, versus what is true. The standard ruler is the Kullback–Leibler divergence.

true pguess qKL(p ‖ q)≥ 0, and 0 only if p = q
KL is the extra surprise you pay for using the wrong belief

Reading KL

KL(p ‖ q) is the average number of extra bits you waste if you encode data drawn from the true distribution p using a code built for the wrong distribution q. It is never negative, and it is zero only when the two beliefs match exactly.

Two warnings. It is not symmetric — KL(p‖q) and KL(q‖p) differ, and choosing which way to measure is a real modeling decision. And it is infinite wherever q gives zero probability to something p allows: a guess that says “impossible” about something that then happens is punished without limit.

Mutual information

A close cousin answers a different question: if I learn X, how much do I learn about Y? That is mutual information, the amount by which uncertainty about one variable drops once you see the other. It is zero exactly when the two are independent.

And perplexity — the number language models are actually scored on — is just 2^(cross-entropy). A perplexity of 20 means the model is, on average, as unsure as if it were choosing among 20 equally likely next tokens.

Three rulers for the same idea

divergence, information, surprise

KL tells you how far a belief is from the truth. Mutual information tells you how far one variable is from another. Cross-entropy tells you how much a fixed guess costs. All three are built from the same logarithm.

the bridgeThis is the last piece of pure information theory we need. From here on, we use it as a tool: to score models, to measure learning, and to define the objective everything else optimizes.
KLbelief vs. truth — asymmetry and all
MUTUAL INFOX and Y — how much they share
PERPLEXITY2^loss — the model's confusion
the hingeThat closes the substrate. We can now measure surprise, compress it, protect it, and compare beliefs. Next we turn to the instruments that change beliefs: probability, vectors, and gradients.
introduces →KL divergencemutual informationperplexity
← previousChannels: Pushing Through Noisenext →Belief, Updated