Distance Between Beliefs
We often need to compare two probability distributions: what a model believes, versus what is true. The standard ruler is the Kullback–Leibler divergence.
Reading KL
KL(p ‖ q) is the average number of extra bits you waste if you encode data drawn
from the true distribution p using a code built for the wrong distribution q. It is
never negative, and it is zero only when the two beliefs match exactly.
Two warnings. It is not symmetric — KL(p‖q) and KL(q‖p) differ, and choosing
which way to measure is a real modeling decision. And it is infinite wherever q
gives zero probability to something p allows: a guess that says “impossible” about
something that then happens is punished without limit.
Mutual information
A close cousin answers a different question: if I learn X, how much do I learn about Y? That is mutual information, the amount by which uncertainty about one variable drops once you see the other. It is zero exactly when the two are independent.
And perplexity — the number language models are actually scored on — is just
2^(cross-entropy). A perplexity of 20 means the model is, on average, as unsure as
if it were choosing among 20 equally likely next tokens.
Three rulers for the same idea
KL tells you how far a belief is from the truth. Mutual information tells you how far one variable is from another. Cross-entropy tells you how much a fixed guess costs. All three are built from the same logarithm.