AI Foundationspredict · compress · act
act II

The Substrate

One formula measures how uncertain a situation is — and every loss function in machine learning is a version of it.

04

Entropy: The Price of Surprise

before this →The Bit: A Yes or a No

If you know the odds of each outcome, you can average the surprise. That average is entropy. It is the single most reused formula in this whole guide.

fair coin50%50%H = 1 bitloaded coin95%5%H = 0.29 bits
entropy is highest when all outcomes are equally likely, and zero when one is certain

The formula, in words

Entropy asks two things of every possible outcome: how likely is it, and how surprising is it if it happens. Surprise is measured as -log p — a rare event (small p) gives a big surprise. Entropy averages that surprise over all outcomes, weighted by how often each occurs.

The result is a single number in bits. It says: on average, how many yes/no questions do I need to learn the outcome?

Cross-entropy: the loss function hiding everywhere

Entropy measures surprise under the true odds. But a model never knows the true odds — it has a guess. When you score a guess against reality, you get cross-entropy: the average surprise the model felt while watching the truth unfold.

If a model is confident and right, cross-entropy is low. If a model is confident and wrong, cross-entropy is enormous.

This is exactly the loss function used to train language models. When a paper says “the model minimizes cross-entropy on the next token,” it is saying: the model is being punished in proportion to how surprised it was by what actually came next.

Training is surprise reduction

the same idea, three dresses

Entropy, cross-entropy, and perplexity are one concept at different temperatures. Perplexity is just two-to-the-cross-entropy: “how many equally likely options was the model effectively choosing between?”

the takeawayA language model is a machine for lowering cross-entropy on text. Every architectural advance is judged by how far it pushes that number down.
ENTROPYsurprise under the true odds
CROSS-ENTROPYsurprise under the model's guess
PERPLEXITY2^loss — effective number of guesses
the threadWe now have a way to score a guess. Next: how to make messages as short as the entropy allows — coding.
introduces →entropysurprisecross-entropy
← previousThe Bit: A Yes or a Nonext →Codes: Saying More with Less