Entropy: The Price of Surprise
If you know the odds of each outcome, you can average the surprise. That average is entropy. It is the single most reused formula in this whole guide.
The formula, in words
Entropy asks two things of every possible outcome: how likely is it, and how
surprising is it if it happens. Surprise is measured as -log p — a rare event
(small p) gives a big surprise. Entropy averages that surprise over all outcomes,
weighted by how often each occurs.
The result is a single number in bits. It says: on average, how many yes/no questions do I need to learn the outcome?
Cross-entropy: the loss function hiding everywhere
Entropy measures surprise under the true odds. But a model never knows the true odds — it has a guess. When you score a guess against reality, you get cross-entropy: the average surprise the model felt while watching the truth unfold.
If a model is confident and right, cross-entropy is low. If a model is confident and wrong, cross-entropy is enormous.
This is exactly the loss function used to train language models. When a paper says “the model minimizes cross-entropy on the next token,” it is saying: the model is being punished in proportion to how surprised it was by what actually came next.
Training is surprise reduction
Entropy, cross-entropy, and perplexity are one concept at different temperatures. Perplexity is just two-to-the-cross-entropy: “how many equally likely options was the model effectively choosing between?”