Chapter XIII · 24 min
Entropy and Information
Surprise, divergence, and why models train
“Information is the resolution of uncertainty.”
A probability distribution is a complete story about uncertainty. Often you need a number that says how much story is there: how surprised a source is on average, how different two stories are, how much one variable tells you about another. Those numbers are entropy, Kullback–Leibler divergence, and mutual information. They show up wherever a forecast is scored — weather offices, compression, exams, and, among other trades, learning.
Surprise and entropy
If an event of probability occurs, a natural measure of surprise is . Certain events () are unsurprising; rare events are surprising. The logarithm makes independent surprises add. (The base chooses the unit: base 2 is bits, the unit of a fair coin; base is nats, convenient for calculus. A bit is “how much surprise a fair coin-flip carries.”)
The entropy of a discrete random variable is expected surprise:
Entropy is non-negative, zero only for a sure (deterministic) variable, and maximised — among distributions on a fixed finite set — by the uniform. A fair coin has entropy 1 bit; a coin with has less, because it is easier to guess.
For a continuous variable the same idea yields differential entropy , which can be negative and is not invariant to changes of units. Treat it as a relative tool, not as a “number of bits in a real number” — that last quantity is infinite.
Joint, conditional, chain
Joint entropy is the surprise in the pair. Conditional entropy is the leftover surprise in once is known. The chain rule is the same shape as probability’s:
Conditioning cannot increase entropy: . Information does not make the world noisier.
Mutual information
The reduction in surprise about upon seeing is the mutual information
It is symmetric, non-negative, and zero if and only if and are independent (Chapter V). Unlike correlation, mutual information sees nonlinear dependence: a perfect circle of pairs can have correlation zero and mutual information large.
KL divergence and cross-entropy
Kullback–Leibler is two surnames. Divergence here means “how different two distributions are,” not the divergence of vector calculus. Given a true distribution and a model , the extra surprise of using to encode samples from is
It is non-negative (Gibbs’ inequality / Jensen), and zero if and only if on a set of probability one. It is not a metric: it is asymmetric, and it does not obey the triangle inequality. is large when puts mass where does not; the other direction punishes the reverse.
Cross-entropy is expected surprise under the model, with data from the truth:
For a fixed true distribution , minimising cross-entropy in is exactly minimising KL, which is exactly maximum likelihood. That is why a scored forecast, a multiple-choice exam, or a weather office’s log-score is doing this arithmetic, whether or not the write-up says so.
Forward KL, reverse KL, and mode behaviour
Maximum likelihood minimises — forward KL. The expectation is under the data, so the model is punished for missing mass that the data has: it tends to cover all modes, even if it smears. Minimising the reverse is the variational / GAN-adjacent mood: the model is punished for putting mass where the data has none, and is allowed to drop modes. When a generated sample looks too average, or too sharp and incomplete, you are watching this asymmetry.
Foundations studio: make the idea yours
This extended studio deliberately slows the pace. It is for a first-time learner who wants to recognize the idea in a new story, not merely reproduce a formula. Work with pencil and paper. Predict before calculating; redraw the pictures; and finish every numerical answer with a sentence in ordinary language.
A mental map before more algebra
Information turns probability into a logarithmic scale of surprise. Entropy is the expected surprise under a distribution; cross-entropy and KL compare coding or predictive distributions.
Do not treat the arrows as a theorem. They are a study aid. A strong probability habit is to move back one box whenever a formula feels unmotivated: ask what the experiment is, what information is available, and what quantity the question actually requests.
Three formulas worth being able to narrate
Read this line from left to right and explain what every symbol refers to in the experiment. If a symbol has no story, the model is not finished.
Now read the statement backwards: what would have to be known to use it? Backwards reading is often the difference between recognizing a formula and knowing when it applies.
Test the expression at an edge case or simple symmetric case. Probability formulas should survive sanity checks before you trust the arithmetic built on them.
Worked example ladder
A small experiment you can actually do
What usually goes wrong
When you notice this mistake, do not merely correct the final number. Return to the first line where the model became ambiguous. Probability errors are often representation errors wearing arithmetic clothing.
Questions beginners are right to ask
Why logarithms?
Independent probabilities multiply, while we want independent information amounts to add; logarithms are the natural bridge.
Can KL divergence be negative?
No, though it is asymmetric and therefore not a metric.
Which log base?
Base 2 gives bits, base e gives nats. The mathematics differs only by a constant scale factor.
Where the abstraction earns its keep
For each application, ask what counts as an outcome, what the model treats as random, and which assumptions are approximations. This is how probability becomes a modelling language instead of a catalogue of formulas.
Connections: do not store chapters in separate boxes
Problem-solving clinic: from recognition to fluency
There is a stage where every worked example looks clear but a fresh problem still feels foreign. The cure is not another formula; it is practice choosing the representation. Before equations, do a sixty-second scan: identify the experiment, what is known, what remains uncertain, the quantity being asked for, the assumption doing the heavy lifting, and one impossible answer that gives you a sanity bound.
Case clinic A: Fair vs loaded coin
A fair coin has more uncertainty than a coin that almost always lands heads. Entropy captures this without caring whether we call heads 0 or 1.
Case clinic B: Twenty questions
Each balanced yes/no question can reveal about one bit. Good questions split the remaining probability mass roughly in half.
Solve or reason about it twice: once exactly and once with a rough estimate, simulation, or symmetry argument. If the two approaches disagree dramatically, investigate before trusting the more sophisticated calculation.
Debug a confident wrong answer
Two questions to answer without notes
Why logarithms? Independent probabilities multiply, while we want independent information amounts to add; logarithms are the natural bridge.
Can KL divergence be negative? No, though it is asymmetric and therefore not a metric.
A notebook protocol for proficiency
Give this chapter one notebook page divided into four quadrants: picture, formula, example, mistake. Redraw the main visual from memory, narrate one formula in English, invent a fresh story using the same mathematics, and record the most tempting wrong move. Revisit the page after two days and again after a week.
A mastery check before you move on
Try these without looking back. If one item feels slippery, return to the corresponding example and rebuild it rather than memorizing the answer.
- Give a one-minute explanation of the chapter title to a curious teenager without a formula.
- Invent a tiny example with at most six elementary outcomes and solve it completely by enumeration.
- State one assumption that would make your example invalid and identify exactly which line would break.
- Draw the mental map from memory and connect at least two boxes to an earlier or later chapter.
- Write one question whose answer you still do not know. Good questions show that the concept has become active rather than passive.