Skip to text
Measure of Chance

Chapter XII · 24 min

Likelihood and Models

Fitting a story to data

“A communication system is a system for transmitting information; inference is a system for recovering it.”

— After Claude Shannon, A Mathematical Theory of Communication (1948)

Having a model and some data, how do you turn the knobs? A clinic fitting a prevalence, a pollster fitting a turnout, a biologist fitting a capture–recapture count, and — among other trades — a learning algorithm fitting a network are running the same two procedures: maximise a likelihood, or maximise a posterior. This chapter names them. The theorems you now own are why any of it is more than a recipe.

Likelihood

A parametric model is a family of distributions with a knob (or a few knobs). The knob is the parameter — a coin’s bias, a Poisson’s rate, a line’s slope. Write for “the chance of the data, if the knob is set to this value.” Having seen , the likelihood is that same function, now viewed as a function of . It is not a probability distribution on parameters — it need not even integrate to one in . Chapter IV already used the word; here the knob may be continuous.

The maximum likelihood estimator is any maximiser

For i.i.d. observations — independent repeats of the same experiment — the joint likelihood multiplies, and one maximises the log to turn the product into a sum. A fair-looking coin that showed 7 heads in 10 flips has Bernoulli likelihood , maximised at .

p = 0p = 1MLE at 7/10
Bernoulli likelihood after 7 heads and 3 tails. The curve is a function of p, peaked at the observed frequency. MLE reads off that peak.

MAP and the return of the prior

MAP is short for maximum a posteriori: pick the parameter that maximises the posterior instead of the likelihood alone.

The evidence can be dropped because it does not depend on . A prior that prefers small knobs is the Bayesian name for what some fitters call a penalty: taking logs, MAP for a bell-shaped prior on the knob is “likelihood minus a square of the knob.” Regularisation — shrinking wild estimates toward zero — is not a hack that probability is silent about. It is a prior, written in the loss.

Naïve Bayes, again

A classifier needs . Bayes’ theorem says

Naïve Bayes estimates the label prior from frequencies and estimates as a product of per-feature distributions. Prediction is argmax of the product. The independence is false; the decision often works. You can now state, precisely, which assumption you are buying.

Cross-entropy is negative log-likelihood

For a discrete label, the log-likelihood of a model that emits probabilities on a dataset is

up to a convention on the constant. Scoring a forecast with cross-entropy (Chapter XIII) is maximum likelihood under a categorical model — a model whose values are named types, not numbers. Calibration — whether a predicted 0.9 is right nine times in ten — is a separate question about the same object.

  • A monsoon desk scoring rain/shine with a log-score: MLE for a two-type (categorical) likelihood.
  • Fitting a line by least squares (Stranger in LA (Stranger in LA, X) ): MLE if the vertical scatter is modelled as a bell of constant width.
  • Among other trades, a classifier that turns scores into chances that add to 1 (a “softmax”) and then uses cross-entropy is doing this same MLE.

Generative and discriminative

A generative classifier specifies and uses Bayes’ theorem at prediction time. Naïve Bayes, linear discriminant analysis, and most language models (as next-token distributions) are generative. A discriminative classifier specifies directly — logistic regression, softmax networks — and never claims a density on . Both are probability. They answer different questions. If you need to sample features, detect outliers, or handle missing inputs, you want a joint. If you only need a label, the conditional is the shorter story, and often the better fit, because it does not spend capacity on .

Logistic regression is the discriminative twin of a naïve Bayes with Gaussian (or Bernoulli) class-conditionals: the posterior as a function of is a sigmoid of a linear score. Same decision boundary family, different training objective.

Calibration and the evidence

A predicted probability is calibrated when, among all cases given score , a fraction about are positive. Cross-entropy training encourages this but does not guarantee it, especially after a temperature or a stop-gradient. Reliability diagrams and proper scoring rules (log loss, Brier) are how you check. Accuracy is silent on calibration: a model that says 0.51 on every true item and 0.49 on every false one can be perfectly accurate and useless as a probability.

The evidence is the normaliser of Bayes’ theorem and a model-comparison score. A more flexible model can always fit, but it spreads prior mass thinner, and the evidence notices. This is Occam’s razor as arithmetic — MacKay’s book, in Further Reading, is the long form.

Overfitting is a confusion of sample and world

A model that interpolates the training set has made the empirical distribution certain. The true risk is an expectation under the world, not under the sample. When the hypothesis class is rich enough, the empirical minimum need not track the expectation — the LLN was stated for a fixed function, and you chose the function after seeing the data. Generalisation theory repairs this with uniform laws of large numbers. Regularisation, held-out sets, and cross-validation are the practical repairs. All of them are ways of remembering that the sample is one outcome in a sample space of datasets.

Foundations studio: make the idea yours

This extended studio deliberately slows the pace. It is for a first-time learner who wants to recognize the idea in a new story, not merely reproduce a formula. Work with pencil and paper. Predict before calculating; redraw the pictures; and finish every numerical answer with a sentence in ordinary language.

A mental map before more algebra

Probability treats parameters as fixed and data as random before observation. Likelihood takes the observed data as fixed and views the same formula as a function of the unknown parameter.

modelparameter thetaobserved datalikelihoodfit/compareWhen a formula feels unmotivated, move one box to the left.
A working map for Likelihood and Models. Cover the labels and reconstruct the chain from memory.

Do not treat the arrows as a theorem. They are a study aid. A strong probability habit is to move back one box whenever a formula feels unmotivated: ask what the experiment is, what information is available, and what quantity the question actually requests.

Three formulas worth being able to narrate

L(\\theta;x)=p_\\theta(x)

Read this line from left to right and explain what every symbol refers to in the experiment. If a symbol has no story, the model is not finished.

Now read the statement backwards: what would have to be known to use it? Backwards reading is often the difference between recognizing a formula and knowing when it applies.

\\hat\\theta_{MLE}=\\arg\\max_\\theta L(\\theta)

Test the expression at an edge case or simple symmetric case. Probability formulas should survive sanity checks before you trust the arithmetic built on them.

Worked example ladder

A small experiment you can actually do

PredictRepresentComputeCheckExplainThe arithmetic is the middle of the loop, not the whole of it.
A five-step habit for every example in this chapter.

What usually goes wrong

When you notice this mistake, do not merely correct the final number. Return to the first line where the model became ambiguous. Probability errors are often representation errors wearing arithmetic clothing.

Questions beginners are right to ask

Why use logs?

Products become sums, numerical underflow is reduced, and maxima are unchanged because log is monotone.

What is a model?

A family of probability distributions indexed by parameters or structures, together with assumptions about how data arise.

Can a high likelihood prove a model true?

No. It only says the observed data are relatively compatible under that model; misspecified models can still fit well.

Where the abstraction earns its keep

For each application, ask what counts as an outcome, what the model treats as random, and which assumptions are approximations. This is how probability becomes a modelling language instead of a catalogue of formulas.

Connections: do not store chapters in separate boxes

Problem-solving clinic: from recognition to fluency

There is a stage where every worked example looks clear but a fresh problem still feels foreign. The cure is not another formula; it is practice choosing the representation. Before equations, do a sixty-second scan: identify the experiment, what is known, what remains uncertain, the quantity being asked for, the assumption doing the heavy lifting, and one impossible answer that gives you a sanity bound.

PredictRepresentComputeCheckExplainThe arithmetic is the middle of the loop, not the whole of it.
The expert loop returns every calculation to the original story.

Case clinic A: Seven heads in ten tosses

For a Bernoulli model, the likelihood is proportional to p^7(1-p)^3 and is maximized at p=.7. The likelihood is not a probability distribution over p unless a prior and normalization are added.

Case clinic B: Gaussian mean

With known variance, maximizing the normal likelihood for the mean is equivalent to minimizing squared error, revealing why least squares appears so widely.

Solve or reason about it twice: once exactly and once with a rough estimate, simulation, or symmetry argument. If the two approaches disagree dramatically, investigate before trusting the more sophisticated calculation.

Debug a confident wrong answer

Two questions to answer without notes

Why use logs? Products become sums, numerical underflow is reduced, and maxima are unchanged because log is monotone.

What is a model? A family of probability distributions indexed by parameters or structures, together with assumptions about how data arise.

A notebook protocol for proficiency

Give this chapter one notebook page divided into four quadrants: picture, formula, example, mistake. Redraw the main visual from memory, narrate one formula in English, invent a fresh story using the same mathematics, and record the most tempting wrong move. Revisit the page after two days and again after a week.

A mastery check before you move on

Try these without looking back. If one item feels slippery, return to the corresponding example and rebuild it rather than memorizing the answer.

  1. Give a one-minute explanation of the chapter title to a curious teenager without a formula.
  2. Invent a tiny example with at most six elementary outcomes and solve it completely by enumeration.
  3. State one assumption that would make your example invalid and identify exactly which line would break.
  4. Draw the mental map from memory and connect at least two boxes to an earlier or later chapter.
  5. Write one question whose answer you still do not know. Good questions show that the concept has become active rather than passive.