Skip to text
Measure of Chance

Chapter XIV · 24 min

Estimation

From a sample to a claim about the world

“The object of statistical methods is the reduction of data.”

— R. A. Fisher, Statistical Methods for Research Workers (1925)

A sample is one outcome in a space of datasets. An estimator is a function of that outcome, aimed at a property of the world: a mean, a weight vector, a whole distribution. This chapter makes the aim precise — bias, variance, consistency — and then returns to maximum likelihood as a method with theorems, not just a slogan. Generalisation, in the last pages, is estimation of risk.

Estimators

Let be i.i.d. from a distribution indexed by , and let be a quantity you care about (often itself). An estimator is a random variable . It is a function of the sample, hence random; the thing it estimates is not.

  • The sample mean estimates .
  • The empirical risk estimates the true risk .
  • A trained network is an estimator of a prediction function — a much larger .

Bias, variance, mean squared error

The bias is the systematic error, . Unbiased means the expectation hits the target; it does not mean any one sample is close. The variance is . They combine in the mean squared error:

This is the same algebra as Chapter 7, now read as a design tradeoff. A more flexible model class lowers bias (you can hit more targets) and typically raises variance (the sample’s idiosyncrasies get fit). Regularisation, early stopping, and smaller architectures are variance-control. The popular “bias–variance tradeoff” in learning is this identity, applied to a predictor’s risk rather than to a scalar parameter.

Consistency

An estimator is consistent for if in probability as . The sample mean is consistent for the mean (LLN). Consistency is a large-sample promise: with enough data, the estimator forgets its start. It says nothing, by itself, about .

Maximum likelihood, more carefully

Under regularity (the support — the set of values the data can take — does not jump when moves; the log-likelihood is differentiable; different knobs give different distributions; …), the MLE is consistent, asymptotically unbiased, and asymptotically normal:

where is the Fisher information — named for R. A. Fisher, not for fishing. It is a number (or a matrix, if there are several parameters) that says how much the data, on average, distinguish nearby values of :

Read the second form as expected curvature of the log-likelihood (Chapter XII). A sharp peak (large ) means the data distinguish nearby parameters well, so the MLE’s variance is small. The Cramér–Rao bound says that, among unbiased estimators, you cannot beat in variance. MLE attains this bound in large samples. That is the precise sense in which “likelihood is efficient.”

MAP, regularisation, and misspecification

MAP is MLE with a prior. As grows, a continuous prior is usually washed out and MAP and MLE agree. At small , the prior is doing the variance-control of the MSE decomposition. Weight decay is a Gaussian prior; an penalty is a Laplace prior. The vocabulary is interchangeable.

All of the efficiency theorems assume the model is well specified: some really generated the data. If not, MLE still minimises KL to the nearest member of the family (Chapter 13). That nearest member may be a useful lie, or a disaster. Model checking is the act of asking which.

Estimating risk: holdout and concentration

For a fixed predictor , the empirical risk on an i.i.d. test set of size is an unbiased estimator of true risk, with variance shrinking as . For bounded losses, Hoeffding’s inequality gives a tail:

(for loss in ). This is why a test set of a few thousand examples already makes accuracy a stable number — for one fixed . If you selected by looking at the same data, the estimator is no longer unbiased for the selected model’s risk. That is the reason for a held-out set, and the reason a training loss is not a reportable number.

Foundations studio: make the idea yours

This extended studio deliberately slows the pace. It is for a first-time learner who wants to recognize the idea in a new story, not merely reproduce a formula. Work with pencil and paper. Predict before calculating; redraw the pictures; and finish every numerical answer with a sentence in ordinary language.

A mental map before more algebra

An estimator is itself a random variable because a different sample would produce a different estimate. Good estimation therefore studies the sampling distribution, not only the one number we observed.

sampleestimatorsampling distributionbias/varianceuncertaintyWhen a formula feels unmotivated, move one box to the left.
A working map for Estimation. Cover the labels and reconstruct the chain from memory.

Do not treat the arrows as a theorem. They are a study aid. A strong probability habit is to move back one box whenever a formula feels unmotivated: ask what the experiment is, what information is available, and what quantity the question actually requests.

Three formulas worth being able to narrate

Read this line from left to right and explain what every symbol refers to in the experiment. If a symbol has no story, the model is not finished.

Now read the statement backwards: what would have to be known to use it? Backwards reading is often the difference between recognizing a formula and knowing when it applies.

Test the expression at an edge case or simple symmetric case. Probability formulas should survive sanity checks before you trust the arithmetic built on them.

Worked example ladder

A small experiment you can actually do

PredictRepresentComputeCheckExplainThe arithmetic is the middle of the loop, not the whole of it.
A five-step habit for every example in this chapter.

What usually goes wrong

When you notice this mistake, do not merely correct the final number. Return to the first line where the model became ambiguous. Probability errors are often representation errors wearing arithmetic clothing.

Questions beginners are right to ask

What makes an estimator consistent?

As sample size grows, it concentrates near the true parameter under the assumed data-generating model.

Why care about standard error?

It quantifies how much the estimator would vary across repeated samples and is the scale used by many intervals and tests.

Is MLE always unbiased?

No. MLEs are often asymptotically well behaved, but finite-sample bias can exist.

Where the abstraction earns its keep

For each application, ask what counts as an outcome, what the model treats as random, and which assumptions are approximations. This is how probability becomes a modelling language instead of a catalogue of formulas.

Connections: do not store chapters in separate boxes

Problem-solving clinic: from recognition to fluency

There is a stage where every worked example looks clear but a fresh problem still feels foreign. The cure is not another formula; it is practice choosing the representation. Before equations, do a sixty-second scan: identify the experiment, what is known, what remains uncertain, the quantity being asked for, the assumption doing the heavy lifting, and one impossible answer that gives you a sanity bound.

PredictRepresentComputeCheckExplainThe arithmetic is the middle of the loop, not the whole of it.
The expert loop returns every calculation to the original story.

Case clinic A: Estimating a coin probability

The sample proportion is unbiased for p and has variance p(1-p)/n. The uncertainty shrinks with n, but any one observed proportion still fluctuates.

Case clinic B: Biased variance estimator

Dividing the sum of squared deviations by n rather than n-1 creates a small downward bias when estimating population variance; Bessel's correction repairs it.

Solve or reason about it twice: once exactly and once with a rough estimate, simulation, or symmetry argument. If the two approaches disagree dramatically, investigate before trusting the more sophisticated calculation.

Debug a confident wrong answer

Two questions to answer without notes

What makes an estimator consistent? As sample size grows, it concentrates near the true parameter under the assumed data-generating model.

Why care about standard error? It quantifies how much the estimator would vary across repeated samples and is the scale used by many intervals and tests.

A notebook protocol for proficiency

Give this chapter one notebook page divided into four quadrants: picture, formula, example, mistake. Redraw the main visual from memory, narrate one formula in English, invent a fresh story using the same mathematics, and record the most tempting wrong move. Revisit the page after two days and again after a week.

A mastery check before you move on

Try these without looking back. If one item feels slippery, return to the corresponding example and rebuild it rather than memorizing the answer.

  1. Give a one-minute explanation of the chapter title to a curious teenager without a formula.
  2. Invent a tiny example with at most six elementary outcomes and solve it completely by enumeration.
  3. State one assumption that would make your example invalid and identify exactly which line would break.
  4. Draw the mental map from memory and connect at least two boxes to an earlier or later chapter.
  5. Write one question whose answer you still do not know. Good questions show that the concept has become active rather than passive.