Chapter VII · 28 min
Expectation and Variance
A balance point, not a hope
“The expectation of the sum is the sum of the expectations, whether the variables are independent or not.”
A distribution is a whole shape. Often you need a number: the fair price of a game, the centre of a cloud of data, the average error of a model. Expectation is that number when “average” is taken with respect to chance. Variance is how widely the distribution refuses to sit there.
The English word is a saboteur. “I expect the die to show 3.5” is nonsense: the die has no such face. “I expect to win” is a mood. is neither. It is a balance point — the place a seesaw would rest if you piled probability as mass on the number line. Slow pictures first; the algebra is a caption.
Not “what you expect to happen”
Three things people mean by “expected”, only one of which is ours:
- Typical outcome. The mode, or a value with a lot of mass. A die’s typical face is 1, 2, 3, 4, 5, or 6 — not 3.5.
- Hope. What you would like. Irrelevant.
- Centre of mass. Weight each possible value by its probability and add. That is . It need not be a possible value, and it need not sit where most of the mass sits.
A weighted average, as a seesaw
If is discrete, taking values with masses ,
provided the sum converges absolutely. Physically: put a point mass at each number on a rigid rod. The fulcrum that balances the rod is .
If takes values on directly, the same quantity is . For equally likely finite , it is the ordinary arithmetic mean of the list of values. A long-run reading, for later: if you roll the die many times, the average of the faces you saw will settle near 3.5. That is the law of large numbers, Chapter XI. Expectation is the number it settles to — defined now, justified then.
When the masses are unequal
Let and . Most of the time you see a 1. The mean is not 1:
This is why “the average salary” can sit well above what most people earn, and why a model’s average error can look calm while a few examples are disasters. Rare large values have leverage. The seesaw is the whole lesson.
Linearity, with no independence
Independence would let you factor a product: . Linearity is about sums, and sums do not ask permission. Mixing those two sentences is the second popular trap of this chapter.
Indicators are the bridge back to events: . Many “how many” questions are sums of indicators, hence sums of probabilities. If you can name the things you are counting, you can often name the expectation without naming the law.
Variance and spread
The mean says where the seesaw balances. It is silent about whether the masses sit next to the fulcrum or at the two far ends. Variance is the expected squared deviation from the mean,
It is always non-negative, and zero only when is almost surely constant. The standard deviation restores the original units — “how many pips away”, not “pips squared”. Unlike expectation, variance is not linear: , and the covariance term vanishes when the variables are independent (more precisely, uncorrelated). Scaling: . Shifting a random variable never changes its spread; stretching it by 2 stretches the variance by 4.
Foundations studio: make the idea yours
This extended studio deliberately slows the pace. It is for a first-time learner who wants to recognize the idea in a new story, not merely reproduce a formula. Work with pencil and paper. Predict before calculating; redraw the pictures; and finish every numerical answer with a sentence in ordinary language.
A mental map before more algebra
Expectation is a probability-weighted balance point; variance measures squared distance from that balance. Both are operators on distributions, not promises about what one trial will produce.
Do not treat the arrows as a theorem. They are a study aid. A strong probability habit is to move back one box whenever a formula feels unmotivated: ask what the experiment is, what information is available, and what quantity the question actually requests.
Three formulas worth being able to narrate
Read this line from left to right and explain what every symbol refers to in the experiment. If a symbol has no story, the model is not finished.
Now read the statement backwards: what would have to be known to use it? Backwards reading is often the difference between recognizing a formula and knowing when it applies.
Test the expression at an edge case or simple symmetric case. Probability formulas should survive sanity checks before you trust the arithmetic built on them.
Worked example ladder
A small experiment you can actually do
What usually goes wrong
When you notice this mistake, do not merely correct the final number. Return to the first line where the model became ambiguous. Probability errors are often representation errors wearing arithmetic clothing.
Questions beginners are right to ask
Why square deviations?
Squaring makes deviations nonnegative, penalizes large deviations, and interacts beautifully with algebra. Other spread measures exist, but variance is mathematically central.
Does E[g(X)] equal g(E[X])?
Usually not. Equality is special; Jensen's inequality tells us how convexity controls the direction.
When can variances be added?
For independent variables, or more generally when covariance terms vanish.
Where the abstraction earns its keep
For each application, ask what counts as an outcome, what the model treats as random, and which assumptions are approximations. This is how probability becomes a modelling language instead of a catalogue of formulas.
Connections: do not store chapters in separate boxes
Problem-solving clinic: from recognition to fluency
There is a stage where every worked example looks clear but a fresh problem still feels foreign. The cure is not another formula; it is practice choosing the representation. Before equations, do a sixty-second scan: identify the experiment, what is known, what remains uncertain, the quantity being asked for, the assumption doing the heavy lifting, and one impossible answer that gives you a sanity bound.
Case clinic A: A game with no possible mean payoff
A game pays 0 or 10 with equal probability. Its expectation is 5 even though 5 is impossible on any single play. Expectation is a long-run balance, not a forecast of one outcome.
Case clinic B: Indicator trick
For X=sum of indicators of success, linearity gives E[X] by summing success probabilities, even when the indicators are dependent. This is one of the most useful shortcuts in elementary probability.
Solve or reason about it twice: once exactly and once with a rough estimate, simulation, or symmetry argument. If the two approaches disagree dramatically, investigate before trusting the more sophisticated calculation.
Debug a confident wrong answer
Two questions to answer without notes
Why square deviations? Squaring makes deviations nonnegative, penalizes large deviations, and interacts beautifully with algebra. Other spread measures exist, but variance is mathematically central.
Does E[g(X)] equal g(E[X])? Usually not. Equality is special; Jensen's inequality tells us how convexity controls the direction.
A notebook protocol for proficiency
Give this chapter one notebook page divided into four quadrants: picture, formula, example, mistake. Redraw the main visual from memory, narrate one formula in English, invent a fresh story using the same mathematics, and record the most tempting wrong move. Revisit the page after two days and again after a week.
A mastery check before you move on
Try these without looking back. If one item feels slippery, return to the corresponding example and rebuild it rather than memorizing the answer.
- Give a one-minute explanation of the chapter title to a curious teenager without a formula.
- Invent a tiny example with at most six elementary outcomes and solve it completely by enumeration.
- State one assumption that would make your example invalid and identify exactly which line would break.
- Draw the mental map from memory and connect at least two boxes to an earlier or later chapter.
- Write one question whose answer you still do not know. Good questions show that the concept has become active rather than passive.