Chapter VII · 28 min
Expectation and Variance
A balance point, not a hope
“The expectation of the sum is the sum of the expectations, whether the variables are independent or not.”
A distribution is a whole shape. Often you need a number: the fair price of a game, the centre of a cloud of data, the average error of a model. Expectation is that number when “average” is taken with respect to chance. Variance is how widely the distribution refuses to sit there.
The English word is a saboteur. “I expect the die to show 3.5” is nonsense: the die has no such face. “I expect to win” is a mood. is neither. It is a balance point — the place a seesaw would rest if you piled probability as mass on the number line. Slow pictures first; the algebra is a caption.
Not “what you expect to happen”
Three things people mean by “expected”, only one of which is ours:
- Typical outcome. The mode, or a value with a lot of mass. A die’s typical face is 1, 2, 3, 4, 5, or 6 — not 3.5.
- Hope. What you would like. Irrelevant.
- Centre of mass. Weight each possible value by its probability and add. That is . It need not be a possible value, and it need not sit where most of the mass sits.
A weighted average, as a seesaw
If is discrete, taking values with masses ,
provided the sum converges absolutely. Physically: put a point mass at each number on a rigid rod. The fulcrum that balances the rod is .
If takes values on directly, the same quantity is . For equally likely finite , it is the ordinary arithmetic mean of the list of values. A long-run reading, for later: if you roll the die many times, the average of the faces you saw will settle near 3.5. That is the law of large numbers, Chapter XI. Expectation is the number it settles to — defined now, justified then.
When the masses are unequal
Let and . Most of the time you see a 1. The mean is not 1:
This is why “the average salary” can sit well above what most people earn, and why a model’s average error can look calm while a few examples are disasters. Rare large values have leverage. The seesaw is the whole lesson.
Linearity, with no independence
Independence would let you factor a product: . Linearity is about sums, and sums do not ask permission. Mixing those two sentences is the second popular trap of this chapter.
Indicators are the bridge back to events: . Many “how many” questions are sums of indicators, hence sums of probabilities. If you can name the things you are counting, you can often name the expectation without naming the law.
Variance and spread
The mean says where the seesaw balances. It is silent about whether the masses sit next to the fulcrum or at the two far ends. Variance is the expected squared deviation from the mean,
It is always non-negative, and zero only when is almost surely constant. The standard deviation restores the original units — “how many pips away”, not “pips squared”. Unlike expectation, variance is not linear: , and the covariance term vanishes when the variables are independent (more precisely, uncorrelated). Scaling: . Shifting a random variable never changes its spread; stretching it by 2 stretches the variance by 4.