Chapter IX · 26 min
Continuous Chance
Density, not mass — and which curve
“The error to which our observations are subject follows a law which can be represented by a curve.”
Not every chance lives on a list. A waiting time, a measurement error, a model’s predicted probability — these sit on a continuum. The new object is a density: a function whose integrals, not values, are probabilities. The value is not . That last probability is zero.
Density is not mass
A random variable is continuous (absolutely continuous, if one is being careful) when there is a non-negative integrable function such that
The cdf is still , and now at continuity points. Intervals of the same length in a tall part of the density are more probable than those in a low part — that is all the picture says.
| Word | Symbol | What it is | How to use it |
|---|---|---|---|
| pmf | p(k) | a mass at a point, for discrete X | sum p(k) over a set |
| density | f(x) | a height; not a probability | integrate f over an interval |
| cdf | F(x) | P(X ≤ x), always | P(a < X ≤ b) = F(b) − F(a) |
| support | where f > 0 | the values that can actually happen | uniform lives on [a,b]; exponential on [0, ∞) |
A field guide
| Family | Story | Lives on | Mean |
|---|---|---|---|
| Uniform | equally likely on an interval | [a, b] | (a + b) / 2 |
| Exponential | wait with no memory | [0, ∞) | 1/λ |
| Normal | sum of many small independent bits | all of ℝ | μ (spread σ) |
Three more names you will meet later, and can ignore until you need them: Beta lives on and is a prior on a coin’s ; Gamma generalises the exponential to waiting for several ticks; Student-t is a normal with a heavier tail, used when the variance is estimated from a small sample. They are not required for the rest of this book.
Uniform and exponential
The uniform law on has density there and zero elsewhere. It is the continuous analogue of “equally likely”: equal-length subintervals are equally likely. The standard uniform is the raw material of simulation: any cdf that can be inverted yields a random variable as .
The exponential law with rate has density on , mean , and the same memorylessness as the geometric: given survival to time , the remaining wait is a fresh exponential. It is the continuous waiting time for a Poisson process to tick once.
The normal law
The normal (Gaussian) family has density
Mean , variance . The standard normal is the case ; any other is a scaled shift, . To turn a question about into a question about , write
That fraction is the only linear algebra this family needs: subtract the centre, divide by the spread. The empirical rule: about 68% of the mass sits within one standard deviation of the mean, 95% within two, 99.7% within three. These are integrals of this curve, not commandments of nature — but many measurements are close enough that the numbers are useful.
Why this curve, among all bell-shaped functions? Because sums of many small independent contributions, suitably scaled, become normal. That is the central limit theorem, Chapter 11. It is also why Gaussian noise is the default in so many models: not because the world is Gaussian, but because leftover error is often a sum.
Try the bell lab.
Foundations studio: make the idea yours
This extended studio deliberately slows the pace. It is for a first-time learner who wants to recognize the idea in a new story, not merely reproduce a formula. Work with pencil and paper. Predict before calculating; redraw the pictures; and finish every numerical answer with a sentence in ordinary language.
A mental map before more algebra
In a continuous model, probability is spread as density. The density is not itself a probability; probability is area under the density over a set.
Do not treat the arrows as a theorem. They are a study aid. A strong probability habit is to move back one box whenever a formula feels unmotivated: ask what the experiment is, what information is available, and what quantity the question actually requests.
Three formulas worth being able to narrate
Read this line from left to right and explain what every symbol refers to in the experiment. If a symbol has no story, the model is not finished.
Now read the statement backwards: what would have to be known to use it? Backwards reading is often the difference between recognizing a formula and knowing when it applies.
Test the expression at an edge case or simple symmetric case. Probability formulas should survive sanity checks before you trust the arithmetic built on them.
Worked example ladder
A small experiment you can actually do
What usually goes wrong
When you notice this mistake, do not merely correct the final number. Return to the first line where the model became ambiguous. Probability errors are often representation errors wearing arithmetic clothing.
Questions beginners are right to ask
Why is P(X=x)=0?
A single point has zero width, hence zero area under an ordinary density. This does not mean the value is impossible in ordinary language.
Does the density determine the cdf?
Yes by integration, and where differentiable the cdf determines the density by differentiation.
Uniform over all real numbers?
No finite constant density on the entire real line can integrate to one.
Where the abstraction earns its keep
For each application, ask what counts as an outcome, what the model treats as random, and which assumptions are approximations. This is how probability becomes a modelling language instead of a catalogue of formulas.
Connections: do not store chapters in separate boxes
Problem-solving clinic: from recognition to fluency
There is a stage where every worked example looks clear but a fresh problem still feels foreign. The cure is not another formula; it is practice choosing the representation. Before equations, do a sixty-second scan: identify the experiment, what is known, what remains uncertain, the quantity being asked for, the assumption doing the heavy lifting, and one impossible answer that gives you a sanity bound.
Case clinic A: Uniform interval
For X uniform on [0,1], every exact point has probability zero, yet any interval has probability equal to its length. Uncountably many zero-probability points can form an event with positive probability.
Case clinic B: Exponential waiting
A constant hazard rate leads to an exponential waiting time. It is the continuous counterpart of the geometric memoryless distribution.
Solve or reason about it twice: once exactly and once with a rough estimate, simulation, or symmetry argument. If the two approaches disagree dramatically, investigate before trusting the more sophisticated calculation.
Debug a confident wrong answer
Two questions to answer without notes
Why is P(X=x)=0? A single point has zero width, hence zero area under an ordinary density. This does not mean the value is impossible in ordinary language.
Does the density determine the cdf? Yes by integration, and where differentiable the cdf determines the density by differentiation.
A notebook protocol for proficiency
Give this chapter one notebook page divided into four quadrants: picture, formula, example, mistake. Redraw the main visual from memory, narrate one formula in English, invent a fresh story using the same mathematics, and record the most tempting wrong move. Revisit the page after two days and again after a week.
A mastery check before you move on
Try these without looking back. If one item feels slippery, return to the corresponding example and rebuild it rather than memorizing the answer.
- Give a one-minute explanation of the chapter title to a curious teenager without a formula.
- Invent a tiny example with at most six elementary outcomes and solve it completely by enumeration.
- State one assumption that would make your example invalid and identify exactly which line would break.
- Draw the mental map from memory and connect at least two boxes to an earlier or later chapter.
- Write one question whose answer you still do not know. Good questions show that the concept has become active rather than passive.