Chapter XVI · 14 min
Further Reading
Where the subject opens out
“I have always imagined that Paradise will be a kind of library.”
A first course is a door, not a room. What follows is a guided shelf: not every book that exists, but the ones this text was written to hand you to, with notes on tone, difficulty, and what each does that we did not. Read against your grain. If this volume felt computational, take a dose of Feller; if it felt chatty, take Bertsekas; if it felt silent on belief, take Jaynes.
First courses
Introduction to Probability
Joseph K. Blitzstein and Jessica Hwang
The book this one is in debt to. Stories first, then the theorem that makes the story inevitable. Indicator variables, symmetry, and conditioning are treated as methods rather than as definitions you recite. The problem sets are the real text; Stat 110, the accompanying Harvard course, is the spoken version. If you want a second pass through the same arc at greater length, start here and do the exercises in order.
A First Course in Probability
Sheldon Ross
The standard undergraduate workhorse. Combinatorial problems in great quantity, clean statements, less conversation. Excellent as a drill companion: when you finish a chapter here, try the matching problems in Ross until the counting feels automatic. The later chapters on limit theorems and simulation are more terse than Blitzstein; use them as a checklist, not as a first explanation.
Introduction to Probability
Dimitri P. Bertsekas and John N. Tsitsiklis
Compact, engineering-flavoured, and unusually honest about modelling. The authors write as if you might have to put a probability on a real system next week. Discrete and continuous are developed in parallel rather than in sequence, which is bracing if you have just come from this book’s ordering. The MIT OCW lectures pair with it.
Rigorous next steps
An Introduction to Probability Theory and Its Applications, Volume I
William Feller
The desert-island book. Feller’s Volume I stays discrete and still contains more ideas per page than most full curricula: generating functions, branching processes, fluctuation theory, the arc-sine law. The style is famous and not gentle — he expects you to follow a computation and a joke at the same time. Read slowly, with paper. Volume II is a different, more analytic mountain; do not start there.
Probability and Random Processes
Geoffrey Grimmett and David Stirzaker
The British thoroughbred: measure-aware without being a measure-theory course, with random processes taken seriously. Markov chains, Poisson processes, and stationary processes sit where this book stops. The exercises range from computational to sly. If you want one hard book after the first course, this is the usual correct answer in the UK; Feller is the usual correct answer everywhere else. You could do worse than both.
Bayesian and scientific reasoning
Probability Theory: The Logic of Science
E. T. Jaynes
A manifesto that treats probability as extended logic — Cox’s theorems, Laplace resurrected, priors as the record of information. It is opinionated, occasionally unfair to frequentists, and unmatched at making you feel that Bayes’ theorem is a moral duty. Read it after you can compute, not before. The unfinished later chapters are still worth the published ones.
Information Theory, Inference, and Learning Algorithms
David J. C. MacKay
The rare book that is simultaneously about bits, Bayes, and neural networks, and is good at all three. MacKay’s Monte Carlo chapters and the treatment of the evidence (Occam factors) are the natural sequel to our Chapter 12. The whole volume is freely available from the author’s site; work the exercises, especially on typical sets and on variational ideas in embryonic form.
Information and estimation
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas
The cleanest second course after Chapter 13. Entropy, typical sets, KL, mutual information, and the source and channel theorems, written so that the proofs feel inevitable. Chapter 2 of Cover & Thomas is the usual next week after this book’s entropy chapter. Not an ML text; it is the reason ML’s losses have names.
Statistical Inference
George Casella and Roger L. Berger
The standard mathematical-statistics sequel to Chapter 14: exponential families, completeness, UMVUE, likelihood ratio tests. Heavier than this book, and the right place to go if you want to know what “efficient estimator” means with hypotheses written in full. Skip to the likelihood and interval chapters first if you are coming from applications rather than from a statistics degree.
Probability in the world
Pattern Recognition and Machine Learning
Christopher M. Bishop
Not a probability text, but the probability chapters — 1 and 2 — are a complete second course in the distributions that models actually use: Beta-Binomial, Dirichlet-Multinomial, Gaussians and their conjugates, the exponential family. Graphical models later in the book are the organised form of our remarks on conditional independence. Dense, notation-heavy, worth owning as a reference even if you never read it cover to cover.
Causality
Judea Pearl
The book that made “correlation is not causation” into a calculus instead of a shrug. do-calculus, graphs, interventions versus observations. Our Chapter 10 stops at the shrug; Pearl starts there. It is not easy, and it is not optional if you want to talk about “the effect of” anything a model did not randomly assign. The slimmer Book of Why is the public-facing version; this is the one with theorems.
Problem books
Fifty Challenging Problems in Probability
Frederick Mosteller
A thin classic. Each problem is a story (the three-cornered duel, the vanished square, the cab problem), and each solution is a short education in modelling. Work them without looking, then look. Several are famous because the wrong sample space is so tempting — a useful vaccine after our Chapter 1.
Online, and a caution
Harvard Stat 110 (Blitzstein)
Lectures, a well-designed problem sequence, and a community of solutions. If you want this book with a voice and more exercises, Stat 110 is the continuation. Do not binge. One lecture, then problems, then the next.
3Blue1Brown, “Probability”
Visual intuition for Bayes, the binomial, and the central limit theorem that is clearer than most textbooks’ diagrams. Use it as a preview or a review, not as a substitute for computing a single by hand. (Grant Sanderson will be the first to say so.)
Thinking, Fast and Slow
Daniel Kahneman — with a caveat
This is psychology, not a probability text. It is unmatched on how humans fail the clinic problem of Chapter 4, base rates, and small samples. It will not teach you to compute, and it sometimes treats a heuristic failure as more universal than later replication would like. Read it to understand your own first guesses, then return to the axioms to repair them.