Flip a coin and everyone tells you the same thing: the probability of heads is one half. It is the most familiar number in all of mathematics, and almost nobody can say what it actually means. It does not mean this next flip is somehow half-heads — the coin lands one way, fully. It does not mean that two flips give you exactly one head — flip twice and you'll often get two heads, or none. So here is the question this whole chapter turns on, the one most courses skip on their way to the formulas: what is that ½ a statement about, and where does it come from? Here is the short answer. Probability is not a mysterious force at all. It is a fraction you can work out by counting the ways a thing can land, resting on one quiet assumption that we'll find, poke, and eventually break. We'll build it in your hands. We count a coin, then two coins, then a pair of dice down three different roads to one number. We watch the count lie the instant the coin is biased. We see why "random" almost never means what you think. And we finish by turning the whole map around to find statistics — and machine learning — hiding on its back. There's no magic in any of it, just careful counting of a world too big to watch all at once.
01A probability is a fraction
Let's start with the smallest thing that has a probability at all: one coin, one flip. Every probability question begins with the same move, and we make it before computing anything. We list what can happen. A coin can land two ways: heads or tails, so the list has exactly two entries. That complete list of everything that could occur has a name we'll use all book long: the sample space, written Ω. Each single item in the sample space is an outcome. Flip the coin below and watch the two outcomes fill in. The thing we want — say, heads — gets tinted, and the probability is just how much of Ω that want takes up.
counted Ω: 1 of 2 heads → P = ½
The number came from counting Ω, not from any flip. Flip all day — P stays ½.
Fig. 1. A probability is just favourable outcomes ÷ total outcomes. Here the total is the sample spaceΩ = {heads, tails} — the two things that can happen. One of them, heads, is the outcome you want, so P(heads) = 1 ⁄ 2. Flip the coin as many times as you like: each flip lands somewhere, but the number on the right never moves. That's the quiet surprise — you knew it was ½ before you flipped even once, because you got it by counting the possibilities, not by watching the events.
One favourable outcome, two outcomes in Ω, so the probability of heads is 1/2. Notice we never watched a single real flip to get that number. We counted the possibilities instead. That is the entire engine, and it has exactly three parts.
Those three parts deserve their real names, because we'll say them thousands of times. The sample spaceΩ is the set of all possible outcomes. An event is any subset of Ω you care about: "it's heads," "it's an even roll," "at least one five." An event is literally just a handful of the outcomes lassoed together. And the probability of that event is a simple ratio, as long as every outcome is equally likely. Count the outcomes in the event, then divide by the count in Ω. Here are the three, colour-bonded, so the words and the picture lock together.
Ω — all 6 outcomes the die can land on
An event is nothing scary — it's just a handful of outcomes lassoed out of Ω, and its probability is the share of Ω it covers.
Fig. 2. Three words, one picture. Step (or hover) the buttons: Ω the sample space lights every outcome the die can produce — six of them. An event is just a subset: the gold loop lassoes the even faces {2,4,6} out of Ω. And the probability is only ever a share — how much of Ω the event covers, |E| / |Ω| = 3/6 = 0.5. Each term wears one colour and owns one region of the same picture, so the vocabulary and the geometry lock together.
We just lassoed a single event out of Ω. Reading an event someone hands you is easy. Building your own is the real skill, and it means turning any English phrase into the exact set of outcomes it stands for, chosen from the six faces of a die. A phrase like “at least a four” trips almost everyone. Does the four itself belong in the set? Read each phrase below, decide which faces it covers, and catch the ones that fool you.
Tap every face that is an EVEN number
Tap the faces (click or Enter), then check — the fraction ticks as you go.
Tap the faces, then Check.
Fig. 3. An event is nothing but a subset of Ω — so read a phrase and tap the exact faces it names. Try “at least a 4”: the honest event is { 4, 5, 6 }, three faces, so |E|/|Ω| = 3/6 = 0.50 — and the die flashes green on the 4 the instant you drop it, because “at least” quietly swallows its own boundary. Flip to Set→phrase and the lit { 1, 2, 3, 4, 5 } is “not a 6” (5/6 ≈ 0.83), never “less than a 5”. The event is never the words — it's the precise faces they point at.
Now scale it up by one notch to feel the machine work. Flip two coins. Most people's gut says the outcomes are "two heads, one of each, two tails," three in all. But that isn't the sample space, because "one of each" secretly happens two ways: heads then tails, or tails then heads. List them honestly and Ω has four equally-likely members: HH, HT, TH, TT. Pick an event below, say "at least one head," and watch the favourable outcomes light up and the fraction assemble itself.
pick an event — favourable outcomes light up
everything but TT → 3 of 4
Gut says 3 outcomes: two heads, one of each, two tails. But “one of each”
happens two ways — HT and TH are different orders — so Ω has 4 members.
That’s why it’s always /4, and why one-of-each is ½, not ⅓.
Fig. 4. Flip two coins and the honest sample space Ω (the set of everything that can happen) has four equally-likely members, not three: HH, HT, TH, TT — each a 1-in-4 outcome. Pick an event from the chips and its favourable outcomes light up while the fraction assembles. The trap is “one of each”: the gut lumps it into a single bucket and guesses ⅓, but heads-then-tails (HT) and tails-then-heads (TH) are two distinct outcomes — so it happens two ways, and its probability is 2/4 = ½. That extra way is exactly why Ω has four members, and why every probability here is counted over /4.
Three of the four outcomes contain a head — HH, HT, TH — so the probability is 3/4. The two-way "one of each" is exactly why it is three and not two. That's the discipline the whole subject rewards: enumerate, name, count, and only then divide. Get the sample space honest and the arithmetic is trivial.
So far every Ω has been handed to you, already listed and ready to count. The real skill is building one yourself, and that is exactly where beginners quietly lose an outcome. Here is a fresh experiment you have not met. Flip a coin and roll a die at the same time. List every outcome that pair can produce. If you miss one, the panel catches you, exactly as a wrong count would.
coin
die
Pick a coin side, then a die face — a chip drops into Ω.
You’re building Ω as a set: every outcome, none missed, none twice. Order-twins (BG ≠ GB) are different outcomes. Once the list is honest, the probability is only division.
Fig. 5.You build the sample space. Tap a coin side then a die face and a chip drops into Ω — the counter tallies what you’ve found but hides the true total. Press “I’m done” too early and it surfaces one outcome you overlooked; try to add a chip twice and it’s refused, because a set never lists a member twice. Switch to two children and the order-twins bite: claim “boy, one of each, girl” and it catches the missing GB — “one of each” happens two ways, so once Ω is honest, P(one of each) = 2/4 = 0.50. Setting up the problem is building Ω; the counting was always the whole job — the probability is only division.
One more thing to fix in place before we move on. A probability is always a number between 0 and 1. An event that can never happen has probability 0, because no outcomes favour it. An event that is certain has probability 1, because every outcome favours it. Everything real lives in between. Drag an event across the range below and watch where its fraction lands on the line of certainty.
3 of 6 → p = 1/2: dead centre, a coin flip
Why 0 and 1 are hard walls. A probability is just
favourable ÷ total. You can't have fewer than 0
winning faces, or more than all 6 — so the fraction is
trapped on the line: never below 0, never above 1.
Fig. 6. A probability is just favourable ÷ total, so it can only ever live on the line from 0 to 1. Drag the knob to choose how many of a fair die's six faces count as a win: none is the left wall (impossible), all six is the right wall (certain), and three of six lands dead centre — exactly a coin flip.
So a probability is a fraction pinned to a number line. One end, 0, is where the impossible event sits. The other end, 1, is where a certain event sits. Our coin's ½ lands dead centre between them. It's simple, but simple is not the same as shallow, as the next roll of the dice will show.
02Many roads, one number
Real understanding shows up when different roads to the same number agree. We just built a probability by listing the sample space, but listing is only one road. So let's take a slightly bigger question and attack it three separate ways. Roll two dice. There are 36 equally-likely outcomes, six faces on the first die times six on the second, laid out as a grid below. Here's the question, and I want you to commit to a guess before you look. In how many of those 36 cells does at least one die show a five?
Roll two dice. How many of the 36 outcomes show at least one 5? Commit a guess:
Pick a guess — then watch the fives light up.
The (5,5) corner sits in the row and the column. Add both sixes and it gets counted twice — so subtract it once.
Fig. 7. Every roll of two dice is one of 36 equally-likely cells — die A down the side, die B across the top. Guess how many show at least one 5, then watch it build: the 5-row lights up (6 cells), the 5-column adds 6 more — and for a beat the count reads 12. But the (5,5) corner was just counted twice, because it lives in the row and the column. Flag it, subtract one, and the honest total is 6 + 6 − 1 = 11, so the probability is 11/36. That single shared corner is the whole reason it's 11, not 12 — the first taste of "don't double-count the overlap."
Most people guess six, or twelve. The honest count is 11: the six cells of the "5-row" plus the six of the "5-column," minus the corner where they overlap, since (5,5) would otherwise be counted twice. Eleven favourable cells over thirty-six gives 11/36, about 30.6%. Hold onto that number. We arrive at it again without ever drawing the grid.
That grid held 36 cells, and we counted every one as equally likely. But why 36, and not fewer? The sum of two dice runs from 2 to 12, which is only eleven different outcomes to list. So surely the chance of a seven is just one in eleven, about 9%? Choose a sample space below and let it roll. Only one of the two lists is telling the truth about equal weight.
Two spaces, one question — roll to break the tie.
A 7 can be built six ways (1-6, 2-5, … 6-1); a 2 only one (1-1).
The tidy 11-tile list looks even, but the tiles carry unequal weight — so
favourable ÷ total is only honest on the 36 pairs.
Fig. 8. The question is fixed — P(sum = 7) — but the toggle picks the sample space. The tidy 11 sum-tiles look even and say 1/11 ≈ 0.091; the 36 ordered pairs say 6/36 ≈ 0.167. Press Roll to actually throw two dice twelve thousand times: the rolled bar homes onto ≈0.167, landing on the green 6/36 tick and missing the red 1/11 one. The inset shows why — a 7 is six pair-ways, a 2 is only one, so the sum-tiles never carried equal weight. Equally-likely is a property you earn by symmetry, never a label you assume.
Here's a second road that never lists a single cell. Think of it as two steps. Either the first die is a five, with probability 1/6, or it isn't a five, with probability 5/6, and then the second die rescues us with a five (another 1/6). Those two ways cannot overlap, so they are disjoint and we simply add them: 1/6 + (5/6)(1/6). Step through it and watch it collapse to the same fraction.
step 1 / 5
Split by the first die: is it a 5, or not?
Why may we just add? The two cases share no single outcome — the first die
can't be a 5 and not-a-5 at once. Disjoint pieces simply sum:
P(A or B) = P(A) + P(B).
Fig. 9. The same 11/36 down a second road — no 6×6 grid, just cases. Step through it: split on the first die. Case A — it's a 5 (probability 1/6 = 6/36); the event is already true no matter the second die. Case B — it isn't a 5 (5/6) but the second one is (1/6), worth 5/6 × 1/6 = 5/36. The two cases can never both happen — the first die is either a 5 or it isn't — so they overlap in zero outcomes, which is exactly the licence to add the pieces: 6/36 + 5/36 = 11/36. That's the whole reason splitting into disjoint cases works.
That comes to 6/36 + 5/36 = 11/36, identical to the grid count and reached by splitting into cases instead of listing every cell. Two roads, one number. Now for the third road. This one doesn't reason at all. It just rolls.
We don't have to trust the arithmetic. If the probability really is 11/36, then rolling two dice a hundred thousand times should turn up "at least one five" in about 30.6% of them, roughly 30,600 rolls. So let's simulate it and let the empirical frequency settle onto the theoretical value in front of us.
rolls 0 · with a five 0
press Run — roll two dice, watch the fraction
No formula in here — just rolls. Yet the blind average crawls onto the 11/36 we reasoned out. That's why you can trust the arithmetic.
Fig. 10. A third road — one that never reasons. Hit Run and the machine just rolls two dice, over and over, counting how often at least one shows a 5. Early on the running fraction lurches — a single lucky roll swings it wildly. But as the rolls pile into the thousands, the wobble dies and the blue line crawls onto the gold one: 11/36 ≈ 0.306, the exact value we got from 1/6 + (5/6)(1/6). Nobody told the dice that number. That is why you can trust the arithmetic — the world, rolled blindly, reproduces what you computed on paper.
The simulated fraction wobbles at first, then homes in on 11/36 ≈ 0.306 as the rolls pile up. Three roads: list the outcomes, split them into cases, simulate them by rolling. Every road lands on the same value, and that agreement is the tell that the number is real.
Why does that agreement matter so much? Because it means the probability is a property of the world, not of your clever method. You didn't invent 11/36. You uncovered it, and any honest path arrives at the same place.
Pick a road — or run all three.
Agreement across independent roads is the signature of a real number — the probability belongs to the dice, not to your method. That's why we cross-check.
Fig. 11. Roll two dice; what's the chance of at least one 6? Three roads that share nothing: ① List every outcome and count (11 of the 36 cells), ② Split it with the complement. The only way to miss is not-6 twice, which is (5/6)² = 25/36. Everything else is a hit, so the chance is 11/36 — and ③ Simulate, rolling thousands of times and watching the tally settle on 0.306. Replay each, or run all three. They land on the same number because the probability is a property of the dice, not of the method you used to find it — and that agreement across independent roads is exactly why cross-checking is trustworthy.
Keep this habit: when a probability is hard one way, come at it from another and check. Checking a hard count against a second road is the single most reliable move in the whole subject, and we'll lean on it in every chapter to come.
03The beam nobody mentions
Everything so far has quietly leaned on four words: "when every outcome is equally likely." That phrase is the load-bearing beam under "favourable over total," and like a good beam you never notice it until it fails. So let's make it fail on purpose. Take our coin and bias it, weighting it so heads comes up more often than tails. The naive count still says one favourable outcome out of two, still 1/2. But drag the bias below and watch the coin's true long-run frequency peel away from that count.
Fair coin — ½ and the flips agree
The arithmetic never broke: ½ is still ½. What broke is the formula's hidden promise that both faces are equally likely — false the instant you load the coin.
Fig. 12. Probability's oldest formula — favourable outcomes over total — quietly assumes every outcome carries equal weight. Drag the bias: the count can only see two faces, so it stays pinned at ½ forever, while the coin's true long-run frequency (the value a million flips would centre on) peels away to match the bias. The formula didn't miscalculate; its hidden "equally likely" assumption simply stopped being true.
The count says 1/2. The world says 0.7, which is seven heads in every ten flips. The formula didn't break because the arithmetic failed. It broke because its hidden assumption was false. "Favourable over total" only equals the probability when the outcomes carry equal weight. Bias one outcome, and counting the two of them like equals is simply wrong.
There's a second, cleaner way to break the count, one where counting doesn't just lie: it becomes meaningless. Spin a pointer on a circle and ask for the probability it lands on exactly 90°. The sample space is now infinite, since every real angle between 0 and 360 is an outcome. So "favourable over total" is one over infinity, which is 0, and that holds for every single angle. Yet the pointer does land somewhere. Counting has run out of road.
each of 4 slots weighs a real 1/4
∞ exact angles, each weighing 0 — and 0+0+0+… still stacks to 1.
That’s why you can’t count a continuum: there’s nothing to count one of.
Fig. 13. Hit Spin and the pointer clearly stops on some angle θ — a real number with endless decimals. Now ask the honest question: what were the odds it landed on exactly that angle? Pretend the dial has only N equally-likely slots and each is worth 1/N; drag the resolution up and watch that green slice — on the dial and on the bar — shrink toward nothing, while the whole bar stays pinned at total = 1. Then hit ∞ for the real spinner: every exact angle now weighs precisely 0, yet the pointer always lands. That’s the wall counting hits — with infinitely many outcomes, each single one carries weight 0, and adding up zeros can never make 1. A continuum has to be measured, not counted.
So counting a discrete, fair sample space is not the definition of probability. It's the easiest special case of something more general. When fairness fails, or when the outcomes go continuous, we'll have to replace "count the outcomes" with "measure the set." That means assigning weight to regions of Ω directly: area under a curve rather than a tally of points.
Flat weight — tally = area = 0.40
Counting favorable outcomes only works when every outcome carries equal weight — the flat setting. Tilt the weight and the tally goes blind, but the area follows. That's why Chapter 3 defines probability as a measure on Ω: area, not tally. Counting was only its equally-likely corner.
Fig. 14. From count to measure — the same sample space Ω, read two ways. In tally mode we split Ω into equally-likely slices and read P(A) as favorable / total. In area mode we read P(A) as the weight (area) sitting over the region A. With the weight flat, the two answers are identical — but tilt the weight and the tally freezes at 0.40 while the true probability climbs: counting can't see uneven weight, only area can. That is exactly why Chapter 3 rebuilds probability as a measure on Ω — area, not tally — and reveals that counting equally-likely outcomes was only its special, flat-weighted corner.
That upgrade, from a count to a measure, is the whole job of Chapter 3, and now you know exactly why we'll need it. For the rest of this chapter we'll stay in the fair, countable world where the ratio holds. I just wanted you to see the beam before we lean on it.
04What "random" really means
Time to confront the word we've been using without defining it: random. A coin obeys Newtonian mechanics, full stop. It's tempting to think a coin flip is random the way a dice-god decides it, as if the universe genuinely doesn't know the answer until it lands. That's almost never what's happening. If you knew the exact upward force, the exact spin, the air, and the table, the outcome would be completely determined before the coin left your thumb. Set those knobs below and the "random" flip becomes a predictable one.
air 0.86 s · 12.0 half-turns · flip line 0.50 away
Locked in — 12 half-turns → Heads
Nothing is random here. Force and spin decide the answer before release. It only feels like chance because no thumb repeats the same spin twice.
Fig. 15. A flip isn't magic — it's Newton. Force sets how long the coin is airborne; spin sets how fast it turns; multiply them and you get a fixed number of half-turns. Land on an even count and it comes up the way it started (Heads); odd, and it's Tails. The answer is decided the instant it leaves your thumb. So why does it feel random? Watch the gold error band: a real thumb can't hold the spin tighter than about ±0.2 rev/s, and near a gold flip-line that wobble reaches clear across the boundary. Determined in principle, uncallable in practice — that is all "random" ever means for a coin.
Here's the reframe that changes how you'll read this entire book. For a coin, "random" doesn't mean undetermined. It means deterministic but too complex to follow. The flip only feels random because those initial conditions are absurdly sensitive, and we can't measure them precisely enough to track. Probability is the tool we reach for precisely when the bookkeeping defeats us.
Once you see it that way, the real power of a probability distribution shows up: it's a compression. Take a balloon of gas, around 10²³ molecules, which is a hundred billion trillion of them, each one hurtling and colliding. No computer will ever store that many positions and velocities. So we give up on tracking the molecules and describe the whole swarm with three numbers instead: pressure, volume, temperature. Slide the molecule count up below and watch an impossible ocean of detail collapse onto a handful of honest dials.
Few molecules — easy to track, but dials jitter
A probability distribution isn’t ignorance. It’s deliberate lossy compression — throw away which molecule is where, keep the three numbers you can actually use.
Fig. 16. Slide the molecule count from a handful toward ~10²³. The per-molecule ledger you would need — six numbers each (position + velocity) — explodes as 6N, while the swarm’s honest description stays just three dials: pressure, volume, temperature. That collapse is a probability distribution: a deliberate lossy compression, and one that grows more trustworthy (to ±1/√N) exactly as the raw detail becomes impossible to hold.
That's the north-star idea of the subject. A distribution is not a confession that we're ignorant. It's a deliberate, lossy compression of overwhelming complexity into something we can hold and reason about. The gas didn't stop being deterministic. We just stopped pretending we could follow every particle, and we got a far more useful description in return.
Now, is anything truly random, or is it complexity all the way down? Here we have to be honest, because the answer splits in two. A coin is deterministic chaos: predictable in principle, hopeless in practice. But when a single radioactive atom decays, quantum mechanics says there is no hidden clock and no initial condition we're failing to measure. The decay time is irreducibly random, as far as physics can tell. The same maths describes the coin and the atom, but the metaphysics underneath is completely different.
On top: a fair ½ you can't predict
Same ½ on top, opposite worlds beneath: the coin hides a clockwork we can't track; the atom hides nothing. Probability describes the uncertainty either way — that's why it's universal.
Fig. 17. Two cards, one description. Up top the maths is identical: P(outcome) = ½, a 50/50 you can't call in advance. Peel each card and the metaphysics splits. Under the coin sits a hidden clockwork — the exact flick force, spin and air are all fixed, merely untracked; know them and you could predict the toss. It's deterministic chaos. Under the atom there is nothing — quantum mechanics says no secret clock decides when it decays. It's irreducibly random. Same ½, opposite worlds underneath — and probability describes the uncertainty in both. That's why it's the universal language of chance: it doesn't care whether the randomness is our ignorance or the universe's own dice.
Here's what's lovely about that split: probability doesn't care which kind you hand it. Whether the randomness is chaos we can't track or true quantum dice, the same fraction, the same distribution, and the same machinery all work. Probability is a language for uncertainty of any origin. That generality is exactly why it shows up in physics, finance, biology, and every machine that learns.
05The mirror: statistics
We've been travelling in one direction the whole chapter: from a known model, such as a fair coin or honest dice, forward to what the data should look like (½, 11/36). That direction isprobability. Given the rules of the world, predict the observations. But there's an obvious mirror question, and it's the one that actually pays your bills. You almost never know the model. You get handed the data and have to work backwards to the model that produced it. Flip the map below and watch the arrow reverse.
Rule P(H)=0.5 → predict ~5 heads in 10
Flip the arrow (or tap the map). Statistics isn't a new subject — it's this exact same map read backwards: same coin, same Ω, same fractions. You just start from the other end.
Fig. 18.The two-directional map. One picture holds the whole relationship. On the left, the model — the coin’s rule, P(H)=0.5, over its outcome set Ω = {H, T} (Ω is just “every result a flip can land on”). On the right, the data — one run of 10 flips that came up 6 heads. Probability reads the map left→right: given the rule, predict the data (a fair coin should give about n×p = 5 heads). Statistics reads the exact same map right→left: given the data, infer the rule (6 in 10 → best-guess p ≈ 0.6, the estimate written p̂). Flip the arrow and watch the same coin, the same Ω, and the same fractions get re-read from the opposite end — that’s why statistics isn’t a new subject, it’s this one map run backwards.
That reverse trip, from data back to model, is statistics. It's not a different subject with a different toolkit. It's this exact map read in the opposite direction. Probability asks "given a fair coin, what flips will I see?" Statistics asks "given these flips, is the coin fair?" Same coin, same sample space, same fractions. The arrow just points the other way.
Let's make the backwards trip concrete. Suppose a coin has some unknown bias p, the true long-run fraction of heads, hidden from you. You can't see p, but you can flip the coin. Flip it a handful of times and your estimate of the bias is shaky. Keep flipping and watch that estimate converge onto the true, hidden value.
A coin has a hidden biasp (its chance of heads) you can't see. You can't read p off the coin — but you can flip it and count.
flips 0 · heads 0 · tails 0
flip the coin — your estimate will settle
The shaded band is every bias that could plausibly have produced your tally. Each flip kills off more of them — that narrowing is inference.
Fig. 19. The coin has a bias p — its true chance of heads — that you cannot see. So you flip and tally: your estimate p̂ (heads ÷ flips, the blue line) lurches wildly at first, then settles. The blue band is every bias that could plausibly have produced what you've seen (roughly a 95% range); watch it funnel as flips pile up — fewer and fewer models survive the data. Hit Reveal and the hidden gold line drops in: your data cornered it without ever being told. That's why more data helps — each flip narrows the set of models that could have made your tally.
You just learned a model from data. You inferred the coin's hidden parameter p by watching the flips it produced. And if that phrase rings a bell, it should, because it's the whole of machine learning in one sentence. Training a model is estimating hidden parameters from observed data. It is this chapter's backwards arrow, run at enormous scale.
One hidden number, guessed from what you saw. Slide the data, then engage the lock.
Slide the data, then engage the lock →
Fig. 20. Flip a coin you don't understand and count heads: p̂ = heads ÷ n is your best guess at the wire's hidden p, and the ± band is how unsure you still are. Slide data collected and watch the guess sharpen as √n does its work — then engage the lock. The right panel was the same picture all along: a model is just parameters θ fit to data, its loss the very same uncertainty. Training GPT is this — data → parameters — with billions of θ instead of one. That's the whole trick, run enormous.
So hold the two directions together as one object: probability going forward, statistics coming back, and every data-driven field you've heard of living on that return trip: ML, experimental science, quality control, medical testing. Learn the forward direction cold, which is what this book does first. Then the backwards direction becomes a change of perspective, not a new mountain.
06You are here
Before we go, let's zoom all the way out. This course can look like a zoo of unrelated formulas, a dozen distributions and a pile of theorems. It is actually one family tree, and every branch on it grows from the single idea you built today. Here's the whole map, with our current rung lit up.
You are here — the root ratio
THE ROOTProbability is one ratio: favourable outcomes ÷ total outcomes. Everything else grows from it.
Every branch is the same ratio, counted smarter — so you never memorize a distribution, you re-derive it from this one idea.
Fig. 21. The entire course is one family tree, and its root is the only thing you truly have to hold: probability is the favourable fraction, count / total. Combinatorics is that count done cleverly; distributions are it done repeatedly; conditioning is it done inside a subset; Bayes is it done backwards; expectation is it done with values attached. Tap any branch and watch its path light up straight back to the root — that is why you needn't memorize a single distribution. Every one of them is smarter counting of this one ratio, re-derivable from the tree. Right now, you are standing on the root.
The tree just promised something bold: everything above the root is only smarter counting of the same fraction. That is easy to say and easy to doubt, so let us feel it just once. Take four coin flips. That is sixteen outcomes, two to the fourth, still small enough to write out by hand. Count the ones with exactly two heads, then watch a shortcut reach the same number without listing a thing.
Scan all 16 sequences. Tap the ones with exactly two heads (blue = a head).
tap the sequences with exactly two heads
The shortcut never touched the ratio — it just counts the top a smarter way: pick which 2 of the 4 positions hold heads, C(4,2)=6, the same 6 you found by scanning.
Fig. 22. Tap the sequences with exactly two heads — scanning all sixteen by hand lands on 6, so the probability is 6/16 = 0.375. Now hit show the shortcut: instead of reading rows, it picks which 2 of the 4 positions hold the heads — the pairs {12,13,14,23,24,34}, C(4,2)=6 — the identical six, counted without ever writing the list. Then push the flips slider: at ten flips the list is 1024 rows, hopeless to write, yet choose 2 of 10 = 45 still answers at a glance. The shortcut never moved the favourable/total ratio — it only counted the top a smarter way, which is the whole of the counting to come, previewed in one move.
We are standing on the ground rung: count the favourable fraction of equally-likely outcomes. Everything above that rung is smarter counting of that same ratio — counting engines, measures, conditioning, the Bayes flip, the whole family of distributions, the great limit theorems. If you can keep this tree in your head, you never have to memorize a distribution again. You re-derive it.
The very next rung is already forced on us by a wall we can see coming. Counting by hand was fine for two coins and two dice. But ask about 100 coin flips and the sample space has 2¹⁰⁰ outcomes, which is about 1.27 × 10³⁰ sequences. Listed at a billion a second, writing them all out would take some 40 trillion years. That is a list no one will ever write down. Watch the outcome count explode as we add flips.
write all of them by hand — then just count.
The method never breaks: P = favourable ÷ total. What breaks is building total by listing. Chapter 2 counts it — without ever writing the list.
Fig. 23. Every coin flip doubles the list of possible outcomes, so n flips give 2ⁿ of them. At n = 2 you can write all four (HH HT TH TT) and count by hand — that is the method: probability is favourable outcomes ÷ total outcomes. Hit + add a flip and watch the count blow past the people on Earth, past the seconds since the Big Bang, past the stars in the sky. At n = 100 it reaches 2¹⁰⁰ = 1,267,650,600,228,229,401,496,703,205,376 — about 1.27 × 10³⁰. Listed at a billion a second, that takes ~40 trillion years, roughly 2,900 times the age of the universe. The method isn't wrong; the list is impossible. That is exactly why Chapter 2 is combinatorics — a way to count the favourable outcomes without ever writing a single one down.
So we've hit the first honest wall of the book. The method is right, but brute enumeration doesn't scale. We need a way to count the favourable outcomes without ever listing them, and that is a counting engine. The engine is combinatorics, and building it is where we go next.