Last chapter ended with something small and unsettling. We were working out the chance that the second card off a deck is red. The answer was a flat 2 in 5, and it felt fixed. Then someone turned card one face up. It was red — and the number moved. Sit with how strange that is. The cards did not change. The deck did not change. The only thing that changed was what you know, and the probability shifted anyway. That is the crack this whole chapter climbs through. Up to now, every P(A) has been a single number that just is. But that is not how real reasoning works. A blood test comes back, a witness speaks, the first card flips — and the odds of everything downstream lurch. So here is the honest question the back half of this course is built on: what exactly happens to a probability the moment you learn something? The answer is almost embarrassingly physical. You shrink the universe by throwing away every outcome the clue just ruled out. Then you re-measure your event inside the smaller world that is left. Do those two things and you have conditional probability, P(A|B) = P(A∩B)/P(B). It comes with a quiet twin, independence, which is the special case where the clue tells you nothing and nothing moves at all. No new axiom. Just one new verb: condition.
01Shrinking the universe
A clue does one physical thing: it deletes part of the world. Let's begin exactly where the last chapter left us, because that jumping number is the whole idea in miniature. A probability like P(A) was always a statement made inside a fixed world. That world is the sample spaceΩ from Chapter 3 — every outcome that could happen, each carrying its share of the weight. Now you are told "B happened." Every outcome where B is false is instantly impossible: gone, zero weight, off the board. What a clue leaves behind is a smaller Ω, and your event A has to be re-measured in there. Watch it happen. Pick an event, then hand the machine a clue, and see the probability move — up, down, or not at all — the instant information lands.
before the clue: P(A) = 1/2
the die never changed — only what you know about it did
Fig. 1. Pick event A on a fair die, then flip the clue — "the roll is ≥ 4" — on and off. The die itself never changes; only the outcomes still standing do. Watch P(A) slide up, slide down, drop to zero, or sometimes not move at all — four different clues in disguise, one slider. That's the whole engine behind conditioning: a witness's tip or a test result doesn't touch the world, it only narrows what you know about it — and every probability downstream has to lurch to match.
That movement is the entire subject of the chapter, so let's pin the mechanism down precisely instead of by feel. Here is the picture from the metal up. Draw Ω as a rectangle of area 1, which is the total probability. Your event A is a region inside that rectangle, and so is the clue B. When you learn B, you erase everything outside B. Grey it out, because it cannot happen now. B becomes your new, smaller universe. The only part of A that survives the erasing is the sliver sitting inside B, the overlapA∩B. So the new probability of A is that overlap measured against the shrunken world: P(A|B) = P(A∩B) / P(B). Drag B around the sample space below. Watch the outside go grey, the overlap light up, and the conditional recompute live.
9/30 survivors in A → 0.30
Drag B across Ω. The outside greys out — those outcomes can't happen now. So we divide A's surviving slice by B's size, not by Ω.
Fig. 2. Conditioning is shrinking the universe. Ω holds 96 equally likely outcomes; the blue region is event A. Drag the gold rectangle B around and resize it — the moment you condition on B, every outcome outside it greys out (it simply can't happen now), so B becomes the whole world. What's left of A is the green slice A∩B, and its share of that shrunken world is exactly P(A∩B)/P(B) — count the green cells, divide by the cells inside B. That fraction is P(A|B): A's part of the surviving world, measured against how big that world now is.
Read that formula out loud in plain words, because the notation hides how simple it is: "the chance of A given B is the chance of both, divided by the chance of the clue." The numerator P(A∩B) is measured in the old, full world, so it is the overlap as a fraction of everything. But a fraction of everything is not what we want. We want A's share of the new world, which is only as big as B. So the denominator /P(B) has exactly one job, to renormalize. It scales the surviving slice back up so the shrunken universe sums to 1 again, the way any honest probability space must. Step through the rescaling and watch a slice worth 0.1 of everything become 0.4 of what remains.
Step 1 / 4
B = 10% of everything; A∩B = 4%
step through: B gets lifted out, then stretched by 1/P(B) until its own world fills back up to 1.
Fig. 3. Step through it: B is lifted out of everything and, for a moment, its own little world is still sized like the slice it was — worth only 0.10, not 1. The ×1/P(B) stretch is what fixes that, scaling the whole slice back up until it fills the width again — and the green A∩B sliver rides along, growing from 4% of the old everything into 40% of the new one. That's P(A|B). The division by P(B) was never ceremony — it's the one operation that lets a shrunken universe still add up to 1.
Figure 3 showed the stretch fixing the total, but you never watched what breaks without it. So switch the ÷P(B) step off yourself and add up the survivors. They sum to P(B) instead of 1, which means the result is not a legal probability at all. Turn the division back on and watch axiom II snap into place.
B's 4 weights add to P(B) = 0.40. Before flipping, guess: once you renormalise, what does Σ become?
not a probability: total must be 1
With the /P(B) step off, the survivors sum to 0.40 — a broken space. Flip it on: every weight × 2.5, the bar fills the ceiling, Σ snaps to 1.00.
Fig. 4. The shrunken world B holds four outcomes whose raw weights add to P(B) = 0.40. Leave renormalise OFF and look at what you actually have: the stacked bar stops well below the 1.00 ceiling, the red gap says short by 0.60, and Σ = 0.40 — that is not a probability space at all, because axiom II demands the total be 1. Guess where Σ lands, then flip renormalise ON: every weight is multiplied by 1/P(B) = 2.5, the bar grows to fill the ceiling exactly, and Σ snaps to 1.00. The single event A = o₁ jumps with it, from a raw 0.16 (its share of all of Ω) to 0.40 (its share of B) — which is exactly P(A|B) = 0.16/0.40. Dividing by P(B) isn't bookkeeping ceremony; it is the one move that keeps the zoomed-in world a legal probability space.
So conditioning is two moves welded together. First, delete the outcomes the clue killed. Then rescale what is left so it is a full probability again. That division is not bookkeeping ceremony. It is the step that makes the answer a real probability in the smaller world. With the machine defined, let's point it at something concrete. We will watch conditioning do three completely different things to the same event, depending only on which clue you feed it.
02A lot, a little, or nothing at all
Take a standard 52-card deck — the same measure we built in Chapter 3, with every card worth 1/52. Fix one event and never change it: A = "the card is a spade." With no clue at all, that is a clean 13/52 = 1/4. Now we hand the machine a clue and watch what conditioning does to that 1/4. The first clue is "the card is black." Before you look, guess. Does knowing the card is black push the chance of a spade up, down, or leave it alone? Commit, then reveal.
Event A = spade (13 of 52). Learn the card is black — does P(A) rise, fall, or hold?
P(spade) = 1/4 — guess, then reveal
Fig. 5. Event A = the card is a spade — 13 of 52, so P(A) = 1/4. Pick a guess for what happens once you learn the card is black, then hit reveal: the 26 red cards (hearts, diamonds) fall away, and every one of the 13 spades survives inside the smaller 26-card world — so P(A) doesn't move a little, it doubles to 1/2. The clue wasn't neutral: it deleted outcomes that were entirely against A (the reds) while sparing every outcome for A, and that lopsided deletion is exactly what drags the probability upward.
The chance of a spade doubles, from 1/4 up to 1/2, and the shrink-the-universe picture tells you exactly why with no algebra. "Black" throws away the 26 red cards, leaving a world of only 26 black ones. All 13 spades survive that cut, because every spade is black. So inside the new world spades are 13 out of 26 = 1/2. The clue was partly relevant, and the probability moved partway. Now take the same event and hand it the opposite clue: "the card is red." Watch the overlap this time.
♠ = spade, the event we're tracking · ♥ ♦ = red, the clue
P(spade) = 13/52 = 1/4
Spades are always black — the red world contains none of them, so the overlap A∩B is empty before you even divide.
Fig. 6. Same 52-card grid as before, same event A = spade (gold outline, top row) — only the clue changed. Click apply ‘red’: the two red rows become the surviving world (26 cards), both black rows fade out, and the spade row fades with them — it's entirely black, so zero of it survives into the red world. The overlap A∩B is the empty set ∅, so P(spade|red) = 0/26 = 0. That's the flip side of Fig 5: a clue doesn't just reweigh a probability, it can erase an event completely if the clue and the event never coexist.
This time the chance of a spade annihilates: P(spade|red) = 0. Shrink to the red world and there is not a single spade left in it. The overlap A∩B is empty. So a clue can drive a probability all the way to zero, if it rules the event out entirely. Conditioning has now moved our 1/4 up to 1/2 and down to 0. Here is the third case, and it is the one that seeds everything after. Hand the machine a clue that has nothing to do with suits at all: "the card is a queen."
spade share: 13/52 = 1/4
The clue named a rank (queen) — it said nothing about suit. That's why the spade share can't move.
Fig. 7. Click queen clue: the deck's 48 non-queens fade out, a gold box seals off the Q column, and the world shrinks to just 4 cards — one queen per suit, the spade one still glowing green. Count inside that box: 1 spade queen out of 4 queens, so P(♠ | Q) = 1/4 — the exact same value as the baseline P(♠) = 13/52 = 1/4 among all 52. The big fraction on the right never even twitches. That's the phenomenon: a clue about rank carries zero information about suit, so conditioning on it leaves the spade probability standing exactly still.
Nothing happens. P(spade|queen) = 1/4, dead equal to the unconditioned answer. Shrink to the 4 queens and exactly 1 of them is a spade, which is 1/4 — the same fraction as the whole deck. The clue was informative about rank but silent about suit, so it left P(spade) exactly where it found it. That "the clue changed nothing" case is not a boring edge. It is a named, load-bearing concept, and it is what the middle of this chapter is about.
Three times now you have watched the deck renormalise, but every card weighed the same 1/52. Real clues land on lopsided tables where the cells are not equal. So here is a table with uneven counts, and this time you drive it. Pick the row to condition on, then read P(A|B) straight off as the cell over the row total.
click a row or column header to condition
set your guess, then click the Rain row to reveal.
condition on a row (÷ its total) — or a column
Fig. 8. Your turn — the counts are lopsided, so eyeballing won't save you. Drag the slider to commit a guess for P(Bus | Rain), then click the Rain row: the other row greys out and each surviving count divides by that row's total of 80, so 0.275 + 0.575 + 0.150 = 1.000 and the readout writes it exactly — P(Bus | Rain) = 46 / 80 = 0.575. Now click the Bus column: the same cell 46 stays on top, but the denominator switches to the column's 74, giving P(Rain | Bus) = 46 / 74 = 0.622. Conditioning was never a card trick needing a tidy deck: pick your band, divide every cell by that band's total — the denominator, not the cell, decides which question you asked.
03Independence: the clue that changes nothing
When learning B leaves the probability of A untouched — P(A|B) = P(A) — we say A and B are independent. That is the whole definition: the clue tells you nothing about the event. But "queen and spade don't affect each other" can feel like a coincidence of this particular deck. It is not a coincidence. There is a clean geometric reason, and once you see it you will recognise independence on sight. Lay the deck out as a 13×4 grid, with 13 ranks down and 4 suits across. Now "spade" is one column, a vertical stripe taking 1/4 of the width. And "queen" is one row, a horizontal stripe taking 1/13 of the height. Watch what conditioning on the queen-row does to the spade-column's share.
a full deck: 13 rows × 4 columns
Grab Q row: restricting to the queen row leaves the spade share at 1/4 — that's independence.
Fig. 9. A deck is a 13×4 grid: pick a suit and you pick a column (spades are 1 of 4, so 1/4); pick a rank and you pick a row (queens are 1 of 13). Slide to the Q row and the spade share stays 1/4 — the row can't distort the column. That's what independence means: the attributes are perpendicular, so the joint cell factorizes, 1/52 = (1/13)(1/4) — the very same move as separation of variables in a PDE.
There it is: restricting to the row does not change the column's fraction. A column takes up 1/4 of a single row exactly as it takes up 1/4 of the whole grid. That is what independence is — perpendicular attributes. Suit runs one direction, rank runs the other, and slicing along one axis cannot distort proportions along the other. And the intersection is one specific cell, the queen of spades, with probability 1/52. That is exactly (1/13)·(1/4), the row-fraction times the column-fraction. That little multiplication is not an accident. It is a rule we can derive, right now, from the definition we already have.
Start from conditioning, P(A|B) = P(A∩B)/P(B), and feed in the one thing independence asserts, P(A|B) = P(A). Set the two right-hand sides equal and clear the denominator. Step through it — three honest lines — and watch the multiplication rule drop out.
always true — no assumption yet
That's why "multiply the probabilities" is a privilege you earn by proving independence — never a default you may assume.
Fig. 10. Click next and watch the derivation write itself. Line 1 is just the definition of conditional probability — true for any A and B, no strings attached. Line 2 is the one move that costs something: it swaps in P(A|B) = P(A), the actual meaning of "A and B are independent" (the second bar fades in at exactly the same length, because knowing B happened told you nothing new about A). Only after that swap does line 3 — the boxed P(A∩B) = P(A)·P(B) — fall out as pure algebra. Step back and the box vanishes with it: no independence, no shortcut.
You just derived the rule, and Figure 9 showed it holding on a perfect rectangle. But you have never watched that same rule fail. So take a table and bend it. Hold the margins fixed and drag one cell. The moment P(A∩B) parts from P(A)·P(B), independence is gone.
Predict first: the margins are pinned — can you still break independence?
pinned margins: P(A)=0.50P(B)=0.40
independent — cells = row × col
At coupling 0 the A/¬A split is flat across both columns — a perfect product rectangle, so P(A∩B) lands exactly on P(A)·P(B) = 0.20. Drag up and watch the answer.
Fig. 11. The deck was independent because its grid was a perfect rectangle; here you get to bend one. Drag coupling and only the inside of the 2×2 joint moves — the margins stay pinned at P(A)=0.50 and P(B)=0.40 the entire time. At coupling 0 the A/¬A split line is flat across both columns and the face-off reads P(A∩B)=0.20 = P(A)·P(B)=0.20 — the chip is green, independent. Push it up and P(A∩B) climbs to 0.40 while the product stubbornly stays 0.20; the split lines pull apart, the ¬A∩B cell drains to 0.00, and the chip flips red. The aha: independence lives in the cells, not the margins — keep every margin fixed, nudge a single cell, and one broken "cell = row-share × column-share" makes the whole table dependent.
So P(A∩B) = P(A)·P(B) holds when — and only when — A and B are independent. This is worth flagging loudly, because people memorise it as if it were handed down on a tablet. It is not a fourth axiom. It is a consequence of the definition of conditioning plus the claim that the clue is irrelevant. "Multiply the probabilities" is a privilege you earn by establishing independence, never a default. (Have you met separation of variables in differential equations? A solution splits as f(x,y) = X(x)·Y(y) exactly when the two directions do not couple. This is the same move in a probabilist's coat: independence is the joint factorizing.) And now the most dangerous confusion in the whole topic, one that trips up nearly everyone. Two events that can't both happen feel as though they must be independent. They are the exact opposite. Watch.
slide A between overlapping B and sitting entirely outside it
independent — P(A|B) = P(A)
Disjoint isn't "unrelated" — it's the MOST dependent case: once B happens, A becomes flat-out impossible. That's why the two words are opposites, not synonyms.
Fig. 12. 16 equally likely outcomes, each with probability 1/16 — this is the whole sample space. The dashed green box is event A, the shaded blue box is event B; a dot's colour shows which event(s) it belongs to, and gold means it's in A∩B (both). P(A|B) — "the probability of A, given B already happened" — means: throw away every outcome outside B, and ask what fraction of what's left is still in A. Hit overlap: A and B share 4 outcomes out of B's 8, so P(A|B) = 0.50, exactly P(A) — learning B told you nothing about A, the inform-nothing definition of independence. Hit disjoint: A slides to the top rows, A and B now share zero outcomes, so P(A|B) crashes to 0 — the overlap-nowhere definition of mutual exclusivity. Same P(A) = 0.5 both times; the only thing that moved is how much B tells you about A. That's the whole point: independence and mutual exclusivity aren't shades of the same idea, they're opposite ends of it — mutual exclusivity is about as dependent as two events can get.
If A and B are mutually exclusive — disjoint, so they never co-occur — then learning that B happened tells you A definitely didn't: P(A|B) = 0. That is about as far from P(A|B) = P(A) as you can get. Disjoint events are maximally dependent, not independent, because knowing one nails down the other. The two words sound cousinly and mean opposite things. Mutually exclusive is about outcomes that overlap nowhere. Independent is about clues that inform nothing. Keep those apart and half the classic mistakes never happen. Now let's spend independence on something it was born to do: building things that do not break.
04Building reliable things
Independence multiplies, and that single fact is the entire engineering theory of reliability. Picture a system built as a series chain: n components in a row, where the system works only if every one of them works. Think of a string of old holiday lights, where one dead bulb kills the whole strand. If each part works independently with probability p, then "all work" is a chain of ands, so P(system works) = pⁿ. Slide n and p and watch how brutally that product falls. Even excellent parts, chained long enough, make a terrible whole.
6 good links, still 53% overall
Click a dead link to re-roll it — independent means one part's fate never touches another's, yet the chain still needs every single one to survive.
Fig. 13. Drag n and p — every part rolls its own independent coin, but the chain only works if all of them do, so P(works) = pn falls fast even when each part is quite good. Click any link to re-roll it and watch one dead part snap the whole chain.
Ten parts at 99% each give a system at only 0.99¹⁰ ≈ 90%. You lost nine points of reliability just by having a long chain. Series is fragile by construction. The weakest link is not a metaphor here, it is the maths. So how do you build something robust? You flip the wiring to parallel: n redundant copies, where the system fails only if all of them fail at once. And look at the shape of that question. "At least one still works" is the complement move from Chapter 4, back on stage. P(fails) = (1−p)ⁿ, so P(system works) = 1 − (1−p)ⁿ. Slide the knobs and watch redundancy climb toward certainty.
triple-nines: parts barely matter now
3 parts at 90% each must ALL die together to kill the system — 1-in-1000, not 1-in-10. Parallel wiring turns mediocre parts into three-nines.
Fig. 14. Drag n and p: each of the n copies works on its own with probability p, wired side by side between IN and the big system node. That node only goes dark if every copy fails at once — Ch 4’s P(≥1) = 1 − P(none) again, just re-dressed as P(works) = 1 − (1−p)ⁿ. Hit re-roll and watch three 90%-parts almost never die together: you’ve wired the failures so they must all agree before the system does.
Three parallel copies at a shaky 90% each give a system at 1 − 0.1³ = 99.9%. You bought three-nines reliability out of mediocre parts, just by wiring them so their failures have to agree. This is why spacecraft carry triple-redundant computers and your data lives on three disks. But notice the load-bearing word I slipped past twice: those failures have to be independent. Every number on this screen assumed exactly that. And that assumption is where careful engineers get killed.
05The trap: independence is a physical claim
Here is the sentence to tattoo somewhere. Independence is a claim about the physical world, not a mathematical convenience you may assume when it's tidy. When we wrote P(all fail) = (1−p)ⁿ for the parallel system, we quietly asserted that the three drives fail independently — that one dying tells you nothing about the other two. But suppose all three share a power supply, or sit in one rack, or ride the same building's wiring. Now a single lightning strike takes out all three at once. The failures are not independent anymore. They have a common cause. Toggle that shared cause on below and watch your beautiful three-nines redundancy collapse.
3 independent backups · 99.9% reliable
Predict first: three backups, each 90% reliable, wired in
parallel — the system only dies if all three die. Multiply and you get a
glorious 99.9%. Now flip on the one wire the brochure forgot.
The maths was never wrong. The independence it assumed was.
Fig. 15. Three backups in parallel, each 90 % reliable. Treat their failures as
independent and you multiply: the system only dies if all three die at once, so
0.1 × 0.1 × 0.1 = 0.1 % failure — a dazzling 99.9 %. But wire in one
shared cause — a single power feed a lightning surge can take out — and that failure mode is
never cubed. It passes straight through at its full 10 %, dragging real failure to ~10 % and reliability
back to a single component’s 90 %. The three backups bought you almost nothing. The maths was flawless; the
independence assumption feeding it was the lie — and one hidden wire overstated your safety by ~100×.
With one shared cause switched on, three redundant drives stop behaving like three and start behaving like one, because when the surge hits they all go together. The 99.9% you thought you had bought evaporates back toward the reliability of a single component. That is not a rounding error. That is overstating your safety by orders of magnitude, and it is the mechanism behind a startling number of real catastrophes: correlated mortgage defaults in 2008, redundant systems felled by one flood, "independent" sensors sharing one bad calibration. The maths of multiplication was never wrong. The assumption feeding it was. So before you ever multiply, you owe yourself one question: is there a wire — physical, causal, or hidden — connecting these events? When there genuinely is not, independence is a gift. Let's watch it hold in the cleanest possible case, and prove it by simulation.
A fair coin has no memory and no wires to anything. Each flip is independent and identically distributed, or IID — the atom of the whole rest of this course. So the chance of all heads in N flips is just (1/2)ᴺ, one clean factor per flip. Let's not take that on faith. Let's run it, and also break it by slipping in a shared bias, so you can see the product hold when the flips are independent and lie when they are not.
coins: independent, or wired?
IID = independent, identically-distributed. p is a coin's chance of heads — "shared bias" makes one hidden p drive every flip in the run.
hit run — independent, N=4
one wire between flips is all it takes to break the product.
Fig. 16. Each run flips N = 4 coins and checks whether all four landed heads. In independent mode every flip is a fresh, unrelated p = 0.5 coin — run it 20,000 times and the empirical fraction settles right on the gold line, (½)⁴ = 0.0625, exactly what the product rule predicts. Now flip to shared bias: one hidden draw — p = 0.9 or p = 0.1, a coin-flip's chance each — sets the SAME p for all 4 flips in that run. Any single flip is still fair on average (0.5×0.9 + 0.5×0.1 = 0.5), yet the run's empirical fraction rockets to ≈0.33, over five times the formula's claim. That's why IID — independent, identically distributed — is the clean atom of this course: with no wire between flips the product is exact; wire one flip's outcome to the rest and it fails, just like the drives.
With independent flips, the simulated fraction of "all-heads" runs lands right on (1/2)ᴺ, and multiplication tells the truth. Couple the flips with a shared hidden bias and the empirical number drifts off the product — the same failure the drives showed, now on coins. That brings us to where conditioning matters most in ordinary life: reading a test result. Getting these ideas exactly right is the difference between a sensible decision and a panic.
06The language of tests
Conditioning hands us the precise vocabulary for any diagnostic test, and the vocabulary is where people — including, occasionally, textbooks — get sloppy. A test for some condition has two quality numbers. They are separate conditional probabilities, and both point in the same direction: from the truth to the result. The first is sensitivity = P(test positive | you have it). Of the people who genuinely have the condition, what fraction does the test correctly flag? The second is specificity = P(test negative | you don't have it). Of the healthy people, what fraction does the test correctly clear? Drive the two knobs independently and watch the four boxes of true and false positives and negatives fill in.
sens 90% · spec 85% · acc 87%
Two knobs on the SAME truth column, not one. Hit trap: sensitivity hits 100% — it "catches" every sick person — only because it flags everyone. Specificity craters to 0%, so most of the false-alarm block is the healthy majority.
Fig. 17. The population splits by truth first (columns, fixed: 30 have the condition, 70 don't) and then by test result (rows) — and that second split is two separate conditionals: sensitivity = P(test + | has condition) governs the left column only, specificity = P(test − | no condition) governs the right column only. Drag them — each moves only its own column. Hit trap: always +: sensitivity rockets to 100% while specificity collapses to 0%, and the "false alarm" block swallows the entire healthy majority. A single "accuracy" number can't show you that — which is exactly why the two are always quoted as a pair.
Notice that these are two knobs, not one. You cannot collapse them into a single number called "accuracy." A test can be 100% sensitive and still be useless: just return "positive" for everyone, and it catches every sick person while flagging every healthy one too. So sensitivity and specificity have to be quoted as a pair. Hold that pair firmly, because now comes the confusion that the whole next chapter exists to destroy. It is a swap so natural that it has sent innocent people to prison. Both of the numbers above condition on the truth. But the thing you actually want, once your result comes back positive, points the other way.
P(+|sick)=90% ≠ P(sick|+)=48%
⚖ same trap, the courtroom
A DNA "match": P(matches | innocent) ≈ 1-in-a-million. Swap it for P(innocent | matches) — which depends on how many suspects got tested, and can be a thousand times bigger.
Fig. 18. The same 1,000 people, split two ways. Forward: start from the 50 sick people (coral) — a fixed 90% of them test positive, and that number never moves. Backward: start from whoever tests positive — drag false-positive rate and watch the positive crowd fill up with well people (blue) who were wrongly flagged, dragging P(sick | +) down while P(+ | sick) sits still. Same symbols, reversed order, two different numbers — that's the whole trap: P(+|has) reads "positive, given sick"; P(has|+) reads "sick, given positive," and nothing forces them to agree.
Figure 18 turned the arrow from result back to truth, but there is a stranger reversal left. Card one is dealt, then card two, both face down. Now flip card two and predict what happens to the odds on card one. Your gut says nothing changes, because card one already happened. Your gut is wrong.
A tiny deck: 2♠ · 2♥. Card 1 is dealt first and stays face down — it is physically fixed. Predict what flipping the later card does to P(card 1 = ♠):
Now flip card 2 — try both suits:
P(card 1 = ♠) = 1/2 — predict, then flip
the card under your finger never changed — only what you know did
Fig. 19. A four-card deck — 2♠ and 2♥ — both cards face down, so P(card 1 = ♠) = 1/2. Card 1 was dealt first; commit to a prediction (most readers pick "it already happened, so it stays"), then flip the later card. Reveal card 2 = ♠ and card 1's spade-odds fall to 1/3; reveal card 2 = ♥ instead and they rise to 2/3. The card under your finger never turned over — it is physically fixed, the past does not change. What changed is your knowledge: conditioning updates belief in both directions of time, so a later clue can rewrite an earlier card's probability. That is the lede's thesis at its sharpest, and the first hint of Bayes reasoning from effect back to cause.
Sit with the gap. P(positive | you have it) is sensitivity, a property of the test. It is emphatically notP(you have it | positive), which is the thing you actually care about and a fact about you. The test is not the disease. Swapping those two conditionals is so common that it has a name in two separate fields. In law it is the prosecutor's fallacy — confusing "the chance an innocent person's DNA would match" with "the chance this matching person is innocent." In medicine it is the base-rate fallacy. The two conditionals feel identical and can differ by a factor of a thousand. The machine that correctly flips P(result|truth) into P(truth|result) is Bayes' theorem, and it is the entire next chapter. For now, just hold the discomfort of knowing the two are different. That discomfort is the hook.
07You are here
Before we flip anything, here is one honest accounting of what we added. Conditioning composes. You can apply a clue, then another, then another, each one measured in the world its predecessors left behind. That is the chain rule: P(A∩B∩C) = P(A)·P(B|A)·P(C|A,B). Peel it one factor at a time and watch each conditional live in a smaller universe than the last.
60 suspects, nothing known yet
each clue's odds are measured only in the smaller world the clues before it left behind.
Fig. 20. Sixty suspects, three clues. Click +A and the pool halves — only 30 match the first clue, so P(A)=0.50. Click +B and the second clue isn't judged against all 60 anymore — it's judged only inside the 30 who already had A, so P(B|A)=18/30=0.60, the smaller box shrinking around exactly that group. Click +C and the third clue narrows again, this time inside the 18-suspect world A and B already carved out — P(C|A,B)=9/18=0.50. Chain the three fractions straight through — 0.50 × 0.60 × 0.50 — and you land on exactly the 9 suspects who match all three. The chain rule isn't a new law; it's conditioning applied again and again, each factor measured in the world its predecessors left behind.
And the reason you are allowed to keep conditioning is quietly profound: a conditional probability is itself a full probability. Once you have shrunk to B, the function P(·|B) obeys all three axioms we built in Chapter 3. It is non-negative, the whole shrunken space has probability 1, and disjoint events still add. So it is not a weakened cousin of probability. It is a genuine probability measure on a smaller stage, and that is what lets the machine recurse.
click each axiom to verify
That's why conditioning composes: P(·|B) is a real probability on the shrunken world — so you can condition again, and again, each time obeying the same three rules.
Fig. 21. Ω holds 20 equally-likely outcomes; B (dashed blue) keeps 8 of them — the world you're left with once you condition. Click each axiom: ① a sub-event of B (cyan) still measures out non-negative; ② B measured against itself fills the whole shrunken space, P(B|B) = 1; ③ the two disjoint pieces A1 and A2 that make up B still add — .375 + .625 = 1. All three of Chapter 3's rules hold, unchanged, on the smaller stage.
So step back to the family tree and see the true size of what happened. We did not add a new axiom or a new distribution. We took the measure from Chapter 3 and learned a new verb to perform on it — condition: shrink the universe to what you know, then re-measure inside it. Everything else fell out of that one verb: independence, the multiplication rule, reliability, sensitivity and specificity. Find our rung, lit gold.
Hover — or tap — any rung, including today's, for its one-line role.
hover any rung — trace the spine up
Fig. 22. The whole spine so far, today lit gold: ch1 gave you the ratio, ch2 the machine that counts it, ch3 promoted both into a measure, ch4 taught that measure to handle "at least one." This chapter added no sixth axiom — it taught the measure one new verb, condition: shrink the universe to what you know, re-measure inside it. Hover any rung — the cyan arc from ch3 is the tell: conditioning reaches straight back to the measure itself, not to whatever came right before it.
Here is the seam into the next chapter, and it is the most beautiful sleight of hand in the subject. Everything we did ran in one direction: from a cause to an effect, from the truth to the result. That is the direction that is easy to measure. A lab can tell you P(positive | sick) by testing a thousand sick people. But you never get to stand in the world of "sick." You stand in the world of "I got a positive" and you want P(sick | positive) — effect back to cause. Watch what happens when we lay the two conditionals side by side, because they share a single term.
two worlds, two different fractions
Both readouts have the same numerator, P(A∩B) — only the denominator, which world you're standing in, differs.
Fig. 23. Two worlds, same two facts. Left: inside world A (the cause happened), a slice B is also true 75% of the time — P(B|A) = 0.75, easy to measure directly. Right: inside world B (the effect happened), a slice A is also true 60% of the time — P(A|B) = 0.60, the one you actually want. Click shared term: both slices glow, because they're the same group of outcomes, P(A∩B) = 0.15, just measured against a different-sized world. Click flip → Bayes: set the two products equal, divide by P(B), and the known arrow A→B turns into the wanted arrow B→A — that flip is Bayes' theorem, not a new law.
Both conditionals are built on the same overlap, P(A∩B) — the chance of both the truth and the result together. That overlap does not care which order you name the two events in: P(A∩B) = P(B∩A). So P(A|B)·P(B) and P(B|A)·P(A) are two expressions for the identical quantity. Set them equal, divide, and you have turned the arrow around. You have flipped a probability you cannot measure into one you can. That flip is Bayes' theorem. It is not a new law, just this one shared term solved two ways, and it is where we go next. Turn the page.