Quantitative Finance — the Mathematics of Markets · chapter 10
10Conditioning, Independence & Bayes
Chapter 9 built probability out of three plain rules about weight. One unit of it, spread across a sample space, and a function P that weighs whatever region you circle. Every number on that page was measured against the whole of Ω. This chapter is about what happens when a fact arrives. Someone tells you the roll came out even, and half the sample space is suddenly dead. Nothing about the die changed and no force reached across the table. Your world got smaller, and the weights left standing no longer add to one. We are going to take that single move apart: shrink the world, then rescale it until it is whole again, and watch most of probability fall out of it. The multiplication rule is that move read backwards. Independence is that move doing nothing. The law of total probability is that move run over every branch at once, and Bayes' theorem is that move run in reverse. By the end you will not be memorising Bayes. You will be able to rebuild it in three lines on a napkin, and you will be able to look at a test that is right 99% of the time and tell it, correctly, that it is about 9% right.
Before we roll anything, here is the seam we are about to cross. Chapter 9 gave every event a weight, and every one of those weights was measured against the whole sample space. That is the quiet assumption this chapter breaks. Tap a downstream node in the panel below and watch how much of this course is waiting on the repair.
The course spine — and the exact seam where Ω (the whole sample space) stops being the right yardstick, and B (the fact you just learned) takes over.
Tap a blue chapter on the spine — a gold road traces back to this page and names the exact debt it owes conditioning.
what Ch 9 handed us — then learn one fact:
six weights sum to 1.000 — world whole
What you're looking at — the map so far, and the one repair Ch 10 owes the rest of the course.
Ch 10 · you are here — every gold road downstream traces back to this page
a downstream tool that can't work until conditioning lands (tap it)
weights that still sum to 1 — measured against Ω, the whole world
learn ≤ 3 and only 0.500 survives — they stop summing to 1; that gap is what P(A|B) divides back out
Fig. 1. The seam between two chapters. Chapter 9 handed you a fair die whose six weights each read 1/6 and sum to 1 — every event measured against Ω, the whole sample space. Now learn one fact — the roll came out ≤ 3 — and three outcomes vanish: the survivors add to just 0.500. That refusal to sum to one is the whole problem, and P(A | B) = P(A and B) / P(B) is nothing but the fix — shrink the world to where B held, then re-measure inside it. Tap any blue chapter to see how much of the course is waiting on that one repair.
Three of those roads deserve naming now. Chapter 24 scores a classifier with precision and recall, and both of them are conditional probabilities read straight off a table. Chapter 19 argues about priors, which is a fight you cannot follow until you know what a prior is. And Chapter 33's Kalman filter is this page applied over and over, once per tick of a clock.
01The world shrinks
There is a die on the table, face down, and neither of us has looked at it. You want the probability that it shows a six, and Chapter 9 answers that in one line. One unit of weight, spread evenly over six faces, so P(6) = 1/6.
Now a friend who can see the die says four words. It came up even. She adds nothing else and she does not touch the die, so the physical situation on the table is identical to what it was a second ago. What is the probability of a six now?
Fig 2 · The world shrinks — a friend says four words and never touches the die
▸ pick a clue — watch outcomes go dark
asking about face:
WORLD = {1,2,3,4,5,6}
favourable = 1
P(6 | no clue) = 0.1667
full world — each face is 1 in 6
Conditioning edits your world, not the world — the die never moved.
the 1/6 weights — physical; a clue never touches them
the face you ask about (the favourable outcome)
the other outcomes still in play — the new, smaller world
💡 three outcomes just stopped being live → P(6) climbs 1/6 → 1/3 with no force across the table
Fig. 2. A fair die lies in the open. A friend says four words — “it’s even” — and never touches it. Watch what happens: the ruled-out faces go dark, their 1/6 weights struck through but unchanged, and the still-live ribbon collapses from six boxes to three. Nothing about the die moved; only the set of outcomes still in play did — so P(6) climbs from 1/6 to 1/3 with no force crossing the table. Change the target face, or the clue: with “prime” your 6 is ruled out entirely and its probability drops to 0. That is all conditioning is — it shrinks the world to where the evidence held, then re-measures inside it.
The answer is 1/3, and the reason is worth going slowly over. Nothing physical happened. The die did not move, the weights on its faces did not change, and no force crossed the table. What changed is the set of outcomes still in play, because three of them are now known to be false. Three faces remain, and the six is one of three rather than one of six.
Here is how I have come to think about conditioning, and it is the sentence the whole chapter hangs on. Learning that B happened does not change the world. It shrinks yours to B, and then you ask your original question again inside what is left. We write that as P(6 | even).
That vertical bar is where a lot of readers quietly go wrong, so let's deal with it before it does any damage. It is not a division sign, and it is not shorthand for "and". It is the single English word given, and the thing to its right is the new world you are standing in. P(A | B) is read aloud as "the probability of A, given B."
If that sounds like a fuss over nothing, the two wrong readings give genuinely wrong numbers, and I would rather you watch them fail than take my word for it. Take A as "the roll is bigger than 3", which is {4, 5, 6}, and keep B as "the roll is even", which is {2, 4, 6}. Commit to an answer for P(A | B) before you touch the panel.
A = 'bigger than 3' = {4,5,6} · B = 'even' = {2,4,6} · a fair die. Which reading is really P(A|B)?
Tap a card: which one is P(A|B)?
one fair die judges all three readings
The bar is one word: given. It shrinks the world to B, then re-measures A inside it.
What you're looking at — the bar is not an operator; it is the word 'given'
the two misreadings — '÷' and 'and'
B shrinks the world to {2,4,6}
4 and 6 survive — the honest recount
0.667 — confirmed by the tally
Conditioning shrinks the world to where B held, then re-measures A inside it. On the very same fair die, 'A÷B' returns 1.000 and 'A and B' returns 0.333 — both wrong. Only the recount in the even world, 2 of {2,4,6}, gives 0.667. That is what the vertical bar means.
Fig. 3. 3 — Three readings of P(A|B) on one fair die. Read the bar as '÷' and it returns 1.000; read it as 'and' and it returns 0.333 — both plausible, both wrong. Only the recount inside the even world {2,4,6}, where 4 and 6 beat 3, returns 0.667. The vertical bar is not an operator; it is the single word given.
Reading the bar as a division of the two probabilities returns 1.000, which is a certainty, and that is absurd for an event a two can defeat. Reading it as "A and B" returns 0.333. The honest recount inside the even world returns 0.667, because two of the three surviving faces are bigger than three. One reading of a symbol, three different answers, and only one of them is what actually happens.
02★★ Why there is a division
So we have a procedure that works: delete the outcomes B rules out, then recount inside what survives. It worked only because the die was fair, which is what made counting legal in the first place. Now for the general formula. The honest way to find it is to ask what that recount was really doing.
Look at the even world again, but this time leave the old weights written on the faces. The two carries 1/6, the four carries 1/6, the six carries 1/6. Before you read the next line, answer one question in your head. These three faces are the entire world now, so what must their probabilities add up to?
One. A whole world sums to one, which is Kolmogorov's second axiom, and Chapter 9 squeezed half its results out of it. Now actually add the three surviving weights: 1/6 + 1/6 + 1/6 = 1/2. That is not one. The shrunk world, exactly as it stands, is illegal.
The renormalizer — restrict to the evens, watch the world break, then divide it whole again
walk the 3 steps — watch the total break, then heal
② the 3 survivors still read 0.1667 — they add to?
step ① — restrict the world to the evens
What you're looking at — conditioning is re-normalisation, felt by hand
gold = the value to watch: P(6), and the divisor /P(B) that repairs the world
green total = a legal world (sums to 1.000)
red total = 0.500, not yet a world
Restrict to B = even and the survivors keep their old 1/6 weights, so they add to only 0.500 — B is not yet a legal world. Dividing every survivor by that total is the one repair, and the total is P(B). That is all P(A|B) = P(A∩B)/P(B) says: the denominator is the price of promoting B to the new world.
Fig. 4. Six fair faces each carry weight 1/6 and the world total reads 1.000. Hit learn: it's even and the odd faces fall away — but the three survivors still show 0.1667, so the total collapses to 0.500 and glows red: a world must sum to 1. Predict what the survivors add to, commit, and watch the numbers refuse you (0.1667×3 = 0.5000). Then press renormalize ÷ 0.500: every survivor is divided by that total, each climbs to 0.3333, the counter heals to 1.000, and P(6) visibly grows from 1/6 to 1/3. The formula writes itself underneath — the division you just did by hand.
There is exactly one way to repair it, and it is the move the panel makes when you press renormalize. Divide every surviving weight by their total. That total is 0.500, and 0.500 is P(B), the weight of the even faces measured back in the old world. Each face goes from 1/6 to (1/6)/(1/2) = 1/3, and the counter turns green at 1.000.
Now write down what just happened, in general. The weight A had inside B was P(A ∩ B), since only the part of A that survived the shrink counts for anything. Every surviving weight got divided by P(B). So the probability of A in the new world is
P(A | B) = P(A ∩ B) / P(B)
And that is the answer to the question almost nobody asks out loud. Why is there a division in there at all? Because B has to become a whole world, and dividing by P(B) is the one and only number that makes it one. The denominator is not a rule somebody decided on. It is the price of promoting B to the new world, and you just paid it by hand.
Check the die against it. P(6 ∩ even) is the weight of {6}, which is 1/6, and P(even) is 1/2. The formula returns (1/6)/(1/2) = 1/3, which is what we got by counting faces. Good, but counting faces was only ever legal because the die was fair. The formula is not so fussy, and that difference matters more than it looks.
Load the die — conditioning is re-weighing, not re-counting
six = 0.167 · the other five share 0.833 evenly (0.167 each)
count 0.3333 · weigh 0.3333
What you're looking at — one even-world, measured two ways as you load the six
COUNT freezes at 0.3333 — "three even faces, the six is one of them." It never reads the weights.
RE-WEIGH divides each survivor by P(even), the new smaller world — and they still sum to 1.
Watch P(6 | even): hit LOADED and it climbs to 0.5556 while count still insists 0.3333. The gap is the whole lesson — conditioning re-weighs, it never needed fair faces.
Fig. 5. Load the die — conditioning is re-weighing, not re-counting.
Load the die so the six carries 1/3 of the sand and the other five faces share the rest evenly at 2/15 apiece. Now the even world weighs 2/15 + 2/15 + 1/3 = 0.600. Renormalize by that and the six comes out at (1/3)/(3/5) = 5/9 ≈ 0.5556, while the two and the four settle at 2/9 each. The three still add to exactly one.
Naive counting would have insisted on 1/3, because there are still three even faces and the six is still one of them. It is wrong by a wide margin. Conditioning is not re-counting, it is re-weighing, and the two only agree when every weight happens to be equal.
We just claimed to have made B into a new world. That is a strong claim. Is the thing we built actually a probability, in the full Chapter 9 sense of the word? Everything later leans on the answer.
Re-audit the three axioms — this time inside the world where B happened
The audit passes: P(·|B) obeys all three of Chapter 9’s axioms — so it is a probability in its own right, and every result built on them (like the complement rule) transfers for free.
the gold-ringed faces are B — the shrunk world; each shows its re-normalised weight wᵢ / P(B).
A₁ (blue) and A₂ (violet) are events you paint inside B; overlap them and lamp ③ goes red with the overshoot printed.
a green lamp = that axiom holds for P(·|B). AHA: three green → nothing from Ch 9 has to be re-proven inside B.
Fig. 6. Point Chapter 9’s axiom rack at the conditioned world. Pick the evidence B and the world shrinks to its faces, each re-weighted by dividing through P(B). The three lamps re-check Kolmogorov’s axioms for P(·|B): it is never negative, it still sums to 1, and disjoint events still add — paint two overlapping events and lamp ③ turns red with the overshoot. Load the die so it can’t roll 1, then condition on {1}: P(B)=0 and the panel goes dark, because you cannot condition on the impossible. All three lamps stay green — which is exactly why the complement rule, and the rest of Chapter 9, works inside B with no new proof.
It is, and you can check all three lamps yourself. Every conditional probability is a non-negative number divided by a positive one, so Axiom 1 holds. P(Ω | B) = P(Ω ∩ B)/P(B) = P(B)/P(B) = 1, so Axiom 2 holds and the new world really is whole. Disjoint events still add, because their intersections with B are disjoint too, so Axiom 3 holds.
That is a bigger deal than it sounds. It means nothing from Chapter 9 needs rebuilding inside the shrunk world. The complement rule still works there, so P(Aᶜ | B) = 1 − P(A | B) without any further argument. Conditioning does not take you somewhere stranger. It takes you to a smaller copy of the same place.
03Read the same rule backwards
The definition we derived has a second life, and getting to it costs one line of algebra. Multiply both sides by P(B) and the division disappears:
P(A ∩ B) = P(B) · P(A | B)
Read that in words and it stops being algebra. To land in both A and B, first land in B, which happens with probability P(B). Then land in A given that you are already standing inside B, which happens with probability P(A | B). Multiply the two and you have walked a path. This is called the multiplication rule, and it is conditioning read as a journey rather than as a measurement.
The natural picture for a journey is a tree, and there is a wrinkle in it worth meeting head on. The same intersection can be reached in two different orders, so P(A ∩ B) also equals P(A) · P(B | A). Both are correct, and beginners freeze at exactly that point, wondering which one is the real one.
Two mirror trees, one leaf — the multiplication rule read in either order
New info arrives — what happens to the probability? Pick the event, then read either tree.
both red: 3/5 × 2/4 = 0.300
tip: tap a leaf to read its outcome
What you're looking at — the same "both red" answer, reached in two orders
a red ball — 3 sit in the bag
a blue ball — 2 sit in the bag
the path you walk, and the product it prints
Left asks the 1st draw first, right asks the 2nd first — same shrink 3R2B→2R2B, same 0.300. So P(A∩B)=P(A)P(B|A)=P(B)P(A|B): use whichever conditional you actually know.
Fig. 7. Landing in "both balls red" is a two‑step walk, and you may take the steps in either order. The left tree conditions on the first draw — 3/5 that it is red, then the bag has shrunk to 2 red / 2 blue so the second is red with probability 2/4. The right tree conditions on the second draw first — also 3/5 (tap why? to see all 20 ordered pairs, the 12 with a red second ball lit green) — then the first is red given that, with probability 1/2. Both routes print 0.300. Flip the target to one of each and each tree lights two routes that must be added, not just multiplied, to reach 0.600.
Take a bag holding 3 red balls and 2 blue, and draw two without putting the first one back. Down the left tree, the first draw is red with probability 3/5, and then the bag holds 2 red out of 4, so the second is red with probability 1/2. Multiply along the path and both-red is 0.300.
The mirror tree lands on the same leaf by the other route. The probability that the second ball is red, before you know anything about the first, is also 3/5. Drive the panel until that stops feeling like a trick. Then the first is red given the second was red, which works out to 1/2, and the product is 0.300 again. Neither order is more real than the other. You use whichever conditional you actually know.
And once you see it as a walk, there is no reason to stop at two steps. Each step conditions on everything that has already happened, so a three-step path is just a longer product. Deal three cards off a shuffled deck and ask for three hearts.
The shrinking deck — deal hearts one at a time and watch the naive answer drift wrong
up to five — watch the fractions shrink
deck 52 · hearts 13
deal a heart to begin
What you're looking at — every new card is drawn from a smaller deck
the hearts we're chasing — one leaves the deck each deal
true odds: a product of shrinking fractions 13/52×12/51×11/50
naive odds: 13/52 every time — pretends the deck never shrinks
put each card back and the naive answer becomes right
Fig. 8. The two-step walk had no reason to stop at two. Hit Deal a heart and each card is drawn from a smaller deck, so a longer hand is a longer product of shrinking fractions: 13/52 × 12/51 × 11/50 ≈ 0.0129, not the reflex 13/52 cubed = 0.0156. The red bar shows the naive answer drifting further wrong with every card — because it pretends the deck never shrinks. Now flip to with replacement: put each card back, the shrinking stops, and the naive product turns green because it is finally right.
The first card is a heart with probability 13/52. Given that, the deck now holds 12 hearts among 51 cards, so the second is a heart with probability 12/51. Given both, the third is 11/50. The whole hand comes to 13/52 × 12/51 × 11/50 ≈ 0.0129, so about 1.29% of the time. Every fraction in that product is a world that has already shrunk.
04★ When the shrink does nothing
Now for the special case that everybody meets first and almost nobody meets honestly. Flip a coin, then roll a die, and your friend tells you the coin came up heads. What is the probability the die shows a six?
Still 1/6, obviously. The sample space had 12 equally likely pairs, learning the coin killed six of them, and exactly one of the six survivors shows a six. The world shrank by half and the answer did not move by a hair.
The no-op meter — independence is the shrink that moves your odds by nothing
Tracking A = the die shows a 6. A clue arrives and the world shrinks — does A's probability move? Predict, then click.
condition on the coin
or on the die
pick a clue — does the shrink move A?
The coin halves the world but moves A's odds by nothing. Predict which clue is different.
What you're looking at — conditioning shrinks the 12-outcome world; the needle asks whether A's odds actually budged
the six-column is event A — the two outcomes we track. P(A|B) = gold cells ÷ lit cells, read straight off the grid.
gold = a six that survived the clue (the numerator) and the needle pointing at that fraction.
dimmed = ruled out by the clue B — the new, smaller world you re-measure inside.
green ✓ = needle didn't move, so P(A|B)=P(A): INDEPENDENT, and P(A∩B)=P(A)P(B) is just that swap. A clue that moves it is DEPENDENT (red ✗).
Fig. 9. Tracking one event: A = the die shows a six. New information arrives, and conditioning does exactly one thing — it shrinks the world to the outcomes where the clue held, then re-measures A inside it. Learn the coin came up heads and the world halves from 12 to 6, yet the needle does not budge off 0.1667: one six in six is the same odds as two sixes in twelve. That frozen needle is independence — not "unrelated", not "separate causes", just the flat statement P(A|B) = P(A). Swap the clue to die is even and the whole six-column survives, so the odds double to 0.3333; swap to die is odd and every six is ruled out, so they collapse to 0 — both DEPENDENT. Watch the strip below: the multiplication rule P(A∩B) = P(B)·P(A|B) is always true, but only when the needle holds still may you replace P(A|B) with P(A) — and that single substitution is the whole of the famous product rule P(A∩B) = P(A)P(B).
That is the whole definition, so let's name it before it gets dressed up. Two events are independent when the shrink does nothing, which is to say P(A | B) = P(A). Learning B tells you nothing at all about A. It is a statement about information, not about causes or wires or physical separation.
Now put that condition into the multiplication rule and watch the famous formula assemble itself. We had P(A ∩ B) = P(B) · P(A | B). If A and B are independent, then P(A | B) is just P(A), so
P(A ∩ B) = P(A) · P(B)
That product rule is the thing most courses hand you as the definition of independence, on day one, with no reason attached. It is not the definition — it is a consequence of the shrink doing nothing, and you just watched it fall out in one substitution.
Which raises a question worth being blunt about. If independence is a single equation, how sturdy is it really? People talk about independent events as though it were a property of the story, like two things being unrelated. It is not a story. It is a numerical coincidence between four numbers.
The joint table — independence is one equation, and it is a knife-edge
start from a preset
Drag any block ↕ to move its mass — the other three re-absorb it, the total stays 1.000. A preset that reads INDEPENDENT breaks the instant you nudge.
What you're looking at — independence is one equation, ad = bc, sitting on a knife-edge.
the A row (top two blocks) — the event under study
the Aᶜ row (bottom two)
the divider you watch — level across both columns ⇒ independent
the step = how far ad has drifted from bc: the violation
Fig. 10. A 2×2 joint distribution, drawn so its four blocks' areas always sum to 1.000. Independence is not a story — it is one equation, ad = bc, and the right-hand panel shows it wearing three disguises that always pass or fail together: the joint vs the product, the conditional vs the marginal, the two cross-products. When the gold divider runs level across both columns, the events are independent; grab any block and drag, or hit nudge +0.01, and the line steps apart into a red gap — dependent. Try lopsided (independent without any symmetry) and disjoint (mutually exclusive is maximally dependent). That fragility is the whole point: independence survives only while two products stay exactly equal, so it is an assumption you impose, almost never a fact the world hands you.
Drag mass around that 2×2 joint table and watch the test flicker on and off. Independence survives only while the two diagonal products stay equal, meaning the top-left times the bottom-right matches the top-right times the bottom-left. Nudge any cell and the equation breaks immediately. Independence is a knife-edge, and real markets sit on it about as often as a coin lands on its rim.
There is one more confusion to clear, and it is the most notorious break in the whole chapter. Two events are mutually exclusive, or disjoint, when they cannot both happen, so P(A ∩ B) = 0. Both phrases, mutually exclusive and independent, sound like they mean "these two have nothing to do with each other." Your gut fuses them.
So commit to an answer first. Roll one die, take A as "it's a 1" and B as "it's a 2". These clearly cannot both happen, so are they independent?
Fig 11 · The collision — disjoint events are maximally dependent, not independent
Predict first — then learn B. ↓
One die. A={1}, B={2}. These two events are…
commit a guess above ↑
disjointindep.A⊆B
Mutually exclusive is the OPPOSITE of independent, not a synonym for it.
A={1} — the event under study; its prior is P(A)=.167
B (the evidence) & P(A|B) — conditioning shrinks Ω to B, then re-measures
disjoint: learning B annihilates A → P(A|B)=0. Product test FAILS.
independent is the lone notch where the bar sits ON the prior — shift=0
Fig. 11. Predict, then learn B: disjoint events don’t leave A alone — they collapse it to zero. Turn the dial to find the lone independent notch.
They are as far from independent as two events can get. Learn that the roll is a two and P(A | B) is 0, where before it was 1/6. Conditioning did not leave A alone. It annihilated it. Run the product test and you get P(A ∩ B) = 0 against P(A) · P(B) = 1/36 ≈ 0.0278, which fails.
Here is the sentence to keep. Disjoint events are maximally dependent, because standing inside one tells you with certainty that you are not in the other. That is a large piece of information, not the absence of one. Disjoint events shout at each other. Independent events say nothing.
05Every path that lands on A
We can now walk down a tree and multiply. The next move is to walk down all of the branches and add. That one addition is the machine that makes Bayes possible, and it earns a section of its own.
Two urns sit on a table. Urn 1 is 70% red balls, Urn 2 is 20% red, and you pick an urn with a fair coin before drawing a single ball from it. What is the probability the ball is red?
The denominator machine — collect every path that lands on red, multiply along it, add across
Tap each red leaf → collect its path
PARTITION ●DISJOINT ●
Tap the red leaves to collect the paths
What you're looking at — a coin picks an urn, then you draw a ball; every route to a red ball is one path to collect.
P(Uᵢ) — the branch: how the coin splits the world (must sum to 1)
P(red|Uᵢ) — the red fraction inside that urn
a collected path — multiply along it
P(red) — the paths added across: the answer
Fig. 12. A red ball can arrive by more than one road — through Urn 1 or through Urn 2 — so its probability is not one number but a sum over paths. Walk a path and you multiply along it: the chance the coin sent you to that urn, times the red fraction waiting there (0.50×0.70 = 0.350; 0.50×0.20 = 0.100). Tap each red leaf to drop its product in the tray; collect them all and the total 0.450 locks in and the formula writes itself — P(A) = Σ P(A|Bᵢ)P(Bᵢ), a weighted average before it is a symbol. You may add the paths only because the branches are disjoint (the coin can't land both ways); and the pieces must be a real partition — add a 3rd urn and over-fill the split, and the PARTITION lamp catches the probability leaking out the back. This is the machine that will build Bayes' denominator.
There are two paths to a red ball and they cannot both happen, so you walk each one and add. Urn 1 then red is 0.5 × 0.7 = 0.35. Urn 2 then red is 0.5 × 0.2 = 0.10. Add them and P(red) = 0.450. That is the whole computation, and it is nothing but the multiplication rule used twice with an addition on the end.
Written in general it looks heavier than it is. Cut the world into pieces B₁, B₂, … that do not overlap and together cover everything, which Chapter 1 called a partition. Then
P(A) = Σᵢ P(A | Bᵢ) · P(Bᵢ)
That is the law of total probability, and the sigma is doing less work than it appears to. Read it as a weighted average. How likely A is inside each piece, weighted by how likely each piece is. The pieces must be disjoint, which is what licenses the addition, and they must cover everything, which is what stops you losing probability out the back.
Being an average carries a consequence worth seeing rather than being told, and it is a free sanity check on every total probability you will ever compute.
The caged average — slide the mix, drag the branches, and try to escape the shaded region
P(red) = 0.450, caged in [0.20, 0.70]
What you're looking at — a weighted average is trapped between its branches
the blue posts are the branch conditionals P(red∣urn) — drag them to resize the cage
the shaded CAGE is the span from the smallest to the largest branch
the gold marker is P(red), the mix — it physically cannot leave the cage
0.80 sits outside every branch, so no mixing weight can ever reach it
Fig. 13. The law of total probability re-assembles the branches into one number: P(red) = w(0.70) + (1−w)(0.20) = 0.20 + 0.50w — a straight line in w whose endpoints are the two posts. Slide the mixing weight and the gold marker slides with it, but it stays pinned inside the shaded region; push w to either end and it parks exactly on a post (0.200, 0.700). That is the free sanity check: a mixture is a convex combination, so it always lands between the smallest and largest conditional. Hit try to reach 0.80 and the marker slams against the top post and stops — impossible, 0.80 exceeds every branch. The only way to get there is to change a branch: drag a blue post above 0.80, or add a 3rd branch further out, and watch the cage grow to the new min–max span — now the target is inside, and reachable. Mixing can move a probability around; it can never manufacture an extreme.
Slide the coin's bias in that panel and P(red) moves, but it never leaves the interval between 0.20 and 0.70. It cannot. A weighted average of two numbers is caged between them, and only reaches an end when all the weight sits on that branch. So if you ever compute a total probability that lands outside every one of its conditionals, you have made an arithmetic error, and you now know it before anyone tells you.
06★ Bayes, in three lines
Everything so far ran forwards. You knew which urn, then asked about the ball. Real problems arrive the other way round, and that is where this chapter turns useful. You drew a red ball. Which urn did it come from?
You know P(red | Urn 1), because that is printed on the urn. You want P(Urn 1 | red), which is the same conditional pointed backwards. Nothing we have built so far runs backwards, so we have to make it, and it takes three lines.
Peel one intersection twice — and watch Bayes assemble itself from parts you already own.
1 / 5
P(A∩B): you are at Urn 1 and you draw red.
one block — nothing has been split yet.
One block: the overlap A∩B
◀ prev · next ▶ — step through the five lines. Watch the block on the right get cut two ways.
What you're looking at — one overlap, factored two ways, gives Bayes with no new machinery.
the block is P(A∩B) — the overlap being peeled
blue = the A-side factors: P(A), P(A|B)
cyan = the B-side factors: P(B), P(B|A)
gold = P(A|B), the posterior you're after
Fig. 14. The most famous theorem in the subject, derived rather than decreed. One overlap — being at Urn 1 and drawing a red ball — is a single area, P(A∩B). Peel it left-to-right and it reads P(A)·P(B|A); peel the same area top-to-bottom and it reads P(B)·P(A|B). Two names for one block, so set them equal and divide by P(B): P(A|B) = P(A)P(B|A)/P(B). Fill the denominator from the total-probability sum and the urn numbers finish it at 0.35/0.45 = 7/9 ≈ 0.7778. No new machinery anywhere — step through and watch it assemble.
Line one: the multiplication rule peels the intersection two ways, and both are true at once, so P(A) · P(B | A) = P(B) · P(A | B). Line two: divide both sides by P(B). Line three: there is no line three, because it is already done:
P(A | B) = P(B | A) · P(A) / P(B)
That is Bayes' theorem, first published in 1763 and running in production right now inside spam filters, fraud alerts and radar trackers. There is no new machinery in it anywhere. It is one intersection, written two ways, and divided once. The P(B) on the bottom is exactly the denominator the total probability law builds for you.
Run it on the urns, where P(Urn 1 | red) = (0.7 × 0.5) / 0.45 = 0.35/0.45 = 7/9 ≈ 0.7778. Before the draw you had no idea which urn it was, at 50/50. One red ball dragged you to about 78% confidence in Urn 1. Pull a blue ball instead and the same formula gives 0.15/0.55 = 3/11 ≈ 0.2727, so the evidence pushes the other way.
Those three quantities have names. I'd rather bolt the vocabulary onto machinery you already own than hand you a glossary.
The three slots — hover the anatomy, then feed it a ball and watch belief move
① hover / tap a slot in the picture
② observe a ball → belief moves
③ which one is the likelihood?
Pick the card you think is the likelihood.
What you're looking at — one Bayes computation, its four slots named and colour-coded
prior P(U1) — what you believed before the ball
likelihood P(red|U1) — a red ball's chance if it's Urn 1
evidence P(red) — a red ball's chance at all (normalizer)
posterior P(U1|red) — what you believe now
The one that misleads is the likelihood: it conditions on the hypothesis (Urn 1) and asks about the evidence (a red ball) — P(red|U1), not P(U1|red). Read it backwards and you have already made the inversion the rest of this chapter is about. The same four slots return as the Beta prior in Ch 12 and as the measurement update in Ch 33.
Fig. 15. The same urn computation, dissected into its four named slots — prior P(U1)=0.500, likelihood P(red|U1)=0.700, evidence P(red)=0.450, posterior P(U1|red)=0.778 — each on its own colour with its worked number. Hover a slot and the matching part of the two-urn picture lights up: the likelihood lights Urn 1's red balls, the evidence lights every red ball in both urns. Observe a red ball and the belief bar slides 0.500→0.778; a blue ball pulls it 0.500→0.273. The trap is the one that reads backwards: the likelihood is P(red|U1), the chance of the evidence given the hypothesis — not P(U1|red), which is the posterior.
P(Urn 1) is what you believed before the ball, which is the prior. P(red | Urn 1) is how well a red ball fits that guess, which is the likelihood. P(Urn 1 | red) is what you believe after, which is the posterior. In one breath: belief before, times fit of the evidence, divided by a normalizer, gives belief after.
The word likelihood is a trap and it is worth flagging plainly. In everyday English it means the same thing as probability, so it sounds like it should be the chance the hypothesis is true. It is not — it is the chance of the evidence, assuming the hypothesis. Getting those two backwards is not a beginner's slip, and the rest of this chapter is what happens when a whole profession does it.
07★★ Where is the denominator?
Start with a question that has nothing to do with medicine. Is the chance of rain given clouds the same as the chance of clouds given rain? Almost all rain arrives with clouds, so one of those is close to certain. Most cloudy hours are perfectly dry, so the other is not.
The flip and its price — why P(A|B) and P(B|A) are different numbers, and by exactly how much
Same 99 hours — counted from two worlds. Pick which world you stand in:
pick a side to shrink the world
Same law, three domains where it bites:
What you're looking at — the same 99, measured from two different worlds
blue = the rare thing (rain / sick / a crash)
violet = the common thing (cloudy / a positive test / a signal)
gold = the shared overlap — the identical 99 that never moves
green LOCK: flip ÷ flip always equals base-rate ÷ base-rate
Fig. 16. New information arrives — so shrink the world to where it held, then re-measure. Stand in the rain and it's cloudy 99 times in 100 (0.990) — obvious. Now flip it: stand in the cloudy hours and how many are rainy? Guess before you reveal — almost everyone says 0.990, and the counts refute them: it's 0.2475, because the same 99 now sits in a world of 400, not 100. The two directions are not interchangeable. Divide one flip by the other and the shared 99 cancels, leaving exactly P(rain)/P(cloudy) = 100/400 = 0.25 — a known exchange rate, and it locks green in every domain: a signal that catches 90% of crashes still means a crash only 9% of the time, and a 99%-accurate test on a rare disease still leaves you 9% likely sick. That gap is the most expensive confusion in applied statistics.
Take a thousand sampled hours where 100 were rainy, 400 were cloudy, and 99 were both. Then P(cloudy | rain) = 99/100 = 0.990 and P(rain | cloudy) = 99/400 = 0.2475. Same 99 hours, same table, and the two directions differ by a factor of four.
That factor is not a coincidence, and the panel lets you drive it. Divide one conditional by the other and the shared intersection cancels out, leaving P(A)/P(B) exactly. So the exchange rate between the two directions is the ratio of the two base rates, and here that is 0.1/0.4 = 0.25. Rare thing given common thing is small, always, and by a known amount.
Now the version that actually hurts. There is a disease that affects 1 person in 1000. There is a test for it that is right 99% of the time in both directions: it catches 99% of sick people, and it clears 99% of healthy people. You take the test, and it comes back positive.
Write your gut answer down before you go any further, because the guess is the point of the exercise. Most people, including a famous share of practising doctors, land somewhere near 99%.
Predict first: if you test positive, how likely are you sick? Then stop using percentages and count the actual people.
A test is 99% accurate; the disease hits 1 in 1,000. You test positive — what's the chance you're actually sick?
—
drag the slider, then commit
What you're looking at — the same 100,000 people, counted not percented
gold = the 99 truly sick who test positive (the numerator)
red = the 999 healthy false positives that swamp them
grey = the 99,900 healthy ocean — the base rate you can't see
The crowd is an exact arithmetic partition of 100,000 people, not a simulation: 1 in 1000 → 100 sick, 99,900 healthy; 99% catches 99 of the sick; 1% mislabels 999 of the healthy. All 1,098 positives share one tray, and only 99 are real — so the false-positive crowd is the denominator, and the denominator was the whole answer. '99% of sick people test positive' and 'you tested positive so you're 99% sick' are different sentences with wildly different counts.
Fig. 17. The crowd of 100,000 — a 99% test that is 9% right. First drag a guess for “if you test positive, what's the chance you're sick?” and press COMMIT; your guess pins as a red marker on the scale. Then step four beats: ① the field splits into 100 sick and 99,900 healthy — so few they need a lens; ② the test lights 99 of the sick, missing 1; ③ the test lights 999 of the healthy, who fly in and swamp the 99; ④ everything not positive fades, leaving 1,098 in one tray — 99 gold beside 999 red — and the answer prints 99/1098 = 0.0902. The footer runs the identical numbers through Bayes and lands on the same 0.0902: the false-positive crowd is the denominator.
Drop the percentages and count people instead, because a crowd is something a human brain can actually hold. Line up 100,000 of them. Of those, 100 have the disease, and the test catches 99 of them. The other 99,900 are healthy, and the test wrongly flags 1% of them, which is 999 people.
So look at everyone holding a positive result. There are 99 + 999 = 1098 of them, and only 99 are actually sick. Your probability of being sick is 99/1098 ≈ 0.0902, which is about 9%. Not 99%. The other 91 people out of every 100 who get this frightening result are perfectly healthy.
Notice what your gut did, because the mistake has a name and a shape. It anchored on the test's 99% accuracy, which is P(positive | sick), and reported that as the answer to P(sick | positive). That is the inversion we just priced. It also never counted the healthy people, and those 999 false positives are the denominator. This is base-rate neglect, and it is forgetting the prior and the denominator in one move.
The arithmetic agrees with the crowd, which is the useful cross-check. Bayes says P(sick|+) = (0.99 × 0.001) / (0.99 × 0.001 + 0.01 × 0.999), and the bottom is the total probability law assembling P(positive) out of the two branches. That is 0.00099 / 0.01098 ≈ 0.0902, the same 9%.
Which leaves a fair question: if a 99% test is this weak, what would actually make it strong?
The prevalence dial — sweep the prior and the error rates, and watch the one number the test actually hinges on.
LR = 99.0 · break-even 1 in 100
at 1 in 1,000, buy ONE upgrade — which one rescues it?
tap an upgrade — then predict which moves the answer.
drag the dials — watch the answer swing
What you're looking at — the answer hangs on the false-positive rate, not on sensitivity
gold = P(sick | +), the posterior you actually care about — the curve, the marker, the number up top
cyan = the break-even prevalence 1/(1+LR): the curve crosses 50% exactly here, and it is the whole design in one number
blue = sensitivity — feels decisive, but sharpening it barely nudges the answer for a rare disease
red = specificity, i.e. the false-positive rate — the lever that actually drags a rare disease from 9% to ~50%
Fig. 18. One live curve: P(sick | +) against how rare the disease is. Park it at 1 in 1,000 and a 99%/99% test reads a mere 9% — because the healthy millions throw off far more false positives than the few sick throw true ones. Now run the duel: sharpening sensitivity 99→99.9% barely twitches the answer, while tightening specificity the same amount hauls it from 9% to ~50%. The single number the whole design turns on is the break-even prevalence1/(1+LR): at 99/99 that is 1/(1+99) = 1 in 100, which is exactly why the middle snap-point lands on 50.0%. And it is not sensitivity.
Slide the prevalence and the posterior swings enormously. At 1 in 1000 you get 9%, and at 1 in 100 the same test lands on exactly 0.500, because the break-even prevalence is 1/(1+99) = 0.01. Now leave prevalence alone and tighten the false-positive rate from 1% to 0.1%. The same rare disease jumps from 9% to about 0.498.
That is the design lesson hiding in the panel, and it transfers well beyond medicine. When you are hunting something rare, sensitivity barely matters and specificity is nearly everything, because the false positives are drawn from a vastly larger crowd. Fraud detection, alarm systems and strategy backtests all live on that same knife.
One honest footnote before we leave the clinic. A single positive is weak, but evidence compounds. Take a second, independent positive test and the 9% posterior becomes the new prior, and Bayes returns about 0.9075. That is why doctors retest rather than despair, and it is the same update loop Chapter 33 runs on a moving target.
08When a streak means something
We close on an apparent contradiction. You are now holding two ideas that seem to fight. Independence says past flips tell you nothing about the next one. Bayes says you should update on evidence, and a run of data is evidence. Which is it?
Here are two coins, and both have just landed heads five times in a row. The first is a coin you inspected yourself and know to be fair. The second came out of a drawer and might be an ordinary coin or might be two-headed, and you have no idea which. For each one, what is the probability the next flip is heads?
Two coins, one streak — HHHHH on both. Predict the next flip for each, then reveal.
Coin A
Coin B
Predict both, then press Reveal
What you’re looking at — the same five heads, judged two ways
Coin A is known fair, so P(next H) is nailed at ½ — the streak is noise.
Coin B is a drawer coin (½ fair, ½ two-headed); each head drags P(next H) up toward 65/66.
the belief it’s the two-headed coin — five heads pushes it to 32/33.
a tail has probability 0 under “two-headed” — one kills the belief and snaps back to ½.
Fig. 19. Both coins just landed HHHHH. Guess P(next heads) for each, then Reveal. Coin A you inspected — its per-flip probability is known and fixed at ½, so the answer is a flat 0.500 and the run tells you nothing (press bet tails, flip 200× and watch the tally stay stubbornly at half — that’s the gambler’s fallacy). Coin B came from a drawer that’s half fair, half two-headed — its probability is unknown, so every head is evidence about the coin: the two-headed belief climbs 2/3, 4/5, 8/9… to 32/33 and the prediction rises to 65/66 ≈ 0.985. Now flip a tail: it has probability zero under “two-headed,” so one is enough to collapse the belief and snap you back to ½. Same five heads, two answers — the discriminator is simply whether the per-trial probability was known before the flipping started.
The fair coin answers 0.500, and it will answer that forever. The flips are independent, so conditioning on the streak does nothing at all, and tails is not due. Coins have no memory — and no obligation to balance the books. Believing otherwise is the gambler's fallacy, and it is treating independent trials as though they were dependent.
The mystery coin answers 65/66 ≈ 0.985, and it is not a different rule, it is Bayes. A two-headed coin produces five heads with probability 1, while a fair one manages it with probability 1/32. So the streak drags the posterior for "two-headed" to 32/33 ≈ 0.9697, and the next flip inherits that belief. Same five heads, opposite correct answers.
Here is the discriminator, and it is the thing to carry out of this section. Ask whether the per-trial probability is known and fixed. If it is, the trials are independent and the past is irrelevant to the next one. If it is unknown, then every trial is evidence about the mechanism, and the data should move you.
That distinction is not a party trick, and in markets it is close to the whole job. A strategy has just made money nine months running. Is that a fixed edge you can now trust, or a fair coin having an ordinary run? Chapter 16 onwards is largely about answering that question honestly, and this chapter is the machinery it uses.
The router — quote it a real claim and it names the confusion hiding inside, then points downstream
① pick a real claim someone told you
Pick a claim — the router traces it to the exact confusion it hides, then the map shows where that machinery runs next.
② the one move, four names — tap to recall each
Which way? And whose denominator?
What you're looking at — one router that names four confusions, and the roads they run down
wrong direction — P(A|B) read off P(B|A); the flip has a base-rate price
missing denominator — the false-positive crowd nobody counted
not independent — the joint cell beats P(A)·P(B)
fixed vs unknown — a run is noise, or evidence about the coin
Carry one question out of this chapter: when someone quotes you a conditional probability, ask which direction it runs and where its denominator came from. A test that is 99% accurate yet 9% right is not a paradox — it is a denominator nobody counted.
Fig. 20. Everything on this page was one move — shrink the world, renormalize inside it, multiply out, flip — re-used four times. Quote the router a real claim and it names which confusion the claim forgot: wrong direction (P(A|B) read off P(B|A)), missing denominator (the false-positive crowd), not independent (the joint cell beats P(A)·P(B)), or fixed vs unknown (a run is noise, or evidence about the coin). The map on the right lights Ch 10 and runs roads out to where this same machinery reappears — tap a chapter to read the debt it owes this page. The one question to carry: which direction does the conditional run, and where did its denominator come from?
So let's collect the chapter into one move, because it really is only one. Learning B shrinks your world to B and forces you to rescale it, and that rescale is the division in P(A | B) = P(A ∩ B)/P(B). Multiply it out and you get a walk down a tree. Set the walk's two orders equal and you get Bayes. Notice a case where the shrink changes nothing and you have independence.
And carry the one question that catches the most expensive errors in this field. When somebody quotes you a conditional probability, ask which direction it runs, and ask where its denominator came from. A 99% accurate test that is 9% right is not a paradox. It is a denominator nobody counted.
Next we attach a number to every outcome instead of a truth value. A die roll stops being a face and becomes a payoff, and the weights we have been shuffling around become something you can average. That average is expectation, and it is where probability starts paying rent.