◈ quant roadmapPart 1 · Ch 09/45
Quantitative Finance — the Mathematics of Markets · chapter 09

09Probability Axioms & the Sample Space

Chapter 8 handed us an arrow that finds the bottom of any landscape, and with it Part 0 closed. Every tool the field quietly assumes is now built. This chapter points the whole toolbox at the thing it was sharpened for, which is putting an honest number on an event that has not happened yet. You already own one such number. Favourable over total is the rule you were handed at about nine years old, and you have never had cause to doubt it. We are going to break it in the first two minutes. Then we rebuild probability from three plain rules about weight, and out of those three rules the entire familiar list falls: the complement rule, the empty event, monotonicity, and the addition law with its subtracted overlap. Somewhere in the middle the school formula walks back in, wearing a different name. It turns out to be a theorem that quietly assumes every outcome is equally likely, which is a claim about the physical world and not a fact of mathematics. By the end you will be able to take a probability question you have never seen, build its sample space, decide whether counting is even licensed, and finish with 23 people in a room and a result that ought to feel impossible.

Before the first die is rolled, here is where this page sits in the course. Part 0 is behind us and Part 1 begins right here, at the root of everything that follows. Tap a downstream node in the panel below and watch the road run back to this page.

You are here — Part 0 closes; from here on, everything is weight spread on a set
PART 0 MODEL ESTIMATE LEARN PRICE ACT certain → uncertain THIS PAGE HANDS IT → P(A) : one unit, Σ=1 Ch 11 Ch 12 Ch 19 Ch 27 Ch 30 CH 09 you are here what Ch 8 built: the gradient — it points to the one answer ∇f certain: one arrow, one destination — nothing here is uncertain the same unit, now spread on a set: one weight → six equal ⅙ lumps uncertain at last: ⅙ on each — this is what probability IS
① tap a blue node above — trace what it owes this page
tap a node to see its debt
② what Ch 8 handed us
the machinery is built — nothing it touched was uncertain yet
What you're looking at — the page where certainty ends and weight begins
gold = Ch 09 (here), where weight-on-a-set is first defined
blue = later chapters that all spend that same weight
cyan seam = Part 0 was certain; from Ch 09 on, it isn't

Probability is nothing but a function defined on sets — which is why Part 0 spent eight chapters on sets, functions and counting. Its very first theorem, favourable / total, turns out to be a counting problem. Expectation, every named distribution, the birthday shock, an option's price, a Kelly bet — all of them are just this one unit of weight, spread and re-weighed.

Fig. 1. The course spine, Part 0 through Act, with Ch 09 lit gold as home and a seam marking where certain maths becomes uncertain probability. Tap any blue node — Ch 11's expectation, Ch 12's distributions, Ch 19's birthday collision, Ch 27's risk-neutral price, Ch 30's Kelly bet — and a gold road traces home while the inset names the debt: every one of them is this page's single unit of weight, worn as a costume. Below, flip what Ch 8 handed us and watch the deterministic gradient arrow morph into a die whose six faces each carry an equal lump — the same machinery, aimed at uncertainty for the first time.

Three of those roads are worth naming now. Chapter 11 turns an event's weight into expectation, which is just that weight used to average a payoff. Chapter 27 prices an option as a discounted expected payoff, so the whole of derivatives pricing is this chapter's machinery run at altitude. And Chapter 16 admits the uncomfortable part: in real markets nobody hands you the weights. You estimate them from a finite, noisy sample.

01The formula that breaks

Let's start with the one probability fact everybody in the room already owns. Count the outcomes you want, divide by the total number of outcomes, and that is the probability. A fair coin has one head out of two sides, so heads is 1/2, and a fair die has three even faces out of six, so even is 3/6. It works, it is quick, and almost nobody is ever told what it assumes.

So here is a coin. It has been bent in a vice until it lands heads seven times in ten, and we have thrown it ten thousand times to be sure. It still has exactly one head, and it still has exactly two sides. Commit to an answer before you touch the panel: what does favourable over total say the probability of heads is, and what is it really?

The bent coin — where "favourable over total" quietly lies
the coin, seen edge-on heads ▲ bend: strong THE FORMULA'S LEDGER favourable (heads) = 1 total (sides) = 2 P = 1 / 2 = 0.500 what the coin actually does 1.0 0.5 0.0 formula 0.500 heads tails 0.700 0.300 gap throws: 0 commit a guess ↓ to run 10,000 throws
Now commit — is the school formula right?
pick a bend, then commit a guess
drag to bend · then commit a guess
What you're looking at — the count is flawless, the answer is still wrong
the bent coin — one head, two sides
the formula, frozen at 0.500
what the coin really does — heads 0.700

Favourable / total counts outcomes and assumes they're interchangeable — but a head and a tail aren't. Slide back to flat and the red line meets the truth: the formula is right only when the two sides are equally likely. That hidden assumption is exactly why probability needs a real definition.

Fig. 2. Bend the coin, then commit a guess before the reveal. The formula's ledger stays frozen at 0.500 while 10,000 throws settle at 0.700 — the count was perfect; it just can't see that a head and a tail aren't interchangeable. Slide back to flat and the lie disappears: the formula is right only when the two sides are equally likely.

The formula answers 0.500, because one head out of two sides is all it can see. The coin answers 0.700. Nothing about the counting was done wrong. The formula was asked a question it has no equipment to handle: it counts outcomes, and this coin's two outcomes are not interchangeable.

You might reasonably reply that bent coins are a stunt. Fair enough. Here is the same failure with perfectly fair equipment and no trickery at all. Roll two ordinary dice and add the faces. The possible sums are 2 through 12, so there are eleven outcomes on the list. How often does a seven come up?

Two fair dice, eleven sums — flip between counting SUMS and counting DICE, then let 2,000 real rolls pick a side.
P( sum = 7 ) favourable 1 ÷ total 11 = 0.0909 every slot returns 0.0909 — that flatness is the tell FORMULA 1/11 = 0.0909 vs TRUTH 6/36 = 0.1667
what I'm counting ↓
tap any sum on the stage to price it
flip to DICE — watch the same click change its answer
What you're looking at — the same two fair dice, counted two ways. Listing outcomes is not listing equally-likely outcomes.
SUMS — eleven look-alike slots. The school formula favourable/total reads 0.0909 for every sum. It can't tell 7 from 2 — that sameness is the lie.
DICE — each slot grows into its ordered rolls (the 36 truly equally-likely atoms). 7 has six ways, 2 has one — so 6/36 vs 1/36.
2,000 real rolls pile onto the towers, never the flat row. Both dice were fair and all 11 sums were possible — summing them just destroyed the equal weighting the formula silently needs.
Fig. 3. The bent coin was a stunt, so here is the same failure on honest gear. Two fair dice make eleven possible sums, so favourable/total confidently claims 1/11 = 0.0909 for a seven — and, tellingly, the very same 0.0909 for every other sum. Flip what I'm counting to DICE and each slot explodes into the ordered rolls that feed it: seven stands six blocks tall, two stands one, and the ledger recomputes to the real 6/36 = 0.1667. Hit roll 2,000 times and the cyan tallies land on the towers, never the flat row. Tap the two and the scorecard shows the formula overshooting (0.0909 vs 0.0278); tap the seven and it undershoots (0.0909 vs 0.1667). Both dice were fair and every listed sum was genuinely possible — but summarising two dice into one number destroyed the equal weighting the formula silently relies on. Listing outcomes is not the same as listing equally-likely outcomes.

Favourable over total says one sum out of eleven, which is 0.0909. Roll the dice and a seven arrives about 0.1667 of the time, which is nearly double. Both dice were fair, and every sum on the list was genuinely possible. The list was still the wrong thing to count, and that is the crack we spend this chapter sealing.

Notice what actually went wrong, because it is the same fault twice. In both cases we listed the outcomes correctly and then treated them as though they were interchangeable. The bent coin's head and tail are not. A sum of seven and a sum of two are not, because seven can be built six different ways and two can be built only one. We need a definition of probability that never needed them to be interchangeable in the first place.

02Every event is a subset

Rebuilding starts with something almost insultingly simple. Take an experiment, meaning any procedure with an uncertain result, and write down every distinguishable result it can produce. For one roll of a die that list is {1, 2, 3, 4, 5, 6}. That set has a name: it is the sample space, written Ω, and its elements are called outcomes or atoms.

Now the move that most courses spend half a sentence on, and that is worth a great deal more than half a sentence. Take a statement about the experiment, something like "the roll is even". Ask which outcomes make that statement true, and the honest answer is 2, 4 and 6. So the statement has quietly handed you a set, {2, 4, 6}, and that set is what we call the event.

The word event fights you here, and it is worth naming why. In English an event is something that happens in time, so your mind pictures the die tumbling and coming to rest. Put that picture down. In this subject an event is a static subset of Ω, chosen before anything is rolled. Drive the panel below and watch sentences turn into sets.

An event isn't a happening in time — it's the set of outcomes that make a plain-English sentence true. Watch the sentence pick it out.
Ω every way the die can land picks → the event sentence — pick a statement — true faces the set
Tap a statement. I sweep Ω face by face, stamp each true or false, then lasso the true ones into a set.
each sentence names a set — that set is the event
What you're looking at — a sentence is a machine that picks a subset of Ω, and that subset IS the event
Ω (the blue box) is the sample space — all six outcomes the die can produce, before anything is rolled
the green faces are the outcomes that make the sentence true — asked one at a time, in the sweep
the gold lasso gathers them into the event: a set like {2,4,6}. Same three things, printed below the drawing
aha — the word fights you: in English an event happens in time. Here it is a static subset of Ω, chosen before the roll. "the roll is a 7" lassoes nothing → , the impossible event; "between 1 and 6" lassoes everything → Ω, the certain event. Flip to build a set and the map runs backwards — every subset you draw has a sentence.
Fig. 4. An event is not something that happens — it's the set of outcomes that make a sentence true. Tap a statement and watch it sweep Ω, stamp each face, and lasso the true ones into a set like {2,4,6}. Two are deliberately extreme: "the roll is a 7" catches nothing → , the impossible event; "between 1 and 6" catches everything → Ω, the certain event — the smallest and largest events, arriving for free. Flip to build a set and the map runs backwards: every subset you draw has a sentence waiting for it.

Two extremes in that panel are worth pausing on. "The roll is a seven" makes no outcome true, so it hands you the empty set , the impossible event. "The roll is between 1 and 6" makes every outcome true, so it hands you Ω itself, the certain event. Impossible and certain are not special cases bolted on afterwards — they are simply the smallest and largest subsets you can name.

And now the payoff for Chapter 1. Once events are sets, the ordinary words of logic stop being vague. "A or B" is the union A ∪ B, because an outcome makes the compound sentence true exactly when it lies in one set or the other. "A and B" is the intersection A ∩ B. "Not A" is the complement Aᶜ. The set algebra you learned in Chapter 1 turns out to be the logic of events, unchanged and pointed somewhere new.

The event workbench — "or", "and", "not" are just Chapter 1's union, intersection and complement, aimed at chance
A = even B = at least 4 disjoint? they share faces — A ∩ B ≠ ∅ outside · in neither lasso Ch.1 · unchanged A B
Tap a die face to move it in / out of lasso:
operation · paints the region
operationA ∪ Bunion
englisheven or ≥ 4
set{2, 4, 5, 6}
they overlap — share 2 faces
What you're looking at — a compound event is nothing but a set of faces
A = the even faces {2,4,6}
B = at least 4 {4,5,6}
gold = the faces the operation keeps
lamp green = disjoint: A ∩ B = ∅ (Axiom 3 will need this)
Aha: a roll makes "A or B" true exactly when it lands in one lasso or the other — and that is the definition of union. The English word and the set operation are the same object. Nothing new was invented; Chapter 1's set algebra just got pointed at chance.
Fig. 5. Compound events are Chapter 1's set operations, now pointed at chance. With A the even faces and B the faces ≥ 4, drive ∪, ∩, ᶜ, ∖ and watch the English sentence and the set of faces move as one thing. Drag the two lassos apart until they share no face and the disjoint lamp turns green — A ∩ B = ∅ — the exact condition Axiom 3 will lean on when disjoint events are allowed to simply add.

Take A as "even", which is {2, 4, 6}, and B as "at least 4", which is {4, 5, 6}. Then A ∩ B = {4, 6} and A ∪ B = {2, 4, 5, 6}. One more word to bank while we are here. Two events with no outcome in common are called disjoint, or mutually exclusive, which just means A ∩ B = ∅. That word is about to do an enormous amount of work.

03One unit of weight

We have the sets. Now we need the number, and the shape of the object that supplies it is where beginners genuinely stall. Every function in Part 0 took a number and returned a number. This one takes a set. So P(A) is not P multiplied by A, and it is not a typo — it is a function being fed a subset of Ω.

Here is the picture, and we keep it for the rest of the chapter. Imagine one unit of weight, a single kilogram of sand, and you get to spread it over the six faces of the die however you like. Spread it evenly and the die is fair. Pile it all on the six and the die is loaded. Whatever you do, the total on the table stays one kilogram.

The amount sitting on a single outcome is that outcome's probability mass, and the little table of "how much on each atom" is the probability mass function, or pmf. Now circle any event, meaning any subset, and ask the obvious question. How much weight is inside the circle? That total is P(A). P is a weighing machine for regions.

The weighing machine — one kilogram of sand, and P weighs the circle you draw
drag a pile ↕ · tap a die to lasso it Σ = 1.000 0 ½ 1 0.167 P(event) one kilogram of sand — the six piles always total 1.000 0.167 = 0.167
Drag a pile to reweight it — the other five shrink so the sand still totals 1 (weight is conserved). Tap a die to draw it into your circle.
P({4}) = one pile = 0.167
What you're looking at — P is a weighing machine, not a formula waiting to be applied.
each blue pile is the weight P puts on one face — its pmf; the six always total 1
green = the sand that lies inside the circle you drew — the favourable weight
the dial reads the total sand inside: that sum is the event's probability
Fig. 6. There is exactly one kilogram of sand — spread it however you like over the six faces by dragging a pile; the other five always shrink to keep the total at 1, so you can feel weight being conserved long before Axiom 2 says it. The height of a single pile is that face's probability, its pmf value: P({4}) is just one pile. Tap dice to draw your circle, and the dial weighs the sand inside it — that total, printed longhand, is the event's probability. That is all P ever is: a machine that weighs the circle you drew.

Drag the sand around in that panel and the definition becomes something you can feel. The probability of an event is the sum of the masses of the outcomes inside it, and nothing more mysterious than that. On a fair die each face carries 1/6, so the event {2, 4, 6} weighs 1/6 + 1/6 + 1/6 = 0.500.

One more hard stare at P before we constrain it, this time at its type. Our Ω has six atoms, so the number of subsets you could hand to P is 2⁶ = 64. Sixty-four legal inputs, from the empty set up to Ω itself, and every one of them comes back as a single number between 0 and 1.

Sixty-four doors, one rail — P's whole domain and codomain, laid out at once
1 Ω · all of it 0 ∅ · no sand P(A) P set ↦ number FAIR: 64 doors → 7 heights height = k / 6, k = faces in A
domain: the 64 events → codomain: [0,1]
tap any door in the grid ↖
tap a door → see its P‑value
What you're looking at — P's whole type, drawn once: a SET goes in, a NUMBER comes out
blue dots = the faces inside one event A (the SET you hand to P); each of the 64 glyphs is a legal input
gold rail = the codomain [0,1]; every event lands on it — ∅ (no sand) pinned at 0, Ω (all of it) pinned at 1
cyan marker = the value flying from a door to its P‑height; the shelf's width at a height = how many events share it
FAIR collapses all 64 onto just 7 heights, k/6 — the keystone you can see now but can't yet name; LOADED SIX scatters them
Fig. 7. Here is the whole reason P(A) is not a type error, drawn once. Chapter 1 taught events as sets; a die has six faces, so there are exactly 2⁶ = 64 subsets you may legally hand to P — every one is a door in the 8×8 grid, ordered by size so the empty set sits at the top‑left and the whole space Ω at the bottom‑right. Tap any door and a cyan marker flies to its landing point on the single gold rail from 0 to 1: domain = the 64 events, codomain = [0,1]. Two doors are pinned by meaning, not by choice — holds no sand so it must land on 0, Ω holds all of it so it must land on 1. Hit spray all 64 and watch the crowd settle. Under FAIR the 64 points fall onto only seven heights — 0, 1/6, 2/6, … 1 — a foreshadowing of the coming keystone P(A)=|A|/6 you can already see but not yet name; flip to LOADED SIX and those seven neat shelves shatter, because when the outcomes aren't equally likely the tidy ratio simply isn't there.

That panel is the whole domain of P laid out at once. Click any of the 64 doors and the answer lands on the rail. Two of them are pinned by meaning rather than by arithmetic: the empty set weighs 0 because it contains no sand, and Ω weighs 1 because it contains all of it. Hold onto those two, because they are about to stop being observations and become axioms.

04★ Three rules and nothing else

In 1933 Andrey Kolmogorov wrote down what a probability has to be, and the striking thing is how little he needed. Three rules. Everything else in this chapter, and a fair amount of the rest of the course, is squeezed out of them. Read each one as a statement about sand rather than as a commandment.

Axiom 1, non-negativity. For every event A, we require P(A) ≥ 0. You cannot place a negative amount of sand on a face, because there is no such physical thing — and where there is no such thing, there is no such probability.

Axiom 2, normalization. P(Ω) = 1, meaning the whole pile is exactly one unit. This is the axiom that fixes the scale, and it is the reason probabilities are comparable at all.

Axiom 3, additivity. If A ∩ B = ∅, then P(A ∪ B) = P(A) + P(B). Two piles that share no grain of sand combine by plain addition, because nothing gets counted twice. This is the workhorse, and every theorem below is really this axiom applied to a cleverly chosen cut.

That is the entire definition of probability. Notice what is not in it. There is no mention of coins, of fairness, of equal likelihood, or of dividing anything by anything. A bent coin satisfies all three axioms perfectly well. Try to build a P that does not.

Build an illegal P — the axioms as weight behaving like weight
A1 · NON-NEGATIVITYevery mass ≥ 0
A2 · NORMALIZATIONP(S) = 1.00
A3 · ADDITIVITYdisjoint A,B: P(A∪B)=P(A)+P(B)
press “test a disjoint pair” to check A3 live →
dare the machine to break a rule:
legal P — a valid spread of sand
What you're looking at — six piles of sand you can pile any way you like
each bar = the mass on one die face — drag it ↕
the running total P(S) — it must be 1 (A2)
an axiom breaking — there is no negative sand (A1)
Try “break A3”. You can’t — because P(an event) is defined as the sand sitting inside it, so two disjoint pieces already add. Additivity is not an extra assumption bolted on; it is what “total weight in the circle” already means. That is why THESE three axioms and no others.
Fig. 8. Six unconstrained piles of sand — the mass on each face of a die, free to run from −0.4 to +1.4, and deliberately not auto-normalised. Drag a bar below zero and A1 flips red — there is no negative sand. Push the total off 1 and A2 flips red with the running P(S). Press test a disjoint pair and watch A3 hold to three decimals every time; dare break A3 and after three failures the machine explains why it cannot: additivity is not an extra rule, it is what "total weight inside the circle" already means. Repair renormalises your mess back to a legal P and tells you what it changed. That is why THESE three axioms and no others — they are the weight picture written as rules.

Push the weights to the extremes in that panel and watch which lamp goes red. Drag a face below zero and Axiom 1 fails on the spot. Pile up more than one unit and Axiom 2 fails. What you cannot do is find a spread of sand that feels wrong and yet satisfies all three. The axioms are not a filter placed on top of the weight picture. They are the weight picture, written as rules.

Axiom 3 deserves its own experiment, because its condition is the one people forget under pressure. It only licenses addition when the two events are disjoint. Watch what the sum does when you let the circles touch.

Slide event B across the die, one face at a time, and watch where P(A)+P(B) stops equalling P(A∪B).
Ω = A 1 2 3 4 5 6 B 1.0 0.5 0 P(A)+P(B) P(A∪B) overshoot 0.000 = P(A∩B)
P(A)+P(B)1.000
P(A∪B)1.000
overshoot0.000
A3 applies
Disjoint — addition is exact.
What you're looking at — the exact price of an overlap.
P(A) — faces 1,2,3 (fixed)
P(B) — the sliding window
P(A∪B) — weighed directly
overshoot = weight of A∩B, counted twice
Fig. 9. Axiom 3 says probabilities add — but only for disjoint events, and that clause is the one people drop under pressure. Start disjoint: the left tower P(A)+P(B) and the directly-weighed right tower P(A∪B) are the same height, and the lamp reads A3 applies. Now slide B one face into A. The instant they share an outcome, the left tower runs tall by a red block, and that block is never a mystery: it is always, to the last decimal, the weight of the region they share — watch the readout, overshoot always equals P(A∩B). Flip to LOADED and the same equality holds on an unfair die, so it is no artefact. That overshoot is the axiom marking the edge of its own licence — and the reason the fix is to subtract the shared region once: P(A∪B) = P(A)+P(B)−P(A∩B).

Keep the events apart and P(A) + P(B) tracks P(A ∪ B) exactly. Slide them together and the sum runs high, and the excess is precisely the weight in the shared region. That overshoot is not a flaw in the axiom. It is the axiom telling you, in numbers, where its permission stops, and we cash that in as a theorem when we prove the addition law.

05Everything else is a theorem

Here is where three lines start paying rent. Every law you were ever handed about probability can be derived from those axioms, and each derivation follows the same recipe. Rewrite the event you care about as a union of disjoint pieces, apply Axiom 3, and use Axiom 2 wherever Ω shows up. No new assumptions enter anywhere.

Start with the complement rule, where an event A and its complement Aᶜ share nothing and together cover everything. So A ∪ Aᶜ = Ω and A ∩ Aᶜ = ∅. Axiom 3 then gives P(A) + P(Aᶜ) = P(Ω), and Axiom 2 says that right-hand side is 1. Rearrange and you have P(Aᶜ) = 1 − P(A), which will do more work in this course than almost any other line.

The law rack — four everyday theorems, one recipe, one axiom per line
THE RECIPE — rewrite → add → use Ω NUMERIC CHECK — fair die, A = {2,4,6} 1 2 3 4 5 6 AXIOMS PICKED UP SO FAR A1 ×0 A2 ×0 A3 ×0 only three tools, ever
COMPLEMENT · step 1 of 4
press next line to reveal each step
algebra / set logic — no axiom yet
What you’re looking at — four everyday laws, each derived from the same three axioms
A1 — every probability is ≥ 0 (nothing is rarer than impossible)
A2 — the whole space is certain: P(Ω) = 1
A3 — disjoint events add: P(A∪B) = P(A)+P(B)
alg — ordinary set logic & arithmetic, no axiom spent
The ceiling P(A) ≤ 1 is forced by two axioms and one cut — nobody decreed it. A model that reports 1.4 has broken an axiom upstream, not found an exotic case.
Fig. 10. Four everyday laws sit in a rack — the complement rule, P(∅) = 0, monotonicity, and the ceiling P(A) ≤ 1 — and not one of them is an extra rule to memorise. Pick a card and press next line: every card runs the same recipe. Rewrite the target event as a disjoint union, add the pieces with A3, and wherever the whole space Ω appears, replace P(Ω) with 1 using A2; when a leftover piece must be non-negative, that is A1. The chip on every line names the exact tool, the picture on the right shows the cut, and the die keeps a live number on each abstract step (A = {2,4,6}, so P(A) = 0.5). The tally at the bottom counts every axiom you spend across all four derivations — and it never climbs past three tools. That is the payoff: the ceiling at one is not a decree, it is forced by A2 and a single cut, which is why a model that ever reports 1.4 has broken an axiom upstream rather than found an exotic case.

Step that rack of derivations and you get three more for free. Take A = Ω in the complement rule and out drops P(∅) = 0, so the impossible event carries no weight. For monotonicity, cut a larger event B into the disjoint pair A and B ∖ A. Then P(B) = P(A) + P(B ∖ A). Axiom 1 makes that last term non-negative, so A ⊆ B forces P(A) ≤ P(B). Put B = Ω in that and you get the ceiling, P(A) ≤ 1.

Sit with that for a second, because it is the substance of the chapter: nobody decreed that probabilities top out at one. It is forced, by two axioms and one cut. If a model of yours ever reports a probability of 1.4, you have not found an exotic case; you have broken an axiom somewhere upstream.

That leaves the addition law, which is the one people half-remember from Chapter 1's counting. You cannot just write P(A ∪ B) = P(A) + P(B), because Axiom 3 refuses to apply to overlapping events. So we have to earn it, and the way you earn it is with scissors.

Fig 11 · The legal cut — Axiom 3 won’t touch overlapping events, so you cut first, then subtract
Axiom 3 requires A ∩ B = ∅. It is not. A3? 2 5 4 6 A B A∖B A∩B B∖A A B A∩B counted: 2×
Beat 1 · The illegal move
tap Next ▶ to run the proof
Illegal — A3 needs disjoint sets
Ch 1 counted |A∪B| = |A|+|B|−|A∩B|. Same shape — here proved.
What you’re looking at — Axiom 3 is only licensed on pieces that touch nowhere
A∖B — outcomes in A only: {2}
A∩B — the shared lens {4,6}: the block that gets double-counted
B∖A — outcomes in B only: {5}
cutting first makes the pieces disjoint, so A3 finally applies — and the lens comes off once
Fig. 11. Beat 1 writes the tempting move — P(A∪B) = P(A) + P(B) — and a red [A3?] chip flags it: Axiom 3 only adds events that touch nowhere, and these overlap on {4,6}. So beat 2 cuts: the union splits into three disjoint pieces — A∖B, A∩B, B∖A — and the chip goes green. Beats 3–4 rebuild A and B from those pieces and the gold lens lands in the pile twice; beat 5 lifts one off, closing P(A∪B) = P(A) + P(B) − P(A∩B). The die strip proves it live — 1.000 − 0.333 = 0.667, the exact weight of {2,4,5,6} — and the FAIR/LOADED toggle re-runs it on lopsided weights: same picture, different numbers, identity holds. It’s Chapter 1’s count, now proved from an axiom for any weights at all.

The cut in that stepper is the whole proof. Slice A ∪ B into three pieces that touch nowhere: the part of A outside B, the shared lens A ∩ B, and the part of B outside A. Now Axiom 3 is legal on all three, so their weights add. Reassemble P(A) + P(B) and you find the lens has been counted twice, so subtract it once. The result is P(A ∪ B) = P(A) + P(B) − P(A ∩ B).

Check it on the die. With A = {2, 4, 6} and B = {4, 5, 6} on fair weights, P(A) + P(B) = 0.5 + 0.5 = 1.0, which is already absurd for an event that misses the faces 1 and 3. The lens {4, 6} weighs 2/6, and taking it off leaves 4/6 ≈ 0.667. That is exactly the weight of {2, 4, 5, 6}, counted directly.

This looks like Chapter 1's inclusion–exclusion, and it should, because it is the same picture. The difference is the status of the claim. In Chapter 1 we counted and saw the overlap double-counted. Here we proved it, from an axiom, for any weights at all, fair or bent. Same shape, sturdier foundation.

06★★ Where favourable-over-total comes from

We have built quite a lot without ever dividing favourable by total. So where does the formula you have used your whole life actually live? It is hiding inside Axiom 3, with one extra assumption strapped to its back, and that assumption is doing all the work.

Take a fair die and add up the weight of the even faces the slow way. The event {2, 4, 6} is the disjoint union of the three single-outcome events {2}, {4} and {6}, so Axiom 3 applies three times over. That gives P(even) = 1/6 + 1/6 + 1/6 = 3/6. Look hard at where those two numbers came from before you go on.

The keystone — watch favourable / total assemble itself out of Axiom 3, then load the die and catch it lying
A fair die — add the even faces with Axiom 3 11/6 21/6A3 31/6 41/6A3 51/6 61/6A3 ☝ tap faces 2, 4, 6 — one at a time P(even) = 1/6 + 1/6 + 1/6 = 3 6 the 3 = how many even faces there are = |A| the 6 = only because every face weighs the same 1/6 = 1/|Ω| so «favourable / total» was never a rule — it fellstraight out of adding equal weights.
Beat 1 of 3 · assemble
Tap the even faces (2, 4, 6). Each click adds one 1/6 by Axiom 3.
Tap the even faces to build P(even).
What you’re looking at — where «favourable over total» comes from, and where it breaks
green = the favourable even faces {2,4,6}, and the TRUE probability that adds their real masses
gold = the number to watch: the 3/6 you built, and the count/weight it’s made of
red = the formula |A|/|Ω|, frozen at 0.500 — confidently wrong once the die is loaded
The 3 is a count; the 6 only appears because every face carried the identical 1/6. Load the die and P(even) climbs to 0.600 while the formula still insists on 0.500 — yet non-negativity, total-one, and disjoint-add stay green. Favourable/total is a theorem (equal weights + Axiom 3), never a definition.
Fig. 12. The keystone of the chapter, in three gated beats. Assemble: tap the even faces of a fair die one at a time and P(even) builds itself — 1/6 + 1/6 + 1/6 = 3/6 — with an A3 chip on every click, because disjoint faces add. The gold brackets then name what you actually wrote: the 3 is a count of even faces (|A|), and the 6 is there only because every face weighed the identical 1/6 (so 6 = 1/|Ω|). That is the whole reveal — «favourable over total» is a theorem that falls out of Axiom 3 plus the unspoken equal-weight assumption, not a definition. Commit: before touching anything, predict whether loading the die keeps the ratio right. Reveal: slide P(6) from 1/6 up to 1/3 and the other five rescale to 2/15 apiece — the TRUE probability, recomputed from the real masses, climbs to 0.600, while the formula stays frozen at 0.500 and turns red. Yet the three axiom lamps stay green the entire way and the weights still sum to 1.000. Nothing in the arithmetic complains — which is exactly why a confident wrong answer needs no warning bell. Reset to fair to run it again.

The 3 is nothing but the number of even faces, which is |A|. The 6 appears only because every face happened to carry the identical weight 1/6, which is 1/|Ω|. Write it in general and the collapse is immediate: if all |Ω| outcomes carry the same mass, then P(A) = |A| × (1/|Ω|) = |A|/|Ω|. Favourable over total is additivity, plus one substitution.

That substitution has a name, and it is worth saying out loud, because nobody ever states it. It is equal likelihood, the claim that every outcome in Ω carries exactly the same weight. Not a theorem but a physical claim about the equipment. And equipment can be bent.

So that panel asks you to commit before it reveals anything. Load the die so the six carries 1/3 of the sand and the other five faces share the rest evenly, at 2/15 apiece. Does favourable over total still work? Most people say yes, and the recomputation says otherwise. The true weight of the even faces becomes 2/15 + 2/15 + 1/3 = 0.600, while the formula sits there insisting on 0.500.

Now check the three lamps on that loaded die. Every weight is still non-negative, so Axiom 1 holds. They still total exactly one, so Axiom 2 holds. Disjoint events still add, so Axiom 3 holds. The axioms survive a loaded die without a scratch, and the school formula does not. That is the whole reason we bothered with three rules instead of one fraction.

Which leaves you with a question to ask before you ever write |A|/|Ω| again. Not "how many outcomes are there", but "does this experiment actually hand every outcome the same weight?" Some do, honestly and by construction. Others only look like they do.

Does this Ω earn it? — audit six sample spaces one at a time, then watch the true weights
① one fair die the sample space, as listed: Ω = {1, 2, 3, 4, 5, 6} the true weight of each outcome ? cast your vote to reveal
① vote  ② watch the weights  ③ ▶ next
1 / 6
is every outcome equally likely?
judged 0/6 · 0 right
your call — EQUAL or NOT?
What you’re looking at — the same test every “favourable / total” count quietly assumes.
the gold line is 1/n, every outcome’s fair share — equal likelihood means every bar lands exactly on it
green & flat = the equipment earns it, by symmetry
red & ragged = broken by weighting or by summarising — the urn shows the same experiment can do either, depending only on how you write Ω
Fig. 13. Six experiments, each handed to you with its sample space Ω already written out. Before the answer shows, you vote — is every outcome equally likely? — and the widget then draws the true weight of each outcome as bars against the gold 1/n line, every outcome’s fair share. Flat and on the line means the equipment earned it by symmetry (the die, the shuffled deck); ragged means it was broken — by weighting (the bent coin, the lopsided spinner) or by summarising (the sum of two dice, where seven has six preimages and two has one). The urn is the twist: written as {red, blue} it fails, but hit REPAIR and the same three-red-one-blue urn, relabelled {r1, r2, r3, b1}, snaps flat and earns it. Same experiment, opposite verdict — decided only by how you chose to write Ω. That is the question to ask before ever writing |A|/|Ω| again.

Work through those cases and a pattern shows itself. Equal likelihood is earned by physical symmetry, meaning the outcomes are related by something that genuinely cannot tell them apart, like the faces of a machined cube or the cards in a shuffled deck. It is destroyed by bending, by weighting, and above all by summarising, which is what the eleven dice-sums did to us in the opening.

07★ Counting, done honestly

Suppose the assumption is earned. Then the whole problem changes character, and that is the good news for the rest of Part 1. Computing P(A) becomes a pure counting exercise. And counting you already know, from Chapter 1: the multiplication principle, factorials, combinations.

The failure mode is no longer the arithmetic. It is building the sample space wrong in the first place. Let's go back to the two dice and do it properly, because the classic error lives right there.

Fig 14 · Thirty-six, not twenty-one — a sample space you are allowed to count
d₂ (second die) → d₁ (first die) ↓ ? how many outcomes? rolling two dice 3 / 21 = 0.1429 6 / 36 = 0.1667 how many of the 36 ordered cells feed each surviving atom
Beat 1 of 6 · Commit to a number
Commit first — then we build Ω honestly and check you.
Pick 11, 21 or 36 above ↑
Fold (2,5) onto (5,2) and the arithmetic survives — but the equal weights do not.
the 36 ordered cells (d₁,d₂) — the atoms that really are equally likely
the sum-7 event: 6 cells → 6/36 = 0.1667, true
after folding, a mixed pair weighs 2, a double weighs 1 — the atoms stop matching
3,000 real rolls: even over 36, refuse to be even over 21
Fig. 14. Rolling two dice does not have 21 outcomes — it has 36, once you keep them ordered. Commit to a guess, then watch Ω get built the Chapter-1 way: 6 choices then 6, filling a grid whose every cell is an equally-likely pair (d₁,d₂). Sum-seven lights 6 cells, so P = 6/36 = 0.1667 — pure counting. Now fold the grid on its diagonal so (2,5) lands on (5,2): the count falls to 21 and the arithmetic still runs, but a mixed pair now arrives from two cells while a double arrives from one — the bars go ragged, the atoms stop weighing the same. Apply |A|/|Ω| to the 21 anyway and it prints 3/21 = 0.1429 in red beside the true 6/36 = 0.1667. Roll 3,000 times: the tallies stay even across the 36 and refuse to even out across the 21. The counting was correct; the sample space was not.

Step that grid and the honest Ω appears. Each die independently shows one of six faces, so the multiplication principle gives 6 × 6 = 36 outcomes, laid out as a grid your eye can count. Six of those cells sum to seven, so P(sum = 7) = 6/36 = 0.1667. That is the number the dice actually produce.

Now watch the stepper fold the grid. Treat the dice as indistinguishable, so that (2,5) and (5,2) become one outcome, and Ω shrinks from 36 to 21. Every count in sight is still correct. But the 21 atoms no longer weigh the same, because a mixed pair like {2,5} arrived from two grid cells while a double like {3,3} arrived from one. Apply the formula anyway and a seven comes out as 3/21 = 0.1429, which is wrong.

A fair objection arrives here, and it is a good one. We fold outcomes together all the time in counting problems, and it usually works fine. So when exactly is folding legal? The rule is sharper than most courses admit.

Fold two problems the same way — one collapse is legal, one is a quiet lie
PICK 3 OF 20 · 5 will beat P(all three beat)? SUM OF TWO DICE P(sum = 7)? 5 · 4 · 3 = 60 20 · 19 · 18 = 6840 0.008772 ways to make 7 = 6 ordered outcomes = 36 0.1667
▸ step through the four beats
Both counts honest so far
The same-size test — folding is legal ⟺ every collapsed group has the same size
stocks: every triple folds from exactly 3! = 6 orderings — uniform, so the 6 cancels top and bottom and ordered = unordered to the digit
dice sums: the 7 folds from 6 outcomes, the 2 from just 1 — ragged, no single divisor cancels, so 6/36 ≠ 1/11
run it anywhere: (1) list the groups you collapse (2) count each one's size (3) all equal → fold; ragged → don't
Fig. 15. Two audits in parallel. Collapsing "which came first" is safe only when every group you fold has the same size — then the constant cancels top and bottom. Picking 3 of 20 folds uniformly by 3! = 6, so ordered and unordered agree to the last digit; folding dice outcomes into sums is ragged, so the same shortcut lies with total confidence.

Folding is safe when every group you collapse has the same size. Take the panel's example: 20 stocks, 5 of which will beat the market, and you pick 3 at random. Counting ordered picks gives 20 × 19 × 18 = 6840 in the denominator and 5 × 4 × 3 = 60 on top. Counting unordered hands gives C(20,3) = 1140 and C(5,3) = 10. Both routes return 0.0088, or about a 0.88% chance.

They agree because every unordered triple folds from exactly 3! = 6 orderings, top and bottom alike, so the sixes cancel. The dice sums had no such luck. A seven folds from six ordered cells and a two folds from one, so the groups were ragged and the equal weights died in the fold. That single test, are the groups all the same size, is what separates a clean simplification from a broken sample space.

08Twenty-three people

Time to spend all of it at once. You are in a room and you want the probability that at least two people share a birthday. Assume 365 equally likely days and no twins, which are the honest simplifications we will name and keep. How many people do you need before that probability passes one half?

Commit to a number now, out loud if you can, because the guess is the point. Most people reason that there are 365 days, so you need somewhere near half of them, and they land around 180. The real answer is 23. Before we explain why, here is how you compute it.

Attacking "at least one shared birthday" head-on is a nightmare of overlapping cases, because two people might match, or three, or two separate pairs might match. So flip it. The complement of "at least one shared" is "all distinct", and that one is a single clean count. The first person may have any birthday, so 365/365, and the second must dodge one taken day, so 364/365. The third dodges two, so 363/365, and so on down.

Predict first, then run the code — how many people until a shared birthday beats even odds?
DAYS = 365 p_distinct = 1.0 # everyone still free for n in range(1, 251): p_distinct *= (DAYS - (n-1)) / DAYS p_shared = 1 - p_distinct # 1 - P(all distinct) print(n, round(p_shared, 4)) 1.0 0.5 0 P(shared) 50 100 150 200 250 people (n) → even odds — ½ stdout — the printed evidence: press ▶ Run the code to execute
1 · your guess — n until P(shared) > ½
2 · run it
days in the year (the sample space) — crossing ∝ √days
guessed 183 — press Run to check
assumes: 365 equally likely days · no twins · no leap years
What you're looking at — "at least one shared" solved by its own complement
blue = P(at least one shared) = 1 − (365/365 × 364/365 × …), the shrinking product, executed live
gold = the ½ line and the crossing — 23 people is already better than even odds
grey = your guess on the same axis, so the miss is visible; most guesses land in the flat "basically certain" zone
no new machinery: "all distinct" is one multiplication count over an equally likely space; change days → crossing moves like √days
Fig. 16. The finale, computed honestly. "At least one shared birthday" is a tangle of overlapping cases, so take the complement: P(shared) = 1 − P(all distinct), and P(all distinct) = 365/365 × 364/365 × … — the shrinking product of still-free fractions. Commit a guess for the crossing, then hit ▶ Run and read the stdout verbatim: n=22 → 0.4757, n=23 → 0.5073 straddle ½; it hits 0.7063 by 30 and 0.9901 by 57. Your guess drops onto the same axis as a grey tick, so the miss is plain — most of us aim near 183, half of 365, and land deep in the certain zone. Toggle the days to 100 or 1000 and the crossing slides roughly as √days. No new machinery: 1 − P(none) is the complement rule proved from the axioms, and all distinct is one multiplication-principle count over an equally likely space.

That is real code, executed, and the printed table is its output. At 22 people the chance of a shared birthday is 0.4757. Add one more person and it jumps to 0.5073, crossing the line. Keep going and it climbs fast: 0.7063 at 30 people, 0.9704 at 50, and 0.9901 at 57.

Every ingredient in that computation is something we built on this page. "At least one" became 1 − P(none) by the complement rule from the axioms. "All distinct" was a multiplication-principle count over an equally likely sample space. No new machinery anywhere. That is exactly the claim this chapter has been building toward.

Computing the answer is not the same as understanding it. The shock deserves an explanation, not a shrug. Your intuition missed by so much because you were quietly asking the wrong question: "how many people must there be before someone matches me?" That question really does need about 253 people to reach a half.

Two hundred and fifty-three pairs — drag n, then run the room from scratch
every pair is one chord 23 people 253 pairs how the two counts grow 253 23 people in the room, n → pairs = n(n−1)/2 people = n
seed 42 · press run
simulated share w/ a match
complement rule predicts0.5073
drag n — watch the pairs explode
What you're looking at — people grow in a straight line, their pairs grow in a parabola
blue — the "match me" reading: n people, n comparisons. Nearly flat.
gold — the pairs, C(n,2)=n(n−1)/2. At n=23 that's 253 collisions in waiting.
green — a room built from scratch (uniform birthdays, seed 42) lands on the same ≈0.507 the complement rule got.
Same trap wearing a lab coat: scan 100 strategies for one that "beats the market" and you're testing C(100,2)=4,950 pairs — a lucky coincidence is almost forced. That's the failure Chapter 19 names and corrects.
Fig. 17. Executed with seed 42: the complement rule gives P(match among 23) = 0.5073, and a from-scratch Monte-Carlo of 60,000 random rooms (birthdays drawn uniformly over 365 days) lands on ≈0.507 — two independent roads, one number. Drag n to watch people rise as a line while their pairs, C(n,2), sweep up as a parabola: 23 people already carry 253 pairs, which is why the intuition that "counts people" missed the collision.

The room does not care about you. A collision needs a pair, and 23 people contain C(23,2) = 253 distinct pairs, each one an independent chance to collide. Your gut counted people, which grows linearly. The collisions live among pairs, which grow like n²/2. That mismatch between a line and a parabola is the entire paradox, and once you see it the result stops being a trick and becomes obvious.

This matters well beyond party games. Scan a hundred trading strategies for one that beats the market and you are not running one test. You are running the pair-count's cousin, a hundred chances for a fluke to look like a signal. Chapter 19 gives that failure its proper name and its correction.

The router · pick a probability question you've never seen and watch which gates it must pass through
BUILD Ω outcomes as a set · · EQUALLY LIKELY? the loaded-die reveal · · ‘AT LEAST ONE’? the complement law · · OVERLAP? the legal cut · incl–excl · · ANSWER axioms 1 · 2 · 3 · ·
tap an unseen question ↓
pick a question from the rack →
What you're looking at — the router for a probability question you've never seen
BUILD Ω — write the outcomes as a set
each gate that fires on your question lights up
no symmetry → skip counting, sum the pmf masses
the answer falls out at the bottom
Fig. 18. Pick any question you've never solved and the router lights the exact path it must take. The bent coin skips counting entirely — its ½ was the trap; with no symmetry you read P(H) straight off the pmf. The ‘at least one’ questions flip to the complement first; the OR question subtracts the overlap once. Every gate here is a theorem this chapter proved, standing in the one place where skipping it hands you a confident wrong answer: the equally-likely gate came from the loaded-die reveal, the complement gate from the law rack, the overlap gate from the legal cut. That's the template — carry it out the door.

So here is the chapter as one procedure you can carry out the door. Write down Ω, the set of distinguishable outcomes. Write the event as the subset that makes your sentence true. Ask whether the outcomes really carry equal weight, and only then reach for counting. If the question says "at least one", flip to the complement first. If two events overlap, subtract the lens once.

And hold onto the shape underneath it, because it never changes after this. Probability is one unit of weight spread over a sample space. P is a function that weighs whatever region you circle. Three rules govern that weight: none negative, the whole pile is one, separate piles add. Every law in this chapter came out of those three, including the school formula, which turned out to be a theorem carrying an assumption it never declared.

Which sets up the next question rather neatly. Everything we have weighed so far was measured against the whole of Ω. But information arrives. Someone tells you the roll came out even, and now half the sample space is dead. What happens to the weights that are left?

Shrink the world — tell the die something, watch the surviving sand renormalise to one
P(A) — against all six faces 0.500 P(A | B) — against survivors 0.500 Σ surviving sand = 1.000 Axiom 2 holds the whole way you are told: the roll is at least 4 P(A|B) = |A∩B| / |B| = 2/3 = 0.667 that operation has a name: conditional probability
▸ pick what you're told, then drag the dial
Full Ω · every face weighs 1/6 · total 1.000
Full world — P(A) = 0.500
Re-weighing A inside the smaller world IS conditional probability — the axioms force it
A = even = {2,4,6} — the event we keep re-weighing; its blue faces never change, only the denominator does
B = the survivors — the faces your fact leaves alive; their sand grows so the total returns to 1.000, because Axiom 2 demands it
A∩B, gold-ringed — even faces that survive: the numerator. P(A|B) = |A∩B| / |B|
swept away — faces outside B; their weight drains to zero and pours into the survivors
Fig. 19. Every weight so far was measured against the whole of Ω. Now a fact arrives, part of the sample space dies, and the surviving sand must renormalise to one — watch the total stay pinned at 1.000 the entire drag. Re-weigh A inside that smaller world and 0.500 becomes 0.667: for "at least 4", two of the three survivors are even. Writing it as a fraction, |A∩B| over |B|, is the whole of the next chapter's definition — arrived at with your own hand before it is named. This is P(A|B), conditional probability, and nothing new was needed to ask for it.

Drag that panel and you can watch the answer take shape. The surviving region gets renormalised so its weight is one again, because the axioms demand it, and every event inside is re-measured against the smaller world. That operation is conditional probability, and Chapter 10 is where it stops being a picture and becomes the sharpest tool in the course. Run it backwards and it is Bayes' theorem.

iolinked.com
Written by Ajai Raj