◈ quant roadmapPart 1 · Ch 13/45
Quantitative Finance — the Mathematics of Markets · chapter 13

13Symmetry, Order Statistics, the LLN & the CLT

Chapter 12 grew a whole family of distributions out of a single coin flip, and every branch arrived with a formula you could sit down and evaluate. That is a real skill, and it has a ceiling. Two kinds of question walk straight past it. The first is the tangled one, where the pieces of a problem lean on each other so hard that no clean pmf is available. The second is worse, because nobody hands you the shape at all and you are still expected to say how wrong your answer is. This chapter builds the two engines that handle both. The first engine is symmetry, and it says that interchangeable pieces sharing a known total must each get the same share, with no independence assumed anywhere. The second engine is averaging, and it says that independent pieces, once averaged, are forced onto the mean at a rate you can compute, with a leftover wobble that can only be one shape. Between them they answer the questions where computing is either impossible or beside the point. And they leave you holding one small expression, σ/√n, that quietly prices every research decision a quant will ever make.

Look at what this page is standing on. Chapter 11 handed us three tools that carry most of the weight here: E[X] as a probability-weighted average, the identity Var = E[X²] − E[X]², and linearity of expectation, which splits a sum without ever asking how the pieces are related. Chapter 12 gave us the Normal, the Uniform, and one unpaid debt. Chapter 9 gave us the axioms, and the third one, that probabilities of disjoint events add, does more work on this page than anywhere else in the course.

The seam — Chapter 12 handed you a dozen shapes; press AVERAGE THEM and watch how many survive.
CH 12 · the parents CH 13 · what is left average Binomial Poisson Uniform Exponential lopsided μ σ one shape, whatever went in ▶ press AVERAGE THEM PARENTS IN 5 SHAPES OUT ? ▲ you are here · Ch 13 Ch 16 Ch 17 Ch 18 Ch 31 Ch 34 tap a Ch chip → what it takes from this page

Predict first. Five different parents go in. How many different shapes come out?

Two questions the tree above cannot answer — tap a dark card.
guess, then press AVERAGE
What you're looking at — five parents in, one shape out, and where that leaves you
blue = Ch 12's parents (Binomial, Poisson, Uniform, Exponential, a lopsided pmf). Nothing here is bell-shaped.
gold = the one Normal they all land on. Only μ (the centre) and σ (the spread) survive — the parent's shape is forgotten.
violet = the two questions the tree can't answer, which is why this chapter also teaches symmetry.
Fig. 1. Chapter 12 handed you a shelf of shapes. Average enough of any of them and the shelf collapses to one curve — only a centre and a spread survive. That collapse is this chapter; the two dark cards are the questions the shelf never could answer.

That is the seam. Chapter 12 kept saying that sums of many independent things go Normal, and you watched it happen without ever being told why it must. This chapter pays that debt, states the conditions honestly, and then shows you the case where the machine breaks.

01The answer you refuse to compute

I want to start with a question, and I want you to commit to a number before reading on. No calculating and no hedging, just a number you would defend.

A plane has 100 seats and 100 passengers, and each passenger holds a ticket for one specific seat. The first passenger has lost his boarding pass, so he picks a seat completely at random and sits down. Every passenger after him takes his own seat if it is free. If it is taken, he picks uniformly at random from whatever seats remain. What is the probability that the 100th passenger ends up in his own seat?

Most people say 1 in 100, and a few insist it must obviously depend on how big the plane is. Both answers feel reasonable and both are wrong.

Commit a bet you will not get to check — then look at the two roads that lead to the number
1 · THE PLANE · 100 seats passenger 1 lost their boarding pass → takes a RANDOM seat. everyone after: own seat if free, else a random free one. P( passenger 100seat 100 ) = ? — no bet placed yet — pick one on the right, then SEAL IT no answer appears here — by design. 2 · TWO ROUTES TO ONE NUMBER ROUTE A · COMPUTE depth 0 / 100 0 paths ROUTE B · CONSTRAIN ? · undecided a property of the whole still blank so — is Route B an approximation? AHA · every part stays unknown — the whole still forces one exact number. SEAL A BET TO UNLOCK commit first · that is the rule
1 · place your bet — required
pick a bet · nothing revealed here
Your guess is sealed, not marked. It gets opened in §13.1.
What you're looking at — a question you have committed to, and the two roads that reach it
blue = passenger 100 and seat 100 — the only thing the question asks about
red = passenger 1's lost pass, and Route A: the randomness it injects is exactly what makes the branch tree explode
violet = Route B — one padlock, one constraint on the whole, no arithmetic
gold = your sealed bet — stamped, and opened in §13.1, not on this page

Route A is honest and hopeless: unfold it and the path count runs past a hundred billion before you are six passengers in, and the real tree is a hundred deep. Route B never touches a single branch — it puts one constraint on the whole arrangement and the number is forced. That is why "answer it without computing" is a method, not a dodge.

Fig. 2. Before anything else, commit. A plane, 100 seats, 100 passengers; passenger 1 has lost their boarding pass and sits at random, and everyone after takes their own seat if it is still free. What is the chance passenger 100 ends up in seat 100? Pick 1/100, 1/2, "depends on n", or dial your own — then SEAL IT. Your guess goes into an envelope stamped opens in §13.1; nothing is revealed here, and that is deliberate. Sealing unlocks the stance card: two roads to the same number. Route A computes — unfold it and watch the branch tree grey into a thicket while the path counter climbs past 858 billion by passenger six, with 94 more passengers still to go. Route B does not compute at all: one padlock, a property of the whole, and its output slot stays blank on this page. Press IS THIS APPROXIMATION? and the badge flips: the honest answer is EXACT. That is the whole stance — a constraint on the whole can force an exact number while every single part stays unknown.

The answer is exactly 1/2, and it does not depend on the 100 at all. Change the plane to 3 seats or to 3 million and the answer is still one half, and I am not going to prove that yet.

Here is what I want you to sit with instead. You have spent four chapters learning one model of how a probability gets computed. Build the sample space, count the favourable outcomes, divide. Or write the pmf and grind out the sum. Both are honest, and both are the wrong tool here, because the branching in this problem is genuinely horrible.

There is another way to pin a number, and it is not approximation. You find a constraint on the whole thing that forces the answer, and then you never look at the parts at all. The result is exact — there is no arithmetic anywhere in it, and that combination is what makes people distrust it the first time they meet it.

This chapter builds two engines of that kind. The first is symmetry, which asks what interchangeable pieces are forced to share. The second is averaging, which asks what many independent pieces are forced to become. The plane is now a debt we owe you, and it gets paid before this chapter is half done.

02★ The licence: exchangeability

Every source I have read writes the phrase "by symmetry" and then moves on. That phrase is doing an enormous amount of work, and what it hides is the condition actually being used. Let's dig it out.

Take the cleanest possible setup: a deck of two cards, one ace and one blank, shuffled so that either order is equally likely. Ask for the chance that the first card is the ace and the answer is plainly 0.5000. Ask for the chance that the second card is the ace and it is 0.5000 as well, because relabelling which card you look at first cannot change anything.

Now push on it with Chapter 10's conditioning. Given that the first card is the ace, what is the chance the second card is the ace?

Two 2×2 joint tables with the same edge totals — mirror the labels, then condition on card 1.
REAL DECK  2 cards, 1 ace card 1 card 2 ace blank sum ace blank 0.0000 0.5000 0.5000 0.0000 0.5000 0.5000 sum 0.5000 0.5000 1.000 ruled out P(2nd = ace | 1st = ace) ? FAKE  independent, same edges card 1 card 2 ace blank sum ace blank 0.2500 0.2500 0.2500 0.2500 0.5000 0.5000 sum 0.5000 0.5000 1.000 ruled out P(2nd = ace | 1st = ace) ? swapping labels can't tell them apart — conditioning can
marginals: same ✓
joint: not yet tested
predict, then CONDITION
What you're looking at — one pair of marginals sitting on top of two different joints
REAL deck (2 cards, 1 ace): all the mass sits off the diagonal — two cards can't both be the ace.
FAKE: an independent pair built on purpose to have the very same edges — 0.2500 in every cell.
the marginals (edge sums): 0.5000 everywhere, identical in both, before and after the mirror.
the conditional P(2nd = ace | 1st = ace): 0.0000 real vs 0.5000 fake — that's the joint talking.
Fig. 3. Two cards, one ace. Mirroring the labels leaves both tables exactly where they were — so the swap test cannot tell them apart, and both decks are exchangeable. Conditioning can: rule out the row where card 1 is blank, renormalise the survivors, and the real deck's P(2nd = ace | 1st = ace) collapses to 0.0000 while the independent fake sits at 0.5000. Identical marginals, opposite joints — exchangeable is not independent.

It is 0.0000. There is only one ace in the deck, so knowing where that ace landed tells you everything about the other card. Two cards, identical marginal probabilities, and about as much dependence as it is possible to have.

So here is the sentence that has to be said out loud, because the whole first half of this chapter rests on it. Exchangeable means the pieces are interchangeable. It does not mean they ignore each other.

The formal statement is short: a collection of random variables is exchangeable when relabelling them leaves the joint law unchanged, so you can swap any two names and every joint probability comes out identical. The joint law is the full description of how the pieces move together, not just how each one behaves alone.

That distinction is the tear worth guarding for the rest of the course. The marginal is what one piece looks like on its own, and the joint is how they move together. Equal marginals can sit on top of completely different joints, which is exactly what the two tables in that figure show, and nothing in Chapters 9 through 12 ever forced you to hold them apart. Almost every distribution there was assembled out of independent atoms.

One more thing to keep straight, since the words all sound alike. Independent is a strong condition, and it implies exchangeable whenever the pieces also share a distribution. Exchangeable is far weaker and far more common, which is precisely why it is worth having as a separate idea. Cards from a shuffled deck, balls drawn without replacement, and the pieces of a broken stick are all exchangeable, and not one of them is independent.

03★ Equal shares from a fixed total

Now the payoff for that care: there is one lemma underneath every symmetry argument in this chapter, and it is three lines long. Almost nobody writes those lines down, which is why "by symmetry, each is T/n" feels like a convention you are expected to nod at.

You have n pieces X₁, …, Xₙ which are exchangeable, and they are known to add up to a fixed total T. Fixed means a genuine constant here, the same number in every outcome, not a value that merely happens to be the average.

Four aces cut a 52-card deck into five gaps holding 48 non-aces between them. Walk the three lines that force every gap to average T/n — and then try to break them.
48 non-aces, cut into 5 gaps by 4 aces. T = 48 · CONSTANT TOTAL 48 non-ace cards T/n = 9.6000 Σ Lᵢ = 48.00 linearity: still valid ✓ E[Lᵢ] = T/n = 48/5 = 9.6000 lines ② and ③ do the work; ① supplies the number. the gold frame is T — watch it refuse to move.
beat 1 — one fixed total. Press ② when you are ready.
① T = 48 · CONSTANTthe fixed thing shares add to
② linearity · ΣE[Lᵢ]=E[T]splits it · ignores dependence
③ symmetry · all E[Lᵢ] equalrelabelling can't change it
T = 48. Fixed. That is line one.
What you're looking at — the two ingredients of "by symmetry", separated.
The gold frame is T, the total the five gaps must add to. Drag the dial: the blue pieces lurch, the frame never twitches — that is line ① plus ②.
Linearity is blind to dependence, so the badge stays green at every setting. It splits the total but says nothing about which gap is which.
Only ① can break. Make T random and the exact 9.6000 downgrades to E[T]/n — you now owe a second calculation before you may speak.
Fig. 4. "By symmetry, each gap averages T/n" is three lines of work compressed into two words — so here they are, one beat at a time. Four aces cut the deck into five gaps holding T = 48 non-aces between them; that total is a constant, and the gold frame is it. Split it, then drag the dependence dial: the blue pieces lurch wildly — one gap swallowing the deck while the rest starve — yet Σ Lᵢ = 48.00 at every setting and the badge never leaves green. Linearity never asks how the pieces relate. Hold RELABEL: the same five lengths get shuffled between rows, and the five running averages — scattered at first — slide onto one gold line at 9.6000. Nothing was computed; relabelling simply cannot favour a row. Two lines bought you E[Lᵢ] = T/n = 48/5 = 9.6, and the first ace therefore sits at 9.6 + 1 = 10.6 = (n+1)/(k+1). Now press BREAK IT: make T random and line ① dies alone — ② and ③ stand, but the crisp 9.6000 downgrades to E[T]/n, a promise you can only keep after a second calculation. That is precisely what "by symmetry" spends.

Line one: since T = X₁ + … + Xₙ is a constant, its expectation is itself, so E[T] = T. Nothing has happened yet.

Line two, and this is the line doing the heavy lifting. Linearity of expectation splits that sum into E[X₁] + … + E[Xₙ] = T. Stop and notice what was not required. Chapter 11 proved linearity with no independence assumption anywhere, so it does not care that these pieces may be violently dependent, and that is the whole power of the lemma.

Line three: exchangeability says all n of those expectations are the same number, and n equal numbers summing to T are each T/n. That is arithmetic, and the lemma is finished.

Read the receipt carefully, because knowing which line paid for what is the difference between owning a method and repeating a trick. Linearity did the splitting and is blind to dependence, while exchangeability did the equalising. Neither one needed a distribution and neither one needed independence.

The condition that can break is the fixed total. If T is itself random, line one collapses and you get E[Xᵢ] = E[T]/n instead, which is a weaker statement about an average rather than about a constant you actually know. That is the failure mode to watch for when you aim this at a new problem.

One fixed total, n interchangeable parts — the same three lines settle resistors, trading desks and deck gaps, and never once mention what the parts are made of.
four identical resistors, one supply TOTAL current T = 2.4 A · fixed today's split T/n adds to 2.400 A — exactly T ✓ day 1 · cyan tick = running avg THE SAME THREE LINES 1 · the four resistors are identical — swap any two, the circuit is unchanged. 2 · so none can expect more current than another: E[I₁] = … = E[I₄] = m 3 · and the four currents must add to the supply: 4m = 2.4 A m = 2.4 / 4 = 0.600 A 🔒 ALL IT USED: a conserved total, and parts you can swap. Nothing else.
1 · pick a world   2 · step the three lines   3 · crank the noise
each resistor gets 0.600 A
What you're looking at — one argument, wearing three costumes. T is the fixed total; n is how many parts share it.
Blue is what actually happened today — the split of T across the n parts. Crank the noise and the bars go wild, but the top bar stays exactly full: the total is conserved.
Gold is each part's expectation m = T/n. It never budges — the three lines fix it before a single day is drawn, and they never ask whether the parts are amps, lots or cards.
Cyan is each part's running average over the days you press. It wanders at first, then crawls onto the gold line — the same lemma that puts the first ace at 9.6 + 1 = 10.6.
Fig. 5. Same argument, three costumes. Pick a world, then step the panel on the right: (1) the parts are interchangeable, (2) so none can expect more than another, (3) and their shares must add to the fixed total T — so each one is worth m = T/n. Nothing in those lines knows whether the parts carry amps, lots or cards, which is exactly why the lemma transfers. Now crank the noise: today's blue split scatters as wildly as you like, the top bar stays exactly full, and the gold expectation marker does not move a pixel. Press ↻ new day a few times and the cyan running averages crawl onto it — the same 9.6-blanks-per-gap that puts the first ace at position 10.6.

The same three lines work anywhere the two hypotheses hold, and that is worth seeing outside probability once. Identical resistors across one supply split a fixed current equally. Symmetric desks running the same strategy split a fixed daily volume equally. The argument never mentions what the pieces are made of.

04Symmetry on events: the plane and the circle

So far we ran symmetry on expectations, and the reason it worked is that expectations add. Chapter 9 handed us another quantity that adds, which is the probability of disjoint events. Run the identical argument with events in place of numbers. The move itself does not change at all.

The event version. If a relabelling maps the whole setup onto itself, any two events it swaps must have equal probability. And if those equal events are disjoint and cover everything between them, each one's probability falls out by division.

Now the plane, and we are going to attack it the cowardly way. Shrink it until the answer is undeniable, then hunt for the thing that did not depend on the size.

Shrink the plane until the answer is undeniable — draw every branch for n = 2, 3, 4, 5 and watch 1/2 refuse to move
EVERY BRANCH, BY HAND n = 3 · 4 branches THE ENVELOPE OPENS 2 / 4 = 1/2 · exactly winning leaf weights: 1/3 + 1/6 = 1/2 2 nodes · one ✓ + one ✗ at every single one WALK · which seat first? seat 1: 0 seat 3: 0 no runs yet AHA · seat 1 & seat n are BOTH on offer at every node, and taking either one ENDS it.
4 branches · 2 win → 1/2
Drag the dial → through 2, 3, 4, 5 and count the branches by hand. Then press WALK IT, or click any leaf chip to trace one branch.
What you're looking at — the whole plane, small enough to count, and the one thing that never changes
green leaf = the last passenger does get their own seat
red leaf = seat n was taken first — the story ends the other way
blue = seat 1, the other exit; MARK shows every node offers both
gold = the count, favourable / total — and it is 1/2 at every n

The branches are not equally likely — look at the weights, 1/3 next to 1/6. It does not matter. Every fork puts seat 1 and seat n side by side in the same list, and whichever is taken, the process stops. So the leaves come in twins that sit at the same node with the same weight, one green, one red. Half the weight each. That is why no 100 ever appears in the answer.

Fig. 6. The envelope from the stance panel opens here — by shrinking the plane until you can count the whole thing. Drag the dial to n = 2, 3, 4, 5 and the widget draws every branch: each fork is one displaced passenger choosing among the free seats, each chip is the seat they took, green ✓ when the last passenger still gets their own seat and red ✗ when they do not. Count them by hand: 1/2, 2/4, 4/8, 8/16 — the fraction never budges. The branches are not equally likely (read the weights: 1/3 sits next to 1/6), so press MARK THE TWO SEATS and the reason appears: at every single node, seat 1 and seat n are both on the menu, and taking either one ends the story. The leaves therefore come in twins hanging off the same node with the same weight — one green, one red — so the two halves must weigh the same. Push the dial past 5 and the tree refuses to draw: 32 branches is already unreadable, and n = 100 would need 299 of them — which is precisely why you must not compute. WALK IT runs one honest realisation as a chain of choices, RUN ×100 piles them up, and the tally bar settles on the dashed 50% line. Click any leaf chip to trace its branch and read its exact weight.

With 2 seats it is a coin flip, so the answer is 1/2. With 3 seats you can enumerate the whole thing by hand in about a minute. Passenger 1 takes his own seat with chance 1/3 and everything then runs clean. He takes seat 3 with chance 1/3 and passenger 3 is sunk. He takes seat 2 with chance 1/3, and passenger 2 then picks between seats 1 and 3 with chance 1/2 each. Adding the two winning roads gives 1/3 + 1/6 = 0.5000.

At that point you know it is not a coincidence, and you can stop hunting for the 100. So what is the invariant?

Watch any displaced passenger make a choice. From his point of view, seat 1 and the last seat are perfectly interchangeable. Both are simply free seats he has no claim on, and nothing in the rules ever treats one of them differently from the other. The whole process ends the instant one of those two is taken, because taking seat 1 unblocks everybody afterwards and taking the last seat ends the story. Two interchangeable seats, and exactly one of them gets hit first. That is a coin flip, and it is why the 100 never appears in the answer.

Now the second classic, which leans on the disjointness half of the move. Drop 3 points independently and uniformly on a circle. What is the chance that all three lie within some single semicircle?

The trouble with that sentence is the word "some". A floating semicircle is genuinely ambiguous, and people who try to count it directly end up double-counting configurations. The fix is to anchor it to one of the points.

Three dots on a circle. Sweep a half-circle by hand — then let every dot grow its own half-circle.
FLOATING — sweep it by hand drag any dot · drag the gold handle ONE EVENT, MANY WITNESSES every possible placement 360° placements that work: 40° of 360° you have found 1 that works a whole ARC of them works — so which one would you count? P(Aₓ) = (1/2)² = 0.2500 each 3 × 0.2500 = 0.7500 P(all 3 in ONE half-circle) they overlap ⇒ you may NOT add
the half-circle is…
← press ANCHORED next
sweep it: φ = 200° floating mode only
dots on the circle: n = 3 n / 2ⁿ⁻¹ is the answer
covers! — but so do many others
What you're looking at — one vague event, or n sharp ones
Anchored: dot i owns the half-circle that starts at it and runs forward. Event Aₓ = “everyone else is inside mine.” P(Aₓ) = (1/2)n−1, since each other dot lands in my half with probability 1/2, independently.
At most one chip is ever green — only the dot at the start of the cluster can see all the others ahead of it. Disjoint, so Axiom 3 says add: n × (1/2)n−1. That is the step no textbook writes down.
Floating, an arc of placements works, so counting witnesses double-counts one event. Complement = the dots surround the centre. (Two dots exactly coincident fire two anchors — that costs zero probability.)
Fig. 7. Drag the gold handle and the half-circle covers the three dots from many different starting angles — a whole arc of witnesses for a single event, which is exactly why you cannot count them. Press ANCHORED and the ambiguity dies: each dot claims the half-circle running forward from itself, and only the dot at the back of the cluster can see the other two ahead of it — so however you drag, at most one chip is green. Three disjoint events, each with probability (1/2)2 = 0.2500, and only now is 3 × 0.2500 = 0.7500 an application of Axiom 3 rather than a card trick. Push n to 5 and the same argument gives n/2n−1 = 0.3125.

For each point, define the event that the other two both lie in the half-circle running clockwise from it. Each of those events has probability (1/2)² = 0.2500, since the two other points are independent and each has an even chance of landing in that half.

Here is the step no textbook writes down, and it is the one that turns this into a theorem instead of a card trick. Those three anchored events are disjoint. Only one point can be the clockwise-most of a clustered trio, so two of the events cannot both hold. And if the three points do fit inside a semicircle, one of them must hold. Axiom 3 then adds them: 3 × 0.25 = 0.7500.

One honesty note, in plain words. Configurations where two points land at exactly the same spot would break the "only one can be clockwise-most" claim. For continuous points those have probability zero, so they cost us nothing. The general answer, by the identical argument with n anchors, is n/2ⁿ⁻¹.

Flip the question over and you get the version a trader actually cares about. The chance that the three points surround the centre is 1 − 0.75 = 0.2500, which is the same question as whether three independent bets leave you exposed in every direction at once.

05Where symmetry pays: gaps and order statistics

Now we point the lemma at something that looks like it needs real machinery. Shuffle a standard 52-card deck, and ask how far down you have to go on average before you hit the first ace.

The instinct is to write the distribution of that position and sum it out. You can, and it is unpleasant. Instead, look at the deck the other way around.

There are 4 aces and 48 other cards, so forget the aces for a second and lay the 48 blanks out in a row. The four aces then sit in the spaces between those blanks. Count the spaces, not the cards.

48 blank cards, 4 aces you can drag — the aces cut the row into five gaps, and the two at the ends are the ones the eye forgets.
k = 4 aces → gaps = 5 Σ = 48 ✓ GAPS 10 · 9 · 10 · 9 · 10 drag an ace ◀▶ DEPENDENCE · the same five sizes as bars, one fixed total Σ is nailed to 48 — stretch one gap and another must give way. That’s where the +1 lives. Before the first ace and after the last: real gaps, even at 0. 2 cuts → 3 pieces 2 of the 3 are ends
law unchanged ✓
or drag an ace along the row
5 gaps · Σ 48 · 0 empty
What you’re looking at — 4 cuts in a row of 48, and the 5 pieces they leave behind
the aces — the k markers you drag. They are the cuts.
the two end gaps: before the first ace, after the last. Push an ace to the edge and one hits 0 — still a gap, still counted. That is the +1.
the interior gaps, alternating tints so each piece reads separately.
Σ = 48 always — the five sizes are dependent, and the lemma doesn’t care: each has the same expected size 48/5 = 9.6, so E[first ace] = 9.6 + 1 = 10.6 = 53/5 = (n+1)/(k+1).
Fig. 8. Four aces dropped into a row of 48 blanks leave five gaps, not four — and the fifth is not hiding anywhere clever. It is the pair at the ends: the run before the first ace and the run after the last one. Drag an ace to the very front and that end gap falls to 0, which is exactly when the eye stops counting it — so the widget keeps its tick and its 0 badge, and the count stays at 5 while the sizes churn. MIRROR reverses the row and RELABEL permutes the aces’ names; neither can change the joint law, so the five gaps are exchangeable and each has the same expected size, 48/5 = 9.6. The strip below shows those five sizes as bars that always total 48 — visibly dependent, and the argument does not care. Hence E[position of the first ace] = 9.6 + 1 = 10.6, i.e. (n+1)/(k+1) = 53/5, with no series summed.

Count them carefully, because this is where the classic off-by-one lives. Four aces cut the row of blanks into five gaps, not four. There is a gap before the first ace and a gap after the last one, and both of those end gaps are frequently empty. Empty is exactly why the eye skips them. In general, k markers make k+1 gaps.

Those five gaps are exchangeable: relabel which ace is which, or mirror the whole deck end to end, and the joint law of the gap sizes does not move. They are also obviously dependent, since the five of them must add to exactly 48. That dependence is the precise thing the last section built the lemma to survive.

So apply it unchanged — five exchangeable pieces sharing a fixed total of 48 give E[gap] = 48/5 = 9.6000 blanks each.

The first ace sits one card past the first gap, so its expected position is 9.6 + 1 = 10.6000. Write that in general and it is (N−k)/(k+1) + 1, which tidies to (N+1)/(k+1), and for the deck that is 53/5 = 10.6.

Now separate the two +1s in that formula, because memorised together they mean nothing. The +1 in the denominator is the number of gaps, and it came straight from the picture. The +1 in the numerator is the ace itself, sitting one card past its gap. Completely different origins.

Take a sample and sort it. The sorted values are new random variables called the order statistics, written X₍₁₎ ≤ X₍₂₎ ≤ … ≤ X₍ₙ₎, and the first ace's position was one of them. Those parentheses matter. X₍₃₎ is the third smallest value, while X₃ is the third value you happened to draw, so the subscript has quietly changed meaning from identity to rank.

And they really are random. Sorting feels like a tidying-up operation that removes randomness, which is a comfortable picture and a wrong one. Shuffle again and X₍₁₎ lands somewhere else.

100,000 real shuffles — guess where the first ace lands, then read the actual stdout
python first_ace.py · real stdout the answer, split in two E = 52 + 1 4 + 1 = 10.60 9.60 blanks in gap 1 48 blanks ÷ 5 gaps + 1 the ace itself one step past gap 1 10.60 = (N+1)/(k+1) = 53 ÷ 5 100,000 real shuffles measured — predict to run — ▲ this is the deck it ran on: 52 / 4 ? 13  ·  10.6  ·  9.6
Guess where the first ace lands:
▶ tap one — then the code runs
# this really ran ↓
random.seed(20260813)
for _ in range(100_000):
  shuffle(deck)
  pos.append(first_ace(deck))
print(mean(pos))  # 10.6042
predict, then the code runs
What you're looking at — the two +1s in (N+1)/(k+1) are doing two different jobs
the gap — the +1 underneath: k aces cut the deck into k+1 gaps, not k. So 48 blanks share 5 gaps → 9.60 blanks expected before the first ace.
the ace — the +1 on top: N+1 = (N−k) + (k+1), so dividing leaves a whole +1. We want a position, not a count of blanks — step one card past the gap.
10.60 = 9.60 + 1, and 100,000 executed shuffles measured 10.6042.

Nothing here was computed. The four aces are exchangeable with each other and with the 48 blanks — relabel them and the shuffle looks identical — so the five gaps they carve out must have the same expected length. Five equal gaps sharing 48 blanks: 9.6 each. Then take one more card and you're standing on the ace. That's the whole derivation, and it's why the aces land on an even ladder 10.6 · 21.2 · 31.8 · 42.4 — press all 4 and read the measured column.

Fig. 9. A genuine CodeRun. The left panel is the real stdout of a Python script that seeds its generator at 20260813, shuffles a 52-card deck 100,000 times, and records where the first ace falls: the empirical mean is 10.6042 against the closed form (52+1)/(4+1) = 10.60 — a difference of +0.0042 — while the naive “a quarter of 52 = 13” misses by 2.4 cards. all 4 shows the same run's means for every ace (10.6042, 21.2454, 31.8280, 42.4019 against 10.60, 21.20, 31.80, 42.40; worst gap 0.0454), and gaps shows the five stretches of blanks the aces carve out (9.6042, 9.6412, 9.5825, 9.5739, 9.5981, summing to exactly 48) — the measurement of exchangeability itself. The right panel splits the answer into its two pieces and colours each +1 in (N+1)/(k+1) to the piece it pays for: blue below, because k aces make k+1 gaps; cyan above, because a position is one card past the gap. The N and k dials re-derive the formula for any deck, while the executed numbers stay honestly pinned to the 52-card, 4-ace deck the code actually ran on.

That figure is a genuinely executed run rather than a drawing. It shuffles 100,000 decks and reports the mean position of the first ace, and the empirical answer lands on 10.6. The formula stops being a claim and becomes a measurement you took.

The continuous version is the same picture with the dots replaced by a stick. Drop k points uniformly on the interval from 0 to 1. They cut it into k+1 pieces, the pieces are exchangeable, and they share a fixed total length of 1.

Same argument, dots removed: 4 random points cut a stick into 5 pieces — step the real run and watch the five pieces become equal.
order_stats.py — real stdout one real sample · k = 4 5 pieces, total = 1.0000 running mean of the 5 gaps longest − shortest = 0.1696 analytic: k+1 equal pieces k = 4 → 5 pieces of 0.2000 exactly the run above ✓ k points → k+1 pieces each 1/(k+1) → U(j)=j/(k+1)
10 samples — nothing is even
N = 10 · next → 100
import random
k, N = 4, 100_000
s = [0.0]*k
for _ in range(N):
  u = sorted(random.random()
             for _ in range(k))
  for j in range(k): s[j] += u[j]
What you're looking at — the deck argument with the dots taken out
k = 4 random points land on a stick of length 1 and cut it into k+1 = 5 pieces. One real draw, drawn to scale — the pieces are wildly uneven, but they always total exactly 1.0000.
The running mean of those 5 gap lengths over the executed run. At N = 10 it is ragged; press ×10 and the raggedness dies — the five gaps are exchangeable (relabelling them cannot change the answer), so they must share the 1 equally.
The forced answer: each piece averages 1/(k+1), so the j-th smallest point sits at E[U(j)] = j/(k+1). Evenly spaced — and now with a reason under it, not a hunch. Drag k to regenerate the ladder; the printed numbers stay tied to the k = 4 run that was actually executed.
Fig. 10. The continuous twin of the deck argument, executed rather than asserted. Real Python draws k = 4 independent uniform points on [0, 1], sorts them into the order statistics U(1) < U(2) < U(3) < U(4), and averages each one over 100,000 samples (seed 20260813); the stdout on the left is that run's genuine output, stepped through the checkpoints N = 10, 100, 1,000, 10,000, 100,000 — the same run read at five moments, not five runs. The four empirical means close on 0.2000, 0.4000, 0.6000, 0.8000, and the five gap means all close on 0.2000. No integral was computed: the k points cut the stick into k+1 pieces whose joint law is unchanged by permuting them, so no piece can be longer on average than any other, and since the pieces always total exactly 1 each must average 1/(k+1). Sum the first j of them and you get E[U(j)] = j/(k+1) — the even spacing your gut wanted, now forced. The k slider regenerates that analytic ladder for k = 2 to 8; the printed numbers deliberately stay pinned to the k = 4 case that was actually run, so the executed claim and the general claim never get confused for one another.

So each piece has expected length 1/(k+1), and the j-th smallest point sits at E[U₍ⱼ₎] = j/(k+1). With k = 4 that gives 0.2000, 0.4000, 0.6000 and 0.8000, which is evenly spaced. Your gut wanted that answer, and now it has a reason for it.

06Counting the stories: a bound for free

Before we leave the first engine, one more move in the same spirit. Constrain the whole thing again, but with a different conserved quantity. Not a total this time, but a count of outcomes.

The classic setup: you have 12 balls that look identical, and exactly one of them is a different weight. Heavier or lighter, and you do not know which. You have a balance scale and 3 weighings, and you must find the odd ball and say whether it is heavy or light.

Almost everybody starts hunting for a clever first weighing. That search has no stopping rule, so before searching, ask a completely different question. Could three weighings ever be enough?

Before you design a single weighing: count the answers you can possibly give, and count the things you must tell apart.
THE BUDGET · answers THE LEDGER · states tap a ball ▸ 3¹ = 3 answers 12 balls × 2 = 24 states BUDGET STATES 3¹ = 3 24 one weighing gives 3 answers — 24 states cannot fit 3 answers cannot name 24 states — push w to 3
1 · drag w — the tree fills   2 · at w = 3 tune k   3 · tap a ball: it is two states
3 answers < 24 — IMPOSSIBLE
What you're looking at — a budget of answers weighed against a ledger of states, before any weighing is designed.
The budget. One weighing has exactly 3 outcomes — left tips, balances, right tips — so w weighings end in at most 3w distinguishable ways: 3, 9, 27. That is the tree.
The ledger. Any one of the 12 balls could be odd, and could be heavy (solid half) or light (pale half): 12 × 2 = 24 things a strategy must tell apart.
States past the cut can never be named, so 3 and 9 fail outright. At 27 the slack is only 3 — no weighing may waste an outcome, and that alone forces k = 4. Switch to 13 balls: 26 ≤ 27 goes green, yet no strategy exists. A counting bound proves impossibility, never possibility.
Fig. 11. Twelve balls, one is odd, and you may weigh three times — but before designing a weighing, count. A weighing has three outcomes, so w weighings can end in at most 3w distinguishable ways: drag the dial and the tree fills to 3, 9, 27. Against it stands the ledger: each ball could be odd and could be heavy or light — tap one and it splits — so 24 states must be told apart. At two weighings the bar runs red past the cut: 9 < 24, PROVABLY IMPOSSIBLE, and no cleverness can rescue it. At three the budget is 27 with a slack of just 3 — so no weighing may waste an outcome. Tune k, the size of the first weighing: a tip leaves 2k states, a balance leaves 24 − 4k, and both must fit under the remaining ceiling of 9. Only k = 4 does. The bound designed the opening for you. Then press 13 balls: 26 ≤ 27, the budget goes green — and still no strategy exists. Counting proves impossibility, never possibility.

Count both sides of the ledger. Each weighing has three possible results, which are left heavy, right heavy, and balanced, so three weighings produce at most 3³ = 27 distinguishable outcomes. Meanwhile the number of states you must tell apart is 12 balls times 2 possibilities each, which is 24. That is the whole ledger: 24 states against 27 answers.

Two weighings would give only 3² = 9 outcomes against those 24 states, so two are provably impossible. No cleverness rescues them, because a procedure that cannot produce 24 different answers cannot always name the right one.

Now use how narrow that margin is, because there are only 3 spare outcomes. Almost no weighing can afford to waste one, and that alone forces the famous opening. Weigh k against k. If it tips, the live states drop to 2k. If it balances, they drop to 2(12 − 2k). Each branch has two weighings left, so each branch must fit inside 9.

Those two requirements are 2k ≤ 9 and 24 − 4k ≤ 9, giving k ≤ 4 and k ≥ 3.75. So k = 4 exactly. The four-against-four opening is not a flash of insight at all, it is the only number the budget permits.

And now the honest half, which matters far more than the puzzle. A counting bound proves impossibility. It never proves possibility. Having found that 24 fits inside 27, you have built nothing, and you still owe an actual strategy.

The case that settles it is 13 balls. The budget says 26 ≤ 27, so the count is perfectly happy. No strategy exists that both finds the odd ball and names it heavy or light. Room in the budget is a permit, not a construction, and reading it as one is a habit that will misfire for years.

07★★ The turn: where the √n is born

Everything so far deliberately refused to assume independence, and that refusal bought us a lot. It also has a hard ceiling. Symmetry can pin a mean and it can pin a probability, and it can never say one word about spread.

Independence is the one assumption we have not spent yet, so let's spend it and see what it buys. The first thing it touches is precisely the quantity symmetry could not reach, because Chapter 12 showed that variances add for independent pieces.

Here is where nearly everyone slips, so predict before you read on. Suppose X₁, …, Xₙ are independent with variance σ² each, and you form the average X̄ = (X₁ + … + Xₙ)/n. What happens to the variance?

The gut says divide by n, because that is what dividing does. It divides by , and that single squared factor is the entire √n. The reason the gut cannot correct itself is that we quietly picture variance as living in the same units as the variable. It does not, and its own definition E[(X−μ)²] says so.

One class of six, measured twice — step from centimetres to metres and watch which number falls by 100 and which falls by 10,000.
beat 1 of 3 — the class, in centimetres unit: cm view A — the six people, drawn to fit mean −11 −5 −1 +1 +5 +11 the SIDE of each square is one deviation its AREA is that deviation, squared view B — the same numbers on ONE fixed ruler 10 5 0 SD × SD areas: 121 25 1 1 25 121 mean area = 49.0000 = the gold square SD — a spread, a LENGTH 7.0000 cm = √49.0000 Var — mean sq. dev., an AREA 49.0000 cm² = mean of the 6 areas SD is a length: cm. Var is an area: cm². You stand 7 cm off the mean — never 49 cm² off it. SD(aX) = |a| · SD(X) Var(aX) = a² · Var(X) ×1 ×1.00 ×1.00 the metre flip was just a = 1/100 → SD ×1/100, Var ×1/10,000 now set a = 1/n and average n independent copies of X the scale factor (1/n)² = 1/16 × the sum's variance n·σ² = 196 = variance of the mean σ²/n = 12.2500 SD(mean) = 3.5000 cm = 7/√4 — there is the √n σ = 7.0000 cm, from beat 1 — same class, averaged n times
Var lives in cm², SD in cm
Six students. Each deviation is how far that person sits from the mean height, in cm.

Square each one, average the six squares → that average is the variance. Its unit is cm × cm.

The gold square has side SD, so its area is exactly the mean blue area.
1 cm = 0.01 m, so every deviation is divided by 100. Both dials, then reveal.
3 beats · you predict in beat 2
What you're looking at — one sample, two rulers, and the factor that gets squared
Six deviations −11, −5, −1, +1, +5, +11 cm. Each blue square has one of them as its side, so its area is that deviation squared. View A is drawn to fit and never moves; view B plots the same numbers on a fixed ruler, which is where the unit shows.
SD is a length: 7.0000 cm. Flip to metres and it just loses two decimal places — 0.0700 m. Scale by any a and it scales by |a|. That is the whole of it.
Var is an area: 49.0000 cm² → 0.0049 m², down by 100² = 10,000. So Var(aX) = a²Var(X). Put a = 1/n and the average picks up 1/n², which meets the sum's nσ² and leaves σ²/n — SD σ/√n. The √n is born right there.
Fig. 12. The rung that everyone skips. One class of six, deviations −11, −5, −1, +1, +5, +11 cm about the mean, so SD = 7.0000 cm and Var = 49.0000 cm² — and the variance is drawn as what it literally is: six squares whose sides are the deviations, averaged into the gold square of side SD. Flip the ruler to metres and nothing physical happens; every side is divided by 100, so every area is divided by 100² = 10,000, and the readouts land on 0.0700 m and 0.0049 m². Predict both before pressing reveal — most readers divide the variance by 100 too, and the squares are what shows them why not. The third beat replaces 1/100 with any factor a: SD scales by |a|, Var by a², and the two gauges cross at a = 1. Then set a = 1/n, which is exactly what averaging does: the outside factor contributes (1/n)² = 1/n², the sum of n independent copies contributes nσ², and the two collide into σ²/n — so SD(X̄) = σ/√n. That is the entire origin of the √n in the central limit theorem and of every annualise-by-√252 rule later in this course: not a mystical constant, just the bookkeeping of squared units. It is also why volatility is quoted as σ in percent and never as variance in percent-squared — a percent² is not a thing anyone can stand next to.

Measure a set of heights in centimetres and the standard deviation comes out at 7 cm, which makes the variance 49 cm². Now measure exactly the same people in metres. Every number shrinks by a factor of 100, so the deviations shrink by 100 while their squares shrink by 10,000.

The variance is therefore 0.0049 m² and the standard deviation is 0.07 m. Nothing about the people changed. The square is bookkeeping, not convention.

One line of algebra confirms it using Chapter 11's identity: Var(aX) = E[a²X²] − (aE[X])² = a²Var(X), and taking the square root gives SD(aX) = |a|·SD(X). That is the rung the rest of this chapter stands on, and it appears in neither Chapter 11 nor Chapter 12.

It also explains a habit you have probably noticed. Finance quotes standard deviation and essentially never quotes variance. A sigma is in percent, which you can feel. A variance is in percent squared, which nobody can.

Now put the two rules together. This is the centre of the chapter.

The right angle — bet first, then watch two errors refuse to queue up
n = 2 · ρ = 1.00 2.0000 σ√(2+2ρ) total error of the SUM downstream costumes: SE of the mean 1.0000 95% CI ±1.96 SE ±1.9600 √252 day→year 15.8745 if they QUEUED end to end: n·σ = 2.0000 braced at a right angle: σ√n = 2.0000 You average 100 noisy readings instead of one. How many times more accurate is your answer? no computing. just bet — dial it on the right → ×100 then press LOCK IN MY BET
place your bet, then lock it in
What you're looking at — errors that brace against each other instead of queueing
each blue leg = one reading's error, length σ (sigma, the typical size of one mistake), joined tip to tail
the gold arrow = the total error you actually get — the hypotenuse, σ√n
red = the myth: if errors queued up end to end the total would be
√252 = 15.8745 — the same right angle wearing its volatility costume
ρ (rho, the correlation) is the cosine of the angle between the two error arrows. ρ = 1 → 0°, they lie on top of each other and queue up. ρ = 0 → 90°, and the total is a hypotenuse. √n was never a convention — it is Pythagoras, n times over.
Fig. 13. Bet before you can compute. Almost everyone answers ×100 — the truth is ×10, and the gap is the anaesthetic. Then the picture: two error arrows of length σ joined tip to tail. At ρ = 1 they lie flat along each other and the total reads 2.0000; drag the dependence dial to 0 and the second arrow swings up until it is perpendicular, the total sliding to √2 = 1.4142 — a right-angled triangle with both legs σ and the hypotenuse read off the gold arrow. Raise n and each new error comes in perpendicular to the running total, a spiral of right angles whose resultant is exactly σ√n: 2.0000 at n = 4, 5.0000 at n = 25, 10.0000 at n = 100. Press ÷n and the same number turns into the error of the average, σ/√n: 0.5000, 0.2000, 0.1000. That single hypotenuse is every standard error, every confidence interval, and the √252 in every annualised volatility quote for the rest of the course.

You average 100 noisy readings instead of taking one. How many times more accurate is your answer? Almost everyone writes 100. It is 10.

The reason is a picture, and it is one you already own from Chapter 6. Draw the first measurement's error as an arrow, then attach the second error to it. If the two errors were locked together, so that knowing one told you the other, they would point the same way and their lengths would simply add. That is the gut's picture, and it is what dependence looks like.

Independence means the second error carries no information whatsoever about the first. Geometrically that is a right angle. So the two errors do not queue up end to end, they brace against each other, and the combined length is the hypotenuse √(σ² + σ²) = 1.4142σ rather than .

Stack n of them, each coming in perpendicular to all the others, and the total error length is σ√n. Then divide by n to turn that sum into an average, and the scaling rule from the units panel carries it to σ/√n.

The algebra is now confirmation rather than argument: variance adds, so Var(X₁ + … + Xₙ) = nσ². Scaling by 1/n multiplies variance by 1/n², so Var(X̄) = nσ²/n² = σ²/n, and therefore SD(X̄) = σ/√n.

The mean, meanwhile, does not move: E[X̄] = μ. That is the equal-shares lemma read in the independent case, since n interchangeable pieces are sharing a total of .

σ/√n is the standard error of the mean, and it is how wrong your average is likely to be. From Chapter 16 onward we use it constantly.

One word about noise before we move on, because it is two different things wearing one coat.

One slider, one √n, two opposite answers — the SUM's spread climbs while the AVERAGE's spread falls, from the very same draws.
n = 4 · √n = 2.00 one √n does both: ×2.00 ÷2.00 20 5 1 SD of the SUM = σ√n = 2.0000 1 0.2 0.05 SD of the AVERAGE = σ/√n = 0.5000 1 400 n — how many draws → 4 25 100 running SUM — fans OUT spread × √n running AVERAGE — funnels IN spread ÷ √n
n=4 → sum ×2.00, avg ÷2.00
HOW MUCH MORE DATA?
÷2.0 the error ⇒ ×4.00 the data
×2.0 the data ⇒ ÷1.4142 the error
at n = 4 ⇒ 16 draws needed
What you're looking at — two questions, one √n, opposite answers
The SUM of n draws. Write σ (sigma) for the spread — the standard deviation — of a single draw; here σ = 1. Add n independent draws and the spreads add in squares (Ch11), so SD = σ√n. It climbs: 2, 5, 10 at n = 4, 25, 100. The dot cloud beneath is real running sums — they fan out.
The AVERAGE of the same n draws. Divide that sum by n and the spread divides by n too: σ√n / n = σ/√n. It falls: 0.5, 0.2, 0.1 at the same n = 4, 25, 100. The lower cloud is the identical draws, funnelling in on the true mean — that is the Law of Large Numbers, drawn.
Both bands are the same √n, multiplying above and dividing below — each value axis is its own, so read the numbers, not the heights. Hence the price: to halve your error you need the data; doubling the data buys only 1.4142×. Same √n behind the √252 annualising rule you'll meet later.
Fig. 14. The word "noise" is doing two opposite jobs at once, and this is the picture that separates them. Write σ for the spread of a single draw (here σ = 1). Because independent variances add (Ch11), the sum of n draws has SD = σ√n — it grows without limit — while the average, which is that same sum divided by n, has SD = σ/√n — it shrinks to nothing. Drag n, or tap the landmarks on the axis: at n = 4, 25, 100 the sum's spread reads 2.0000, 5.0000, 10.0000 while the average's reads 0.5000, 0.2000, 0.1000. Each band has its own value axis, so compare the printed numbers, not the heights — the point is that one √n multiplies the top and divides the bottom. The two dot clouds are the honest version of the claim: the same simulated draws, accumulated once as running sums (fanning out inside the dashed ±σ√n envelope) and once as running averages (funnelling in inside ±σ/√n), so nobody can say the two pictures came from different data. Press new draws and the individual paths change while the two envelopes do not. The HOW MUCH MORE DATA dial prices the consequence: error scales like 1/√n, so cutting error by a factor f costs f² times the data — halve the error, quadruple the sample; double the sample, and you buy a mere 1.4142×. That brutal exchange rate is why estimates get expensive, and it is the same √ law behind annualising volatility by √252 later in the course.

The spread of the sum grows like σ√n. The spread of the average shrinks like σ/√n. Both are true, both come from the same √n, and people who never see them on one panel spend years quietly unsure whether more data makes things noisier or cleaner.

Two consequences are worth carrying out of here. Doubling your data buys a factor of 1.4142, not 2. And to halve an error bar you need four times the data, which is the arithmetic that prices every research decision you will ever make.

A footnote on what independence was really needed for. In the variance addition, the cross term died because E[XY] = E[X]E[Y], and that condition is called uncorrelated, which is weaker than independent. The cross term itself has a name, covariance, and Chapter 18 does the work.

08The centre locks: the law of large numbers

We now know the average's spread heads to zero. It is tempting to declare victory and say the average converges to μ, but there is a join here that most treatments slide across without comment.

"The standard deviation goes to zero" and "the probability of being far from μ goes to zero" are different statements. The second one is what you actually want, and crossing from the first to the second takes one honest tool.

Start with something that looks almost too simple to be useful: let X be a random variable that is never negative, and pick a level a. The definition of expectation says E[X] = Σ x·p(x), and every term in that sum is non-negative.

A balance rail where the mean is locked — so every block you push out past a has to be paid for out of E[X]
8 blocks · each carries mass ⅛ · the mean is bolted at μ = 2 drag a block ←→ (arrow keys work too) a = 3.0 0 2 4 6 8 10 μ locked E[X] 2.000 ① keep x ≥ a 0.000 ② a·P(X ≥ a) 0.000
① throw away every block left of a.
E[X] (locked)2.000
① keep x ≥ a0.000
② a·P(X ≥ a)0.000
tail mass P(X ≥ a)0.000
budget E[X]/a0.667
Every block parked at a costs a×⅛ out of a fixed budget E[X]. Run out and it snaps back.
E[X] ≥ a·P(X ≥ a) — always
What you're looking at — Markov's inequality as a balance you cannot cheat, then the same balance aimed at (X̄ − μ)².
The blocks are probability mass, ⅛ each. The fulcrum is E[X] = 2, bolted down: drag one block right and the rest slide left to pay for it.
Level a. Two honest discards — throw away everything left of a, then drag survivors back to exactly a — give E[X] ≥ a·P(X ≥ a), i.e. P(X ≥ a) ≤ E[X]/a.
Chebyshev = the same two lines pointed at (X̄ − μ)², whose mean is σ²/n. Bound σ²/(nε²) is loose — and still collapses to 0.
Fig. 15. Markov's inequality, built as a machine you can push against. Eight blocks of probability mass (⅛ each) sit on a rail whose mean is bolted at E[X] = 2 — that is the counterweight, and it is the whole trick: drag one block out to the right and every other block slides left to pay for it. Push too far and there is nothing left to pay with, so the drag snaps back. Now set a level a and run the proof as two physical acts. ① Throw away every block left of a: the mean can only fall, so E[X] ≥ E[X·1{X ≥ a}]. ② Drag each survivor back to exactly a: it falls again, and what is left is literally a × (the mass beyond a). Two discards, and out drops E[X] ≥ a·P(X ≥ a), i.e. P(X ≥ a) ≤ E[X]/a — no distribution, no integral, nothing computed. The second tab re-aims the identical machine at the squared deviation (X̄ − μ)², whose mean is σ²/n; taking a = ε² gives Chebyshev, P(|X̄ − μ| ≥ ε) ≤ σ²/(nε²). Watch that bound against a live sampler: it is embarrassingly loose (at n = 100, ε = 0.1 it promises only "≤ 1", which is no promise at all) and yet it collapses on the ladder 1.0000, 0.2500, 0.0400, 0.0100 as n goes 100, 400, 2500, 10,000 — and that collapse is the weak law of large numbers. This is the join every treatment slides across: a shrinking standard deviation is not the same claim as a shrinking probability of being far off. Markov is the two-line bridge between them.

Throw away all the terms with x < a and the sum can only shrink, so E[X] ≥ Σ x·p(x) over the values at least a. Every surviving x is at least a, so replace each one by a and it shrinks again to E[X] ≥ a·P(X ≥ a). Rearranged, that is Markov's inequality, P(X ≥ a) ≤ E[X]/a.

In words, mass far out to the right has to be paid for out of the mean. Now aim it at the squared deviation (X̄ − μ)², which is never negative and has expectation σ²/n.

Being at least ε away from μ is the same event as that squared deviation being at least ε², so Markov gives P(|X̄ − μ| ≥ ε) ≤ σ²/(nε²). That is Chebyshev's inequality, and it is the entire law of large numbers.

Put numbers on it with σ = 1 and ε = 0.1. At n = 100 the bound is 1.0000, which says nothing at all. At n = 400 it is 0.2500, at n = 2500 it is 0.0400, and at n = 10,000 it is 0.0100.

Notice that the bound is loose, meaning far more pessimistic than the truth, and notice that it does not matter. It heads to zero for any ε you care to name, and that statement is the weak law of large numbers.

Now for the thing this law is most often accused of saying. Chapter 10 named the gambler's fallacy. This is where we execute it, because we finally have the machinery to show what happens instead.

The phrase "the law of averages" makes people picture a correcting force. Ten extra heads must somehow be repaid by ten extra tails later. Nothing pays them back.

A live coin sampler — one flip stream drawn twice: the LEAD as a raw count, the GAP as a proportion. Force ten heads and see which one forgets.
the LEAD · heads − tails (a count) lead +0 · ±1 SD 0.0 ±√n = 10.0000 ±√n = 100.0000 +140 0 −140 same +10 heads — only divided by a bigger n the GAP · heads/n − ½ (a proportion) gap +0.0000 · ±1 SD 0.0000 ±0.5/√n = 0.0500 = 0.0050 +0.5 0 −0.5 1 10 100 1,000 10,000 flips n = 0 excess 0 heads · press ⚡ ÷ n = 0.0000
1 · Run   2 · hit ⚡ mid-run   3 · predict
seed (same seed → same run)
does the run pay it back?
Press Run — flip a fair coin.
What you're looking at — one flip stream, drawn twice.
your run — top: the lead, a count; bottom: the gap, a proportion.
the same run with the forced ten taken back out.
the head-start those ten heads bought: same width up top, pinched shut below.
±1 SD of chance: ±√n for the lead, ±0.5/√n for the gap. x = n, log scale.
Fig. 16. A live fair-coin sampler, up to 10,000 flips. The top panel plots the lead — heads minus tails — against the analytic ±√n envelope, which flares open: at n = 100 it is ±10.0000, at n = 10,000 it is ±100.0000. The bottom panel plots the gap — the head fraction minus ½ — against ±0.5/√n, which pinches shut: ±0.0500, then ±0.0050. Same flips, both panels. Press ⚡ Force 10 extra heads mid-run, predict whether the coin pays them back, then watch: the red band up top keeps its exact width to the last flip, while below it closes to nothing. The excess never left — it was only divided by a bigger n.

Flip a fair coin and watch two quantities at once. The absolute lead, meaning heads minus tails, has standard deviation √n. It is 10 at a hundred flips and 100 at ten thousand, so it grows.

The proportion gap, meaning the head fraction minus one half, has standard deviation 0.5/√n. It is 0.0500 at a hundred flips and 0.0050 at ten thousand, so it shrinks.

Same √n, two curves, opposite directions. The excess heads never come home. They simply get divided by a bigger and bigger n, which is why deviations are diluted, never repaid.

09★ The wobble's shape: standardizing and the CLT

The law of large numbers has done something slightly inconvenient. It squashed the average onto a single point, so if you plot the distribution of for growing n on a fixed axis, you get a spike and nothing left to look at.

A spike is not the end of the story. Structure is hiding inside it. To see structure in something that is shrinking you have to zoom in, and the only question is by how much.

The zoom dial — one shrinking average, three magnifications. Only one of them keeps a picture.
the picture is shrinking off the frame: 0.0% −4 −2 +2 +4 one unit = σ ±1 sd = 0.316 units shape vs n=5: 21.9% Z n = X − μ σ / √n tap the gold √n to clear the fraction →
no zoom → shrinking
Predict first: at n = 200, which of the three still shows a bell? Try all three — two of them break.
What you're looking at — one fixed picture frame, one shrinking average, and a magnifying glass with three settings.
blue = the density of the average of n draws from a lopsided parent (waiting times — no bell anywhere in it).
gold = ±1 standard deviation of what is drawn. No zoom shrinks it like 1/√n; ÷σ/n grows it like √n; ÷σ/√n pins it at 1, forever.
violet dashes = the same picture one step of n ago. When it hides exactly under the blue, the shape has stopped changing.
red = mass that has left the picture. Zoom too hard and most of the curve is outside the frame you're looking at.
Fig. 17. The law of large numbers says the average collapses onto μ. That is a problem for anyone who wants to look at it: photograph a shrinking thing at fixed magnification and you end up with a dot. So drag n with no zoom and watch the blue curve squeeze into a single column that punches out of the top of the frame — the badge lights NOTHING LEFT TO SEE. Over-correct with ÷σ/n and you have zoomed too hard: the curve leaves through both walls and the readout counts how much of it is now outside the picture (at n = 200, over three quarters of it). Only ÷σ/√n matches the shrink rate exactly, and the reward is startling — at n = 10, 50 and 200 the drawn curve is the same curve, the violet ghost of the previous n hiding underneath it and the shape-change meter falling toward zero. That is the whole reason the standardising factor is σ/√n and nothing else: to photograph something that is shrinking, you must zoom by exactly its shrink rate. Tap the gold √n to slide it up across the fraction bar — (X̄−μ)⁄(σ/√n) and √n(X̄−μ)⁄σ are one object, not two. (The parent here is a lopsided waiting time, not a bell; the curve is bell-shaped anyway once n is large — that is the central limit theorem doing its work, and the small residual lean at n = 10 is it still finishing.)

Try three settings and let the picture decide. With no zoom at all you get a dot. Multiply by n and the picture explodes off the screen, because you magnified faster than the thing was shrinking. Multiply by √n and the image settles and stops changing as n grows.

That is the whole justification, and it is worth stating plainly. To photograph something that is shrinking, you zoom by exactly its shrink rate, no more and no less. We already know that rate. It is σ/√n.

So define the standardized average by dividing the deviation by precisely that: Zₙ = (X̄ − μ)/(σ/√n). Clear the fraction and it reads √n(X̄ − μ)/σ, which is the form textbooks print. It is the same object in less frightening clothing, and the √n on top is not fighting the one underneath. It is the one underneath, moved.

The zoom finally makes the real question askable. What does the picture actually look like?

A live sampler with two histograms under one n slider — drag it and watch which panel moves
① every single DRAW (the data) ② AVERAGE of n=1 → same picture −3.0 μ = 1.00 5.0 −4σ/√n μ +4σ/√n THIS PANEL NEVER MOVES view = μ ± 4σ = ±4.00 skew 2.0000 frozen skew 2.0000 = 2.00/√1 ex-kurt 6.000 frozen ex-kurt 6.000 = 6.00/n
which sentence is TRUE?
Two sentences. One is false.
What you're looking at — one word, "Normal", pinned to two different objects. Only one of them obeys.
Left = the raw draws. Its shape is the parent's shape, and n never touches it. Skewness γ₁ (lopsidedness, 0 = symmetric) and excess kurtosis γ₂ (0 = bell-tailed) sit frozen.
Right = the average of n draws. Exactly: γ₁/√n and γ₂/n. At n = 30 from the cliff that is still 0.3651 — the honest answer to the folklore "n > 30 is Normal".
The cyan bell is the Normal with the same mean and spread. The coin's γ₁ is already 0 and it is still two spikes — one number is not a shape.
Fig. 18. The most-repeated false sentence in statistics, killed by a panel that refuses to move. Pick a parent that is emphatically not a bell — a fair coin (two bare spikes), an exponential (a cliff at zero with a long right tail), or a hand-drawn three-lump mess — and then drag the single control both panels share: n. At n = 1 the two histograms are the same picture, because the average of one draw is the draw. Now push n up. The left panel — a histogram of individual draws — does not so much as twitch, and its two shape numbers stay frozen: skewness γ₁ (how lopsided: 0 means symmetric) and excess kurtosis γ₂ (how un-bell-like the tails are: 0 means Normal-tailed). The right panel — a histogram of the average of n draws — walks onto the cyan bell and its numbers fall on an exact schedule: the skewness of a mean of n i.i.d. draws is γ₁/√n and its excess kurtosis is γ₂/n. Those are equalities, not hand-waving, and they are the whole content of "n > 30 is enough": from the cliff, γ₁ = 2, so at n = 30 the average still carries skewness 2/√30 = 0.3651. Not zero. Not close to zero for a risk number that lives in the tail. Press the folklore button and read it off. Note also the coin, whose γ₁ is already 0 while it remains two spikes — one number is not a shape, and symmetry is not Normality. So: collecting more data never makes the data Normal; it makes your estimate of its average Normal. That distinction is why a diversified book of many small independent bets has a near-Gaussian P&L while any one position in it does not, and it is the reason the √n bookkeeping here (spread of the sum grows like √n, spread of the mean shrinks like 1/√n) reappears later as the annualise-by-√252 rule. The two hypotheses doing the work — independence and finite variance — are not decoration; the next figure removes the second one and watches both laws die.

Before the answer, kill the corruption. The most common false sentence in all of statistics is that the central limit theorem says data becomes Normal if you collect enough of it. It says nothing of the kind.

The theorem is about the average, or equivalently about the sum. It is never about the data. The left panel of that figure is a histogram of raw draws from a deliberately ugly parent, and sliding n does not move it one pixel. The right panel is a histogram of averages, and it is the only thing that changes.

With that separation held, here is the statement. Take any parent distribution with a finite mean μ and a finite variance σ², draw n independent values, and standardize the average. As n grows, the distribution of Zₙ approaches the standard Normal, whatever the parent happened to be.

Read what that costs the parent. Its shape is forgotten entirely, and only μ and σ survive the trip. That is why two completely different businesses can end up with identically shaped uncertainty about their averages.

While we are here, the n > 30 rule you may have met is folklore and not a theorem. There is no proof anywhere behind it. A symmetric parent is close to Normal by n = 5, and a badly skewed one is still visibly skewed at n = 100. The skew of the average falls like skew/√n, so for an exponential parent at n = 30 it is still 0.3651.

The last question is the one usually answered with a shrug. Why the Normal, of all possible shapes?

Exact arithmetic, not a sampler — convolve the coin with itself, then test which shape survives ÷√2
exact · no sampling 1 coin · 2 outcomes ½ × the shape + ½ × it, shifted → the two-spike coin · press CONVOLVE sum the overlaps ↓ that is the new shape roughness Σp² 0.5000
no inverse — un-smoothing needs negative mass
exact values (k / 2ⁿ)
0.5000 0.5000
Each extra coin averages the shape against a copy of itself shifted one step. An average can only smooth.
the coin · 0.5000 · 0.5000
What you're looking at — one more independent piece is one more averaging of the shape against a shifted copy of itself, and exactly one shape survives that treatment.
Blue = half the shape where it stands. Convolving with a fair coin means: keep half here, slide half one step right.
Violet = that shifted half. Every result bar is blue + violet stacked — the overlap, summed. Averaging only ever smooths, so Σp² can only fall and UNDO is impossible.
Gold = the shape after adding two copies and rescaling by √2. If a limit exists it must be a fixed point of that move — and only the normal has zero mismatch. That it exists needs characteristic functions, not this figure.
Fig. 19. The bell is not chosen — it is what is left. Start with the least bell-shaped object there is: a fair coin, two spikes of 0.5000 and 0.5000, and nothing in between. Adding one more independent coin does exactly one thing to that shape: keep half of it where it is, slide half of it one step right, and add the overlaps. That is convolution, and the widget prints every number it produces — 0.2500 / 0.5000 / 0.2500, then 0.1250 / 0.3750 / 0.3750 / 0.1250 — computed as exact dyadic fractions k/2ⁿ, not sampled, so there is no randomness hiding the argument. Watch the roughness meter Σp²: it falls at every press and never rises, because averaging a shape with a shifted copy of itself is a smoothing operation and smoothing has no inverse inside the probabilities — that is why UNDO SMOOTHING is greyed out for good, and it is the whole reason the process has somewhere to go. But smoothing alone does not say where. The second tab supplies that. If the rescaled sum settles on some shape at all, that shape must be unmoved by the move that generates it: take two independent copies, add them, and divide by √2 to undo the widening — a genuine fixed point. Try a uniform: it comes back as a triangle-ish hump, mismatch nonzero. Try a three-humped spiky thing: it comes back smeared, mismatch large. Try the normal: mismatch 0.00%, exactly, at every scale. That is the answer to "why the bell and not something else" — nothing else survives being added to itself. Two honest caveats, stated rather than hidden. First, this argument pins the destination assuming a destination exists; the proof that one does uses characteristic functions, the machinery this figure deliberately does not have. Second, it assumes the pieces have a finite variance to rescale by — strip that hypothesis away, as a Cauchy parent does, and the √2 rescaling is meaningless, the fixed point is a different shape entirely, and both the CLT and the law of large numbers walk away. That is not a footnote for a quant: it is most of what fat-tailed return data does to a naive estimate.

There are two honest parts to the answer. The first is that adding an independent piece convolves the shape against itself, and convolution is a smoothing operation. Every bump gets averaged against a shifted copy of the whole shape, and it can never sharpen back. Start with a coin, which is two bare spikes, and three convolutions later a bell is visibly assembling out of nothing but spikes.

The second part pins which smooth shape it has to be. Whatever the standardized average settles into must reproduce itself. Add two independent copies, rescale by √2, and you have to land back on the shape you started with. The Normal is that shape.

Now the honesty note. That argument identifies the destination and assumes a destination exists. Proving one exists needs machinery we have not built, namely characteristic functions, the Fourier transforms of the distribution. That debt is named, not hidden.

10The fine print

No new concept arrives in this section. What arrives is the difference between somebody who has heard of the CLT and somebody who can be trusted with money.

Every statement in the last section arrived with hypotheses attached: independent pieces, a finite variance, and a limit in n. So walk back through those three and ask what happens when each one fails.

Three ways the bell curve fails — and the one number, σ/√n, that the whole rest of the course is still built on.
① the tails converge last — P(Z > z) for the average of 30 Exponential draws 99% VaR 99.99% VaR ×1.70 true tail ÷ Normal tail 0 1 2 3 4 z = 0 to 1: agree to 1% · at z = 2.33 the Normal understates ×1.70 THE PAYOFF LEDGER — what σ/√n still buys you standard error = σ/√n (Ch 16) · 95% CI half-width = 1.96 σ/√n (Ch 17) annualise: √252 = 15.8745 · ×2 data ⇒ ÷1.4142 error · ÷2 error ⇒ ×4 data live: at z = 2.33 that 1.96 σ/√n band misses the true tail by ×1.70
z=2.33 → tail is 1.70×
pick a failure mode ↓
true   P(>z) 0.017032
Normal P(>z) 0.010000
ratio        ×1.70
drag out to 4σ →
What you're looking at — the CLT's three hypotheses, each one broken on purpose
What you assume. The Normal approximation — the bell the CLT promises. Write σ (sigma) for the spread of one draw and n for how many you average.
What is actually true. The exact answer, computed without the approximation. Card 1: identical in the middle, badly wrong in the tail.
Where it breaks. The shaded error, the Cauchy trace that never settles (no finite σ, so no CLT and no LLN), and the widening ±bar when draws cluster.
The bill, still payable. σ/√n survives all three — it is the standard error, the CI width, the √252 rule, and the price of accuracy.
Fig. 20. The Central Limit Theorem is usually handed on with its conditions filed off, so here are the three places it breaks — on one frame, with the bill still attached. ① The tails converge last. Average 30 Exponential(1) draws — a parent with no bell shape whatsoever — and compare the true probability of landing more than z standard deviations above the mean with the Normal approximation the CLT promises. Drag z. At z = 1 they read 0.157465 against 0.158655, a ratio of 0.99: indistinguishable. At z = 4 they read 0.000386 against 0.000032 — the true tail is 12.20× fatter. Convergence in the middle is fast; convergence in the tail is slow, and the tail is exactly where a 99% or 99.99% VaR lives (both marked). ② No variance, no theorem. One stream of random numbers, read through two parents: a Normal one and a Cauchy one. Predict whether the Cauchy running average settles, then press run. The Normal average funnels into its ±1/√n band; the Cauchy average keeps jumping at n = 3000 exactly as it did at n = 3, because with infinite variance both the Law of Large Numbers and the CLT simply do not apply. ③ Dependence eats your n. Raise ρ, the correlation between neighbouring draws, and watch the dots stop scattering and start moving in runs: 250 clustered observations carry only n₋ₑₑ = n(1−ρ)/(1+ρ) independent ones, and the error bar widens to match. Underneath, the ledger: whatever breaks, σ/√n remains the number the rest of the course is built from — the standard error of Ch 16, the confidence interval of Ch 17, the √252 = 15.8745 that turns daily volatility into annual, and the brutal exchange rate that makes doubling your data worth only 1.4142× while halving your error costs 4×.

Start with the limit, because it fails in the most dangerous place. The CLT converges in the middle first and in the tails last, and that is worth seeing with real numbers.

Take the average of 30 independent Exponential(1) draws, where the exact answer is computable. At one sigma out, the true tail probability is 0.157465 against the Normal's 0.158655, so the error is invisible. At three sigma it is 0.004164 against 0.001350, already off by a factor of 3.08. At four sigma it is 0.000386 against 0.000032, wrong by a factor of 12.20.

Read that ledger, then read a risk report. A 99% VaR is a question about the far tail, and the far tail is the last thing to converge. So the one place you most want the CLT is the one place it helps you least.

Now kill the finite variance. A Cauchy distribution looks like a perfectly ordinary bell-ish curve, and its tails are heavy enough that the variance does not exist. Take a running average of Cauchy draws and it never settles. It jumps forever, no matter how much data you feed it, because the hypothesis that failed is one the whole derivation quietly needed.

Then kill independence. Market returns cluster in a crisis: a bad day makes another bad day more likely. So the pieces stop being independent exactly when you need the machinery most. Your effective n is far smaller than the n you counted, and your error bars are too narrow.

That crack is not a footnote, it is a door. Chapter 34 walks through it with fat tails and volatility clustering, and Chapter 31 builds risk measures that do not lean on a Normal tail.

Now cash the chapter out in plain numbers. That one expression σ/√n is the standard error of Chapter 16, the width of every confidence interval in Chapter 17, and the √252 = 15.8745 that turns a daily volatility into an annual one.

It is also the reason a Sharpe ratio needs years of data before anybody should believe it, and the reason doubling your sample buys you 1.4142 rather than 2.

Carry three things out of here. First, "by symmetry" is a method with two named parts: linearity splits a fixed total without asking about dependence, and exchangeability makes the shares equal. Second, an average is wrong by about σ/√n, and that is a hypotenuse rather than a formula to recite.

Third, and this is the sentence the whole of statistics rests on. Averaging does not remove uncertainty. It shrinks it, at a known and provable and disappointingly slow rate.

Chapter 14 changes the question. Everything on this page described a static pile of draws, and the moment you ask how long until the first two heads in a row, or who goes broke first, a static distribution has nothing to say. The answer is to write the expectation as one step plus the same expectation again, and then solve.

iolinked.com
Written by Ajai Raj