◈ probability mapProb · Ch 11/16
Probability from the coin flip up · chapter 11

11Rare Events & First Successes

We ended the last chapter on a failure. We had just built the bell out of the binomial, and then we broke it on purpose. Drag p toward zero and the fitted normal slides its left tail across the 0 line, handing out real probability to negative counts. You cannot get −2 successes, so that red slab is nonsense, and it is the wall this chapter climbs. But the wall is not where most books tell you it is, and I want to be honest about that first. They say the normal "breaks for rare events", and it simply does not. Ten thousand components, each failing today with probability 0.002, is about as rare as anything you'll meet, and the bell fits those ten thousand components beautifully. So rareness is not the test. The real test is a single number you can compute in your head, and finding it is the first thing we do. That number is λ, the expected count, and it turns out to be the only thing that matters. Here's the plan. We'll pin λ and slice the world finer and finer, and everything about n and p will burn off except λ. That's the Poisson, the law of rare counts. The price of that collapse is a deadlock between two opposing forces, and the number they deadlock on is e. Not e imported from calculus. e manufactured right here, in a question about broken parts. Then we'll take the same yes/no atom and turn it sideways. Instead of "how many successes?", we'll ask "how long until the first one?" That's the geometric, and its formula has no coefficient at all — for a reason you'll see rather than memorize. And at the end the two halves will turn out to be the same sentence read twice.

01Where the bell really breaks

We'll test that folk wisdom against two situations side by side, both of them unmistakably rare. The left panel holds 10,000 components, each failing today with probability p = 0.002. The right panel holds just 20 components, each failing with p = 0.05. Both have a tiny p. Both describe something that hardly ever happens to any single part. So if "rare breaks the bell" were the real rule, both of these fits should come out as garbage. Fit the normal to each panel and look at what actually happens.

n = 10,000 GOOD · 0.0000 ? predict first n = 20 BROKEN · 0.1525 ? predict first k — number of successes, each panel's own scale
predict: which fit dies?
predict — then click to reveal
blue bars — the real binomial, P(X=k)
gold line — the normal curve fitted to it
red — fitted mass sitting below k=0, impossible
p(left) 0.0020 · p(right) 0.0500
aha — same rareness, opposite fates: p was never the test.
Fig. 1. Two binomials, both unmistakably rare: 10,000 trials at p = 0.002 on the left, 20 trials at p = 0.05 on the right — the blue bars are the real P(X=k), the gold line is the normal fitted to it. Predict which fit dies, then reveal: the left fit is dead-on (red area 0.0000) and the right fit spills a fat red slab below the impossible k = 0 wall. Drag rareness and both p's shrink together, in lockstep — yet the verdicts never swap. Same rareness, opposite fates: p was never the real test. What actually decides it, the next figure hands you a number for.

The fit on the ten-thousand-component panel is excellent. The fit on the twenty-component panel is rubbish. Same rareness, opposite verdicts. So the folk rule is worse than useless: it makes you distrust a fit that is perfectly fine, and it tells you nothing about the one that failed. What you need is a number you can compute rather than a feeling, and we already own the two facts that produce it. Chapter 10 gave us the binomial's center, μ = np, and its spread, σ = √(npq). Now think about what actually goes wrong when the bell fails: it puts mass below zero. So the only question that matters is how far zero sits from the center, measured the way we always measure distance on a bell — in standard deviations. The center sits at np, and we'll call that number λ, the expected count: on the left panel it is 10,000 × 0.002 = 20 failures expected today. When the event is rare, q is nearly 1, so the spread σ = √(npq) lands very close to √λ. On the left panel that spread is √(10,000 × 0.002 × 0.998) = 4.47, and √20 = 4.47 as well. Divide, and there it is.

λ 0 √λ = 4.47σ
▲ λ=1 ▲ λ=20
mean μ = λ20.00
spread σ = √λ4.472
λ ÷ √λ → wall4.47 σ
mass below 00.0000
same bell every time — only the wall moves
safe: the wall is far offstage
Fig. 2. The bell is drawn against its own ruler — every baseline tick is one σ. Slide λ (the expected count, np) and the curve never budges; only the wall at 0 walks in. Since the mean is λ and the spread is √λ, the wall sits λ ÷ √λ = √λ sigmas below the peak — that gold caliper is the number. At λ=20 (10,000 items at p=0.002) it's 4.47σ away and nothing spills; at λ=1 it's 1σ and a red slab of density lands on impossible counts. The bell doesn't fear a small p. It fears a small λ.

Zero sits √λ standard deviations below the mean. That's the whole diagnostic, one division of two things we already had. So the bell does not fear a small p. It fears a small λ. Read the division at each scale and you can see why. At λ = 100 the wall at zero is 10σ away, so far out that no bell has any mass there. At λ = 20 the wall is 4.5σ away, still safely offstage. But at λ = 1 the wall is only from the peak, and a fat red slab of the curve is now parked on impossible negative counts. That's why our two rare cases split so cleanly. The ten-thousand-component panel had λ = 20. The twenty-component panel had λ = 1, because 20 × 0.05 = 1. And notice what the failure just handed us, because it wasn't a warning — it handed us λ, and told us that λ is the quantity the rare world is actually organised around.

02Slice the window until only λ survives

If λ is the only thing that matters, we should stop treating n and p as two independent knobs, and build the picture the rare world actually has. Take a physical window: one minute at a call centre, chopped into 60 one-second slices. Each second is now a tiny yes/no question, a Bernoulli trial: a call either arrives in that second or it doesn't. That gives n = 60 trials, each with some small p. But why seconds? Nothing forces that choice. Chop the same minute into milliseconds and now n = 60,000, a thousand times more slices, each carrying a p a thousand times smaller. Chop it into microseconds and n runs to sixty million. The calls did not change. Only our own bookkeeping changed. Click through the slicing dial and watch.

0s 60s 8 calls caught n = 60 slices p = λ/n ≈ 0.1333 each slice ≈ 1 s
gold — a real call arrival; fixed, never moves
green — the slice that caught it, zoomed below
p≈0.133, n=60 — still 8
If the answer changed sec → ms, is that about the calls, or about us?
pick one, then watch the ruler answer
Fig. 3. Eight calls land in one minute — fixed gold dots on the strip that never move, no matter what we do next. The sec / ms / μs dial re-chops the same minute into n = 60, then 60,000, then 60,000,000 slices: watch p = λ/n collapse toward zero while the magnifier below zooms one slice from a fat 1-second block down to a microsecond hairline. Hit re-chop and the whole grid jumps to a random offset — a different ruler, laid down blind — and the big 8 never twitches. That's the tell: n and p are choices we impose on the timeline, not facts sitting inside it. Answer the question underneath honestly and you've got it — the slicing is our ruler, not the world's. That's exactly why the n→∞, p→0 limit isn't a trick played on symbols: nothing in the world was ever counting slices to begin with.

Here is the point, and it closes the trap before any algebra opens it. The slicing is our ruler, not the world's. The calls land where they land, whether we chop the minute into seconds or into microseconds. So ask yourself: if the answer changed when we switched from seconds to milliseconds, would that be an answer about the calls — or an answer about our chopping? It would be an answer about us. So whatever we compute has to be slicing-proof. Now look at what survives that demand. As we chop finer, n explodes toward infinity and p collapses toward zero, so both of our knobs are running away at once. But their product doesn't move at all. Going from 60 slices to 60,000 multiplies n by a thousand and cuts p to a thousandth, which is exactly what "same minute, finer ruler" means. Watch the three readouts as you drive the dial.

» n=40 0 p=0.5000 λ λ=20.00
predict: which survive every slicing?
predict, then drag the dial
1 slicing-proof quantity
⇒ one-parameter law
n — slice count, races off the right edge
p — per-slice chance, races into zero
np — welded to λ=20, never moves
Fig. 4. Circle a guess in the chips, then drag the slicing dial — this is the invariance argument, in your hands, before a single line of algebra runs. Slice the same 10,000-item, λ=20 setup finer and finer: n rockets off the right edge, p crashes into zero, and every checked box that isn't np gets struck live. The np marker never leaves the middle — it's welded to λ, because p is defined as λ/n on purpose. That's why the Poisson pmf e−λλk/k! carries exactly one parameter: λ is the only quantity a slicing we invented can't touch, so it's the only thing the law is allowed to depend on.

n is meaningless on its own. p is meaningless on its own. Only np = λ is slicing-proof, and it stays pinned at the number we actually care about: the expected count of calls in that minute. So the law of rare events must have exactly one knob, and we haven't computed a thing yet. That isn't a trick of algebra. It's forced on us by the fact that we invented the slicing in the first place. Formally, we're setting p = λ/n in the binomial and letting n be ours to choose: 20 calls expected in the minute, spread over 60,000 millisecond slices, means p = 20/60,000 = 0.00033 for a slice. Every choice of n is a perfectly legal binomial, describing the same minute with the same expected count. Now let's ask that sliced window a question.

03The deadlock that makes e

We'll ask the easiest question there is: what's the chance that nothing happens in the whole window? No calls all minute. For that, every single slice has to miss. The slices are independent, so by the multiplication rule we just multiply n identical factors together. Each slice misses with probability 1 − λ/n, which for our 60,000 millisecond slices is 1 − 0.00033 = 0.99967 per slice. So the answer is (1 − λ/n)ⁿ, and now we chop finer and finer and see where it goes. Before you look, I want you to commit, because there's a real fight inside that expression. Two forces are pulling in opposite directions. Each slice is getting safer: the base 1 − λ/n is racing up toward 1, and one raised to anything is one, so this force says the answer is 1. But there are ever more slices to survive: the exponent n is running to infinity, and survival gets harder every time you have to do it again, so this force says the answer is 0. Pick one. Is it 1, is it 0, or is it something in between? Then run it.

P(nothing) = (1 − 1/n) n →1 →∞ this really runs: lam = 1 for n in (10,100, 1000,10**6): print((1-lam/n)**n) stdout › n=10 n=100 n=1e3 n=1e6 → …
commit first — where does it land?
pick 1, 0, or in between
base 1−λ/n → 1 — each slice is safer
exponent n → ∞ — but more must miss
the deadlock IS e−λ. Nobody imported it.
Fig. 5. Ask the sliced minute the easiest question: what's the chance nothing happens? Every one of the n slices must miss, and each misses with chance 1−λ/n, so the multiplication rule gives (1−λ/n)n — and two forces immediately pull opposite ways. The base climbs toward 1 (each slice is getting safer), while the exponent runs off to ∞ (there are ever more of them to survive). Your gut shouts 1=1. Commit to an answer, then press Run — the code on the left actually executes and prints: 0.34868 · 0.36603 · 0.36770 · 0.367879. Neither force wins. The digits freeze column by column onto an unbudging number, and that number has a name: 1/e. "The value (1−1/n)n stalls at" is not a fact about e — it is what e is. Pin λ=3 and the same tug-of-war stalls on 0.049787 = e−3. Nobody imported a calculus constant into a question about broken components; the deadlock manufactured it. You now own P(X=0) = e−λ — the whole Poisson before the Poisson, and the rest of the law (e−λλk/k!) is just λk/k! bookkeeping stacked on this one number.

Neither force wins. They deadlock, and the number they deadlock on already has a name. Watch the digits freeze column by column as you slice, with the expected count pinned at one: ten slices give 0.34868, a hundred slices give 0.36603, a thousand give 0.36770, and keep slicing and it tightens onto 0.367879 and stops moving. That's not "roughly a third". That's a specific, unbudging number the two forces settle on and never leave again. And the stalemate value is one you already know: it is 1/e, one divided by 2.71828. The number e is defined as the value this tug-of-war lands on, so read the arrow the other way round from how you've probably always read it. Nobody imported e from calculus into probability. We asked an ordinary question about a call centre, and the question manufactured e on the spot. That's why a transcendental constant shows up in a formula about broken components. It was never a guest. It's what "safer and safer, but more and more times" equals.

And it generalizes immediately. Pin λ = 3 instead of the λ = 1 we used above, and the same deadlock lands on 0.0498, which is exactly e^(−3) — one divided by 2.71828 cubed. In general, (1 − λ/n)ⁿ settles on e^(−λ) for whatever λ you pinned. Drag λ and watch the numeric ladder chase the curve.

0 1 2 3 4 5 λ 0 1 P
for n in 10,100,1000,10⁶: v=(1−λ/n)ⁿ
n=10
n=100
n=1000
n=10⁶
e^−λ = —
drag λ — watch the ladder converge
aha — you already own P(zero)=e^(−λ) for ANY λ, before meeting the Poisson.
Fig. 6. Pin any λ you like — drag the slider, type a number, or hit a preset — and the code re-sweeps n = 10 → 100 → 1,000 → 10⁶, printing (1−λ/n)ⁿ at each rung beside e^(−λ) and the gap between them. The gap doesn't just shrink — by n = 10⁶ it reads 0.000000, and that last rung is the gold dot, sitting exactly on the blue curve. And that blue curve is nothing but this same loop, swept across every λ. That's the whole Poisson at k = 0 sitting in your hands: P(zero events) = e^(−λ) — not a fact to memorize later, but the deadlock you just watched, for any window you name.

Fig 6 left you with P(zero) = e^(−λ) for any window you name. So name a bigger window. A call centre takes 20 calls in an average minute — what does it expect over three minutes, and does the chance of a completely quiet stretch change as the window grows?

t=1.00min 0 5min rate: 20/min (fixed) λ = rate×t = 20.0 P(0) = e^(−λ) = 2.06×10⁻⁹
watch 3 min instead of 1 — expected calls is ___?
predict, then drag the window
rate — the phenomenon's real constant; it never moves
λ = rate × window — grows exactly as you stretch it
P(0) = e^(−λ) — quiet gets rarer, fast
same rate, bigger window: λ scales, nothing else does — Ch12 calls this rate × time
Fig. 7. Predict first: stretch the call-centre window from 1 minute to 3, and most guts say the expected count stays 20. Drag the slider and watch instead — the window bar grows, and the λ row grows in exact lock-step with it: λ=20.0 at t=1.00min, λ=60.0 at t=3.00min, on up to λ=100.0 at the full 5-minute window — while the rate row underneath never moves at all. The row you already own, P(0)=e^(−λ), tracks it: an already-tiny 2.06×10⁻⁹ at one minute collapses to a sliver too small to print by three. Flip to “3 typos a page” and the same law runs on friendlier numbers: λ=1.5 at half a page, λ=6.0 across two pages, and P(0) sliding from 0.2231 down to 0.0025. Nothing about the phenomenon changed — only the window you named. λ was never a fact about the world; it's rate × window, and you just picked the window.

So you now own a real result, and you own it before you've met the law it belongs to. P(zero events in the window) = e^(−λ). Put the call centre's 20 calls a minute into it: a silent minute has probability 0.000000002, two chances in a billion. That is the whole Poisson at k = 0, and everything else in the law will turn out to be bookkeeping stacked on top of this one number. Let's go get the rest.

04Building the law from the split

Same window, harder question: what's the chance of exactly k events, not zero? We already have the machinery, because the count of successes in n independent slices is a binomial. So substitute p = λ/n into the binomial PMF and push the same slicing to the limit. Now, this is the exact spot where most readers quietly give up and start memorizing, so let me name the panic before it bites. You're about to watch n run to infinity while k just sits there, unmoved, and that feels illegal. Here's the answer, and it's one line. k is the question we asked — "exactly three calls this minute". n is the ruler we chose, whether that ruler is 60 slices or 60,000. The ruler goes. The question stays. The second thing that usually gets skipped is the regrouping. We are not going to send factors off to their limits one at a time on a whim. We sort first, into four boxes with all the n-dependence exposed, and only then does each box become answerable by inspection. Step through it.

1 · the binomial, as we own it P(X=k) = C(n,k) · pᵏ · (1−p)ⁿ⁻ᵏ k successes in n tries · one fixed p here: k = 3, λ = 2 no limits taken yet — only regrouping n(n−1)(n−2)/n³ 0.9970 λᵏ/k! = 2³/3! 1.3333 (1 − λ/n)ⁿ 0.1351 (1 − λ/n)⁻ᵏ 1.0060 the ruler goes; the question stays k = 3 (frozen) n → ∞ : the ruler streams past
step 1 · nothing substituted yet
n(n−1)(n−2)/n³ → 1 · k ratios, each → 1
λᵏ/k! · not one n inside — it just stays
(1−λ/n)ⁿ → e⁻λ · this IS the definition of e
(1−λ/n)⁻ᵏ → 1⁻³ = 1 · k is finite, so it dies
Fig. 8. the derivation x-ray. Nothing is limited until step 4: we regroup first, so every n sits exposed in its own box. Three boxes then answer by inspection (→ 1, no n at all, → 1); the only real work is (1−λ/n)n → e−λ — which is the definition of e, and the whole reason e shows up in the Poisson pmf e−λλk/k!.

Look at how little actually happened there. Box one is k factors on top, each within k of n, divided by k factors of n. With n = 1000 slices and k = 3 calls that is (1000 × 999 × 998) / 1000³ = 0.997, and as the slices multiply box one goes to 1. Box two has no n in it at all, so box two cannot move. Box three is our deadlock, and it goes to e^(−λ). Box four is a fixed number of factors, each going to 1. So three boxes are trivial, and the one that isn't is the thing we built in the last section. What's left standing is the Poisson distribution:

λ=4 k=3 0.40 0 0 16
e^(−λ) = 0.01832
λ^k = 6.40e+01
1/k! = 1.67e-01
P(k=3)=0.1954 Σ=1.00
Fig. 9. One product, three jobs. e^(−λ) is the price of nothing happening — the deadlock the chapter earned. λ^k is k events actually landing. 1/k! divides out the k! ways those same k events could have been ordered, since order never mattered. Drag the gold arrow across the bars and all three numbers update live. λ^k climbs every step you take, but 1/k! collapses even faster once k passes λ. The exact ratio is P(k)/P(k−1)=λ/k: it equals 1 at k=λ, then drops below it for every k past λ. That crossing is the turnover you see in the bars. Click a colour to isolate that factor's job on the selected bar (the other two dim, in both the key and the race rows below). Press strip e^−λ and the red row vanishes, the bars blow straight through the axis cap, and the total stops being 1 — proof that constant's only job was normalising the tower, not decorating it.

P(X = k) = e^(−λ) · λ^k / k! Read it as three jobs. The e^(−λ) is the price of nothing happening, the deadlock we earned in the last section. The λ^k is the k events landing. The k! divides out the orderings, exactly the way the binomial coefficient's denominator always did. Run all three jobs at λ = 3 and k = 2: 0.0498 × 9 / 2 = 0.224, so something that averages three per window arrives exactly twice about 22% of the time. One parameter, λ. No n, no p. Now let's audit the newborn, because I'm not going to hand you a PMF without checking that it really is one. The first duty of any distribution is that its probabilities sum to 1. Here's the honest argument for this one, and it needs no new tools at all. Every binomial we squeezed summed to 1, and we pinned λ, so no mass ran off to infinity while we squeezed. A limit of laws that each sum to 1 sums to 1. Add the bars up yourself and watch the tank fill.

POISSON λ=6 0 14 k → events landing Σ 0.000000 so far
click bars, or hit pour all
blue = P(X=k), the bar's own height
green = poured; tank's running total
gold = tank locked at 1, binomial's own sum
Honest gap: the closed-form reason Σλk/k! = eλ uses the exponential series, and we haven't handed you that tool yet — so we're not spending it here. The squeeze argument above needs none of it.
Fig. 10. Click bars or hit pour all: each Poisson probability drains into the tank, the total climbing to six decimals — the tail bars are microscopic — pour all 15 and the tank reads 0.998600. The missing 0.0014 is the k>14 tail we cropped off the screen, not mass the law lost. Flip to SQUEEZE and drag n: the bars redraw as the binomial at that n (λ=np pinned at 6), and its tank is already gold and full at 1.000000 — it never once moves, because every binomial you ever squeezed summed to 1 exactly. That's why the normalisation isn't faith or a borrowed series: the mass was pinned at 1 on the way in, and pinning λ meant none of it could escape on the way out.

One honesty note, flagged rather than hidden. There is a closed-form reason the Poisson sums to one: the series Σ λ^k/k! equals e^λ exactly, which cancels the e^(−λ) out front. Set λ = 1 and you can watch that happen: 1 + 1 + 1/2 + 1/6 + 1/24 + ... = 2.71828, which is e itself. The series is genuinely lovely and we'll meet it later in the book. I'm not using it now, because you haven't been handed it and I won't pay for a result with a tool you don't own. The limit argument above is complete on its own. Now the second half of the audit: where is this thing centered, and how wide is it? Ride the same limit through the same door. The binomial's mean is np, and we're holding np = λ, so the Poisson mean is λ exactly. The binomial's variance is npq, which is λ·q. We're driving p → 0, so q → 1, and the Poisson variance is λ too.

λ = 6 (pinned) npq np = λ npq = 6 × 0.55 = 3.30 gap = 6 − 3.30 = 2.70 p = 1−q = 0.45 (p→0 ⇒ Poisson)
gap = 2.70 — npq lags np
blue — the mean side: np, λ
gold — the variance side: npq, σ
q→1 erases the gap — failure becomes free, so npq catches np. mean=variance=λ is Poisson's fingerprint.
Fig. 11. Same knob, two dials. FUSE: with λ=np pinned, drag q toward 1 (p toward 0) and the gold npq ring slides into the blue np ring until they merge — variance catches the mean because failure became free. GROW: drag λ and watch the bars' mean march up as λ while the ±σ whisker only widens as √λ — spread falls behind the center. DIAGNOSE: three summary readouts, all averaging 4.0 events; only the one whose variance also lands near 4 is really Poisson — B's variance of 30.2 means something is clustering, C's 0.8 means it's too regular to be random. Pick a sample, then call it.

Mean = variance = λ. This is the fingerprint, and it deserves a moment because it should feel wrong. For every distribution you've met so far, center and spread were two independent dials. You can slide a bell sideways without changing its width. Here one knob does both. And now you can see exactly why, rather than filing it under trivia. A binomial keeps its two dials apart because np and npq genuinely differ. But in the rare limit q → 1, so the difference between them is erased: on our ten thousand components at p = 0.002, the mean np is 20 and the variance npq is 19.96, the same number to a rounding. The two dials fuse because failure became free. So σ = √λ, which at λ = 20 is 4.47, and that is the same √λ from the very first figure of this chapter — the same number that told us where the bell dies. Carry this fingerprint, because it's how you look at real counts and say "the mean is 4 but the variance is 30, so that is not Poisson — something is clustering."

05Using it, and the thing that isn't in it

Time to close the loop the chapter opened. Ten thousand components, each failing today with probability 0.002, so 10,000 × 0.002 gives λ = 20 failures expected. We now have three laws for this one situation: the exact binomial, its normal limit from Chapter 10, and its Poisson limit from this chapter. Before you run the lab, predict which of the three curves is going to miss the simulated data. Take that guess seriously, then let real code simulate a thousand days and settle it.

P
√λ=4.5
predict which curve misses, then RUN
binomial Δmax
normal Δmax · P(<0)
poisson Δmax
Fig. 12. Every RUN draws 10,000×1,000 real Bernoulli trials right here (not canned data), then measures how far the histogram sits from all three laws. Predict the miss, then let the numbers call it.

None of them miss. All three curves lie on top of each other, and the diagnostic told you that would happen before you ran anything: √20 ≈ 4.5, so the wall at zero is 4.5σ offstage and the bell is perfectly legal here. Now hit the λ = 1 button and re-run. The normal starts bleeding mass across zero. The Poisson doesn't blink. So here's your rule, and notice that it is not a new fact — it's the number you've had since the first figure. Compute λ. Take √λ. If √λ is comfortably big, say 3 or more, which means a λ of 9 or more, the wall at zero is out of sight and the bell is legal and cheap. If √λ is small, you're in the counting-rare-things world and Poisson is the law. The binomial is always exactly right, and usually more work than you need.

3 0 2 4 6 8 √λ = 4.47 λ=20 · n=10k · p=.002 HOLDS
npλ√λεNεP
εN, εP = each law's worst miss vs the real binomial
√λ=4.47 ≥3 → both laws hold
n and p were never observable — only λ ever was. That's why binomial can't even be called on "3 typos a page": there's no n to hand it.
Fig. 13. Click a row — or type your own n, p in the last one — and watch √λ decide the verdict, computed row by row, not eyeballed. Flip to Act 2: "about 3 typos a page" has no trial count. normal_fit and binomial_fit can't even be called on it — they raise, live, right there — while poisson_fit refits instantly from λ alone. Edit λ and watch it happen again.

Fig 13 makes Poisson look universal: hand it a rate and you're done. But independence rode in quietly through those slices, and the real world does not always grant independence. So which rare counts are truly Poisson, and how would you catch a count that only pretends to be one?

radioactive decays / second mean 4.0 var ? — press CHECK to reveal it ● = one event on the strip — even, clumped, or too-regular?
your model → NOT POISSON (needs both)
set your model, then CHECK
mean — the average rate λ
variance — the spread; Poisson needs it ≈ mean
var ≈ mean ⇒ Poisson; var ≫ or ≪ mean ⇒ broken
Fig. 14. Poisson isn't just "rare and counted" — it needs two more things that slipped in silently: the events must be independent and arrive at a steady rate. In BENCH, pick a situation, tick the two boxes to commit a model, then CHECK: decays and rain land var ≈ mean and pass, but goals (var 6.4 vs mean 2.7) and typos (var 8.5 vs mean 3.0) cluster into var ≫ mean, while timetabled buses go var 0.9 ≪ mean 4 — too regular. In MAKE IT BURSTY, drag dependence up: the mean stays pinned at 4 but the gold ±σ band outgrows the blue ±√λ=2 the law allows, variance sliding from 4 to 30 — exactly Fig. 11's "mean 4 but variance 30, so that is not Poisson." Break either assumption and the mean = variance fingerprint snaps.

And now the part that almost nobody says out loud, which is the real reason Poisson exists. Look at the formula again and notice what isn't in it. There's no n. There's no p. One knob. That isn't cosmetic tidiness. It is precisely what the collapse bought us. Somebody tells you a page has about 3 typos, so go ahead and try to use the binomial on that. How many trials were there — keystrokes, words, milliseconds of typing? And what is the per-trial error probability supposed to be? Nobody knows, nobody can measure it, and the question isn't even well posed. There is no n, and there never was one. But λ = 3 is sitting right there, already measured, from the simple act of counting 300 typos across 100 pages. n and p were never observable. λ always is. That's why one single law counts typos per page, decays per second, calls per minute, and packet retries per hour. If you think Poisson is "the binomial's approximation", you will freeze the first time someone hands you a rate. It isn't an approximation. It's the law you can actually fit to the world.

06Turn the atom sideways

We're going to keep the same Bernoulli atom and change nothing about it. All we change is the question. So far we have always asked: how many successes in n trials? Now ask the sibling question: how many trials until the first success? Roll a die until you get a six. Retry a packet until it lands. Interview candidates until one finally says yes. Fire the slots and watch the wait land somewhere different every time.

1 2 3 4 5 6 7 8+
press fire — watch it halt
fail success
p = 1/6 · runs 0 · longest 0
aha: pick rare — most tallies pile into 8+. that's the tail, not a bug.
Fig. 15. Same atom, same p — a different question. Not how many successes in n tries, but how many tries until the first one? The row halts itself, and the slot it halts on is the random variable — that count drops straight into the bin below. Press fire a few times, then run 200: the tally settles into the geometric shape — strictly decreasing, tail that never quite dies.

Now let's compute one: what's the chance the first success lands on trial 4? Your hand is going to reach for a binomial coefficient, because we've spent a chapter and a half welding C(n,k) onto every Bernoulli question. Resist that for one second and just write down what has to happen. Trial 1 fails. Trial 2 fails. Trial 3 fails. Trial 4 succeeds. F F F S. Now here's the question that dissolves the reflex: how many other sequences give a first success on trial 4? Go hunting, and try to find even one.

match rejected arrangement
click any sequence — does it give first success on trial 4?
qualifying sequences: 0
Fig. 16. All 16 length-4 fail/succeed sequences. Click any of them hunting for a second sequence that gives "first success on trial 4" — there isn't one; only FFFS qualifies, and the counter refuses to move past 1. Switch to compare and ask the binomial's question instead ("exactly 1 success in 4 trials"): four chips light at once, C(4,1)=4. Same reader, same chips — the coefficient shows up exactly when arrangements exist. Find FFFS and it expands into q·q·q·p: no C(n,k) out front because there was only ever one arrangement to count.

There aren't any. FSFS? That one succeeded on trial 2, so it's a different event entirely. Rearranging is exactly what you're not allowed to do here, because the word "first" nails down every position at once. The failures aren't scattered anywhere. They're locked into the first three slots, and the success is locked into the fourth. So the coefficient isn't missing. The coefficient is 1. That leaves one plain product of independent factors, q · q · q · p. For a die that product is (5/6) × (5/6) × (5/6) × (1/6) = 0.0965, so a first six on trial 4 happens about a tenth of the time. In general the geometric distribution is P(X = n) = q^(n−1) · p. Take that seriously as a skill, because it transfers: from now on, when a PMF arrives with no coefficient out front, read the absence as evidence that the event has exactly one arrangement.

Two warnings before we use it, and both are places readers get hurt silently. The first is that n quietly changed jobs one paragraph ago. In the binomial, n was fixed — the trials you decided to run — and the count was the random thing. In the geometric, the count of successes is fixed at exactly one, and the number of trials is what's random. Same letter, opposite job, one page apart. The second warning is that "the geometric distribution" is one name for two different distributions. Some books count trials including the success, so X ∈ {1, 2, 3, …} and P(X = n) = q^(n−1)p. That's ours, and we're committing to it. Other books count failures before the success, so X ∈ {0, 1, 2, …} and P(X = n) = q^n p. Both are correct. They're off by exactly one: feed the same die and the same n = 4 to both, and our version returns 0.0965 while theirs returns 0.0804. If you learn one and read the other you'll assume the error is yours. It isn't. Always check the support before you trust a formula.

BINOMIAL n=6 k=? 0 1 2 3 4
n is locked at 6 — count is free
green = success, p=0.4 per trial
gray = a trial that failed
the padlock marks the fixed axis
unlabelled: q^(n−1)p — which support?
pick the support this formula uses
Off by one from the book? Convention, not you — means: 1/p (starts at 1) vs q/p (starts at 0).
Fig. 17. Two traps, one page apart. On the left, the same letter n switches jobs: toggle the panel and watch the padlock slide from the row-count to the success-count. In binomial mode, n=6 is locked — you fire all six rows and the count of greens is what's random. In geometric mode, the success-count is locked at k=1 — you fire row by row and it's the number of rows, n, that's random, stopping the instant a green lands. Same symbol, opposite job. On the right, the second trap: "the geometric distribution" is really two distributions. The formula q^(n−1)p only makes sense on support {1,2,3,…} — try labelling it {0,1,2,…} and the exponent stops matching what actually happened. Same bars, support shifted by exactly one; means 1/p and q/p sit side by side for exactly that reason. Check the support before you trust a formula — an answer off by one isn't your mistake, it's an unlabelled convention.

Now the question you'll actually ask in real life. Not "will it succeed on trial 4?", but "will I still be waiting after m trials?" Your first instinct is to add up the whole tail: P(X > m) is the sum of q^(n−1)p over every n beyond m, forever. That's an infinite series, you don't have the machinery to evaluate it, and you conclude the question is hard. It isn't. Chapter 4 already handed you the answer, because "still waiting" is one thing and one thing only: all m of those trials missed. Watch both roads run side by side.

brute force complement q⁵ convergence →
left: 0 terms · Σ=0.0000
terms left to add: ∞
right: not yet — click step
gap: — (right not run yet)
bonus: P(X≤5)=1−q⁵=0.5981
click step — watch both roads run
Fig. 18. Waiting past m=5 rolls of a die (p=1/6). Hit step: the right road finishes in one move — shade the 5 trials, all must miss, multiply q by itself five times, done — while the left road adds its tail one shrinking term at a time, and terms left to add never leaves , because there truly are that many. Keep hitting step, or jump with →200: the crawling blue marker slides onto the green one that was sitting at the answer since the first click, and the gap readout drains to 0.000000. Same number, one line versus an infinite stack of them — and P(X≤5)=1−q⁵ falls out free. That's the reflex: when "none" is one outcome and "some" is an infinite pile of them, price the complement.

One line: P(X > m) = q^m. The brute-force sum crawls toward that number, and the complement road just arrives there. Put the die in it: (5/6)⁴ = 0.482, so after four rolls you're still waiting almost half the time. This is the complement reflex, worth naming as a habit, not a trick. When the thing you want is a tangle of many ways, price its opposite, which is usually one way. And the CDF comes along free: P(X ≤ m) = 1 − q^m, or 0.518 for those four rolls.

07The weld, the fresh start, and the wait

Now I owe you an explanation. Why are Poisson and geometric in the same chapter? Most books put them together because both are discrete and both smell like Bernoulli, which is a filing decision, not a reason. Here's the real reason. Say this one out loud: "the first success comes after trial m." Now say this one: "there were zero successes in the first m trials." Those are the same sentence. Not similar. They are the same event, described from opposite ends. The first is a wait statement. The second is a count statement. They are true in exactly the same circumstances. Look at one timeline with one shading and ask it both questions.

λ = 20 m 1 n qm m/n e−λt 0.018170 0.200000 0.018316 Δ = 1.5e-4
press a question to read the lamp
blue = shaded stretch, first m slices — empty so far
click a shaded slice to drop the ONE success
λ=20 fixed — the 10,000-item, p=0.002 example
Fig. 19. the weld. WAIT and COUNT read the same lamp: click a shaded slice to drop a success and both flip together. Push slicing finer and qm, e−λ(m/n) converge to 6 dp — that limit is the definition of e.

One shading answered both. So the two halves of this chapter just shook hands, and the algebra confirms what your eye saw. On the sliced window, the geometric tail is q^m = (1 − λ/n)^m. The shaded stretch is a fraction m/n of the window, so the expected count inside it is λ·(m/n) — shade half of our minute and that is 20 × 0.5 = 10 calls expected. Push the same deadlock and you get e^(−λ·m/n), here 0.000045, which is exactly the Poisson P(0) we built in section three. Counts and waits aren't two kinds of question. They're one question read from opposite ends of the same timeline, and they shook hands using e. Hold onto that, because Chapter 12 is going to give it a fancy name and you'll already have derived it.

Next, the most stubborn wrong instinct in all of probability, and it does not yield to a formula. You've rolled nine times and no six. So here's the question: is your six closer now than it was before you started? Commit to an answer first. Your gut says yes, obviously — nine misses is a lot of debt, and debt has got to be paid back. That gut is running a conservation law. It believes the sixes are a finite stock being drawn down, so a long drought must be repaid. Handing you the algebra will not touch that belief, so we're going to drag it instead. Set "trials already failed" to anything you like and watch the remaining-wait curve.

1 · predict before you drag anything you've failed 9 rolls running, no six. is the next six any closer now? closer farther same 2 · drag j — try to move the curve P m = 1…7 rolls after the j failures 3 · the branches that survive j failures P(X>j) = qj qʲ·p q qʲ·q·p q qʲ·q²·p 4 · two sentences, two questions TEN FAILS · SCRATCH q¹⁰ = 0.1615 TRUE NEXT FAIL · GIVEN 9 q = 0.833 TRUE nine failures changed which question you're allowed to ask — not the coin.
predict first — pick closer / farther / same
pick what your gut believes, honestly
the reveal scores you against the maths
then try to break the curve yourself
Fig. 20. the die doesn't know it owes you. Predict first, then drag j — trials already failed — from 0 to 40. The remaining-wait bars and their dashed j=0 ghost never separate; the max-difference readout is genuinely recomputed each drag and stays 0.000000. The tree shows why: every surviving branch already carries a factor of qj, which is exactly what P(X>j) divides out. Nine failures didn't shrink the debt — there was never a debt. They only changed which question you're allowed to ask.

The curve in Fig 20 didn't move, yet the gut still feels owed. That feeling hides two different questions wearing one sentence: a whole streak of failures is genuinely rare, while the next single failure is utterly ordinary — still 5 chances in 6 on a die. Build the streak yourself and watch only one of those two numbers fall.

q 0.8333 run = q1 0.8333 next = q
1 · commit: after nine fails, the next fail is…
commit: pick your gut
From scratch (qk) — the whole streak, judged before you start. Rare, and it plummets toward zero.
From here (q) — the next single fail, judged from where you stand. Ordinary, and it never budges.
One failure so far: both sentences read the same, 0.8333. Press fail again and watch them split.
Fig. 21. First commit: after nine straight failures, is the next failure less likely, the same, or more likely? Your gut says less — a six is surely due. Now press fail again to grow the streak and watch two honest numbers split. The whole run judged from scratch is qk: it starts at 0.8333 for a single fail and plummets — by ten failures it reads 0.1615, the run really is rare looking forward from the start. But the next single fail judged from here is just q = 0.8333, pinned to the gold line, unmoved however long the streak runs. Both true, two different questions: the run's a-priori rarity never leaks into the next step, because the die keeps no ledger of your losses.

It doesn't move. Not flatter, not shifted, not shrunk. Identical, at every setting. Now the algebra lands as an explanation instead of a contradiction, because conditioning is just Chapter 5's move: shrink the world to what you know, then re-measure. So P(X > j+m | X > j) = q^(j+m) / q^j = q^m. Put the die's nine misses through that: the chance of ten misses in a row, 0.1615, divided by the chance of the nine you already have, 0.1938, is 0.8333 — which is q itself, not a penny more. The q^j cancels, and look at why it cancels. Every branch of the shrunken world starts with those j failures, so every branch carries the factor q^j, and that is precisely the number you divide by. It cancels because the failures are in the past, and the past you paid for is not an asset the die is holding for you. That's memorylessness: after any number of failures, the wait ahead of you is a brand-new geometric. Finally, let's separate the two sentences the gambler's fallacy fuses together, because the whole fallacy lives in that collision. "Ten failures in a row from scratch" has probability q¹⁰ = 0.1615, five times longer odds than the single next roll, and it shrinks every time you extend the streak. "The next one fails, given nine already did" has probability q, which is 0.8333 on a die and completely ordinary. Both are true. They are simply different questions. Those nine failures didn't change the die. They changed which question you're allowed to ask about it.

One thing left: how long is the wait, on average? The standard derivation needs an infinite series you won't have for another four chapters, so most books just hand you the answer and ask for trust. We're not doing that. We're going to earn it with the property we just proved, and nothing else. Run the machine a million times and sort the runs by what happened on trial 1. Every single run spent trial 1, no matter what. That's a 1, for everybody. In a fraction p of runs it's over right there — on a die, about 166,667 of the million. In the remaining fraction q, some 833,333 runs, that trial is gone and you're facing an identical brand-new machine with the same average E, by the fresh-start property we just proved. So the average total is E = 1 + qE. And there's the trick: E appears on both sides, because the problem contains a copy of itself.

600 runs charge / run
click to charge trial 1
600 runs enter; trial 1 not yet paid.
aha: trial 1 costs the same 1 for everyone.
Fig. 22. Step through it: every run pays trial 1 (charge 1.000, no exceptions), then trial 1 forks the population by p and q — and the q-branch turns out to be an identical copy of the same machine; zoom in as many times as you like and it never changes. That self-reference is E = 1 + qE, one line of algebra from E = 1/p = 6. Now drag the fulcrum under the geometric PMF: it balances only at n = 6, even though n = 1 is the single most likely roll — a long thin tail carries outsized leverage.

Solve it and you're done: E(1 − q) = 1, so E[X] = 1/p. With p = 1/6 that is 1 divided by 1/6, so a fair die needs 6 rolls for a six, on average. That matches the gut, so let's break the gut with the thing it can't explain. The geometric PMF is strictly decreasing, which means the tallest bar of all sits at n = 1. So the single most likely wait for a six is one roll, with probability 1/6, or 0.167 — higher than any other single number. And yet the average is 6. Both of those are true at once, and here's the resolution. The mean isn't the peak. The mean is the balance point, and this distribution has a long thin tail dragging off to the right. Those rare 20-roll waits carry only about 0.005 of the probability each, but they sit so far out that they haul the fulcrum all the way to 6, even though the tallest bar sits at 1. Most likely is not the same as average. Seen once on a lopsided PMF like this one, that distinction is worth more than the formula.

Step back and look at what we actually did: we never left the Bernoulli atom. Not once. One atom, and two questions. Ask "how many?" and you go up to the binomial hub, where two roads fork out of the same law. The √λ diagnostic is the sign at that junction. Many trials, moderate p, and the wall at zero is offstage: that's the normal. Rare, with λ pinned, and everything but λ burns off: that's the Poisson. Ask "how long?" and you turn the same atom sideways, find that only one sequence answers, and you get the geometric. That's the whole chapter. And please notice that neither law is a thing to remember — you can rebuild both of them from the atom in about twenty seconds.

n? t? λ λ=20→NOR qᵐ→e^(−λt) p n,p N λ Geo Ch12
hover a node to replay its derivation
blue = the atom's territory — one trial, param p
gold ★ = you are here — Poisson & Geometric
grey/dashed = the other path (Normal) & Ch12 ahead
e appears because (1−λ/n)ⁿ→e^(−λ) is the definition of e
Fig. 23. Zoom out and the whole chapter is one atom asked two questions. "How many?" climbs to the binomial hub, which forks on √λ — pin λ=np small and rare, and you fall into Poisson; let λ grow and its own spread flattens it toward Normal instead (type a λ above and watch the correct branch light). "How long?" just turns the same atom sideways into Geometric — no new atom, no new coin. Hover either gold node and watch it get rebuilt from scratch rather than recalled. The grey Ch12 node waits with one dial left at its stop: drag slice→0 and the geometric bars you already trust melt into the exponential curve you haven't met yet — the weld qᵐ→e^(−λt) was already sitting on the Poisson edge, derived, just unlit.

There's the tree, with our two rungs lit gold. And there's exactly one dial we never turned: the slice width. We chopped the window very fine, down to 60,000 millisecond slices, and then we stopped, because we still needed each slice to be a countable trial. Chapter 12 turns that dial all the way down to zero, and both of the things you already own come along with it. The wait stops being "which trial?" and becomes "what time?" That's the exponential. And the weld we built, q^m → e^(−λt), is already the Poisson-exponential duality that chapter is named for. It's derived. It's on the page. We just haven't said its name yet.

iolinked.com
Written by Ajai Raj