11Rare Events & First Successes
We ended the last chapter on a failure. We had just built the bell out of the binomial, and then we broke it on purpose. Drag p toward zero and the fitted normal slides its left tail across the 0 line, handing out real probability to negative counts. You cannot get −2 successes, so that red slab is nonsense, and it is the wall this chapter climbs. But the wall is not where most books tell you it is, and I want to be honest about that first. They say the normal "breaks for rare events", and it simply does not. Ten thousand components, each failing today with probability 0.002, is about as rare as anything you'll meet, and the bell fits those ten thousand components beautifully. So rareness is not the test. The real test is a single number you can compute in your head, and finding it is the first thing we do. That number is λ, the expected count, and it turns out to be the only thing that matters. Here's the plan. We'll pin λ and slice the world finer and finer, and everything about n and p will burn off except λ. That's the Poisson, the law of rare counts. The price of that collapse is a deadlock between two opposing forces, and the number they deadlock on is e. Not e imported from calculus. e manufactured right here, in a question about broken parts. Then we'll take the same yes/no atom and turn it sideways. Instead of "how many successes?", we'll ask "how long until the first one?" That's the geometric, and its formula has no coefficient at all — for a reason you'll see rather than memorize. And at the end the two halves will turn out to be the same sentence read twice.
01Where the bell really breaks
We'll test that folk wisdom against two situations side by side, both of them unmistakably rare. The left panel holds 10,000 components, each failing today with probability p = 0.002. The right panel holds just 20 components, each failing with p = 0.05. Both have a tiny p. Both describe something that hardly ever happens to any single part. So if "rare breaks the bell" were the real rule, both of these fits should come out as garbage. Fit the normal to each panel and look at what actually happens.
The fit on the ten-thousand-component panel is excellent. The fit on the twenty-component panel is rubbish. Same rareness, opposite verdicts. So the folk rule is worse than useless: it makes you distrust a fit that is perfectly fine, and it tells you nothing about the one that failed. What you need is a number you can compute rather than a feeling, and we already own the two facts that produce it. Chapter 10 gave us the binomial's center, μ = np, and its spread, σ = √(npq). Now think about what actually goes wrong when the bell fails: it puts mass below zero. So the only question that matters is how far zero sits from the center, measured the way we always measure distance on a bell — in standard deviations. The center sits at np, and we'll call that number λ, the expected count: on the left panel it is 10,000 × 0.002 = 20 failures expected today. When the event is rare, q is nearly 1, so the spread σ = √(npq) lands very close to √λ. On the left panel that spread is √(10,000 × 0.002 × 0.998) = 4.47, and √20 = 4.47 as well. Divide, and there it is.
Zero sits √λ standard deviations below the mean. That's the whole diagnostic, one division of two things we already had. So the bell does not fear a small p. It fears a small λ. Read the division at each scale and you can see why. At λ = 100 the wall at zero is 10σ away, so far out that no bell has any mass there. At λ = 20 the wall is 4.5σ away, still safely offstage. But at λ = 1 the wall is only 1σ from the peak, and a fat red slab of the curve is now parked on impossible negative counts. That's why our two rare cases split so cleanly. The ten-thousand-component panel had λ = 20. The twenty-component panel had λ = 1, because 20 × 0.05 = 1. And notice what the failure just handed us, because it wasn't a warning — it handed us λ, and told us that λ is the quantity the rare world is actually organised around.
02Slice the window until only λ survives
If λ is the only thing that matters, we should stop treating n and p as two independent knobs, and build the picture the rare world actually has. Take a physical window: one minute at a call centre, chopped into 60 one-second slices. Each second is now a tiny yes/no question, a Bernoulli trial: a call either arrives in that second or it doesn't. That gives n = 60 trials, each with some small p. But why seconds? Nothing forces that choice. Chop the same minute into milliseconds and now n = 60,000, a thousand times more slices, each carrying a p a thousand times smaller. Chop it into microseconds and n runs to sixty million. The calls did not change. Only our own bookkeeping changed. Click through the slicing dial and watch.
Here is the point, and it closes the trap before any algebra opens it. The slicing is our ruler, not the world's. The calls land where they land, whether we chop the minute into seconds or into microseconds. So ask yourself: if the answer changed when we switched from seconds to milliseconds, would that be an answer about the calls — or an answer about our chopping? It would be an answer about us. So whatever we compute has to be slicing-proof. Now look at what survives that demand. As we chop finer, n explodes toward infinity and p collapses toward zero, so both of our knobs are running away at once. But their product doesn't move at all. Going from 60 slices to 60,000 multiplies n by a thousand and cuts p to a thousandth, which is exactly what "same minute, finer ruler" means. Watch the three readouts as you drive the dial.
⇒ one-parameter law
n is meaningless on its own. p is meaningless on its own. Only np = λ is slicing-proof, and it stays pinned at the number we actually care about: the expected count of calls in that minute. So the law of rare events must have exactly one knob, and we haven't computed a thing yet. That isn't a trick of algebra. It's forced on us by the fact that we invented the slicing in the first place. Formally, we're setting p = λ/n in the binomial and letting n be ours to choose: 20 calls expected in the minute, spread over 60,000 millisecond slices, means p = 20/60,000 = 0.00033 for a slice. Every choice of n is a perfectly legal binomial, describing the same minute with the same expected count. Now let's ask that sliced window a question.
03The deadlock that makes e
We'll ask the easiest question there is: what's the chance that nothing happens in the whole window? No calls all minute. For that, every single slice has to miss. The slices are independent, so by the multiplication rule we just multiply n identical factors together. Each slice misses with probability 1 − λ/n, which for our 60,000 millisecond slices is 1 − 0.00033 = 0.99967 per slice. So the answer is (1 − λ/n)ⁿ, and now we chop finer and finer and see where it goes. Before you look, I want you to commit, because there's a real fight inside that expression. Two forces are pulling in opposite directions. Each slice is getting safer: the base 1 − λ/n is racing up toward 1, and one raised to anything is one, so this force says the answer is 1. But there are ever more slices to survive: the exponent n is running to infinity, and survival gets harder every time you have to do it again, so this force says the answer is 0. Pick one. Is it 1, is it 0, or is it something in between? Then run it.
Neither force wins. They deadlock, and the number they deadlock on already has a name. Watch the digits freeze column by column as you slice, with the expected count pinned at one: ten slices give 0.34868, a hundred slices give 0.36603, a thousand give 0.36770, and keep slicing and it tightens onto 0.367879 and stops moving. That's not "roughly a third". That's a specific, unbudging number the two forces settle on and never leave again. And the stalemate value is one you already know: it is 1/e, one divided by 2.71828. The number e is defined as the value this tug-of-war lands on, so read the arrow the other way round from how you've probably always read it. Nobody imported e from calculus into probability. We asked an ordinary question about a call centre, and the question manufactured e on the spot. That's why a transcendental constant shows up in a formula about broken components. It was never a guest. It's what "safer and safer, but more and more times" equals.
And it generalizes immediately. Pin λ = 3 instead of the λ = 1 we used above, and the same deadlock lands on 0.0498, which is exactly e^(−3) — one divided by 2.71828 cubed. In general, (1 − λ/n)ⁿ settles on e^(−λ) for whatever λ you pinned. Drag λ and watch the numeric ladder chase the curve.
Fig 6 left you with P(zero) = e^(−λ) for any window you name. So name a bigger window. A call centre takes 20 calls in an average minute — what does it expect over three minutes, and does the chance of a completely quiet stretch change as the window grows?
So you now own a real result, and you own it before you've met the law it belongs to. P(zero events in the window) = e^(−λ). Put the call centre's 20 calls a minute into it: a silent minute has probability 0.000000002, two chances in a billion. That is the whole Poisson at k = 0, and everything else in the law will turn out to be bookkeeping stacked on top of this one number. Let's go get the rest.
04Building the law from the split
Same window, harder question: what's the chance of exactly k events, not zero? We already have the machinery, because the count of successes in n independent slices is a binomial. So substitute p = λ/n into the binomial PMF and push the same slicing to the limit. Now, this is the exact spot where most readers quietly give up and start memorizing, so let me name the panic before it bites. You're about to watch n run to infinity while k just sits there, unmoved, and that feels illegal. Here's the answer, and it's one line. k is the question we asked — "exactly three calls this minute". n is the ruler we chose, whether that ruler is 60 slices or 60,000. The ruler goes. The question stays. The second thing that usually gets skipped is the regrouping. We are not going to send factors off to their limits one at a time on a whim. We sort first, into four boxes with all the n-dependence exposed, and only then does each box become answerable by inspection. Step through it.
Look at how little actually happened there. Box one is k factors on top, each within k of n, divided by k factors of n. With n = 1000 slices and k = 3 calls that is (1000 × 999 × 998) / 1000³ = 0.997, and as the slices multiply box one goes to 1. Box two has no n in it at all, so box two cannot move. Box three is our deadlock, and it goes to e^(−λ). Box four is a fixed number of factors, each going to 1. So three boxes are trivial, and the one that isn't is the thing we built in the last section. What's left standing is the Poisson distribution:
P(X = k) = e^(−λ) · λ^k / k! Read it as three jobs. The e^(−λ) is the price of nothing happening, the deadlock we earned in the last section. The λ^k is the k events landing. The k! divides out the orderings, exactly the way the binomial coefficient's denominator always did. Run all three jobs at λ = 3 and k = 2: 0.0498 × 9 / 2 = 0.224, so something that averages three per window arrives exactly twice about 22% of the time. One parameter, λ. No n, no p. Now let's audit the newborn, because I'm not going to hand you a PMF without checking that it really is one. The first duty of any distribution is that its probabilities sum to 1. Here's the honest argument for this one, and it needs no new tools at all. Every binomial we squeezed summed to 1, and we pinned λ, so no mass ran off to infinity while we squeezed. A limit of laws that each sum to 1 sums to 1. Add the bars up yourself and watch the tank fill.
One honesty note, flagged rather than hidden. There is a closed-form reason the Poisson sums to one: the series Σ λ^k/k! equals e^λ exactly, which cancels the e^(−λ) out front. Set λ = 1 and you can watch that happen: 1 + 1 + 1/2 + 1/6 + 1/24 + ... = 2.71828, which is e itself. The series is genuinely lovely and we'll meet it later in the book. I'm not using it now, because you haven't been handed it and I won't pay for a result with a tool you don't own. The limit argument above is complete on its own. Now the second half of the audit: where is this thing centered, and how wide is it? Ride the same limit through the same door. The binomial's mean is np, and we're holding np = λ, so the Poisson mean is λ exactly. The binomial's variance is npq, which is λ·q. We're driving p → 0, so q → 1, and the Poisson variance is λ too.
Mean = variance = λ. This is the fingerprint, and it deserves a moment because it should feel wrong. For every distribution you've met so far, center and spread were two independent dials. You can slide a bell sideways without changing its width. Here one knob does both. And now you can see exactly why, rather than filing it under trivia. A binomial keeps its two dials apart because np and npq genuinely differ. But in the rare limit q → 1, so the difference between them is erased: on our ten thousand components at p = 0.002, the mean np is 20 and the variance npq is 19.96, the same number to a rounding. The two dials fuse because failure became free. So σ = √λ, which at λ = 20 is 4.47, and that is the same √λ from the very first figure of this chapter — the same number that told us where the bell dies. Carry this fingerprint, because it's how you look at real counts and say "the mean is 4 but the variance is 30, so that is not Poisson — something is clustering."
05Using it, and the thing that isn't in it
Time to close the loop the chapter opened. Ten thousand components, each failing today with probability 0.002, so 10,000 × 0.002 gives λ = 20 failures expected. We now have three laws for this one situation: the exact binomial, its normal limit from Chapter 10, and its Poisson limit from this chapter. Before you run the lab, predict which of the three curves is going to miss the simulated data. Take that guess seriously, then let real code simulate a thousand days and settle it.
None of them miss. All three curves lie on top of each other, and the diagnostic told you that would happen before you ran anything: √20 ≈ 4.5, so the wall at zero is 4.5σ offstage and the bell is perfectly legal here. Now hit the λ = 1 button and re-run. The normal starts bleeding mass across zero. The Poisson doesn't blink. So here's your rule, and notice that it is not a new fact — it's the number you've had since the first figure. Compute λ. Take √λ. If √λ is comfortably big, say 3 or more, which means a λ of 9 or more, the wall at zero is out of sight and the bell is legal and cheap. If √λ is small, you're in the counting-rare-things world and Poisson is the law. The binomial is always exactly right, and usually more work than you need.
normal_fit and binomial_fit can't even be called on it — they raise, live, right there — while poisson_fit refits instantly from λ alone. Edit λ and watch it happen again.Fig 13 makes Poisson look universal: hand it a rate and you're done. But independence rode in quietly through those slices, and the real world does not always grant independence. So which rare counts are truly Poisson, and how would you catch a count that only pretends to be one?
And now the part that almost nobody says out loud, which is the real reason Poisson exists. Look at the formula again and notice what isn't in it. There's no n. There's no p. One knob. That isn't cosmetic tidiness. It is precisely what the collapse bought us. Somebody tells you a page has about 3 typos, so go ahead and try to use the binomial on that. How many trials were there — keystrokes, words, milliseconds of typing? And what is the per-trial error probability supposed to be? Nobody knows, nobody can measure it, and the question isn't even well posed. There is no n, and there never was one. But λ = 3 is sitting right there, already measured, from the simple act of counting 300 typos across 100 pages. n and p were never observable. λ always is. That's why one single law counts typos per page, decays per second, calls per minute, and packet retries per hour. If you think Poisson is "the binomial's approximation", you will freeze the first time someone hands you a rate. It isn't an approximation. It's the law you can actually fit to the world.
06Turn the atom sideways
We're going to keep the same Bernoulli atom and change nothing about it. All we change is the question. So far we have always asked: how many successes in n trials? Now ask the sibling question: how many trials until the first success? Roll a die until you get a six. Retry a packet until it lands. Interview candidates until one finally says yes. Fire the slots and watch the wait land somewhere different every time.
Now let's compute one: what's the chance the first success lands on trial 4? Your hand is going to reach for a binomial coefficient, because we've spent a chapter and a half welding C(n,k) onto every Bernoulli question. Resist that for one second and just write down what has to happen. Trial 1 fails. Trial 2 fails. Trial 3 fails. Trial 4 succeeds. F F F S. Now here's the question that dissolves the reflex: how many other sequences give a first success on trial 4? Go hunting, and try to find even one.
There aren't any. FSFS? That one succeeded on trial 2, so it's a different event entirely. Rearranging is exactly what you're not allowed to do here, because the word "first" nails down every position at once. The failures aren't scattered anywhere. They're locked into the first three slots, and the success is locked into the fourth. So the coefficient isn't missing. The coefficient is 1. That leaves one plain product of independent factors, q · q · q · p. For a die that product is (5/6) × (5/6) × (5/6) × (1/6) = 0.0965, so a first six on trial 4 happens about a tenth of the time. In general the geometric distribution is P(X = n) = q^(n−1) · p. Take that seriously as a skill, because it transfers: from now on, when a PMF arrives with no coefficient out front, read the absence as evidence that the event has exactly one arrangement.
Two warnings before we use it, and both are places readers get hurt silently. The first is that n quietly changed jobs one paragraph ago. In the binomial, n was fixed — the trials you decided to run — and the count was the random thing. In the geometric, the count of successes is fixed at exactly one, and the number of trials is what's random. Same letter, opposite job, one page apart. The second warning is that "the geometric distribution" is one name for two different distributions. Some books count trials including the success, so X ∈ {1, 2, 3, …} and P(X = n) = q^(n−1)p. That's ours, and we're committing to it. Other books count failures before the success, so X ∈ {0, 1, 2, …} and P(X = n) = q^n p. Both are correct. They're off by exactly one: feed the same die and the same n = 4 to both, and our version returns 0.0965 while theirs returns 0.0804. If you learn one and read the other you'll assume the error is yours. It isn't. Always check the support before you trust a formula.
Now the question you'll actually ask in real life. Not "will it succeed on trial 4?", but "will I still be waiting after m trials?" Your first instinct is to add up the whole tail: P(X > m) is the sum of q^(n−1)p over every n beyond m, forever. That's an infinite series, you don't have the machinery to evaluate it, and you conclude the question is hard. It isn't. Chapter 4 already handed you the answer, because "still waiting" is one thing and one thing only: all m of those trials missed. Watch both roads run side by side.
One line: P(X > m) = q^m. The brute-force sum crawls toward that number, and the complement road just arrives there. Put the die in it: (5/6)⁴ = 0.482, so after four rolls you're still waiting almost half the time. This is the complement reflex, worth naming as a habit, not a trick. When the thing you want is a tangle of many ways, price its opposite, which is usually one way. And the CDF comes along free: P(X ≤ m) = 1 − q^m, or 0.518 for those four rolls.
07The weld, the fresh start, and the wait
Now I owe you an explanation. Why are Poisson and geometric in the same chapter? Most books put them together because both are discrete and both smell like Bernoulli, which is a filing decision, not a reason. Here's the real reason. Say this one out loud: "the first success comes after trial m." Now say this one: "there were zero successes in the first m trials." Those are the same sentence. Not similar. They are the same event, described from opposite ends. The first is a wait statement. The second is a count statement. They are true in exactly the same circumstances. Look at one timeline with one shading and ask it both questions.
One shading answered both. So the two halves of this chapter just shook hands, and the algebra confirms what your eye saw. On the sliced window, the geometric tail is q^m = (1 − λ/n)^m. The shaded stretch is a fraction m/n of the window, so the expected count inside it is λ·(m/n) — shade half of our minute and that is 20 × 0.5 = 10 calls expected. Push the same deadlock and you get e^(−λ·m/n), here 0.000045, which is exactly the Poisson P(0) we built in section three. Counts and waits aren't two kinds of question. They're one question read from opposite ends of the same timeline, and they shook hands using e. Hold onto that, because Chapter 12 is going to give it a fancy name and you'll already have derived it.
Next, the most stubborn wrong instinct in all of probability, and it does not yield to a formula. You've rolled nine times and no six. So here's the question: is your six closer now than it was before you started? Commit to an answer first. Your gut says yes, obviously — nine misses is a lot of debt, and debt has got to be paid back. That gut is running a conservation law. It believes the sixes are a finite stock being drawn down, so a long drought must be repaid. Handing you the algebra will not touch that belief, so we're going to drag it instead. Set "trials already failed" to anything you like and watch the remaining-wait curve.
The curve in Fig 20 didn't move, yet the gut still feels owed. That feeling hides two different questions wearing one sentence: a whole streak of failures is genuinely rare, while the next single failure is utterly ordinary — still 5 chances in 6 on a die. Build the streak yourself and watch only one of those two numbers fall.
It doesn't move. Not flatter, not shifted, not shrunk. Identical, at every setting. Now the algebra lands as an explanation instead of a contradiction, because conditioning is just Chapter 5's move: shrink the world to what you know, then re-measure. So P(X > j+m | X > j) = q^(j+m) / q^j = q^m. Put the die's nine misses through that: the chance of ten misses in a row, 0.1615, divided by the chance of the nine you already have, 0.1938, is 0.8333 — which is q itself, not a penny more. The q^j cancels, and look at why it cancels. Every branch of the shrunken world starts with those j failures, so every branch carries the factor q^j, and that is precisely the number you divide by. It cancels because the failures are in the past, and the past you paid for is not an asset the die is holding for you. That's memorylessness: after any number of failures, the wait ahead of you is a brand-new geometric. Finally, let's separate the two sentences the gambler's fallacy fuses together, because the whole fallacy lives in that collision. "Ten failures in a row from scratch" has probability q¹⁰ = 0.1615, five times longer odds than the single next roll, and it shrinks every time you extend the streak. "The next one fails, given nine already did" has probability q, which is 0.8333 on a die and completely ordinary. Both are true. They are simply different questions. Those nine failures didn't change the die. They changed which question you're allowed to ask about it.
One thing left: how long is the wait, on average? The standard derivation needs an infinite series you won't have for another four chapters, so most books just hand you the answer and ask for trust. We're not doing that. We're going to earn it with the property we just proved, and nothing else. Run the machine a million times and sort the runs by what happened on trial 1. Every single run spent trial 1, no matter what. That's a 1, for everybody. In a fraction p of runs it's over right there — on a die, about 166,667 of the million. In the remaining fraction q, some 833,333 runs, that trial is gone and you're facing an identical brand-new machine with the same average E, by the fresh-start property we just proved. So the average total is E = 1 + qE. And there's the trick: E appears on both sides, because the problem contains a copy of itself.
Solve it and you're done: E(1 − q) = 1, so E[X] = 1/p. With p = 1/6 that is 1 divided by 1/6, so a fair die needs 6 rolls for a six, on average. That matches the gut, so let's break the gut with the thing it can't explain. The geometric PMF is strictly decreasing, which means the tallest bar of all sits at n = 1. So the single most likely wait for a six is one roll, with probability 1/6, or 0.167 — higher than any other single number. And yet the average is 6. Both of those are true at once, and here's the resolution. The mean isn't the peak. The mean is the balance point, and this distribution has a long thin tail dragging off to the right. Those rare 20-roll waits carry only about 0.005 of the probability each, but they sit so far out that they haul the fulcrum all the way to 6, even though the tallest bar sits at 1. Most likely is not the same as average. Seen once on a lopsided PMF like this one, that distinction is worth more than the formula.
Step back and look at what we actually did: we never left the Bernoulli atom. Not once. One atom, and two questions. Ask "how many?" and you go up to the binomial hub, where two roads fork out of the same law. The √λ diagnostic is the sign at that junction. Many trials, moderate p, and the wall at zero is offstage: that's the normal. Rare, with λ pinned, and everything but λ burns off: that's the Poisson. Ask "how long?" and you turn the same atom sideways, find that only one sequence answers, and you get the geometric. That's the whole chapter. And please notice that neither law is a thing to remember — you can rebuild both of them from the atom in about twenty seconds.
There's the tree, with our two rungs lit gold. And there's exactly one dial we never turned: the slice width. We chopped the window very fine, down to 60,000 millisecond slices, and then we stopped, because we still needed each slice to be a countable trial. Chapter 12 turns that dial all the way down to zero, and both of the things you already own come along with it. The wait stops being "which trial?" and becomes "what time?" That's the exponential. And the weld we built, q^m → e^(−λt), is already the Poisson-exponential duality that chapter is named for. It's derived. It's on the page. We just haven't said its name yet.