◈ probability mapProb · Ch 10/16
Probability from the coin flip up · chapter 10

10The Bell as a Limit

The bell curve is not a new object. It is what the binomial becomes when you flip enough coins, and that one sentence is what this whole chapter exists to earn. Last chapter ended at a fork. We had built the binomial — the hub — out of nothing but a stack of coin-flip atoms, and one question decided the road out of it: which regime are you in? Push it toward many trials at a moderate p and the jagged bars fuse into a smooth bell. That is the road we take now. That bell is also the first and cleanest instance of the deepest result in the subject, the Central Limit Theorem, and we get to meet it on ground we already own. Here's the plan, and I want to go slowly, because this is one of the widest jumps in the book. First we feel the pain the bell exists to kill: summing a big binomial by hand. Then we drive an n-slider, watch the bars melt into a curve in front of us, and read that curve's formula not as a spell but as a shape. Then we crack open the one piece everyone memorizes and nobody explains: the 1/(σ√2π) out front. It turns out to be exactly the number that forces the total area to equal 1, which is why a wider bell must be a shorter one. Then we learn where the limit breaks by breaking it: too few trials, or a probability so small the bell spills mass below zero. That spill is exactly what hands us the next chapter. And we finish with the technique the whole edifice runs on: standardizing, the z-score that collapses every bell onto one curve. It turns a 41-term binomial sum into two lookups. One curve, reached by flipping enough coins. That's the chapter.

01The hub, and the wall it hits

Let's stand exactly where Chapter 9 left us. We built the binomial, and we read off its two honest summary numbers with no new work. Stack n Bernoulli atoms, each with success probability p, and the count of successes has mean np and variance npq. So its spread is √(npq). Put numbers in and the symbols come alive: 400 fair flips sit at a center of 200 and a spread of 10, since √(400 × 0.5 × 0.5) = √100 = 10. Hold those two numbers, the center and the width, because the limit we're about to meet hands them right back to us. Here's the binomial again, with its center and spread marked. Drag n and watch the whole pile ride to np and fatten to √(npq).

0 n k · number of successes σ span μ n=12 · p=0.50 μ = np 6.00 σ = √(npq) 1.73 q = 1 − p q=0.50
center 6.0 · spread 1.73
The limit ahead needs no new fitting — it just inherits this center and width.
Fig. 1. Recall the rung under your feet: run n independent trials, each succeeding with chance p (and failing with chance q = 1−p), and the pile of probabilities over k successes is the binomial distribution. Drag n and p and watch two numbers ride along with the pile: its center μ = np (the mean count of successes) and its spread σ = √(npq) (the standard deviation). Push n up and the pile fattens outward but re-centers cleanly on μ every time — those two numbers were never a separate fit, they fall straight out of n and p. That's the whole point of standing here before the limit: the bell curve about to appear inherits this exact center and this exact width, no new fitting required.

Now the wall, and here is its cost up front: one ordinary question about a range takes 41 separate terms. Here's why. The binomial gives you one exact count beautifully — the chance of precisely k successes. But real questions almost never ask for one count. They ask for a range. What's the chance a fair 400-flip run lands between 190 and 230 heads? To answer that from the binomial you have to add up every count in the window, 190, 191, …, 230, and each one of those terms is a giant coefficient C(400,k) times two tiny powers. Watch the summation grind through it.

no terms counted yet — hit Step μ = np = 200 165 235 k = heads out of 400 flips (p = q = ½) 190 230
Count the highlighted window's 41 terms, one giant product at a time.
counted 0 / 41 terms
running sum = 0.00000
0 down · 41 to go — by hand
41 terms, by hand — there's a cheaper way coming.
Fig. 2. The gold band marks the question: heads between 190 and 230 out of 400 fair-coin flips (μ = np = 200 is the peak). Hit Step and watch one monster term — C(400,k)·pk·q400−k — get computed and folded into the running sum; hit Grind and watch all 41 of them grind through, by hand, one after another. That's the wall: 41 separate giant calculations just to answer one range question — exactly why the next figures trade this brute force for a single area read under a smooth curve.

That's 41 products of numbers big enough to overflow a calculator, all to answer one ordinary question. A wider range would be worse. There has to be a better way, and your eye already found it back at the fork. That binomial isn't a random jumble of bars. It has a shape: smooth, symmetric, humped, the kind of thing a curve describes far better than a list does. And reading the area under a curve is a whole lot cheaper than adding 41 bars. So let's chase that curve down.

02The bars turn into a curve

Here's the move, and it's the keystone of the chapter: don't fight the binomial, let it grow. Take the same fair-coin binomial and slide n upward. At n = 6 the bars are chunky and stair-stepped, with obvious daylight above their tops. But as n climbs, the steps get finer, the outline gets smoother, and a single clean curve starts tracing right along the tops of the bars. Before you slide it, commit to a guess. At what n can you no longer see daylight between the bars and the curve? Then drive it and watch the gap meter fall.

-2σ μ +2σ count k, lined up on one ruler: z = (k − μ) / σ flips n = 4 5 bars daylight eye limit 6.0% gap: bell − bar-top (% of peak height)
1 · predict: at what n does the daylight vanish?
Guess n, then slide the dial up
Why it works: the bell isn't a new object — it's the shape the binomial races toward, and it gets there fast.
Fig. 3. The convergence dial. Flip a fair coin n times and count the heads k; the coral bars are the exact binomial probabilities P(exactly k heads). To compare every n on one picture, each count is placed on a shared ruler — z = (k − μ) / σ, its distance from the mean measured in standard deviations, where μ = np is the average number of heads and σ = √(npq) is the spread (here p = q = ½). Each bar is drawn 1/σ wide and centred on k — it owns the strip from k−½ to k+½ (the ±0.5 continuity correction), and its height is scaled by σ so the total area stays 1. Measured that way, every binomial is held up against the same fixed blue standard bell, φ(z) = e^(−z²/2) / √(2π) — a smooth density, not a probability. Commit a guess for the n at which the bar-tops stop showing daylight beneath the curve, then drive the dial: the meter tracks the largest gap between any bar-top and the bell, and it falls past the eye-limit line by n ≈ 30. That is the whole point — the bell was never a new shape to memorise; it is simply the thing the binomial races toward, and it arrives surprisingly fast.

Wherever you guessed, the honest answer surprises most people: by around n = 30 the curve and the bars are already almost indistinguishable, and past that it only gets tighter. That's the phrase to hold: the binomial reaches the bell surprisingly quickly. So the smooth curve isn't a loose approximation we're settling for. It's a destination the binomial is racing toward. Which raises the real question: what is that curve, exactly? It's not a freehand sketch. It's a specific, named law, and we don't have to guess its center or its width, because the binomial already told us both. The fitted curve is the normal distribution with the same center np and the same spread √(npq). Watch it drop into place.

reused: μ=np · σ=√(npq) — no refit P(X=k) μ=0 σ 0 0 count of heads out of n flips, k
n=20 · p=0.50
μ=10.0 · σ=2.24
guess the shape, then flip the curve on
Fig. 4. Every bar is the exact binomial probability P(X=k) for n flips of a coin that lands heads with chance p — drag n or p and the dashed line marks μ=np, the bars' own center. Flip + curve on and a normal curve appears, built from nothing but that same μ and σ=√(npq) — no eyeballing, no separate fit. Keep dragging: the curve never falls behind, because it was never a sketch. It's a formula riding the bars' own numbers, which is exactly why the limit works for any n and p.

Read what just happened: the limit handed us the parameters for free. We didn't fit anything. We reused μ = np and σ = √(npq), the two numbers Chapter 9 already gave us — for 400 fair flips, a center of 200 and a width of 10 — and the normal curve with those two numbers lands right on the bars. That's the normal distribution, also called the Gaussian, and it's the single most important continuous distribution in all of probability. Time to look the curve itself in the eye.

03The bell, and the number out front

The normal's formula has a reputation for being scary, so let's defuse it by reading it as a shape, one piece at a time. The whole engine is a single exponent: e raised to −½ times ((x − μ)/σ)². That inner piece, (x − μ)/σ, is just "how many σ away from the center am I?" It is the distance from the peak, measured in widths. Put our 400-flip numbers in: at 230 heads, with center 200 and width 10, that distance is (230 − 200)/10 = 3 widths. Now square it so left and right count the same, flip the sign, and exponentiate: −½ × 3² = −4.5, and e^(−4.5) ≈ 0.011. The curve out there stands at about 1% of its peak. At the center the distance is zero, the exponent gives 1, and from there it falls off fast. That exponent IS the bell. Everything in front of it is bookkeeping, so hover each piece and watch its job light up.

f(x) = 1/(σ√2π) · e^( −½ z ² ) where z = (x−μ)/σ 0.399 0 0 (μ) pick a piece to see its job hover the formula or tap below
hover or tap a piece of the formula
pick a piece to see its job
aha: the exponent alone makes the bell shape — the constant just sets its height.
Fig. 5. The formula has exactly five moving parts. Hover — or tap a button — over z, , −½, e^z, or the constant and watch its job light up on the curve: z measures how far a point sits from the center in σ-units; squaring makes both sides equal; the minus pushes the exponent more negative the farther out you go; the exponential turns that into the actual curved ratio — 1 at the center, ≈0.325 of the peak at z=1.5 — and that ratio is the whole bell shape. The constant out front, 1/(σ√2π), never touches the shape at all. It only stretches the peak up to exactly 0.399 for this standard bell (σ=1), so the total area comes out to 1. Fig. 8 shows that peak sagging as σ grows. One exponent draws the bell; everything else is bookkeeping.

So the shape is controlled entirely by the two parameters inside that exponent. μ says where the peak sits. σ says how wide the fall-off is. That's the whole vocabulary of the curve, so go and feel both: move μ and the entire bell slides along the axis with its shape unchanged. Move σ and it widens or narrows around a fixed center.

density 0 4 7 10 13 16 x → σ μ
drag μ or σ — see what stays fixed
aha — every bell is this exact curve, just slid by μ and stretched by σ. Nothing else can change its shape.
Fig. 6. The curve is the normal density: height here is not a probability — only the area under a stretch of it is. Drag the gold peak dot and the whole bell slides sideways; that's μ, the mean — it moves the center and touches nothing else. Drag either green square — they always sit at exactly 60.65% of the peak height, no matter how wide the bell is — and the curve widens or narrows around a fixed center while its peak drops or rises to keep the total area equal to 1; that's σ, the standard deviation. Two knobs, and nothing else, produce every normal curve there is.

A probability on this curve is an area: the width of a strip times its height. Fig. 6 warned that the curve's height on its own is not a probability, and it left the real question hanging. That strip is the answer, and it's one you can grab with your hands. Now narrow the strip toward a single exact value and watch the number fall all the way to zero.

0.3989 width = 1.20σ −2σ −1σ μ (the point)
predict: as the strip narrows to a single line, the probability…
area = P(strip)
0.451
the probability
height at μ
0.3989
pinned
predict, then shrink the width
gold strip = area = probability; blue post = height. Shrink the gold and only it vanishes.
Fig. 7. The gold strip stands over the peak; its area is the probability of landing in that range. Predict what happens as it narrows, then drag strip width down: the area readout starts near 0.451, slides straight through the height’s value and keeps falling — while height at μ holds at 0.3989 the entire way. Hit collapse to a single line and the width goes to , the area to 0.000: a single exact value has probability zero. The height never was the chance — a point has no width, and only width carries probability. (Widen back to and the area reads 0.683 — the same 68% inside ±1σ from the empirical rule.)

Here's my favorite thing in this entire chapter, so take it first. The 1/(σ√2π) is nothing but the area-normaliser. The constant 1/(σ√2π) out front, the one piece of the formula everyone memorizes and nobody understands, is doing bookkeeping and nothing else. Here's the whole story. A probability density has one non-negotiable job: the total area under it must equal 1, because something has to happen. That's Chapter 3's normalization axiom, now in continuous form. The bare exponent doesn't have area 1 — it has some other area, and that area depends on σ. So you divide by exactly that area to fix it, and the divisor works out to σ√2π — for a width of 10 that's 10 × 2.5066 ≈ 25. Now, did you notice something odd while dragging σ? Making the bell wider also made it shorter: the peak sagged. The normaliser is exactly why. A wider bell, meaning a bigger σ, covers more ground, so to keep its area pinned at 1 it has to be shorter. Widen it and it must sink. Narrow it and it must spike. Watch the area readout hold dead still at 1.000 while the height and width trade off.

the value x · centred at the mean μ peak height 1/(σ√2π) 0.399 width at half-height 2.355 total area under the curve 1.000 measured by summing the curve height ⇄ width — a see-saw tall wide
σ=1 — the standard bell, area 1.000
The 1/(σ√2π) out front isn't decoration — it's the exact number that forces the shaded area to 1. Spread the bell wider and that number shrinks, so the peak must sag. That's why a wider bell is always a shorter one.
Fig. 8. Drag σ, the width of the bell. Widen it and the curve flattens — the peak height 1/(σ√2π) sags while the width at half-height grows: a see-saw, one up means the other down. And the whole time, the total area — summed straight off the curve — never budges from 1.000. That's the point of the constant out front: it is precisely the number that pins the area at 1, which is the real reason a wider bell must be a shorter one.

Sit with that, because it turns a memorized constant into something you could have derived yourself. The number out front was never magic. It's the price of making the curve a legal distribution, and that price is exactly why height and width are locked in a see-saw. That's the normal, understood from the ground up. Now let's find out where it stops working. A law you only ever see succeed is a law you don't really understand.

04Learning the limit by breaking it

The normal is the binomial's limit, and the word "limit" is a promise about large n. So the first way to break it is the obvious one: make n small. Dial it back down and the smooth curve stops fitting. At n = 4 or 5 the binomial is a handful of fat, chunky bars, and the bell drawn through them is a bad trace. A fit-error meter puts a number on how bad. The lesson isn't that the approximation is fragile. It's that the bell is a destination, and you have to travel far enough to arrive.

0.40 0 0 30 k — number of heads, out of n flips μ=np n = 30 fit error 0.8% Σ|bar−curve| good fit
n=30 — smooth, curve traces bars
drag n down toward 4–5: the bars go fat and stair-stepped and the gold curve stops tracing them — the bell is a destination you reach only after enough trials, not a shortcut you get for free.
Fig. 9. One dial: n, the number of coin flips (p fixed at 0.5, so μ=np always sits dead-center). The gold curve is the normal fitted with that same μ and σ²=npq — the binomial's own numbers, handed to the limit for free. Drag n down toward 4–5 and the bars go fat and stair-stepped while the curve stops tracing their tops; the fit-error meter climbs as the mismatch grows. That's the whole point: the normal isn't a stand-in you get automatically. It's a destination the binomial only reaches once n is large enough. That's why the working rule is np, nq ≥ 5: enough trials that the count is no longer coarse.

The second failure mode is sneakier and far more important, because it's the one that writes the next chapter. The normal curve is symmetric, and it stretches out to both sides forever, including into negative numbers. Usually that's harmless. For 400 fair flips the center np is 200 and one σ is 10, so the tail reaching below 0 is twenty widths out and carries essentially no mass. But now make the event rare: drag p toward 0. At p = 0.01 the center np drops to 4 while the width is still about 2, so the 0 line sits only two widths below the peak. The bell's left tail slides straight across it, and the curve starts assigning real probability to negative counts. That's nonsense, because you cannot get −2 successes. Watch the impossible region shade red.

0 μ k — number of successes (n = 25 trials) p = 0.50 μ = np = 12.50 σ = √(npq) = 2.50 spill ≈ 0.0% ↓ impossible negative counts
safe — bell sits clear of the wall
Rare events need a different limit — meet the Poisson next.
Fig. 10. Drag p toward zero. The mean μ = np slides left with it, and the bell — always perfectly symmetric — drags its left tail along for the ride. Past a certain point that tail slips under k = 0, the shaded wall, and starts crediting probability to a negative number of successes: an event that cannot happen. A symmetric curve has no way to know a hard floor is there. That is the tell that the normal approximation quietly fails for rare events — and exactly the gap the Poisson limit, next, is built to close.

That red spillover is not a bug in our figure. It's the normal approximation honestly telling you it's the wrong tool for rare events. A symmetric bell simply cannot describe something squashed hard against zero. So the binomial's common-and-many road leads to the normal. Its rare-and-many road needs a different limit, one that lives only on 0, 1, 2, … and never spills below. That limit is the Poisson, and it's Chapter 11. For now, let's pin down exactly when the normal is safe to use, so you're never caught applying it where it lies.

n < 10 — too coarse np, nq ≥ 5 p→0 spills p→1 spills 10 50 100 200 n — number of trials (log scale) 1 .5 0 p — chance of success SAFE at this (n,p) → Ch 11 the Poisson rescue
hover or tap a zone on the map
moderate p, big n — bell is legal
Both walls are one idea twice: np is the average success count, n(1−p) the average failure count — each needs room (≥5) for the bell's spread to fit before it hits the 0 or n edge.
Fig. 11. A map of every (n, p) pair: n trials along the bottom (log scale), p — the chance of success on each one — up the side. Hover or tap a zone. Left, n is under 10 and the count is too coarse for a smooth bell no matter what p is. Top and bottom, p sits too close to 0 or 1 and the bell's tail spills past a wall that isn't there — that's the small-p wedge pointing at Ch 11. Only the widening middle, moderate p with big enough n, is where the normal is actually legal. Knowing where you stand on this map is half of owning the tool.

Faced with a fresh n and p, an expert does not squint at the zone map in Fig. 11. They run two sums: compute np and nq, and demand that both clear 5. It is an and, never an or, and our 400 fair flips clear it easily: np = 200 and nq = 200. But a giant n paired with a rare p still fails: at p = 0.01, np = 4.

THE CASE — dealt n = 1000 p = 0.002 raw n, p only — call it, then reveal np = n × p np = ? 5 AND nq = n × (1−p) nq = ? 5 ? tap LEGAL or NOT LEGAL →
call it: legal or not?
It's an AND, not an OR: np guards too few successes, nq guards too few failures. A huge n with a rare p still fails if either one dips under 5.
Fig. 12. Deal a case, then call it before the machine does: tap LEGAL or NOT LEGAL. The reveal computes np and nq live, lights each one against the line of 5, and stamps a verdict. The opening deal is the trap the whole habit is for: n = 1000 feels enormous, but at p = 0.002, np = 2.0 — illegal, full stop. Deal again and you'll meet n = 8, p = 0.5: a perfectly fair coin, yet np = nq = 4.0, both under 5. The connector between the two lights says it plainly — it's an AND: both np and nq must clear 5, so one green light and one red is still a fail. Legality was never a feeling about a big n. It's two numbers, checked together.

There's the map, three zones. Moderate p, big n and the bell is dead-on. Tiny n and it's too coarse to fit. Small p and it spills past zero. Knowing when a tool is legal is half of owning it, so from here on assume we're in the safe zone. Now let's collect on the promise that started the chapter: turning that ugly 41-term sum into something you can do in two lookups.

05One curve to rule them all

Here's a practical problem you might not have seen coming. Every choice of μ and σ is a different bell, with its own center and its own width, and there are infinitely many of them. Back before computers, people answered normal questions by reading printed tables of areas. But you obviously can't print a separate table for every possible μ and every possible σ. That wall — too many bells to tabulate — is the reason the single most useful trick in the subject was invented. Look at the shelf of different normals and feel the problem.

the shelf of possible bells a table for each one 0 5 10 value of X shelf full — ∞ more exist
drag μ, σ then shelve it
each bell needs its own printed table of areas — there's no end to how many (μ,σ) pairs exist.
Fig. 13. Drag μ and σ — the white dashed curve previews the bell those two numbers make. Click shelve it and that bell locks in with its own colour, and a matching card appears on the right: a stand-in for the printed area table someone would have to make just for that (μ,σ). Keep going and the shelf fills, then keeps replacing its own cards while the counter climbs — because there is always one more (μ,σ) pair to try. You cannot tabulate infinity. That wall is exactly why the next figure collapses every bell onto one.

The trick is standardizing, and it's almost too simple. Take any value X from any normal. Instead of asking "what number is it?", ask "how many σ is it from the center?" That question has a formula: Z = (X − μ) / σ. Subtract the center to move the peak to zero, then divide by the width to measure what's left in σ-units. On our 400-flip bell, 230 heads becomes (230 − 200) / 10 = 3, so 230 is simply "three widths above the center". Do that to every point of any bell and the whole thing collapses onto one canonical curve: the standard normal, N(0, 1), centered at zero with width one. Drag X on a generic bell and watch it map to z on the standard one. Watch the shaded area stay identical on both.

N(μ=50, σ=10) N(0, 1) — the one curve x = 50.0 shaded area = 0.500 z = (x−μ)/σ = 0 shaded area = 0.500 z = 0.00
same area both sides · P = 0.500
why it works → re-scale the generic bell in σ-steps from μ and its shape vanishes: every normal lands on one standard curve, and the shaded probability is unchanged. That's why a single z-table answers them all.
Fig. 14. Two bells, one truth. On the left, a generic normal N(μ,σ) — pick any centre μ and any width σ. Drag the value x (or tap a preset) and it maps through z = (x−μ)/σ — the distance from the centre measured in σ-steps — onto the right, the single standard curve N(0,1). Notice what refuses to change: the green shaded area — the probability P — is the same number on both, no matter what μ and σ you chose. The z-map slides and stretches the bell until it becomes the one standard shape, and it does it without touching the area. That is the whole reason one function, one table, answers every normal question ever asked.

That's the whole game: the shaded probability is unchanged by the map. So a question about any normal becomes the same question about the one standard normal, and that single curve is all anyone ever had to tabulate. The area under the standard normal, from the far left up to some z, has a name and a symbol. It's the CDF of the standard normal, written Φ(z), the Greek capital phi. Read it as "the fraction of the bell lying to the left of z." So Φ(0) = 0.5, because exactly half the bell lies left of the center. Drag z and watch Φ fill.

N(0,1) — the standard normal curve -3 -2 -1 0 1 2 3 z shaded = Φ(z) Φ(z) = P(Z ≤ z) 0.5000 area to the left of z z = 0.00
Φ(0.00) = 0.5000 — half the area
Every normal reduces to this one curve — that's why only Φ needs a table.
Fig. 15. Drag the white dot across the standard normal curve, N(0,1) — the one bell every normal problem gets rescaled onto. Everything to the left of your dot shades in, and the readout is Φ(z), read "cap phi of z": the total area under the curve up to that point, which is exactly P(Z ≤ z), the probability of landing at or below z. Drag from the far left to the far right and watch Φ climb from 0.0000 (almost nothing is that far left) to 1.0000 (everything is at or below that far right) — passing through 0.5000 dead at z = 0, because the bell is perfectly symmetric there. Because every normal variable can be rescaled onto this exact curve, one table of Φ(z) values answers every normal-probability question there is.

So Φ is just area-to-the-left, and every normal probability is a lookup of it. Before we use it, let's build the ruler you'll carry for the rest of your life, because these three numbers come up constantly. How much of a bell's mass sits within 1 σ of the center? How much of it sits within 2 widths, and how much within 3? Commit to a guess for each one before you look.

-3σ -2σ -1σ μ μ ± 1σ tail outside ±1σ: —
you guessed 50% — hit Reveal to check
almost all the mass sits inside ±3σ — that's why a 4–5σ reading is a shock.
Fig. 16. Pick a band — ±1σ, ±2σ, or ±3σ — and the gold lines snap straight to it on the bell. Drag your guess to predict what share of the mass sits inside that band, then hit Reveal: the region shades green and the true figure appears — 68.27% within 1σ, 95.45% within 2σ, 99.73% within 3σ. Watch the tail outside readout as you climb the bands: it collapses from 31.73% to 4.55% to just 0.27% left over. That's the whole ruler in one picture — and it's why a value sitting out at 4–5σ isn't just "unusual," it's a genuine shock: almost the entire bell already lives inside 3σ, so there's barely anything left to explain a reading that far out.

Fig. 16 handed you the ruler for the whole bell, so now point it at a single value. Any measurement, whether a test score or a height, sits some number of σ from the centre, and that number is its z-score. Read Φ(z) and you know the exact fraction it beats: its percentile. A score one width above the mean has z = 1 and Φ(1) ≈ 0.84, so it beats about 84% of everyone. That works in any units at all.

score ~ N(μ=100, σ=15) — an IQ-style test drag the dot, or use the slider ↓ -3σ -2σ -1σ μ 55 70 85 100 115 130 145 z = +1.00 · Φ(z) = 0.8413 beats 84.13% of the population
you guessed 50% — hit Reveal to check
Same z, same rarity: a score, a height, a delivery time — whatever the units, z rulers away from μ always means the same Φ(z).
Fig. 17. An IQ-style score, μ = 100, σ = 15. Drag the dot — or the your score X slider — and two numbers compute live: z = (X−μ)/σ, your place in σ-rulers from the centre, and Φ(z), the fraction of the population you beat. At +1σ (115) that's 84.13%. Before you look, predict: a score above μ (130) beats what fraction? Then hit Reveal — it lands at 97.72%, the same Φ(2) number no matter what you're measuring. Now drag all the way to +3.5σ (152.5) and watch the dot flare: you're beating 99.98% of everyone — only ~0.02% score higher. That's the transferable instinct: a value's z-score is its rarity, so σ is the one ruler that makes any measurement, in any units, instantly comparable.

Here's that ruler, and it's famous enough to have a nickname: the 68–95–99.7 rule. About 68% of a bell's mass sits within one σ of the center. About 95% of it sits within two widths. About 99.7% sits within three. So almost everything a normal ever does happens within of the center, which is why a value 4 or 5 σ out is treated as a genuine shock. Now the mechanic that pays off the chapter. A real question is a range: what's the chance X lands between a and b? On the standard curve that's the area between two z-lines. And an area between two points is just the left-area up to the far one minus the left-area up to the near one. So the answer is Φ(b) − Φ(a): two lookups, any interval. Drag the two edges and read it off.

density — φ(z) drag the two markers ↔ Φ(a) = 0.1587 Φ(b) = 0.8413 a = -1.00 cumulative — Φ(z) = P(Z ≤ z) 1 .5 0 Δ z — standard units
Φ(1.00) − Φ(-1.00) = 0.6827
That's why any interval is only two lookups — an area between points is a subtraction of left-areas.
Fig. 18. Drag the blue marker a and the green marker b along the bell: the gold band between them is the chance X lands in that range. Below, the same two points sit on the S-shaped cumulative curve Φ(z) — and the gold bracket on the right shows their vertical gap is exactly that band's area. Try ±1σ: Φ(1)−Φ(−1) = 0.6827, the familiar "68%" rule, now visibly just one subtraction of two left-areas.

You have now seen every piece of the machine: standardize an edge to z, read Φ, subtract the two. But watching a pipeline run is not the same as running it yourself. So here you drive it, on fresh numbers you have never met. Compute each z yourself, read the two Φ values, and subtract for the answer.

YOUR PROBLEM given  μ = 30    σ = 4 find  P( 32 < X < 38 ) espresso pour · seconds 1 z_a = (a − μ) / σ ( 32 − 30 ) / 4 = +0.50 2 z_b = (b − μ) / σ ( 38 − 30 ) / 4 = +2.00 3 read Φ off the table → Φ(z_a) = 0.6915   Φ(z_b) = 0.9772 4 answer = Φ(z_b) − Φ(z_a) 0.9772 − 0.6915 = 0.2857 Φ — standard-normal table z Φ(z) 0.50 0.6915 1.00 0.8413 1.50 0.9332 2.00 0.9772 2.50 0.9938 for −z:  Φ(−z) = 1 − Φ(z)
compute z_a and z_b to begin
Fig. 19. No longer watching — you run it. Hit new numbers for a fresh case (an espresso pour with μ=30, σ=4: find P between 32 and 38 s). Do the four moves yourself: type z_a=(a−μ)/σ+0.50, then z_b+2.00 — each step turns green only when your arithmetic checks. The table then lights the two rows it hands you, Φ(z_a)=0.6915 and Φ(z_b)=0.9772; predict the answer and click subtract — 0.9772−0.6915 = 0.2857. Cycle to the commute case (both z's negative: z_a=−2.00, z_b=−0.50, and Φ(−0.50)=0.3085 via 1−Φ(0.50)) so the sign has to earn its place. That's the whole normal machine: two subtractions for the z's, two table reads, one final subtraction. You own the workflow now, not just the story.

One honest wrinkle before we cash in, and I won't skip it. The binomial is discrete: it lives on whole numbers, drawn as bars. The normal is continuous: one smooth curve. When you lay the curve over the bars, the bar sitting at count k isn't a spike at k. It's a block that really spans from k − 0.5 to k + 0.5 — the bar for 230 heads runs from 229.5 to 230.5. So if you want the normal to capture the bar for k, you take the area under the curve from k − 0.5 to k + 0.5, not just up to k. That half-unit fudge is the continuity correction: widen your range by 0.5 on each side before you standardize. Toggle it and watch the approximation tighten.

4 5 6 7 8 9 10 11 12 13 14 15 16 discrete outcome k μ = 10 n = 20 p = 0.5 μ = 10 σ = 2.236 k = 12 bar: [11.5, 12.5] cutoff x = 12.0 z = +0.894 exact ΣP = 0.8684 normal ∫Φ = 0.8146 error = -0.0538
chops the bar in half — misses mass
a bar at k is a half-unit block, spanning [k−0.5, k+0.5] — not a spike sitting exactly at k. Stop the curve at k and you slice that block clean in two.
Fig. 20. The bars are the exact Binomial(n=20, p=0.5) distribution; the curve is the normal that approximates it, with μ=np=10 and σ=√(npq)=2.236. Pick a bar with k, then toggle the ±0.5 correction: with it off, the shading stops dead in the middle of the k-bar — the normal estimate misses the exact binomial sum by several percent. Switch it on and the cutoff snaps to the bar's true right edge, k+0.5, the shading finally covers the whole block, and the error collapses toward zero. That's the whole trick: a discrete bar at k is a half-unit-wide block, not a spike sitting exactly at k.

A small correction that makes a real difference, especially near the tails and for smaller n. With that in hand you now have the entire toolkit: fit the normal, standardize to z, add the ±0.5, and subtract two Φ-values. Four steps. Let's finally kill the monster from the top of the chapter.

06The payoff — and you are here

Back to the wall. We wanted the chance a fair 400-flip run lands between 190 and 230 heads, and the binomial demanded a 41-term sum of monstrous numbers. Watch it collapse. The center is μ = np = 200, and the spread is σ = √(npq) = 10. Standardize both edges with the ±0.5 correction: 190 becomes (189.5 − 200)/10, which is z = −1.05, and 230 becomes (230.5 − 200)/10, which is z = 3.05. The answer is one subtraction, Φ(3.05) − Φ(−1.05) = 0.9989 − 0.1469 ≈ 0.852. Two lookups instead of forty-one products.

binomial path μ=200, σ=10 · P(190≤X≤230) terms: 41 189.5 230.5 normal path z = (x − 200) / 10 lookups: 0 z=−1.05 z=3.05 Φ=.1469 Φ=.9989
Φ(z) = the area under the bell to the left of z — a number you look up, not compute.
step 1 / 5
μ=200, σ=10 → 41 raw terms
keep stepping to reach the payoff →
Fig. 21. Same question, two routes. Step through it: the binomial path needs all 41 raw terms — one for every count from 190 to 230 — while the normal path corrects the edges by ±0.5, standardizes them to z with μ=200, σ=10, and reads off just two table values. Φ(3.05) − Φ(−1.05) = 0.8520, the identical answer the 41-term sum would have given — for the price of two lookups and a subtraction. That collapse is why standardizing is the technique: it's what makes any normal problem tractable at all.

Forty-one terms down to two. That's the payoff the entire chapter was built for, and it's why standardizing is the technique of the subject. It isn't a convenience. It is the thing that makes normal problems tractable at all. But I never want you to take a shortcut on faith, so let's prove it. Below, real code computes the exact 41-term binomial sum, then the normal estimate without the correction, then the estimate with it, and we see how close each one lands.

the calculation — it actually runs: exact = sum(binom(400,k,.5) for k in range(190,231)) mu, s = 200, 10 lo = norm.cdf(230,mu,s) - norm.cdf(190,mu,s) hi = norm.cdf(230.5,mu,s) - norm.cdf(189.5,mu,s) check: |hi-exact| < |lo-exact| » press Run to execute → stdout › exact — 41 terms, k=190..230 normal, no ±0.5: Δ — normal, ±0.5 corr: Δ — which one lands closer?
predict, then press Run
Same 400 flips, two normal estimates — only one hands the curve the same 41 outcomes the sum counted.
Fig. 22. Proof in code, not just a picture. n = 400 flips, p = 0.5, so the expected head-count is μ = np = 200 and the spread is σ = √(npq) = 10. The left panel actually runs: exact is the real 41-term binomial sum for landing 190–230 heads — no shortcut, every one of the 41 exact probabilities added up. lo asks the normal CDF the same question by treating the count as continuous. But a coin-flip count is a whole number and the bell is a smooth curve — so hi asks it again after stretching each edge count by half a unit (189.5 to 230.5), the ±0.5 continuity correction that hands the curve the same integers the sum counted. Predict which wins, then press Run: the corrected estimate lands roughly 477× closer to the exact sum. That's why the normal curve isn't a fudge for the binomial — it's the truth those 41 exact terms converge to, and ±0.5 is the small tax that makes a continuous curve honest about counting whole numbers.

The corrected normal lands right on the exact binomial, to the decimal. So the approximation isn't a fudge. It's the truth in the limit, and the ±0.5 earns its keep. Now step back, because we've found something much bigger than a shortcut for coin flips. Why does the binomial become a bell? Because the binomial is a sum of many small independent thingsn atoms added up — and it turns out that being a sum is the only thing that matters. Add up a big pile of small independent effects, of almost any kind, and their total piles into a bell. That's the Central Limit Theorem, and it's why the normal is everywhere: measurement error, human heights, the sum of a hundred dice. Every one of those is a sum of many small pushes.

1 6 sum of the N pushes N = 1 push summed μ=3.5 σ=1.7 why dials do this too a misread ruler, a scale that drifts, a jittery clock — each is many tiny independent nudges summed. enough small parts, any shape → always land on one bell. the Central Limit Theorem, live in front of you.
N=1: still lumpy, not a bell
same 3 shapes, every time — only N changes. Watch the bars fuse into the gold line.
Fig. 23. Pick the shape of ONE small push — a fair die (flat), a skewed nudge (lopsided), or a two-point jolt (two separated spikes, as far from a bell as a shape gets). Now drag N, the number of pushes summed: at N=1 the bars are just the raw shape you picked. Push N up and watch the bars fuse toward the gold curve — the normal curve with μ=N·(mean of one push) and σ²=N·(its variance), the same μ=np, σ²=npq machine from the binomial, generalized to any small piece. Switch shapes — die, skewed, two spikes — and N still wins every time: the Central Limit Theorem never cared what shape you started with, only that you kept adding many small, independent pushes. That's why a ruler's misread, a scale's drift, a clock's jitter — any real measurement error, built from countless tiny independent nudges — always ends up Gaussian.

That's the cross-domain lock. The next time you see a bell in a lab, a spreadsheet, or a physics paper, you'll know what it almost always is: a sum of many independent little things. And you met that mechanism here, in its cleanest form, on coin flips. So let's fix our position on the map.

aha — that's why the fork mattered: one road reached the bell, the other needs Poisson Binomial (n, p) recall: μ=np, σ=√(npq) the fork moderate p · big n small p · rare YOU ARE HERE Normal N(np, npq) Poisson next · Ch 11 one hub, two roads — only one has reached the bell so far
Hover or tap the fork to trace both roads out of the Binomial hub.
normal reached — poisson still waits
The moderate-p, big-n road already reached the bell. The rare-p road is still dashed — that's next chapter's Poisson.
Fig. 24. The whole chapter, on one map. On the left sits the Binomial hub from Chapter 9 — recall its two numbers, μ=np and σ=√(npq), because the limit inherits them for free. At the fork, one question splits the road: moderate p, many trials, or rare p? The top road, already lit gold, is the one this chapter walked — it reaches the Normal, marked YOU ARE HERE, the binomial's common-and-many limit and the first instance of the Central Limit Theorem. The bottom road stays dashed and faint: it's rare p, and that's exactly where the Normal spilled probability below zero back in Fig. 10 — the road the bell can't finish. Hover or tap the fork: both threads light up, tracing back to the hub and forward to the still-dashed Poisson. That's the whole point of the fork — one road reached the bell; the other one is Chapter 11.

There's our rung, lit gold. We took the common-and-many road out of the hub and found the Normal: not a new invention, but the binomial's limit, and the first instance of the Central Limit Theorem. We can fit it, read its formula as a shape, explain the constant out front, name exactly where it breaks, and standardize any question down to two Φ-lookups. And the very place it broke, spilling mass below zero for rare events, is the door to the next chapter. The binomial's other road, the rare-and-many one, leads somewhere the bell can't follow: the Poisson. Turn the page.

iolinked.com
Written by Ajai Raj