◈ probability mapProb · Ch 12/16
Probability from the coin flip up · chapter 12

12The Waiting-Time Family

Last chapter ended on a map with one control still at a stop. We had chopped the window into finer and finer slices, and then we deliberately stopped chopping. We had to stop, because every slice still had to be a countable trial you could point at. This chapter turns that dial all the way down to zero. Once the slots dissolve, the geometric question "which trial?" becomes the question "what time?", and the answer is the exponential distribution. Almost nothing new gets imported. You already derived the hard part in Chapter 11, when the geometric tail q^m turned into e^(−λt) and the two halves of that chapter shook hands. That handshake has a name: the duality. Here's the plan. We build the survival function first, because it is the thing you can actually photograph, and we let the density fall out of it as a slope. Then we find that the whole law has one knob. One knob means one length, and that forces the mean and the spread to be the same number. Then comes the strange part. This thing does not age. A component that has run for a thousand hours has a thousand hours left, on average, exactly like a fresh one. Your gut will fight that, and it should, because your gut is right about real bulbs. One instrument sorts out which world is which, and it is the hazard rate. Then the keystone, which is the deepest thing here: no memory doesn't merely describe this curve. It builds it, and leaves room for nothing else. At the end we read four famous laws off one timeline of dots, without generating a single new random number.

01When the slots dissolve

Turn the dial we left alone. In Chapter 11 the window was chopped into slices, each slice a little Bernoulli trial, and the wait was a count of slices. The geometric lived on that grid. Now shrink the slices toward zero width. The grid disappears, and the bars have nowhere left to stand, so watch what happens to the picture.

density 0 1 2 3 trial number k time t (hours) λe^(−λt)
slots / hour1
slice width1.000 hr
one slot = one hour — pure geometric
AHA: same p, same q — the dial was Δt all along.
What you're looking at — the dial Ch 11 left untouched
bars = P(first arrival lands in this slice)
curve = λe^(−λt): λ = arrival rate, t = time
dots = a few fixed arrivals, just for scale

Fig-01 uses the exact per-slice p=1−e^(−λΔt) so the bars sit on the curve at every Δt; Ch 11's sliced p=λ/n is the same thing to first order and lands on the same limit (§02). Watch the axis flip: slot count k becomes time t. With this exact p, q^m always equals e^(−λt) exactly, at any Δt.

Fig. 1. One dial — the one Ch 11 left standing at "one hour." Drag it toward zero: the same p and q redraw the geometric bars over the same fixed arrivals, and the bar tops climb up to meet λe^(−λt) as their widths collapse. The axis flips from counting slots (k) to reading a clock (t), and the q^m tail you derived last chapter turns out to have been converging on e^(−λt) the whole time.

The bars fuse into a curve, and that looks like a cosmetic swap. It isn't one. Something breaks when the grid dissolves, and it breaks quietly. On the grid, "the wait is 3" meant slot three, and P(X = 3) was a real positive number you could compute. In flowing time, ask for the probability that the wait is exactly three hours and the answer is zero. Not small. Zero. Chapter 8 already told you why. A point carries no probability. Only an interval does, and a density is not a probability. So a reader carrying the discrete habit walks up, asks for P(T = 3), gets nothing back, and concludes the model is broken. The model is fine. The question was broken. Watch both readouts as you narrow the slot from an hour to a second.

P(this slot) 0.233 zeros: 0 POINT → dies P(first hour) 0.632 never moves INTERVAL → lives 0h 1h 2h 1 hr slot width: hour
slot still real — about 0.23 of it
What you're looking at — two questions racing one shrinking ruler
blue slot = "the wait lands in THIS exact interval." Narrow it far enough and the number collapses to 0 — a single point in flowing time carries no probability (Ch 8)
green band = "the wait lands somewhere in the first hour." That's an interval question, and it stays at 0.632 no matter how fine the ruler gets
the geometric already asked an interval question — P(X>k)=qk, "the first k slots all missed." Shrink the slots and that exact fact becomes S(t)=P(T>t)=e−λt
Fig. 2. Drag slot width from an hour down to a millisecond: the blue slot sitting on the 1‑hour mark keeps shrinking toward zero, while the green "somewhere in the first hour" region never moves off 0.632. The number dies; the area doesn't — "the probability of the wait" was always an area, and only the interval question survives the slots dissolving.

The probability of any one named slot slides to zero, while the probability of "somewhere in the first hour" doesn't move at all. The point survives nothing. The interval survives everything. So when the slots dissolve, only questions about intervals are left standing, and that tells us which Chapter 11 fact to carry across. We proved P(X > k) = q^k there, and that statement was already about an interval: it says the first k slots all missed. It is the one geometric fact built to survive a dissolving grid, so it is the one we take with us. In flowing time it becomes S(t) = P(T > t), the chance you are still waiting at time t. That is the survival function, and for waiting problems it is the main character. That claim probably sounds like bookkeeping, since S(t) = 1 − F(t) is just the CDF read from the other end. So let me show you why the survival function is the thing you can actually measure.

SEE COMPUTE 1 0 0 t=4 alive —/120 S(t)≈— λ 0 no photo yet expected noise ≈ ±—
drag t, then take a photo
One rack, counted twice
still-glowing bulbs, counted straight off the rack — SEE
two photo-counts, subtracted and divided — COMPUTE, and jumpy
rack dark (t=4): S(t) = P(T>t) = 1 − F(t), the CDF read backwards. Pays off 3 more times:meanthe racethe gamma
Fig. 3. Drag t and watch the rack — that's S(t), read straight off a photograph, no formula required. Press take 2 frames and the right panel differences two of those counts to guess the density: it's second-hand, and it jitters. Shrink Δt and watch the noise readout climb; only hundreds of photos, averaged, settle it near the dashed true value.

Here is the whole reason: you can photograph the survival curve, but you have to compute the density. A hundred and twenty bulbs, all lit at t = 0, and a camera pointed at the rack. S(t) is the fraction still glowing, and you count it straight off the photograph — sixty bulbs alight out of a hundred and twenty means S = 0.5. Now ask the same rack for the density. You'd have to take two frames, subtract the counts, and divide by the gap between them. That is noisier, it is derived, and it sits one step away from anything you actually saw. So the survival curve is the measurement and the density is the computation. Almost every textbook opens on the density, which is exactly backwards, and it costs the reader real money later. We are going the other way, and that choice pays off three more times before this chapter ends: at the mean, at the racing clocks, and at the gamma.

02Build the stream, and the law falls out

Now we build the thing that produces the events, by hand, rather than accept it as a word. Take a rate of λ = 2 arrivals an hour. Here's the first thing nobody checks: λ is measured in events per hour, so it can be 5 or 500, while a probability can never exceed 1. So the slice probability p cannot be λ. Chop the hour into n slices instead, and hand each slice its own independent coin. There is exactly one honest way to spread λ events across n slices: each slice takes p = λ/n, so that n slices times p per slice gives the rate back. That fraction is forced, not chosen. Try n = 2 and you get p = 1, which means two arrivals guaranteed, one per half-hour, and no randomness left at all. Chop the hour into minutes and p = 1/30. Chop it into seconds and p = 1/1800. Drag the slicing and watch the product.

the hour, sliced into n pieces n=— p=— np=— left half right half 0 min 60 min
click "try p = λ" to begin
What you're looking at — a rate forced into a probability
blue bar height in each slice = p, that slice's hit chance
gold dot = a real Bernoulli(p) draw for that slice, computed live
solid green at n=2 = p=1, every slice guaranteed to hit
aha — λ=2 is a rate (events/hour), never a probability, so "p=λ" fails on units. p=λ/n is the only split where n·p stays welded to λ as n climbs. Drag n high enough and the slices become separate coins — reroll the left half and the right half can't feel it. That's independent increments, earned by construction, not assumed.
Fig. 4. Try p = λ first — it fails, because 2 events/hour isn't a probability. Then drag slices n: p = λ/n collapses while n·p stays welded at 2.000. At n=2 every slice is a guaranteed hit; at n=60000 you're standing in Ch 11's Poisson limit again. Reroll left half to see the right half never move — that's independence, built, not assumed.

np never leaves 2. You have built this limit before: n → ∞, p → 0, with np = λ pinned. That is Chapter 11's Poisson setup, word for word. Same limit, new question. Poisson asked "how many?" and we are asking "how long?" The stream itself has a name, the Poisson process, and it is just a constant-rate run of events along a timeline. Now notice the gift sitting inside the construction, because it is worth more later than it costs now. Different slices are different coins, tossed independently, so what happens in one stretch of the timeline tells you nothing about a stretch that does not overlap it. That property is called independent increments, and we never had to declare it. We got it by building it.

Now ask the stream our surviving question: what is the chance you are still waiting at time t? Write it yourself before any algebra. In t hours there are nt slices. The wait exceeds t exactly when every one of those nt slices missed, and the slices are independent, so their probabilities multiply. Cut the hour into 3600 one-second slices and a three-hour wait is 10,800 misses in a row. You have just written q^k with k = nt. That is the geometric's own line, and the only new thing is what the exponent counts. Step it through.

STEP 0 · predict t hours → nt slices, and every one missed 0 t = 2.0 h nt nt = 6 × 2.0 = 12 slices only the first hour commit below — no algebra until you do S(t) = (1 − λ/n)^(nt) (1−λ/n)^(nt) = [ (1−λ/n)^n ]^t [ (1−λ/n)^n ]^t [ e^(−λ) ]^t S(t) = e^(−λt) the geometric's own line, wearing a new hat: qⁿᵏ₀ᵏ₁ qⁿᵏ₀ᵏ₁ ↑ you built this last chapter n = 6 n = 600 n = 60000 e^(−λ) 0.33490 0.36757 0.36788 0.36788 λ=1.0, t=2.0 → e^(−λt) = 0.13534
pick one — no algebra until you do
In t hours there are nt slices. T > t means every one of them missed. Write that probability — then we will.
What you're looking at — e^(−λt) is not a new formula, it is qᵏ with k counting the slices in t
blue — one slice of time. n = 6 per hour in this picture; the algebra's n runs to ∞
gold — the exponent nt: not decoration, just the COUNT of blue boxes in [0, t]
red — your wrong pick, drawn: the stretch that answer actually covers
pinned honesty: the limit is not re-proved here, it is re-used — same e, same reason
Fig. 5. qᵏ wearing a new hat. Commit before you read: in t hours there are nt slices, and T > t means every one missed — so S(t) = (1−λ/n)nt, which is just qk with q = 1−λ/n and k = nt. Regroup to [(1−λ/n)n]t, send the bracket alone to its Chapter 11 limit e^(−λ), and S(t) = e^(−λt) falls out. The e isn't there because exponentials are natural; it's there because of one limit you already watched freeze, digit by digit.

There it is: S(t) = (1 − λ/n)^(nt) = [(1 − λ/n)^n]^t → [e^(−λ)]^t = e^(−λt). Look at how little that cost. The bracket in the middle is the deadlock from Chapter 11 — the same tug-of-war, the same e, standing on the same rung a second time. So the e in e^(−λt) is not there because "exponentials are natural". It is there because of one specific limit you already watched freeze digit by digit. The law of waiting in a constant-rate stream is S(t) = e^(−λt), and at two arrivals an hour that says 13.5% of waits run past the first hour (e^(−2) = 0.135). The random variable T that obeys this law is called the exponential distribution, written Exp(λ). Don't take my word for the limit. Let's build the sliced stream in code and measure the fraction still waiting.

0 1 2 3 t 0 .5 1 S
while(Math.random()>=p) slices++;
t=.25
t=.5
t=1
t=2
pick λ, n — then run the trials
What you're looking at — a machine that never once evaluated an exponential
blue — the loop's own measured fraction still waiting, S(t)
gold dashed — the analytic e−λt, computed only for comparison
→ the row's "measured" column comes from counting sorted coin-flip trials, never from calling exp(). At n=4 the gap column reads red because the slicing is coarse; at n=2000 it locks green because the slices have thinned into (almost) continuous time — that thinning IS the limit Fig. 5 derived on paper.
Fig. 6. The loop above is deliberately dumb: for each of 20,000 trials it flips a p=λ/n coin, slice after slice, until the first success, and records the wait as slices·(1/n), the left edge of the successful slice — scan it, there is no exp() anywhere. Hit run and the table fills: the measured fraction of trials still waiting past each t lands beside e−λt, and the gap locks green when they agree. Flip to n=4 and the gap turns red — the gold grid marks the four coarse slices per hour, and the blue curve visibly steps instead of falling smoothly, because (1−λ/n)nt at a small n is not yet the limit. Push to n=2000 and the steps vanish, the gap drops under a thousandth, and the blue curve lies on top of the gold one. That is Fig. 5's algebra, checked by a program that only knows how to flip biased coins — it never evaluated an exponential, because the exponential is just what enough thin coins do.

The measured fraction lands on e^(−λt) to four decimals, and the machine only ever flipped biased coins. It never once evaluated an exponential. That is a receipt, not an argument. The argument was the algebra; the code just confirms nobody cheated.

03One knob is the whole distribution

We have the survival curve, so the density costs exactly one derivative. You already own f = F′ from Chapter 8 and S = 1 − F from ten minutes ago. Differentiate, and f(t) = −S′(t) = λe^(−λt). At two an hour, the density one hour in is 2 × 0.135 = 0.271. Now I want to stop on the λ out front, because it is the exact inch where readers start memorizing. Every book hands you λe^(−λt) whole, so you parse it as two glued facts: there is an e^(−λt) shape, and there is also a λ scale factor. That is not what it is. The λ is the chain rule pulling the exponent's coefficient out front, and it appears whether you want it or not. So drag along the survival curve and watch its steepness trace out the density.

S(t) — drag the white dot 1 0 |S′(t)| = 0.000 t = 0.00 ↓ the density, traced from that steepness 2 0 0 1/λ t → area = 0.000000
drag the dot to find the slope
Drag the white dot along S(t). Its tangent's tilt is read off as a numerical slope — not asserted — and each reading drops a dot into the panel below.
What you're looking at — the λ falls out, it isn't bolted on
blue — S(t) = e^(−λt), the survival curve; its steepness at each t is dragged out by hand
pale dots — each |slope| you read off, dropped straight down as a point of the density
green — f(t) = λe^(−λt) fading in on “match?”, landing exactly on your own dots
orange (drop λ) — e^(−λt) alone; the sweep tops out at 1/λ, not 1
Ch 10's callback — the bell needed 1/(σ√(2π)) bolted on to force its area to 1. Here nothing is bolted on: λ is the chain-rule factor of −S′(t), and the very same stroke that produces it is what forces this area to 1. Strip the λ (step “drop λ”) and the area drops to 1/λ — it stops being a density.
Fig. 7. the λ out front is the chain rule. Drag the dot along S(t) = e^(−λt) (λ = 2 here) and its tangent's numerical slope drops a dot below — that dot IS f(t), because |S′(t)| = λe^(−λt) by the chain rule, not a separate fact glued to the exponential. “match?” fades in the closed form on top of your own trace. “area” sweeps under it while the integral climbs to exactly 1 — that λ is what forces it. “drop λ” strips the front factor from e^(−λt): same curve shape, but the sweep only reaches 1/λ = 0.5 — it is no longer a density.

The density is nothing but the steepness of the thing you photographed on the bulb rack. Now the callback that closes the loop. In Chapter 10, the number out front of the bell was 1/(σ√(2π)), and its job was to force the total area to 1. The λ here does that same job, and you can check it on screen: ∫₀^∞ λe^(−λt) dt = 1 exactly. Drop the λ and the area comes out as 1/λ instead, which at two an hour is 0.5, not 1. Here is the difference. Chapter 10's constant had to be bolted on. This one falls out of a derivative. A term that could have been inevitable usually gets handed to you as arbitrary. Not here.

Next, the mean — and I want to run the units before the integral. λ is a rate, so its units are 1/time, which makes 1/λ a time. At two arrivals an hour that is half an hour, thirty minutes. Now look in the box, where λ is the only parameter there is. If you want this distribution to hand you a time — the mean, the SD, the median, the 90th percentile, anything — there is nothing else with units of time to build it from. Every one of those numbers is forced to be a plain number times 1/λ. That is the one-scale law, and it lets you predict before you compute: the mean and the SD must be proportional. Same constant, though? Run the integrals and see.

STEP 1 · the parameter box build a TIME out of what's in the box λ [ 1/time ] tap a chip → — test the tray to find it — mean variance ≈ 0.000 ≈ 0.000 max |Δ| = 0.000000 exponential: 1 knob · normal: 2 knobs (μ,σ) shape parameter: — (the gamma's r fills this, §7)
test a chip — is it a time?
What you're looking at — the only length λ can build
blue box — λ alone, stamped [1/time]; a RATE, not a time
gold tray — λ, λ², √λ get struck out; only 1/λ has units of time
green fills — the two integrals actually running, both landing on the constant 1
aha — λ is the only thing in the box, so every time this law makes is forced to be a number × 1/λ; the integral only had to say which number, and it said 1 twice.
Fig. 8. one knob, one length. λ sits alone in the box, stamped [1/time] — test λ, λ², √λ from the tray and the units line strikes each one out; only 1/λ survives, because it is the only combination with units of time. That forces every time this law can produce to be (a number) × 1/λ, mean and SD included — before any integral runs. Predict same-or-different, then watch the two integrals actually run: both constants land on 1. Drag λ in the stretch and a ghost of the λ=1 curve stays glued underneath at every setting — max deviation frozen at 0.000000. One knob, no shape parameter: that empty slot is exactly what the gamma's r fills in section 7.

You can now write the mean as 1/λ. But a rate carries a clock with it, and that is where the first trouble bites. A rate of 2 per hour is the same stream as 1/30 per minute: the arrivals never changed, only the clock did. The mean flips the units too, because λ counts events per unit time while 1/λ is time per event. That one stream has a mean wait of 0.5 hours, which is the same thing as 30 minutes. Switch the clock and both numbers must move together, or your answer turns to nonsense. Set a rate below, change the clock, and watch what stays fixed.

PREDICT · relabel the clock hours → minutes: what actually moves? λ · the rate 2 /hour E = 1/λ · mean wait 0.5 hour ▼ number got smaller ▲ number got bigger |← mean wait →| 0 0.5 h 1 h same physical length in every clock — only the labels move |← STILL the same length →| 30 min (real) 0.5 “min” — 60× too short
hr → min, does λ go…
…and the physical wait?
pick the clock
λ is per HOUR. Read the wait in minutes — but forget to convert λ.
2/hour pinned. Predict.
What you're looking at — one bus stream told on four clocks
blue — λ, the RATE [1/time]; flip the clock and its number scales one way (2/hour → 0.0333/min → 48/day)
gold — E = 1/λ, the mean WAIT [time]; the inverse flip sends its number the other way (0.5 h → 30 min → 1800 s)
red trap — a per-hour λ read as per-minute: the answer is 60× too short, and the algebra never complains
aha — the blue segment never moves: the wait is a fixed length of real time. E=1/λ genuinely inverts a rate into a time, so λ and its unit must always match.
Fig. 9. one rate, two clocks. A single fact is pinned — 2 buses per hour — and the highlighted mean-wait segment sits on a real timeline. First gut-call it: switching hours→minutes, does λ go up or down, and does the physical wait change? Reveal shows the trap: λ goes down (2 → 0.0333/min) while E=1/λ goes up in number (0.5 → 30) — yet the blue segment never moves, because the wait is a fixed length of real time. Flip the clock through sec/min/hour/day and watch λ and 1/λ co-transform (48/day ↔ 0.0208 day) with the segment glued in place. Then mix units: read a per-hour λ as per-minute and E collapses to 0.5 min — 60× short of the true 30 min, with no symptom in the algebra. That is the rung: E=1/λ genuinely inverts a rate into a time, so your λ and your clock must always agree.

E[T] = ∫₀^∞ t·λe^(−λt) dt = 1/λ, and the variance integral gives 1/λ², so SD = 1/λ. At two an hour the mean is 0.5 hours and the SD is 0.5 hours as well. Both constants came out as 1. Be clear about what did which job here. The proportionality was forced by the units, and you knew it before any integral ran. The equality is the exponential's own fingerprint, the same kind of signature as the Poisson's mean = variance = λ. And the one-scale law is something your eye can check directly. Drag λ and the entire picture just stretches sideways. Nothing else about it ever changes, because there is no second knob. The normal had μ and σ. This has one number and no shape parameter at all. Hold that hole in your hand. When the gamma turns up at the end of the chapter, its r is exactly the knob that is missing here.

Now some numbers that will fight you, and they fight you because of where you have just been. Chapter 10 was a symmetric bell, where the mean sits on the median and everything is tidy, so you imported that equality without noticing you imported anything. This curve is skewed. Solve S(t) = ½ and you get e^(−λt) = ½, so t = ln2/λ ≈ 0.693/λ, which at two an hour is about 21 minutes. That median sits to the left of the mean at 1/λ, half an hour. Here's the question that catches almost everyone. A bulb's mean life is 1000 hours, so what fraction of bulbs last longer than 1000 hours? Commit to a number before you look.

μ mean 1/λ = 1000 hr t½ = 693 hr · λ=1.00e-3/hr
you said 50%
drag, then lock to reveal
What you're looking at — why "average life" isn't the halfway point
gold = still burning past the MEAN (1/λ) — always 36.8%, no matter the bulb
green tick = the MEAN — the curve's balance point, hauled right by its long thin tail
blue handle = the MEDIAN = the half-life (ln2/λ) — grab it, λ updates live
the verdict box stamps your locked guess against the true 37% — and keeps it there
computed here, not asserted: e−1 = 0.3679 · that's the gate above (37%). e−0.5 = 0.6065 · 61% outlive half the mean life — not the "≈70%" a quick eyeball guesses.
Fig. 10. Chapter 10's bell trained you that mean = median. This curve is skewed: lock a guess, then watch 63% of bulbs already be dark at their own "average" lifetime — because the mean sits right of the median, hauled there by the long thin tail. Drag the half-life handle and λ=ln2/t½ updates the whole curve live.

Not half. S(1000) = e^(−1) = 0.368, so only 37% of bulbs outlast the average, and nearly two-thirds are already dark by 1000 hours. The long thin tail hauls the mean to the right while most of the mass sits early, which is the same balance-point argument that made the geometric's average six while its most likely wait was one. Run it the other way and it becomes a working tool. A physicist measures a half-life, drags that marker, and reads off λ = ln2/t½: a half-life of 1000 hours gives λ = 0.693/1000 = 0.000693 per hour. That inversion is how a decay constant gets into a textbook in the first place. One more, because we should fix it out loud: e^(−1/2) is 0.607, so half a mean life leaves 61% still alive, not the "about 70%" you'll hear people eyeball. Loyalty is to the subject.

Those numbers hide an everyday lesson worth feeling in the body. The median sits to the left of the mean, and only 37% of items outlast the average. Put it in human terms. Told the average wait is 10 minutes, your gut hears "usually about ten minutes". It is not. Most waits are short, a handful are enormous, and those rare monsters quietly own the average — about 63 people in a hundred are already served inside those ten minutes. So send a hundred people to wait, and watch where they actually land.

0 — min mean med
of 100 people, how many wait LONGER than the 10-min average?
lock a guess, then send 100 people
above avg: — · below: —
longest wait this round: —
What you're looking at — 100 lived waits, not a curve
blue dot = one person's actual wait, dropped where it really lands on the time axis
orange dot = a monster wait, far enough out that it's pinned at the chart's right edge
green line = the MEAN (1/λ) — the balance point the rare monsters haul rightward
dashed blue = the MEDIAN (ln2/λ) — splits the 100 dots into two equal halves
computed here, not asserted: e−1 = 0.3679 · ln2 = 0.6931 — so ~37% outlast the mean and the median sits at ~69% of it, for any λ.
Fig. 11. Send 100 people to wait, guessing first how many outlast the 10-minute average — your gut says near 50, but the count that comes back hovers close to 37, because a few monster waits (orange, off the right edge) haul the mean past most of the crowd. Re-roll and watch the bulk stay left every time while the longest single wait swings wildly. Drag λ and the mean and median rescale together, always in the same ~69% ratio; toggle "what most feel" to dim the mean and light up the median — the line that actually splits the crowd in half.

04The thing that doesn't age

Here's the property this whole family turns on, and it is the one your gut will refuse. Your bulb has a mean life of 1000 hours, and it has already burned for 1000 hours. How much life does it have left, on average? Every gut in the room says much less, because it is old, it is tired, it is closer to death. The answer is 1000 hours. Exactly a fresh bulb, with no discount for the thousand hours it already served. Don't take that from me. Grab the handle and try to make the remaining-life curve move.

1 0 0 1000 2000 5000 mean life 1000h · already run 1000h your guess: — h true answer: — h
already run 1000 h — how many more?
ran 1000h already — guess what's left
What you're looking at — a curve with no memory, and one that ages
blue = the remaining-life curve, recomputed live from wherever the handle sits
grey dashed = the untouched original curve, pinned underneath for comparison
blue turns orange only for the real bulb — its remaining-life curve genuinely sags
Aha: the exponential curve can't move because nothing in its formula records how long it's already run. Your fridge isn't exponential for the opposite reason — a filament's odds of failing climb as it ages, so its hazard genuinely remembers. Both are true; now you know which box each one lives in.
Fig. 12. A bulb that's already run 1000 hours — guess its remaining life, then check: the truth is welded at 1000.0 h. Drag s and the exponential's live curve never leaves its own grey ghost outline — the difference is 0 by algebra, no matter how hard you pull. Switch specimens: the real bulb visibly sags (its remaining life keeps shrinking), while uranium-238 and the call-centre's arrival gaps stay flat. Your gut was right about the bulb — that's exactly why real bulbs aren't exponential.

It never changes shape at any setting you choose — not flatter, not shifted, not shrunk. The algebra behind that is three lines, and Chapter 5 already wrote them for us. Conditioning shrinks the world to what you know, and there is one sub-step worth saying out loud that books skip. If you have outlasted t+s, then you certainly outlasted s, so the event {T > t+s} sits inside {T > s}. Their intersection is therefore just {T > t+s} itself. So P(T > t+s | T > s) = S(t+s)/S(s) = e^(−λ(t+s))/e^(−λs) = e^(−λt) = S(t). Put numbers in it: at two an hour, a wait already 3 hours old passes 4 hours with probability e^(−8)/e^(−6) = e^(−2) = 0.135, the same 0.135 a fresh wait has of passing one hour. The exponents subtract, and the past cancels out of the algebra exactly as literally as it cancels out of the physics. That is memorylessness: the thing does not age.

Now the honest part, because your intuition is not stupid and I am not going to overrule it. Your intuition is correct about real bulbs. Filaments wear out, bearings wear out, people wear out. So when a book says "the exponential is memoryless" as though that were a fact about bulbs, it is telling you something false about your fridge. Memorylessness is a fact about a constant-rate stream, not about hardware. What genuinely does not age is a radioactive nucleus. A uranium-238 atom that is five billion years old is statistically identical to one made this morning. Arrivals at a call centre do not age either, because nothing in the stream knows a call just happened. So keep both intuitions, sorted into the right boxes. The question is which world you are in, and there is an instrument that tells you.

FLAT — no wear t = 380h · alive S(t) = 68.4% each dot ≈ 40 of the 1000 bulbs ÷ ALL 1000 — f(t) 0.07% ÷ SURVIVORS — h(t) 0.10% h(t) = λe^(−λt) / e^(−λt) h(t) = λ — flat, forever tap the formula — watch the e's cancel
hazard flat: h(t) = λ, forever
Of the survivors alive at t, the fraction dying this instant is always λ — hour 10,000 faces the same risk as hour one. That flatness is memorylessness, made mechanical.
What you're looking at — one rack, divided two different ways
blue bar — deaths this instant ÷ the ORIGINAL 1000: f(t), which decays as the rack empties
gold bar — deaths this instant ÷ the SURVIVORS still alive: h(t)=f(t)/S(t), the hazard
the rack itself — gold dots still glowing, grey dots already dark, counted off directly
aha — divide by S(t), not the original count: for the exponential the e's cancel to a flat λ, so hour 10,000 faces exactly the risk of hour one
Fig. 13. the hazard, built from a rack of survivors. Two ratios race off the same rack: deaths this instant ÷ the original 1000 (f(t), which decays as the rack empties) against deaths this instant ÷ the survivors still alive (h(t)=f(t)/S(t)). For the exponential the e's in f(t) and S(t) cancel exactly, leaving a flat λ — tap the formula to watch it happen. Toggle FLAT/RISE/BATH to see hazards that genuinely age, then try the mystery specimen.

Go back to the rack and ask the question a survivor would ask. At hour t, how many bulbs die in the next minute? That is f(t)·dt·1000. But a bulb that is still glowing doesn't care about that number, because that number is watered down by every bulb that already burned out. The living bulb asks a sharper question: of us survivors, what fraction goes dark this minute? That is f(t)dt / S(t). The division isn't exotic — it is Chapter 5's conditioning, with the survivors as the new sample space. That quantity has a name, the hazard rate, and for the exponential it works out as λe^(−λt)/e^(−λt). The e's cancel. It's λ, flat, forever. For the 1000-hour bulb that is 1/1000 = 0.001 per hour, at hour one and at hour ten thousand alike. The bulb at hour ten thousand faces precisely the risk it faced in its first second. That is what "no memory" looks like when you can touch it. Now flip the toggle, because flat is only one curve in a world of possible curves. Rising hazard means the thing ages, and the bathtub curve — high, then flat, then climbing — is the real shape for most manufactured things. So here is your diagnostic, and it is the most transferable thing in this chapter: check the hazard. If it isn't flat, the exponential is a lie, and you now know that before you fit anything.

05No memory leaves no room

Fine — the exponential is memoryless. Surely it is just one of many, so go and build me a memoryless law that isn't exponential. I'll even give you the test as a physical gesture, so you don't need any algebra to run it. Take a survival curve and cut it at time s. Throw away everything to the left, then stretch what is left vertically by 1/S(s) — at two an hour, cutting at one hour means stretching by 1/0.135, about 7.4 times. That gesture is conditioning. It is exactly "given it survived to s", drawn instead of written. If the law is memoryless, the rescaled tail must land dead on top of the original curve, for every s you pick. Try it on a straight ramp, try it on a bathtub, then try it on e^(−λt). Then read what your own hand just did.

s = 1.80 S(t) = 1 − t/6 1 0 0 t = 6 drag the gold dot: cut here, stretch the tail
max gap 0.300000 — the tail misses
What you're looking at — one survival curve S(t) = P(wait > t), and the physical test for “no memory”: cut it at s, throw the left away, stretch what's left by 1/S(s). That stretch is the word “given”.
the original law, S(t) — the curve the tail must land on
the cut, stretched tail: S(s+t)/S(s) = the wait given it already survived to s
the red gap is memory — the curve telling you how long it's been waiting
failed attempts so far: 0 — that number climbing is the whole point
Fig. 14. Every book says it can be shown that the only continuous memoryless law is the exponential, and walks on. Here is the “only”, in your hands. Grab the gold dot at s: the curve is cut there, everything left of it is thrown away, and what remains is stretched upward by 1/S(s) — that stretch is not decoration, it is literally the word given, renormalising the survivors back to probability 1. Memoryless means the stretched tail lands exactly on the grey original. The ramp comes back steeper and its gap is exactly s/6 — the further you cut, the more it remembers. The bathtub fails the other way: having survived its infancy, its tail is flatter than the newborn curve. Draw your own and it fails too. Only e^(−λt) reads 0.000000, and stays there for every s you drag to. Then read back what your hand did — S(t+s)/S(s) = S(t), i.e. S(t+s) = S(t)·S(s) — commit to a guess, and watch six lines close the door: a = S(1) forces a², a³, every aⁿ, then √a, then every rational power, and continuity fills the rest. There was never a second law; there was never room for one. Your own choice of a is the λ, wearing a logarithm.

I just told you the formula is now disposable, that two words rebuild it. That is a strong claim, so do not take it on faith. Take the two words — constant rate and no memory — and build the whole distribution back yourself, one honest rung at a time. Pick the right next piece at each step and its reason locks in beside it. Pick a wrong one and you will see exactly why it cannot stand. By the end you will have rebuilt e^(−λt), its density, and its mean, with no formula in front of you.

constant rate λ no memory + 1 2 3 4 S, f, and the mean — from two words
Rung 1 of 4 — pick the answer below.
Two words are on the shelf. Rung 1: what does “no memory” force S(t) to obey? Pick one — a wrong pick tells you why it fails.
You are not watching a derivation — you are performing it. From the two words alone, choose each rung and the whole exponential falls out: its shape, its density, its mean. Own it, don't memorise it.
the two words — constant rate λ and no memory — the only inputs
a locked rung: correct, with the one line that makes it true
a wrong pick flashes red and tells you why — the distractors teach too
the target: S=e−λtf=λe−λtE[T]=1/λ
Fig. 15. The payoff promised you could regenerate the whole distribution from four words — constant rate λ, no memory. Prove it in your own hands: click your way down the four rungs, with no formula on screen to lean on. Rung 1 turns “no memory” into S(t+s)=S(t)·S(s) (adding time → multiplying probability); rung 2 forces the only continuous answer, S(t)=e−λt; rung 3 differentiates to the density f(t)=λe−λt; rung 4 integrates to the mean E[T]=1/λ. Every wrong tile teaches back — pick 1−λt and it reminds you a survival curve can’t go negative; pick e−λt² and it makes you re-check that (t+s)² spills a cross term and fails to split. Land all four and the verdict reads “four rungs, two words, no formula in sight.” You now own the exponential rather than remember it.

"The cut tail lands on the original" is S(t+s)/S(s) = S(t), which rearranges to S(t+s) = S(t)S(s). That is the whole of memorylessness, written as an equation about a curve. Read what it demands: the curve has to turn adding time into multiplying probability. How many functions can do that — dozens, a family, an infinite zoo? Exactly one of them can, and the proof runs to six lines with no calculus anywhere in it. Let a = S(1), some number strictly between 0 and 1; at two an hour it is 0.135. Then S(2) = S(1+1) = S(1)S(1) = a², which is 0.135² = 0.018, and S(3) = a³, and S(n) = aⁿ for every whole n. Now go the other way. S(½)S(½) = S(1) = a, so S(½) = a^(1/2). The same trick gives S(m/n) = a^(m/n) for every rational. The fractions m/n crowd in arbitrarily close to every point on the line, and S is an unbroken curve, so the in-between values get pinned too, with no freedom left anywhere. Therefore S(t) = a^t = e^(t·ln a) = e^(−λt), with λ = −ln a > 0 — and −ln(0.135) is 2, the rate we started from.

Sit with that for a second, because it is the deepest sentence in the chapter and it is usually thrown away in half a line. There was never a second law. There was never room for one. Memorylessness and the exponential are not two facts about one distribution. They are the same statement in two costumes. And the λ isn't even a real ingredient — it is your own choice of a = S(1) wearing a logarithm, since λ = −ln a. So you can throw the formula away now. Give me four words, constant rate and no memory, and you can rebuild e^(−λt) from scratch, plus the density, plus the mean, plus every number in this chapter. That is a thing you cannot lose.

06One stream, two readings

Now let's collect the debt Chapter 11 left on the table. You have been carrying T, the wait for the first event, from this chapter, and N(t), the count of events inside a window, from the last one. They feel like two different random variables from two different formulas. They aren't. There is only one random thing in this room, and it is dots on a timeline. T and N(t) are two questions asked of the same dots, so watch what happens when you drag one.

λ = 1.00 t = 0.60 0–1 1–2 2–3 3–4 4+ 0 2 4 6 8 0 1 2 3 4+ Poisson k=0: e^-λt(λt)⁰/0! = 0.548812 slicing limit → (1-λt/n)ⁿ = 0.548811
press a sentence to compare
What you're looking at — two roads, one number
blue bars — the real gaps between today's dots (top instrument)
green bars — how many dots a width-1 window catches, scanned along the whole line (bottom instrument)
gold outline — what Exponential(λ) and Poisson(λ) predict for those same bars; drag a dot and watch blue/green chase gold
gold band [0,t] — "the first dot lands after t" and "zero dots in [0,t]" light the identical shading, which is why e^(−λt) shows up in both formulas
Fig. 16. There is only one random thing on this timeline: the dots. Drag any dot, or click empty track to add one — the blue gap-histogram above and the green count-histogram below both twitch on the same tick, because they are two instruments reading the same dots, not two separate worlds. Press either sentence below: "T > t" (the first dot lands after t) and "N(t)=0" (no dots land in [0,t]) light the identical gold band and agree every time — click inside that band to drop a dot and watch both flip false together. The strip at the bottom runs the Poisson formula at k = 0 and the exponential's slicing-limit separately, and they land on the same six decimals: two roads, one number, because gaps and counts were always the same e−λt.

Both readouts twitch, because both instruments point at the same stream. Read the gaps between consecutive dots and the histogram fills in the exponential. Slide a window along and count the dots inside, and that histogram fills in the Poisson. Now shade [0, t] and say both sentences out loud. "The first dot lands after t." "No dots in [0, t]." Point at each one on the timeline. They are the same shaded region — not equivalent, identical. So P(T > t) = P(N(t) = 0), and that's the duality, sitting there as a matter of looking rather than deriving. Then check it against a formula you already own. Put k = 0 into Chapter 11's Poisson PMF with mean λt: e^(−λt)(λt)⁰/0! = e^(−λt). At two an hour over one hour that reads 0.135 × 1 / 1 = 0.135, the exact number the survival curve gave in section two by slicing the timeline and taking a limit. Two roads, one number. The duality isn't a slogan. It is a checkable fact, and you just checked it.

One more property before we build upward, and it is the cheapest rung in the chapter. Every book says "the gamma is the sum of r independent exponentials", and both of the important words in that sentence get smuggled past you for free. Why should the second gap be independent of the first? And why the same λ, when a dot just landed a moment ago? We already own both halves of that sentence. Identical: the hazard is flat forever, so the stream cannot know a dot just happened. Independent: different slices are different coins, which is the independent increments we built in section two. Put them together and the inter-arrival gaps are iid Exp(λ) — at two an hour, every gap has the same mean of 0.5 hours, whatever came before it. Click anywhere on the timeline and check it yourself.

choose where to stand click anywhere below to stand there slice A slice B 0 time here 0 forward wait ~ Exp(λ) Δ = 0 by algebra since last dot — units 0 2+
click the timeline to choose your spot
Click anywhere on the timeline — mid-gap or right on a dot. The forward wait is always the same Exp(λ), no matter where you stand.
What you're looking at — the two halves of "iid," cashed as receipts
blue dots — a fixed stream; click anywhere to stand there
gold marker — "you are here"; the left curve never moves from it
IDENTICAL / INDEPENDENT — click a receipt to replay the figure that paid for it
Fig. 17. the stream starts fresh at every instant. Stand anywhere on the timeline — mid-gap or on a dot — and the forward wait is still Exp(λ): deviation from the fresh-start curve is 0 by algebra, always. IDENTICAL comes from the flat hazard h=λ (Fig. 13: the stream can't know); INDEPENDENT comes from disjoint slices being separate coins (Fig. 4: they share nothing). Both words, now paid receipts: the inter-arrival gaps are iid Exp(λ).

That restart claim carries a quiet assumption: the rate λ is the same at every instant. Real arrivals often break it. A call centre is swamped at 9am and idle at 3am, so the gaps run short in the rush and long at night. Nothing here is aging. No part wears out, and the hazard check from before would happily pass. What is moving is the rate itself, and that is a second, different way for the model to be wrong. Pool every gap, fit one exponential, and see whether it holds.

rate λ(t) across one day · drag the rush up λ̄ pool every gap → fit one Exp(λ̄)
flat rate → one λ fits ✓
The gut says yes, the fit says no — then split it
white line — the rate λ(t) over the day; the dashed line is its average λ̄, held fixed as you drag
blue bars — the histogram of every inter-arrival gap; gold is the single Exp(λ̄) fitted to their mean
red caps — the excess: too many tiny gaps (rush) AND too many huge gaps (night). One λ misses both

Predict, then drag. At rush 0% the stream is a clean Poisson process and one exponential fits its gaps perfectly. Now drag the rush up: the same average rate λ̄ is unchanged, so your gut says the pooled gaps are still just random waits — they should still fit one exponential. They don't. Pooling averages two worlds into one number: the histogram becomes over-dispersed (CV climbs above 1.0), with red excess at both ends. Hit split by time of day and the mess resolves into two clean exponentials, each at its own λ. The family was never wrong about arrivals — a single constant rate was. Second diagnostic: the hazard can be flat (memoryless passes) while the rate is not steady over the window (this fails). Check both.

Fig. 18. when one rate is a lie. Drag rush intensity from 0% up: the dots bunch at the 8h and 18h peaks and thin out overnight, while the average rate λ̄ = 10/hr stays fixed. Watch the pooled gap histogram — at 0% it fits one Exp(λ̄) cleanly (CV ≈ 1.0), but crank the rush and it turns over-dispersed, CV climbing past 1.5, with red caps of excess at both tiny gaps (rush) and huge gaps (night). Then hit split by time of day: the exact same gaps regroup into a steep rush exponential (λ ≈ 25) and a shallow night one (λ ≈ 4), each snapping close to its own line. The hazard passes the whole time — nothing ages; it's the rate that isn't steady. Before trusting one exponential, check both.

Stand mid-gap, or stand right on top of a dot. Either way the clock resets to zero, and the forward view is statistically identical every single time. That's the restart property: the stream begins again at every instant, including the instant an event just happened. Nothing new was assumed there, because both halves were already paid for.

07Waiting for the rth, and the family closes

Let's ask a bigger question of the same dots. Stop waiting for the first event and wait for the rth one instead; call that time T_r. This is where readers quit, because the answer contains a t^(r−1) that looks conjured out of nothing, sitting next to an (r−1)! that looks like a fudge bolted on to make an integral behave. Neither is true. We are going to assemble the formula out of parts you already hold, exactly the way Chapter 9 read C(n,k)pᵏq^(n−k) factor by factor. Ask when the rth dot lands in the sliver [t, t+dt]. Two things must happen, on disjoint stretches of timeline, so independent increments lets us multiply them.

STEP 1 · the bridge when is the rth dot later than t? dt 0 t time e^(−λt)·(λt)²/2! λ·dt disjoint ⇒ independent increments ⇒ MULTIPLY Ch 11's shelf — the Poisson(λt) PMF P(k events in [0,t]) = e^(−λt)(λt)ᵏ/k! P(T₃ > t) = P(N(t) ≤ 2) f(t) = λ³ ·t² ·e^(−λt) /2!
duality: T₃ > t ⇔ N(t) ≤ 2
The rth dot is later than t exactly when at most r−1 dots landed in [0,t] — the Fig. 16 duality sentence, with r in place of 1.
What you're looking at — the gamma density, assembled from parts you already own
blue — the stretch [0,t] and the 2 dots that had to land first
gold — the sliver dt, where the rth dot lands right now: λ·dt
green — Ch 11's shelf: the Poisson(λt) PMF, borrowed whole, not re-derived
f(t) = λⁿ · tⁿ⁻¹ · e^(−λt) / (r−1)! for t ≥ 0 — the wait for the rth dot, mean r/λ
Fig. 19. the gamma, x-rayed factor by factor. Ask when the rth dot lands in the sliver [t, t+dt]: r−1 dots must already sit in [0,t] (Ch 11's Poisson answer, e^(−λt)(λt)r−1/(r−1)!) and one dot must land in the sliver (λ dt). Disjoint stretches, so independent increments lets you multiply — and out drops f(t) = λrtr−1e^(−λt)/(r−1)!. Set r = 1 and it collapses to λe^(−λt).

First, exactly r−1 dots landed in [0, t], and that's Chapter 11's Poisson PMF with mean λt, read straight off the shelf: e^(−λt)(λt)^(r−1)/(r−1)!. Second, one dot lands in the sliver, with probability λ dt — the rate itself. Multiply the two and you have it: f(t) = λ^r t^(r−1) e^(−λt) / (r−1)!. That's the gamma distribution, Gamma(r, λ). Every factor now has a name and nothing is arbitrary. The t^(r−1) is the r−1 events that had to already happen. The (r−1)! is Poisson's own denominator, not a fudge. The leading λ is "one event, right now", the same rate factor as before. Set r = 1 and the whole thing collapses back to λe^(−λt). And one name defused in a line: for whole r, the gamma function Γ(r) is just (r−1)!, so Γ(3) = 2! = 2. The capital Greek letter is what lets r go fractional later, and it is nothing else.

The gamma is not a new animal. It is the exponential with r turned up, and its mean is r/λ, which is just r gaps of mean 1/λ laid end to end: three gaps at two an hour give 3 × 0.5 = 1.5 hours. Chapter 15 will turn "averages add" into a theorem, so I'm flagging that I use it early. And remember the hole from section three, where the exponential had one knob and no shape. Turn r and watch the hole get filled.

λ=1.20 shape: — 0 t → mode=0.00 gaps, laid end to end → mean=r/λ=0.83 0
r=1: mode jammed at t=0 (the wall)
What you're looking at — the shape knob Fig. 8 promised, now turning
t^(r−1) → pushes the bump RIGHT: the r−1 earlier events must land first
e^(−λt) ← pulls it back LEFT: everything still decays at rate λ
mode=(r−1)/λ is where the two tugs balance; at r=1 there's no t^(r−1) term at all, so the mode is nailed to t=0
f(t)=λrtr−1e^(−λt)/(r−1)!, mean r/λ. Press verify: replot in units of λt and the curve sits exactly on a fixed ghost — deviation 0.000000, proof λ only stretches. Press lock to freeze r=1 and watch it collapse onto the exponential: Γ(1,λ)=Exp(λ), overlaid exactly.
Fig. 20. the shape knob that was missing. Section 3's exponential had exactly one dial, λ, and it could only stretch the picture — drag it here and the curve slides, but never reshapes. Drag r and watch the mode, (r−1)/λ, peel off the wall: at r=1 it's nailed to t=0 (waiting for one dot that lands instantly is the likeliest thing there is), by r=2 a mode has appeared, and by r=5 it's a fat bump — because five dots all arriving at once is genuinely unlikely. Press verify to replot in units of λt: the curve lands exactly on a fixed ghost, deviation 0.000000, proving λ never touched the shape. Press lock to freeze r=1 and confirm Γ(1,λ)=Exp(λ) — the family closes on the law you already knew.

At r = 1 the density piles up against zero, because the most likely wait for the very first dot is no wait at all. Turn r up and the pile lifts off the wall and becomes a bump, because waiting for three dots and having all three arrive instantly is genuinely unlikely. That's the shape parameter doing its job. There is one more trap here, and it hides inside a single word.

The word is "three". Three of what? Three events you must live through, or three servers racing to help you? The same number attaches to two opposite structures, and the formula alone will not tell you which one you are in. Only the picture will. Worse, your gut carries a hard prior that more means faster, which is right for one of these and exactly backwards for the other. Look at both before you write anything down, and commit to a guess.

draw it — no formulas yet TOP BOTTOM RING
Q1 — the SUM is
Q2 — FASTER is
draw it — press ▶ run it
no algebra yet — just watch
What you're looking at — the same three exponentials, wired two different ways
blue — TOP, the SUM: live through all three, Gamma(3,λ), mean 3/λ — 3× SLOWER
gold — BOTTOM, the MIN: three racing, first one wins, Exp(3λ), mean 1/(3λ) — 3× FASTER
the repair: "open three lanes" names the MIN, not the sum — lanes and gamma point at different pictures
same λ, same "three" — nine times apart, and only the picture tells you which is which
Fig. 21. which "three" did you mean? TOP lays three waits end to end and walks a marker through all three — that is the SUM, the thing you must live through in full. BOTTOM starts three clocks together and stops the instant one rings, abandoning the other two mid-tick — that is the MIN, the fastest of three racers. Guess which is which, then watch the means land: the sum is Gamma(3,λ) with mean 3/λ, the min is Exp(3λ) with mean 1/(3λ) — nine times apart, same λ, same "three." Race 200 draws of the same three waits, wired both ways, and the two histograms of finishing times land in almost entirely separate territory. The min needs one line: "min > t" means all three exceed t, and independence lets the survivals multiply — [e^(−λt)]³ = e^(−3λt) — the payoff for building S(t) before f(t) back in Fig. 3, sitting right beside the greyed density route you'd otherwise have to differentiate your way to. The common slip: saying "open three lanes" (a MIN) while deriving the gamma (a SUM) — keep the words and the algebra pointing at the same picture.

Three gaps end to end is series: you are waiting for the third dot and you must sit through all three. That's the gamma, mean 3/λ, three times slower. Three clocks started together, stopping the moment the first one rings, is parallel: that's a minimum, mean 1/(3λ), three times faster. At two an hour, the series wait averages 1.5 hours and the parallel wait averages 10 minutes. Same three, same λ, and a factor of nine between them. And here's the survival function paying for itself one last time. "The min exceeds t" means all three clocks exceed t, and they're independent, so the survivals multiply: S_min(t) = [e^(−λt)]³ = e^(−3λt), which is Exp(3λ), read straight off section two. The minimum of exponentials is exponential again — still memoryless, just faster. Try that with densities and you'll be integrating for an hour. With survival it is one multiplication, and that is precisely why we opened this chapter on S and not on f.

So let's close the family, and let's do it with code rather than a claim. Simulate one constant-rate stream and nothing else: draw iid Exp(λ) gaps, then cumulative-sum them into dot times. That array is the entire random world of this chapter. Now interrogate it four ways, without generating a single new random number.

gaps ~ Exp(λ) mean=0.00 theory=0.00
λ 2.0
r 3
hazard flat (slope ≈ 0) ⇒ every law below fits at once.
lawmeantheorymaxΔ
constant hazard → all four laws fit
What you're looking at — one array, sliced four ways
blue bars — the SAME 600 exponential gaps, read four different ways, never redrawn
gold curve — each law's density, computed fresh from λ and r, not fit to the bars
red — what a constant hazard was buying you: lose it, and all four fits die together

That's why “which formula do I use?” is the wrong question — there was only ever one array. Gaps, counts per window, sums of r, minima of r: four named laws, read off one stream with zero new random draws (line 2 is the only place chance touches this figure). Toggle make it age and every one of them breaks at the same instant, because none of the four was ever a separate fact — they were one fact, viewed four ways.

Fig. 22. one array, four laws. Line 2 draws 600 exponential gaps — the ONLY randomness in this whole figure — and every button below just re-reads that same array: gaps is it raw, counts bins the cumulative dot-times into unit windows, Σr sums every r gaps in a row, minr takes their minimum. Four overlays, four exact matches, zero new random draws. Drag λ or r and the whole table re-runs live. Then press make it age: the stream is rebuilt from a rising-hazard generator instead, same mean wait, wrong shape — and every maxΔ column turns red on the same click, because the four "laws" were never four facts.

Histogram the gaps and the exponential PDF lands on them. Count dots per unit window and the Poisson PMF lands on that, with mean and variance both λ, exactly the Chapter 11 fingerprint. Sum every three consecutive gaps and Gamma(3, λ) lands, mean 3/λ, which is 1.5 hours at two an hour. Take the min of each triple and Exp(3λ) lands, mean 1/(3λ), ten minutes. One array, four laws, and the code introduced no new probability at all. Every one of those was proven above, so this is a receipt, not an argument. Then break it on purpose. Replace the constant hazard with a rising one and every fit fails at once, because the thing now ages and no member of this family will ever describe it.

Step back and look at what you are carrying out of here. Not four distributions — one stream, read four ways. Point yourself at any constant-rate process, whether it's decay or arrivals or failures or packets, and you can read it from either end knowing the two readings are the same fact. Given a half-life you can recover λ. Given λ you can state the mean, the SD, the median and the fraction still alive from memory, because there's only one length in the box: at two an hour, a mean of 0.5 hours, an SD of 0.5 hours, a median near 21 minutes. You can answer "it has run a thousand hours, now what?" without flinching. Better, you can say exactly when that answer is a lie. Check the hazard, and if it climbs, the thing ages and this family is the wrong family. Here's the map.

n→∞ atom bin nrml Poisson geo stream Exponential Gamma Ch 13 flat = in · rising = out S(t+s) = S(t)·S(s) ⇒ e^(−λt) is the only solution same “three,” 9× apart series 3/λ parallel 1/(3λ) S → f, here (fig·07) F_Y → f_Y, Ch 13 same chain-rule factor → the Jacobian
click a law — trace it back to the stream
or open a side panel
atom→geo→p=λ/n→q^(nt)→e^(−λt)
What you're looking at — every law in this chapter is one stream, read a different way
blue = the binomial family: one atom, the binomial hub, and its two limits (normal, Poisson)
gold = you are here — Poisson, Exponential, Gamma, three readings of one node
green ring = the constant-rate stream — the ONE random object every gold node hangs off
dashed grey = Chapter 13, parked — press → Ch 13 to see the road already paved
Fig. 23. The whole chapter is one map. Blue is the binomial world you already owned: one atom, forced into a binomial hub, which has exactly two limits — normal and Poisson. The Exponential is reached a different way: the atom turned sideways into the geometric, then the slice-width dial run to zero. Click Poisson, Exponential, or Gamma and watch its derivation light up gold, edge by edge, back to the green stream ring — the one random object all three keep hanging off, drawn live because it never stops. The side panel opens three more receipts: memoryless replays S(t+s)=S(t)·S(s) collapsing to e^(−λt) and nothing else; series/parallel shows the same three exponentials landing nine times apart depending only on the picture; → Ch 13 reruns fig‑07's move — differentiate a survival curve — and stamps the arrow that keeps that trick working: the Jacobian.

And the road out is already paved. Twice in this chapter we got a density by differentiating a survival curve, rather than by substituting into a formula. Do that once and it's a trick. Do it twice and it's a method. Chapter 13 generalizes it into the full change-of-variables machine: to get the law of Y = g(X), you route through the CDF and then differentiate, and you never substitute a density into g. The chain-rule factor that pops out of that differentiation has a name, the Jacobian, and you'll recognise it on sight. It's the same factor that has been standing in front of the exponential this whole chapter.

iolinked.com
Written by Ajai Raj