Chapter 13 handed you a machine with one moving part, and it never failed you. Point it at Y = g(X), chase the event, collect the Jacobian, and a new density falls out the other end. Then it hit a wall. We wanted χ²(k) = Z₁² + Z₂² + … and we could not prove it, and I told you plainly why. The machine re-rules one axis, so adding needs two variables standing in the same room, and we had no room. So here is the gap, said out loud. Everything in this book so far has been one variable at a time. Real problems have several variables moving together, and your instinct already has an answer ready for what that means. Two variables, two distributions, so just do Chapter 8 twice. That instinct is wrong. It is not slightly wrong. It is missing an entire object, and a fair die is going to prove that to you in the first two minutes. Here is where we come out. Two variables at once is not two distributions. It is one pile of probability mass spread over a plane instead of a line. And every question you can ask about the pair is one of exactly two physical moves on that pile. You can project the pile onto an axis, which is the law of total probability, drawn. Or you can slice the pile and rescale the cut face to area 1, which is Chapter 5's shrink the world and re-measure, finally drawn. And independence turns out to be exactly the case where every slice has the same shape.
01Two numbers off one roll
We are going to build nothing new here. Chapter 8 told you that a random variable is a function from outcomes to numbers, written X : Ω → ℝ, and Chapter 13 made us read that definition again because it had quietly rotted. So take the machine you already own and run it twice on the same outcome. One experiment happens, one outcome lands, and two readers each report a number about that one outcome. That is not new theory. It is a second reading of a thing that has already happened.
runs: 1 — clicking never adds one
click a die — read it twice
the question nobody has asked· · · try three different outcomes
What you're looking at — one ω, run through TWO readers at once
Ω (left) — the six real outcomes; click one, nothing is re-rolled
X — reads the face showing, lands on the gold rail
Y — reads the hidden opposite face (a die's faces sum to 7), lands on the green rail
the orange dot — both readings at once, plotted as one point (X,Y)
Fig. 1. Click any outcome in Ω and two arrows fire from the same click — X reads the face showing onto the gold rail, Y reads the hidden opposite face onto the green rail, and one orange dot lands at (X,Y). Nothing was reshuffled: it's the identical die, read twice. Toggle one reader and the second arrow disappears — that collapse is exactly Chapter 13's picture. The run counter never leaves 1, no matter how many times you click. Read three different outcomes and a dimmed question lights up.
There is Ω, the same rack of real outcomes it has always been, and there is the same P measuring it. What is new is that two arrows leave each outcome now instead of one, and they land on a pair of numbers, (X, Y). Nothing was re-randomised. The experiment was not re-run. So here is the one question nobody has asked you in thirteen chapters. Do the two outputs remember each other? You have never needed to ask it, because you have never had two readings of one outcome on the table at the same time. Let me show you why the question is not academic, using the smallest example I can build.
Two worlds, both made out of a fair die. In World A you roll two dice: X is the first die, Y is the second. In World B you roll one die: X is the die, and Y = 7 − X. Now look at the two 1D pictures: in both worlds X is uniform on 1..6. In both worlds Y is uniform on 1..6 as well, because 7 − X runs over exactly those six values, each with probability 1/6. The two pictures are not merely similar. They are identical, digit for digit, and here they are printed side by side. So commit to an answer before you scroll: what is P(X + Y = 7) in each world?
A guesses P(X+Y=7)
B guesses P(X+Y=7)
pick A's guess, then B's
What's printed — Cell 1's marginals, both worlds
World A (two dice, X=first, Y=second): X and Y each print 0.1667 ×6 — 36 equally likely outcomes
World B (one die, X=die, Y=7−X): X and Y each print 0.1667 ×6 — only 6 outcomes
the same six numbers, twice over — yet P(X+Y=7) lands at 1/6 in A and 1.000 in B
Fig. 2. Two worlds, built only from Ch 9's fair die. World A: roll two dice, X = first, Y = second. World B: roll ONE die, X = the die, Y = 7−X. Cell 1 counts both sample spaces and the key above prints the result: in both worlds X is uniform on 1–6 and Y is uniform on 1–6 — not similar, identical, 0.1667 six times over. Lock in a guess for P(X+Y=7) in each world, then run 200,000 real trials per world: World A's running fraction settles near 0.1667, while World B's pins to 1.0000 from trial one and never wavers, because Y was built to always finish the sum at 7. The exact counts confirm it — 6 of 36 cells favour A, 6 of 6 favour B. Press why?: the same six diagonal cells matter in both grids, but World A spreads 1/36 onto each of 36 cells while World B piles 1/6 onto only those six — identical shadows, and a joint that the two marginals alone could never have told you.
Your gut multiplied. It said 1/6 in both worlds, because that is the only move a pair of separate 1D tables permits. World A really is 1/6, so half the gut was right. World B is 1.000. Dead certainty. The pair in World B always sums to seven, because we built Y out of X to make it so. Sit with what just happened, because nothing was hidden from you and nothing mystical occurred. The two worlds have identical shadows and they behave in opposite ways. So the two 1D tables simply do not contain the answer, and no amount of staring at them will produce it. Something real lives between the two variables, and it is invisible from either axis alone. That gap is why this chapter exists.
02One pile of mass on a plane
So let's write down the thing that lives between the two variables. I want to build it before I name it, because the name is going to sound like a new axiom and it is not one. Go back to the two dice worlds and just count. Thirty-six equally likely outcomes sit on a 6×6 grid, with die one down the side and die two across the top. Fill in each cell with the probability that the pair lands in that cell. No theory, just counting.
0/36 counted · Σ = 0.000
What you're looking at — the same 36 outcomes, counted into two tables
hover a cell to decode its comma
blue row = the event {X = x} · gold column = {Y = y}
green ring = their intersection — the one cell p(x,y)
new probability theory added: none — Ch 3 has measured intersections since Part I.
Fig. 3. Drag count it through all 36 equally likely (D1,D2) pairs. World A: X=D1, Y=D2 — each outcome lands in its own cell, so p(x,y)=1/36 everywhere. World B: X=D1, Y=7−D1 (the opposite face of the same die) — all six D2's pile onto one cell, so p(x,y)=1/6 only on the anti-diagonal, 0 off it. Hover any cell: the comma unpacks into P({X=x} ∩ {Y=y}) — nothing Chapter 3 didn't already measure.
You have just built two grids with the same edges and completely different insides. But you only counted what was handed to you, so pin those four edge totals down and try to move the mass around yourself. How many different grids can hide behind one pair of shadows?
Predict — with all four edge totals frozen, how many different tables exist?
degrees of freedom (r−1)(c−1) = 1
Commit a guess above to unlock the table.
What's printed — one line of joints, one pair of shadows
Free cell p(0,0): you set it anywhere in 0.00–0.50. This ONE number is your whole choice.
Forced cells: the other three auto-fill (0.50−p and p) — no freedom left, the margins claim them.
Margins: all four sums stay 0.50 for every slider value — the shadows can't tell these joints apart.
Fig. 4. Freeze all four edge totals at 0.50, then commit: how many tables fit? The gut says one — it still believes the margins are the distribution. Drag the single free cell p(0,0) from 0.00 to 0.50 and the other three auto-fill live, yet every row and column sum stays 0.50, unmoved. The dependence readout p(0,0)−pₓ(0)pᵧ(0) swings + (diagonal), through 0 at 0.25 (independent, World A flat), to − (anti-diagonal, World B) — one pair of shadows shared by a whole line of joints. Click any forced cell and it refuses a second freedom: df = 1. Bump to 3×3 and the counter reads 4 — you watch (rows−1)(cols−1) numbers live between the two margins, exactly why the shadows could never determine the joint.
World A puts 1/36 in every cell, flat as a table. World B puts 1/6 down the anti-diagonal and 0 in the other thirty cells, because those thirty pairs simply cannot happen. Both grids still hold all the mass: thirty-six cells at 1/36 total 1, and six cells at 1/6 total 1 as well. Look at the edges of the two grids and you get the same totals, exactly as promised. Look at the interiors and you are looking at completely different worlds. What the shadows lost is now something your eye can just see. Now the name. That grid of numbers is the joint PMF, written p(x,y). And read the comma out loud, because it is doing more work than it looks: p(3,4) means P(X=3 AND Y=4). That comma is Chapter 3's ∩ and nothing else.
Which means the joint added no probability theory at all. Watch how cheap it was. {X=3} is an event, because Chapter 8 built it that way, and {Y=4} is an event for the same reason. Their intersection is an event too, by Chapter 3, and Chapter 3 has known how to measure events since Part I. So p(x,y) is a probability you have been able to compute since the second week of this book. The grid is not a new object. It is only where we choose to write the numbers down. And the cells are disjoint, and together they cover everything, so the mass on the grid sums to exactly 1, which is Chapter 3's total mass, unmoved.
Now let the cells shrink, exactly the way Chapter 8 shrank them to cross from a PMF to a PDF. That was the sum-to-integral fork, and we are about to run the same fork one dimension over. But two things are about to change meaning underneath you, and nobody says them out loud, so they cause real damage. Chapter 8 drilled in that probability is the area under the curve. Promoted to two dimensions, that sentence is now wrong. And the word density is about to stop meaning per unit length and start meaning per unit area. That switch usually happens silently. So let me say it out loud instead, and let the units do the proof.
N=3 · Δ=0.667 m · 18 cells
What you're looking at — the same fork, one dimension up
Ch 8's sentence, promoted, is now wrong. Area became volume. Density is now per unit AREA.
f(x)·dx = 0.270 mass/m · 0.667 m = 0.180 massAREAf(x,y)·dx·dy = 0.135 mass/m² · 0.667 m · 0.667 m = 0.060 massVOLUME
Σ over all cells (both panels): 1D = 1.000 · 2D = 1.000 — never leaves 1
left = Ch 8's 1D curve; the gold box is ONE sliver, width dx
right = the same idea one dimension up; the gold tile is that SAME cell, now dx by dy
Fig. 5. One shared cell size slider, two panels locked together. Left is Ch 8's sliver: mass = f(x)·dx = height×width = AREA. Shrink the same cell to dx by dy on the right and its mass is f(x,y)·dx·dy = height×tile = VOLUME — that's why "probability is the area under the curve" had to die: the density is now priced per unit AREA, and a region's probability is a volume. Both running totals still read exactly 1.000, at every grid size — only the one highlighted cell's own slice shrinks toward zero.
In one dimension the mass in a thin sliver was f(x)·dx, which is height × width, which is an area. Now take a grid cell here and shrink it to dx by dy. The mass in that cell is f(x,y)·dx·dy, which is height × tile, which is a volume. That is the whole promotion, and it is arithmetic rather than convention. So f(x,y) is the joint PDF, and it is priced per unit area. The probability of a region is the volume of mass sitting over that region. The total is ∫∫f dx dy = 1. The grid just became a landscape. Now let's make sure Chapter 8's oldest sin didn't sneak back in wearing a new costume, because the picture is new and your gut is going to try it again.
mass0.000
density0.00
drag the box · shrink the gold corner
What you're looking at — the colour is a rate, not a probability
the gold box is the region — drag its body to move it, its corner dot to resize it
dark→blue→gold is how thickly the mass is piled there — never call it "the probability at a point"
green fires when mass hits 0.000 while density holds — the picture of P(X=x,Y=y)=0
toggle "probability" to see the tempting label turn red, then watch collapse strike it out
Fig. 6. Drag the gold box across the landscape and both readouts move together, honestly, as mass accumulates. Then grab its corner and shrink it toward a point: mass counts down to 0.000 while density — the raw height under the box — never moves, because f(x,y) prices mass per unit AREA and needs room before it means anything. Toggle the legend to "probability" and watch it turn red; hit collapse and watch the lie get struck through: P(X=x,Y=y)=0, always.
Drag the region around and the enclosed mass ticks up honestly, accumulating as volume. Now shrink that region toward a single point and watch two readouts at once. The mass runs to 0. The colour under your finger never changes. So P(X=x, Y=y) = 0, exactly as in Chapter 8, and the height was never a probability. The height is a rate: it tells you how thickly the mass is piled at that spot, and it only becomes a probability once you multiply it by some room to sit in.
03Project it — the marginal
Now we have the pile, so let's make the first of the two moves you can make on it. Go back to World B's grid and ask for P(X=3), where you simply do not care what Y did. Everyone's instinct is to add up the row, all six of its cells, and everyone's instinct is right. But nobody ever tells you why that is legal. The licence is a theorem you already own and already paid for.
ask: what is P(X=3)?
What you're looking at — one row, flying into its margin
gold = the row {X=3} — six cells {Y=1}…{Y=6}
green = the landed margin cell, p_X(3) = Σ_y p(3,y)
aha — the six {Y=y} cells in a row are disjoint and cover it, so they're a partition, and summing across them is Chapter 6's law of total probability, run once per row.
Fig. 7. World B's grid, and one question: P(X=3). Step 1 lights the row. Step 2 runs the Chapter 6 check on the actual cells — the six {Y=y} cells are disjoint (they flash one at a time, no overlap), they cover the row (the gold sweep fills it edge to edge), so {Y=1},…,{Y=6} is a partition. Step 3 writes the law of total probability twice, symbol under symbol: P(X=3)=Σ_y P(X=3,Y=y) is exactly p_X(3)=Σ_y p(3,y) — same line, different alphabet. Step 4 detaches the row and flies it into one cell in the margin, where it adds to 1/6. Step 5 says the word: that green cell sits in the physical margin of the table — that's the entire etymology, nothing more. Replay on World A: six different cells, 1/36 each, land on the same 1/6 — identical margins, different interiors.
Here is the argument, and it is Chapter 6 word for word. Y must take some value. The six cells {Y=y} in that row are disjoint, and together they cover the row entirely. So they are a partition, and the law of total probability says P(X=3) = Σ_y P(X=3, Y=y). That is not a definition to memorise. It is Chapter 6 with the partition being "what did Y do?", and that is the whole justification. Watch the row's six cells detach and fly into a single cell out at the side of the table. Now the name, and the name is embarrassing once you see where it comes from. Row totals get written in the margin of the table, so we call p_X(x) = Σ_y p(x,y) the marginal PMF. The word is a place on the page. It never meant unimportant, which is roughly the opposite of what this operation does.
And those marginals are the shadows. They are the two identical 1D pictures the two dice worlds handed us. So the two 1D pictures that failed us back in the dice worlds finally have a proper name, and that name is the marginal.
Run the same move through Chapter 8's fork. The row becomes a column of the landscape, and the sum becomes an integral: f_X(x) = ∫ f(x,y) dy. People call this projection, and the word plants a picture in your head that is wrong. You imagine a silhouette, as if you walked around the side of the surface and traced its outline against the wall. It is not an outline at all. It is an accumulation: you pour the entire column of mass straight down and let it stack up into a bar. So predict before you look. Which projects higher, the tall thin spike or the long low ridge?
Predict: which pours MORE at its own x?
pick spike or ridge, then pour ↓
What you're looking at — a map, not a skyline
gold block — the spike: height h=3, but only L=1 deep in y
blue block — the ridge: height h=1, but stretched L deep in y (drag the slider)
green bars — the TRUE marginal: pour the column down, h×L at each x
red dashes — the FALSE silhouette: just h, the skyline you'd have guessed
→ the red curve never moves when you drag L — a silhouette can't see depth. Only the pour can: at x₂ the column is short but wide-in-y, so it out-pours the tall skinny spike the moment L passes 3.
Fig. 8.the pour, not the outline. The spike stands h=3 but is only L=1 deep in y; the ridge stands h=1 but you drag its depth L from 1 to 9. Predict, then hit pour: the beam sweeps to each column and its mass falls straight down into a bar — that bar's height is h×L, the real marginal, not the block's height. At the default L=8 the ridge pours 8.0 against the spike's fixed 3.0 — more than double, from the shorter block. Toggle silhouette and a red skyline appears tracing only h: spike 3, ridge 1 — exactly backwards, and it never moves when you drag L, because a silhouette can't see depth. Slide L down to 3.0 and watch the green bars tie exactly — below it the spike wins for real, above it the ridge does; the red line never notices either way.
The silhouette model says the spike, because the spike is taller. The accumulation says the ridge, because there is more of the ridge to pour. Being wrong about that once is worth more than any sentence I could write, and the silhouette should now be dead. But a second question is sitting here, and every textbook walks straight past it. Why is ∫f(x,y)dy a density and not a probability? You have spent five chapters watching ∫f dxbe a probability. Integration now feels like the operation that produces probabilities, so watching it hand back a density instead looks like an inconsistency, and most readers quietly file that under "don't ask".
mass⁄mx·my
drag to integrate out y
What you're looking at — before you drag: guess, is what survives a probability or a density?
blue landscape = the joint f(x,y) = 0.25 mass/(m·m), constant over the 2m×2m square — click it to probe what it is
gold curve = the marginal fX(x) = 0.500 mass/m, once y is fully swept out — click it to probe what it is
if ∫f dy were already a probability, fX would be unitless and ∫fX dx would carry units of length — but it lands on 1.000, a pure number: contradiction.
Fig. 9. The joint f(x,y) is priced at mass/(m·m) — a constant 0.25 over the 2 m×2 m square, so total mass is 0.25×2×2 = 1.000. Sweep the slider and watch the "per unit y" get spent: the landscape flattens into the marginal fX(x) = 0.25×2 = 0.500 mass/m — still a density, never a probability, because it still needs one more ×dx to become a pure number. Audit it: ∫fX dx = 0.500×2 = 1.000, exactly.
So let's ask, because the answer is arithmetic on the units: f(x,y) is mass per (unit x)(unit y). Integrate over all of y and you spend the "per unit y". What is left over is mass per unit x, and that is a 1D density. That is not a convention someone chose, just the units doing their own bookkeeping. Watch the label on the vertical axis change as the integral runs, then check the derived marginal the way you would check any density: ∫f_X dx = 1, still. Nothing leaked.
04Slice it — the conditional
That was the first move. Here is the second, and it is the one you will use for the rest of your life. Instead of pouring the columns down, hold x still. Fix x = x₀ and read f(x₀, y) as a function of y alone. What you get is a cut face through the landscape, and it looks like the answer to "how does Y behave when X is x₀". Its shape is exactly right. Its size is wrong, and that gap is the whole trap. The shape looks so right that fixing the size is going to feel like fussy bookkeeping.
area = 0.50 — not 1
What you're looking at — the slice IS the marginal, seen face-on
the grid is the joint density f(x,y) = (x+y)/8 — gold cells are denser
gold readout + gold marker are the SAME number, f_X(x₀) — drag and watch them stay tied
Fig. 10. Drag the line and the cut face pops out beside it, area live: it wanders (0.25 to 0.75 here) and never hits 1 — the hunt button scans the whole range and can't find one either. The gold readout in that panel and the gold marker riding the marginal curve below are the same number, moving together, because they're the same definition: f_X(x₀) ≡ ∫ f(x₀,y) dy ≡ the slice's area. Drag rotate and the shape swings from a flat bar (end-on — just the total, no detail) to the true slanted slice (face-on — the density varying with y) without its area ever changing, because it's one object, not two facts that happen to agree. That's why the conditional's denominator, f_X(x₀), isn't a rabbit from a hat — it's the column you already built, seen from the side.
Drag the vertical line and watch the area readout on the cut face. It is never 1, at any x₀, ever. So the cut face is not a density, whatever it looks like. Now here is the part that defuses everything coming next. That area is exactly the height of the marginal bar at x₀, the bar we built a moment ago by pouring columns down. Of course it is. We defined the marginal at x₀ as the total mass in that column, and this slice is that column, seen face-on instead of end-on. So the slice's area is f_X(x₀). That is not a new fact, just the same object looked at from a different angle. Remember it, because a denominator is about to appear and it is going to look like a rabbit out of a hat.
Except we cannot get to it yet, because a contradiction is sitting in the road and I am not going to drive around it. Chapter 5 defined P(A|B) = P(A∩B)/P(B). Chapter 8 proved that P(X=x₀) = 0 for a continuous variable. Put those two sentences together and "condition on X = x₀" reads 0/0. That flatly contradicts two things this book taught you explicitly. Nearly every course writes f(y|x) = f(x,y)/f_X(x) anyway and hopes you don't notice. If you felt that, you were right, and here is the honest fix, with nothing swept under the rug.
Ch 5 and Ch 8, combined honestly. The answer is 0/0.
0/0 — and you were right to balk
What you're looking at — the joint density f(x,y)=x+y on the unit square, and the ratio that survives a limit its own top and bottom do not
blue — the landscape f(x,y)=x+y (darker = denser; it integrates to exactly 1 over the square). Its shadow at x₀=0.4 is the marginal fₓ(0.4)=∫₀¹(0.4+y)dy=0.90 — the strip mass per unit dx
red — the zero-width line X=0.4, and the two masses that collapse: P(strip)=0.90dx+dx²/2 and P(tile)=0.05·dx·(1+dx/2), both exact integrals of x+y, both → 0. Their bars are a log gauge whose floor is the slider's smallest dx
gold — the strip 0.4 ≤ X < 0.4+dx and the tile inside it (y within 0.025 of 0.6). At the dx values here the strip is thinner than a pixel, so it is drawn at a minimum visible width; the dx tokens are the ones that annihilate
green — the survivor: f(y|strip) = (0.4+y+dx/2)/(0.90+dx/2). The dx already cancelled; what is left reads 1.111 at y=0.6 and encloses area 1.000 for every dx on the slider
Fig. 11.the strip — watching the dx cancel. Ch 5 defines P(A|B)=P(A∩B)/P(B); Ch 8 proved P(X=x₀)=0. Put them together and conditioning on X=x₀ is 0/0 — so the sharp reader who stalled here was not slow, the ground really was broken. The repair: condition on the strip 0.4 ≤ X < 0.4+dx, a genuine event with P ≈ fₓ(0.4)·dx = 0.90dx > 0, and Ch 5 applies with nothing bent. Assemble the ratio and one dx sits on top (tile mass f(x₀,y₀)·dy·dx) and the same dx sits below (strip mass fₓ(x₀)·dx) — they cancel, leaving 1.00·0.05/0.90 = 0.0556, and per unit y, f(0.6|0.4) = 1.111. Now squeeze: over the slider's whole range both masses fall by a factor of about eight thousand, to 9.000e−7 and 5.000e−8, while f(0.6|x₀) reads 1.111 and the area reads 1.000 at every setting. That is the aha: conditioning on a probability-zero point is legal because it is the limit of honest conditionings on strips of positive probability — and the ratio survives a limit that its own numerator and denominator do not.
Don't condition on a point. Condition on a thin strip, x₀ ≤ X < x₀+dx, which is a real event with a real, positive probability of about f_X(x₀)·dx. Now Chapter 5's definition applies with nothing bent: the top of the fraction is the little tile's mass, f(x₀,y)·dy·dx. The bottom is the strip's mass, f_X(x₀)·dx. The dx sits in both, and it cancels on screen. Then let dx → 0 and watch the answer. It never moves. The ratio survives the limit even though the top and the bottom both run to zero, and that is why conditioning at a probability-zero point is legal. We did not dodge the division by zero. We showed that the ratio was never the problem in the first place.
So we have earned it: f(y|x₀) = f(x₀, y) / f_X(x₀). Now look at what that division physically does, because as algebra it is forgettable and as a picture it is permanent. The slice has area f_X(x₀). Divide the whole slice by its own area and the new area is exactly 1, by the plain definition of dividing a thing by itself. That is Chapter 13's renormalise discipline, verbatim. A density must be rescaled so the total mass stays 1.
divide by: __ × its own area
raw slice — area 0.510, never 1
What you're looking at — one slice, forced to sum to 1
dashed outline — the target shape, once the slice is divided by its own area
green (or red) wedge — the live slice f(x₀,y): green only when its area is 1
gold line — x = x₀, the cut through the joint density that the slice comes from
The raw slice encloses area 0.510 — that number is f_X(x₀), the marginal density at x₀, because f_X(x₀) = ∫f(x₀,y)dy is exactly the area under this curve. Divide the slice by that exact number and the area is forced to 1 — not by luck, but because dividing anything by its own size always gives 1. Divide by half or double that number instead and the area lands on 2.000 or 0.500, never 1, which is the proof that 1.000 was never a magic constant. And the line itself is not new: it is Chapter 5's clue B, shrunk from a rectangle down to a single vertical line — condition on it, and the rescale you just watched is the re-measure.
Fig. 12. The raw slice f(x₀,y) sits in its panel with area 0.510 — never 1. Press ÷ by its own area and the wedge inflates in place until it exactly fills the dashed target and the readout snaps to 1.000, while peak y=2.00 and width ratio 1.00× never move — proof that only the height changed. Try 0.5× or 2× the true area instead and the readout lands on 2.000 or 0.500, never 1. Press shrink the world and the plane greys to a single gold line — that line is Chapter 5's B, shrunk from a rectangle to one x, and dividing by its own area is the re-measure.
That was a button on a smooth landscape. Now make the same move on numbers you can add up by hand. Here is a small joint table, read straight from a story. Grab one row, divide it by its own total, and you have Y's law for the case where X is fixed.
predict: same bar tallest after?
What you're looking at — dividing a row by its own total, live
blue bars — the raw joint p(x,y), read straight off the story table
green cells — p_X(x), each row's own margin, already summed down the side
gold — the row you picked, rescaled live by whatever you chose as the divisor
Pick a row and divide by its own total and every entry lands where Fig. 12 said it would: shape untouched, size forced to 1. Try the wrong divisors — 1, or a neighbouring row's own total — and the sum drifts off to 0.300 or 0.750, proving the denominator isn't a free choice: it has to be that row's own f_X(x₀), because dividing anything by its own size is the one division guaranteed to land on 1.
Fig. 13. Click Sun and its row divides by its own total, 0.30: 0.12→0.400, 0.10→0.333, 0.08→0.267 — the bars inflate in place, Σ snaps to 1.000, and Walk is still the tallest bar, ratios untouched. Switch the divisor to 1 and Σ lands on 0.300; switch it to the neighbor's total (0.40) and Σ lands on 0.750 — never 1, because only Sun's own total is Sun's own total. Click Rain instead: same move, a different shape falls out — 0.03,0.09,0.18 rescale to 0.100, 0.300, 0.600, and now Car is tallest. That's dependence, felt as one division you could do by hand.
The cut face inflates in place. The shape is untouched, the height changes, and the area readout snaps to 1.000. So the conditional PDF takes its shape from the joint and its size from the renormaliser, and that is the entire definition. And now the sentence this chapter has been walking toward since the very first roll. This is Chapter 5's shrink the world and re-measure, finally drawn. The world shrank to a single vertical line, and the re-measuring is the rescale you just watched. Conditioning was never a formula. It was always this picture.
05When the shadows are enough
Time to put a real landscape on the screen and turn a knob on it. Take two standard bells from Chapter 10, stand them up in two dimensions, and you get the 2D Gaussian. It comes with a dial called ρ. At ρ = 0 its contours are perfect circles. Turn ρ up and the blob tilts into an ellipse. Now let me be straight with you about that symbol, because I am not going to hand you a token you cannot cash. ρ is a knob, and it tilts the blob. Its real name and its real meaning arrive in Chapter 16, and until then I am not going to pretend you own it.
ρ≈0 → the peak never moves (flat)
What you're looking at — a knob that tilts the blob, and a peak that slides because of it
the ellipse is the joint's shape — ρ tilts it from a circle
gold dashed line = x₀ — drag its slider to slice there
green line = the conditional peak, growing as y = ρ·x₀
grey dashed ghost = the x₀=0 slice, frozen for comparison
ρ ("rho") is just a knob here — its real name and formula are Chapter 16's job. We're not pretending you own it yet.
Fig. 14. Drag x₀: the peak lands at ρ·x₀ — flat at ρ=0, sliding as ρ climbs toward 0.95. ρ's real name arrives in Ch 16.
And I am not going to tell you that tilted means dependent either, because that would make dependence a visual vibe. Let's earn it with the slicer instead. Set ρ = 0.8 and drag the slice. The cut face's peak slides as x moves, which means learning X moved Y's whole picture. That is Chapter 5's definition of dependence, in motion, in your hand. Now set ρ = 0 and drag the slice again. The peak sits dead still. So dependence is not a number yet. It is a thing you just watched move.
Which brings us to the question the whole chapter has been for. When are the two shadows enough? There is an obvious answer available, and I want you to commit to it first. On the screen is the tilted ρ = 0.8 blob. Beside it are its own two marginals, f_X and f_Y, computed live from that exact blob, in front of you, by pouring its columns down. Now multiply the shadows. Build the surface f_X(x)·f_Y(y). What comes back?
pick one — then the button unlocks
What you're looking at — one blob, its two shadows, and the world you can rebuild from those shadows alone.
p(x,y) — the blob: 24×24 cells summing to 1
p(x) — the shadow on x, each column summed
p(y) — the shadow on y, each row summed
mismatch: orange = product too low, blue = too high
↳ The knife: every row and every column of the mismatch map sums to exactly 0 — dependence lives entirely in the part the shadows cannot see. That is why factorisation is a theorem, not a tautology. (Drag across either panel to run the slicer. Brightness is on a √ scale so the tails show.)
Fig. 15. Same shadows, different worlds — step 1's two dice all over again, and only one of them factors. The two shadows never move as you turn ρ, so they cannot possibly be holding the tilt. Drag ρ to 0 and Σ|Δ| hits 0.000: the single setting where the product is the joint.
Your gut said the same tilted blob, and it said so with total confidence. Obviously the marginals came out of that blob, so the blob must be what is in them. Press the button. A perfectly circular blob appears, the tilt is gone, and you have just watched your own reconstruction lose something real. Two different worlds, identical shadows: this is the two dice worlds again, in a smarter costume. Now turn on the mismatch map, which is just joint − product, and it glows exactly where the product lied. Positive along the tilt. Negative across it. That glow is the dependence that projection destroyed, and it was never in the shadows to begin with.
Now take the ρ dial and kill the mismatch with your own hand. Turn ρ toward 0 and the mismatch drains to black, hitting exactly zero at ρ = 0. That is the one setting where the product IS the joint. So here is the keystone, and the last two minutes were spent earning it: it is a real theorem, not a tautology. X and Y are independent exactly when learning X never changes Y's picture, meaning f(y|x) = f_Y(y) for every value of x. And that is algebraically identical to f(x,y) = f_X(x)·f_Y(y). Both sentences describe the same surface.
Let's turn that crank, because it takes one line and it uses nothing you don't already have. Say every slice has the same shape g(y). Then the surface is f(x,y) = f_X(x)·g(y), because a fixed shape scaled by a number that depends only on x is exactly what that sentence draws. Now integrate over x and let the marginal machine run: f_Y(y) = g(y)·∫f_X(x)dx = g(y). A marginal integrated over its own variable is 1, so that integral simply vanishes and the common slice shape must be the marginal. We did not assume that anywhere. It was forced on us. Multiply back up by f_X(x) and you have f(x,y) = f_X(x)f_Y(y).
Two things fall out of that for free. First, the equation is symmetric in x and y, and a product does not care which letter you solve for. So if X tells you nothing about Y, then Y tells you nothing about X, and independence is mutual with no second argument required. Second, it is a claim about every pair of values at once, not about any single cell. Checking one cell, seeing a product, and declaring independence is the classic way to be wrong. Put it on the dice grids. In World A the cell at (3,4) holds 1/36, and its row total 1/6 times its column total 1/6 is 1/36 as well, so that cell factors, and so does every other cell. In World B that same cell holds 1/6, while the same product is still 1/36, six times smaller. Section one's belief was that the marginals are the story, and to that belief f = f_X·f_Y reads as a triviality: of course the joint is the product of its marginals, where else would they have come from? But this is exactly what factorization is not. You have now watched the product fail on screen. Factorization is the exact condition under which the two flat shadows contain the whole solid, and most joints do not meet it.
One more thing before we leave factorization, because this structure is not a probability trick and I don't want you filing it away in the probability drawer.
toggle the reading
the bet: independence — else it lies
What you're looking at — one ansatz, filed under two names
Both columns posit a shape before checking it: something fixed (blue) times a number depending only on the other variable (gold). Every row above is the same curve, just rescaled by the gold number on its right.
the joint f(x,y) — X and Y's shared density
posited: f(x,y) = fX(x)·g(y)
BETfails one cell → wrong world
the rod u(x,t) — temperature at x, time t
posited: u(x,t) = X(x)·T(t), e.g. T(t)=e-0.6t
BETno separation → no solution
"a fixed shape in y, scaled by a number depending only on x."
blue = the fixed shape — the same curve every row
gold = the scaling number — the only thing that changes
That's why independence IS separation of variables — both sentences say the same thing: a fixed profile in one variable, scaled by a number depending only on the other. Toggle above: the two stacks don't just resemble each other, they're the same claim wearing a different variable's name.
Fig. 16. Two fields, one sentence. Left: the independence ansatz f(x,y) = fX(x)·g(y) — a fixed shape in y (blue), rescaled by a number depending only on x (gold). Right: the heat equation's separation ansatz u(x,t) = X(x)·T(t) — the identical structure, a fixed shape in x rescaled by a number depending only on t. Toggle the reading and the sentence below swaps its two variable names while the matching stack lights up — same claim, same bet, different field. Both are guesses about shape: independence can fail on a single cell, and separation of variables can fail to find a solution at all.
Physicists make this exact move, and they have a name for it. When one of them solves the heat equation, the first step is to posit that u(x,t) = X(x)·T(t), and the name of that step is separation of variables. It only works when the solution genuinely has that shape: a fixed profile in x, scaled by a number that depends on t. Put that next to what we just drew: a fixed shape in y, scaled by a number that depends on x. Those are the same sentence, written by two different fields about two different things. And in both of them it is a bet rather than a trick. The thing either factors or it doesn't. When it doesn't, the product quietly hands you the wrong world.
The mismatch map lit up on its own. But independence is a promise about every cell, and the glow never told you which cell broke it. So test the cells yourself: clear them all, or catch the single one that fails.
click a cell to test it
What you're doing — running the independence test yourself, one cell at a time. Independence means p(x,y)=p_X(x)·p_Y(y) at EVERY cell.
top number = the joint p(x,y); below it = the margin product p_X·p_Y
✓ green = that cell factors (they match)
✗ red = that cell breaks (they differ) — one is enough
blue edges = the two margins — they never move
↳ The knife: the button unlocks only when all nine cells are green. In the trap table five cells factor perfectly — and it is still dependent, because a break lives in the corner. Independence is not a majority vote. (Load the independent table, clear all nine, then nudge one cell: exactly the four corner cells flip — the margins force breaks into a ± block, so you can never break just one.)
Fig. 17. Run the test with your own hand: click each cell to compare its joint p(x,y) against the margin product p_X·p_Y. The trap table looks tame — five of nine cells factor perfectly (cell x₀y₀ reads 0.26 against =0.20 only when you reach the corner) — yet the DECLARE INDEPENDENT button never unlocks, because a green majority is not independence. Load the independent table, clear all nine, then drag the nudge to Δ=0.02: the two margins do not flinch, and exactly four corner cells flip red at once — the smallest break the margins will permit. One failing cell out of a hundred still kills it.
06The trap in the shape of the region
Now let me show you the way this goes wrong in practice, because you are about to walk into it and almost everybody does. Throw a dart uniformly at a round dartboard of radius one. The density is f(x,y) = 1/π on the disc, which is about 0.32 of mass per unit area at every single point of it, so the density is a constant. It is the most symmetric object we have drawn all chapter. A constant obviously factors, since you can write it as (a constant) × (1), so X and Y are independent. That reasoning is flatly wrong, and the test that convicts it is the one you built ten minutes ago.
guess: independent, or dependent?
What you're looking at — a constant that still won't split
blue = the support, the region where f(x,y) is nonzero — here, the disc
gold = the slice at one x, and the bar showing how dense y is there
green = the rectangle (same area) where the split actually works
the aha: the support is part of the density — a non-rectangular one couples x and y no matter how constant the formula looks
Fig. 18. A constant density on the unit disc looks independent — flat, symmetric, nothing to factor wrong. Drag x and the slice tells a different story: the plateau at x = 0 (range −1 to 1, height 0.500) shrinks to a sliver at x = 0.99 (range −0.141 to 0.141, height 3.544). The shape changes, so X and Y are dependent — and the factor-attempt panel shows why: the indicator 1[x²+y²≤1] can't split into an x-part times a y-part, because the disc's boundary ties the two together. Swap the support for a matched-area rectangle and every slice becomes the same width — the indicator drops cleanly into 1[a≤x≤b]·1[c≤y≤d].
Run the slicer. At x = 0 the slice is a wide flat plateau, spanning y from −1 to 1. At x = 0.99 the slice is a narrow sliver pinned near zero, spanning about −0.14 to 0.14. That half-width is the circle doing the work: the square root of 1 − 0.99² is 0.141. The shape changed. So by the keystone they are dependent, and you can feel exactly why with no algebra at all. Learning X = 0.99 tells you a great deal about where Y can possibly be. Now let the algebra confess what it was hiding. What you actually wrote down was f(x,y) = (1/π)·1[x²+y²≤1], and that indicator is part of the density. It cannot be split into a function of x times a function of y.
The region where the mass is allowed to live has a name, and it is the support. Here is the rule to carry out. Factorization has to hold at every point of the plane, indicator and all. The range of y is allowed to depend on x, and that range is part of the density. So if the support is not a rectangle, the pair is not independent, no matter how clean the formula looks. The support was doing the coupling the whole time. It was invisible only because nobody had said the word.
07Independence lives in the coordinates
Keep the dart exactly where it is. We are going to describe the same dart a second way, and something strange is going to happen. But first I owe you a rung, and without that rung the ending is a rabbit. Chapter 13 taught you that changing coordinates makes a density pick up a Jacobian, and in one dimension that factor was |dx/dy|, a length ratio. In two dimensions the conservation is the same and one thing honestly changes: the factor becomes an area ratio. A gut error sits right underneath, and it is Chapter 13's original sin in a new costume. It says that changing coordinates just relabels the same points, so the density should carry over unchanged. Let's just draw the grid and look.
dr = 0.60 fixed
arc r·dθ = 0.628
AREA = 0.3770
slope dr·dθ = 0.314
mass = 0.0800 locked
density = 0.2122
area = r · dr · dθ · climbs with r
What you're looking at — a polar tile is dr across and r·dθ around, so its area IS r·dr·dθ
the polar grid: rings every dr = 0.6 (a step outward), rays every dθ = 30° = π/6 rad (a step around). Same dr, same dθ everywhere — and the cells out at r=3 are visibly fatter than the slivers at r=0.3.
blue — the tile's radial span. It is dr wherever you drag it: constant.
green — the tile's arc span, r·dθ. An angle only becomes a length once you multiply it by the radius, so this one climbs with r.
gold — the tile. area = (radial span)×(arc span) = r·dr·dθ, plotted right: a dead-straight line through 0, slope dr·dθ = 0.314. Pin the two tiles and the area ratio is 10.0× for a radius ratio of 10.0× — the factor is r itself, not merely something that grows.
coral — one lump of probability mass, 16 specks worth 0.0800, welded into the tile. Dragging cannot change it: renaming a tile's corners touched no outcomes. So the specks just spread.
AHA — the ledger writes itself. mass = density × area. Area is multiplied by r, mass is pinned, therefore density is divided by r. Run it the other way — from (x,y) to (r,θ) — and the density must be multiplied by r: that is the r in f(r,θ) = f(x,y)·r. Ch 13's |dx/dy| rescale, verbatim, with one word swapped: LENGTH ratio → AREA ratio.
Fig. 19.the polar tile — a Jacobian you can see. Same dr, same dθ everywhere, yet the cells out at r=3 are fat and the ones at r=0.3 are slivers: a tile is dr across but r·dθ around, so its area is r·dr·dθ — drag it out and the area trace is a dead-straight line through the origin of slope dr·dθ = 0.314, with area ratio 10.0× for radius ratio 10.0×. Now the half that does the real work: the lump of mass welded into the tile cannot change while you drag, because renaming a tile's corners touched no outcomes. Mass pinned, area ×r — so density ÷r, and that is exactly where the r in f(r,θ) = f(x,y)·r comes from. It is Ch 13's |dx/dy| rescale one dimension up, with one word swapped: LENGTH ratio → AREA ratio.
Overlay a polar grid on the plane, and the cells out near r = 3 are visibly fatter than the cells at r = 0.3. Put a number on it: at ten times the radius, a tile has ten times the area. A tile spans dr in the radial direction and an arc of length r·dθ across, so its area is r·dr·dθ, not dr·dθ, and dragging a tile outward makes its area climb linearly in r. Now finish it with the conservation argument you already believe. The mass in a tile cannot change when you rename its corners, because renaming touched no outcomes. So if the tile's area changed, the mass per unit area must change to compensate. That is Chapter 13's rescale, verbatim, one dimension up, and the exchange rate is the factor r.
Now back to the dart, and this time commit to an answer first. You have watched the disc's slices change shape, so (X,Y) are dependent, and you felt it. Same dart, same disc. But now describe that dart by how far out and which way, as R and Θ. Does knowing the angle tell you anything at all about the radius?
x = 0.50
y = 0.35
r = 0.61
θ = 35°
f(x,y) = 1/π = 0.318
f(r,θ) = r/π = 0.194
drag the dart, then guess: θ→r?
What you're looking at — the same dart, read through two different grids
coral — the one dart. Same physical point, drawn twice: through a square grid on the left, a polar one on the right. Drag it in either panel, or spin θ, and both update — nothing about the dart ever changes, only the pair of questions we're asking about it.
blue — the (X,Y) slice. Fix x and ask what y is allowed to be: near x=0 that's almost the full height of the disc; drag toward the rim and the range visibly shrinks to a point. The shape of that slice depends on x. That's dependence.
green — the (R,Θ) slice of the same dart. Fix θ and ask what the density does as r runs 0→1: it's always the identical ramp, faint near the centre, brightest at the rim — spin θ and the ramp's shape never changes. That's independence.
gold — joint minus (marginal×marginal). It glows across the (X,Y) grid, because the product lies almost everywhere there, and reads a flat 0.000 across the (R,Θ) grid. f(r,θ) = (1/π)·r = (2r)·(1/(2π)) — a function of r times a function of θ, and each factor audits to 1.000 on its own.
AHA — independence was never a property of the dart. It's a property of the pair of questions you ask about it — (X,Y) or (R,Θ) — and the r that makes the second pair factor isn't a fudge. It's the Jacobian you already watched, with your own hand, dragging the tile in Fig. 19.
Fig. 20.the same dart, untangled. One point, two honest descriptions. Read through (X,Y), a fixed-x slice visibly narrows toward the rim — dependent. Read through (R,Θ), the same dart, a fixed-θ slice is the identical 2r ramp every time — independent, exactly what your gut guessed when asked whether θ tells you anything about r. The mismatch map says it without a formula: gold glow across (X,Y), an honest 0.000 across (R,Θ). The reconciliation is one line — f(r,θ) = (1/π)·r = (2r)·(1/(2π)) — and that r is not a fudge, it's Fig. 19's Jacobian, each factor auditing to 1.000 on its own. The tangle was never in the dart. It was in the square grid we insisted on measuring it with.
Staring at a symmetric disc, your gut says no, and this time your gut is right. Here is the arithmetic, in one line, using the rung we just built. f(r,θ) = (1/π)·r = (2r) · (1/(2π)). That is a function of r, times a function of θ. It factors. Check it at a point: at r = 0.5 the left side is 0.5/π ≈ 0.159, and 2 × 0.5 times 1/(2π) is 0.159 too. And the r that made it factor is not a fudge someone slipped in. It is the Jacobian, the area ratio you just watched with your own eyes. So the same dart is tangled in (X,Y) and free in (R,Θ).
That should feel like a cheat, and I want to say plainly why it isn't. Chapter 5 hammered that independence is a physical claim, not a default, and the dartboard slicer had you feel this exact dependence in your hands. So how can it depend on the coordinate system? The resolution is that "the two things" changed: X and Y are not R and Θ. Independence was never a property of the dart at all. It is a property of the pair of questions you decided to ask about the dart. The tangle in (X,Y) was never in the dart. It was in the square grid we insisted on measuring the dart with.
Let's finish with a receipt instead of a formula. One array of darts, sampled uniformly on the disc once, and that array is the entire random world of this section. Every number from here on is the same darts wearing different labels, and we never call the generator again.
Cell 1/5 — draw the darts, once
sampled once — never again below
What you're looking at — one 200,000-dart sample, five relabellings, zero re-rolls
gold heatmap (cell 1) — the actual sampled density: bright near the centre, a hard zero outside the circle. That's the proof the disc really is uniform, before anything else runs.
red cells — f_X·f_Y over-predicts the true joint there (the product thinks there's mass the disc doesn't have — e.g. the rounded corners of the square, outside the circle).
blue cells — f_X·f_Y under-predicts (near the axis edges, where the disc still has mass the flat product misses). Red and blue alternating IS the four-lobe glow — cell 3's grid, same maths, same darts, goes flat instead.
cell 4 checks the two derived marginals against closed-form theory (2r ramp, flat 1/2π); cell 5 slices the raw data by hand at two x's and two θ's and prints the empirical half-widths — a conditional probability estimated from a thin band, not a zero-width conditional density.
AHA — nothing about the darts changed between cells 2 and 3, only their names (x,y → r,θ). So whichever coordinate glows carries the dependence, full stop: in (x,y) the width of the y-slice shrinks as |x| grows toward the disc's edge (that shrink IS the dependence); in (r,θ) every ray from the centre crosses the same full disc, so which θ you're on tells you nothing about r — independence, read straight off the sampler.
Fig. 21.the receipt — one array of darts, relabelled. Cell 1 throws 200,000 darts at the unit disc by rejection and stops: the gold heatmap is that run's own sampled density, bright at the centre and a hard zero past the circle — the header's call counter reads 1 and stays 1 through every button below. Cell 2 histograms those exact darts in (x,y), builds f_X and f_Y by summing bins, multiplies them, and subtracts the true joint: the mismatch map glows in alternating red/blue lobes and the printed total sits well clear of zero. Cell 3 relabels the same arrays to r=hypot(x,y), θ=atan2(y,x) — same histogram-and-subtract recipe, same colour scale — and the grid goes flat, its total landing an order of magnitude down (in general this gap only widens as N grows, since it's sampling noise against a real structural effect). Cell 4 checks the two derived marginals against closed form (2r, flat 1/2π) and they sit right on the curves. Cell 5 slices the raw darts by hand: near x≈0 the y-band is wide, near x≈0.99 it's pinched thin — x really does move y's spread — but two different θ slices print r half-widths that land on each other, because a ray from the centre crosses the same disc no matter which way it points. Same darts, never re-rolled; only the coordinates changed, and only one pair of them could still tell the two variables apart.
So the margins are strictly less than the joint. That leaves an honest question this chapter never closed: what is exactly enough? Keep one shadow, add the conditional, and watch the whole pair come back with nothing lost.
pick a pair — then the button unlocks
The bench — pick two ingredients, multiply cell-by-cell, and compare the result to the true joint on the left.
p(x,y) — the true joint, and the reconstruction (rebuild mode)
p_Y(y) — the forgetful shape the slice flattens toward (dashed)
the two shadows p_X, p_Y — they never move as you flatten
mismatch: orange = recon too high, blue = too low
↳ The identity: p(x,y) = p_X(x)·p(y|x) is just the conditional's own definition read backwards — so one shadow plus how the slice bends is the whole joint, for every world. Flatten the conditional until it forgets x and margin×margin becomes enough: independence is that one special case, watched live.
Fig. 22. Pick a pair of ingredients and press multiply: only p_X·p(y|x) returns the true joint — mismatch dead black, Σ|Δ| = 0.000 — while p_X·p_Y gives Fig. 15's untilted lie with the mismatch glowing. Now drag flatten to f = 1.00: the slice forgets x, and the two-margin reconstruction's Σ|Δ| drains to 0.000. Independence is exactly that one case — the conditional stops depending on x, so the margin alone becomes enough.
Histogram the darts in (x,y), build the two marginals from that data, multiply them, and subtract: the mismatch mapglows. Now relabel the same darts as (r,θ), do the identical thing, and the mismatch comes out as flat noise around zero. Nothing was re-randomised between those two panels. The rule did all the work, which is exactly the claim, checked by you, on data.
So what you carry out of here is not a set of formulas for joint distributions. It is a pile of mass and two moves. Project the pile and you get the marginal, which is the law of total probability wearing an integral. Slice the pile and rescale the cut face to area one and you get the conditional, which is Chapter 5 drawn at last. You can now look at any joint heatmap and read dependence off it by eye: drag a slice, and ask whether the shape changes. You can re-derive f(y|x) instead of recalling it. And you can say exactly why conditioning on a probability-zero point is legal, because you watched the dx cancel.
But one thing sits above all of it: the marginals are strictly less than the joint. That is not a slogan you have to take on trust, because the mismatch map is a picture you have watched glow. Here is where you are, and here is where that fact is about to get paid.
click 15 or 16 above
Two forward arrows, one fact: the marginals are strictly less than the joint. Click 15 or 16.
What you're looking at — the whole chapter compressed to one fact, paying off twice
grey dots = Ch 1–13, one variable at a time; the dashed line is where the plane begins
gold JOINT tile = you are here; the two mini-diagrams beneath it are the chapter's whole content — PROJECT (sum out a variable → the marginal) and SLICE (cut one row, rescale → the conditional)
blue/gold cells = the mismatch map, joint − marginal product; click 15 for the sum-survives payoff, 16 for the covariance payoff
ΣZᵢ² badge = Ch 13's stranded formula — reachable now that this chapter has sums of TWO variables
Fig. 23. The spine so far, with the joint lit gold — click 15 or 16. The sum survives dependence because it only ever touches the marginals; covariance can read a flat 0.00 while the mismatch map underneath it is still loudly glowing.
It pays off twice, almost immediately. Chapter 15 is going to tell you that E[X+Y] = E[X] + E[Y] even when X and Y are wildly dependent, which sounds like magic. It isn't, and you can already see why. The sum only ever touches the marginals. It never touches the dependence structure at all. Then Chapter 16 will warn you that zero correlation does not imply independence, and you are already halfway there. Covariance is going to be one number summarising a whole mismatch map, and one number can be zero while the map is not. Next, though, we do something simpler and just as sharp. We take a distribution, joint or not, and squeeze it down to two honest numbers: its center and its spread.