15The Center and the Spread
Chapter 14 left you holding one hard fact. The marginals are strictly less than the joint, and you watched the mismatch map glow in exactly the place where the two flat shadows had lost the world. So we can now describe two variables honestly, and that is a real thing to own. But a description is a whole picture, and a picture is not what you carry around. Here is the gap. If you could keep two numbers about a distribution and throw the rest away, which two would you keep, and how would you compute them? You have been writing down for six chapters without ever being told what they are. Since Chapter 9 you have typed N(np, npq). Since Chapter 10 you have divided by σ on faith. Today both debts get paid in full, from the ground up. And here is where we come out. There is really only one idea in this chapter, wearing several hats. Averaging is a probability-weighted sum, sums pass straight through sums, and that makes E[·] linear. So everything linear is free. It slides through addition. It slides through scaling. It even slides through dependence, because a sum only ever touches the marginals, and you already proved the dependence is not in the marginals. Then variance is that same operator aimed at a square, and the square is the one thing in the room that refuses to cooperate. Every price this chapter charges you is the square's toll.
01The two numbers you already write
Something slightly embarrassing is sitting in your own notebook, and it is not your fault. Go back and look at what you have been writing. In Chapter 9 we wrote the binomial's normal twin as N(np, npq), and you nodded. In Chapter 10 you standardized with Z = (X − μ)/σ, and you have been dividing by σ ever since. Now here is the embarrassing part. Neither μ nor σ has ever been defined in this book. They were labels on a curve, and you got very good at sliding them around. That is the quiet danger in notation, and it is worth saying out loud. Because you can type a symbol, you feel like you know it.
you can TYPE μ and σ² fluently after six uses — that fluency feels exactly like knowing what they are, and it isn't. This chapter finally defines the two numbers.
So those are the two un-cashed tokens sitting in your own past chapters, and today we cash them. I want to do it without handing you a definition from the sky. The standard opening does exactly that. Nearly every course starts with Definition: E[X] = Σ x·p(x), and you can execute that sum on day one and still never learn why the probabilities are the multipliers. Notice what is missing there: a reason. So we start somewhere you cannot possibly get lost. Not with probability at all. With data.
Roll a fair die sixty times and write down what happened. Say you got ten 1s, twelve 2s, nine 3s, eleven 4s, eight 5s and ten 6s. Those counts total sixty rolls, and they are a perfectly ordinary lopsided run of luck. Now take the average the way a child takes it: add up all sixty numbers and divide by sixty. Nothing clever has happened yet. Then we do one thing only to that expression, and it is algebra, not theory. Group the identical values together, factor each count out front, and push the division by sixty inside each term instead of leaving it at the end.
Watch what fell out of the algebra. The expression (1·10 + 2·12 + 3·9 + …)/60 became 1·(10/60) + 2·(12/60) + 3·(9/60) + …, and the answer never moved off 3.417. Those multipliers are relative frequencies: 10/60 is 0.167, the share of the sixty rolls that came up 1. I did not introduce them and I did not choose them. They fell out of an average you already knew how to take, and the shape they left behind is Σ value × weight. That is the whole engine. You built it yourself out of one fraction and one line of factoring.
Now the one honest step, and I will name it as a step rather than slide it past you. Those weights are counts over sixty, and counts over sixty are what probabilities look like once you have actually rolled the thing. So swap them. Put p(x) where the relative frequency stood, and the definition writes itself: E[X] = Σ x·p(x), the probability-weighted sum. I want to be exact about what just happened, because this is a choice and not a derivation. We are defining the expectation to be that sum. Nothing is being smuggled in. I am declaring the swap out loud so you can watch me make it.
So E[X] is a number the die can never show, which raises the real question. What is it for? Its oldest job is to set a price: E[X] is the fair price of a gamble, the most you should rationally pay to play it. Weigh each payoff by its chance, total them, and you are holding exactly what the bet is worth. Set your own price first, then watch the fair one appear.
Compute it for the fair die and you get 3.5, and there is your first shock. The die can never show 3.5. Look at the marker on the axis: there is manifestly no bar underneath it, because E[X] lives between the bars. It was never claiming to be an outcome. The name is what misleads you here, because "expected value" invites the everyday meaning, what you expect to happen, and that is a different thing wearing the same words. E[X] is a property of the whole distribution. It is not a prediction about a roll.
While we are being careful, let me refuse a line you will meet everywhere else. Most courses define E[X] as "the long-run average of many rolls". That definition is circular, and the circle is worth seeing. The claim that the long-run average settles onto E[X] is the Law of Large Numbers. That is a theorem, it arrives in Chapter 17, and its proof uses E[X]. So a definition leaning on it is a hole with a rug over it. Here is what we do instead. We define E[X] as the weighted sum, because the weighted sum is computable today, from the picture in front of us. "It is also the long-run average" is an IOU I am writing in front of you, and Chapter 17 pays it.
02The rail that balances
Now lay the number line down as a physical rail and glue a lump of mass p(x) at each value x. The more likely the value, the heavier its lump. That rail balances at exactly one point, and that point is E[X]. Every textbook says roughly this and then walks away, leaving you a nice picture you cannot compute with. An analogy you cannot compute with is a souvenir, not a tool. So let's ground the rail in its actual mechanism, because the mechanism is one line of physics you already know.
Torque is mass times signed lever arm. Mass to the left of the fulcrum pushes one way, mass to the right pushes the other, and the arm is the signed distance (x − μ) from the fulcrum at μ. On the die, with the fulcrum at 3.5, the value 6 sits at arm 2.5 and the value 1 at arm −2.5. The rail balances exactly when all those torques cancel, which is the equation Σ p(x)·(x − μ) = 0. Go and drag the fulcrum yourself. Put it in the wrong place and the arrows visibly fail to cancel, and the rail tips over. Slide it to 3.5 and they cancel to nothing.
Now do the algebra you just felt, and split that sum into Σ x·p(x) − μ·Σ p(x). The masses total 1, so the second piece collapses to plain μ. Set the whole thing to zero and it rearranges into μ = Σ x·p(x). Read that line twice, because it is the point of the section. The balance condition and the definition are the same equation with one term moved across. The fulcrum is not a metaphor for the mean. It is the mean. And the picture is now a machine you can run: a symmetric distribution balances on its own axis of symmetry, so you can name its mean with no algebra at all. Physicists call E[X] the first moment, and that is not a coincidence. It is literally the same sum.
The mean is one of three things people call a "center", and the three disagree. The median is the half-mass point: as much mass sits to its left as to its right. The mode is just the tallest bar. Both are read straight off the same rail, and neither needs a new idea. Now take a small dataset and drag one point far out to the right. Commit to a guess about the mean and the median first, then watch the two readouts.
The mean chases the outlier without limit. The median does not move at all. Everyone is told the median is robust, and almost nobody is told why. The reason is already in your hands, and it takes one sentence. The mean weights each mass by its lever arm, so a point far out has enormous leverage and can drag the fulcrum as far as it likes. The median counts nothing but which side of it the mass falls on. Distance is literally invisible to the median. Move that outlier a mile out, or a lightyear out, and the median changes by exactly zero.
There is a second thing your gut believes about the mean, and it is also wrong. The gut says the average is a typical value. It is not. The mean is a balance point, and a balance point can sit in vacuum. Morph the bell into two humps and watch the fulcrum.
E[X]=Σx·p(x) is a weighted average, not a "typical value" — it's the point where the two humps balance like weights on a seesaw. Slide the weights apart and the balance point doesn't move; only the weight sitting under it disappears.
The fulcrum parks itself in the empty valley between the two humps, at a value the variable essentially never takes. That mean is real, correct and useful, and it sits exactly where nothing ever happens. Hold on to that, because it is the honest reason we will need a second number before this chapter is out.
Continuous variables change nothing structural here, and I want to show you that rather than assert it. Chapter 8 told you the mass sitting in a sliver at x is f(x)dx. So the torque that sliver contributes is x·f(x)dx: its mass times its lever arm. Adding up slivers is integration. That is the whole move: E[X] = ∫ x f(x) dx. Same rail, same fulcrum, cells shrunk to nothing. Halve the bar width again and again, and watch the sum become the integral.
Two things to notice while the picture is still in front of you. First, the dx is not decoration. Chapter 8 was blunt that density is not probability, and here is that same fact doing work: f(x) on its own is a height, and the only reason the sliver has any mass to weigh is the dx multiplying that height. Second, and this one is an honest flag most courses never plant, that integral is written everywhere as though it always produces a number. It doesn't. If the mass reaches far enough out with enough weight behind it, the torque sum has no finite total. Then E[X] does not exist. Not "is infinite". Does not exist. Every distribution having a mean is a comfortable belief, and it is false. We meet the failure properly later, at the triangle of blocks.
03The average of a function
Here is the question that runs the rest of the chapter. What is the average of a function of X? Not X itself, but X², or ½mX², or the Fahrenheit version of a Celsius reading. The answer is outrageous in the best way: E[g(X)] = Σ g(x)·p(x). Push each value through g and keep the old weights. You never build the distribution of g(X) at all.
You should be annoyed by that rule, and the tradition jokes about it instead of curing it. Its standard name is "the law of the unconscious statistician". That name points at your discomfort and then walks off. The discomfort is legitimate, because the rule collides head-on with the chapter you just finished. Chapter 13 spent itself teaching you that you cannot push a density through g. You must chase the event through the CDF and pay the Jacobian |dx/dy|, and squaring a PDF gives a bogus curve whose area is not 1. Now we push values through g, keep the weights, and it works. So either Chapter 13 was theatre, or something here is genuinely different. If nobody tells you which, you quietly stop trusting the subject.
Here is the difference, in one sentence you already own: mass is conserved, price-per-inch is not. Chapter 13 was asking for a density, which is mass per unit of y. When g stretches the y-axis, the exchange rate between an inch of y and an inch of x is exactly what you have to pay for, and that payment is the Jacobian. LOTUS asks only for a total: mass times value, added up. Relabelling an outcome from x to g(x) does not create or destroy a single grain of mass. It only changes the number written on the grain. There are no inches in a total, so there is no exchange rate to pay. Now watch the regrouping that proves it.
LOTUS pushes every value through g, keeps the old weights, and totals. So a tempting shortcut whispers at you. Why not push the mean through g instead, and read off g(E[X]) in one step? That shortcut is wrong whenever g is not a straight line. The reason is a picture you can bend with your own hands, so go and bend it.
Same tokens on the table, two bookkeeping orders. Sort them into y-piles and total y × (pile mass): that is Chapter 13's route. Or leave them in x-order, write g(x) on each token, and total g(x) × p(x). Same tokens, same total. Written out it is three lines, and they start from Σ_y y·P(Y=y). Expand P(Y=y) into the sum of p(x) over the outcomes that map to y, which is legal because Chapter 3 says disjoint events add. Then regroup by x and you have Σ_x g(x)·p(x). Two different x's landing on the same y need no apology, because they simply fall in the same pile.
One more rung before the big theorem, and it is the rung everyone forgets to build. What is the average of a function of two variables? Look hard at X+Y. It is not a function of X. It is not a function of Y. It is a function on the pair. You read one outcome, you get the (x,y), and that pair hands you one number, x+y. So there is exactly one distribution in the room that can weigh it, and that is the joint.
Click any cell of the 6×6 grid and it tells you three things: its coordinates, its mass p(x,y), and the number x+y written across it. That is the setup, and there is nothing else to it. Each of the 36 cells is one outcome, every outcome hands you one number, so weigh each number by that cell's mass and total. Then name it. 2D LOTUS: E[g(X,Y)] = ΣΣ g(x,y)·p(x,y). The proof is the same regrouping argument you just ran, with the outcome label x replaced by the pair (x,y), on a grid instead of a line. I am labouring this because textbooks write E[X+Y] within a page of defining E[X], having defined the expectation of nothing but a function of one variable. The reader then has no idea which distribution the weights come from, and the next theorem becomes unstateable.
04The one rule that never asks
Start with the easy half, which costs one line. Put g(x) = ax + b into LOTUS and split the sum. The a pulls out of the first piece as a constant. The second piece is Σ b·p(x) = b·Σ p(x) = b, because the masses total 1. So E[aX+b] = aE[X] + b. Take the die and the map 2x + 10: its mean goes from 3.5 to 2(3.5) + 10 = 17. You can read that two ways at once. Every value doubles, so the total doubles. And every value shifts up by 10, but the weights add to one, so the total shifts by exactly 10, not by 60.
Put a box around Σ b·p(x) = b. It looks like nothing. It is the first time in this book that "total mass is 1" has done real algebraic work instead of sitting in an axiom list, and we will need it twice more before the chapter is out. Now the theorem.
Here it is: E[X+Y] = E[X] + E[Y]. Always. No hypothesis, no fine print, no independence. X and Y can be welded together so tightly that one determines the other, and the sum rule does not flinch. Its mirror image, E[XY] = E[X]E[Y], is false in general. That one has to buy independence before it is true. Two rules wearing the same shape, and only one of them has a price. Before you read on, go and bet.
Chapter 14's two worlds, side by side. On the left, two independent dice. On the right, one die X and its partner Y = 7 − X, not merely correlated with X but welded, X relabelled. Their shadows are pixel-identical, both uniform on 1..6, so E[X] = 3.5 and E[Y] = 3.5 in both worlds. The bet was whether the welding breaks the sum rule, and the room votes that it does, because welding is the tightest dependence the plane allows. Now look. In the welded world, X + Y = 7. Not on average. Every single time, because we built Y out of X to make it so. The whole distribution of the sum is a single spike, so E[X+Y] = 7 = 3.5 + 3.5. Total dependence, and the sum rule did not flinch.
Do not let that cool off, because the mechanism is on the grid you already own. Write it with 2D LOTUS: ΣΣ (x+y)·p(x,y). Split it into ΣΣ x·p(x,y) + ΣΣ y·p(x,y). Now do the inner sum first and watch the cells collapse. Σ_y p(x,y) is the row total of the grid, which is p_X(x), Chapter 14's marginal, the shadow. That is the entire proof, and the entire reason. The sum never looked at the joint. It only ever touched the two marginals. And you proved in Chapter 14, with these exact two worlds, that the dependence is not in the marginals. Dependence lives in the mismatch map, and projecting onto a margin destroys that map. So dependence cannot reach E[X+Y]. The sum's arithmetic throws the map away before it ever reads it.
Now turn the same handle on the product and let it break in your hands. Write ΣΣ xy·p(x,y) and try to split it. You can't. The xy stays welded to the joint mass, and the only thing that can prise it apart is Chapter 14's factorization, p(x,y) = p_X(x)p_Y(y). So the product rule fails exactly where you would now predict. In the welded world the true average of the product is E[XY] = 28/3 ≈ 9.33, while the product of the two averages is E[X]E[Y] = 12.25. Same pair, same afternoon, same grid. One rule lives, one dies, and you can point at the precise line where they parted. The sum's inner Σ collapsed onto a margin. The product's didn't. That is your condition audit forever. Sums touch only the marginals. Products read the whole joint.
The sum rule generalises with no extra work, because you just iterate the two-term version: E[aX + bY + c] = aE[X] + bE[Y] + c, and E[ΣXᵢ] = ΣE[Xᵢ] for any number of pieces, dependent or not. This is linearity of expectation, and it is the highest-leverage tool in the subject. Now watch it stop being a fact and become an engine.
What is E[X] for a Binomial(n,p)? The brute-force route is Σ k·C(n,k)p^k q^(n−k), a genuinely unpleasant sum of factorials. Instead, remember what a binomial is. Chapter 9 built it as n stacked Bernoullis. So write X = X₁ + X₂ + … + Xₙ, where Xᵢ is 1 if trial i succeeds and 0 if it doesn't. Each atom is trivial: E[Xᵢ] = 1·p + 0·q = p. Then apply linearity and you are done. E[X] = np, in three lines, with no factorials anywhere, so twenty tosses of a fair coin average 20 × 0.5 = 10 heads.
See the inversion, because it is the whole art and nobody teaches it. The move is not "I have a sum, so I may split it". The move is "I want a hard expectation, so let me manufacture a sum of trivial pieces and split that". The arrow points backwards. And now the twist that makes it permanent. Switch the sampler to draw the n balls without replacement. The trials are now flagrantly dependent, because every ball you draw changes what is left. On the left, the factorial sum becomes a hypergeometric nightmare. On the right, the derivation does not change by a single character. Each Xᵢ still has E[Xᵢ] = p, so the answer is still np. Linearity works on the marginals, and each atom's marginal never noticed the other draws.
05How far from the middle
One number is not enough, and here is the proof. Two distributions sit side by side: a tight cluster around 3.5, and a coin flip that pays 1 or 6. Both have E[X] = 3.5. The fulcrum cannot tell them apart, and apart from that one number they have nothing in common. So we need to measure the spread as well. The obvious first try is to average the deviation, X − μ. Go and try to beat it.
↳ Why it can't break: μ = E[X] = Σx·p(x) is defined as the balance point, so Σp(x)(x−μ) = Σx·p(x) − μ·Σp(x) = E[X] − μ·1 = 0 — always, for any weights, not just the equal ones this meter uses. That's linearity of expectation, not luck: "the negatives cancel" isn't roughly true, it's exactly true by construction, which is exactly why raw deviation is useless as a spread measure. Flip to |Δ| or Δ² above and the same meter finally moves — mean absolute deviation is real, honest, and used elsewhere. We reach for the square anyway, for two reasons the rest of this book leans on: it's differentiable everywhere (|Δ| has a sharp corner at Δ=0 — no slope to chase there), and it expands — (a−b)² = a² − 2ab + b², three clean terms an algebra can grab, while |a−b| has no such expansion; it just stays stuck as itself.
The average deviation gives zero. Not approximately. Not usually. Exactly zero, for every distribution that has a mean. Drag the masses wherever you like, make the shape as lopsided and vicious as you can, and the meter stays nailed at 0.000. You cannot beat it, because you are fighting the fulcrum itself. Signed deviations cancelling is the balance condition from the rail, and one line of linearity kills the idea for good: E[X − μ] = E[X] − μ = 0. The standard line here is "we square so the negatives don't cancel", and that line is quietly poisonous, because it implies they roughly cancel. They cancel identically and structurally. The failure of the obvious idea is a theorem, not a nuisance.
So repair it. Kill the signs with the absolute value, |X − μ|, or with the square, (X − μ)². Flip the switches on the meter and both come alive. And I want to be honest that this is a fork, not a fact. Most courses present the square as inevitable, and it is a choice. Mean absolute deviation is a real statistic that real people really use. We take the square for two reasons. The square is differentiable everywhere, while the absolute value has a corner exactly at the interesting point. And decisively, the square has an algebra: it expands. Everything that follows in this chapter, and in the next three, is that expansion cashing out. The absolute value expands into nothing, so it goes nowhere.
Define it: Var(X) = E[(X − μ)²], the expected squared deviation. Notice that we needed no new machinery to get there. It is just LOTUS with g(x) = (x − μ)², weighing a new number against the original masses, where μ is the number we already computed on the rail.
μ is the mean, already pinned down back in Ch 9 — we don't recompute it here. LOTUS says: to get Var(X)=E[(X−μ)²], plug g(x)=(x−μ)² into the same weighted sum that gave you E[X]. That sum is literally a moment of inertia, Σ m·r² from mechanics — mass times squared distance from a pivot — with probability standing in for mass.
Aha: Var comes out in cm² because it's built from squares, so it can't be marked on the cm rail it lives on top of — try it, it fails. σ=√Var undoes the square and lands back in cm, which is exactly why both symbols exist: Var is the algebra's currency, σ is the only one you can point to on an axis. And it's why every "σ²" you've written for a Normal(μ,σ²) since Ch 9 was already this same object, six chapters ahead of being checked.
Draw each deviation as a segment, then build the actual square on it. "Weighted by the square of the lever arm" stops being words at that moment: a deviation of 2 draws a tile of area 4, a deviation of 4 draws a tile of area 16, and the area explodes as a mass moves out. That lands the first consequence immediately. The units come out squared. A variance of heights measured in cm is a number in cm², and you cannot draw cm² on a rail. So we take the root: σ = √Var, the standard deviation, back in the original units. On the fair die the variance is 2.9167 in squared pips, and the standard deviation is √2.9167 ≈ 1.71 pips. That is the division of labour, and it is why both quantities exist. Variance is the thing you do algebra with. σ is the thing you put on the axis. One is an area, one is a length, and only a length can be marked on a rail. While we are here, the name over-promises. Your gut reads "standard deviation" as the typical distance from the mean, and it is not that. It is the root-mean-square deviation, and squaring gives distant points extra weight, so σ is always at least the average distance.
Then the physics, and this one is not a gesture. E[X] weighted each mass by its lever arm and handed you the center of mass. Var weights each mass by the square of its lever arm. Now look at what a physicist writes for moment of inertia: Σ m·r². And we are computing Σ p·(x−μ)². That is the same equation with p in place of m, so variance is literally the moment of inertia of the probability mass about its own center of mass. The first moment says where the rail balances. The second central moment says how hard the rail is to spin. Feel a far mass cost you a fortune on the spin-up readout, and you will never mis-remember which of the two squares.
And now Chapter 9's debt gets paid on the spot, in two lines. For N(μ, σ²), the density is symmetric about μ, so the fulcrum argument alone gives E[X] = μ, with no integration at all. And the second line, ∫(x−μ)²f(x)dx = σ², comes out by parts. So yes: the μ and the σ² you have been typing since Chapter 9 really are the mean and the variance of that curve. That was a promise made by suggestive notation, and nobody ever checks it, because it looks too obvious to mention. Looking too obvious to mention is exactly how a belief goes six chapters without once becoming knowledge.
06The algebra of the square
Variance has a second face, and it is the one you will actually compute with: Var(X) = E[X²] − μ². Expand the square inside the bracket, let linearity walk through the pieces, and this form falls out. But I want you to read what it says, not just use it. E[X²] is square-then-average. E[X]² is average-then-square. Two pipelines, same start, opposite order.
You now hold two formulas for one number, and they look nothing alike. The definition squares each deviation and then averages: Σp(x)(x−μ)². The shortcut squares first and then subtracts the mean's square: E[X²]−μ². The two routes share not one middle number, yet they must agree, because one is the other with the square expanded. Run both on a tiny gamble and watch them lock.
Belt A squares each die value first, 1, 4, 9, 16, 25, 36, and then averages, landing on 91/6 ≈ 15.17. Belt B averages first to 3.5, then squares, landing on 12.25. Subtract belt B's answer from belt A's and the gap is 2.9167, which is 35/12, which is the variance of a fair die. So E[X²] − μ² is not a shortcut to memorise. It is a measurement of the square's refusal to commute with the average. If squaring were linear that gap would be zero, and there would be no such thing as spread. The gap also makes the sign unforgettable. Belt A must win, because the amount it wins by is a variance, and variances are built out of squares. So a free theorem falls out on the way: E[X²] ≥ E[X]², forever, for everything.
Two traps in that derivation, and both are worth stopping on. The first trap is notation. E[X²] and E[X]² differ by the position of one bracket, and your eye will glide straight over it. On the die, square-then-average gives 15.17 and average-then-square gives 12.25, nowhere near each other. People compute E[X]², call it E[X²], get a variance of zero, and stare at the page. That is not carelessness. The notation genuinely does not signal that these are two different pipelines, which is why we drew them as belts.
The second trap arrives mid-derivation. We write E[(X−μ)²] = E[X² − 2μX + μ²], then pull the μ out of the bracket, and something in you objects. How can μ come out of the expectation? μ came from X. It is made of X, so surely it is random too? The answer is trivial once said and invisible until said. μ is a number you already computed and wrote down. For the die it is literally 3.5. It has no randomness left in it, because it was averaged away. Run the derivation with that number written in and the objection cannot even form: E[X² − 7X + 12.25] = E[X²] − 7E[X] + 12.25, and nobody wonders whether 12.25 is random. Then the last two terms fuse, −2μ·μ + μ² = −μ², and you have it.
Next question. How does spread respond to an affine map? Everybody's height doubles overnight. What happens to the variance? Commit to an answer first.
Var is Σ p·r² — an AREA, not a length. Double every r and every square's area quadruples; nothing else can happen. Sliding moves the fulcrum and every mass together, so no r changes and b never reaches Var. The two rules combine to pay the z-map's debt: a=1/σ, b=−μ/σ gives Var(Z)=σ²/σ²=1 and E[Z]=0, in one line.
Most people say the variance doubles. It quadruples. The gut treats Var as a length, because "spread" is a length-word and we draw σ on the axis as a length. But Var is an area. And the physics settles it faster than the algebra. Variance is Σ p·r², so double every radius and every single term picks up a factor of 2² = 4. It cannot possibly do anything else. Now the shift. Slide the whole apparatus ten metres to the left. The fulcrum slides with it, every lever arm is untouched, and the inertia cannot have changed. So b is gone for a physical reason before it is gone for an algebraic one. Put the two together: Var(aX+b) = a²Var(X), and SD(aX+b) = |a|·σ. That absolute value is not pedantry. Setting a = −1 flips the picture left-to-right, which cannot change how spread out it is, and a formula returning a negative spread is broken.
Try the rules on a real map. F = (9/5)C + 32. A mean of 20 °C becomes (9/5)(20) + 32 = 68 °F, so the mean shifts and scales, because E feels both knobs. The variance ignores the 32 entirely and picks up (9/5)² = 81/25 = 3.24. And now Chapter 10's debt comes due, five chapters late. Look at Z = (X − μ)/σ and see what it actually is: an affine map with a = 1/σ and b = −μ/σ. So run it through the two rules. E[Z] = (μ − μ)/σ = 0. And Var(Z) = (1/σ²)·σ² = 1. That is it. The z-map you have used on faith since Chapter 10 really does land at mean 0 and variance 1, and now you have proved that instead of assuming it. Until this moment σ was a ritual, the thing you divide by. It is a quantity now.
Which brings us to the chapter's last trap, and you are going to walk straight into it. E just handed you an unconditional sum rule with a memorable slogan. So does variance have one too? Is Var(X+Y) = Var(X) + Var(Y)? Vote first, then watch it die.
Set Y = X. Then Var(X+Y) = Var(2X) = 4Var(X), by the rule you learned one minute ago. The proposed rule says the answer should be 2Var(X). So the proposed rule is dead, and no counterexample had to be invented: your own new rule shot it. Expand honestly instead. Apply the definition to S = X+Y, note that S − E[S] = (X−μ_x) + (Y−μ_y), and square that binomial. You get Var(X+Y) = Var(X) + Var(Y) + 2·E[(X−μ_x)(Y−μ_y)]. That third term is a real quantity. It is not zero in general, and it vanishes exactly when X and Y are independent. So "variance adds" is true, but only conditionally, and never for the reason people think.
This matters more than any other error in this chapter. Assuming the pieces do not move together is how people underestimate portfolio risk, error bars and queueing delays, and the assumption is a pattern-match off a slogan. Now bring back the two worlds one last time and look at what that third term is for. Both worlds have E[X+Y] = 7, identical, because their shadows are identical. But World A, the independent dice, has Var(X+Y) = 35/6. World B, the welded pair, has Var(X+Y) = 0, because its sum is a constant 7 and never moves at all. The two worlds are indistinguishable to the mean and screamingly different to the variance. So the variance is reading something the mean cannot see. It is reading the joint, and the cross term is that reading.
I am going to leave that cross term unnamed, and the silence is deliberate. It has a name, that name is the whole of the next chapter, and you are now holding the exact question it answers.
07Counting the same pile sideways
One last thing expectation can do, and it looks at first like a misprint. For a random variable taking values 0, 1, 2, 3 and so on, E[X] = Σ_(k≥1) P(X ≥ k). Read that again. There is no value multiplying anything, and it is a sum of tails. Every expectation you have ever computed had the shape "something times its weight", and this formula has no multiplication in it at all. The weights are not even masses, but the cumulative quantities you associate with CDFs. So where did the value go?
The standard proof answers by swapping the order of a double sum, and you can verify every line of it and learn absolutely nothing. So let's do something else. The whole mystery is hiding inside the ×k, so take that apart. k·p(k) is k copies of p(k) added together, because multiplication by a whole number was always repeated addition. For k = 3 that is p(3) + p(3) + p(3), three copies and nothing more. So don't draw one bar of height k·p(k). Draw k stacked unit blocks, each worth p(k).
Now you have a triangle of blocks, and the whole trick is that one pile can be counted two ways. Row k has k blocks in it, each worth p(k). Read the pile row by row and you are summing k·p(k), which is the definition of E[X]. Same blocks, nothing moved, nothing added. Now read the same pile column by column. Column j contains exactly one block from every row k ≥ j, so its total is Σ_(k≥j) p(k), and that is P(X ≥ j). Column 3, for instance, collects one block from row 3, one from row 4, one from row 5, and so on upward. The tail was never a new object. It is a column of the very triangle you were already summing. The value did not go anywhere. It was always k separate ones, and the tail-sum is what you get by counting them column-wise instead of row-wise. One honest note on the licence. Every block has non-negative area, which is exactly why we may rearrange the pile at will. And if the tail area is infinite, the pile has no total and E[X] does not exist, the same flag we planted at the slivers, now with a picture under it.
Then watch the column trick earn its keep. For the geometric, P(X ≥ k) = q^(k−1), and you read that straight off the definition: the first k−1 trials all failed. So E[X] = Σ q^(k−1) = 1/(1−q) = 1/p. One line, no series differentiation, and with p = 1/6 that is 1/(1/6) = 6, the six rolls you wait for a six. The continuous twin is the same picture with the cells shrunk: E[X] = ∫₀^∞ S(x)dx, the area under the survival curve. So E[Exp(λ)] = ∫₀^∞ e^(−λt)dt = 1/λ. One line, no integration by parts. Two ugly pages collapse into two lines, because you counted the same pile sideways.
08Temperature is a second moment
The payoff comes from reading one equation backwards. We taught Var(X) = E[X²] − μ² as a way to get the variance, and you filed it under exactly that. Now move one term across: E[X²] = Var(X) + μ². Read that way, the equation hands you the second moment for free, out of a variance you already know. On the die it returns 2.9167 + 12.25 = 15.17, which is the square-then-average number from the belts. It is a one-bracket move you would never make on your own, because nothing in the exercises ever asks for it, and every physical application in this chapter runs through it.
So point that move at a gas. A molecule's velocity components v_x, v_y, v_z are each N(0, σ²), and its kinetic energy is ½m(v_x² + v_y² + v_z²). Watch how little work this takes. Linearity walks straight through the sum. LOTUS handles each square without ever constructing the distribution of v². And each E[v_x²] = Var(v_x) + 0² = σ², by the equation we just read backwards. Three components, each contributing ½mσ², and 3 × ½ gives E[KE] = (3/2)mσ², in four lines.
Sample a million molecules, compute ½m|v|² for each, average them, and the simulation lands on the number the algebra predicted before any simulation was run. Now set that prediction against the thermodynamic (3/2)kT and read what falls out: mσ² = kT. Say it plainly, because it is the reason this chapter exists. Temperature, the most physically immediate quantity in the course, the thing your skin measures, is not a substance and not a force. It is the second central moment of a distribution. It is the variance of the velocity distribution, wearing a different unit.
And one last thing, because this is the misconception's last hiding place. Look back at that sum of three squares and ask the obvious question. Are v_x and v_y independent? We don't know. We didn't check. And it changed nothing. That is what linearity bought you.
So here is what you carry out, and it is much smaller and much stronger than a page of formulas. One operator and one warning. E[·] is a probability-weighted sum, and sums pass straight through sums, so it is linear. That is not a fact you memorised. It is a move you can audit, because you watched the inner sum collapse onto a marginal with your own eyes. That single sight settles which rules cost independence and which are free, permanently. And you can now do what almost nobody who merely passes this course can do: reach for linearity backwards. The warning is the mirror of the operator. The square is the one nonlinear thing in the room, and every price you were charged was the square's toll. So you can re-derive rather than recall. It all comes back from expand-the-square plus linearity, and the fulcrum and the moment of inertia let you predict the answer before you compute it.
Aha: that's why Ch 16 isn't arbitrary — you're already holding a real number that reads the joint, that the mean never saw, and that you were never allowed to name.
Look where that leaves you. You are holding an unnamed cross term that you know reads the joint, and you have seen exactly why you need it: two worlds, one shared mean, two wildly different spreads. So Chapter 16 is not arbitrary. It is inevitable, and it opens by giving that cross term its name. After that, Chapter 17's Markov and Chebyshev will bound the tails of any distribution using nothing but the two numbers you built today, and it will pay the IOU I wrote you about the long-run average. And Chapter 18's Central Limit Theorem is a question about the spread of a sum, which is a question you now know cannot be answered without the term you were handed and told to hold.