Quantitative Finance — the Mathematics of Markets · chapter 16
16Sampling, Standard Error & Estimators
Chapter 15 closed on a sentence that has been quietly true since Chapter 9. An expected value is never a fact about the world. It is the output of a model somebody chose, with three slots filled in by hand. Kelly ate the number p and never asked where it came from, and we watched what happened when the answer was "a hundred noisy flips". So here is the question this chapter turns on, and the whole of Part 2 grows out of it: who fills those slots in when nobody knows the answer? From this page the arrow runs backwards. You hold the data, and the model is the unknown. That reversal does something strange and wonderful to a number you compute from data. It makes that number random — not the digits on your screen, which are finished, but the procedure that produced them. A number computed from data is itself a random variable, so it has a distribution of its own. The width of that distribution, and not the width of your data, is exactly how much you are allowed to believe yourself. By the end you will be able to draw a cloud you can never observe, using one sample and a square root. You will know why n−1 is a derivation and not a convention. And you will know why a century of market history still cannot tell you what a stock is expected to return.
Look at what this page stands on. Chapter 9 gave us the sample space and favourable-over-total. Chapter 11 gave us E[X], linearity, variance and Jensen's inequality. Chapter 12 gave us the Normal and the 68-95-99.7 rule. Chapter 13 did the heaviest lifting: the √n law, the standardized average, and the central limit theorem, plus two words it used and deliberately did not define. Those two words were standard error and sampling distribution, and this chapter is where that debt gets paid in full.
Where Part 2 begins — flip the DIRECTION switch, then re-run the world and watch which number moves.
Flip the arrow, then re-run.
THE TWO OBJECTS
θ the truth—
θ̂ your estimate—
Right now both are empty: in Ch 9–15 nobody had to estimate anything.
Predict: re-run the world. Which number moves?
flip the switch ▸
What you're looking at — the same machine, run in both directions
the model's three slots — Ω what can happen, P with what chance, μ/σ what it's worth. Ch 9–15 filled these by hand
the data — endless outcomes when the model is known; one sample of n = 100 when it isn't
θ (“theta”), the true value in the slot: fixed · unknown · never observed
θ̂ (“theta-hat”), what you compute from the sample: known · different every single run — so it has a distribution of its own
Fig. 1. Chapter 15 ended on a machine that always answers: hand it Ω, P and a value scale and it returns E[X]. Notice who filled those three slots — you did, from the problem statement. That is the arrow running left to right, and it is every chapter from 9 to 15. Now flip the switch. The slots blank out to question marks, the arrows turn, and the only thing left on the table is one sample of a hundred outcomes. Two objects fall out of that reversal and must never be confused again: θ, the true value sitting in the blanked slot — fixed, unknown, never observed — and θ̂, the number you compute from the sample. Predict which one moves when you re-run the world, then press it. The truth never budges; your estimate lands somewhere new every single time. That is the keystone of the whole of Part 2: an estimate is itself a random variable, so it has a distribution of its own — the sampling distribution — and how wide that is, is the answer to "how wrong is my number?" Chapter 13 said those words in passing. This chapter earns them.
That is the seam. Everything before it ran model → data: here is p, now tell me what the coin will do. Everything after it runs data → model, and the first casualty of that reversal is a habit almost nobody notices they have.
01The arrow turns around
Here is a price series with a hundred daily returns in it, and I want you to tell me what this stock is expected to return on an average day.
You will average them, because that is obviously the right thing to do, and you will get 0.043% a day. Now let me tell you something you could not have known. That series came out of a machine, and I set the machine's true expected daily return to 0.020%. Your answer was off by more than double.
Press the machine again, with the same settings and the same hidden truth, over a completely fresh hundred days. This time the average comes back −0.011%, which does not even have the right sign. Press it once more and you get 0.061%. The truth never moved. Your answer moved every single time.
One machine, one curtain, one big button — press it and watch the hidden truth hold perfectly still while your answer refuses to.
100 days. what do they average?
draw #1 · nothing pinned yet
The machine printed these 100 days. Its dial — the one number that says what it really returns per day — is behind a curtain. Average the days and commit to an answer.
What you're looking at — a truth that never moves, and an answer that never stops moving
The estimand μ — the machine's true expected daily return, 0.020%. It is real, it is fixed, and without the curtain trick you would never see it. Watch the gold marker across every press: it never moves a digit.
The estimate x̄ — the average of the 100 days you happened to get. The bars are that sample; the dot is the one number you squeezed out of it. Press again and a fresh 100 days from the identical machine hands you a different dot.
The miss, in percentage points. It is not a bug, a bad model or a slow day — it is what a finite sample always does. Turn on keep the dots and the misses stop looking like errors and start looking like a shape.
Fig. 2. Same machine, same hidden setting, three presses: 0.043%, then −0.011%, then 0.061%. The truth sat at 0.020% the whole time and never twitched. So x̄ is not μ seen slightly out of focus — the estimand and the estimate are two different objects, and every press proves it again.
That is the phenomenon, and now it earns its two names. The fixed number out in the world is the estimand, written θ, and in this case it is μ = 0.020%. It is real, it is fixed, and you will never see it. The number you computed is the estimate, written θ̂ and read "theta hat", and here it is x̄ = 0.043%. It is known, it is sitting on your screen, and it is different every time.
These are two different objects, and almost everyone arrives believing they are one object seen at two levels of focus. That belief is completely reasonable and it is the thing this chapter has to take away from you. In every problem you have ever solved, "the mean" was handed to you in the question, so computing an average felt like arithmetic rather than estimation. Nothing so far has ever put a number in front of you that is real, fixed, and permanently invisible.
Now the second inch, and this one is aimed at finance readers in particular. The word population sounds like it means a big group of people, so it sounds like something you could in principle finish collecting. Suppose you own every daily return the S&P 500 has printed since 1928. That is about twenty-five thousand numbers, the complete historical record, nothing missing. Surely now you know μ exactly?
Four things a finance reader calls “the data” — sort each into a bin before you press reveal.
tap a bin for each card →
0 of 4 cards placed
What you’re looking at — four things a reader calls “data”, and only one of them is a population.
the population is the process — firms, traders and shocks, printing returns forever. It is a machine, not a bigger file. Only card d.
a, b and c are all samples: finite piles of numbers the machine happened to print. Size does not promote a pile to a population.
the trap: a COMPLETE RECORD feels like everything. It is one draw — and the twentieth century cannot be re-run to get a second.
footnote: more history does shrink your error on μ and on σ. It does not shrink σ itself — the world stays as wild as it was. That first shrink is the standard error, next.
Fig. 3. You own every daily S&P return since 1928 — twenty-five thousand numbers, nothing missing. Surely, then, there is nothing left to sample: the mean of that file simply is μ. Sort the four cards first, before any answer shows, and see where you put card b. Almost everyone drops it in POPULATION, because completeness feels like totality. Press REVEAL and watch three of the four slide the other way. The population is not a bigger dataset; it is the process — firms, traders, shocks — that keeps printing returns, and it never stops printing. The complete record is what came out of that machine on one particular run of history: one draw, of size 25,000, from something that could have produced a different century. Then press REDRAW THE CENTURY, which is the whole chapter in one dead button. It cannot fire. The counter reads 1 of 1 and always will, which is exactly why the mean you computed from that file is an estimate carrying an error, not the truth carrying none.
You do not, and the reason is the whole of Part 2 in one sentence. The population is the process, not a bigger dataset. What generated those twenty-five thousand numbers was a data-generating process — a machine of firms and traders and shocks — and the record is one long draw from it. Chapter 15 already told you that the model is something somebody chose, never something the world hands over. The one thing you can never do is take a second draw of the twentieth century.
So the arrow has turned around, and it has left us holding exactly two objects. There is a θ we want and cannot see. There is a θ̂ we have and cannot trust. Everything else in this chapter is the relationship between them.
02Where the randomness actually lives
Let us make that relationship precise, and the first step is a piece of notation that is switched silently on every page of every statistics textbook. Once your data is on the screen it is not moving. Those hundred returns are finished. So what, exactly, is random?
The randomness is in the design, not in the numbers. Before you look, the sample is X1, X2, …, Xn — n random variables, each an independent draw from the same population. After you look it is x1, x2, …, xn, which are n perfectly ordinary numbers. Capital letters are the plan. Lowercase letters are what the plan happened to produce this once.
One population, one button. Press it and every lowercase number changes — and not one CAPITAL symbol ever moves.
PROBE — click a statement ↓
press DRAW again · watch what does NOT move
blue = the machine (capitals) gold = one press of it (lowercase)
What you’re looking at — the same population twice: as a plan you can ask questions about, and as one finished press.
the density and every CAPITAL symbol are the procedure — “draw 8 daily returns from this world.” Nothing here moves, ever, however many times you press.
the tray and every lowercase number are one press: x₁ = a value, x̄ = its mean. Press again and all eight change — that flicker is the randomness.
P(x₁ > 0) and E[x̄] are refused: those numbers are already determined. The probability that +1.42 exceeds 0 is 1, not 0.516.
Fig. 4. A population of daily returns sits on top — μ = +0.04%, σ = 1.00% — and it is described entirely in CAPITALS: X₁, X₂, …, X₈ is not eight numbers, it is a recipe, the plan for a press you have not made yet. Hit DRAW. Eight values fall out of the density and land in the tray, and there they become x₁, x₂, …, x₈ — lowercase, printed, finished. Press again: all eight lowercase numbers change and the mean x̄ jumps with them, while above, not one capital symbol so much as flickers. Hold the second button and 40 presses fire in a couple of seconds — the tray becomes a blur and the blue curve does not move a pixel. That is the whole lesson: the randomness was never in the numbers on your screen. It is in the press. Now use the PROBE. Ask P(X₁ > 0) and the panel answers — it shades the area under the plan to the right of zero and reads 0.516, an answer that was available before you ever touched the button. Ask E[x̄] and it answers μ, highlighting the DRAW button itself, because E averages over presses of that button. Now ask the same two things in lowercase. P(x₁ > 0) is refused, and the offending number is circled in the tray: x₁ is a specific printed value, so the probability it exceeds zero is 1 or 0, never 0.516. E[x̄] is refused for the same reason — you already computed it; there is nothing left to average over. Keep this straight and the rest of Part 2 is readable: when we write E[X̄] = μ or ask how much X̄ scatters, we are asking about the machine, not about the one number your spreadsheet is showing you.
Press the button and dots fall out. Press it again and different dots fall out, from a machine that never changed. Every probability statement in statistics is a statement about that press. It is never a statement about the dots that already landed.
This matters more than notation usually does. When you read E[X̄], you are not being asked for the expectation of a quantity that is already determined, which would be either nonsense or a tautology. You are being asked what the averaging procedure does across the presses you did not make. Hold that and the rest of the chapter parses. Miss it and every sentence from here reads as word salad.
Now the assumption hiding inside "each an independent draw from the same population". Those three words get compressed into the letters iid, and they are usually muttered before the real work like a spell nobody unpacks. They are the load-bearing assumption of everything that follows, and in market data they are false.
One dial, φ — how much each draw remembers the last one. Turn it, watch iid break, then watch what the break costs you.
STICKINESS φ — HOW MUCH EACH DRAW REMEMBERS THE LAST
φ = 0.00 · no memory at all
THE FOUR TABULATED φ
PREDICT — AT φ = 0.50, THE TRUE SPREAD vs σ/√n
pick one — the dial will jump to 0.50 and show you.
φ = 0.00 · iid holds · bar honest
What you're looking at — the same 100 draws, made stickier, and the bill that arrives for it.
the draws in the order they arrived. φ is how much draw i remembers draw i−1. At φ=0 the strip flickers; wind it up and it grows runs — long blocks on one side.
what the formula assumes: the population's own shape, and the bar σ/√n = 1/√100 = 0.1000 it hands you. Neither moves — the world never changed.
your one sample: its histogram and its mean. Sticky draws make it lopsided and gappy, and its mean wanders outside the ±0.10 you would have quoted.
the true spread of that mean across 200,000 re-runs. The blue block inside it is all your quote covers. σ/√n is not wrong — it answers a question about a world where each draw ignores the last.
Fig. 5. Every formula in this chapter rests on four letters — iid, independent and identically distributed — and the honest thing to do is break them in front of you rather than let you meet the word cold. The dial is φ, the stickiness: how much each draw remembers the one before it. The 100 shocks underneath never change, and the population never changes — its spread is pinned at exactly 1.0 at every setting, so nothing you see is the world getting wilder. At φ = 0 the strip flickers, the histogram sits on the population like a decent tracing, and the sample mean lands inside the ±0.10 you would have quoted. Now wind it up. The strip grows runs — twenty in a row above the line, then thirteen below — because a sticky world does not hand you 100 fresh facts, it hands you a handful of facts repeated. The histogram goes lopsided and gappy: the same population, badly represented. And panel 3 is where it costs money. σ/√n = 1/√100 = 0.1000, always, because that formula never looked at the ordering — it cannot see what you are looking at. But re-run this world 200,000 times and the mean actually scatters by 0.1721. The blue block sitting inside the red one is the whole of your quoted uncertainty; everything outside it is error you promised was not there. That is 72% too narrow, and the closed form says why with no simulation at all: √((1+φ)/(1−φ)) = √3 = 1.732. Read the last line carefully, because it is the point: σ/√n is not wrong. It is a perfectly correct answer to a question about a world in which every draw ignores the last one. Market returns are not that world — overlapping windows, volatility clustering and autocorrelation all put φ above zero — so when you annualise by √252 or quote a mean return with an error bar, you are quietly betting on φ = 0. Chapter 21 is where we stop betting and start measuring it.
Turn on the dial that makes each draw lean toward the one before it, and watch the sample stop looking like a fair picture of the population. You get clusters, you get runs, and you get a mean that wanders further than it should. That leaning has a name — autocorrelation — and Chapter 21 gives it a full working. Meet it once here as a picture, so that its name later lands on something you have already seen.
And take the honest cost with you now rather than later. At a stickiness of 0.5, the spread of the sample mean measured across 200,000 runs is 1.72× what the independent-draws formula predicts. So the error bar you would have quoted is 72% too narrow, and you would have quoted it with total confidence. Everything in this chapter assumes the button was pressed cleanly. Markets do not press it cleanly, and Chapters 21 and 32 exist because of that. The formula is not wrong, by the way, and that is the uncomfortable part. It answers a question about a world where each draw ignores the last one.
03An average is a noun
A statistic is any number you compute from the sample. The mean is one of them, and so are the variance, the maximum, the median, and a Sharpe ratio. Feed random variables into a function and what comes out is a random variable, and that is not a new theorem at all.
It is Chapter 11's definition, read on a bigger sample space. A random variable was only ever a rule that takes an outcome and returns a number. Here the outcome is the whole sample and the rule is "add them up and divide by n". The definition already covers this case, and Chapter 9 built a sample space out of one roll rather than out of a hundred rolls at once. Nothing in that construction ever objected to the larger version. Saying so out loud is the entire move, and almost no course says it.
Which means X̄ has to stop being a verb. Right now it is something you do: add, divide, done. It has to become a noun — an object with a distribution, a mean, a variance and a shape. As long as it stays a verb, the phrase "the variance of the sample mean" is unparseable, because verbs do not have variances.
Here is the smallest example that forces the change, and you already own every piece of it. Roll two dice. But instead of taking the sum, take the mean of the two faces.
Five beats that turn “average them” from a verb into a noun — then count all 36 two-dice samples by hand.
THE FIVE BEATS — JUMP ANYWHERE
PREDICT · Var( add them and divide by 2 ) = ?
Pick one. The panel will not move until the question is fixed.
beat 1 · a verb, not a noun
What you're looking at — one rule, re-read as a random variable, then its entire law counted out of 36 cells.
the verb: “add them and divide by 2” is an instruction. Instructions have no spread, so beat 1 refuses to answer.
the noun: Ch 11 says a random variable is a rule from an outcome to a number. Let the outcome be the whole sample (d1, d2) and the rule qualifies.
the 36 cells, one per equally likely sample. Brightness = how common that mean is; tap any cell (or bar) to spotlight every cell sharing its x̄.
the value on the spotlight — its count out of 36 is its exact probability. No limit theorem, no simulation: 1.2076147 twice over.
Fig. 6. Here is the sentence that quietly breaks people, and it is worth breaking it slowly. Somebody writes Var(x̄) and you are supposed to nod — but look at what x̄ was when you met it: add the numbers up and divide by how many there are. That is an instruction. It is something you do. Asking for the variance of an instruction is like asking how tall a recipe is, which is why beat 1 refuses to answer and just sits there. Beat 2 is the whole repair, and it is one line of Chapter 11: a random variable is a rule that takes an outcome and returns a number. Nothing in that sentence says the outcome has to be a single roll. Let the outcome be the entire sample — both dice, (d1, d2), one experiment — and “average them” is a perfectly good rule from an outcome to a number. The nameplate flips from VERB to NOUN and the question stops being broken. From there it is just arithmetic you can do on paper. Two dice have 36 equally likely outcomes, so beat 3 draws all thirty-six and colours each one by its mean; tap any cell and every cell sharing that mean lights up, with the count running beside it. Beat 4 sweeps that grid into the shape it was always hiding: eleven values from 1.0 to 6.0 in half-steps, with counts 1 2 3 4 5 6 5 4 3 2 1 out of 36, mean exactly 3.5, and P(x̄ = 3.5) = 6/36 = 0.166667. If it looks familiar it should — it is the identical triangle Chapter 9 drew for the sum, the same 36 cells with the axis relabelled. And beat 5 is the payoff. Take the spread straight out of those 36 counts: Σ(v−3.5)²·count = 52.5, divide by 36, take the root, and you get 1.2076147. Now take Chapter 13's route instead: one die has σ = √(35/12) = 1.7078251, and σ/√n with n = 2 gives 1.2076147. Same seven decimals, stacked digit over digit so you cannot look away from it. Notice what never appeared: no central limit theorem, no simulation, no “for large n”. The √n law is not an approximation that shows up when samples get big — it is an exact consequence of how variances add, and here it is at n = 2, verified by counting on your fingers.
There are 36 equally likely samples, so lay them out in a 6×6 grid and colour each cell by its mean. Sweep the cells into a histogram and a familiar triangle appears. It is the same triangle you met in Chapter 9 as the distribution of the sum, and it turns out to have been a sampling distribution the whole time, wearing different axis labels.
Read the numbers straight off it. There are 11 possible values, from 1.0 to 6.0 in steps of 0.5. The mean of those values is 3.5, exactly μ. And the standard deviation is 1.2076, which is where this gets good.
Check that against Chapter 13's √n law. One die has σ = √(35/12) = 1.7078. Divide by √2 and you get 1.7078 / 1.4142 = 1.2076. It matches to seven decimal places. We did not assume the √n law here and we did not approximate it. We counted 36 cases with Chapter 9's favourable-over-total, and the previous chapter's promise came true inside the answer.
Notice what this kills. Most readers assume the distribution of X̄ needs the central limit theorem, or a simulation, or some heavy machinery arriving later — and so they file it away as an approximation they will have to take on trust. It is not an approximation at all: for a small enough case it is exactly computable by counting, and you have just done it.
04The sampling distribution
That histogram has a name, and it is the centre of the entire subject. The sampling distribution is the law of a statistic across all the samples you could have drawn.
Sit with the strangeness of that for a second, because it is genuinely strange. You will only ever draw one sample. The distribution is spread over repetitions that never happen. With the dice it felt safe, because all 36 samples were visible at once on the grid. With real data, 35 of them are invisible forever.
And the miracle this whole chapter is built on is that you can describe that distribution anyway, from the single sample you actually hold. We will get there in a moment. First we have to clear up a collision that causes more wrong error bars in finance than any other single confusion.
Three different objects now wear the word "distribution". There is the population's distribution, which is a fact about the world. There is your sample's own histogram, which is the only one you will ever hold in your hands. And there is the statistic's distribution, the cloud we just met. Textbooks draw all three with the same axis label, at different scales, on different pages. Readers fuse the sample with the population, which is survivable, and then fuse the sample with the statistic, which is not. That second fusion is the error behind almost every misquoted error bar in finance.
Three histograms, one locked ruler. Drag n and watch which of them narrows — and which one flatly refuses to.
WHICH ONE NARROWS AS n GROWS?
tick one — that unlocks the n slider.
n — DRAWS IN YOUR ONE SAMPLE
n = 10 · s = 1.20%
RE-RUN THE WORLD · THEN TAKE IT AWAY
tick a prediction to unlock n
What you're looking at — three different distributions, drawn for once at the same scale, on the same ruler.
the population: every daily return the market will ever produce. Its spread is σ = 1.20%. Nothing you do moves it.
your one sample: n draws, and the tick is their mean x̄. Its bracket s hovers at 1.2% forever — more data makes it smoother, never narrower.
the cloud of x̄: where that tick would land if you re-ran the world. Its bracket is SE = σ/√n, and it is the only thing on screen that collapses. The dashed ghost is its width at 4n — exactly half.
the ruler is hard-locked across all three. That is the whole trick: textbooks draw these on three pages at three scales, which is why the two “spreads” get confused.
Fig. 7. Open any textbook at this chapter and you will meet three histograms on three different pages, drawn at three different scales, and every one of them called a “spread”. That is where minds tear. So here they are stacked on one hard-locked ruler — same axis, same pixels per percent, no rescaling allowed — and the confusion has about three seconds to live. The top panel is the population: the whole distribution of daily returns, spread σ = 1.20%, a real number for a broad equity index. You never see this panel in your life; it is drawn only so you know what the other two are aiming at. The middle panel is your one sample — n actual days, actually recorded — and the bottom panel is the cloud of x̄: not data, but the set of means you would have got had the world been re-run. Hit RESAMPLE and that is exactly what happens: the world re-runs, the middle histogram reshuffles, and one more dot drops into the pile at the bottom. The pile is the sampling distribution, and notice that you built it — nobody defined it at you. Now the whole chapter, in one drag. Push n from 5 to 2000. The middle bracket s wobbles at first and then settles at 1.2% and stays there: more data does not make the market calmer, it just draws the same market more smoothly. The bottom bracket falls off a cliff — 0.38% at n=10, 0.12% at n=100, 0.027% at n=2000 — because it is SE = σ/√n, and Chapter 11's variance rule hands you that with no new machinery: n independent draws sum to a variance of nσ², dividing the sum by n divides the variance by n², so Var(x̄) = σ²/n and the spread is its square root. The dashed ghost bracket is the width you would get at 4n, and it is exactly half — that is the exchange rate, and it is brutal: to halve your error you buy four times the data, not twice. Say the two quantities out loud so they never merge again: σ is a fact about the world, and it does not care how long you watch. SE is a fact about your knowledge, and it is the only thing your effort can buy. Then press REAL LIFE. Panels 1 and 3 vanish, because they were always imaginary — you were never handed the population, and you never get to re-run the year. All you hold is the middle panel, one finite sample, dimmed and alone. Which raises the question the next figure exists to answer: if the cloud is invisible, how do you know how wide it is? (You estimate σ from that one sample — and that estimate has a bias of its own, which is where Bessel's n−1 walks in.)
So here they are on one screen, with the axes locked to the same ruler and a slider on n. Panel one is the population, and it is wide, fixed, and completely unmoved by anything you do to the slider. Panel two is one sample. It jitters whenever you resample, and here is the part to watch: it stays about as wide as panel one no matter how large you make n. Panel three is the cloud of x̄ values. It jitters too, it sits far narrower, and it visibly collapses as n rises.
Your eye settles the whole confusion in about three seconds. Panel two's width is a fact about the world and it refuses to shrink. Panel three's width is a fact about your knowledge and it shrinks like 1/√n. Chapter 13 separated data from average with two histograms; this is that picture with the middle panel put back where it belongs.
Now take panels one and three away, and leave only panel two. That is the real situation. Always. You have one sample and no way to see the cloud your estimate came from.
I am about to draw a fresh 100 days from the identical world. Write down what its mean will be — then seal it, and watch what happens.
BEAT 1 · THE NEXT SAMPLE'S MEAN WILL BE (%)
BEAT 2 · MORE SAMPLES NEEDED TO GET THE WIDTH?
BUILT FROM THE ONE SAMPLE
① s = 1.20% the DATA's spread
② ÷ √n = √100 = 10
③ SE = 0.12% the MEAN's spread
step 1 of 10 · commit
write a number — any number
What you're looking at — one sample on top, and underneath it the axis its mean lives on, zoomed in ten times.
everything read off your one sample: its mean x̄, and the curve those two numbers alone predict — drawn before a single extra sample exists.
s = 1.20%, the spread of the DATA — how much one day differs from another. Divide it by √n and it becomes SE, the spread of the ESTIMATE. Two different quantities, one word.
500 means from samples you will never draw. In real life this heap is invisible — you get one dot from it and no more.
your sealed prediction. Not the answer, and never was: one dot among five hundred.
Fig. 8. Here is the sentence this whole part of the course stands on, and it is worth reading twice: a statistic is a function of random variables, so it is itself a random variable. You computed a mean from 100 days — x̄ = +0.043% — and that number feels like a fact. It is not. It is a draw. So before you scroll, do the thing the figure asks: write down what the mean of the next 100 days from the identical world will be, and seal it. Most readers write 0.043 or something snuggled up against it; a few refuse and say it cannot be predicted — and that refusal is already the right answer, it just has not been given a shape yet. Draw again: +0.061%. Again: −0.008%. Nothing about the world changed between those draws; only the days did. Now let 500 of them rain down and the shape arrives by construction, not by decree: a heap, centred somewhere near the truth, with a definite width. That heap has a name — the sampling distribution — and the number you had been treating as THE answer is lit up inside it as one dot among five hundred. It was never anything else.
Then comes the part that makes statistics possible at all, and it is easy to miss how strange it is. Wipe the 500 away. In real life they were never there: you have one sample, and no budget for another. So ask the question honestly — how many more samples do you need before you can say how wide that heap was? Everyone answers with a number. The answer is zero. Look at what the width actually is. By Chapter 11's variance rule, the sum of n independent draws has variance nσ², so the mean — that sum divided by n — has variance σ²/n, and its spread is σ/√n. Every quantity in that formula is either something you chose (n = 100) or something your one sample already estimates (σ, estimated by s = 1.20%). One number and one square root: SE = 1.20% / 10 = 0.12%. That is the whole of it. The curve gets drawn from a single sample, with no second sample anywhere in the building — and when the 500 real means are dropped back on top of it, they land inside.
Two warnings, because this is exactly where minds tear. First, s and SE are not the same thing wearing different clothes. s = 1.20% is a fact about the world — how much one day differs from another — and it does not shrink when you collect more data, ever. SE = 0.12% is a fact about your knowledge, and it shrinks like 1/√n. That is why the top panel and the bottom panel need a 10× zoom between them to be drawn on the same page at all. Second, read the exchange rate off the formula and let it sting: √n, not n. To halve your error bar you need four times the data; to shrink it tenfold, a hundred times. This is why an expected return takes decades to pin down while a volatility is estimable from months — and why the modest-looking curve you just drew is the same object as every confidence interval, every p-value, every t-stat and every error bar on a Sharpe ratio for the rest of this course.
Run it in two beats, and commit out loud on both. Beat one: your hundred returns average 0.043%, and I am about to draw a fresh hundred from the identical world. Write down what the new average will be. Most people write 0.043 or something near it, and some protest that it cannot be predicted — and both of those answers are the admission we need. The new average comes back 0.061%. Draw again and it is −0.008%. Let 500 fresh samples rain down as dots and pile into a heap, and the number you had been treating as the answer is one dot in that heap. It never was anything else.
Beat two is the stroke. Wipe the 500 dots off the screen and put back your one sample, which is the real situation. Now, how many more samples do you need before you can say how wide that heap was? Every honest answer is a number. A few. A hundred. As many as possible.
The answer is zero. The heap's width is σ/√n, and the one sample in front of you already estimates σ. Draw the predicted curve from that single sample — one number and one square root — and then drop the 500 real dots back on top of it. They land inside a curve that was drawn without them.
That is the moment statistics becomes possible, and it is worth being explicit about what just happened. The cloud you can never observe was drawn from the one sample you did. Every confidence interval, every p-value, every t-statistic and every error bar on a Sharpe ratio for the rest of this course is that curve, drawn that way.
So the cloud is the object, and it has exactly three properties worth asking about. Where is it centred? How wide is it? What shape is it? Those three questions are called bias, standard error and the central limit theorem, and the next three sections take them in order.
05The centre — bias, and what it does not promise
The distance from the cloud's centre to the truth is the bias: bias(θ̂) = E[θ̂] − θ. For the sample mean that distance is exactly zero, and the proof is one line.
Take the expectation of the average. E[X̄] = (1/n)·ΣE[Xi] = (1/n)·(nμ) = μ. That is linearity of expectation from Chapter 11 and nothing else. Notice what the argument never used: independence appears nowhere in it. Each of the n draws contributes μ/n to the answer, and the n copies add straight back up to μ. This is Chapter 13's equal-shares lemma standing up again in new clothes.
So the sample mean is an unbiased estimator of μ. Now let me take most of that word away from you, because it buys far less than it sounds like it buys.
Two words, four shooters — then the cloud itself, perfectly centred on the truth, with one dot in it that is badly wrong.
THE TWO STAGES
TAP A WORD, THEN TAP AN AXIS SLOT
One word is about WHERE the group sits. One is about HOW SPREAD OUT it is. Which axis is which?
PICK A DOT — CENTRE → EDGE
dot 1/240 · 0.00 SE out
INDEPENDENCE IN THE PROOF
Independence is a real assumption elsewhere in this chapter. Watch: this one proof never touches it.
tap a word, then tap a slot
What you're looking at — the same 2×2 every textbook draws, and then the cloud it was a cartoon of.
the truth — the bullseye on every board, and in stage 2 the fixed gold line at μ = 0.020%. It never moves and you never see it.
the estimates — each shot, each dot, one x̄ from one re-run of the world (n = 100 days, σ = 0.35%, so SE = σ/√n = 0.035%). The dashed blue line is where those 240 dots average out.
your dot's own error — the red bar from the gold line out to the one dot you picked. Unbiased says nothing about this bar.
the aha — blue lands on gold exactly, by linearity alone, with independence struck out of the proof. And roughly a third of those same dots are more than a whole SE from the truth. Centred is not accurate.
Fig. 9.Unbiased is the most oversold word in statistics, so let us buy it carefully. Stage one is the picture every book draws and no book makes you earn: four shooters, nine shots each, the gold bullseye the same on all four. Drag — tap — the words onto the axes and notice which way each one runs. Going down a column, the group slides off the gold as a lump: the centre has moved, and that is bias. Going across a row, the centre stays put and the shots just fly apart: that is variance. Two entirely separate complaints, which is precisely why one word cannot cover both. Stage two dissolves the cartoon into the real thing. Those are 240 re-runs of the same 100 trading days, each handing you its own x̄, piled into the sampling distribution you built by hand three figures ago. The gold line is the truth, μ = 0.020%, and the blue dashed line is where the 240 dots average out — it lands on the gold exactly, and the strip underneath shows why in one line: E[x̄] = E[(1/n)(X₁+…+Xₙ)] = (1/n)(E[X₁]+…+E[Xₙ]) = (1/n)(n·μ) = μ. Every step is Chapter 11's linearity of expectation, which holds whether or not the draws are related, so hit DELETE IT and watch independence get struck out of the assumptions while the proof stands there untouched. Now the sting. Slide the picker outward and the red bar grows: a dot at +0.115% against a truth of 0.020% is nearly six times the number it is estimating, and it came out of a cloud that is perfectly centred. Roughly a third of these dots sit more than a whole SE from the truth — and one of them is your Tuesday. That is the whole trade: unbiased is a promise about the average of re-runs you will never take, and it says nothing whatsoever about the dot you are standing on. Which is why the next thing we build is not a better centre but a width — an interval honest enough to admit where the dot might be.
"Unbiased" is a compliment in English before it is a definition in statistics, so it gets heard as accurate, correct, trustworthy. It means none of those things. It means the cloud is centred on the truth, averaged over infinitely many samples you will never take. It makes no promise whatsoever about the one estimate you are holding.
Look at the dot near the edge of the cloud in that panel. That dot is somebody's Tuesday. Their estimator was unbiased the entire time it was lying to them, and every honest quant has shipped one. An unbiased estimate can be wildly wrong; unbiasedness is a property of the procedure, not of your answer.
There is a second trap here, and it detonates years later if we leave it. Readers assume unbiasedness passes through functions. If s² is unbiased for σ², then surely s is unbiased for σ?
s² is unbiased for σ². Push every value through √ and it stops being unbiased for σ — drag n and watch the fee.
PREDICT — s² HITS σ² ON AVERAGE. DOES s HIT σ?
pick one — then watch the curve settle it.
SAMPLE SIZE n — DRAG IT
n = 5 · sd(s²) = 0.71 σ²
THE FOUR YOU MEET IN TABLES
SWAP THE FUNCTION — CHANGE THE BEND
n = 5 · E[s] is 6.00% below σ
What you're looking at — one honest cloud, sent through a bend, and the bill that comes back.
the s² cloud along the bottom — the sample variances you would get by re-running the world. Written as multiples of σ².
the truth: it balances exactly on 1, so E[s²] = σ². Straight up from there, on the curve, sits √(σ²) = σ — what s is supposed to hit.
the s cloud up the left — every one of those values pushed through √. Same information, new shape.
the fee: the chord joining any two mapped points hangs below the curve, so their average does too. The band is how far below — and it never reaches zero.
Fig. 10. Here is a sentence that sounds like it must be true and is not: s² is unbiased for σ², so s is unbiased for σ. Bessel's correction bought you the first half fairly — divide by n−1 and the cloud of sample variances balances exactly on σ², which is the green tick on the bottom axis. Now take the square root of every one of those values, which is what you do the moment you quote a volatility. Each blue dot travels up to the curve and across to the left, and the s cloud that lands there does not balance on σ. It balances lower, every time, and the red band is how much lower. The mechanism is drawn rather than asserted: pick any two of your s² values, one either side of σ². The straight chord joining their two mapped points is the average of that pair — and because √ bends downward, the chord hangs under the curve. Its midpoint sits directly below the point on the curve at the same place. Averaging can only ever land on a chord; the truth lives on the curve; the curve is above. That gap has a name and an exact size, c₄(n): at n = 2 your s is 20.21% too small on average, at n = 5 it is 6.00% too small, at n = 102.73%, and at n = 100 just 0.25%. Drag the dial and watch why it shrinks: the fee is not charged by the square root alone, it is charged by curvature times spread, and the spread of s² is σ²√(2/(n−1)) — it collapses as n grows, so a wide cloud on a bent curve pays dearly and a tight cloud barely notices. Hit x² and the sign flips: the same cloud through a curve that bends the other way overshoots instead. So it was never anything about variances or square roots in particular — it is the direction of the bend, and nothing else. Which is exactly what Chapter 15 already told you in different clothes: ln E[G] ≥ E[ln G] was the same inequality collecting the same fee, and the volatility drag you paid there is this bill under another name. Two practical consequences you will meet for years. First, the annualised volatility on your screen is a touch low if it came from a short window, which is why control charts carry c₄ tables. Second, and much more important: unbiasedness does not survive a function. If you ever need E[g(θ̂)] to equal g(θ), you must check the bend — you cannot inherit it.
It is not, and the tool that settles it is one you used two chapters ago on volatility drag. The square root is concave, so Jensen's inequality says E[√(s²)] ≤ √(E[s²]) = σ. The sample standard deviation is biased low, even though the sample variance is not.
Put numbers on it so it stops being abstract. For Normal data, E[s] sits 20.21% below σ at n = 2, 6.00% below at n = 5, and 0.25% below at n = 100. It is the curvature of the square root doing all of that work. It is the same curvature that made ln E[G] ≥ E[ln G] in Chapter 15, charging the same kind of fee for the same reason.
06The width — the standard error
The second question about the cloud is how wide it is. The standard deviation of a sampling distribution has its own name — the standard error — and Chapter 13 already handed us the formula for the sample mean.
Var(X̄) = σ²/n, so SE(X̄) = σ/√n. That is the number that answers "how much should I believe my own estimate", and it is a completely different quantity from the σ that describes the data. Chapter 13 named the words and promised a working; this is the working.
Now the most expensive confusion in applied statistics, which is a naming failure rather than a conceptual one. σ and SE are both standard deviations. Both get written as a ± after a number. Both are printed in the same units. Nothing in the words tells you which random variable each one belongs to, and that is the entire distinction.
Two numbers, both standard deviations, both printed as a ± in percent. One belongs in the note and one is wrong by a factor of √n.
PREDICT — AS n GROWS, WHICH RULER MOVES?
pick one — n then sweeps to 10,000 and back, and scores you.
SAMPLE SIZE n — DRAG IT
n = 100 · SE = 0.120%
The ± is blank. Both chips are standard deviations, in the same units. Nothing in the words says which one belongs here.
the ± is still blank
What you're looking at — two standard deviations, same units, same ± sign, and only one of them belongs in the note.
Ruler A, σ = 1.2% — how far one day's return sits from the average. A fact about the market. Drag n to 10,000 and it does not move a pixel.
Ruler B, SE = σ/√n — how far your average could sit from the truth. A fact about your knowledge. It collapses like 1/√n.
0.043% — the estimate itself, the gold dashed line both rulers are pinned to. It is what the ± decorates.
the trap — put σ after the ± and you overstate your uncertainty by exactly √n, and make "more data" look like "a calmer market".
Fig. 11. Here is the single most expensive naming collision in applied statistics, and it hides in plain sight because both sides of it are telling the truth. You ran the numbers on a hundred days of a stock and got a mean daily return of 0.043%. Your note needs a ± after it, and two numbers are sitting on your desk that both deserve the name standard deviation, both measured in percent, both perfectly correct. σ = 1.2% is the spread of the returns themselves — how far a single day tends to land from the average. SE = σ/√n = 0.12% is the spread of your average — how far the number you computed tends to land from the true mean you were trying to measure. Nothing in the words distinguishes them. Only the question they answer does. Ruler A answers how volatile is this market; ruler B answers how much do I know. So before you drop either chip, make the prediction, because it is the one that separates readers who have understood this chapter from readers who have merely read it: as you take more and more data, do both rulers narrow? Most people say yes, and it feels obviously right — more data, less spread, surely. Watch what happens. Sweep n from 10 to 10,000 and ruler B collapses from ±0.38% to ±0.012%, a shrink of thirty-one fold, while ruler A does not move one single pixel. It cannot. σ is a property of the market and you do not own it; collecting data does not make Tuesdays calmer, it only tells you more precisely what Tuesdays are like. That is the whole distinction in one animation, and the moment you feel it you will never write the wrong one again. Now drop a chip. Choose σ and the note reads 0.043% ± 1.2%, which overstates your ignorance by exactly √n — a factor of ten here, and run it backwards it becomes absurd: push n to 10,000 and that same mistake would print ±0.012% and claim, straight-faced, that more data calmed the market. Choose SE and the note is honest, and honesty is brutal: 0.043% ± 0.12% means the error bar is 2.79 times the size of the estimate it decorates. You did not measure a return. You measured a number smaller than its own uncertainty, and only decades of data would fix it — which is exactly why expected returns are famously unestimable while volatility is not. One last honesty, in the bottom panel: a ±1 SE bar is not a fence. Re-run the world thirty times and about ten of those estimates land outside it. It is a typical miss, not a bound — and turning it into an actual guarantee is what Chapter 17 is for.
Two rulers on one screen, locked to the same scale, with one slider. Ruler A is labelled "how spread out the returns are". Ruler B is labelled "how spread out my answer is". Before you drag n from 10 to 10,000, predict which one moves. A good fraction of readers expect both to shrink, and that prediction is exactly the misconception worth extracting.
Ruler A does not move a millimetre, and it never will, because it is a fact about the market. Ruler B collapses towards a point. So σ describes the world and does not shrink with data; SE describes your knowledge and shrinks like 1/√n. Only one of the two is bought with data.
Two live failures follow from mixing them up, and they run in opposite directions. Forwards: a reader with 100 daily returns of standard deviation 1.2% reports their mean return as 0.04% ± 1.2%, which is wrong by a factor of ten. Backwards: a reader believes that collecting more data makes the returns themselves less volatile, as though the market calmed down because they opened a longer spreadsheet.
Now land it in the real numbers, because the numbers are the whole argument. A hundred days of returns with σ = 1.2% a day gives SE = 1.2/10 = 0.12%. The sample mean was 0.043%. The error bar is nearly three times the size of the thing being measured. That single line is why quantitative research is hard, and it is not rhetoric. It is division.
One more honesty note before we move, because readers hear a guarantee where none was offered. A range of ±1 SE is not a bound. Under a Normal cloud it covers about 68% of the time, which means roughly one time in three your estimate lands outside it. It is a typical wobble, not a fence. So when a research note prints an estimate with a ± beside it, the first question is always which of the two numbers that is.
07The σ you do not have
Now look hard at the formula we just wrote down. SE = σ/√n. You have n, obviously; you counted your own data. You do not have σ.
Before you quote an error bar, audit it. SE = σ/√n — tap each symbol and ask the only question that matters: do I actually have this?
PREDICT — OF SE's TWO INGREDIENTS, WHICH DO YOU ALREADY OWN?
pick one — then audit the card.
AUDIT AN INGREDIENT
CANDIDATE FOR THE MISSING σ²
Σ(xᵢ − x̄)² / n
verdict slot — empty
nothing audited yet
What you're looking at — the one symbol in your error bar that nobody ever handed you.
SE — how wrong your mean is. The number you were about to quote.
n — you counted it yourself. Genuinely in hand.
σ — a fact about the population, which is the thing you were trying to learn. The arrow loops back on itself.
s — the substitute you'd reach for. Test it before you trust it.
Fig. 12. Here is the sentence that opens Part 2, and it is uncomfortable: the formula for your own uncertainty contains a number you do not have. You took forty returns, you averaged them, and you want to say how wrong that average might be — so you reach for SE = σ/√n, the exchange rate Chapter 13 handed you. Audit it symbol by symbol, which is what this card makes you do. Tap n and the tag comes back clean: you counted it yourself, it is forty, nobody can dispute it, and it turns green because it is genuinely in your possession. Tap σ and the tag comes back red, because σ is not a fact about your data at all — it is the true spread of the whole population, every return the market will ever produce, and no one has handed you that. Follow it down the supply chain and the problem becomes structural rather than annoying: the arrow that would deliver σ comes out of the population box, and the population box is the very thing your sample mean was hired to learn about. The dashed loop is a real circular dependency, drawn honestly: to say how well you know the population you must already know something about the population. This is why the next few sections exist, and it is the motivation textbooks skip when they simply assert the n−1 rule. There is only one way out — estimate the spread too, from the same forty numbers, using the natural guess in the candidate panel: Σ(xᵢ − x̄)²/n, the average squared distance from the sample's own mean. But notice the button does not say ACCEPT IT. Section 16.4 just built you an instrument for exactly this moment — an estimator is unbiased when its long-run average lands on the truth — so the honest move is to test the guess, not adopt it. Press it and the verdict slot deliberately stays blank: BIASED?, tested in the sections that follow. Hold that question open; when the answer arrives it will be the reason for the strangest-looking division in statistics. And read the second tag before you leave, because it names a debt that will be paid in Chapter 17: if σ had to be estimated, then SE was computed from an estimate, which means your error bar has an error bar of its own. That extra, unadvertised wobble is precisely why the normal distribution is not quite the right ruler for a small sample, and why a slightly fatter one — Student's t — has to exist at all.
That σ is a property of the population, and the population is the thing you are trying to learn about in the first place. So your error bar contains an unknown. Before you can quote your own uncertainty, you have to estimate the spread as well as the mean.
I want to slow down on this step even though nothing hard happens in it, because skipping it is what makes the next three sections feel arbitrary. In most courses, Bessel's n−1 arrives as an unmotivated rule in a box, next to a formula nobody had any reason to want. If you never felt the need for s, then n−1 is a convention — and a convention can only be memorised and then forgotten, never re-derived.
There is a second thing hiding here that most books never say at all. Your error bar has an error bar. You are about to substitute an estimate for σ and carry on as though nothing happened, and something did happen. That something is precisely why the t-distribution exists, and we will meet it later in this chapter rather than having it appear from nowhere in the next one.
So take the obvious candidate and use the sample's own spread: average the squared distances from x̄, dividing by n. It is the natural guess, and we now have a tool for judging it. Ask whether it is biased.
It is, and it is biased low. Not scattered low — shifted low, every time. And the reason has nothing to do with probability at all.
x̄ is the point that minimises the total squared distance to your data. That is what an average is, and one line of Chapter 3 proves it. Differentiate f(a) = Σ(xi−a)² to get −2Σ(xi−a), set that to zero, and a = x̄ falls straight out. The second derivative is 2n > 0, so it is a genuine minimum and not just a stationary point.
So squared distances measured from x̄ are smaller than squared distances from any other point. Including from the true μ. Always, in every sample.
Two data points, one gold μ you can put anywhere. Blue measures scatter from the sample mean x̄, gold from μ. Hunt for a drag that makes blue longer.
CAN'T FIND ONE? LET IT HUNT FOR YOU
the blue dots and the gold ▲ in the picture are draggable — arrow keys work too.
SHOW THE LANDSCAPE — f(a) FOR EVERY a
PROVE IT — WHY a = x̄ IS THE MINIMUM
BLUE < GOLD · gap 16.8200
What you're looking at — the same two points measured twice, and a race that is fixed before it starts.
blue — each point's distance to x̄, the average of the two dots. Squared and added, that is Σ(x−x̄)²: the naive variance's numerator.
gold — the same two distances, but to μ, the true centre you never actually know. Drag μ anywhere; the gold bar is never the shorter one.
the gap is exactly n(μ−x̄)² — a square, so it cannot be negative. That single fact is the whole proof: gold = blue + something ≥ 0.
ohh — THAT'S why: x̄ is defined as the centre that minimises this total. Measuring scatter from it is grading your own homework, so it comes out too small every time, not on average. Bessel's n−1 is the refund.
Fig. 13. Here is the cleanest way to see why the naive variance is broken, and it needs no expectations at all. On the line are two data points you can drag anywhere, and a gold marker for μ, the true centre — the number the world used when it generated the data, and the number you will never actually be told. x̄ (say "x-bar") is just the average of your two dots, so it moves when they move. The figure measures the same scatter twice: once from x̄ and once from μ. Square each distance, add them up, and you get the two bars: Σ(x−x̄)² in blue, Σ(x−μ)² in gold, both printed to four decimals so there is nowhere to hide. Now the invitation: break it. Drag the dots, drag μ, or press random ×25 and let the widget try configurations faster than you can. The counter climbs. The blue bar shrinks, stretches, sometimes nearly catches gold — and never, not once, passes it. Put μ exactly on x̄ and the two totals become equal, which is the closest you will ever get; a hair either side and gold is ahead again. That is not a tendency you would have to average over many samples to notice. It is an inequality, true in the single sample sitting in front of you. Press SHOW THE LANDSCAPE and the reason stops being mysterious. The curve is f(a) = Σ(x−a)²: for every candidate centre a along the line, how much total squared distance that choice costs you. It is a parabola opening upwards, and the whole comparison collapses into reading two heights off one curve — x̄ sits at the bottom of the valley, μ sits somewhere on the wall. Nothing sits below the bottom. That is the entire argument. PROVE IT gives it in one line of Chapter 3's calculus: f′(a) = −2Σ(x−a), set that to zero and you get a = x̄; f″(a) = 2n > 0 confirms it is a genuine minimum and not merely a flat spot. Better still, expand the algebra and the curve has an exact shape: f(a) = f(x̄) + n(a−x̄)². The gold total is the blue total plus a square — and squares are never negative. So the punchline lands where it should. x̄ is, by construction, the point closest to your own data. Measuring your data's spread around it is grading your own homework: the marker and the exam were written by the same hand, so the score comes back flattering, in every sample, by exactly n(μ−x̄)². Divide that too-small total by n and your variance estimate is biased low, systematically, forever — which is precisely the debt Bessel's n−1 repays, and precisely what "you spent one degree of freedom estimating the centre" means in pictures. Keep this parabola in mind: in Chapter 20 regression does exactly this hunt for the lowest point of a squared-error surface, only in many dimensions at once, and the same "fitted values hug their own data" flattery reappears there wearing a different name.
Do not take that on trust — hunt for a counterexample. Two points on a line, a marker for the true μ, and two bars drawn live: the sum of squared distances to x̄ in one colour, the sum to μ in another. Drag the points anywhere you like, through as many configurations as you have patience for, and try to find a case where the x̄ bar is longer. You cannot. The two bars touch only in the instant x̄ lands exactly on μ.
That is bias seen as an inequality, with your own hands, before any algebra. And an inequality is a much stronger fact than a tendency. Readers who accept there is a bias usually assume it is bad luck that could go either way and would average out. It cannot go the other way.
Here is the sentence that makes it permanent. x̄ is the point that best fits your data, so measuring the scatter around it is grading your own homework. And flag the forward link now, because it is the same fact wearing a different hat: "the point minimising squared distance" is exactly what OLS computes in Chapter 20, one dimension at a time.
08Exactly one observation's worth
How much too small, exactly? Predict before you read on, because the surprise lives in the size rather than in the fact. Is the shortfall a fixed fraction of the variance? Is it proportional to n? Does it fade away as the data piles up?
Most people guess proportional. The answer is a constant.
Four lines and one trick — watch the shortfall land on exactly one σ², the same size whether n is 2 or 1000.
STEP 0 — SEAL A CALL. HOW BIG IS THE SHORTFALL?
Three ways the gap could behave. Pick the one you believe — it is sealed until step 4.
WALK THE DERIVATION
0/4 · seal a call
n — THE SIZE DIAL (lands at step 4)
n = 10 · gap = 1σ² · 10.0%
pick one, then walk the 4 steps
What you're looking at — four lines, one trick, and a gap that never changes size.
Σ(Xᵢ−X̄)² — the spread you can actually compute, measured from the sample's own centre X̄.
the free steps — Σ(Xᵢ−X̄) = 0 kills the cross term; Ch 11 and Ch 13 hand you both expectations.
the landing — (n−1)σ², which is why you divide by n−1 to get back to σ².
the missing piece — exactly one σ², identical at n=2 and n=1000. Only its share changes.
Fig. 14.Where n−1 comes from — and why the missing piece is exactly one observation's worth of variance. Before a single line is written, seal a call: when you measure spread around the sample's own centre X̄ instead of the true centre μ, you come up short of nσ² — but short by how much? A shrinking amount that fades away as the sample grows, a growing amount proportional to n, or the same amount every time? Most readers reach for "it fades", because the share does fade and the two facts are easy to merge. Now walk the four lines. Step 1 is the only trick in the entire derivation: write Xᵢ − μ = (Xᵢ − X̄) + (X̄ − μ) — slide the sample centre in, so every deviation from the truth is split into a deviation from the sample mean plus the drift of that mean. Square it and sum, and three terms appear. Step 2 kills the middle one for free: the cross term carries the factor Σ(Xᵢ−X̄), and deviations measured about their own mean always sum to zero — strike it, watch it fall off the line, and what is left rearranges with no work at all into Σ(Xᵢ−X̄)² = Σ(Xᵢ−μ)² − n(X̄−μ)²: the spread you can measure equals the spread you cannot, minus how far the centre drifted. Step 3 takes expectations term by term, and it deliberately stops on the second one, because that is where readers stall. The first is easy — n copies of E[(Xᵢ−μ)²] = Var(Xᵢ) = σ², straight from Chapter 11's definition, giving nσ². The second looks like new work, and it is not: since E[X̄] = μ, the expression E[(X̄−μ)²]isE[(X̄ − E[X̄])²], which is the definition of variance wearing different clothes — it is Var(X̄), and Chapter 13 already told you that equals σ²/n. Flip the two chips in the spotlight and watch the same quantity change costume. So the second term contributes n · σ²/n = σ² — the n cancels, which is the whole reason the answer is clean. Step 4 lands: nσ² − σ² = (n−1)σ², so dividing by n−1 rather than n is not a fudge, it is the exact repair. Then the size dial separates the two facts you must never merge. In the top strip each cell is one σ², and the red cell — the missing one — is the same width at n=2 and at n=1000; only the blue cells multiply. In the bottom bar the total is held fixed, and the red slice shrinks: 50% of the truth missing at n=2, 10% at n=10, 0.1% at n=1000. The gap never changes; your exposure to it does. And notice what was never used anywhere in those four lines: the shape of the population. No bell curve, no symmetry, no independence beyond a finite variance — which is why Bessel's correction is not a normal-theory result but a fact about arithmetic, and why the one degree of freedom you spent estimating the centre costs exactly one observation's worth of variance, every single time.
The derivation is three lines and each one earns its place. Start by writing the deviation from the sample mean in terms of the deviation from the true mean: Xi − X̄ = (Xi − μ) − (X̄ − μ). That is add-and-subtract μ, and it is the only trick in the whole thing.
Square it and sum over i. The cross term is −2(X̄−μ)·Σ(Xi−X̄), and Σ(Xi−X̄) = 0 by the definition of X̄, so that term dies for free. What is left is clean: Σ(Xi−X̄)² = Σ(Xi−μ)² − n(X̄−μ)².
Now take expectations of both sides, and stop on the second term, because this is where every honest attempt at this derivation stalls. The first term is easy: E[Σ(Xi−μ)²] = nσ², straight from the definition of variance. The second term is n·E[(X̄−μ)²], and readers stare at that as if it were a new quantity to compute.
It is not new. Look at it once more: E[(X̄−μ)²]is a variance — it is Var(X̄) written in its long form, because μ is the mean of X̄. And Chapter 13 already told us that Var(X̄) = σ²/n. So the second term is n·(σ²/n) = σ².
Put the two together: E[Σ(Xi−X̄)²] = nσ² − σ² = (n−1)σ². The shortfall is exactly one σ². Not a fraction of the data, not something that fades with n — exactly one observation's worth of variance, stolen by the act of locating the mean. Divide by n−1 instead of n and it is restored, exactly. That is Bessel's correction, and you just derived it. Notice that nothing in those three lines used the shape of the population, so the result holds for dice, for returns, and for anything with a finite variance.
The constant gap and its shrinking importance are two separate facts, so let us keep them separate. The missing σ² is the same size at every n. What changes is its share of the total. At n = 2 the naive rule returns half the true variance, so the correction is +100%. At n = 10 it returns 90% and the correction is +11.1%. At n = 1000 it returns 99.9% and the correction is +0.1%. Same missing piece throughout. It just matters less when there is more of everything else.
Bessel, verified by counting — the same 36 two-dice samples, read again for spread: ÷n averages 1.4583, ÷(n−1) averages 35/12 exactly.
THE AUDIT — TAP A BEAT
DIVIDE BY
tap a bar or a cell to light it
PREDICT — WITH THE FAIR RULE, HOW MANY OF THE 36 LAND BELOW σ²?
six values across 36 cells
What you're looking at — the same 36 samples you already counted, read a second time for spread.
the 36 cells — column = die a, row = die b. Each is stamped with Σ(xᵢ−x̄)², which at n = 2 is just (a−b)²/2.
σ² = 35/12 = 2.9167 — the true spread of one die, and the line every estimate is judged against.
the average of all 36 — the flat bar holding the same area as the jagged ones.
below the truth — 30 of 36 with ÷n, and still 24 of 36 with the fair rule.
Fig. 15. Chapter 13 used this grid to build the sampling distribution of the mean. Here is the same grid, the same 36 equally likely samples, read a second time — this time for spread. Two dice, so n = 2, and every cell carries the one quantity every variance estimate is built from: Σ(xᵢ − x̄)², the total squared distance of a sample's two values from its own mean. At n = 2 that collapses to (a−b)²/2, so only six numbers can ever appear — 0 in 6 cells, 0.5 in 10, 2 in 8, 4.5 in 6, 8 in 4, 12.5 in 2 — and the bars on the right are those 36 cells laid side by side, each one unit wide, so the total area is the total of all 36 estimates and the flat cyan bar of equal area is their average. Now flip the divisor and watch what nobody's algebra ever makes you feel. Dividing by n gives 0, 0.25, 1, 2.25, 4, 6.25 and an average of 52.5/36 = 1.4583: not a whisker low, not a rounding artefact — exactly half of the truth, which is precisely what the factor (n−1)/n = 1/2 predicts at n = 2. Dividing by n−1 = 1 does nothing to the sums at all, and those untouched sums average 105/36 = 2.9167 = 35/12, which is σ² for a fair die, to the last digit. That is Bessel verified rather than believed: no simulation, no convergence, no "approaches" — 36 cells, one division, an exact match. Then the beat that is the whole reason this figure exists rather than the derivation alone. Turn on the honesty pass and count the red. Under the naive rule, 30 of the 36 samples come in below the truth. Under the correct, unbiased rule, 24 of the 36 still come in below — two thirds of the time your fair estimator hands you a number that is too small. Both facts sit on the same picture and neither is a contradiction: two thirds of the width lies under the gold line, and yet the thin, tall bars — the six samples that rolled a 5-gap or a 6-gap — carry enough area to haul the average back up onto it exactly. That is what unbiased means, and it is all it means: the average over infinitely many re-runs is right. It says nothing whatsoever about the run you actually got.
And because we already have a case small enough to count, let us not believe any of this on authority. Go back to the two dice and read those same 36 samples for spread instead of for centre. The naive rule takes the values 0, 0.25, 1, 2.25, 4 and 6.25 across those 36 cells, and its average is 1.4583. The true variance of one die is 35/12 = 2.9167. The naive rule returns exactly half of it, which is what (n−1)/n predicts at n = 2. Divide the same sums by n−1 = 1 instead and the average becomes 2.9167 — the true variance, to the last digit, by counting.
One honesty note that this counted panel makes unavoidable, and it ties straight back to what bias does not promise. The naive rule lands below the truth in 30 of the 36 samples. The unbiased rule still lands below the truth in 24 of the 36, which is two thirds of the time. Unbiased means the average is right. It has never meant that the typical sample is right.
09Degrees of freedom, defined
Now the phrase everyone repeats can finally be defined, and it deserves the attention because you will meet it again as n−k in regression, as the parameter of the t and chi-square curves, and as a number every software package prints without explanation.
A degree of freedom is a freedom of the residuals. Once x̄ is computed from the data, the n residuals xi − x̄ are no longer n free numbers. They must sum to zero, so tell me any n−1 of them and the last one is forced.
So a sample of n carries n−1 independent pieces of information about spread, and dividing the squared total by n−1 is dividing by how many free numbers actually went into it. Most readers can happily say "we used up a degree of freedom estimating the mean" and, if pressed for what a degree of freedom is, have nothing. That is a bad state to be in for something this reusable.
A degree of freedom is a freedom of the residuals — drag one bar and watch another move by itself, because once x̄ is fixed the bars are no longer allowed to disagree.
n — HOW MANY OBSERVATIONS
n = 5df = 4
VIEW
One observation. What does dividing by n−1 give?
drag a bar — the red one follows
What you're looking at — n numbers, one rule, and the last number is no longer yours to choose.
the mean line x̄ and the SUM gauge — the constraint. Compute x̄ from the data and the residuals are pinned: they must add to 0.0000, always.
a residual ri = xi − x̄, the gap from a point to the line — still free, you may drag it anywhere.
set — a freedom you have spent. In the geometry view the green sheet is the same idea: every allowed r lives in a 2-D plane.
forced — not data any more, arithmetic. It refuses the drag. That is why df = n − 1, and why n = 1 gives 0/0.
Fig. 16. Everyone can recite it — divide by n − 1, not n — and almost nobody can say what the missing 1 is. Here it is, and you can feel it with your finger. A residual is just the gap from one observation to the mean, ri = xi − x̄, drawn as a bar hanging off the mean line. Now notice what computing x̄ did to you: x̄ is the balance point, so by construction the residuals add to zero, and the SUM gauge never moves off 0.0000 no matter how hard you pull. Drag a bar and you are not moving one number, you are moving two — the red bar repays every unit you take, instantly, because it has no choice. Spend the freedoms one at a time and watch the counter run down: with n = 5 you get four genuinely free choices, and when the fourth goes green the fifth is no longer data, it is arithmetic. Try to drag it and it tells you so. That is a degree of freedom — not a mystical correction, just a count of how many of your numbers were still allowed to vary once you used the sample to locate its own centre. The GEOMETRY view says the same sentence in the language you will need for regression: stack the three residuals into one vector, and the rule r1+r2+r3 = 0 is exactly the statement that this vector is perpendicular to the all-ones direction(1,1,1) — the dot product prints 0, so the vector cannot leave the 2-D plane drawn under it. Three numbers, one constraint, two dimensions of wiggle: sample variance is a squared length divided by the number of directions that length was actually free to live in. And then slide n down to 1, which is where the two formulas stop agreeing and start telling on themselves. One point sits exactly on its own mean, so its residual is forced to be 0 — not small, zero, with no information in it at all. Divide that by n and you get a variance of 0.0000: a confident, fluent, completely fabricated claim that returns never move. Divide by n − 1 and you get 0/0 — undefined, which looks like a failure and is in fact the only honest answer available, because a single observation genuinely tells you nothing about variability. Bessel's correction is not a fudge factor bolted on to nudge a number upward. It is the formula agreeing to count only the freedoms you still had. Fit k quantities from the data instead of one and you will pay k: df = n − k, and every regression you meet in Part 5 is that same arithmetic with a bigger bill.
So pull on the constraint until you can feel it. Draw n dots with residual bars hanging off the mean line, and drag any bar you like. As you drag, another bar moves by itself to keep the total at zero. Drag n−1 of them into any configuration you please, and the last bar is no longer draggable. It is being held in place by the others.
Then take it to n = 1, which settles the whole thing. One data point, one residual, and that residual is forced to be exactly 0. Bessel's formula returns 0/0 — undefined — and that is the correct answer. One observation genuinely tells you nothing about variability, and the honest formula refuses to answer. Divide by n instead and it reports "variance = 0", which is a confident lie about a question that has no answer.
There is a geometric version of this, and it is worth having because Chapter 20 will need exactly this picture. Stack the residuals into a vector in n dimensions. The sum-to-zero constraint says that vector is perpendicular to the all-ones direction (1,1,…,1), using Chapter 6's dot product. So it is confined to an (n−1)-dimensional subspace, and its squared length is built from n−1 independent directions rather than n. That is what the number counts. And it is why a regression that fits k coefficients will divide by n−k.
10The shape — and the two places it lies
Centre and width are settled, so the last property of the cloud is its shape. Chapter 13 proved that an average of iid draws is driven towards the Normal. Read on this side of the course, that says something much stronger than it did there.
It says the sampling distribution of X̄ has a known shape. And a known shape is what converts a distance into a probability. "My estimate is 1.8 standard errors from that value" becomes a number you can act on, which is the entire bridge into hypothesis testing. In Chapter 13 the CLT described a mathematical limit; here it is a licence to attach probabilities to your own error, and that licence rests on its hypotheses being true of your data.
Three things need measuring rather than arguing about, so this one is a real run.
CodeRun — one Normal population, 200,000 samples per n, seed 20260813. Standardise with the true σ, then with s, then take the maximum instead of the mean.
WHAT GETS STANDARDISEDtap
SAMPLE SIZE ndrag →
PREDICT FIRSTpick one
With the TRUE σ: how much mass lands beyond ±1.96?
Drag n. The blue bars never leave the violet curve — and the red tail stays at 5%.
TAIL 0.0504 · TARGET 0.0500
What you're looking at — the same population and the same seed three times over; only the recipe for the number changes. A known shape is what converts a distance into a probability.
the standard Normal — the known shape, drawn dashed in all three panels so you can see what is being compared to what.
σ known: z lies ON the curve at n = 5 as at n = 100; tail 0.0504 vs the 0.0500 promised.
s plugged in: s is itself a random draw, so the ratio wobbles twice — lower peak, fatter tails. That curve has a name: t, with n−1 degrees of freedom. Cutting a true 5% at n = 5 needs 2.7764, not 1.9600 — and "use n > 30" is just this table, rounded.
where it breaks: the mass beyond ±1.96, and the sample MAXIMUM — skew RISES 0.3043→0.8675 toward the Gumbel limit 1.1395, because the CLT is a theorem about averages. Fine print from Ch 13: the middle converges first and the tails last, and every line here assumes independent draws.
Fig. 17. The same Normal population, standardised three ways: with the true σ the histogram sits exactly on the standard Normal at every n and 1.96 really does cut 5% — but plug in s and the tail fattens to 0.1216 at n = 5 (that curve is the t, and it needs 2.7764), while the sample maximum, which is not an average at all, grows more skewed as n grows.
Panel one standardizes with the trueσ. The histogram lies on the standard Normal at n = 5 and at n = 100 alike, and the fraction of runs beyond ±1.96 comes out 0.0504 and 0.0502 against the 0.05 it should be. Panel two standardizes with s from the same samples. At n = 100 you can barely tell the difference: the tail fraction is 0.0535. At n = 5 it is 0.1216, which is more than double what it should be.
That is the debt from the plug-in step coming due, and it now has a name. Once you divide by s instead of σ, the standardized quantity is no longer Normal, because the denominator is random too — and sometimes it is too small, which throws the ratio out into the tail. That fatter curve is the t-distribution, and Chapter 17 does the working. The extra weight in its tails is the exact price of not knowing σ, and it is a price that falls away as the sample grows. Put concretely: at n = 5 you need 2.776 standard errors for 95% coverage, not 1.960. The folklore rule "use n > 30" is standing in for this whole story, and at n = 30 the honest number is 2.045.
Panel three is the honest counterexample, and it matters more than it looks. Having learned that sampling distributions are Normal, readers apply that to every statistic — medians, maxima, variances, correlations, Sharpe ratios. The CLT is a theorem about averages.
So take the sampling distribution of the sample maximum, drawn from a perfectly Normal population, and measure its skewness as n grows. A Normal has skewness 0.0000. The maximum gives +0.30 at n = 5, then +0.55 at 30, +0.66 at 100, +0.79 at 1000, and +0.87 at 10,000. It is not converging to Normal. It is walking away from Normal, toward a different limit entirely, and more data makes it worse. That is exactly why Chapter 32 needs a separate theory for tails.
And keep the fine print you already met in Chapter 13. The CLT converges in the middle first and in the tails last, which is a problem precisely because a risk number lives in the tail. Everything in this section also assumed independent draws, which is the hypothesis this entire chapter rests on.
11What you actually pay
One last question about estimators, and it is the one that stops "unbiased" from hardening into a superstition. Two errors matter, not one: how far the cloud's centre sits from the truth, and how wide the cloud is.
Put them together and you get the mean squared error: E[(θ̂−θ)²] = Var(θ̂) + bias². The derivation is the same add-and-subtract move as before, this time around E[θ̂], with the cross term dying for the same reason. Bias and variance are measured in the same squared units here, which is exactly what lets you add them and get one score. And MSE is the only thing you actually pay.
So here is a rigged contest. I offer two estimators of the mean daily return: the sample mean, and the sample mean multiplied by 0.9. Which is better? Everyone picks the sample mean, because the other one is obviously biased.
A rigged contest. Two estimators of the same mean — one honest, one deliberately shrunk by 10%. Call the winner before you press SCORE.
YOUR CALLwhich wins?
SHRINKAGE DIALc = 0.90
HONESTY CARD
c* is computed from the true μ — the one number nobody has. So it is a ceiling, not a recipe: it says how much is there to win, not how to win it. Closing that gap is what Ch 19 and Ch 22 are for.
tap a card — which one wins?
What you’re looking at — the only score anyone actually pays: MSE = variance + bias².
variance — how far the estimate scatters run to run. Shrink by c and it scales by c²: it collapses fast.
bias² — how far its centre sits from the truth: (1−c)²μ². μ is tiny, so this stays a sliver.
MSE — their sum, the total error you pay. At c = 0.9 it is 81.0% of the honest estimator’s.
three words readers use as synonyms, and shouldn’t:unbiased = centred on the truth, at this n. consistent = the whole cloud collapses as n grows. efficient = the tighter of two unbiased clouds.
Fig. 18. Here is a contest you are meant to call wrong. Two people are trying to estimate the same thing — the true average daily return μ of a stock — from the same 100 days of data. The first uses the sample mean, x̄: add the hundred numbers up, divide by a hundred. It is the honest estimator, and it is unbiased, which means that if you re-ran the world forever its answers would centre exactly on the truth. The second person does something that looks like vandalism: they compute the same sample mean and then multiply it by 0.9, deliberately dragging it toward zero. That estimator is biased — on average it lands 10% short. Pick your winner before you press SCORE, because almost everybody picks the honest one, and on this data almost everybody is wrong. The reason is that “unbiased” was never the score anyone pays. The score is mean squared error — the average of (estimate − truth)² — and it splits cleanly into two pieces: MSE = variance + bias². Variance is how far your answer scatters when the world is re-run; bias is how far its centre sits from the target. The sample mean has bias 0 and variance σ²/n = 1.44e−6, so its MSE is its variance and its typical error is 0.1200% — six times the size of the 0.020% it is trying to measure, which is the real scandal of this chapter. Multiply an estimate by 0.9 and, by the scaling rule from Chapter 11, its variance is multiplied by 0.9² = 0.81 — a 19% cut — while the bias you take on is only 0.1μ = 0.002%, and it is the square of that which enters the score: 4e−10, a rounding error next to 1.44e−6. Total: 1.1668e−6, or 81.0% of the honest estimator’s MSE. You bought a large cut in variance for a negligible bill in bias². Now drag the dial and watch the two layers fight: the blue variance layer falls like c², the red bias² layer climbs like (1−c)², and the gold line on top is the sum you actually pay. The magnifier exists because on a linear scale the bowl’s floor is invisible: the true minimum sits at c* = μ²/(μ² + σ²/n) = 0.027, with a typical error of 0.0197% — and yes, that says the best pure-MSE estimator of this mean is almost the number zero, because with 100 days of data the noise so overwhelms the signal that guessing “no drift” is nearly optimal. Which is exactly why the honesty card is pinned there and cannot be dismissed: c* is computed from the true μ, the one quantity nobody has, so it is a ceiling — it tells you how much there is to win, not how to win it. Estimating a shrinkage factor from the data itself, without cheating, is the whole business of Chapter 19’s intervals and Chapter 22’s bias–variance tradeoff, and it is the seed of shrinkage estimators, ridge regression and every Bayesian prior you will meet later — none of which are cheating, all of which are just paying a little bias² to buy a lot of variance. Finally, three words this chapter refuses to let you blur: unbiased is a statement about the centre of the cloud at a fixed n; consistent is a statement about the whole cloud collapsing onto the truth as n grows; efficient is a comparison — the tighter of two unbiased clouds. An estimator can be biased and consistent, unbiased and useless, and, as you have just watched, biased and better.
Score them both on our running sample — 100 days, σ = 1.2%, true μ = 0.020%. The sample mean has zero bias and a variance of σ²/n, giving MSE = 1.44×10⁻⁶, so its typical error is 0.120%. The shrunken one carries a little bias and a lot less variance, giving MSE = 1.167×10⁻⁶ and a typical error of 0.108%. The biased estimator carries 81% of the unbiased one's MSE. It wins.
Turn the shrinkage dial and you can watch the trade directly. As the multiplier c falls from 1, the squared bias climbs from zero while the variance falls like c², so the total dips below the unbiased point before turning back up. That dip is the same picture Chapter 22 will call the bias-variance tradeoff, and it is why regularization, shrinkage and Bayesian priors will read as sensible engineering rather than as cheating.
The honest limit of the demonstration, stated plainly: the best multiplier here is c* = μ²/(μ² + σ²/n) = 0.027, which would cut the typical error to 0.020%. But that formula uses the true μ, which nobody has. It is a ceiling, not a recipe, and closing the gap between the two is what Chapters 19 and 22 are for.
Two more words finish the vocabulary, and they are routinely used as synonyms when they are not. Consistency means the whole cloud collapses onto θ as n grows, which is Chapter 13's law of large numbers seen from this side. Efficiency means that among unbiased estimators, yours has the narrower cloud. Unbiasedness is about the centre at a fixed n. An estimator can have either property without the other.
Now cash the whole chapter out where it decides real money. Estimating a mean return and estimating a volatility are not equally hard. They are not even close, and the reason is one line of Chapter 5.
Sample the same year 93× faster and the mean's error bar does not move — because n is nowhere in σ/√T.
STEP 0 — SEAL IT. SHARPE 1.0: HOW MANY YEARS?
Sharpe 1.0 = one σ of edge per year. How long until you'd believe it?
WALK THE FOUR STEPS
0/4 · seal a number
THE TWO DIALS — BAR SIZE (STEP 2), YEARS (STEP 3)
daily · n = 252 · SE 0.2029
pick a horizon, then walk it
What you're looking at — two error bars that are told apart by one missing symbol.
SE(μ̂) = σ/√T — the mean's error bar. T is elapsed calendar years. Sampling faster cannot touch it.
SE(σ̂) ≈ σ/√(2n) — the vol's error bar. n IS in it, so it collapses 0.0089 → 0.0009.
the band — the true 8.00%/yr drift ±1 SE. A full century still leaves ±2.00% a year.
what is NOT in the formula — no n, no bar frequency. That absence is the whole lesson.
Fig. 19.μ takes decades, σ takes weeks — and t = SR·√T. Before anything runs, seal a number: a strategy shows a Sharpe ratio of 1.0 — its average annual return equals exactly one annual σ of its own noise — so how many years of track record before you would believe the edge is real? Most readers reach for one year, because a year of daily data feels like 252 pieces of evidence. Step 1 is two lines and it quietly demolishes that feeling. Over T years the log returns add, so by Chapter 11 the T-year total has mean μT and variance σ²T; the annual mean is that total divided by T, so its standard error is √(σ²T)/T = σ√T/T = σ/√T. Now click the cyan box and read the answer backwards: there is no n in that expression. No bar frequency, no Δt, nothing whatsoever about how often you looked — only σ and T measured in calendar years. Chopping the same year into finer pieces gives you more numbers, but each one is correspondingly smaller, and the two effects cancel exactly. Step 2 is the receipt, printed from a real run (seed 20260813, true drift 8.00%/yr, true vol 20.00%/yr): the same single year sampled daily (252 bars, σ per observation 0.012599), hourly (1,764, 0.004762) and by the minute (23,400, 0.001307). Flip between them. The blue bar — the error on the mean — sits at 0.2029, 0.1984, 0.2023: it refuses to move while n rises ninety-three-fold. The gold bar — the error on the volatility — falls 0.0089 → 0.0033 → 0.0009, an order of magnitude, because the vol's error does carry n: it is roughly σ/√(2n). That single contrast is the chapter's finance grounding in one picture. Step 3 runs the other axis instead: hold the bars at daily and stretch the calendar. At T = 1, 4, 25 and 100 years the measured SE comes out 0.1996, 0.0988, 0.0407 and 0.0200, against a theory of 0.2000, 0.1000, 0.0400, 0.0200 — textbook √-law, and 25× the time buying exactly 5× the precision. Watch the cyan band, which is what you actually know about the true 8.00%/yr drift at ±1 SE. After one year it spans −12% to +28% and you cannot even sign the drift; after four it still swallows zero; only past twenty-five years does zero fall outside; and after a full century the answer is 6.0% to 10.0% — ±2.00% a year, on a quantity that is about 8%. Nobody alive has a century of any one strategy. Step 4 cashes it out and unseals your guess. The t-statistic is t = μ̂/SE(μ̂) = μ̂/(σ/√T) = (μ̂/σ)·√T = SR·√T: the σ cancels, and the question "is this edge real?" collapses into a question about elapsed time and nothing else. Demanding t = 2, a Sharpe of 1.0 needs 4.0 years, a Sharpe of 0.5 needs 16.0, and a Sharpe of 0.3 needs 44.4 — longer than most funds exist. Two honest footnotes close it. First, the volatility side deserves its own care: the relative error of σ̂ is about 1/√(2n) for Normal data — 15.43% from 21 days, 4.45% from 252 — and deriving that properly needs fourth moments, which are owed to Chapter 29. Second, the two conventions you have been using all along now close their own loop: variances add over time, which is exactly why annualising a daily σ multiplies by √252 rather than by 252; and finance quotes the standard deviation rather than the variance because a percent is something a human can feel, while a percent-squared is not.
Before the numbers, commit to an answer. A strategy shows a Sharpe ratio of 1.0, meaning its annual return is one annual standard deviation. How many years of track record would you want before you believed it was not luck? Guesses land anywhere from a few months to a lifetime.
Now derive it, and start from the one fact that makes it work. Over T years, log returns add, so the total has mean μT and variance σ²T. The estimated annual mean is that total divided by T, so SE(μ̂) = σ√T / T = σ/√T. Look at what is missing from that expression. There is no n in it. Sampling frequency does not appear at all.
The run confirms it and the table is worth staring at. One year of data sampled daily, hourly and by the minute gives standard errors on the annual mean of 0.2029, 0.1984 and 0.2023 — all of them σ/√T = 0.20, unmoved. Over the same three rows the standard error of the estimated volatility collapses from 0.0089 to 0.0033 to 0.0009. Sampling faster does nothing for the mean and almost everything for the vol.
That asymmetry is the punchline, and it is genuinely surprising the first time. The precision of a mean is set by the calendar. The precision of a volatility is set by the count. So μ takes decades and σ takes weeks — one year of daily data pins a 20% volatility to about ±0.9 percentage points, while pinning the mean to ±2% a year takes a century.
That century is not a figure of speech. With σ ≈ 20% a year for equities, a hundred years of data gives SE(μ̂) = 20/10 = 2% a year, and that is the best anyone will ever do with that dataset. Read against an expected return of roughly 8%, a 95% interval spans about four points either side. People hold "we have a hundred years of data" as though it settles things. It settles the risk. It barely touches the return.
And now the sentence a quant lives by falls out in one more step. The t-statistic is the estimate over its own standard error, so t = μ̂/SE = (μ/σ)·√T = SR·√T. A Sharpe of 1.0 needs four years to reach two standard errors. A Sharpe of 0.5 needs sixteen. A Sharpe of 0.3 needs forty-four. That is how long an edge must survive before anyone should believe in it, and it is why so many track records prove nothing at all.
Two conventions you have been using now have reasons instead of customs. Variances add over time, so annualizing a daily volatility means multiplying the variance by 252 and taking the root — that is Chapter 13's √252, and it is not a fudge factor. And finance quotes σ rather than variance because σ is in percent, which a human can feel, while variance is in percent-squared, which nobody can.
One honesty note on the volatility side, because I would rather flag a debt than over-claim. The relative error of an estimated σ is about 1/√(2n) for Normal data, which is where the ±0.9 points came from. Deriving that properly needs fourth moments, and those arrive with kurtosis in Chapter 29. Take the number now and the derivation then.
Where Chapter 16 sits: one object, the five places it gets spent — and the two links it cannot hold.
WHERE IT GETS SPENTtap
FIVE THINGS YOU OWNdrag →
1 of 5 · estimand vs estimate
One object sits under everything downstream: the distribution of your estimate. Tap a destination to see which of its three properties that chapter spends.
one object · five places it is spent
What you’re looking at — one object, everything downstream, and the two assumptions holding it up.
The sampling distribution — not the data, but the pile your estimate makes when you re-run the world. Chapter 16 built it by stacking sample means.
Its three properties, each with a live number: centre (x̄ = 0.043%), width (SE = σ/√n = 1.20%/10 = 0.12%), shape (≈ normal, by the CLT).
Where it gets spent. Every later idea is one more sentence about this object — a tail of it (Ch 17, 19), its width for a slope (Ch 20), it built by brute force (bootstrap), its width on a Sharpe (Ch 30).
The two snapped links. Press THE CRACK: every result here assumed independent draws (Ch 21 — autocorrelation inflates the true SE by 1.72×) and a finite variance (Ch 32 — fat tails).
Fig. 20. This is the closing figure of the chapter, so it is a map rather than an argument — and the thing in the middle of it is the only object the chapter ever built. Not the data. The sampling distribution: the pile your estimate makes when you re-run the world, which you constructed by hand earlier on this page by drawing a sample, marking its mean on an axis, and doing it again until the dots became a shape that was never in the data at all. It has exactly three properties, tagged here in gold. Its centre, which is where your estimate sits on average — here x̄ = 0.043% a day, from a hundred days of returns. Its width, which is the standard error, SE = σ/√n = 1.20%/√100 = 0.12% — and notice that this is a fact about your knowledge, not about the world: σ = 1.20% is how variable the market is and it does not move when you collect more days, while 0.12% is how wrong your number probably is, and it shrinks only as √n. Four times the data to halve the error; there is no better exchange rate on offer, ever. And its shape, which the central limit theorem hands you as approximately normal almost regardless of what the parent looked like — which is precisely why one formula covers so much ground. Now tap the destinations, because the claim this figure makes is a strong one: every later idea in the ESTIMATE and ACT pillars is one more sentence about this same object. Tap 17: a p-value is nothing but a tail area of this curve, and a confidence interval is its width read backwards, x̄ ± 1.96·SE = ±0.24% — the same object, read two ways, and you can watch which properties light up. Tap 19: run twenty strategies that have no edge whatsoever and this curve alone guarantees that one of them lands past two standard errors, 1 − 0.9520 = 64% of the time; multiple testing is sampling enough clouds until a lucky dot appears. Tap 20: a regression slope has this object too, its width written SE(β̂) with n − k underneath, because each of the k parameters you fit spends a degree of freedom exactly the way the sample mean spent one to earn the n − 1. Tap BOOT: when your statistic has no formula, you build the object by brute force — resample your own hundred days ten thousand times, recompute, and read the width straight off the pile. Tap 30: the Sharpe ratio, quoted everywhere without an error bar, finally gets one, SE ≈ √(1/T), so a Sharpe of 0.50 measured over four years is 0.50 ± 0.50 and its t-statistic is 1.0 — a claim exactly as large as its own noise. Then press THE CRACK, because a map that only shows what works is a lie. Every result on this page rested on two assumptions, and they are drawn here as what they are: snapped links. The first is independent draws — the √n in the denominator comes from Chapter 11's rule that variances add for independent terms, and market returns are not independent. Momentum, overlapping windows, and volatility clustering all mean today leans on yesterday, so a hundred days buys you fewer than a hundred independent facts; tap that link and watch the honest curve widen by 1.72×, which is to say the error bar you were about to print was 42% too narrow. That repair is Chapter 21. The second is finite variance — every step of the derivation divided by a σ that was assumed to exist and settle down; tap that link and the curve grows tails that never let the variance converge, at which point σ/√n is not merely inaccurate, it is undefined. That is Chapter 32. Both broken links are the ones markets break first, and that is not a coincidence: they break precisely where the money is. So that is where this sits and what it still owes. Everything from here is one more sentence about a single object, and the handover to the next chapter is a single question — if a statistic has a sampling distribution, how surprising is your particular value under a null?
Look at what you can do now that you could not do at the top of the page. Shown any number computed from data, you reach automatically for three properties of the cloud behind it: where it is centred, how wide it is, and what shape it has. You know all three can be had from the single sample in front of you, which is the trick the whole of statistics is built on. You will never again confuse σ with SE, because you watched one refuse to move while the other collapsed.
You can derive n−1 rather than reciting it, and you can answer the question almost no practitioner can. A degree of freedom is a freedom of the residuals, which must sum to zero, so n numbers carry n−1 independent pieces of information about spread. And when n = 1 the formula honestly refuses to answer instead of confidently lying.
Everything from here is one more sentence about the sampling distribution. Chapter 17's p-value is a tail probability of it, and its confidence interval is its width read backwards. Chapter 19's multiple-testing problem is what happens when you sample enough clouds to find a lucky dot. Chapter 20's regression standard errors are this same object for a slope, with n−k in place of n−1. The bootstrap is this object computed by brute force when no formula exists, and Chapter 30 finally gives the Sharpe ratio the error bar it has been missing.
And the crack is already visible, which is the honest part. Every result on this page assumed independent draws and a finite variance. Markets break both, and they break them exactly when it matters most. That is the door Chapter 21's autocorrelation and Chapter 32's fat tails walk through.
One sentence to carry out of here, smaller and more useful than any formula on the page. Before you believe a number you computed, ask how wide the cloud behind it was.