Concept

Sampling variation — where it appears

The movement in a computed quantity from one possible sample to the next, which is what a standard error measures. One realisation of it is an anecdote, which is why a figure showing a single run here is usually a figure showing twenty.

Named by 14 essays across 10 fields — each of them below, with the objects they name alongside it.

One wrong model, four designs, four slopes. The slope a straight line converges to when the truth is a quadratic, under four covariate distributions, by two routes: the population projection in closed form, and the mean of 2500 fitted slopes at 200 rows apiece. The even spread over [0, 2] gives 1.6000 and the same spread moved to [1, 3] gives 2.6000, while widening it to [0, 4] gives 2.6000 — the same number as the shifted one, because a symmetric design's target is the truth's tangent slope at the design's own mean and does not read the spread at all. An exponential spread with the SAME mean as the first gives 2.6000. So two studies of one world, each fitting the same wrong model, honestly report slopes 1.0000 apart, and neither is making an error.

What a wrong model estimates

A straight line fitted to a curved truth converges on the tangent at its own design's mean. Two honest studies of one world, fitting the same wrong model, report 2.600000 and 1.600000, and neither is in error.

sandwich · Misspecification
How often a randomised trial reports the reversal, advantage 6 points. Simple randomisation against randomisation stratified by group, 4,000 trials at each size. The simple design reverses on 3.40% of trials at its worst size and 0.50% at 1280 units; the stratified design reverses on none of them, at any size.

The reversal a coin cannot prevent

Randomisation removes Simpson's reversal in expectation, which is not the same as removing it. A correctly randomised trial of eighty units, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials — and stratifying the randomisation takes that to zero at every size.

reversal · Simpson
Two 95% bands for a quantile plot of 40 points. The outer band is left by 5% of genuinely normal samples — which is what a reader is using a band for. The inner one holds each point separately at 95%, which is what software draws, and 45.0% of genuinely normal samples step outside it. The outer is the inner widened by a factor of 1.502.

The band the eye was standing in for

The confidence band software draws on a quantile plot holds each point at 95%, and a genuinely normal sample of forty has forty chances to leave it — so 45.0% of them do. The band a reader is actually using is that one widened by a factor of 1.502, and nothing draws it.

lineup · Qq
Where a replication's estimate lands against a 95% interval, replication the same size. The chance that a 95% interval contains a replication's estimate is 95.00% when the original landed on the truth, 82.99% one standard error away and 48.40% two away. Averaged over where originals land it is 83.42%, and 5.00% of originals capture a replication less than half the time.

Five times in six

A 95% interval is read as a 95% chance that a replication's estimate will land inside it. With the spread known and a replication of the same size, the chance is 83.42% — five times in six — because both estimates are uncertain. An original that landed two standard errors from the truth captures a replication 48.40% of the time; a replication a tenth the size lands inside 44.54% of the time; and among significant originals from studies with 17% power, 66.94%.

alongside · Repetition
Unrepresentative in every respect but the one that matters. Three properties of the complete cases as the chance of being observed leans harder on the regressor, in closed form, at 35.0% of outcomes missing throughout. The mean of the regressor among the rows kept climbs from 0.0000 to 0.5528 against a population mean of zero, and the mean of the outcome from 0.0000 to 0.3980 above its own. The bias in the fitted slope is exactly zero at every one of the ten settings, because selection acting on the regressor alone leaves the conditional law of the outcome given the regressor untouched and least squares conditions on exactly that. The sample is wrong about almost everything and right about the one quantity being estimated.

Dropping the incomplete rows

Push the missingness until the rows that survive have a covariate mean of 0.543905 against a population zero and a variance of 0.5041 against one, and the fitted slope is still exactly right. Where the rule reads the outcome instead, the same sweep takes coverage to 2.42% at eight hundred rows.

missing · Missingness
One forecaster, one sample, and the bins it was read in. A reliability diagram of 500 forecasts in 10 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.002262 and its expected calibration error 0.0358, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.

A curve that is a binning

A forecaster with no miscalibration in it at all reads 0.001429 at five bins and 0.014100 at fifty, on the same five hundred forecasts. The closed form is K/n times the forecaster's own irreducible score, and subtracting it returns zero.

calibrate · Calibration
What the plot says, and what the interval does, at n = 40. For each source: how often a quantile plot of the data leaves its pointwise band, and how often the 95% t interval for the mean misses. The two-lump source leaves the band on 100% of samples and its interval covers 94.80%; the t on three degrees of freedom leaves it on 57% and covers 95.73%, the best of the five.

The plot is about the wrong quantity

A t interval needs the sampling distribution of the mean to be normal, not the data. A two-lump source leaves its quantile band on 100% of samples of forty and its interval covers 94.80%; a t on three degrees of freedom leaves it on 57% and covers 95.73%, the best of five sources.

lineup · Qq
What each instrument costs to read. The number of draws each instrument needs to separate a rectangular block from a trapezoidal one at two standard errors, at a block length of 20 and 120 rows — measured from each instrument's own spread on the same draws. The implied variance needs 7.0 and the 95% point needs 20.2, a factor of 2.90 at this block length. There is a closed form beside it and it does not depend on either the scale or the size of the gap: the standard error of a p-quantile is √(p(1−p))/f(q) over √B where a standard deviation's is σ/√(2B), which at the 95% point of a nearly normal reference distribution is 3.30 times as many draws for the same statement. And the quantile route needs every one of those draws resampled, where the variance route needs none.

Measuring a variance rather than a quantile

A resample's implied long-run variance can be computed from the sample with no resampling in it at all. A critical value cannot, and the difference is a factor of three in the draws before any of the resampling is counted.

crossing · Bootstrap
Twenty 95% intervals for a proportion that really is 0.35. 2 of the twenty miss the true value. The 95% is a property of the procedure across repetitions — no single interval has a 95% chance of anything, because it either contains 0.35 or it does not.

Twenty intervals and one expected miss

The 95% belongs to the procedure, not to the interval in front of you. Twenty intervals from twenty samples make that visible in a way no definition does, and the one that misses is not a mistake.

intervals · Repetition
Twenty samples of 40, every one of them genuinely normal. Each panel is a quantile-quantile plot of 40 draws from a normal distribution. The worst point in the worst panel sits 0.87 standard deviations off the line. Anything a reader would reject here would be a false alarm.

What normal actually looks like

A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.

normal · Qq
A filled value is not an observation. What a 95% interval for the slope actually covers after each way of handling 35.0% missing outcomes, counted over 4000 studies of 200 rows. Dropping the incomplete rows covers 95.93%. Filling with the observed mean covers 13.85%, because the estimate itself has moved. Filling with a fitted value covers 80.85% against a closed prediction of 79.73%: the estimate is right and the reported standard error is short by a factor of 0.6567 against a predicted 0.6500, because the residual sum of squares is divided by the whole sample's degrees of freedom. Adding residual noise recovers the spread and covers 85.78% against a predicted 84.62%, since the interval still ignores the variance of having imputed at all.

One imputation is not an observation

Three ways of filling a missing outcome, under a mechanism that makes dropping the rows beyond reproach. Filling with the observed mean covers 13.85%, filling with a fitted value covers 80.85%, adding noise covers 85.78%, and the thing all three were meant to improve on covers 95.93%.

missing · Missingness
How long an honest forecaster looks broken for. The calibration error shown by a forecaster with no miscalibration in it at all, at five record lengths, drawn against one over the square root of the length so that the closed form is a straight line through the origin. It is 0.1252 at fifty forecasts and 0.0090 at ten thousand, against a closed form of √(2K/πn) times the mean root bin variance which gives 0.1257 and 0.0089. The threshold drawn across it is 0.02, a figure routinely read as evidence that something is wrong; the mean falls under it at 1976 forecasts and the 95th percentile at about 4111. Below that, an honest forecaster and a miscalibrated one are being told apart by a statistic that is mostly the sample size.

The miscalibration a perfect forecaster shows

A forecaster whose true reliability is exactly zero shows a calibration error of 0.1252 on fifty forecasts and 0.0090 on ten thousand. Every one of 1,200 blameless hundred-forecast records exceeds the 0.02 routinely read as evidence of a problem, and the mean does not fall under it until 1,976 forecasts.

calibrate · Calibration
The rows are held fixed; only the clusters move. Counted coverage of four 95% intervals for a slope, at five cluster counts with the row count held at 300 throughout and a within-cluster correlation of 0.1, over 6000 draws apiece, with the sizes equal. The interval that counts rows covers 53.42% at 5 clusters — a second closed form says 2Φ(z/√D) − 1 = 54.44% for a design effect of 6.900, and reads nothing about clusters at all. The cluster-robust interval read against a normal covers 74.43% there and 94.20% at 100 clusters; read against a t on G − 1 it covers 85.08% and 94.47%. The number of independent things is the cluster count, and every quantity here is blind to how many rows were typed.

The count that is not the rows

Three hundred rows in five clusters of sixty carry 6.9000 times the variance an independent-rows calculation reports, and the interval that counts rows covers 53.42%. The same five unequal sizes laid out two ways give design effects of 9.3158 and 5.4652.

sandwich · Misspecification
Two companions on one simulation, two hundredfold apart. How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand.

The check worth more than the check

The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.

method · Routes

Named alongside it

The objects these essays reach for when they reach for this one.

Closed formConfidence intervalCoverageLeast squaresMonte CarloNormalityQ–Q plotSample sizeStandard errorBinningComplete-caseCritical value

All concepts