The collection

Every essay

One idea per essay, ordered so that the earlier ones set up the later ones — but nothing here depends on being read in sequence.

The distribution itself

Sums converge on it, which is the most-quoted theorem in the subject. What the quoting leaves out is the rate: the middle converges quickly and the tail does not, and the tail is where the approximation is actually read.

00.1000.2000.3000.400-2024standardised sumdensityskew 0.7472/√n = 0.70740,000 sums, one seed eachthe rate is predicted, not just the shape

Sums of almost anything

The theorem says sums converge on one shape whatever they are sums of, which is remarkable and true. Watching it happen from a one-sided skewed source, with the rate of convergence predicted in advance, is more convincing than watching the shape appear.

6 figures
sample sizerelative error10204080160320640128010%1%at the mediantwo sigma outthree sigma outexact binomial against its normal approximationthe tail converges last

The tail converges last

The central limit theorem is usually shown as a shape arriving. What the demonstration leaves out is the rate — and the rate is wildly different in the middle and in the tail, which is where every approximation in the subject is actually read.

6 figures
20 independent samples, 40 points eachall twenty are normalthis is what the noise looks like

What normal actually looks like

A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.

6 figures
00.1000.2000.3000.400-4-202standard deviations from the meandensity68.27%95.45%99.73%the bands are integrated, not recalledsigma = 1.00

The shape, and where its mass is

68, 95, 99.7 is recited more often than any other set of numbers in the subject. They are integrals of a specific curve, they are worth computing rather than remembering, and the third one is the one people misuse.

6 figures

Intervals, counted

An interval that claims 95% is making a checkable statement about a procedure. Build every possible sample and count. The interval taught first fails, the failure is worst where proportions are most often reported, and more data does not monotonically help.

0.4000.6000.80010.2000.4000.6000.800the true proportionactual coverage of a nominal 95% intervalnominal 95%Wald — the textbook oneWilsonAgresti–CoullClopper–Pearsonsummed over all 31 outcomesno interval is 95% until it is counted

What the 95% refers to

An interval that claims 95% is making a checkable statement about a procedure, not about the interval in front of you. Build every possible sample and count, and the interval taught first turns out to cover 87.6% of the time.

6 figures
0.8000.8500.9000.950255075100sample sizecoverage at a true proportion of 0.15n = 19n = 20WilsonWaldexact coverage at every n from 10 to 120more data is not automatically better here

More data is not monotonically better

Coverage of an interval for a proportion does not improve smoothly as the sample grows. It oscillates, and there are larger samples that cover materially worse than smaller ones — a sample of twenty covers twelve points worse than a sample of nineteen.

6 figures
the truth, 0.350.00.20.40.60.80 of 20 missedthe 95% belongs to the procedure, not to one interval

Twenty intervals and one expected miss

The 95% belongs to the procedure, not to the interval in front of you. Twenty intervals from twenty samples make that visible in a way no definition does, and the one that misses is not a mistake.

6 figures
Wald0.241covers 94.2%Wilson0.247covers 96.5%Agresti–Coull0.259covers 96.5%Clopper–Pearson0.273covers 98.3%expected width, and what it buysorange: fails its nominal levelthe shortest interval is the one that misses

The shortest interval is the one that misses

Four intervals for the same data, with their widths and their coverage measured together. The narrowest is the one that fails its stated level, which is exactly why it looks the most appealing.

6 figures
the sample mean93.6%nominal 95%the sample maximum0.0%nominal 95%3,000 samples of 40the same procedure, two statisticsresampling cannot see past the data

Where the bootstrap lies

Resampling is the most generally useful trick in the subject and it has a failure mode that is easy to state: it cannot see past the data. For a statistic that lives at the edge of the sample, coverage collapses from 95% to almost nothing.

6 figures
00.1000.2000.3000.400-4-2024standard errors from the meandensityt 2.57z 1.96solid: t · dashed: normal31% wider at 5 df

The correction for not knowing the spread

The t distribution exists because the standard deviation is estimated rather than known. At eight observations, using the normal instead makes every interval 12% too short — and the coverage that follows can be measured rather than argued about.

6 figures

Tests, and the second number

A p-value alone cannot be read: the same 0.04 means different things at different sample sizes, and nothing at all without knowing how many analyses were available. Every figure here carries the number that makes it interpretable.

05001e+30.2000.4000.6000.800p-valuecountflatKS distance 0.0165a p-value that is not uniform is not a p-value

A p-value that is not flat is not a p-value

Under a true null, p-values are uniform. That is stronger than saying the test rejects 5% of the time, it constrains the whole distribution rather than one point of it, and it catches implementation errors that a rejection rate sails past.

6 figures
00.2500.5000.750100.5001true effect, in standard deviationsprobability of rejecting80%, the usual target56% at d = 0.54,000 simulated studies per pointthe second number a p-value needs

What a p-value does not say

The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.

6 figures
02505007501e+3-0.50000.5001effect the study reportsstudiesthe truth, 0.3what gets published, 0.61shaded: reached p below 0.05power 20%, inflation 2.04×

The winner's curse

Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.

6 figures
00.2000.4000.6000.8005101520analyses available to the researcherchance of finding something significantthe nominal 5%if the analyses were independentcorrelated, as real ones areno effect present anywhereevery individual analysis is correct

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 58% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

6 figures

Reversals that are not errors

Simpson's reversal, the base rate, regression to the mean. Each is normally taught with one famous table. A table is a point; these are swept, so how much of the space behaves that way and how large it can get both have answers.

What makes it checkable

A seeded generator, so a figure is the same on every build. A closed form beside every simulation, so there are two routes to each number. And p-values checked for uniformity, which catches errors no rejection rate would reveal.