A method that claims to be right 95% of the time is making a statement anyone can count. Almost nobody counts it.
The interval taught first in every introductory course covers 87.6% of the time when it says 95%, and the failure is worst exactly where proportions are usually reported — a rare event in a small sample. That is not an estimate. For a proportion the sample space is finite, so the coverage is a sum over every outcome that could have occurred, and the number is exact. The same discipline applies to everything here: a test's p-values are checked for flatness, a simulation carries its seed, and every claim that a picture makes has been run across many seeds rather than one.
Start anywhere
19 essays
What the 95% refers to
An interval that claims 95% is making a checkable statement about a procedure, not about the interval in front of you. Build every possible sample and count, and the interval taught first turns out to cover 87.6% of the time.
The distribution itselfSums of almost anything
The theorem says sums converge on one shape whatever they are sums of, which is remarkable and true. Watching it happen from a one-sided skewed source, with the rate of convergence predicted in advance, is more convincing than watching the shape appear.
Tests, and the second numberA p-value that is not flat is not a p-value
Under a true null, p-values are uniform. That is stronger than saying the test rejects 5% of the time, it constrains the whole distribution rather than one point of it, and it catches implementation errors that a rejection rate sails past.
Reversals that are not errorsSimpson's reversal is a region, not a table
The treatment wins in both groups and loses overall. That is normally shown with one famous table, which cannot answer the two questions a reader has — how often, and how large. Swept, it turns out to occupy 31% of the allocation space.
What makes it checkableThe seed is part of the figure
Every other site in this fleet draws from a deterministic rule, so a figure either is or is not what it claims. Here the figures are samples, and a sample can be right by luck. That changes what a figure has to carry.
Intervals, countedMore data is not monotonically better
Coverage of an interval for a proportion does not improve smoothly as the sample grows. It oscillates, and there are larger samples that cover materially worse than smaller ones — a sample of twenty covers twelve points worse than a sample of nineteen.
The distribution itselfThe tail converges last
The central limit theorem is usually shown as a shape arriving. What the demonstration leaves out is the rate — and the rate is wildly different in the middle and in the tail, which is where every approximation in the subject is actually read.
Tests, and the second numberWhat a p-value does not say
The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.
Reversals that are not errorsWhat a positive test is worth
A test that is 90% sensitive and 95% specific sounds accurate. For a condition affecting one person in a thousand, 98% of its positive results are wrong, and a worse test on a commoner condition beats a better test on a rare one.
What makes it checkableTwo routes to every number
A site about probability that only simulates has one route to each answer and no way to tell a right one from a plausible one. Every important number here is computed twice, by arithmetic that shares nothing, and the two are required to agree.
The distribution itselfWhat normal actually looks like
A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.
Intervals, countedTwenty intervals and one expected miss
The 95% belongs to the procedure, not to the interval in front of you. Twenty intervals from twenty samples make that visible in a way no definition does, and the one that misses is not a mistake.
Tests, and the second numberThe winner's curse
Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.
Reversals that are not errorsRegression to the mean
Select the worst performers, measure them again, and they improve. Select the best and they decline. No intervention is required for either, the size of the apparent effect is predictable from the correlation alone, and it is the reason so many things appear to work.
The distribution itselfThe shape, and where its mass is
68, 95, 99.7 is recited more often than any other set of numbers in the subject. They are integrals of a specific curve, they are worth computing rather than remembering, and the third one is the one people misuse.
Intervals, countedThe shortest interval is the one that misses
Four intervals for the same data, with their widths and their coverage measured together. The narrowest is the one that fails its stated level, which is exactly why it looks the most appealing.
Tests, and the second numberTwenty analyses of nothing
Twenty honest, correct analyses of data with no effect in it find something significant 58% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.
Intervals, countedWhere the bootstrap lies
Resampling is the most generally useful trick in the subject and it has a failure mode that is easy to state: it cannot see past the data. For a statistic that lives at the edge of the sample, coverage collapses from 95% to almost nothing.
Intervals, countedThe correction for not knowing the spread
The t distribution exists because the standard deviation is estimated rather than known. At eight observations, using the normal instead makes every interval 12% too short — and the coverage that follows can be measured rather than argued about.
Threads running through
themes, not chapters
Count it, do not claim it
An interval that says 95% is making a statement about a procedure, and a procedure can be run ten thousand times. Every coverage figure on this site is a count — for a proportion it is an exact sum over the whole sample space, so the answer carries no simulation noise at all.
One run is an anecdote
A figure showing a single simulation is showing one draw from a distribution of figures. Where the claim is about behaviour rather than about one dataset, the figure runs across many seeds and reports what held for all of them — and where it shows one run, it shows twenty of them side by side.
Two routes to a number
Every simulation here has a closed form beside it and every closed form has a simulation. A distribution function and its quantile must invert each other to ten digits; an exact coverage sum and a Monte Carlo count must agree. Neither route can confirm itself, and they share no arithmetic.
The tail is where it is read
A p-value, a control limit and a risk figure are all tail statements, and the tail is exactly where every approximation in this subject is worst. The central limit theorem converges in the middle long before it converges where anyone looks.
The second number
A p-value cannot be interpreted alone. The same 0.04 corresponds to a large effect at ten observations and a negligible one at two thousand, and it means nothing at all until you know how many analyses were available to produce it.
Reversals that are nobody's mistake
Simpson's paradox, regression to the mean and the winner's curse all arise from correct arithmetic applied honestly. Each is shown as a region of a parameter space rather than one famous table, so the questions of how often and how large have answers.