Intervals, counted

Twenty intervals and one expected miss

The 95% belongs to the procedure, not to the interval in front of you. Twenty intervals from twenty samples make that visible in a way no definition does, and the one that misses is not a mistake.

The definition of a confidence interval is short, correct, and almost impossible to absorb from reading it. This figure is the definition, drawn.

Twenty 95% intervals for a proportion that really is 0.350 of the twenty miss the true value. The 95% is a property of the procedure across repetitions — no single interval has a 95% chance of anything, because it either contains 0.35 or it does not.the truth, 0.350.00.20.40.60.80 of 20 missedthe 95% belongs to the procedure, not to one interval
Fig. 1 Twenty samples from the same population, one interval built from each by the same rule. The vertical line is the true value. The intervals that miss it are marked.

What is random and what is not

Before the sample is drawn, everything is random: the data, and therefore the interval that will be built from it. The statement “this procedure produces an interval containing the truth 95% of the time” is a statement about that randomness, and it is true.

After the sample is drawn, nothing is random any more. The interval is a pair of numbers. The truth is a number. Either it is between them or it is not. There is no 95% left anywhere.

That is why the careful phrasing is always about the procedure. It is not pedantry — it is the difference between a statement that can be checked by repetition and one that cannot be checked at all.

Coverage of four nominal 95% intervals, n = 30Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.0.4000.6000.80010.2000.4000.6000.800the true proportionactual coverage of a nominal 95% intervalnominal 95%Wald — the textbook oneWilsonAgresti–CoullClopper–Pearsonsummed over all 31 outcomesno interval is 95% until it is counted
Fig. 2 What the same procedure does across every true proportion, computed exactly rather than sampled twenty times.

What the picture adds

Three things a definition cannot convey.

The intervals have different widths. Each is built from its own sample, so each gets its own estimate of the spread. The narrow ones are not more precise studies; they are samples that happened to look tidy.

The centres wander more than people expect. The point estimate moves around substantially from sample to sample even at a reasonable n, which is the thing that makes a single study’s headline number less informative than it appears.

The misses are not near-misses. An interval that fails usually fails clearly, sitting entirely to one side. That is a consequence of the centre and the width being estimated from the same data: a sample that is unluckily off-centre also tends to look tight, so it produces a confident interval in the wrong place.

Twenty samples of 40, every one of them genuinely normalEach panel is a quantile-quantile plot of 40 draws from a normal distribution. The worst point in the worst panel sits 1.52 standard deviations off the line. Anything a reader would reject here would be a false alarm.20 independent samples, 40 points eachall twenty are normalthis is what the noise looks like
Fig. 3 The same argument for a different object: twenty panels of genuinely normal data, because one panel cannot be read.

One in twenty, and what that licenses

Across twenty intervals one miss is expected. Sometimes it is none, sometimes three; the count itself is random, and with twenty draws at a 5% rate the standard deviation of the number of misses is about one.

So a run of twenty with two misses is entirely ordinary and says nothing about whether the procedure is working. Checking a procedure needs thousands of repetitions, or an exact sum over the sample space — which is exactly why the coverage figures elsewhere on this site are computed rather than eyeballed from a picture like this one.

This figure is for intuition. It is not evidence about the procedure, and it is worth being clear which figures on a site are which.

Coverage against sample size, true proportion 0.15Coverage does not improve monotonically. n = 19 covers 93.8% while the larger n = 20 covers 81.9%. The sample space is discrete, so the endpoints jump as n changes.0.8000.8500.9000.950255075100sample sizecoverage at a true proportion of 0.15n = 19n = 20WilsonWaldexact coverage at every n from 10 to 120more data is not automatically better here
Fig. 4 Coverage as a function of sample size. The procedure has a long-run rate; a single interval has no rate at all.

The question this reframes

Once the 95% is understood as a property of the procedure, a common question dissolves.

“My interval is [2.1, 4.8]. What is the probability the true value is in there?”

Under the frequentist framing the answer is that the question has no content: the probability is 0 or 1 and which one is unknown. That is unsatisfying, and it is the honest answer to the question as asked.

The question people mean usually has a Bayesian answer, which requires a prior and gives a credible interval — a different object with a different guarantee. The two coincide in some standard cases and diverge in others, and the divergence is largest exactly when the prior is doing real work, which is when the data is weak.

What is not available is a frequentist interval reinterpreted as a probability statement about the parameter. That is the most common misreading in applied statistics, and the picture above is the cheapest available correction: the truth is a fixed vertical line, and it is the intervals that move.

Expected width against coverage, n = 30, p = 0.15The Wald interval is the shortest and covers 94.2%. Clopper–Pearson covers 98.3% and is 13% wider. Shortness is not a virtue on its own — an interval of zero width is the shortest of all.Wald0.241covers 94.2%Wilson0.247covers 96.5%Agresti–Coull0.259covers 96.5%Clopper–Pearson0.273covers 98.3%expected width, and what it buysorange: fails its nominal levelthe shortest interval is the one that misses
Fig. 5 Width against coverage. An interval is a summary of what the data is compatible with, and both numbers are part of it.
Coverage of a 95% interval for a mean, n = 8Measured over 20,000 samples. The t interval covers 95.0% and the z interval 91.3%. The difference is the price of pretending the standard deviation was known.t interval, 7 df95.0%± 0.3 at 2 s.e.z interval91.3%± 0.4 at 2 s.e.20,000 samples, one seed eachnominal 95%counted, not assumed
Fig. 6 The same reading for a mean, with the coverage of each procedure counted.

What a single interval is good for

Having said what it is not, it is worth saying what it is, because the negative framing can leave a reader thinking intervals are useless.

A single interval is a summary of what the data is compatible with. Values well outside it are ones the data argues against; values inside it are ones the data does not distinguish. That reading requires no probability statement about the parameter and it is what an interval genuinely supports.

It is also, usefully, an inverted test: the interval is exactly the set of null hypotheses that would not have been rejected at the corresponding level. That equivalence is what the better intervals for a proportion are built from, and it is why they behave so much better than the one obtained by plugging an estimate into a standard error formula.

Why the width varies

The intervals in the figure are not the same length, and the reason is worth drawing out because it produces a second, subtler problem.

Each interval’s width comes from an estimate of the spread computed from its own sample. A sample that happened to look tight produces a small estimated standard error and a narrow interval; one that happened to look spread produces a wide one.

So width is not a measure of study quality. Among studies of identical design, the narrow-looking ones are partly the lucky ones — and, since the centre and the width come from the same data, a sample that is unluckily off-centre also tends to look tight.

That correlation is why the misses in the figure are clear rather than marginal. The intervals that fail are disproportionately the confident-looking ones, which is the worst possible arrangement for a reader judging by appearance.

The interval as an inverted test

The most useful reframing available, and the one that makes the better procedures intelligible.

A 95% interval is exactly the set of null hypotheses that would not be rejected at the 5% level by this data. Every value inside it is one the data does not argue against; every value outside is one it does.

That equivalence has three consequences worth carrying.

It explains where the good intervals come from. Wilson’s interval is constructed by inverting the test directly rather than by plugging an estimate into a standard error, which is why it behaves so much better.

It makes “the interval excludes zero” and “p < 0.05” the same statement, which they are, and stops them being reported as two findings.

And it explains why non-overlap is the wrong test for comparing two intervals. Two 95% intervals can overlap while the difference between the parameters is significant, because the interval for a difference is not built from the two individual intervals. Comparing by eye for overlap is a common and wrong shortcut.

What the interval does not contain

Three things a reader is entitled to expect and will not find in it.

The systematic error. The interval covers sampling variation and nothing else. A biased measurement produces a tight interval around the wrong value, and no sample size fixes it. Most published intervals are narrower than the uncertainty a careful person would report.

The model uncertainty. The interval is conditional on the model being right — the distribution, the independence, the functional form. Uncertainty about the model is not in it.

Anything about a future observation. A confidence interval for a mean is about the mean. A prediction interval for one new value is far wider, and confusing them understates the spread badly.

Reading the figure honestly

A closing note on what this particular picture is for, since it is the most decorative on the site.

It is an intuition pump, not evidence. Twenty intervals cannot establish that a procedure covers at 95% — the count of misses is itself random with a standard deviation of about one, so anything from zero to four misses is ordinary.

The evidence is elsewhere, in an exact sum over the whole sample space that carries no randomness at all.

Keeping those two roles apart matters. A site that showed only this figure would be making a claim its figure could not support, and a site that showed only the exact computation would be correct and would not have conveyed what the number means. Both are needed and they are doing different jobs.

Why the count of misses is itself random

A last observation that the figure quietly demonstrates and that is worth stating.

The number of misses across twenty intervals is a binomial count with n = 20 and p = 0.05. Its mean is 1 and its standard deviation is about 0.97, so zero, one, two or three misses are all unremarkable and four is not surprising.

That has a consequence for reading the figure honestly: no run of twenty could confirm or refute the 95%. A run with three misses is entirely consistent with a correct procedure, and a run with none is consistent with a procedure covering 99%.

Which is exactly why the coverage figures elsewhere on this site are computed by exact summation or by tens of thousands of trials. Twenty is enough to convey what the number means and nowhere near enough to measure it — and being clear about which figures are for conveying and which are for measuring is part of what a site like this owes its reader.

Conditional coverage, and why it is uncomfortable

A subtlety that the picture makes visible and that most treatments skip.

The 95% is an unconditional guarantee: across all samples, 95% of intervals cover. It is not a guarantee conditional on anything about the sample that was drawn.

In particular, the narrow intervals in the figure do not cover 95% of the time and neither do the wide ones. Conditioning on the observed width, the coverage differs — narrow intervals came from samples that looked tight, which correlates with being unrepresentative, so they cover less often than 95%.

That is uncomfortable because a reader always has a specific interval with a specific width, and the guarantee attaching to the average over widths is not obviously the relevant one.

The problem has a name — the reference class problem — and no fully satisfactory frequentist answer. Conditional inference exists and addresses some cases; the Bayesian route sidesteps it by making a statement about the parameter given this data.

It is worth knowing that this is an open seam in the framework rather than a subtlety that resolves on closer reading. The 95% is real, it is a property of the procedure, and the question of what it licenses about the interval in hand is genuinely harder than the definition suggests.

What to write in a paper

The practical upshot, since the essay is otherwise interpretive.

Report the interval and the estimate, not the p-value alone. The interval carries the location and the precision together.

Say what procedure produced it, because the procedures differ materially and “95% CI” does not identify one.

Do not compare two intervals by overlap. Build the interval for the difference.

And do not describe it as containing the parameter with 95% probability, which is the one sentence this whole essay exists to prevent. “Values outside this range are ones the data argues against” is accurate, requires no probability statement about the parameter, and is what a reader wanted anyway.