Twenty intervals and one expected miss
Worth reading first: What the 95% refers to.
The definition of a confidence interval is short, correct, and almost impossible to absorb from reading it. This figure is the definition, drawn.
What is random and what is not
Before the sample is drawn, everything is random: the data, and therefore the interval that will be built from it. The statement “this procedure produces an interval containing the truth 95% of the time” is a statement about that randomness, and it is true.
After the sample is drawn, nothing is random any more. The interval is a pair of numbers. The truth is a number. Either it is between them or it is not. There is no 95% left anywhere.
That is why the careful phrasing is always about the procedure. It is not pedantry — it is the difference between a statement that can be checked by repetition and one that cannot be checked at all.
What the picture adds
Three things a definition cannot convey.
The intervals have different widths. Each is built from its own sample, so each gets its own estimate of the spread. The narrow ones are not more precise studies; they are samples that happened to look tidy.
The centres wander more than people expect. The point estimate moves around substantially from sample to sample even at a reasonable n, which is the thing that makes a single study’s headline number less informative than it appears.
The misses are not near-misses. An interval that fails usually fails clearly, sitting entirely to one side. That is a consequence of the centre and the width being estimated from the same data: a sample that is unluckily off-centre also tends to look tight, so it produces a confident interval in the wrong place.
One miss is the modal outcome, not the required one
A panel of twenty showing exactly one miss invites a reading it does not support, so the whole distribution is worth writing down.
With twenty independent intervals each covering with probability 0.95, the number of misses is binomial:
- no misses: 35.8%
- exactly one: 37.7%
- exactly two: 18.9%
- three or more: 7.5%
So one miss is the single most likely outcome and it is barely ahead of none, and two or more misses happens more than a quarter of the time.
That matters for how the picture should be read. It is not a demonstration that one interval in twenty misses; it is one draw from a distribution whose most likely value happens to be one. A panel with three misses is not evidence that anything is wrong — it happens on one panel in thirteen — and a panel with none is not evidence that the intervals are conservative.
Which is the same point the essay makes about a single interval, applied one level up. Twenty intervals is a sample too, and a count of misses out of twenty carries a standard error of about one.
One in twenty, and what that licenses
Across twenty intervals one miss is expected. Sometimes it is none, sometimes three; the count itself is random, and with twenty draws at a 5% rate the standard deviation of the number of misses is about one.
So a run of twenty with two misses is entirely ordinary and says nothing about whether the procedure is working. Checking a procedure needs thousands of repetitions, or an exact sum over the sample space — which is exactly why the coverage figures elsewhere on this site are computed rather than eyeballed from a picture like this one.
This figure is for intuition. It is not evidence about the procedure, and it is worth being clear which figures on a site are which.
The question this reframes
Once the 95% is understood as a property of the procedure, a common question dissolves.
“My interval is [2.1, 4.8]. What is the probability the true value is in there?”
Under the frequentist framing the answer is that the question has no content: the probability is 0 or 1 and which one is unknown. That is unsatisfying, and it is the honest answer to the question as asked.
The question people mean usually has a Bayesian answer, which requires a prior and gives a credible interval — a different object with a different guarantee. The two coincide in some standard cases and diverge in others, and the divergence is largest exactly when the prior is doing real work, which is when the data is weak.
What is not available is a frequentist interval reinterpreted as a probability statement about the parameter. That is the most common misreading in applied statistics, and the picture above is the cheapest available correction: the truth is a fixed vertical line, and it is the intervals that move.
What a single interval is good for
Having said what it is not, it is worth saying what it is, because the negative framing can leave a reader thinking intervals are useless.
A single interval is a summary of what the data is compatible with. Values well outside it are ones the data argues against; values inside it are ones the data does not distinguish. That reading requires no probability statement about the parameter and it is what an interval genuinely supports.
It is also, usefully, an inverted test: the interval is exactly the set of null hypotheses that would not have been rejected at the corresponding level. That equivalence is what the better intervals for a proportion are built from, and it is why they behave so much better than the one obtained by plugging an estimate into a standard error formula.
Why the width varies
The intervals in the figure are not the same length, and the reason is worth drawing out because it produces a second, subtler problem.
Each interval’s width comes from an estimate of the spread computed from its own sample. A sample that happened to look tight produces a small estimated standard error and a narrow interval; one that happened to look spread produces a wide one.
So width is not a measure of study quality. Among studies of identical design, the narrow-looking ones are partly the lucky ones — and, since the centre and the width come from the same data, a sample that is unluckily off-centre also tends to look tight.
That correlation is why the misses in the figure are clear rather than marginal. The intervals that fail are disproportionately the confident-looking ones, which is the worst possible arrangement for a reader judging by appearance.
The interval as an inverted test
The most useful reframing available, and the one that makes the better procedures intelligible.
A 95% interval is exactly the set of null hypotheses that would not be rejected at the 5% level by this data. Every value inside it is one the data does not argue against; every value outside is one it does.
That equivalence has three consequences worth carrying.
It explains where the good intervals come from. Wilson’s interval is constructed by inverting the test directly rather than by plugging an estimate into a standard error, which is why it behaves so much better.
It makes “the interval excludes zero” and “p < 0.05” the same statement, which they are, and stops them being reported as two findings.
And it explains why non-overlap is the wrong test for comparing two intervals. Two 95% intervals can overlap while the difference between the parameters is significant, because the interval for a difference is not built from the two individual intervals. Comparing by eye for overlap is a common and wrong shortcut.
What the interval does not contain
Three things a reader is entitled to expect and will not find in it.
The systematic error. The interval covers sampling variation and nothing else. A biased measurement produces a tight interval around the wrong value, and no sample size fixes it. Most published intervals are narrower than the uncertainty a careful person would report.
The model uncertainty. The interval is conditional on the model being right — the distribution, the independence, the functional form. Uncertainty about the model is not in it.
Anything about a future observation. A confidence interval for a mean is about the mean. A prediction interval for one new value is far wider, and confusing them understates the spread badly.
Reading the figure honestly
A closing note on what this particular picture is for, since it is the most decorative on the site.
It is an intuition pump, not evidence. Twenty intervals cannot establish that a procedure covers at 95% — the count of misses is itself random with a standard deviation of about one, so anything from zero to four misses is ordinary.
The evidence is elsewhere, in an exact sum over the whole sample space that carries no randomness at all.
Keeping those two roles apart matters. A site that showed only this figure would be making a claim its figure could not support, and a site that showed only the exact computation would be correct and would not have conveyed what the number means. Both are needed and they are doing different jobs.
Why the count of misses is itself random
A last observation that the figure quietly demonstrates and that is worth stating.
The number of misses across twenty intervals is a binomial count with n = 20 and p = 0.05. Its mean is 1 and its standard deviation is about 0.97, so zero, one, two or three misses are all unremarkable and four is not surprising.
That has a consequence for reading the figure honestly: no run of twenty could confirm or refute the 95%. A run with three misses is entirely consistent with a correct procedure, and a run with none is consistent with a procedure covering 99%.
Which is exactly why the coverage figures elsewhere on this site are computed by exact summation or by tens of thousands of trials. Twenty is enough to convey what the number means and nowhere near enough to measure it — and being clear about which figures are for conveying and which are for measuring is part of what a site like this owes its reader.
Conditional coverage, and why it is uncomfortable
A subtlety that the picture makes visible and that most treatments skip.
The 95% is an unconditional guarantee: across all samples, 95% of intervals cover. It is not a guarantee conditional on anything about the sample that was drawn.
In particular, the narrow intervals in the figure do not cover 95% of the time and neither do the wide ones. Conditioning on the observed width, the coverage differs — narrow intervals came from samples that looked tight, which correlates with being unrepresentative, so they cover less often than 95%.
That is uncomfortable because a reader always has a specific interval with a specific width, and the guarantee attaching to the average over widths is not obviously the relevant one.
The problem has a name — the reference class problem — and no fully satisfactory frequentist answer. Conditional inference exists and addresses some cases; the Bayesian route sidesteps it by making a statement about the parameter given this data.
It is worth knowing that this is an open seam in the framework rather than a subtlety that resolves on closer reading. The 95% is real, it is a property of the procedure, and the question of what it licenses about the interval in hand is genuinely harder than the definition suggests.
What to write in a paper
The practical upshot, since the essay is otherwise interpretive.
Report the interval and the estimate, not the p-value alone. The interval carries the location and the precision together.
Say what procedure produced it, because the procedures differ materially and “95% CI” does not identify one.
Do not compare two intervals by overlap. Build the interval for the difference.
And do not describe it as containing the parameter with 95% probability, which is the one sentence this whole essay exists to prevent. “Values outside this range are ones the data argues against” is accurate, requires no probability statement about the parameter, and is what a reader wanted anyway.
How many misses to expect, and how surprised to be
The figure shows twenty intervals with one miss, and one is the expected number. The next question a careful reader has is how far from one the count can reasonably fall, and it has an exact answer that makes the picture much less suggestive than it looks.
Counting misses across twenty thousand repetitions of the whole twenty-interval experiment:
- 36.3% of the time, no interval misses at all
- 37.2% of the time, exactly one misses
- 7.6% of the time, three or more miss
and the mean number of misses is 0.997, which is the 1.0 the procedure promises.
The closed form agrees, as it must: the number of misses is binomial with twenty trials and a one-in-twenty chance, giving 35.8%, 37.7% and 7.5% for the same three quantities. Two routes, no shared arithmetic, the same answers to within the simulation’s own error.
The striking number is the first. A run of twenty intervals with no misses at all happens more than a third of the time, and a reader who saw that version of the figure would reasonably conclude the procedure was conservative. A run with three or more misses happens one time in thirteen, and a reader seeing that one would reasonably conclude the procedure was broken.
Both readers would be wrong, and neither could have known from the picture in front of them. Twenty is simply not many repetitions, and the count of misses in twenty repetitions is a noisy estimate of a rate — which is the same lesson the whole site keeps returning to, applied now to its own illustration.
What the figure can and cannot establish
This deserves to be stated plainly, because the figure is the most persuasive thing on the page and persuasion is not the same as evidence.
The figure shows what the definition means. Intervals vary, the parameter does not, and the 95% counts intervals rather than describing any one of them. For that purpose one panel of twenty is ideal, and no amount of extra repetition would make the idea clearer.
The figure does not establish that the coverage is 95%. It cannot, and no picture of twenty intervals could. The count it displays is compatible with a true coverage anywhere from about 75% to 100% — a single run of twenty pins the rate down hardly at all.
That is why the number quoted in the prose comes from somewhere else entirely: from an exact sum over the sample space where one is available, and from tens of thousands of repetitions where it is not. The figure illustrates the definition; the arithmetic establishes the value; and keeping those two jobs separate is the difference between a demonstration and a proof.
The general form of the caution applies to every figure on this site that shows a handful of runs. A panel of twenty is chosen because twenty is what a reader can take in, not because twenty is enough to measure anything, and wherever a panel like that appears the measured claim beside it was computed at a scale the panel does not show.
The width varies as much as the position
One feature of the figure is easy to overlook because attention goes to which intervals miss. The intervals are not the same length, and the variation is substantial.
The width depends on the sample standard deviation, which is itself an estimate and wanders from sample to sample. At the sample sizes used here the widest interval in a run of twenty is routinely half again as wide as the narrowest, from the same procedure applied to data generated the same way.
That has a consequence for reading a single published interval. A narrow interval is not evidence of a careful study; it may be a study whose sample happened to have a small spread. And the narrow ones are more likely to miss, for the same reason the Wald interval fails at the edges — an interval built from an underestimate of the spread is both narrower and more likely to be wrong, and the two failures arrive together rather than independently.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A tenth as wide, and both of them right — both name confidence interval, coverage, sample size, standard deviation
- A block size that changes — both name confidence interval, coverage, sample size
- A schedule that reads the mean — both name confidence interval, coverage, sample size
- An interval that covers and says nothing — both name confidence interval, coverage, frequentist interpretation
- Robust is not free — both name confidence interval, coverage, sample size
- The correction for not knowing the spread — both name coverage, sample size, standard deviation
Named objects
A flat tag is an object no other essay names yet.
Confidence intervalCoverageFrequentist interpretationSample sizeSampling variationStandard deviation