Twenty intervals and one expected miss
The definition of a confidence interval is short, correct, and almost impossible to absorb from reading it. This figure is the definition, drawn.
What is random and what is not
Before the sample is drawn, everything is random: the data, and therefore the interval that will be built from it. The statement “this procedure produces an interval containing the truth 95% of the time” is a statement about that randomness, and it is true.
After the sample is drawn, nothing is random any more. The interval is a pair of numbers. The truth is a number. Either it is between them or it is not. There is no 95% left anywhere.
That is why the careful phrasing is always about the procedure. It is not pedantry — it is the difference between a statement that can be checked by repetition and one that cannot be checked at all.
What the picture adds
Three things a definition cannot convey.
The intervals have different widths. Each is built from its own sample, so each gets its own estimate of the spread. The narrow ones are not more precise studies; they are samples that happened to look tidy.
The centres wander more than people expect. The point estimate moves around substantially from sample to sample even at a reasonable n, which is the thing that makes a single study’s headline number less informative than it appears.
The misses are not near-misses. An interval that fails usually fails clearly, sitting entirely to one side. That is a consequence of the centre and the width being estimated from the same data: a sample that is unluckily off-centre also tends to look tight, so it produces a confident interval in the wrong place.
One in twenty, and what that licenses
Across twenty intervals one miss is expected. Sometimes it is none, sometimes three; the count itself is random, and with twenty draws at a 5% rate the standard deviation of the number of misses is about one.
So a run of twenty with two misses is entirely ordinary and says nothing about whether the procedure is working. Checking a procedure needs thousands of repetitions, or an exact sum over the sample space — which is exactly why the coverage figures elsewhere on this site are computed rather than eyeballed from a picture like this one.
This figure is for intuition. It is not evidence about the procedure, and it is worth being clear which figures on a site are which.
The question this reframes
Once the 95% is understood as a property of the procedure, a common question dissolves.
“My interval is [2.1, 4.8]. What is the probability the true value is in there?”
Under the frequentist framing the answer is that the question has no content: the probability is 0 or 1 and which one is unknown. That is unsatisfying, and it is the honest answer to the question as asked.
The question people mean usually has a Bayesian answer, which requires a prior and gives a credible interval — a different object with a different guarantee. The two coincide in some standard cases and diverge in others, and the divergence is largest exactly when the prior is doing real work, which is when the data is weak.
What is not available is a frequentist interval reinterpreted as a probability statement about the parameter. That is the most common misreading in applied statistics, and the picture above is the cheapest available correction: the truth is a fixed vertical line, and it is the intervals that move.
What a single interval is good for
Having said what it is not, it is worth saying what it is, because the negative framing can leave a reader thinking intervals are useless.
A single interval is a summary of what the data is compatible with. Values well outside it are ones the data argues against; values inside it are ones the data does not distinguish. That reading requires no probability statement about the parameter and it is what an interval genuinely supports.
It is also, usefully, an inverted test: the interval is exactly the set of null hypotheses that would not have been rejected at the corresponding level. That equivalence is what the better intervals for a proportion are built from, and it is why they behave so much better than the one obtained by plugging an estimate into a standard error formula.
Why the width varies
The intervals in the figure are not the same length, and the reason is worth drawing out because it produces a second, subtler problem.
Each interval’s width comes from an estimate of the spread computed from its own sample. A sample that happened to look tight produces a small estimated standard error and a narrow interval; one that happened to look spread produces a wide one.
So width is not a measure of study quality. Among studies of identical design, the narrow-looking ones are partly the lucky ones — and, since the centre and the width come from the same data, a sample that is unluckily off-centre also tends to look tight.
That correlation is why the misses in the figure are clear rather than marginal. The intervals that fail are disproportionately the confident-looking ones, which is the worst possible arrangement for a reader judging by appearance.
The interval as an inverted test
The most useful reframing available, and the one that makes the better procedures intelligible.
A 95% interval is exactly the set of null hypotheses that would not be rejected at the 5% level by this data. Every value inside it is one the data does not argue against; every value outside is one it does.
That equivalence has three consequences worth carrying.
It explains where the good intervals come from. Wilson’s interval is constructed by inverting the test directly rather than by plugging an estimate into a standard error, which is why it behaves so much better.
It makes “the interval excludes zero” and “p < 0.05” the same statement, which they are, and stops them being reported as two findings.
And it explains why non-overlap is the wrong test for comparing two intervals. Two 95% intervals can overlap while the difference between the parameters is significant, because the interval for a difference is not built from the two individual intervals. Comparing by eye for overlap is a common and wrong shortcut.
What the interval does not contain
Three things a reader is entitled to expect and will not find in it.
The systematic error. The interval covers sampling variation and nothing else. A biased measurement produces a tight interval around the wrong value, and no sample size fixes it. Most published intervals are narrower than the uncertainty a careful person would report.
The model uncertainty. The interval is conditional on the model being right — the distribution, the independence, the functional form. Uncertainty about the model is not in it.
Anything about a future observation. A confidence interval for a mean is about the mean. A prediction interval for one new value is far wider, and confusing them understates the spread badly.
Reading the figure honestly
A closing note on what this particular picture is for, since it is the most decorative on the site.
It is an intuition pump, not evidence. Twenty intervals cannot establish that a procedure covers at 95% — the count of misses is itself random with a standard deviation of about one, so anything from zero to four misses is ordinary.
The evidence is elsewhere, in an exact sum over the whole sample space that carries no randomness at all.
Keeping those two roles apart matters. A site that showed only this figure would be making a claim its figure could not support, and a site that showed only the exact computation would be correct and would not have conveyed what the number means. Both are needed and they are doing different jobs.
Why the count of misses is itself random
A last observation that the figure quietly demonstrates and that is worth stating.
The number of misses across twenty intervals is a binomial count with n = 20 and p = 0.05. Its mean is 1 and its standard deviation is about 0.97, so zero, one, two or three misses are all unremarkable and four is not surprising.
That has a consequence for reading the figure honestly: no run of twenty could confirm or refute the 95%. A run with three misses is entirely consistent with a correct procedure, and a run with none is consistent with a procedure covering 99%.
Which is exactly why the coverage figures elsewhere on this site are computed by exact summation or by tens of thousands of trials. Twenty is enough to convey what the number means and nowhere near enough to measure it — and being clear about which figures are for conveying and which are for measuring is part of what a site like this owes its reader.
Conditional coverage, and why it is uncomfortable
A subtlety that the picture makes visible and that most treatments skip.
The 95% is an unconditional guarantee: across all samples, 95% of intervals cover. It is not a guarantee conditional on anything about the sample that was drawn.
In particular, the narrow intervals in the figure do not cover 95% of the time and neither do the wide ones. Conditioning on the observed width, the coverage differs — narrow intervals came from samples that looked tight, which correlates with being unrepresentative, so they cover less often than 95%.
That is uncomfortable because a reader always has a specific interval with a specific width, and the guarantee attaching to the average over widths is not obviously the relevant one.
The problem has a name — the reference class problem — and no fully satisfactory frequentist answer. Conditional inference exists and addresses some cases; the Bayesian route sidesteps it by making a statement about the parameter given this data.
It is worth knowing that this is an open seam in the framework rather than a subtlety that resolves on closer reading. The 95% is real, it is a property of the procedure, and the question of what it licenses about the interval in hand is genuinely harder than the definition suggests.
What to write in a paper
The practical upshot, since the essay is otherwise interpretive.
Report the interval and the estimate, not the p-value alone. The interval carries the location and the precision together.
Say what procedure produced it, because the procedures differ materially and “95% CI” does not identify one.
Do not compare two intervals by overlap. Build the interval for the difference.
And do not describe it as containing the parameter with 95% probability, which is the one sentence this whole essay exists to prevent. “Values outside this range are ones the data argues against” is accurate, requires no probability statement about the parameter, and is what a reader wanted anyway.