A tenth as wide, and both of them right
Worth reading first: The shape, and where its mass is · What the 95% refers to.
Three bands can be drawn around the same sample mean. All three are written “95%”. At twenty observations their half-widths are 0.468, 2.145 and 2.752 sample standard deviations, and every one of them covers what it claims to cover at very nearly exactly 95%.
The three questions are different and only one of them is usually asked out loud.
Where is the mean? — the parameter, a fixed unknown number. Half-width .
Where will the next observation be? — one future draw, which has the population’s own spread in it as well as the uncertainty about where the centre is. Half-width t·s·.
Where is 95% of the population? — a region, and a promise to have covered that region. Half-width k·s, with k the tolerance factor.
The ratio is exact and has one term in it
The first two bands differ by
with the t and the s cancelling entirely. The ratio does not depend on the confidence level, on the population’s spread, or on anything estimated. It is a function of the sample size and nothing else, and at the sizes people work at it is large: 3.32 at ten observations, 4.58 at twenty, 10.05 at a hundred, 20.03 at four hundred.
That is the wrong direction for intuition, and it is worth saying plainly because it inverts what “more data” is expected to do. Collecting more data makes the two bands further apart, not closer. The mean’s interval shrinks like and the prediction interval does not shrink at all — it converges to the population’s own 95% band, which is a fixed width that no amount of data reduces. At four hundred observations the prediction half-width is 1.968 standard deviations, and the population’s is 1.960.
What each one costs the other way
Reading the wrong band is not a rounding error, and the size of the mistake is countable.
The mean’s interval read as a prediction. At ten observations it covers a new draw 48.6% of the time. At a hundred, 15.7%. At four hundred, 7.9%. The claim gets worse with more data, for exactly the reason above: the band the reader is looking at is shrinking and the thing it is being read about is not.
That is the figure worth carrying, because it is the direction nobody expects. A large study reporting its mean with a tight error bar is reporting a tight statement about the mean. If the bar is read as “where the values are”, the reading is less accurate than it would have been from a small study — a hundred observations gives a reading that is wrong six times in seven.
The prediction interval read as a statement about the mean. It covers μ essentially always — at every sample size measured it covered on 100.00% of twenty thousand samples. That failure is not symmetric with the first and it is not harmless. An interval that covers 100% of the time when it claims 95% is not conservative in any useful sense; it is an interval that has stopped distinguishing between hypotheses, so any test built from it has no power at all.
The mean’s interval read as a tolerance interval. This is the first failure again, in the form it usually takes. A reference range, a specification, a “typical values” statement — all of them are questions about the population, and answering them with a standard error gives a band an order of magnitude too narrow.
The notation that causes it
The standard error is the standard deviation divided by , which means the two are the same symbol divided by a number, and the number is usually not printed.
| symbol | what it is the spread of |
|---|---|
| s | the data |
| the sample mean |
At twenty-five observations the second is a fifth of the first. At a hundred it is a tenth. A figure whose error bars are one standard error and a figure whose error bars are one standard deviation are different figures by a factor of ten, and they look identical.
The shape’s own essay makes the same point about ±1σ against 95%, and the two errors compound: a bar drawn at one standard error, read as two standard deviations of the data, is out by a factor of twenty at a hundred observations. The habit that prevents it is one sentence — say what the bar is — and the reason it needs saying is that no convention exists across fields.
Why the third band is not the second one widened
The prediction interval and the tolerance interval are both about the data rather than the parameter, so it is natural to read the tolerance interval as a slightly wider prediction interval. It is a different construction, and the difference matters when n is small.
A prediction interval covers one future observation with probability 95%, averaged over samples and over the future observation together. A tolerance interval covers 95% of the population with probability 95%, and those two probabilities are over different things: the inner one is over the population given the sample, the outer one is over samples.
At twenty observations the prediction half-width is 2.145 and the tolerance half-width is 2.752 — 28% wider. At four hundred they are 1.968 and 2.084, 6% apart. They converge, because both are heading for the population’s own 1.96, but they converge at different rates: the prediction interval is already within half a per cent of its limit at four hundred, and the tolerance interval is still 6% above it, because the confidence half of its promise never becomes free.
The distinction has a clean operational test. Ask how many future observations the statement is about. One: prediction interval. A stated share of all of them: tolerance interval. All of the next m: neither, and the right band is wider than both and depends on m.
What the t is doing in two of the three
Both the mean’s interval and the prediction interval carry the same t quantile, which looks like an accident of notation and is not.
The mean’s interval needs t because (x̄ − μ)/() is a t: a normal over an independent root. The prediction interval needs t because is also a t, for the same reason — the numerator is normal with variance , the denominator is the same independent s, and the ratio standardises to the same distribution on the same n − 1 degrees of freedom.
That is why the two bands share a t and why their ratio has no t in it. The tolerance interval does not standardise to anything, which is why its factor has no closed form: its event is about a difference of two normal integrals rather than about a single standardised quantity.
The shared t also explains a small fact that surprises people. The prediction interval’s factor approaches 1.96 from above and does so slowly at small n — 2.373 at ten observations against 1.96 — and all of that excess is the t rather than the . At ten observations = 1.049, which is 5% of the width, and t(9) = 2.262 against 1.96, which is 15%. So the prediction interval at small samples is mostly a statement about not knowing σ, and only slightly a statement about the extra draw.
Only one of the three is protected by the limit theorem
Everything so far assumes a normal population. Dropping that assumption separates the three bands again, and it separates them along a line nobody draws in advance.
The mean’s interval is protected. x̄ is an average, so the limit theorem applies to it and its distribution becomes normal whatever the population was. Drawing from an exponential — one-sided, skewed, nothing like a bell — the 95% t interval for the mean covers 89.8% at ten observations, 92.9% at thirty, 94.1% at a hundred and 94.8% at four hundred. It is converging on its label, slowly and reliably, and the convergence is the theorem doing its work.
The prediction interval is not protected and does not converge. On the same draws it covers 93.1%, 94.2%, 94.6% and 94.5% — and it stops there. No amount of data helps, because the quantity it is about is one observation and nothing about one observation is an average of anything. The population’s own shape is the thing being predicted, and the interval is built out of a mean and a standard deviation, which do not describe that shape.
The total is not the reading. Split the misses by side and at a hundred observations the interval misses 0.0% below and 5.4% above. Not 2.5% and 2.5%: zero and more than five. The lower limit sits below zero, where an exponential has no mass at all, so it cannot miss low — and every miss the interval is entitled to arrives on the upper side, at twice the rate a reader would price. The two errors are of opposite sign, the sum reads 94.6%, and a coverage table with one column would report that as a success.
That is the general shape and it is worth stating in one line. Averaging is what the limit theorem buys, and a prediction is not an average. A band about a parameter inherits the theorem; a band about a value inherits the population, including everything about the population that the mean and the standard deviation cannot see. The same asymmetry is why an approximation that is excellent in the middle is wrong in the tail, arriving here as a statement about which of two bands can be trusted on a shape nobody checked.
The same three bands around a fitted line
The distinction is drawn most often in regression, where all three bands exist and two of them are routinely plotted on the same axes.
A fitted line has a confidence band for the mean response at each x, with half-width proportional to , and a prediction band for a new observation at each x, with the same expression plus 1 inside the root. The extra 1 is the population’s own variance, exactly as the 1 in is here, and it dominates everything else the moment n is past about twenty.
Two consequences follow and both are visible in any plotted pair.
The confidence band pinches at x̄ and flares at the ends, because the second term under the root vanishes at the centre of the design. The prediction band barely pinches at all: the 1 is constant and the rest is small, so it is nearly parallel to the line. A reader who has seen only the confidence band has seen a picture whose shape is an artefact of which question was asked.
And the vertical gap between them is the same factor this page is about. At a hundred observations, at the centre of the design, the prediction band is about ten times the confidence band — and the confidence band is the one that gets drawn, because it is the narrower and it looks like a statement about the fit. Four datasets with identical fits makes the neighbouring point about what a summary of a line cannot show; this one is about what a band around it is a statement about.
Three counted coverages, and one exact identity
Each band is counted against the quantity it is about, which is the only comparison that can distinguish them — a band counted against the wrong quantity would come out right or wrong for reasons unconnected to its construction.
The mean’s interval covers μ on 95.2% of twenty thousand samples at twenty observations. The prediction interval covers a fresh draw on 95.1%. The tolerance interval holds at least 95% of the population on 95.1% of samples. Three constructions, three quantities, one nominal rate, and the counts agree with it to within the simulation’s own error.
Beside them sits the one thing on this page that is not counted at all. The ratio of the first two half-widths is to nine decimal places at every setting the slider takes, because it is an identity rather than a measurement: the t and the s cancel algebraically, so a band built from a different quantile or a different standard deviation would break it immediately. That is the cheapest statement separating the two bands, and it catches a transcription error that no coverage count would — a coverage of 95% is what both bands give, so drawing the prediction interval twice would look correct in every counted number on the page.
Which band the question actually wants
The three constructions are easy to keep apart once the questions are written down, so here are the questions in the form they usually arrive in.
“Is the treatment effect different from zero?” The mean’s interval. The object is a parameter, the answer is about where that parameter is, and the band shrinks with data because the parameter becomes better known.
“What result should this patient expect?” The prediction interval. One future value, with the population’s spread in it, and no amount of data narrows it below that spread.
“What range counts as normal on the report form?” The tolerance interval. A share of the population, with a confidence attached, and the widest of the three at every sample size.
“How wrong is the reported average likely to be?” The mean’s interval again — and this is the one that is most often answered with a standard deviation, because the standard deviation is the number printed beside the mean.
The test that separates them in one step is to ask what would make the statement false. If a single future observation outside the band would falsify it, the band is a prediction interval. If only a systematic excess of observations outside it would, it is a tolerance interval. If nothing observable would — because the statement is about a quantity that is never observed — it is the mean’s interval, and its 95% is a property of the procedure rather than of the interval in front of the reader.
Still open: the band that has to hold for all of the next m
The operational test above ends on a case none of the three bands answers. A statement about the next ten observations, all of them inside is neither a prediction interval nor a tolerance interval, and it is the statement a warranty, a batch release or a monitoring rule actually makes.
The naive route is to take the prediction interval and demand ten successes at 95% each, which gives 60% and is wrong in both directions: the ten future draws are independent of each other but they are not independent given the interval, because they share the same x̄ and s. A sample that drew a small s gives a narrow band that all ten draws are likely to miss together. Conditioning is what makes simultaneous statements hard, and it is the same structure that makes a family of tests not a product of individual ones.
What such a band costs, whether it is closer to the tolerance interval than to the prediction interval, and how its width behaves as m grows — these have answers, and they are not any of the three numbers on this page.
Two things can be said about the answer in advance and neither settles it. As m grows the band has to approach a tolerance interval for the whole population, since holding all of infinitely many future draws is holding the population; so the tolerance interval is the ceiling, and the question is how fast the ceiling is reached. And at m = 1 the band is the prediction interval exactly, so the whole family runs between two of the three widths above. What is not obvious is whether the approach is quick — whether ten future observations already cost most of what all of them cost — and that is a measurement rather than an argument, which is the form every other claim here has been made to take.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name confidence interval, coverage, sample size, standard error
- An interval that carries its scale — both name confidence interval, coverage, estimated variance, standard error
- The same draws for both methods — both name confidence interval, coverage, sample size, standard error
- Twenty intervals and one expected miss — both name confidence interval, coverage, sample size, standard deviation
- What studentising costs — both name confidence interval, coverage, estimated variance, standard error
- A block size that changes — both name confidence interval, coverage, sample size
Named objects
A flat tag is an object no other essay names yet.
Confidence intervalCoverageEstimated variancePrediction intervalSample sizeStandard deviationStandard errorTolerance interval