The interval that holds observations, not a mean

A tenth as wide, and both of them right

The interval for a mean and the interval for one future observation are both labelled 95%, and at a hundred observations one is 10.05 times the other — exactly the square root of n + 1. Read the narrow one as the wide one and it covers a new value 15.7% of the time.

Worth reading first: The shape, and where its mass is · What the 95% refers to.

Three bands can be drawn around the same sample mean. All three are written “95%”. At twenty observations their half-widths are 0.468, 2.145 and 2.752 sample standard deviations, and every one of them covers what it claims to cover at very nearly exactly 95%.

Three bands called 95%, at n = 20Half-widths in sample standard deviations: 0.468 for the mean, 2.14 for one future observation, 2.75 to hold 95% of the population. The first two differ by exactly the square root of n + 1, which is 4.58 here.where the mean is0.468 scovers it 95.2%where the next value is2.145 scovers it 94.9%where 95% of the population is2.752 sholds it 95.1%the mean's band is 1/4.58 of the observation'shalf-widths in sample standard deviationsn = 20
Fig. 1 The three half-widths at twenty observations, each beside the counted rate at which it covers its own quantity. Nothing here is a correction to anything else: they are three answers to three questions.

The three questions are different and only one of them is usually asked out loud.

Where is the mean? — the parameter, a fixed unknown number. Half-width ts/nt\,s/\sqrt{n}.

Where will the next observation be? — one future draw, which has the population’s own spread in it as well as the uncertainty about where the centre is. Half-width t·s·1+1/n\sqrt{1 + 1/n}.

Where is 95% of the population? — a region, and a promise to have covered that region. Half-width k·s, with k the tolerance factor.

The ratio is exact and has one term in it

The first two bands differ by

ts1+1/nts/n  =  n+1,\frac{t\,s\sqrt{1 + 1/n}}{t\,s/\sqrt{n}} \;=\; \sqrt{n+1},

with the t and the s cancelling entirely. The ratio does not depend on the confidence level, on the population’s spread, or on anything estimated. It is a function of the sample size and nothing else, and at the sizes people work at it is large: 3.32 at ten observations, 4.58 at twenty, 10.05 at a hundred, 20.03 at four hundred.

That is the wrong direction for intuition, and it is worth saying plainly because it inverts what “more data” is expected to do. Collecting more data makes the two bands further apart, not closer. The mean’s interval shrinks like 1/n1/\sqrt{n} and the prediction interval does not shrink at all — it converges to the population’s own 95% band, which is a fixed width that no amount of data reduces. At four hundred observations the prediction half-width is 1.968 standard deviations, and the population’s is 1.960.

Three bands called 95%, at n = 4. Half-widths in sample standard deviations: 1.591 for the mean, 3.56 for one future observation, 6.40 to hold 95% of the population. The first two differ by exactly the square root of n + 1, which is 2.24 here.
Fig. 2 Four observations, where the two bands are closest: the mean’s is 1.591 standard deviations wide and a new draw’s is 3.558, a factor of 2.24. This is as similar as they ever are.
Three bands called 95%, at n = 400. Half-widths in sample standard deviations: 0.098 for the mean, 1.97 for one future observation, 2.08 to hold 95% of the population. The first two differ by exactly the square root of n + 1, which is 20.02 here.
Fig. 3 Four hundred observations, where they are twenty times apart. The wide band has barely moved since the frame above; the narrow one has shrunk by a factor of sixteen.

What each one costs the other way

Reading the wrong band is not a rounding error, and the size of the mistake is countable.

The mean’s interval read as a prediction. At ten observations it covers a new draw 48.6% of the time. At a hundred, 15.7%. At four hundred, 7.9%. The claim gets worse with more data, for exactly the reason above: the band the reader is looking at is shrinking and the thing it is being read about is not.

That is the figure worth carrying, because it is the direction nobody expects. A large study reporting its mean with a tight error bar is reporting a tight statement about the mean. If the bar is read as “where the values are”, the reading is less accurate than it would have been from a small study — a hundred observations gives a reading that is wrong six times in seven.

The prediction interval read as a statement about the mean. It covers μ essentially always — at every sample size measured it covered on 100.00% of twenty thousand samples. That failure is not symmetric with the first and it is not harmless. An interval that covers 100% of the time when it claims 95% is not conservative in any useful sense; it is an interval that has stopped distinguishing between hypotheses, so any test built from it has no power at all.

The mean’s interval read as a tolerance interval. This is the first failure again, in the form it usually takes. A reference range, a specification, a “typical values” statement — all of them are questions about the population, and answering them with a standard error gives a band an order of magnitude too narrow.

The notation that causes it

The standard error is the standard deviation divided by n\sqrt{n}, which means the two are the same symbol divided by a number, and the number is usually not printed.

symbol what it is the spread of
s the data
s/ns/\sqrt{n} the sample mean

At twenty-five observations the second is a fifth of the first. At a hundred it is a tenth. A figure whose error bars are one standard error and a figure whose error bars are one standard deviation are different figures by a factor of ten, and they look identical.

The shape’s own essay makes the same point about ±1σ against 95%, and the two errors compound: a bar drawn at one standard error, read as two standard deviations of the data, is out by a factor of twenty at a hundred observations. The habit that prevents it is one sentence — say what the bar is — and the reason it needs saying is that no convention exists across fields.

What x-bar plus or minus 2 sample standard deviations holds, at n = 20. The content of the band is a random variable. Across 20,000 normal samples of 20 it averages 93.5%, its fifth percentile is 84.6%, and it falls short of 95% on 55.0% of samples. The band that is drawn to show where 95% of the data lies.
Fig. 4 Why the third band is a different kind of object. The share of the population inside a band drawn from the sample is itself random, so a promise about it needs a confidence as well as a coverage.

Why the third band is not the second one widened

The prediction interval and the tolerance interval are both about the data rather than the parameter, so it is natural to read the tolerance interval as a slightly wider prediction interval. It is a different construction, and the difference matters when n is small.

A prediction interval covers one future observation with probability 95%, averaged over samples and over the future observation together. A tolerance interval covers 95% of the population with probability 95%, and those two probabilities are over different things: the inner one is over the population given the sample, the outer one is over samples.

At twenty observations the prediction half-width is 2.145 and the tolerance half-width is 2.752 — 28% wider. At four hundred they are 1.968 and 2.084, 6% apart. They converge, because both are heading for the population’s own 1.96, but they converge at different rates: the prediction interval is already within half a per cent of its limit at four hundred, and the tolerance interval is still 6% above it, because the confidence half of its promise never becomes free.

The distinction has a clean operational test. Ask how many future observations the statement is about. One: prediction interval. A stated share of all of them: tolerance interval. All of the next m: neither, and the right band is wider than both and depends on m.

What the t is doing in two of the three

Both the mean’s interval and the prediction interval carry the same t quantile, which looks like an accident of notation and is not.

The mean’s interval needs t because (x̄ − μ)/(s/ns/\sqrt{n}) is a t: a normal over an independent χ2\chi^2 root. The prediction interval needs t because (xnewxˉ)/(s1+1/n)(x_{\text{new}} - \bar{x})/\bigl(s\sqrt{1 + 1/n}\bigr) is also a t, for the same reason — the numerator is normal with variance σ2(1+1/n)\sigma^2(1 + 1/n), the denominator is the same independent s, and the ratio standardises to the same distribution on the same n − 1 degrees of freedom.

Student's t on 19 degrees of freedom, against the normal. The two-sided 95% critical value is 2.093 for t(19) and 1.960 for the normal — 7% wider. Using the normal at this sample size makes every interval too short by that much.
Fig. 5 The distribution both bands standardise to. The same nineteen degrees of freedom, because the same sample standard deviation is in both denominators.

That is why the two bands share a t and why their ratio has no t in it. The tolerance interval does not standardise to anything, which is why its factor has no closed form: its event is about a difference of two normal integrals rather than about a single standardised quantity.

The shared t also explains a small fact that surprises people. The prediction interval’s factor approaches 1.96 from above and does so slowly at small n — 2.373 at ten observations against 1.96 — and all of that excess is the t rather than the 1+1/n\sqrt{1 + 1/n}. At ten observations 1+1/10\sqrt{1 + 1/10} = 1.049, which is 5% of the width, and t(9) = 2.262 against 1.96, which is 15%. So the prediction interval at small samples is mostly a statement about not knowing σ, and only slightly a statement about the extra draw.

Only one of the three is protected by the limit theorem

Everything so far assumes a normal population. Dropping that assumption separates the three bands again, and it separates them along a line nobody draws in advance.

The mean’s interval is protected. x̄ is an average, so the limit theorem applies to it and its distribution becomes normal whatever the population was. Drawing from an exponential — one-sided, skewed, nothing like a bell — the 95% t interval for the mean covers 89.8% at ten observations, 92.9% at thirty, 94.1% at a hundred and 94.8% at four hundred. It is converging on its label, slowly and reliably, and the convergence is the theorem doing its work.

The prediction interval is not protected and does not converge. On the same draws it covers 93.1%, 94.2%, 94.6% and 94.5% — and it stops there. No amount of data helps, because the quantity it is about is one observation and nothing about one observation is an average of anything. The population’s own shape is the thing being predicted, and the interval is built out of a mean and a standard deviation, which do not describe that shape.

The total is not the reading. Split the misses by side and at a hundred observations the interval misses 0.0% below and 5.4% above. Not 2.5% and 2.5%: zero and more than five. The lower limit xˉts1+1/n\bar{x} - t\,s\sqrt{1 + 1/n} sits below zero, where an exponential has no mass at all, so it cannot miss low — and every miss the interval is entitled to arrives on the upper side, at twice the rate a reader would price. The two errors are of opposite sign, the sum reads 94.6%, and a coverage table with one column would report that as a success.

That is the general shape and it is worth stating in one line. Averaging is what the limit theorem buys, and a prediction is not an average. A band about a parameter inherits the theorem; a band about a value inherits the population, including everything about the population that the mean and the standard deviation cannot see. The same asymmetry is why an approximation that is excellent in the middle is wrong in the tail, arriving here as a statement about which of two bands can be trusted on a shape nobody checked.

The same three bands around a fitted line

The distinction is drawn most often in regression, where all three bands exist and two of them are routinely plotted on the same axes.

A fitted line has a confidence band for the mean response at each x, with half-width proportional to 1/n+(xxˉ)2/(xxˉ)2\sqrt{1/n + (x - \bar{x})^2 / \sum (x - \bar{x})^2}, and a prediction band for a new observation at each x, with the same expression plus 1 inside the root. The extra 1 is the population’s own variance, exactly as the 1 in 1+1/n\sqrt{1 + 1/n} is here, and it dominates everything else the moment n is past about twenty.

Two consequences follow and both are visible in any plotted pair.

The confidence band pinches at x̄ and flares at the ends, because the second term under the root vanishes at the centre of the design. The prediction band barely pinches at all: the 1 is constant and the rest is small, so it is nearly parallel to the line. A reader who has seen only the confidence band has seen a picture whose shape is an artefact of which question was asked.

And the vertical gap between them is the same factor this page is about. At a hundred observations, at the centre of the design, the prediction band is about ten times the confidence band — and the confidence band is the one that gets drawn, because it is the narrower and it looks like a statement about the fit. Four datasets with identical fits makes the neighbouring point about what a summary of a line cannot show; this one is about what a band around it is a statement about.

Three counted coverages, and one exact identity

Each band is counted against the quantity it is about, which is the only comparison that can distinguish them — a band counted against the wrong quantity would come out right or wrong for reasons unconnected to its construction.

The mean’s interval covers μ on 95.2% of twenty thousand samples at twenty observations. The prediction interval covers a fresh draw on 95.1%. The tolerance interval holds at least 95% of the population on 95.1% of samples. Three constructions, three quantities, one nominal rate, and the counts agree with it to within the simulation’s own error.

Beside them sits the one thing on this page that is not counted at all. The ratio of the first two half-widths is n+1\sqrt{n + 1} to nine decimal places at every setting the slider takes, because it is an identity rather than a measurement: the t and the s cancel algebraically, so a band built from a different quantile or a different standard deviation would break it immediately. That is the cheapest statement separating the two bands, and it catches a transcription error that no coverage count would — a coverage of 95% is what both bands give, so drawing the prediction interval twice would look correct in every counted number on the page.

Twenty 95% intervals for a proportion that really is 0.35. 2 of the twenty miss the true value. The 95% is a property of the procedure across repetitions — no single interval has a 95% chance of anything, because it either contains 0.35 or it does not.
Fig. 6 What the counting means, for the first of the three. Twenty intervals from twenty samples, of which the ones that miss are not mistakes.

Which band the question actually wants

The three constructions are easy to keep apart once the questions are written down, so here are the questions in the form they usually arrive in.

“Is the treatment effect different from zero?” The mean’s interval. The object is a parameter, the answer is about where that parameter is, and the band shrinks with data because the parameter becomes better known.

“What result should this patient expect?” The prediction interval. One future value, with the population’s spread in it, and no amount of data narrows it below that spread.

“What range counts as normal on the report form?” The tolerance interval. A share of the population, with a confidence attached, and the widest of the three at every sample size.

“How wrong is the reported average likely to be?” The mean’s interval again — and this is the one that is most often answered with a standard deviation, because the standard deviation is the number printed beside the mean.

The test that separates them in one step is to ask what would make the statement false. If a single future observation outside the band would falsify it, the band is a prediction interval. If only a systematic excess of observations outside it would, it is a tolerance interval. If nothing observable would — because the statement is about a quantity that is never observed — it is the mean’s interval, and its 95% is a property of the procedure rather than of the interval in front of the reader.

Still open: the band that has to hold for all of the next m

The operational test above ends on a case none of the three bands answers. A statement about the next ten observations, all of them inside is neither a prediction interval nor a tolerance interval, and it is the statement a warranty, a batch release or a monitoring rule actually makes.

The naive route is to take the prediction interval and demand ten successes at 95% each, which gives 60% and is wrong in both directions: the ten future draws are independent of each other but they are not independent given the interval, because they share the same x̄ and s. A sample that drew a small s gives a narrow band that all ten draws are likely to miss together. Conditioning is what makes simultaneous statements hard, and it is the same structure that makes a family of tests not a product of individual ones.

What such a band costs, whether it is closer to the tolerance interval than to the prediction interval, and how its width behaves as m grows — these have answers, and they are not any of the three numbers on this page.

Two things can be said about the answer in advance and neither settles it. As m grows the band has to approach a tolerance interval for the whole population, since holding all of infinitely many future draws is holding the population; so the tolerance interval is the ceiling, and the question is how fast the ceiling is reached. And at m = 1 the band is the prediction interval exactly, so the whole family runs between two of the three widths above. What is not obvious is whether the approach is quick — whether ten future observations already cost most of what all of them cost — and that is a measurement rather than an argument, which is the form every other claim here has been made to take.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Confidence intervalCoverageEstimated variancePrediction intervalSample sizeStandard deviationStandard errorTolerance interval