The interval that holds observations, not a mean

All of the next ten

A warranty, a batch release or a monitoring rule promises something about every one of the next ten observations, not about one. From a sample of ten, the band that holds all ten with 95% probability reaches 3.716 sample standard deviations either side of the mean — already wider than the 3.382 of a tolerance interval for 95% of the population — and it keeps widening: 4.942 for a hundred, 6.008 for a thousand, with no ceiling. A 95% prediction interval, read as the answer, holds all ten 67.9% of the time: more than 0.95 to the tenth power, because the ten succeed and fail together.

Worth reading first: The shape, and where its mass is.

A tenth as wide, and both of them right separated three bands that share a label: the interval for the mean, the prediction interval for one future observation, and the tolerance interval for a stated share of the population. It ended on a statement none of the three makes — all of the next ten observations will be inside — and on two guesses about it. At one future value the band is the prediction interval. As the number grows it “has to approach a tolerance interval for the whole population”, which would make the tolerance interval its ceiling.

The first guess is right. The second is not, and the way it fails is the most useful thing about the band.

The factor that holds all of the next m observations with 95% probability, from a sample of 10From 10 observations, the band for one future value has factor 2.371, for ten 3.716, for a hundred 4.942 and for a thousand 6.008. The 95%-content tolerance factor is 3.382 and is passed by m = 10; the Bonferroni stretch of the prediction factor reaches 7.567 at a thousand.1101001,00002468future observations the band must holdfactor k in x̄ ± k stolerance, 95.0% content: 3.38tolerance, 99.0% content: 4.44tolerance, 99.9% content: 5.68all m held, 95%Bonferroniexact expectation over the sample's mean and spreadno finite band holds the population
Fig. 1 From a normal sample of ten, the factor k — the band’s half-width in sample standard deviations — that holds all of the next m observations with 95% probability, against m on a logarithmic axis. The dashed curve is the Bonferroni stretch of the prediction factor; the horizontal rules are tolerance factors for 95%, 99% and 99.9% of the population. The slider changes the sample size.

The statement being asked for

A batch of ten units will be released if every one of them meets a specification; a warranty covers the next ten installations; a monitoring rule raises no alarm if the next ten readings stay inside a band drawn from the last ten. Each is a statement about the maximum deviation among ten future values, not about any single one.

From a sample of nn with mean xˉ\bar x and standard deviation ss, the band is xˉ±ks\bar x \pm k s, and the question is the kk at which

E[(Φ ⁣(zˉ+ks)Φ ⁣(zˉks))m]  =  0.95E\left[\left(\Phi\!\left(\bar z + k s\right) - \Phi\!\left(\bar z - k s\right)\right)^{m}\right] \;=\; 0.95

in units of the population’s spread, the expectation taken over the sample’s own mean and spread. Given the sample, each future value lands inside independently with probability equal to the band’s content, so all mm land inside with probability equal to the content raised to the mm-th power; averaging over samples gives the chance the band holds them all. There is no closed form, but the expectation is a two-dimensional integral over a normal and a χ2\chi^2, and it can be computed to four decimals and checked against a count.

The factor, as the promise grows

From a sample of ten:

future observations held factor k Bonferroni stretch
1 2.371 2.373
2 2.783 2.816
5 3.319 3.408
10 3.716 3.870
20 4.102 4.348
50 4.591 5.014
100 4.942 5.549
1,000 6.008 7.567

At one future value the factor is the prediction interval’s, t0.975,91+1/10t_{0.975,\,9}\sqrt{1 + 1/10} = 2.373, to the precision of the integration. Holding two future values costs 17% more width; ten cost 57% more; a hundred more than double it.

The tolerance factor for 95% of the population, at 95% confidence, is 3.382 from a sample of ten — and the band for all of the next ten is already wider, at 3.716. The band for five is just inside it and the band for ten is past it. The tolerance factor for 99% of the population, 4.445, is passed between twenty and fifty future values; the one for 99.9%, 5.678, between a hundred and a thousand. None of them is a ceiling.

Why the tolerance interval is not the limit

The guess that the band approaches a tolerance interval rests on an intuition: holding all of infinitely many future draws is holding the population, and a tolerance interval is a statement about the population. The flaw is in “holding the population”. A tolerance interval holds a stated share of it — 95%, 99%, 99.9% — and a normal population has no finite interval holding all of it. Holding every one of mm future draws means holding the largest of them, and the largest of mm normal draws grows without bound, roughly as 2lnm\sqrt{2\ln m} standard deviations.

So the right comparison for the band for mm values is not a tolerance interval but a tolerance interval whose content grows with mm: to hold all of mm draws with high probability, the band has to hold roughly all but a fraction 1/m1/m of the population. That is why it passes the 95%-content factor between five and ten future values, the 99%-content factor between twenty and fifty, and the 99.9%-content factor between a hundred and a thousand. The content the band implicitly promises is set by how many values it must hold.

The practical lesson follows directly. A specification that promises “all units in the batch” is a promise about a tail whose depth depends on the batch size, and a band drawn at a fixed factor — two standard deviations, or a tolerance factor looked up for 95% — holds a small batch and fails a large one, with no change in the process.

The factor that holds all of the next m observations with 95% probability, from a sample of 100. From 100 observations, the band for one future value has factor 1.994, for ten 2.874, for a hundred 3.601 and for a thousand 4.238. The 95%-content tolerance factor is 2.233 and is passed by m = 2; the Bonferroni stretch of the prediction factor reaches 4.263 at a thousand.
Fig. 2 The same factors from a sample of a hundred. Every curve is lower, because the sample’s mean and spread are better known, and the tolerance factor for 95% of the population is passed by the band for only two future values.

From a sample of a hundred the factors are 1.994 for one future value, 2.874 for ten, 3.601 for a hundred and 4.238 for a thousand, and the 95%-content tolerance factor, 2.233, is already passed at two. A larger sample lowers every factor, because less of the width is spent on not knowing the population’s mean and spread, and it slows the growth with mm — from one future value to a thousand the factor rises by 2.24 rather than 3.64 — but it cannot stop it, because the part of the growth that belongs to the population’s tail is there however well the sample is known.

A fourth band with the same label

The three bands the earlier essay separated can be set side by side with the fourth.

Three bands called 95%, at n = 20. Half-widths in sample standard deviations: 0.468 for the mean, 2.14 for one future observation, 2.75 to hold 95% of the population. The first two differ by exactly the square root of n + 1, which is 4.58 here.
Fig. 3 Three 95% bands from a normal sample of twenty, in sample standard deviations: the interval for the mean, the prediction interval for one future observation, and the tolerance interval for 95% of the population.

From a sample of twenty the interval for the mean has a half-width of 0.468 standard deviations, the prediction interval 2.14 and the tolerance interval 2.75. The band for all of the next ten is 3.206 — wider than all three, and the widest band that anybody would still describe, loosely, as “where the data will be”. Each of the four is correct for its own statement, and the four statements are progressively stronger: about the mean, about one future value, about most of the population, about every one of a stated number of future values.

The ordering is not fixed. The band for ten future values is wider than the 95% tolerance band at samples of ten and twenty; the band for two is narrower than it at a sample of ten and wider at a sample of a hundred. What a label cannot convey is which statement the band makes, and the widths differ by a factor of seven between the narrowest and the widest from the same twenty observations.

The ten fail together

The obvious shortcut is to take the 95% prediction interval, which holds one future value 95% of the time, and argue that it holds ten with probability 0.95100.95^{10} = 59.9%. That treats the ten future values as independent trials, and they are not.

How often a 95% prediction interval from 10 observations holds all of the next m. The interval holds one future value 95% of the time. It holds all of the next ten 67.9% of the time, against the 59.9% that 0.95 to the tenth power gives, because the future values share the sample's mean and spread and succeed or fail together.
Fig. 4 How often a 95% prediction interval from a sample of ten holds all of the next m observations, counted over the sample’s mean and spread, beside the product 0.95 to the power m.

Given the sample, the future values are independent. Across samples they are not, because they all share the same band. A sample that happened to have a small standard deviation produces a narrow band that many future values will miss together; one with a large standard deviation produces a wide band that all of them clear together. Averaged over samples, all ten are held 67.9% of the time — more than the 59.9% the independent product predicts, because the misses cluster in the samples that drew the narrow bands.

The effect depends on how uncertain the band is. From a sample of five, the prediction interval holds all of the next ten 75.1% of the time; from a sample of a hundred, 60.7%, almost the independent product, because a band estimated from a hundred observations barely varies from sample to sample. The dependence is created entirely by estimating the band, and it disappears as the estimate becomes exact.

This is the structure five times in six found in replications of one study, which miss together because they share one original interval, and that false discoveries that arrive together found in correlated tests. A family of statements that share an estimated quantity has a distribution of failures much wider than independence predicts: more often all right, and more often several wrong at once.

How the sample size enters

Every factor above has two parts: the spread of the population, which the band must cover whatever the sample, and the uncertainty about the population’s mean and spread, which the band must also cover because it is built from estimates.

The factor a band needs to hold 95% of the population, 95% of the time. At ten observations it is 3.382 sample standard deviations, against the 1.96 an interval for the mean uses. The two only converge in the hundreds: at n = 300 the factor is still 2.106.
Fig. 5 The tolerance factor for 95% of a normal population at 95% confidence, against the sample size. It falls steeply at small samples and slowly afterwards, and is still above the known-parameter value of 1.96 at three hundred observations.

The tolerance factor’s curve shows the second part fading. At ten observations the factor is 3.382, at three hundred 2.106, and only in the limit does it reach 1.96, the value for a known population. The band for all of mm future values has the same second part and a first part that grows with mm, so its factor falls with the sample size towards a limit that is itself rising with mm: for a known normal population and a thousand future values, the band that holds all of them with 95% probability is 4.050 standard deviations wide on each side, and no sample size brings the factor below that.

That separation is the useful way to read the table. At a sample of ten most of the factor for a few future values is the cost of estimation, and collecting more data is the remedy; at a sample of a hundred and a thousand future values most of it is the population’s own tail, and no amount of data removes it. A specification that is too wide can be narrowed by measuring more when the first part dominates and cannot be when the second does.

Why Bonferroni overshoots

The Bonferroni stretch in the table splits the 5% evenly across the mm future values and builds a prediction interval at level 10.05/m1 - 0.05/m for each. It guarantees at least 95% for all of them, and the table shows how much it overpays: at ten future values its factor is 3.870 against the exact 3.716, and at a thousand 7.567 against 6.008 — a band a quarter wider than it needs to be.

The overpayment has the same source as the clustering. Bonferroni’s inequality is sharp when the events it adds up rarely happen together, and misses of a shared estimated band happen together far more often than that. The more the future values share — the smaller the sample — the more of Bonferroni’s allowance is spent on overlaps that do not need paying for. At a sample of a hundred the band is nearly known, the misses nearly independent, and Bonferroni’s factor at a thousand future values, 4.263, is within 1% of the exact 4.238.

The same diagnosis applies to a correction for many analyses of correlated data: Bonferroni is exact for independent comparisons and increasingly wasteful as they share more.

What such a band assumes

Every number here is for a normal population, and the band for many future values is more exposed to that assumption than any band before it in this family of bands. Two standard deviations of what and the tolerance intervals after it depended on the population’s shape in its shoulders; a band for a thousand future values depends on its shape six standard deviations out, where no sample of ten or a hundred has any observations — the predicament of a level with no data in it, read for a band instead of a return level.

That is the tail converging last in its sharpest form. The factor of 6.008 for a thousand future values is a statement about a normal tail, and a population whose tail is even slightly heavier than normal — a mixture, an occasional different mechanism — needs a much wider band, which nothing in the sample can reveal. Ninety-three observations and nothing assumed gave the distribution-free alternative for a stated share of the population; for all of mm future values the distribution-free answer is the range of a sample large enough that its extremes reach that deep, and the sample it needs grows with mm rather than with any fixed content.

What to use

For a statement about the next few observations, compute the factor for all of them rather than borrowing a prediction interval or a tolerance factor. For five or fewer from a sample of ten the prediction interval stretched by Bonferroni is within 3% and is a reasonable approximation; beyond that the exact factor is noticeably narrower.

For a statement about an open-ended stream — every future batch, every future reading — no band from a finite sample can hold all of them, and the statement has to be reframed as a rate: a tolerance interval for a stated share, with the expected number of exceedances made explicit.

For any band whose promise reaches far into the tail, state the distributional assumption as the thing being relied on, because at those depths it is the whole of the band’s width.

The same arithmetic in a control chart

Process monitoring runs into the growth with mm every day, usually without naming it. A control chart draws limits at three standard deviations either side of a process mean and raises an alarm when a reading falls outside. Each reading from a stable process falls outside with probability 0.27%, which sounds negligible and is negligible for any one reading taken alone.

Over a thousand readings from a process that never changes, the chance that no reading falls outside three-sigma limits — even with the process mean and spread known exactly — is 6.7%. A stable process run for a thousand readings is expected to raise a false alarm, and does so more than nine times in ten. That is the band for all of the next thousand values read the other way: the limits are a band at a factor of three, and holding a thousand values with 95% probability needs a factor of 4.050.

Control-chart practice deals with this by reading the chart as a rate — an average run length between false alarms — rather than as a promise about every reading, which is exactly the reframing recommended above for open-ended streams. The mistake it avoids is the one this essay began with: reading a band built for one value at a time as though it said something about all of them together.

What the integrals establish, and what they do not

For one future observation the factor is the prediction interval’s, t1+1/nt\sqrt{1 + 1/n}, to within two thousandths at a sample of ten; counting samples and twenty future values at the computed factor holds all twenty 95% of the time, within a point over twenty thousand samples; and for a thousand future observations the factor has passed the 95%-content tolerance factor.

The factor rises with every increase in m and never exceeds the Bonferroni stretch, at both sample sizes drawn.

What does not survive is a 95% prediction interval read as holding all of the next ten observations. From a sample of ten it holds all ten 67.9% of the time.

Not claimed: that the factors are right for any population but the normal. And not claimed that the growth with mm follows 2lnm\sqrt{2\ln m} closely at these sizes — the factor at a thousand includes the sample’s uncertainty about the mean and spread, which inflates it well above the known-parameter value, and the asymptotic rate describes only how the known-parameter part grows.

Still open: a band that is allowed to be wrong about a few

A promise about every one of mm future values becomes unaffordable quickly and unattainable in the limit. The version most applications can actually use is weaker: at most rr of the next mm outside the band. Its factor sits between the prediction interval’s and the all-of-mm band’s, and for large mm it approaches a tolerance factor whose content is 1r/m1 - r/m — so the tolerance interval does become the limit, for a promise about a fixed share of failures rather than about none.

How quickly that happens, how much width allowing a single failure among a hundred saves, and how the clustering of misses changes the answer — since a narrow band tends to fail by several at once, allowing one failure buys less than it seems — are computable by the same integral with a binomial in place of a power, and have not been computed.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BonferroniEstimated varianceExtreme valueNormal distributionPrediction intervalSample sizeSimultaneous inferenceTolerance interval