The interval that holds observations, not a mean

A band allowed a few misses

From ten observations, a band holding every one of the next hundred with 95% probability reaches 4.942 sample standard deviations either side of the mean. Allowing one of the hundred outside brings it to 4.289; allowing five, to 3.378 — the 95%-content tolerance factor, 3.382, to within half a hundredth. The tolerance interval turns out to be the limit of a promise about a share of failures. And the misses arrive together: when this band misses once, it misses again 45.4% of the time, where independent misses at the same rate would do so 12.5%.

Worth reading first: The shape, and where its mass is.

All of the next ten priced the promise a warranty, a batch release or a monitoring rule actually makes: not that one future observation will fall inside a band, but that every one of the next ten will. From a sample of ten, that band reaches 3.716 sample standard deviations either side of the mean, wider than the tolerance interval for 95% content, and the factor grows without limit as the number of future observations does. No finite band holds a normal population forever.

It ended on the weaker promise most applications can actually keep: at most rr of the next mm outside the band. A batch of a hundred units may be released with one allowed out of specification; a monitoring rule may accept a few false alarms a year. Its factor should sit between the prediction interval’s and the all-of-mm band’s, approaching a tolerance factor whose content is 1−r/m1 - r/m as mm grows. How quickly, how much width one allowed failure saves, and whether the misses of a narrow band cluster so that the allowance buys less than it seems, were left to compute.

The band that allows at most r of the next m observations outside it, 95% of the time, from a sample of ten. At a hundred future observations the band that holds all of them has factor 4.942; allowing one outside, 4.289; two, 3.948; five, 3.378. The tolerance factors for 99% and 95% content are 4.445 and 3.382.
Fig. 1 The factor kk in xˉ±ks\bar x \pm k s that, from ten observations, has at most rr of the next mm outside it with 95% probability, against mm on a logarithmic scale, for none, one, two and five allowed outside. The dotted lines are the tolerance factors for 99% and 95% content.

Where the all-of-m band left off

The factor that holds all of the next m observations with 95% probability, from a sample of 10. From 10 observations, the band for one future value has factor 2.371, for ten 3.716, for a hundred 4.942 and for a thousand 6.008. The 95%-content tolerance factor is 3.382 and is passed by m = 10; the Bonferroni stretch of the prediction factor reaches 7.567 at a thousand.
Fig. 2 The factor that holds all of the next mm observations with 95% probability, from a sample of ten, beside the tolerance factors: it passes the 95%-content tolerance factor at a handful of future observations and keeps rising.

The all-of-mm band is the upper edge of the family measured here, and its behaviour frames everything below it. It rises without limit — at a thousand future observations it is 6.008 — because a normal population has no edge: the largest of mm draws keeps growing, roughly as the square root of twice the logarithm of mm, and a band that must hold the largest has to grow with it. The shape and where its mass is is the reason: the 99.7% within three standard deviations leaves three in a thousand outside, and a thousand draws will usually include some of them.

A few-misses band escapes that growth only partly. Allowing a fixed number of misses — one, two, five — still leaves a band that must contain the m−rm - r smallest-magnitude draws, and those also grow with mm, more slowly. Allowing a fixed share escapes it completely, because the share of a normal population beyond any fixed point is fixed. That is the difference between the two families of curves in the hero figure and the flat line in the next one.

The same integral, with a binomial in it

A sample of ten gives a mean xˉ\bar x and a spread ss, and a band xˉ±ks\bar x \pm k s misses any one future observation with a probability qq that depends on how far xˉ\bar x and ss happened to land from the truth. Given the sample, the future observations are independent, so the number of the next mm outside the band is binomial with that qq. The chance of at most rr outside is that binomial’s distribution function, averaged over the sample’s own sampling distribution — the same average over xˉ\bar x and ss that gave the all-of-mm band, with the binomial in place of (1−q)m(1 - q)^m. At r=0r = 0 the two are identical, which ties the new calculation to the old.

The factor that gives 95% probability is found by bisection on kk.

The family runs from one familiar interval to another. With m=1m = 1 and r=0r = 0 it is the prediction interval for one future observation, t1+1/nt\sqrt{1 + 1/n}, 2.373 from ten observations; with r=0r = 0 and mm growing it is the all-of-mm band; and with r/mr/m held fixed and mm growing it is the tolerance interval, as the next sections show. Every band a reader is likely to be offered for future observations is a member of it, and the only thing that distinguishes them is the promise — how many future values, and how many of them are allowed to fall outside.

What one allowed miss saves

At a hundred future observations, from a sample of ten:

allowed outside, of the next hundred factor kk
none 4.942
at most one 4.289
at most two 3.948
at most five 3.378

Allowing one miss in a hundred narrows the band by 13%, from 4.942 sample standard deviations to 4.289; allowing five narrows it by a third. For a batch-release rule that is a large difference in the rejection rate of good batches, and for a monitoring rule a large difference in how early a real shift can be seen, bought by a promise that is only slightly weaker.

At ten future observations the savings are larger in proportion: from 3.716 for none outside to 2.807 for one and 1.341 for five — the last being narrower than the 95% prediction interval for a single observation, 2.373, since a band allowed to miss half of ten can be very narrow indeed. At a thousand the factors are 6.008, 5.489, 5.224 and 4.796, and the absolute saving from one allowed miss is between half and two thirds of a standard deviation at a hundred and at a thousand future observations.

The limit is a tolerance interval

The dotted lines in the hero figure are tolerance factors: the kk for which xˉ±ks\bar x \pm k s contains at least a stated share of the population with 95% confidence. At a hundred future observations the band allowing five outside is 3.378 and the 95%-content tolerance factor is 3.382. The band allowing one outside is 4.289 against the 99%-content factor’s 4.445; two outside, 3.948 against 98% content’s 4.014.

The agreement is not a coincidence. A band that holds all but rr of mm future draws is a band whose miss rate qq is at most about r/mr/m, give or take the binomial’s scatter, and as mm grows the scatter vanishes and the promise becomes exactly “the band’s miss rate is at most r/mr/m”, which is a promise about the band’s content: a tolerance interval with content 1−r/m1 - r/m.

The band that allows at most one in twenty of the next m outside, against m, from a sample of ten. The factor is 3.290 at 20, 3.348 at 40, 3.378 at 100, 3.386 at 200, 3.390 at 400, 3.392 at 1000. The tolerance factor for 95% content with 95% confidence is 3.382.
Fig. 3 The factor for at most one in twenty of the next mm outside, against mm, beside the 95%-content tolerance factor it converges to.

Stated as a fixed share — at most one in twenty outside — the factor runs 3.290 at twenty future observations, 3.348 at forty, 3.378 at a hundred and 3.392 at a thousand, around the tolerance factor of 3.382. It crosses it between a hundred and two hundred and settles on it from above, since at large mm the binomial’s own scatter, which the tolerance interval ignores, is a small extra risk the band has to cover. By a hundred future observations the two are equal to two decimals.

That answers the question the essay on the next ten left: the tolerance interval is the limit of a promise about a fixed share of failures, not of a promise about none. A promise of none has no finite limit; a promise of a share has one, and it is reached quickly. For an application whose real requirement is “no more than one in twenty out of specification over a long run”, the tolerance interval is the right band and was the right band from around a hundred future units onwards.

Misses that arrive together

The allowance buys less than its arithmetic suggests, for a reason that belongs to the sample rather than the future.

How the misses of a band from 10 observations arrive: at most one of the next hundred allowed outsideAt the factor 4.289, which allows one of the next hundred outside 95% of the time, the band misses at least once with probability 11.0% and, having missed once, misses again 45.4% of the time. Independent misses at the same average rate, 0.263 per hundred, would give 23.2% and 12.5%.at least one of 100 outside — this band11.0%independent misses, same average23.2%a second, given a first — this band45.4%independent misses, same average12.5%exact mixture over the sample's mean and spreada narrow band misses in bunches
Fig. 4 At the band that allows one of the next hundred outside, the chance of at least one miss and the chance of a second given a first, beside what independent misses at the same average rate would give. The slider sets the sample size.

At the factor 4.289, the band from ten observations misses on average 0.263 of the next hundred. If those misses were independent, at least one would occur 23.2% of the time and a second, given a first, 12.5% of the time. They are not independent. At least one miss occurs 11.0% of the time — less often — and once one has occurred, a second occurs 45.4% of the time, nearly four times as often as independence would suggest.

The reason is the sample. A sample of ten whose spread came out small produces a narrow band, and a narrow band misses often, for every future observation at once; a sample whose spread came out large produces a wide band that almost never misses. The misses are independent given the sample and strongly dependent across samples, so they arrive in bunches: most samples produce bands with no misses at all, and the few that produce narrow bands produce several. The one allowed miss is spent, in most of the samples that need it, on the first of a cluster.

That is why the factor for “at most one” is not much narrower than it would have to be for a cluster, and why allowing two or five buys more per allowed miss than allowing one does: each additional allowance absorbs more of a cluster. It is also the same mechanism that made the prediction interval hold all of the next ten more often than 0.95 to the tenth power: the future observations share a sample, and what they share is where the band sits.

How much of each factor is the sample

With the population’s mean and spread known exactly, every factor here would be much smaller, and the difference is the price of estimating them from ten observations.

allowed outside, of the next hundred factor from ten observations factor with the population known
none 4.942 3.474
at most one 4.289 2.914
at most two 3.948 2.643
at most five 3.378 2.220

The sample adds between 1.2 and 1.5 standard deviations to every factor, and the addition is largest in proportion where the promise is weakest: a sample of ten makes the five-miss band 52% wider than it would be with the population known, and the no-miss band 42% wider. A weak promise needs a narrow band, and a narrow band is the one most exposed to a spread estimated too small.

The known-parameter factors also show what the allowance saves when the sample is not in the way. With the population known, allowing one of a hundred outside saves 0.56 standard deviations and allowing five saves 1.25 — close to the savings from ten observations, 0.65 and 1.56. The saving is mostly a property of the promise and only a little of the sample; the level of every factor is mostly the sample.

A larger sample, and less clustering

How the misses of a band from 30 observations arrive: at most one of the next hundred allowed outside. At the factor 3.344, which allows one of the next hundred outside 95% of the time, the band misses at least once with probability 18.4% and, having missed once, misses again 27.1% of the time. Independent misses at the same average rate, 0.262 per hundred, would give 23.1% and 12.4%.
Fig. 5 The same comparison from a sample of thirty. The band’s misses still cluster, and less: a second miss follows a first a little more than twice as often as independent misses would make it.

From thirty observations, the band that allows one of the next hundred outside has factor 3.344, and it misses on average 0.262 of the hundred — almost exactly the ten-observation band’s rate. But its misses cluster less. At least one occurs 18.4% of the time, against 11.0% from ten, and a second follows a first 27.1% of the time, against 45.4%. Independent misses would give 23.1% and 12.4%.

The clustering shrinks because it comes from the sample’s spread, and a spread estimated from thirty observations varies less than one estimated from ten. In the limit of a perfectly known spread the misses would be exactly independent, and the band’s promise would be a plain binomial one. The practical lesson is that an allowance of a few misses is worth more the larger the sample the band was built from — which is one more reason the sample size is the lever that matters most.

What the numbers say to someone setting a limit

The practical reading depends on the promise being made, and the table makes the choice explicit.

A promise about every unit in a batch needs the all-of-mm factor, and it is expensive: 4.942 standard deviations for a hundred units from ten observations, and growing without limit with the batch.

A promise about a share needs the tolerance factor, which is 3.382 for 95% content with 95% confidence from ten observations, and that is the right band once the batch is a hundred or more. Using the all-of-mm band for a share promise is waste; using the tolerance band for an every-unit promise is a promise broken far more often than stated.

A promise about a few failures sits between them and costs more per allowed failure at small allowances than at large ones, because the failures cluster. A specification that allows one out-of-limit unit in a hundred saves 13% of width against one that allows none, and one that allows five saves a third. If the business is indifferent between one and five failures in a batch of a hundred that is otherwise shipped, the second specification buys a much tighter limit for nothing.

And in every case the sample of ten is the binding constraint. Two standard deviations of what found that a band drawn from ten observations holds less than its nominal content on most samples, and every factor here is large because ss from ten observations is unreliable. A larger sample shrinks every factor towards its known-parameter value, which for at most five of a hundred is 2.22, and does more for a limit than any choice of promise.

Three-sigma limits from ten points

The commonest few-misses band in practice is not called one. A control chart’s limits are the mean plus and minus three standard deviations, and when the mean and spread are estimated from a short initial run — ten points is not unusual — the limits are exactly xˉ±3s\bar x \pm 3s from ten observations, and the next hundred points plotted against them are the next mm.

The same integral prices them. With nothing wrong in the process, limits set at three sample standard deviations from ten points hold all of the next hundred 52.56% of the time; they let at most one through 69.35% of the time and at most two 78.14%. A chart that its operators believe will raise a false alarm about once in three hundred and seventy points, as three-sigma limits on a known process do, raises at least one in the next hundred on nearly half of all charts set up this way — and, because the misses cluster, the charts that raise one tend to raise several, which is the pattern that gets a real process shut down for a problem it does not have. All of the next ten met the same limits from the all-of-mm side; the few-misses version says that no reasonable allowance rescues them at ten points.

The price of each promise, and the limit a share promise reaches

From ten normal observations, the band holding all of the next hundred with 95% probability has factor 4.942; at most one outside, 4.289; at most two, 3.948; at most five, 3.378. The 95%-content tolerance factor is 3.382.

The band for at most one in twenty of the next mm outside converges to the 95%-content tolerance factor, reaching 3.378 at a hundred and 3.392 at a thousand, so the tolerance interval is the limit of a promise about a share of failures.

The misses cluster: at the one-in-a-hundred factor, a second miss follows a first 45.4% of the time against 12.5% for independent misses at the same average rate.

Every probability is an exact expectation over the sample’s mean and spread — a product rule over two hundred cells of each — with the binomial count of misses computed exactly inside it; at r=0r = 0 it reproduces the all-of-mm calculation of the essay on the next ten to every digit. The factors are found by bisection on kk.

Not claimed: that the population is normal. Every factor here is for normal data, and the tolerance interval’s normal factor fails off normality in ways the distribution-free band does not; a few-misses band for unknown shapes is built from order statistics instead, and its factors are a different calculation. Not claimed either that the future observations are independent of each other apart from the shared sample; a process that drifts produces misses that cluster for a second reason, and the allowance buys less again.

Still open: the few misses without the normal

The distribution-free band — the range of a sample of ninety-three — answered the all-of-one question without assuming a shape, and its few-misses version has a clean form: the chance that at most rr of mm future draws fall outside the range of nn is a beta-binomial probability that does not depend on the population at all.

What that costs against the normal factors here — how large a sample a distribution-free promise of at most one in a hundred needs, and whether the clustering of misses is stronger or weaker when the band is the sample’s own extremes rather than a multiple of its spread — has not been computed, and it decides whether a limit set without assuming normality can afford to allow a few failures at all.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Binomial distributionCoverageEstimated variancePrediction intervalQuality controlSample sizeSimultaneous inferenceTolerance interval