A range allowed a few misses
Worth reading first: The shape, and where its mass is.
A band allowed a few misses priced the promise most specifications actually make — not that every future unit falls inside the limits, but that at most a few of the next hundred fall outside — for a normal population. From ten observations the band lets at most one of the next hundred outside with 95% probability, and the allowance saves a great deal against the 4.942 that holding all hundred requires. Every factor there rested on the population being normal, and a normal factor carried to a population that is not normal fails quietly, the way two standard deviations of what found a textbook band holding less than it advertises more often than not.
The essay ended on the version that assumes nothing. Ninety-three observations and nothing assumed had already shown that the interval between a sample’s smallest and largest values holds a share of the population whose distribution does not depend on the population at all. A few-misses promise built on that interval has an exact answer, the same for every continuous population, and what it costs turns out to be a different kind of price.
Why the answer contains no population
Take observations from any continuous population and the band from the -th smallest to the -th largest. The share of the population outside that band is a random quantity, and its distribution is whatever the population is — because transforming every observation by the population’s own distribution function turns them into uniform draws without changing their order, and the band’s content is a statement about order alone.
Given the share outside, the next draws fall outside independently, so the number of misses is binomial with that share. Averaged over the share, the count is beta-binomial:
Every probability about the band’s misses follows from that line, exactly, and nothing about the population appears in it. A band from the sample’s extremes is the one limit a quality engineer can set with no model at all, and the formula says what it promises.
What each promise costs in observations
For the next hundred observations, with 95% probability:
| promise | from the extremes | from the second extremes |
|---|---|---|
| none outside | 3,850 | 7,750 |
| at most one outside | 637 | 1,204 |
| at most two outside | 300 | 549 |
| at most five outside | 107 | 187 |
| at most ten outside | 50 | 85 |
Holding all hundred future draws needs 3,850 observations — the distribution-free version of the promise the all-of-ten essay found had no finite price for a normal band as the number of future draws grows, arriving here as a price in sample size that grows with instead of a factor that does. Allowing one miss cuts it to 637; five, to 107.
At a hundred future draws, “at most five outside” is the few-misses form of a 95%-content promise, and its 107 sits beside the 93 that the 95/95 tolerance interval needs from the same extremes. The two are close because they are nearly the same statement: a band holding 95% of the population will, with a hundred future draws, usually let about five through. They are not equal because five of a hundred is a count with its own noise. At 93 observations the at-most-five promise holds 92.84% of the time, a little short of 95%.
The normal-theory band keeps every one of these promises from ten observations; it widens its factor instead. The distribution-free band cannot widen — its limits are two of the observations — so the only lever it has is to take more of them, and every promise is priced in the one currency it has.
Why the price grows with the number of future draws
The table’s first row has a simple shape. Holding none of the next ten outside needs 386 observations; none of the next hundred, 3,850; none of the next thousand, 38,495. Each is about thirty-eight and a half times the number of future draws, and the reason is the beta-binomial’s first term: the chance that no future draw falls outside the range of is , which is close to , and keeping that above 95% needs a little under . The all-of- promise without a model is a promise to have about forty times as much past as future.
A promise about a share behaves differently. At most five of a hundred needs 107 observations; at most fifty of a thousand, 95. As the number of future draws grows, “at most one in twenty” stops being a count with its own noise and becomes a statement about the band’s content, and the sample it needs falls towards the 93 of the tolerance interval — the distribution-free version of the limit the normal few-misses band found its factors converging to. All of the next ten is the other end of the same shape: a promise about every future draw has no finite price in the long run, whether it is paid in a factor or in observations.
The count is the same on any population
The claim that no population enters is checkable, and it is worth checking on populations that have nothing in common.
Over twenty thousand samples of ninety-three from each of four populations — normal, exponential, uniform, and Cauchy, which has no mean and no variance — the promise held 92.75% of the time on every one of them, identically, against an exact value of 92.84%. The four counts are not merely close; they are the same number, because each population’s draws were made by pushing the same uniforms through its own distribution function, and a band made of order statistics sees only the order. A normal-theory band counted the same way would have given four different answers, three of them wrong.
What the range costs on normal data
The comparison that decides whether a distribution-free promise is affordable is not the one against ten observations. A study that has 637 observations can also compute the normal band from all of them, and on normal data that band is narrower.
For at most one of the next hundred, the range of 637 normal observations reaches on average ±3.109 standard deviations; the normal-theory band from the same 637 is ±2.934, 6% narrower; from ten it is ±4.289. With the population’s mean and spread known exactly the band would be ±2.914. For at most five, the range of 107 averages ±2.530, against ±2.329 for the normal band from 107 and ±3.378 from ten.
So a study that can afford the observations loses little by refusing the normal assumption: the range is a few per cent wider than the best a normal model could do at the same size, and it keeps its promise on every population. What the assumption buys is the ability to keep the promise from ten observations at all. A process with ten measurements and a normal model can specify a band; a process with ten measurements and no model cannot, and the gap between them is not a width but a factor of sixty in data.
The misses of one range are shared
The normal few-misses band found that its misses cluster: when the band from ten observations lets one of the next hundred through, it lets a second through 45.4% of the time, where independent misses at the same rate would do so 12.5%. The band from a small sample is itself uncertain, and a band that happens to be too narrow is too narrow for every future draw at once.
The range’s misses are shared in the same way, because every future draw is judged against the same two observations, and the beta-binomial measures it.
For the range of 107 the chance of no miss at all among the next hundred is 0.266, against 0.154 if the misses were independent at the same average rate, and the chance of more than eight is 0.007 against almost none. Both tails are heavier because the misses share their band.
Measured the way the normal band’s clustering was, the range of 637 that keeps the at-most-one promise lets a second of the next hundred through 19.7% of the time given a first, against 14.7% for independent misses at the same rate. That is much weaker clustering than the normal band from ten, and for a simple reason: the range of 637 is a far better-determined band than from ten, and it is the uncertainty in the band that makes misses arrive together. At equal sizes the range clusters somewhat more than the normal band — the normal band from 637 gives 17.1% against its own independent 16.3% — because a band set by two observations moves more than one set by a mean and a standard deviation.
There is also an exact version. Given that one particular future draw falls outside the range of , another particular draw falls outside it with probability , against an unconditional — a ratio that tends to one and a half however large the sample. The range never stops sharing its misses; it only makes them rarer.
Trimming to survive a bad value
A band from the very extremes is only as good as the two most extreme observations, and a single recording error at either end widens it without limit. The usual defence is to use the second smallest and second largest values instead. The table prices the defence: every promise needs roughly twice the observations — 1,204 instead of 637 for at most one miss in a hundred, 187 instead of 107 for at most five — because a band from the second extremes has twice the expected share outside it and needs a larger sample to push that share back down.
That is a reasonable trade where the process produces occasional wild values, and an expensive one where it does not. It has the same shape as robust is not free found for a robust standard error: protection against a failure of the assumptions is paid for when the assumptions hold, and here the payment is counted in observations.
The same band under another name
A band made of a sample’s order statistics, with a coverage that holds for any population because only ranks enter, is also the construction behind conformal prediction. Coverage from exchangeability alone showed that a conformal interval’s coverage is a fact about the ranks of the calibration scores and the new one, computable before any data arrive; the range of a sample is the simplest conformal interval there is, with the observations themselves as the scores.
The clustering measured here is the conformal literature’s distinction between marginal and conditional coverage, seen from a quality engineer’s side. A conformal interval promises its coverage on average over calibration samples; any one calibration sample gives an interval whose coverage is a draw from a beta distribution, and every future point judged against it shares that draw. Marginal is not conditional found the same gap between an average promise and the promise one realised band keeps, there across subgroups of the future points and here across the future points of one lot. Both are the reason a few-misses promise has to be computed rather than read off an average coverage.
One qualification lot, two specifications
A concrete case makes the choice visible. A supplier qualifies a process on 120 measured units and has to state what the next hundred will do.
Without a model, the only band it can offer is the range of its 120, and the strongest few-misses promise that range keeps at 95% is at most five of the next hundred outside, which holds with probability 0.965; at most four holds only 0.930. A promise of at most one outside holds with probability 0.568 — the range of 120 lets two or more of the next hundred through more often than not.
With a normal model, the same 120 units support for at most one of the next hundred, and for none of them. Those bands are wider than the range will typically be, but they keep promises the range cannot keep at any width, because the range cannot be widened.
So the supplier’s choice is not between a wide band and a narrow one. It is between a weak promise that holds for any process and a strong promise that holds if the process is normal, and the second is worth exactly as much as the evidence that the process is normal — which, from 120 units, is limited in the tails where the promise is kept. A specification that states which of the two it made, and on what evidence, is one a customer can hold it to.
What the specification should say
If no shape can be assumed, state the band as order statistics and the promise as a count. “At most one of the next hundred outside the range of the qualification lot” is a precise promise with an exact probability, and 637 is the lot size it needs.
If a shape can be assumed, check what it buys. A normal band from ten observations keeps the same promise, and a normal band from 637 keeps it 6% more narrowly. If the normal assumption is doubtful, the range costs a few per cent of width at the sample sizes where it is available; if the sample is small, the assumption is doing all the work, and a tail the sample never saw is what it is doing the work on.
Say how many misses the process can tolerate, because it decides the sample more than anything else. None of a hundred needs 3,850 observations, one needs 637 and five need 107. A specification that could tolerate two failures in a hundred but was written as zero asks for six times the data.
And expect the misses in clumps. A lot whose limits were set from a sample’s extremes lets its failures through together more often than a binomial count suggests; an acceptance plan that treats consecutive failures as independent evidence against the process will overreact to the band’s own uncertainty.
What is exact here and what is counted
The number of the next m draws outside the band between the s-th extremes of n is beta-binomial with parameters 2s and n + 1 − 2s, for every continuous population.
For the next hundred, at 95%: none outside needs 3,850 observations, at most one 637, at most five 107; from the second extremes 7,750, 1,204 and 187.
On normal data the range of 637 averages ±3.109 standard deviations against ±2.934 for the normal band from 637; given a first miss among the next hundred, a second follows 19.7% of the time, against 45.4% for the normal band from ten.
The counts on four populations, twenty thousand samples each, agree with the beta-binomial to within their sampling error, and with each other exactly. The range’s average half-width on normal data is counted over four thousand samples.
Not claimed: that populations are continuous. A discrete or heavily rounded measurement produces ties at the extremes, and a future draw equal to the sample’s largest is then neither inside nor outside in the sense the formula uses; the promise needs restating for that case.
Still open: a band that learns as the lot is used
Every band here is set once, from a qualification sample, and then judged against the next hundred. A process in production keeps measuring, and the natural practice is to widen the reference sample as units arrive — the range of everything seen so far — so that the band moves as the lot is used.
A band updated that way is no longer a fixed band with a beta-binomial count of misses; each future draw is judged against a band that includes every earlier future draw, and the count of misses becomes the number of new records in a sequence, whose distribution is known and quite different. How that changes what “at most r of the next m outside” means, and whether an updated band keeps a few-misses promise more cheaply than a fixed one from a larger initial sample, is a calculation this family of results points straight at and has not made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A tenth as wide, and both of them right — both name prediction interval, sample size, tolerance interval
- A detector built for the ordering — both name distribution-free, exchangeability
- The score is the modelling — both name distribution-free, exchangeability
- What the split costs — both name exchangeability, order statistic
- When the order matters — both name distribution-free, exchangeability
Named objects
A flat tag is an object no other essay names yet.
Binomial distributionDistribution-freeExchangeabilityOrder statisticPrediction intervalQuality controlSample sizeTolerance interval