A range that grows with the lot
Worth reading first: The shape, and where its mass is.
A range allowed a few misses priced the specification a quality engineer can write with no model at all: the band between a qualification sample’s smallest and largest values, and a promise that at most a few of the next hundred units fall outside it. For any continuous population the number outside is beta-binomial, and the promise has an exact price in observations — 637 for at most one of the next hundred, 107 for at most five, 3,850 for none.
Every band in that essay is set once and then left alone. A process in production does not work that way. It keeps measuring, and the natural practice is to fold each unit into the reference as it arrives, so that the band at any moment is the range of everything seen so far. The essay ended by noting that a band updated this way is no longer a fixed band with a beta-binomial count of misses — each unit is judged against a band that already includes every earlier unit — and that the count becomes the number of new records in a sequence, whose distribution is known and quite different. That distribution turns out to be simpler than the one it replaces, and what it says about the practice is less comfortable than the price suggests.
A flag is a record
Start with observations and let the band be their range. The first future unit is flagged if it falls below the smallest or above the largest, and then it joins the reference whether or not it was flagged. The second future unit is judged against the range of , and so on. Unit is flagged exactly when it is the smallest or the largest of the first values — when it sets a new record in one direction or the other.
For an exchangeable sequence from any continuous population, each of the first values is equally likely to hold any rank among them, so the chance that the newest is the smallest or the largest is . And a classical result about records says more than that. The rank of each new value among those before it is independent of how the earlier values ranked among themselves. So the flags are independent events, with probabilities
and the number flagged among the next is a sum of independent Bernoulli draws with those probabilities. Every probability about the learning range’s flags follows from that list, exactly, and like the beta-binomial it replaces, it contains no population at all.
The two counts start from the same place: the first unit is flagged with probability under either band. After that they part. The fixed range keeps flagging each unit at ; the learning range’s chance falls with every unit, because the reference keeps growing.
The same price when nothing may be flagged
The table of prices has one entry the two bands share exactly.
| at most this many of the next 100 flagged | fixed range | learning range |
|---|---|---|
| none | 3,850 | 3,850 |
| one | 637 | 512 |
| two | 300 | 197 |
| three | 190 | 101 |
| five | 107 | 36 |
| seven | 74 | 15 |
| ten | 50 | 4 |
A lot in which no unit is flagged is a lot in which the learning range never moved, because the range of a sample changes only when a new value falls outside it. So the event “no unit flagged” is the same event for both bands, its probability is the same number, , and the price of promising it is 3,850 starting observations whichever way the band is kept. All of the next ten found that a promise about every future unit has no finite price in the long run, whether it is paid in a factor or in observations; learning does not escape that, because the promise of no flag at all is a promise that learning never happens.
That observation has a consequence worth stating in its own right. A reference that grows only from units that were accepted — the conservative practice of never letting a flagged unit into the baseline — is not a learning band at all. It is the fixed band. The range of a sample is moved by nothing that falls inside it, so adding accepted units leaves the band exactly where it was, and all of the learning range’s difference from the fixed one comes from a single decision: to add the flagged units too.
Everything else is cheaper, and much cheaper
Once one flag is allowed, the learning range pulls ahead. At most one of the next hundred needs 512 starting observations against 637; at most two, 197 against 300; at most five, 36 against 107; at most ten, 4 against 50. The saving grows with the allowance because each flag the learning range raises widens it, and the wider band makes the next flag less likely. A fixed band that has been shown too narrow by one unit stays too narrow for all the others.
The distribution of flags from 107 starting observations shows the mechanism. Both bands flag nothing with probability 0.266. The fixed range then spreads its remaining probability over a long tail — its mean is 1.852 flags, and it lets more than five through with probability 0.0492 — while the learning range piles it onto one and two flags, with a mean of 1.315 and more than five with probability 0.0021. The learning range’s tail is short because its flags are independent and each one lowers the chance of the next; the fixed range’s is long because all its misses share one band, the clustering a band allowed a few misses measured for the normal band and the range essay measured for order statistics.
That clustering is gone entirely. The fixed range of 637 lets a second of the next hundred through 19.7% of the time given a first; the learning range from the same start does so 13.8% of the time, which is exactly what its independent flags give, since knowing one unit was a record says nothing about whether another will be. A quality engineer who treats consecutive flags as independent evidence is right about the learning range and wrong about the fixed one.
The rate falls like one over the count
The independence is not the only difference in shape. The two bands answer a question about the next hundred units very differently from a question about the next ten thousand.
From 107 starting observations, the fixed range flags every unit with probability 0.0185, and once any other unit has been flagged, with probability 0.0275 — the shared band made visible, since a flag is evidence that the band came out narrow. The learning range flags its first unit with the same 0.0185, its hundredth with 0.0097, its thousandth with 0.0018. Over a thousand units it raises 4.66 flags on average against 18.52 for the fixed range, and over ten thousand, 9.09 against 185.2.
The learning range’s count grows like the logarithm of the lot’s length, about , where the fixed range’s grows in proportion to it. A few-misses promise from a learning range therefore gets easier the longer the lot, and in the limit it is a promise about records, which arrive ever more rarely in any stable sequence. That is what makes the table’s right-hand column so cheap. It is also the first sign of what the band has stopped measuring.
What a record means and what a miss meant
The fixed range’s flag has a plain meaning: this unit lies outside what the qualification lot showed the process doing. The learning range’s flag means something weaker: this unit is the most extreme the process has produced so far, counting the lot itself. The two agree on a stable process, which is the only case the prices above describe. They disagree on the case a specification exists for.
Suppose the process mean steps up after the twentieth unit of a lot of a hundred, and both bands start from the range of 107 normal observations.
After a step of two standard deviations, the fixed range flags about a third of all units from the step onward, and keeps flagging them at that rate to the end of the lot: 25.03 flags on average among the eighty units after the step. The learning range flags the first few units after the step nearly as often, and each of those flags widens it towards the new level, until the new level is simply inside the band. It raises 3.62 flags after the step, nearly half of them within ten units of it. By the last twenty units of the lot, the fixed range flags at least one unit in 97.8% of lots and the learning range in 25.5% — against 18.4% for the learning range when there was no step at all.
The learning range has not missed the step. It saw it, raised three or four flags, and then took those flags as information about what the process does. That is precisely what it was built to do, and it is what a band for detecting a changed process must not do. Locating a step after the fact is a separate problem, the one a break that was looked for takes up for a series whose regimes are unknown; a band’s job is the earlier one of saying, unit by unit, that something has changed at all.
What the promise catches
The at-most-five promise can be read the other way, as an alarm: a lot with more than five flags is a lot in which the process is declared not to be behaving as qualified.
With no step, the fixed range breaks its promise in 4.7% of lots, as a 95% promise should, and the learning range in 0.3% — it is the more conservative alarm. At a step of one standard deviation the fixed range raises its alarm in 51.6% of lots and the learning range in 2.2%. At two, 97.2% against 17.0%; at three, 99.9% against 33.3%. The learning range needs a step of three standard deviations to raise its alarm in a third of lots, and even then, by the end of the lot it is no more alert than it was before the step.
So the comparison in the table is between two promises that are not the same promise. The fixed range from 107 promises that at most five of the next hundred will lie outside what the qualification lot showed, and it keeps that promise on a stable process and breaks it loudly on a shifted one. The learning range from 36 promises that at most five of the next hundred will be records, which it keeps on a stable process and keeps almost as well on a shifted one. The second promise is cheaper because it says less.
The same choice in a conformal interval
A band made of order statistics, with a coverage that holds for any population, is the simplest case of a conformal interval, as coverage from exchangeability alone established: the calibration scores are the observations and the interval is their range. Marginal is not conditional found that such a guarantee is an average over calibration samples rather than a promise any one band keeps, which is the clustering of the fixed range seen from the conformal side. The learning range is then the conformal procedure that adds each new score to the calibration set as it arrives, and the fixed range is the one that calibrates once.
When the order matters found that breaking exchangeability costs a conformal interval coverage in amounts that run against how quickly a test would notice, and a detector built for the ordering priced the checks that could catch a drifting scale. The step measured here adds the other half of that picture. An interval that recalibrates on everything it sees repairs its own coverage after a change, which is what an online conformal method is praised for, and in doing so it removes the evidence that the change happened. Coverage and detection pull in opposite directions, and a band cannot be tuned to deliver both from one reference.
Why neither band is the wrong one
It would be easy to read the step figure as a verdict against the learning range, and it is not one. A production line whose mean wanders slowly and harmlessly — a tool wearing within its tolerance, a raw material varying between batches — is badly served by a fixed band that treats every such drift as a failure; the fixed range’s 97.2% at a two-sigma step is an alarm about the change, not about the units, and if the new level is fine the alarm is a cost. The learning range is the right instrument for a specification that is about each unit being unremarkable relative to the process as it now is.
The fixed range is the right instrument for a specification that is about the process still being the one that was qualified. Most specifications written for a customer are of the second kind — a supplier is promising that the next hundred units behave like the hundred and seven the customer saw — and that is the promise the fixed range prices. Ninety-three observations and nothing assumed and its successors priced that promise. A learning band cannot keep it more cheaply, because it does not keep it at all.
What goes wrong in practice is the mixture: a band that is described to the customer as the qualification range and maintained on the line as the range of everything seen. The prices then come from the right-hand column and the promise from the left.
A middle course, and what it costs
The two bands are the ends of a range of practices. A reference that grows from accepted units only is, for the range, the fixed band, as shown above; it moves only through a separate decision to requalify. A reference that grows from everything is the learning band. Between them is a reference that admits a flagged unit only after it has been investigated and found to be sound — which is a fixed band with a human in the loop, and whose price is the fixed band’s price plus the cost of the investigations.
For bands that are not ranges — a normal-theory band recomputed from a growing sample, or the second extremes — interior units do move the reference, and the accepted-only practice is genuinely between the two. It widens slowly as the sample grows and the factor falls, and a flagged unit kept out of it keeps its evidential weight — though a normal-theory reference carries its own risk, since its promise is only as good as the shape assumed in tails it has not seen, which is where a tail the sample never saw found the assumption doing all the work. A band allowed a few misses priced that band when it is fixed; how its growing version trades cost against detection is a calculation of the same kind as the one here, without the exact independence that makes the range’s version so clean.
The specification a learning band can honestly state
Say which band is being kept. “At most five of the next hundred outside the qualification range” and “at most five of the next hundred outside the range of everything measured so far” are different promises with different prices — 107 starting observations and 36 — and the cheaper one does not imply the dearer.
If the band learns, say that its flags are records. For a stable process they are independent and arrive with probability , so consecutive flags are not evidence of clustering and a long run with few flags is expected rather than reassuring.
If the purpose is to catch a changed process, keep the band fixed or keep a second, fixed band beside it. The learning range raised its at-most-five alarm in 17.0% of lots after a two-sigma step; the fixed range, in 97.2%. A process monitored only by a band that learns will be declared stable after most of the changes a customer would want to hear about.
And do not buy the learning band’s price with the fixed band’s promise. The saving is real only for the promise the learning band makes.
Exact, and counted
Exact, for any continuous population: the -th future unit is flagged by the range of everything before it with probability , the flags are independent, and the count among the next is their sum. No flag at all has probability for the learning and the fixed range alike, so promising none costs 3,850 starting observations either way for the next hundred.
Exact, from the same sum: at most one of the next hundred flagged needs 512 starting observations against the fixed range’s 637; at most five, 36 against 107; at most ten, 4 against 50. From 107, the learning range flags 1.315 of the next hundred on average against 1.852.
Counted, over four thousand normal lots at each step: after a two-sigma step at unit 20, 25.03 flags from the fixed range and 3.62 from the learning range, and a flag in the last twenty units in 97.8% of lots against 25.5%.
The independence was checked by counting as well as derived: on twenty thousand lots from an exponential population the learning range from 36 kept the at-most-five promise 94.97% of the time against an exact 95.09%, and a Cauchy population drawn from the same uniforms gave the identical count.
Not claimed: that a step is the only way a process changes. A slow trend produces records steadily rather than in a burst, and the learning range should be expected to flag it for longer, though that was not counted here; a change in spread rather than in mean produces records at both ends. The step is the case where the difference between the two bands is sharpest, not the only case where it exists.
Still open: a band that forgets
The learning range remembers everything, and that is why it absorbs a step: after enough units at the new level, the old extremes are still in the reference and the new level’s extremes have joined them. A reference of the last units only — a moving window — forgets as well as learns. It would flag a step when it arrives, absorb it within about units, and then flag a return to the old level as a change in its own right.
A moving window’s flags are no longer records of an exchangeable sequence, so the exact independence above is lost; the count depends on how the window’s extremes turn over, which for a stable process is still free of the population but no longer a sum of independent terms. What window length a few-misses promise needs, whether a window can be chosen to flag a step for long enough to be acted on and then stop, and how it compares with keeping one fixed band and one learning band side by side, have not been worked out here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A tenth as wide, and both of them right — both name sample size, tolerance interval
- The score is the modelling — both name distribution-free, exchangeability
- Two standard deviations of what — both name sample size, tolerance interval
- What the split costs — both name exchangeability, order statistic
Named objects
A flat tag is an object no other essay names yet.
Change pointDistribution-freeExchangeabilityOrder statisticQuality controlSample sizeTolerance interval