The interval that holds observations, not a mean

A range that grows with the lot

A band set from the range of everything seen so far flags a unit only when it sets a new record, and for any continuous population those flags are independent, with the k-th unit flagged with probability 2/(n + k). That makes a few-misses promise far cheaper than the fixed range's — at most five of the next hundred from 36 starting observations rather than 107 — and exactly as expensive when no miss is allowed, because a lot with no flag never moves the band. The saving is bought by believing every flag. After a two-standard-deviation step in the mean, the fixed range flags 25.03 of the eighty units that follow and the learning range 3.62; by the last twenty units it sees a unit outside in 25.5% of lots, against 18.4% with no step at all.

Worth reading first: The shape, and where its mass is.

A range allowed a few misses priced the specification a quality engineer can write with no model at all: the band between a qualification sample’s smallest and largest values, and a promise that at most a few of the next hundred units fall outside it. For any continuous population the number outside is beta-binomial, and the promise has an exact price in observations — 637 for at most one of the next hundred, 107 for at most five, 3,850 for none.

Every band in that essay is set once and then left alone. A process in production does not work that way. It keeps measuring, and the natural practice is to fold each unit into the reference as it arrives, so that the band at any moment is the range of everything seen so far. The essay ended by noting that a band updated this way is no longer a fixed band with a beta-binomial count of misses — each unit is judged against a band that already includes every earlier unit — and that the count becomes the number of new records in a sequence, whose distribution is known and quite different. That distribution turns out to be simpler than the one it replaces, and what it says about the practice is less comfortable than the price suggests.

How many starting observations a range needs to flag at most r of the next 100, fixed at the start or learning as the lot is used95% of the time, for any continuous population. At most 0: fixed 3,850, learning 3,850; At most 1: fixed 637, learning 512; At most 2: fixed 300, learning 197; At most 3: fixed 190, learning 101; At most 5: fixed 107, learning 36; At most 7: fixed 74, learning 15; At most 10: fixed 50, learning 4. With no flag allowed the two are the same, because a lot with no flag never moves the learning range.101001,00001235710at most this many of the next 100 flaggedstarting observations the range needs, log scalerange fixed at the startrange of everything seen so farexact, no population in itthe learning range believes every flag
Fig. 1 How many starting observations the range needs before at most r of the next hundred units are flagged with 95% probability, for a range fixed at the start and for the range of everything seen so far, on a logarithmic scale. No population enters. The slider sets the number of future units.

A flag is a record

Start with nn observations and let the band be their range. The first future unit is flagged if it falls below the smallest or above the largest, and then it joins the reference whether or not it was flagged. The second future unit is judged against the range of n+1n + 1, and so on. Unit kk is flagged exactly when it is the smallest or the largest of the first n+kn + k values — when it sets a new record in one direction or the other.

For an exchangeable sequence from any continuous population, each of the first n+kn + k values is equally likely to hold any rank among them, so the chance that the newest is the smallest or the largest is 2/(n+k)2/(n + k). And a classical result about records says more than that. The rank of each new value among those before it is independent of how the earlier values ranked among themselves. So the flags are independent events, with probabilities

2n+1, 2n+2, 2n+3, …,\frac{2}{n+1},\ \frac{2}{n+2},\ \frac{2}{n+3},\ \ldots,

and the number flagged among the next mm is a sum of independent Bernoulli draws with those probabilities. Every probability about the learning range’s flags follows from that list, exactly, and like the beta-binomial it replaces, it contains no population at all.

The two counts start from the same place: the first unit is flagged with probability 2/(n+1)2/(n + 1) under either band. After that they part. The fixed range keeps flagging each unit at 2/(n+1)2/(n + 1); the learning range’s chance falls with every unit, because the reference keeps growing.

The same price when nothing may be flagged

The table of prices has one entry the two bands share exactly.

at most this many of the next 100 flagged fixed range learning range
none 3,850 3,850
one 637 512
two 300 197
three 190 101
five 107 36
seven 74 15
ten 50 4

A lot in which no unit is flagged is a lot in which the learning range never moved, because the range of a sample changes only when a new value falls outside it. So the event “no unit flagged” is the same event for both bands, its probability is the same number, n(n−1)/[(n+m)(n+m−1)]n(n-1)/[(n+m)(n+m-1)], and the price of promising it is 3,850 starting observations whichever way the band is kept. All of the next ten found that a promise about every future unit has no finite price in the long run, whether it is paid in a factor or in observations; learning does not escape that, because the promise of no flag at all is a promise that learning never happens.

That observation has a consequence worth stating in its own right. A reference that grows only from units that were accepted — the conservative practice of never letting a flagged unit into the baseline — is not a learning band at all. It is the fixed band. The range of a sample is moved by nothing that falls inside it, so adding accepted units leaves the band exactly where it was, and all of the learning range’s difference from the fixed one comes from a single decision: to add the flagged units too.

Everything else is cheaper, and much cheaper

Once one flag is allowed, the learning range pulls ahead. At most one of the next hundred needs 512 starting observations against 637; at most two, 197 against 300; at most five, 36 against 107; at most ten, 4 against 50. The saving grows with the allowance because each flag the learning range raises widens it, and the wider band makes the next flag less likely. A fixed band that has been shown too narrow by one unit stays too narrow for all the others.

How many of the next 100 units are flagged by the range of 107, fixed or learning, for any continuous population. Fixed at the start, the count is beta-binomial with mean 1.852; learning, it is a sum of independent draws with mean 1.315. No flag at all has probability 0.266 for both. Five or fewer: 0.9508 fixed, 0.9979 learning. More than five: 0.0492 against 0.0021.
Fig. 2 The number of the next hundred units flagged by the range of 107 starting observations, fixed at the start and learning. No flag at all has the same probability under both.

The distribution of flags from 107 starting observations shows the mechanism. Both bands flag nothing with probability 0.266. The fixed range then spreads its remaining probability over a long tail — its mean is 1.852 flags, and it lets more than five through with probability 0.0492 — while the learning range piles it onto one and two flags, with a mean of 1.315 and more than five with probability 0.0021. The learning range’s tail is short because its flags are independent and each one lowers the chance of the next; the fixed range’s is long because all its misses share one band, the clustering a band allowed a few misses measured for the normal band and the range essay measured for order statistics.

That clustering is gone entirely. The fixed range of 637 lets a second of the next hundred through 19.7% of the time given a first; the learning range from the same start does so 13.8% of the time, which is exactly what its independent flags give, since knowing one unit was a record says nothing about whether another will be. A quality engineer who treats consecutive flags as independent evidence is right about the learning range and wrong about the fixed one.

The rate falls like one over the count

The independence is not the only difference in shape. The two bands answer a question about the next hundred units very differently from a question about the next ten thousand.

The chance each future unit is flagged by the range of 107, fixed or learning. Fixed at the start, every unit is flagged with probability 0.0185, and 0.0275 once another unit is known to have been flagged. Learning, the k-th unit is flagged with probability 2/(107 + k): 0.0185 for the first, 0.0097 for the hundredth and 0.0018 for the 1,000th, whatever happened before. Over 1,000 units the learning range flags 4.66 on average against 18.52 for the fixed range.
Fig. 3 The chance that each future unit is flagged by a range from 107 starting observations, against its position in the lot on a logarithmic scale: fixed, fixed once another unit is known to have been flagged, and learning.

From 107 starting observations, the fixed range flags every unit with probability 0.0185, and once any other unit has been flagged, with probability 0.0275 — the shared band made visible, since a flag is evidence that the band came out narrow. The learning range flags its first unit with the same 0.0185, its hundredth with 0.0097, its thousandth with 0.0018. Over a thousand units it raises 4.66 flags on average against 18.52 for the fixed range, and over ten thousand, 9.09 against 185.2.

The learning range’s count grows like the logarithm of the lot’s length, about 2ln⁡((n+m)/n)2\ln\left((n+m)/n\right), where the fixed range’s grows in proportion to it. A few-misses promise from a learning range therefore gets easier the longer the lot, and in the limit it is a promise about records, which arrive ever more rarely in any stable sequence. That is what makes the table’s right-hand column so cheap. It is also the first sign of what the band has stopped measuring.

What a record means and what a miss meant

The fixed range’s flag has a plain meaning: this unit lies outside what the qualification lot showed the process doing. The learning range’s flag means something weaker: this unit is the most extreme the process has produced so far, counting the lot itself. The two agree on a stable process, which is the only case the prices above describe. They disagree on the case a specification exists for.

Suppose the process mean steps up after the twentieth unit of a lot of a hundred, and both bands start from the range of 107 normal observations.

How often each unit is flagged after the mean steps by 2 standard deviations, fixed range and learning range. A range from 107 normal observations, a lot of a hundred, the mean stepping up by 2 standard deviations after unit 20; 4,000 lots. After the step the fixed range flags 25.03 units on average and the learning range 3.62. A unit in the last twenty is flagged in 97.8% of lots by the fixed range and 25.5% by the learning range.
Fig. 4 The share of lots in which each unit is flagged, when the mean steps up by two standard deviations after unit 20, for the range fixed at the start and the range of everything seen so far.

After a step of two standard deviations, the fixed range flags about a third of all units from the step onward, and keeps flagging them at that rate to the end of the lot: 25.03 flags on average among the eighty units after the step. The learning range flags the first few units after the step nearly as often, and each of those flags widens it towards the new level, until the new level is simply inside the band. It raises 3.62 flags after the step, nearly half of them within ten units of it. By the last twenty units of the lot, the fixed range flags at least one unit in 97.8% of lots and the learning range in 25.5% — against 18.4% for the learning range when there was no step at all.

The learning range has not missed the step. It saw it, raised three or four flags, and then took those flags as information about what the process does. That is precisely what it was built to do, and it is what a band for detecting a changed process must not do. Locating a step after the fact is a separate problem, the one a break that was looked for takes up for a series whose regimes are unknown; a band’s job is the earlier one of saying, unit by unit, that something has changed at all.

What the promise catches

The at-most-five promise can be read the other way, as an alarm: a lot with more than five flags is a lot in which the process is declared not to be behaving as qualified.

What the fixed and the learning range say about a step in the mean, by its size. Share of lots of a hundred in which more than five units are flagged: step 0, fixed 4.7%, learning 0.3%; step 0.25, fixed 7.3%, learning 0.4%; step 0.5, fixed 15.6%, learning 0.6%; step 0.75, fixed 30.6%, learning 1.1%; step 1, fixed 51.6%, learning 2.2%; step 1.5, fixed 84.6%, learning 7.9%; step 2, fixed 97.2%, learning 17.0%; step 2.5, fixed 99.5%, learning 26.1%; step 3, fixed 99.9%, learning 33.3%. Share with a flag among the last twenty units, at no step and at a step of 3: fixed 29.0% and 99.9%, learning 18.4% and 25.5%.
Fig. 5 Against the size of a step in the mean after unit 20, the share of lots of a hundred in which more than five units are flagged, and the share with a flag among the last twenty units, for the fixed and the learning range from 107 starting observations.

With no step, the fixed range breaks its promise in 4.7% of lots, as a 95% promise should, and the learning range in 0.3% — it is the more conservative alarm. At a step of one standard deviation the fixed range raises its alarm in 51.6% of lots and the learning range in 2.2%. At two, 97.2% against 17.0%; at three, 99.9% against 33.3%. The learning range needs a step of three standard deviations to raise its alarm in a third of lots, and even then, by the end of the lot it is no more alert than it was before the step.

So the comparison in the table is between two promises that are not the same promise. The fixed range from 107 promises that at most five of the next hundred will lie outside what the qualification lot showed, and it keeps that promise on a stable process and breaks it loudly on a shifted one. The learning range from 36 promises that at most five of the next hundred will be records, which it keeps on a stable process and keeps almost as well on a shifted one. The second promise is cheaper because it says less.

The same choice in a conformal interval

A band made of order statistics, with a coverage that holds for any population, is the simplest case of a conformal interval, as coverage from exchangeability alone established: the calibration scores are the observations and the interval is their range. Marginal is not conditional found that such a guarantee is an average over calibration samples rather than a promise any one band keeps, which is the clustering of the fixed range seen from the conformal side. The learning range is then the conformal procedure that adds each new score to the calibration set as it arrives, and the fixed range is the one that calibrates once.

When the order matters found that breaking exchangeability costs a conformal interval coverage in amounts that run against how quickly a test would notice, and a detector built for the ordering priced the checks that could catch a drifting scale. The step measured here adds the other half of that picture. An interval that recalibrates on everything it sees repairs its own coverage after a change, which is what an online conformal method is praised for, and in doing so it removes the evidence that the change happened. Coverage and detection pull in opposite directions, and a band cannot be tuned to deliver both from one reference.

Why neither band is the wrong one

It would be easy to read the step figure as a verdict against the learning range, and it is not one. A production line whose mean wanders slowly and harmlessly — a tool wearing within its tolerance, a raw material varying between batches — is badly served by a fixed band that treats every such drift as a failure; the fixed range’s 97.2% at a two-sigma step is an alarm about the change, not about the units, and if the new level is fine the alarm is a cost. The learning range is the right instrument for a specification that is about each unit being unremarkable relative to the process as it now is.

The fixed range is the right instrument for a specification that is about the process still being the one that was qualified. Most specifications written for a customer are of the second kind — a supplier is promising that the next hundred units behave like the hundred and seven the customer saw — and that is the promise the fixed range prices. Ninety-three observations and nothing assumed and its successors priced that promise. A learning band cannot keep it more cheaply, because it does not keep it at all.

What goes wrong in practice is the mixture: a band that is described to the customer as the qualification range and maintained on the line as the range of everything seen. The prices then come from the right-hand column and the promise from the left.

A middle course, and what it costs

The two bands are the ends of a range of practices. A reference that grows from accepted units only is, for the range, the fixed band, as shown above; it moves only through a separate decision to requalify. A reference that grows from everything is the learning band. Between them is a reference that admits a flagged unit only after it has been investigated and found to be sound — which is a fixed band with a human in the loop, and whose price is the fixed band’s price plus the cost of the investigations.

For bands that are not ranges — a normal-theory band xˉ±ks\bar x \pm ks recomputed from a growing sample, or the second extremes — interior units do move the reference, and the accepted-only practice is genuinely between the two. It widens slowly as the sample grows and the factor kk falls, and a flagged unit kept out of it keeps its evidential weight — though a normal-theory reference carries its own risk, since its promise is only as good as the shape assumed in tails it has not seen, which is where a tail the sample never saw found the assumption doing all the work. A band allowed a few misses priced that band when it is fixed; how its growing version trades cost against detection is a calculation of the same kind as the one here, without the exact independence that makes the range’s version so clean.

The specification a learning band can honestly state

Say which band is being kept. “At most five of the next hundred outside the qualification range” and “at most five of the next hundred outside the range of everything measured so far” are different promises with different prices — 107 starting observations and 36 — and the cheaper one does not imply the dearer.

If the band learns, say that its flags are records. For a stable process they are independent and arrive with probability 2/(n+k)2/(n + k), so consecutive flags are not evidence of clustering and a long run with few flags is expected rather than reassuring.

If the purpose is to catch a changed process, keep the band fixed or keep a second, fixed band beside it. The learning range raised its at-most-five alarm in 17.0% of lots after a two-sigma step; the fixed range, in 97.2%. A process monitored only by a band that learns will be declared stable after most of the changes a customer would want to hear about.

And do not buy the learning band’s price with the fixed band’s promise. The saving is real only for the promise the learning band makes.

Exact, and counted

Exact, for any continuous population: the kk-th future unit is flagged by the range of everything before it with probability 2/(n+k)2/(n + k), the flags are independent, and the count among the next mm is their sum. No flag at all has probability n(n−1)/[(n+m)(n+m−1)]n(n-1)/[(n+m)(n+m-1)] for the learning and the fixed range alike, so promising none costs 3,850 starting observations either way for the next hundred.

Exact, from the same sum: at most one of the next hundred flagged needs 512 starting observations against the fixed range’s 637; at most five, 36 against 107; at most ten, 4 against 50. From 107, the learning range flags 1.315 of the next hundred on average against 1.852.

Counted, over four thousand normal lots at each step: after a two-sigma step at unit 20, 25.03 flags from the fixed range and 3.62 from the learning range, and a flag in the last twenty units in 97.8% of lots against 25.5%.

The independence was checked by counting as well as derived: on twenty thousand lots from an exponential population the learning range from 36 kept the at-most-five promise 94.97% of the time against an exact 95.09%, and a Cauchy population drawn from the same uniforms gave the identical count.

Not claimed: that a step is the only way a process changes. A slow trend produces records steadily rather than in a burst, and the learning range should be expected to flag it for longer, though that was not counted here; a change in spread rather than in mean produces records at both ends. The step is the case where the difference between the two bands is sharpest, not the only case where it exists.

Still open: a band that forgets

The learning range remembers everything, and that is why it absorbs a step: after enough units at the new level, the old extremes are still in the reference and the new level’s extremes have joined them. A reference of the last ww units only — a moving window — forgets as well as learns. It would flag a step when it arrives, absorb it within about ww units, and then flag a return to the old level as a change in its own right.

A moving window’s flags are no longer records of an exchangeable sequence, so the exact independence above is lost; the count depends on how the window’s extremes turn over, which for a stable process is still free of the population but no longer a sum of independent terms. What window length a few-misses promise needs, whether a window can be chosen to flag a step for long enough to be acted on and then stop, and how it compares with keeping one fixed band and one learning band side by side, have not been worked out here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Change pointDistribution-freeExchangeabilityOrder statisticQuality controlSample sizeTolerance interval