A detector built for the ordering
Worth reading first: Coverage from exchangeability alone.
The essay on what breaking exchangeability costs reported a reversal: the departure from exchangeability that costs the most coverage is the one that is hardest to notice, and the one that gets tested for costs almost nothing. It then named the objection to itself. The detectability half of that comparison rested on the power curve of a single rank test — a comparison of the first half of the calibration scores against the second — and a test that throws away the ordering inside each half is an obvious thing to improve on — the same complaint a declustering rule invites when it reduces a sequence to a count. If a better detector existed, the reversal might be a property of the instrument.
Three detectors, each given its own critical value by simulation so that all three sit at 5% exactly, answer it.
A better detector exists and it is better by a factor of two. It does not overturn the conclusion.
The three checks, and why they are ordered as they are
Each detector reads the same two hundred calibration scores and each throws away a different amount of what is in them.
The first half against the second reduces every score to one bit — which half it was in — and compares the two buckets by rank. A drift is a monotone trend in scale, and bucketing discards every distinction between the second observation and the ninety-ninth.
The rank correlation of each score against its position keeps the ordering of positions. A score’s position is no longer “early” or “late” but a number from one to two hundred, and the statistic asks whether large scores tend to sit at large positions.
The largest running departure from the mean keeps the positions and the magnitudes. It is the supremum of a partial-sum process, which is the shape a walk that reaches half its own maximum is built on, used here as a statistic rather than as an object of study. Scores are centred, accumulated, and the statistic is the largest the running sum ever gets. Under no drift the sum wanders back and forth around zero; under a drift it goes down and then up, because the early scores are systematically small.
The three reach four-in-five power at 4.31, 2.65 and 2.12, and that is the ordering by how much of the sample’s arrangement each one uses. It is not a coincidence and it is not deep: a statistic that is a function of less of the data cannot have more power against a departure the discarded part carries.
What the improvement is worth
The useful quantity is not where a detector fires but how much coverage is already gone when it does. That essay’s whole argument was about the size of the region a check cannot see, and a check that halves the departure it misses does not necessarily halve the damage inside that region.
Halving the undetected drift recovers 1.86 points of the 8.36. The reason is the shape of the coverage curve rather than any weakness in the detectors: coverage falls from 94.65% to 90.05% as the growth factor goes from zero to one, and from 90.05% to 86.10% across the whole range from one to nine. Most of the damage happens before any of these checks has meaningful power.
At a growth factor of 1 — the noise scale merely doubling across the sample — the coverage is already 90.05%, which is five points short, and the three detectors fire 33.4%, 43.5% and 49.0% of the time. Improving a detector moves the point at which it becomes reliable; it does not move the point at which the damage starts, and those are different points.
What the departure does to the interval, for reference
The coverage column the detectors are being scored against is that essay’s, recomputed here on the same draws that produce the power curves so that the two cannot come from different worlds.
Reading the two families of curve together is the whole of the argument. The coverage curve does most of its falling between growth factors of zero and one; the power curves do most of their rising between two and five. They are steep in different places, and the gap between those places is the region a practitioner sits in with a clean diagnostic and a broken interval.
The ordering, redone
With the best detector for each departure the comparison that earlier essay made can be made again, and this time neither half of it is hostage to a particular choice of test.
A factor of thirteen. That earlier essay read the same comparison off one detector each and found about nine; the best-detector version is larger, not smaller, because serial correlation’s best detector is very good and drift’s best detector is only somewhat better than the one it replaces.
So the reversal survives, and survives in a stronger form than it was stated in. It is not that practitioners happen to run a weak test for drift and a strong one for dependence. It is that the departure that damages a conformal interval most is intrinsically harder to see than the one that damages it least, at the sample sizes the interval is calibrated on.
It is worth being precise about what “best” means in that comparison, because it is doing real work. For drift it is the running-departure statistic, at 6.50 points. For serial correlation it is the lag-one check — the incumbent, which is already the right instrument for that departure and which the running-departure statistic does not beat. So the comparison is not stacked: each departure is watched by whichever of the four checks does best against it, and the specialist wins one and loses the other.
Why the intrinsic difficulty is intrinsic
That earlier essay gave a mechanism for the reversal and this essay is in a position to check it.
The argument was that a departure is easy to detect in proportion to how much it shows up in a statistic computed from many points at once. Serial correlation is a property of every adjacent pair, so a lag-one statistic on two hundred residuals aggregates two hundred pieces of evidence. A drift in scale is a slow change in a quantity — a spread — that is estimated much less precisely than a correlation.
If that is right, then no detector for drift can do very much better, because the limit is the information rather than the statistic. The measurements are consistent with it: three detectors spanning the range from “one bit per score” to “everything in the sample” separate by a factor of two, and the best is a general-purpose statistic with no special knowledge of the departure. A factor of two between the crudest and the best available summary is what a problem looks like when the statistic is not the binding constraint.
That figure also answers a question that earlier essay did not ask. A general detector exists. The running-departure statistic is nearly as good against dependence as the specialist lag-one check and far better against drift than anything else here, which makes it the one check to run if only one is going to be run — a conclusion that follows from measuring the two departures against the same four statistics and could not have come from either sweep alone.
Two points that were that essay’s and are now measured
The earlier essay made two claims it could not support with the one test it had, and both are now readings rather than arguments.
The cost of a departure and the visibility of a departure are independent quantities. That was stated as a mechanism and is now a measurement over four statistics and two departures: the four checks’ powers against drift span a factor of three and their powers against dependence span a factor of four, and the ordering is different in the two columns. Nothing about how damaging a departure is enters any of those numbers.
The check that matches the assumption is a check for stationarity in the sample’s order. That was a recommendation without an instrument attached. The running-departure statistic is one: it is a test against a positional trend in the scores and it needs no model of what kind of trend. It is not a test for independence, and it beats the test for independence against the departure that matters — which is the distinction between a symmetry and an absence of dependence arriving as a choice of diagnostic.
What a defensible check looks like
Run the running-departure statistic on the calibration scores. It costs a pass over the scores, it needs no model of the departure, and its critical value is a simulation anyone can do: draw the scores’ own values in a random order a few thousand times and take the 95th percentile of the statistic. That last step is what makes it distribution-free in the same sense the interval is, and it is why no table is quoted here.
Prefer a stated power to a verdict. “The calibration scores show no detectable drift” is a sentence two very different situations produce, and the thing that separates them is the drift the check would have found. It is computable before any data arrive, from the calibration size alone, in the same way the p-value a replication can expect is: the sweep above is the whole of the calculation.
Do not treat a clean check as a licence. The best available check is at 49.0% power where the coverage is already five points short. A check that passes rules out a large drift and says very little about a moderate one, and the quantity worth stating is the drift the check would have found rather than the fact that it did not find one.
And keep the two checks separate in what they are for. The lag-one check is nearly useless against drift — 27.2% at a ninefold growth — so running it and nothing else is the arrangement that earlier essay warned about, and it is not repaired by running it more carefully. Two checks, or one general one, is the minimum.
What this means for this field’s one exact repair
One consequence is worth drawing out, because it changes what the field’s repairs are for.
Of the three departures that earlier essay measured, only one has an exact repair — a shift in who is being predicted, which is weighted away, and which the previous essay showed survives an estimated weight. That departure is also the one with the easiest check, because the check needs only covariates.
Drift has no repair and the worst detection. Serial correlation has no repair and the best detection, and costs almost nothing. So the three departures sort cleanly by whether anything can be done about them, and the sorting has nothing to do with how bad they are: the one with a repair and a good check is a middling cost, the one with neither is the worst, and the one with a good check and no repair does not need one.
The practical reading is that the field’s apparatus is well aimed at the departure it can fix and blind to the departure it cannot, and that this is not a coincidence — a departure gets a repair when its structure is known well enough to write a weight for, and the same knowledge is what makes it checkable. A departure nobody can model is a departure nobody can detect, and neither fact protects the interval.
What is claimed here and what is not
The three sweeps share their draws. Each detector is evaluated on the same simulated calibration sets at each growth factor, so the comparison between them is paired and a difference of a few points of power is readable at two thousand draws. The coverage column is computed on those same draws, which is what lets a power and a coverage be read at the same setting rather than matched across two experiments.
Three detectors is not the space of detectors. A statistic built with knowledge of the departure’s shape — a likelihood ratio for a linear trend in log scale, say — would beat all three, and nothing here bounds by how much. What the sweep establishes is that the obvious improvements to the incumbent buy a factor of two, which is the quantity the deferral was about; it does not establish that no larger improvement exists.
All four critical values are simulated and none is asymptotic. Each is the 95th percentile of its own statistic over four thousand draws with no departure, on seeds disjoint from the sweep’s. That is what makes the three power curves comparable — their measured sizes are 5.3%, 4.8% and 4.6% — and it is also why the incumbent’s numbers differ slightly from that earlier essay’s, which read the same statistic against a normal approximation. The difference is under half a point of power and it moves the crossing from 4.05 to 4.31.
Four-in-five is a convention. Reading the crossing at a different power moves all three numbers and does not move their ordering, since the curves do not cross each other anywhere in the sweep. The coverage figures attached to the crossings move with it too, and the comparison between the two departures — thirteen to one — is robust to that because the coverage curves they are read on have very different slopes.
Two hundred calibration points throughout. Every power figure is at that size, and power against a positional departure grows with it — a thousand calibration points would move all three crossings left. The coverage curve does not move the same way, because the damage is a property of how much the scale has grown across the set rather than of how many points are in it, so a larger calibration set improves the detection half of the comparison without improving the damage half. That is a direction the argument would want followed, and it is not followed here.
And the drift is linear in the scale. A scale that grows steadily from one to nine across the calibration set is the simplest possible positional departure and it is the one the running sum is best suited to. A departure concentrated in the last tenth of the sample, or an oscillating one, would reorder the three detectors — the running sum is a test against a monotone alternative and pays for that elsewhere. The claim is about this departure.
What a practitioner is left with
Collecting the three essays in this part of the field that deal with departures gives a short procedure, and it is shorter than the amount of work behind it.
Run the running-departure statistic on the calibration scores, with a permutation critical value. It is the best of the four checks against a drifting scale and nearly the best against serial correlation, so it is the single check to run if only one is run.
Check the covariates of the points about to be predicted against the calibration set’s. That check needs no responses, and the departure it finds is the one with an exact repair — which survives an estimated weight.
And recalibrate rather than repair wherever fresh data allows. A new calibration set drawn under current conditions is exact and needs no ratio, no threshold and no check to have passed. The elaborate repairs is for the case where fresh calibration data cannot be got, which is the case worth naming rather than assuming — the same distinction a split conformal interval draws about its own training data.
Still open: a check whose power tracks the damage
That earlier essay ended by naming the instrument the field would want: a check whose power tracks the damage, so that a departure worth worrying about is a departure that fires it. Nothing here supplies one, and the measurements now say why it is hard rather than merely absent.
The damage curve is steepest at small departures and the power curves are flattest there. Those are the two shapes that would have to be brought together, and a better statistic moves only one of them. What would move the other is a different construction — an interval whose coverage degrades more gently under a positional departure, so that the region no check can see is a region where little is lost.
That is a question about the interval rather than about the diagnostic, and it has a shape: the quantile is taken over the whole calibration set, which is why a scale that has grown makes the test score’s rank non-uniform. A quantile taken over the recent part of the calibration set would inherit less of the early, quieter scores — at the cost of being a quantile over fewer of them, which is the sawtooth this field opened with. What that trade costs, and whether the resulting interval’s damage curve is flat where the detectors are blind, is the measurement this field has not made.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The score is the modelling — both name calibration set, conformal prediction, distribution-free, exchangeability, heteroskedasticity, nonconformity score
- How long a block a multiplier shares — both name autocorrelation, critical value, heteroskedasticity, monte carlo
- The charge that is not a sum — both name critical value, monte carlo, statistical power, supremum statistic
- The cliff that is a slope — both name autocorrelation, critical value, monte carlo, statistical power
- Two defects and one resampling — both name autocorrelation, critical value, heteroskedasticity, monte carlo
- What a multiplier cannot keep — both name autocorrelation, critical value, heteroskedasticity, monte carlo
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationCalibration setConformal predictionCritical valueDistribution-freeExchangeabilityExchangeability testHeteroskedasticityMarginal coverageMonte CarloNonconformity scoreRank correlationStatistical powerSupremum statistic