A quantile over the recent scores
Worth reading first: Coverage from exchangeability alone.
A detector built for the ordering settled one question and sharpened another. The best of three checks for a noise scale that grows across the sample — the largest running departure of the calibration scores from their mean — reaches 80% power at about half the growth the standard check needs. Even so, coverage has already fallen by more than six points before that check fires reliably. The damage curve is steepest where every power curve is flattest. A better statistic moves the power curve, and that essay concluded that only a different interval would move the other one: an interval whose coverage degrades gently under a positional departure, so that the region no check can see is a region where little is lost.
The construction it named is the obvious one. Split conformal prediction takes a quantile over every calibration score, and when the scale has grown, most of those scores come from a quieter time than the test point’s. Taking the quantile over only the most recent scores should inherit less of that quiet. The cost is that a quantile over fewer scores is coarser and noisier, which is the sawtooth this field began with. This essay measures both sides of that trade on the same draws, and then the two things a careful analyst would try next: weighting old scores down rather than dropping them, and switching to the recent scores only when a detector says something has moved.
The setting, and why the windows are 199, 99, 59, 39 and 19
The world is the one the departures were measured in. A straight line with noise; a hundred points to fit the line, two hundred to calibrate the absolute residuals, and one more point to predict. Under drift, the noise scale grows linearly from 1 at the first point to at the last, so the test point is the noisiest in the sample. Each setting is drawn four thousand times, which puts the standard error of a coverage near 95% at 0.34 points.
A window is the last calibration scores, and the interval uses their -th smallest. Any scores and the test score are exchangeable whenever all of them are, so every window keeps the rank guarantee when nothing has drifted. But each window keeps a different version of it. Exact coverage is , which is 95.08% at sixty scores and 96.61% at fifty-eight — so a comparison between windows of convenient round sizes would mix the drift’s effect with each window’s own sawtooth. The windows here are chosen so that is a whole number: 199, 99, 59, 39 and 19 all cover exactly 95% with exchangeable data. Any gap between two of them under drift is drift.
The damage curve flattens
With all 199 scores, coverage falls quickly and keeps falling: 89.83% at a growth of 1, 86.78% at 3, and 85.63% at 8, the largest growth in the sweep. The last 99 scores lose about half as much, bottoming at 90.88%. The last 39 never fall below 93.42% anywhere in the sweep, and the last 19 never below 93.53%. Their curves are nearly flat. The shape of the damage changes as well as its size: the full set’s loss keeps growing with the drift, while a short window’s loss is roughly constant once the drift is underway, because the window’s scores are always close to the test point in time and the gap between their scale and the test point’s grows only in proportion to how far back the window reaches.
That is the property the detector essay asked for. At a growth of 2.12, where the best detector first reaches 80% power, all 199 scores cover 88.11% — about seven points short. The last 39 cover 94.02% at the same growth. Every growth to the left of the vertical line is a drift that no check finds reliably, and there the short window’s coverage stays within about a point and a half of nominal, against more than six for the full set. The region the detectors are blind to has become a region where little is lost. That is what a better statistic could not do.
The mechanism is arithmetic. With all 199 scores, the quantile is drawn from points whose scales run from about 1.34 to 1.997 at a growth of 1, against a test point at 2. With the last 39 they run from about 1.87 to 1.997. The test point is still slightly noisier than every score it is ranked against, so the window does not remove the bias, but it bounds it by the drift over 39 points rather than over 199.
What the window costs when nothing has drifted
The price is paid when the data are exchangeable, which is the case the guarantee was built for and the case an analyst hopes to be in.
All 199 scores give a mean width of 3.996. The last 99 give 4.037, the last 59 4.086, the last 39 4.143 and the last 19 4.329. So the last 39 cost 3.7% of width and the last 19 cost 8.3%. The mean understates the cost, though. A quantile of fewer scores varies more from draw to draw: the width’s spread is 0.271 with all 199, 0.630 with the last 39 and 0.955 with the last 19. A practitioner sees one interval, not the average, and the last 39 make that one interval more than twice as variable as the full set does.
The trade can be read off the two figures together. For a 3.7% widening on average and more than double the variability, the last 39 cap drift’s damage at 1.6 points where the full set allows 9.4. The last 19 buy almost nothing more against drift for more than twice the cost. A window near forty is where the curve bends, in this world, for an interval at 95%. The general rule is not the number but the shape: the coverage gained flattens quickly as the window shrinks, while the width and its spread keep rising.
Each window still covers exactly 95% with exchangeable data, so none of this is paid in validity. That is the difference from the cost of splitting a sample between fitting and calibrating, where the width cost is borne for a better fit. Here it is borne as insurance.
Weighting old scores down instead of dropping them
A window is a step: full weight inside, none outside. The smoother alternative gives every score a weight of to the power of its age — the most recent at , the one before at , and so on — and takes the weighted quantile with the test score carrying weight one, as the weighted repair for a shift did with likelihood ratios. Two decays are measured: 0.99, whose effective sample size is about 199, and 0.97, about 66.
The decayed weightings behave differently from windows, and the difference is in their guarantee. With exchangeable data a window covers exactly 95%. A decayed weighting covers more: 95.30% at a decay of 0.99 and 97.08% at 0.97. The test score’s weight of one is a larger share of a smaller effective set, so it pushes the quantile up. Its guarantee is the bound of nonexchangeable conformal prediction: at least 95% with exchangeable data, less a weighted measure of how far each calibration score is from exchangeable with the test score otherwise. The over-coverage is the slack in that bound when nothing has drifted. At a decay of 0.97 the width is 4.664, 16.7% above the full set’s, which is twice the cost of the last 19 scores.
What the over-coverage buys is a curve that never dips below 95.70% across the whole drift sweep. Under drift, a decay of 0.97 is the safest interval measured here. A decay of 0.99 is nearly the full set again, bottoming at 91.20%. The smooth construction is not better than the window. It is the same trade with a different dial, and the dial it offers is coverage held above nominal rather than at it.
The same windows under serially correlated errors
A short window was chosen against drift, and drift is not the only way exchangeability fails. The departure the detector essay found cheapest to miss was serial correlation in the errors. It was cheap because the full set of 199 scores averages over many stretches of correlated noise. A window averages over fewer.
The prediction holds. Up to an autocorrelation of 0.3 the windows and the full set are indistinguishable, and at 0.6 only the last 19 have begun to slip, to 93.92%. At 0.8 all 199 scores cover 94.17%, the last 39 92.85% and the last 19 92.47%. At 0.95 all 199 cover 93.53%, the last 39 91.30% and the last 19 90.95%. The shorter the window, the more its quantile reflects whatever the most recent errors happened to be. With strongly correlated errors, 39 consecutive scores carry far less information about the spread of the noise than 39 independent ones would, because they are a few long excursions rather than many short ones. The quantile of a short window is then an estimate of a recent excursion’s size, and the test point, which is correlated with that excursion, still falls outside it more often.
A window is therefore not a free repair for positional departures in general. It trades a loss of up to 9.4 points under one departure for an extra loss of 2.2 points under another. Which side of that trade an application is on depends on which departure it faces, and that is the question the detectors were supposed to answer.
Switching only when the detector fires
The natural compromise is to use all the scores unless something looks wrong. Run the best drift detector on the calibration scores; if it fires at 5%, take the quantile over the last 39, and otherwise over all 199. With exchangeable data this costs almost nothing, since the detector fires on 5% of draws, and it seems to get the window’s protection exactly when it is needed.
It does not. Under drift the gated rule covers 93.17% at a growth of 0.5, where the detector fires 21.9% of the time, and 91.67% at a growth of 1, where it fires 48.9% of the time. That is worse than the last 39 used throughout, which cover 93.65% at that growth. The switching rule’s worst coverage sits where the detector is unsure. There, half the draws keep the full set and lose the full set’s coverage, and the other half switch. The gated rule inherits the detector’s blindness exactly where the window was meant to cover for it.
Under serial correlation the switching rule is worse again. The running-departure statistic was chosen because it is nearly the best check for both departures, and that is the problem here: it cannot tell them apart. At an autocorrelation of 0.95 and no drift at all it fires on 97.5% of draws, switches to the window, and covers 91.33%, against 93.53% for never switching. The rule applies the drift repair to a departure that the drift repair makes worse, and it does so most reliably when the departure is strongest.
This is the same failure a test for a flat point run before the interval showed in a different field. Choosing the procedure by a test on the same data is a procedure of its own, with its own coverage, and that coverage is usually worse than either of the procedures it chooses between. The pretest borrows the weakness of whichever branch it takes, at the settings where it is least sure which branch that should be.
Two other ways to put time into the interval
A window is not the only way to tell a conformal interval that the data have an order. Two others already sit in this field, and comparing them shows what kind of repair a window is.
The first changes the score. The score is the modelling showed that dividing each residual by an estimate of its own spread changes how adaptive the interval is while leaving the guarantee untouched. Under drift, the spread that matters is the spread at each point’s time, and a score divided by a fitted scale-in-time would make the scores exchangeable again if the fit were right. That repair models the departure. It works exactly as well as the model of the scale, and a linear drift fitted as linear would be repaired completely. The window models nothing. It assumes only that recent points are more like the next one than old points are, which is weaker and correspondingly less efficient when the stronger assumption holds.
The second changes the guarantee. Marginal is not conditional split the calibration set by group and took a quantile within each, so that coverage held group by group rather than on average. A window is that construction with time as the grouping and only one group kept, the most recent. Its guarantee is the marginal one over draws, as before, but the scores it ranks against are the ones from the test point’s own neighbourhood in time. The price is the same in both cases: a quantile over fewer scores is coarser and noisier, and the floor on how few scores a 95% quantile can be taken over is nineteen.
Neither framing changes the ordering of the three departures by cost. A shift in who is being predicted still has an exact repair by weighting. A drifting scale now has a cheap partial one that needs no detector. Serial correlation, the departure that cost least with all the scores, is the one a window makes more expensive.
What an interval that may face drift should state
Which scores the quantile was taken over. All of them, the last , or weights decaying with age. These are different intervals with different behaviour under the same departure, and a reader cannot reconstruct which was used from the coverage claim alone.
What the window costs when nothing has moved. For the last 39 here, 3.7% of mean width and a spread more than twice as large. That cost is paid every time, and the protection only under drift.
Which departure it is insurance against. A short window protects against a scale that changes smoothly in time and costs coverage under serially correlated errors. A decayed weighting at 0.97 protects against both in this world, at the price of over-covering by two points when nothing has drifted.
Whether the choice was made before looking. A window chosen in advance has the coverage measured here. A window chosen because a detector fired has the gated rule’s coverage, which is worse than either fixed choice.
What is claimed here and what is not
Exact by the rank argument. Every window whose is a whole number covers exactly 95% with exchangeable data, since any calibration scores and the test score are exchangeable when all of them are. The four thousand draws per setting agree with that within their counting error.
Counted. Every coverage and width under drift and serial correlation, for five windows, two decays and the gated rule, on the same draws at each setting, so differences between constructions are not differences between samples. The detector’s critical value is the one the detector essay calibrated on its own null draws.
Not claimed. The drift here is linear in time. A scale that jumps once, or cycles, would reward a window differently: a window entirely after a jump recovers fully, and one that straddles it does not. Nor is a window the only construction; adaptive conformal methods that move the quantile after every miss are a different family and are not measured here.
Still open: a window that knows the departure
The window and the decay each protect against drift and cost something under correlation, and the switching rule that tried to choose between them could not tell the two departures apart. The two departures leave different fingerprints. A drifting scale moves the scores’ level slowly and in one direction. Serial correlation leaves the level unchanged on average but makes neighbouring scores alike. A statistic that separates the two would let a rule take the window under drift and keep the full set under correlation. The lag-one autocorrelation of the scores measured against a running mean, or a statistic computed on differences of neighbouring scores, are the obvious candidates. Whether any of them separates the two departures in the region where both are hard to see, and whether a rule built on it beats the decay of 0.97 that needs no choice at all, is a measurement this field can make and has not yet made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name coverage, monte carlo, statistical power
- A horizon chosen after looking — both name coverage, monte carlo, statistical power
- A simulation that stops when it looks settled — both name coverage, monte carlo, statistical power
- Five times in six — both name coverage, prediction interval, statistical power
- A band allowed a few misses — both name coverage, prediction interval
- A block size that changes — both name coverage, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Calibration setConformal predictionCoverageExchangeabilityMonte CarloNonconformity scorePrediction intervalSerial correlationStatistical power