Coverage without a distribution

When the order matters

Three ways of breaking exchangeability cost 4.93, 11.07 and 1.07 points of coverage, and the ordering by cost is the reverse of the ordering by how soon a test would have caught them. The departure practitioners check for is the cheapest one.

Worth reading first: What the 95% refers to · Coverage from exchangeability alone.

Everything on this ladder has rested on one assumption, and only one. The guarantee survived a wrong model, a misaimed score, a split anywhere in a wide range, and a population with parts it knew nothing about. What it needs is that the calibration scores and the future score be exchangeable — that any ordering of them be as likely as any other.

Three ordinary ways of breaking that, measured on the same machinery, cost 4.93, 11.07 and 1.07 points of coverage. The useful finding is not the sizes but their order, because it is the reverse of the order in which a practitioner would have noticed. Serial correlation, the departure that gets tested for, is the one that costs almost nothing. A noise scale that grows across the sample, which is almost never tested for, costs ten times as much and is still being missed by its own test at a departure that has already cost 9.87 points.

What a scale that grows across the sample costs. What the interval covers when the noise scale grows across the sample, against how far the departure has gone, over 1500 draws at each setting. The coverage runs from 94.47% at no departure to 83.93% at the end of the sweep, a loss of 11.07%. The rank argument needs the 200 calibration scores and the test score to be exchangeable, and this is one of the three ways that fails. A test built for it reaches 80% power at 4.054, where the coverage is 85.13% — so 9.87% of the loss is inside the region such a test would have missed.
Fig. 1 What the interval covers when the noise scale grows across the sample, against how much larger it is at the end than at the start. Coverage falls from 94.47% at no drift to 83.93% at a ninefold growth.

The departure everybody tests for costs a point

Serial correlation is the failure of independence that has a whole field of diagnostics attached to it, and it is the one a careful analyst checks. Swept across first-order autocorrelations from zero to 0.95, the coverage of the interval falls from 95.00% to 93.93%. The largest loss anywhere in the sweep is 1.07 points.

What serial correlation in the errors costs. What the interval covers when the errors are serially correlated, against how far the departure has gone, over 1500 draws at each setting. The coverage runs from 95.00% at no departure to 93.93% at the end of the sweep, a loss of 1.07%. The rank argument needs the 200 calibration scores and the test score to be exchangeable, and this is one of the three ways that fails. A test built for it reaches 80% power at 0.207, where the coverage is 94.01% — so 0.99% of the loss is inside the region such a test would have missed.
Fig. 2 The same reading under serially correlated errors. The whole column spans a point, and the worst setting — an autocorrelation of 0.95, where consecutive errors are nearly identical — covers 93.93%.

That is a striking number beside what dependence does elsewhere on this site. Fifty correlated observations worth about six takes a 95% interval for a mean down to 47% at an autocorrelation of 0.8. Here the same autocorrelation costs under a point.

The difference is what the two intervals are built from. An interval for a mean divides by the square root of the sample size, and that division is a claim about how much independent information the sample carries — a claim dependence falsifies directly. A conformal interval divides by nothing. It takes an order statistic of the calibration scores, and serial correlation does not change what those scores look like as a set: the marginal distribution of each residual is unchanged, and only their arrangement in time is not exchangeable. The rank argument is damaged rather than destroyed, because the thing it counts is close to what it would have counted anyway.

What a 95% interval covers when the observations are dependent, n = 50. Each point is 3,000 series of 50 observations. The ordinary interval covers 94.6% at φ = 0 and 28.1% at φ = 0.92. Dividing by the effective sample size n(1 − φ)/(1 + φ) instead of by n brings it back to 82.1%.
Fig. 3 The same departure priced for an interval for a mean: coverage against the strength of the dependence, with the repaired interval beside it. The scale of the vertical axis there is what the comparison rests on.

The contrast is worth holding onto, because it is the one case on this ladder where the distribution-free construction is not merely as good as the modelled one but dramatically better, and the reason is the same property that made it look weak everywhere else. An interval that reads a variance is exposed to everything that makes a variance the wrong number. An interval that reads only an ordering is exposed only to what disturbs the ordering.

And a test for it works extremely well. The lag-one autocorrelation of the calibration residuals, compared against its own null standard error, reaches 80% power at an autocorrelation of 0.2071 — where the coverage is still 94.01% — and is at 78.6% power at 0.2 and 100.0% at 0.6 and above.

What a test for it would have caught. How often the lag-one autocorrelation of the calibration residuals against its null standard error rejects, against the size of the departure, over 1500 draws at each setting. At no departure at all it fires 4.9% of the time, which is the size of a test run at five per cent. It reaches 80% power at a departure of 0.207, where the coverage is already 94.01%. That is the number to carry: a departure smaller than this one is a departure a practitioner would not have found, so the coverage at it is what undetectable non-exchangeability costs — and how much that is depends entirely on which assumption broke.
Fig. 4 How often the lag-one test on the calibration residuals fires, against the autocorrelation. It reaches four times in five at 0.207, and never misses beyond 0.6.

So the whole of the undetectable region costs about one point, and the whole of the rest is caught. That is the best possible arrangement of a defect, and it is why this departure is the one to stop worrying about — not because dependence is harmless in general, but because this construction happens to be nearly immune to it and the standard check happens to be sensitive to it.

Exchangeability is weaker than independence, and that is why it is not safe

The assumption being broken is worth stating precisely, because it is weaker than the one most diagnostics are built for, and the weakness cuts both ways.

Independent and identically distributed observations are exchangeable. So are many things that are not independent: a sample drawn without replacement from a finite population, a sequence whose members share a common unobserved parameter, any set whose joint distribution is symmetric in its arguments. Exchangeability is a symmetry, not an absence of dependence, and a great deal of dependent data satisfies it — which is why a conformal interval is not required to be run on independent data and why the serial-correlation result above is less surprising than it first looks.

What exchangeability rules out is any structure attached to position. If the index of an observation carries information about it, the orderings are not equally likely and the argument is gone. That is why the drift result is the severe one: a scale that grows with the index is position information in its purest form.

The practical consequence is that the standard diagnostics point at the wrong thing. A test for independence is a test for a condition that is stronger than what is needed, so it can fail on data the guarantee holds perfectly well for; and it can pass on data whose defect is positional but not serial, which is the case that costs eleven points. The check that matches the assumption is a check for stationarity in the sample’s order rather than for independence — the same distinction that separates a series with a level to return to from one without, arriving here as a question about which diagnostic to run rather than about which model to fit.

The departure nobody tests for costs eleven

A noise scale that grows across the sample is a different matter. Coverage falls from 94.47% with no drift to 91.00% when the scale is 1.4 times larger at the end than at the start, 85.80% at four times larger, and 83.93% at nine times — a loss of 11.07 points.

The mechanism is direct. The calibration scores come from the earlier, quieter part of the sample and the test point comes from the later, noisier part, so the test score is systematically larger than the scores it is being ranked against. Its rank is not uniform; it is biased towards the top, and the interval built at the ninety-fifth percentile of a quieter population misses a noisier point far more than one time in twenty.

What a test for it would have caught. How often a rank comparison of the first half of the calibration scores against the second rejects, against the size of the departure, over 1500 draws at each setting. At no departure at all it fires 4.7% of the time, which is the size of a test run at five per cent. It reaches 80% power at a departure of 4.054, where the coverage is already 85.13%. That is the number to carry: a departure smaller than this one is a departure a practitioner would not have found, so the coverage at it is what undetectable non-exchangeability costs — and how much that is depends entirely on which assumption broke.
Fig. 5 The power of a rank comparison between the first half of the calibration scores and the second, against the growth factor. It reaches four times in five only at a growth factor of 4.05, by which point the coverage is 85.13%.

The test built for it is a rank comparison of the first half of the calibration scores against the second — a reasonable thing to run and, on the face of it, well suited: if the scale grows, the second half’s scores are stochastically larger, and Mann–Whitney is the standard instrument for exactly that. It reaches 80% power at a growth factor of 4.0541. At a growth factor of 1 — the scale doubling across the sample — it fires 34.1% of the time, and even at a ninefold growth it is only at 91.6%.

So the loss inside the region such a test would have missed is 9.87 points. A practitioner running that check, seeing it pass, and reporting a 95% interval is reporting an interval covering about 85% of the time, with a diagnostic in hand that was run correctly and came back clean.

Why the two orderings are opposite

The reversal is not a coincidence of these two sweeps and it has a stateable cause.

A departure damages coverage in proportion to how much it makes the test score’s rank non-uniform. Serial correlation leaves each score’s marginal distribution alone and scrambles only the joint arrangement, so a single test score is still a draw from the same population as the calibration scores; its rank is nearly uniform and the coverage barely moves. Drift changes the marginal distribution the test score is drawn from, so its rank is uniform against nothing.

A departure is easy to detect in proportion to how much it shows up in a statistic computed from many points at once. Serial correlation is a property of every adjacent pair in the sample, so a lag-one statistic on two hundred residuals aggregates two hundred pieces of evidence about it. A drift in scale is a difference between two halves of the sample, which is two estimates of a spread from a hundred points each — and a spread is estimated much less precisely than a correlation, so the same amount of departure yields far less evidence.

The two mechanisms are unrelated, and that is exactly why the orderings can be opposite: the thing that makes a departure costly is not the thing that makes it visible. There is no reason to expect the cheap failures and the visible failures to be the same failures, and here they are exact opposites.

A shift in who is predicted, and the ninety per cent floor returning

The third departure is a change in which members of the population arrive to be predicted, with the relationship between covariates and response unchanged. With half the test points drawn from the noisier group — the calibration mix — the coverage is 94.40%. With 70% of them from it, 92.33%. With all of them, 90.07%, a loss of 4.93 points.

That last number is the arithmetic floor of the third rung arriving as a marginal loss. The single interval covers the noisy group 90% of the time whatever the mix; shifting the mix to that group alone makes the marginal coverage equal the noisy group’s conditional coverage, and 90.07% is what the closed form said it would be. The departure did not create a new failure. It exposed one the guarantee had been averaging away.

A two-proportion comparison of the calibration set’s own group mix against a batch of unlabelled test covariates reaches 80% power at a mix of 0.6781, where the coverage is 92.61% — so the loss inside the undetectable region is 2.39 points. That test is the easiest of the three to run, because it needs no responses at all: only the covariates of the points about to be predicted, which are in hand by definition.

The one departure with a repair

Weighting each calibration score by the likelihood ratio between the shifted distribution and the calibration one restores the promise, and it restores it exactly rather than approximately. The weighted interval covers 95.20% at a 60% shift and 94.93% at a complete one — flat across the whole sweep, where the unweighted interval falls five points.

A known shift is repairable and an unknown one is not. What the interval covers when the future observations come from a different mix of the two groups than the calibration set did, over 1500 draws at each setting. Calibrated as though nothing had shifted, the coverage falls from 94.40% to 90.07% — which is the conditional spread of the noisy group arriving as a marginal loss, because the covariate that shifted is the one the coverage was uneven in. Weighting each calibration score by the known likelihood ratio, with the test point's own mass in the weighted quantile, holds the promise throughout: 94.40%, 95.20%, 94.67%, 94.60%, 95.20%, 94.93%. What makes that work is the word known. The weights here were constructed rather than estimated.
Fig. 6 Coverage under a shift in who is being predicted, calibrated as though nothing had shifted and weighted by the known likelihood ratio. The weighted line holds the promise across the whole sweep.

The load-bearing part of the machinery is the point mass. The weighted quantile puts the test score’s own weight into the same total the calibration weights are normalised against, and without it the weights sum to less than one at the largest calibration score and the claim fails at precisely the place the promise is about. That is the same correction as the plus one in a randomisation p-value, in weighted form, and it is easy to drop for the same reason: it changes almost nothing except in the tail, and the tail is the whole statement.

The load-bearing word is known. The likelihood ratio here was constructed rather than estimated — the shift was built by the same code that supplied the weights, so the weights are exactly right by construction. That is the situation of a weighting scheme that needs only a ratio and not the situation anybody deploying this is in.

What could have made these numbers wrong

The powers are properties of three chosen tests, not of the departures. This is the essay’s central claim and it is also its largest qualification: nothing measured here says a better detector for a drifting scale does not exist. A cumulative sum of the calibration scores against position, or a regression of score on index, would very likely do better than a two-sample rank comparison, because both use the ordering rather than throwing it away into two buckets. The honest statement is therefore about the tests a practitioner actually runs, and if a routine drift detector became standard the ordering in this essay would change. That is a measurement worth making and it is not made here.

Coverage understates what drift costs, because the interval also widens. At a ninefold growth the mean width is 26.630 against 4.005 at no drift. So the reported 83.93% is being achieved by an interval more than six times wider than the one the same procedure gives on exchangeable data, and a comparison that held width fixed would show a far larger loss. This is the direction that matters: the drift figure is a lower bound on the damage, not an upper one.

A single detection rate is not a power curve. Each point is 1,500 draws, so a power of 78.6% carries a standard error of about a point, and the growth factor at which 80% is crossed is read by interpolating between two settings rather than by solving anything. The figure 4.0541 should be read as “about four”, and the claim it supports — that the crossing happens well after the coverage has collapsed — is robust to a factor of a half in either direction, since the coverage at a growth of 1 is already 88.87%.

The three sweeps could have differed in something other than the departure. They do not: the same training size, calibration size, miss rate and absolute-residual score are used in all three, and at the baseline setting of each — no drift, no correlation, the calibration mix — the coverage reads 94.47%, 95.00% and 94.40%, all within a standard error or two of the 95.02% the calibration size promises. Those three baseline readings are the calibration of the whole essay, and a sweep whose zero point was wrong would be measuring its own bug.

And the three departures are not on a common scale. A growth factor of 8 and an autocorrelation of 0.95 and a complete shift in the test population are not equally large departures in any sense that could be defined, so the three maximum losses are not directly comparable and nothing above compares them. What is compared is each departure’s cost against its own detectability, which is a comparison within a single sweep and needs no common currency.

What a practitioner is left holding

The three sweeps convert into a short and slightly uncomfortable set of instructions.

Check the ordering, not the independence. A lag-one autocorrelation on the calibration residuals is cheap, powerful and nearly pointless here, because what it finds costs a point. A comparison of the early calibration scores against the late ones is the check that matches what actually damages the guarantee, and it is the weak one — so it should be run at a generous size rather than at five per cent, since the cost of a false alarm is recalibrating and the cost of a miss is ten points of coverage.

Check the covariates of what is about to be predicted. That test needs no responses, which makes it the only one of the three that can be run before any outcome is known, and a shift in who is being predicted is the departure with a repair. It is also the departure whose damage is bounded by something computable in advance: the worst it can do is push the marginal coverage down to the conditional coverage of whichever group the shift favours, and that number is available from the group’s share alone.

And recalibrate rather than repair. The three departures all have the same root — the calibration set no longer represents what is being predicted — and the cheapest response to any of them is a fresh calibration set drawn from the current conditions. That costs observations and nothing else. It requires no likelihood ratio, no model of the drift, and no test to have passed, and it is exact the moment the new calibration set is exchangeable with the new test points. The elaborate repairs above are for the case where fresh calibration data cannot be got, which is the case worth naming rather than the case worth assuming.

Where the ladder stops

Two things this field named and did not measure, both of them at this rung.

An estimated likelihood ratio. The repair above is exact with a known ratio and there is no reason to think it degrades gracefully with an estimated one. Whether a misspecified weight loses coverage smoothly or falls off a cliff is the obvious next measurement, and it decides whether weighted conformal is a practical tool or a demonstration.

A better drift detector. The ordering this essay is about — cheap failures tested for, expensive ones not — rests on the power curve of one rank test. Measuring a detector built for the ordering rather than for a two-way split would move the second finding, and it is the deferral most worth taking, because it is the one that could overturn the essay’s conclusion rather than extend it.

What survives either result is the shape of the argument. A guarantee that holds under one assumption is only as good as the chance of noticing that the assumption has failed, and the two are independent quantities: the cost of a departure and the visibility of a departure have no reason to be related, and here they are inversely related. The instrument the field would want is not a better interval. It is a check whose power tracks the damage — a diagnostic run before the standard error rather than after — and nothing in the construction supplies one, because the construction’s whole strength is that it never had to look at the data’s shape at all.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationCalibration setConformal predictionCovariate shiftCoverageDistribution-freeExchangeabilityExchangeability testLikelihood ratioMarginal coverageModel misspecificationNonconformity scoreStatistical powerWeighted conformal