The weight that has to be estimated
Worth reading first: Coverage from exchangeability alone.
The exact repair for a shift weights each calibration score by the likelihood ratio between the population the test points come from and the population the calibration points came from, and it restores the promise flatly: 95.20% at a partial shift and 94.93% at a complete one, where the unweighted interval falls five points.
The essay that measured it named the word the repair turns on. The ratio was known — built by the same code that built the shift — which is the position a weighting scheme that needs only a ratio is also written from, and is nobody’s situation. What happens with an estimated ratio was left as the measurement that would decide whether weighted conformal is a practical tool or a demonstration, and the two possibilities named were that it degrades smoothly or that it falls off a cliff.
It does neither, and the shape it does have is more useful than either.
Overstating the ratio costs nothing. A weight sixteen times too large still covers 96.13%, which is above the promise. Understating it costs, and it costs steeply once the understatement passes about a half. The curve is smooth, and it is one-sided.
Why the two directions are not symmetric
The asymmetry has a mechanism and it is visible in what the weights do to the quantile.
A weighted conformal interval takes the weighted quantile of the calibration scores: it walks them in order and stops where the accumulated weight first reaches 1 − α of the total. A group whose weight is raised contributes more mass early in that walk, so the walk stops later, and a later stop is a larger score and a wider interval.
The test points here come mostly from the noisier group, and the noisier group’s scores are the large ones. Raising their weight pushes the quantile out into their part of the distribution, which is where the test point is going to land. Lowering it pulls the quantile back towards the quiet group’s scores, which is where the test point is not.
So an overstatement widens an interval that was already wide enough and an understatement narrows one that was not. The same one-sidedness runs through every conservative interval: a promise stated as a floor is broken in one direction only, and an error that overshoots it is invisible to the promise and visible in the width. That is the ordinary direction — a wrong answer that errs towards caution is the safe one — and it is worth stating because the weight is the only dial the method has.
The price of the safe direction is small and it does not run away. From the exact ratio to sixteen times it, the width rises from 11.54 to 12.17 — 5.5% — and the curve is flattening, because past a point the quiet group’s scores have been weighted out of the quantile altogether and raising the loud group’s weight further changes nothing.
Where the failure comes from, when it comes
The left-hand end of the curve is worth looking at directly, because 67.90% is a spectacular failure for a method whose whole claim is a distribution-free guarantee.
At a factor of a thirtieth the loud group’s calibration scores carry almost no weight, so the quantile is taken essentially over the quiet group’s scores alone. The quiet group’s noise scale is a third of the loud group’s, so that quantile is about a third of the right one, and the interval is 5.43 wide against the exact repair’s 11.54. Four test points in five are drawn from the loud group, and an interval built for the quiet one misses them.
That is the same failure the conditional floor describes, reached from the opposite direction. There the interval was correct marginally and wrong for a group; here the weights have been set so that the interval is correct for the wrong group, and the marginal coverage inherits it. The mechanism in both is that a conformal interval is only ever calibrated against the population its scores came from, and the weights are what decides which population that is.
What an estimate actually has to get right
Given that shape, the question about an estimated ratio changes. It is not “how accurate does the estimate have to be” but “how often does it land on the wrong side of a threshold”, and the threshold is visible on the curve: coverage is inside the promise down to about half the true ratio and falls away below it.
Under a shift in the mix of an observed group the ratio is two shares — the test mix and the calibration mix — so a practitioner estimates the first from a batch of unlabelled test covariates and the second from the calibration set. Both are binomial, and the requirement “clear half the true ratio” is a binomial tail.
Five unlabelled covariates are enough 99.328% of the time. That is not because five is a good sample for estimating a proportion — it is a terrible one, with a standard error of 0.18, and an interval built on it would be the kind whose failure has already been measured — but because the estimate does not have to be good. It has to be on the right side of a number, and at a test share of 0.8 a batch of five puts it there unless it draws at most one point from the majority group, which happens 0.672% of the time.
The threshold is also the reason the two possibilities the previous essay named were the wrong two. “Smooth” and “cliff” are both descriptions of how a curve falls, and they share an assumption — that the interesting behaviour is a decline on both sides of the right answer. Here one side does not decline at all, so the question of how fast it declines has no content, and the question that does have content is about a threshold on the other side. Getting the shape of a curve wrong is a cheap mistake to make in advance and an expensive one to keep, because it decides what gets measured.
The measurement
Simulating the whole procedure gives the same answer by counting rather than by arithmetic, which is the check this site applies to anything computed in closed form.
The estimated column sits on the exact column at every batch size in the sweep, and the unweighted column sits two and a half points below both. There is no degradation to report, and the flatness is the finding rather than the absence of one: the two routes to it — a simulated coverage and a binomial tail — agree about where the threshold is and agree that a batch of five clears it.
It also settles the practical question in the direction the previous essay was doubtful about. The repair needs covariates and not responses, so the batch is exactly the thing a practitioner already has: the points about to be predicted. A shift can be repaired before any outcome is known, from data whose whole purpose is to be predicted.
The interval this buys, and what it does not fix
One caution belongs beside the result rather than in the qualifications, because it is a property of the repair and not of the estimate.
Weighting restores the marginal promise under the shift. It does not make the interval conditionally valid, and the group whose coverage was 90% before the shift still has 90% coverage after it. What the weights do is stop the marginal number from inheriting the noisy group’s conditional number as the mix moves towards that group, which is what the unweighted 92.80% is: the marginal coverage sliding down towards a conditional coverage that was always there.
So the repair is honest about what it repairs. A reader who wanted the noisy group covered 95% of the time is not served by any of the intervals in this essay, and the arithmetic floor that says so is the third essay’s and is not moved by any weight.
What this does not extend to
The result is clean and it is about one kind of shift, and the boundary matters more here than usual because the clean case is the one everybody reaches for.
The ratio here is two numbers, because the shift is in the mix of an observed binary group. A shift in a continuous covariate has a ratio that is a function of that covariate, and estimating a function is not estimating a proportion: the estimate can be wrong by different amounts in different places, the threshold argument above applies pointwise rather than once, and nothing in the binomial tail says anything about it. That case is the one the method is usually deployed in and none of this prices it.
What does carry across is the asymmetry, because it is a property of the weighted quantile rather than of the shift — the same way a score that is valid and useless is a statement about the construction rather than about any particular score. A ratio that overstates the weight of the region the test point comes from widens the interval; one that understates it narrows the interval. So a function estimated with a bias towards the flat — a smoothed density ratio, a clipped one, a ratio truncated at a ceiling — errs in the direction that costs width, and a method that clips the ratio from above is trading coverage for width in exactly the wrong direction. Clip from below, not from above is the transferable rule, and it follows from the shape of the curve rather than from the measurement of the binary case.
What a defensible use looks like
Bound the ratio from below rather than from above. Software that implements weighted conformal routinely clips extreme likelihood ratios, and clipping is usually described as a stability measure with no direction attached to it. It has one. Clipping from above caps the weight of the region the test point comes from, which is the understating direction, and costs coverage; clipping from below raises a weight that was small, which costs width. Two clips described by the same word do opposite things here, and only one of them is safe.
Weight even when the ratio is a guess. The unweighted interval covers 92.80% and every weighted one in the sweep covers at least 94.7%. A repair applied with a poor estimate is not worse than no repair unless the estimate is off by more than a factor of two in the dangerous direction, and it is better than no repair everywhere else.
Err upwards. If the ratio has to be rounded, guessed or clipped, take the larger value. It costs width, the width cost is bounded at a few per cent, and the alternative costs coverage with no bound at all.
Use the calibration set’s own mix rather than a design value. The denominator of the ratio is the calibration population’s share, and it is tempting to use the value the sampling design intended — a half, here. The calibration set’s own observed share is a better estimate of what its scores actually represent, costs nothing, and is what the measured column above uses. It is the same preference an estimator that reads the realised assignment expresses in a different field: the quantity that matters is what happened, not what was planned.
And report the batch size the ratio came from. It is the only quantity in the repair a reader can use to price it, and it is free to state. A reader given the weighted interval alone cannot tell a ratio taken from theory, a ratio estimated from five points and a ratio estimated from five hundred apart, and the measurement above says those are three different objects with the same name.
Recalibrate instead, where the data allows it. The previous essay’s last recommendation still dominates everything here: a fresh calibration set drawn under current conditions needs no ratio, no batch and no threshold, and is exact. The whole of this essay is about the case where fresh calibration data cannot be got, which is worth naming rather than assuming — and the finding is that in that case a crude estimate is enough.
What is claimed here and what is not
One shift, at one size. The sweep is at a test population 80% drawn from the noisier group, against a calibration set that is half and half. A larger shift makes the weights more extreme and the unweighted interval worse; a smaller one makes both less interesting. The asymmetry’s direction does not depend on the size, since it follows from what raising a weight does to a quantile, and the position of the half-way threshold does.
Two hundred calibration points. The weighted quantile is a quantile of two hundred scores, and that is what makes it insensitive to moderate changes in the weights. At forty calibration points the same sweep is noisier and the quantile can run off the end of the calibration set entirely, returning an infinite interval — which covers, and says nothing. The sweep is at a size where the interval is finite in every draw.
The floor on the estimated shares is arbitrary and load-bearing. An estimated share of exactly zero or exactly one makes one group’s weight zero or infinite, and the estimates here are clipped to lie between 0.02 and 0.98. Without a clip a batch that happened to draw no points from the quiet group would produce a quantile over the loud scores alone, which at this shift is nearly the right answer and at a smaller shift would not be. The clip is stated rather than tuned, and nothing in the sweep sits near it.
The batch is drawn from the test distribution and is assumed clean. The covariates in it are generated from the same shifted population the test point comes from, so the estimate is unbiased for the quantity it is estimating. A batch collected some other way — an earlier period, a convenience sample, whatever happened to be on hand — estimates something else, and the threshold argument applies to the wrong number. What is measured here is the sampling error in a batch that is the right batch.
And “no coverage cost” means none that this many draws can see. At three thousand draws a coverage difference of half a point is about the size of a standard error, so the flat top of the misspecification curve is flat to within that. What the curve establishes is that the cost of overstating is small and bounded, not that it is exactly zero.
The shape of the result, stated once
Three separate things in this essay are the same fact about a weighted quantile, and it is worth collecting them because the transferable content is the fact rather than the three readings.
Raising the weight on the scores the test point resembles pushes the quantile out; lowering it pulls the quantile in. Coverage is therefore monotone in the weight, and the promise is a one-sided constraint on it rather than a target. Everything else follows: the flat plateau above the exact ratio, the steep fall below it, the threshold that a crude estimate can clear, the rule about which way to clip, and the fact that a batch of five unlabelled points is enough. A method whose error is one-sided needs its estimate to be on the right side of a line, and being on the right side of a line is a much weaker demand than being close to a number.
Still open: the detector deferred earlier
Two deferrals were named at the start of this field and this settles one of them. The other is the one that could overturn a conclusion rather than extend it.
The finding that a drifting scale is expensive and invisible while serial correlation is cheap and visible rests on the power curve of one rank test — a comparison of the first half of the calibration scores against the second. A test that throws away the ordering within each half is the obvious thing to beat, and if a better detector exists the reversal might be a property of the instrument rather than of the departures. Whether it survives the best detector anyone would reach for is the next thing this field has to measure, and it is the measurement that decides whether the previous essay’s conclusion stands.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What the split costs — both name calibration set, closed form, conformal prediction, empirical quantile, exchangeability, interval width, overcoverage
- An interval that covers and says nothing — both name binomial proportion, closed form, interval width, overcoverage
- The arm whose variance is its answer — both name binomial proportion, closed form, monte carlo, plug in estimate
- The check worth more than the check — both name binomial proportion, closed form, interval width, monte carlo
- A coverage table with its own error — both name binomial proportion, closed form, monte carlo
- A lag the sample has less of — both name closed form, monte carlo, plug in estimate
Named objects
A flat tag is an object no other essay names yet.
Binomial proportionCalibration setClosed formConformal predictionCovariate shiftEmpirical quantileExchangeabilityInterval widthLikelihood ratioMarginal coverageMonte CarloOvercoveragePlug in estimateWeighted conformal