The score is the modelling
Worth reading first: What the 95% refers to · Coverage from exchangeability alone.
A nonconformity score is the one place a modeller’s judgement enters a conformal interval, and it is worth knowing what that judgement buys. Six of them, measured on the same draws against a promise of 95.0249%, cover 94.75%, 94.90%, 94.60%, 94.75%, 94.65% and 94.30%. The spread of that column is 0.60 points against a standard error of a difference of 0.69. It is one number, six times over.
Two of those six are not serious proposals. One is aimed five units away from the fitted line, so that its interval is centred where nothing is. One never reads the response at all. Both cover at the same rate as the two that were built with care, because the rank argument does not read the score: it needs the calibration scores and the test score exchangeable, and it needs nothing else.
So the coverage column is not the finding. It is the control, and everything a reader actually wants from an interval has to be somewhere else.
What the guarantee is indifferent to
The construction takes m calibration scores, sorts them, and builds the interval at the ⌈(m+1)(1−α)⌉-th. The only property of the score that entered was that the m+1 numbers — the calibration scores and the future one — could be dealt out in any order with equal probability. Anything computed the same way for every point has that property, whatever it computes.
A score may therefore be a squared residual, an absolute residual, a residual divided by something, a distance from a quantile band, a distance from an arbitrary point in the plane, or the covariate with the response ignored entirely. It may be computed from a model that fits beautifully or from a model fitted to a different dataset. The interval it produces will cover 95.02% of the time, exactly, in every one of those cases.
That indifference is the guarantee’s whole strength and it is the reason the guarantee cannot be the criterion. A statement that is equally true of a careful model and of a random one has no capacity to distinguish them, which means coverage is not evidence that a model is any good — a point worth making sharply, because a conformal interval is often presented as though the exactness of its coverage were a property of the modelling underneath it. It is a property of the sorting.
The comparison to keep beside this is the interval that forgot it had estimated its parameters, which covers 87.3% because the formula it uses was derived for known parameters and computed with estimated ones. That is a guarantee that does depend on the modelling, and it fails when the modelling is wrong. Conformal’s does not depend on it and so cannot fail that way — and cannot report it either.
The plus one is the test point’s own place in the count
One detail of the construction is easy to read as a rounding convention and is doing real work, and it is the same detail that decides the coverage column’s exact value. The interval is built at the ⌈(m+1)(1−α)⌉-th smallest of m calibration scores. The index is computed from m+1 and the set has m things in it.
The m+1 is the test point counting itself. The rank argument is about the ordering of m+1 exchangeable numbers, of which m have been observed and one has not, and the interval covers when the unobserved one’s rank is at or below k. Computing the index from m instead would be computing the rank distribution of a set the test point is not in, and the guarantee would fail on exactly the draws where the test score is the largest of the m+1 — which is the 5% of draws the promise is about.
That is the same arithmetic as a randomisation p-value’s (1 + count)/(1 + B), and the essay that priced dropping it is worth reading beside this one, because it establishes the thing that makes the correction hard to notice: at the values anybody actually uses it changes almost nothing. At 200 calibration points the index is 191 with the plus one and 190 without, and the promise falls by about half a point. At twenty calibration points it is the difference between the twentieth score and the nineteenth, and the promise falls from 95.2381% to a little over ninety.
The general shape recurs wherever a reference set is built from things that could have happened: the observed case is one of them. Sampling a reference distribution has the same correction for the same reason, and in all three settings the correction is invisible on large sets and load-bearing on small ones — which is the worst possible combination, since a routine tested at m = 200 and deployed at m = 20 will look correct throughout its development and be wrong in the field.
Everything a reader wants is in the width
On the same draws, at the same coverage, the mean widths are 8.227 for a residual divided by its own group’s estimated spread, 8.449 for a conformalised quantile band, 10.058 for a plain absolute residual, 17.873 for the score aimed five units off the fit, and 19.424 for the same normalisation with the two groups’ scales exchanged. The widest is 2.361 times the narrowest.
The misaimed score is the cleanest illustration, because nothing about it is wrong. It scores each point by |y − ŷ − 5|, which is a perfectly legitimate function of the point, computed the same way for every point, and therefore exchangeable exactly as required. Its interval is centred five units above the fitted line, and to cover 95% of points from that position it has to be wide enough to reach back across the gap. It is 17.873 wide against the plain residual’s 10.058 — the same coverage for one and three quarter times the interval — and no diagnostic that reads only the guarantee would report anything at all.
Width is the currency here in the same way it is in every other interval comparison on this site, and the difference is that in those comparisons the shortest interval was the one that missed. Here nothing misses more than anything else, so width is not traded against coverage. It is traded against nothing. A score three quarters wider again is three quarters worse, and there is no compensating column.
There is a second thing the width column is doing here, which is measuring how much a model was worth. The plain absolute residual comes from a least-squares line that is correct for this population; the normalised score comes from the same line plus one estimated number per group. That one number is worth 18% of the interval. The quantile band comes from two fitted quantile regressions on a design that includes the group, which is a great deal more machinery, and it is very slightly wider than the score with one number in it — so the elaborate model is not paying for itself here, and the width column is where that shows. Under a differently shaped noise distribution the ordering would go the other way, since a quantile band can be asymmetric and a scaled residual cannot.
From adaptive through flat to backwards
Mean width is still an average over the population, and an average over a population with parts is the statistic this ladder has already been caught by once. The sharper reading is how much wider each score makes the interval for the noisy group than for the quiet one, when the population’s true scale ratio is 3.
The normalised score gives 2.878, close to the truth. A conformalised quantile band gives 2.500. The plain absolute residual gives 0.997 — one interval for both groups, since nothing in |y − ŷ| reads the group. And the exchanged normalisation gives 0.344, which is the ratio upside down: it makes the interval widest exactly where the noise is smallest. The column runs over a factor of 8.377.
The consequence in group coverage is the argument of the previous rung arriving as a consequence of a modelling choice rather than of the guarantee. The plain absolute residual covers 100.00% of the quiet group and 89.73% of the noisy one. The normalised score covers 94.07% and 95.69%. The quantile band covers 95.60% and 93.74%. All three are marginally valid and only the last two are conditionally anything.
So the choice of score is not a refinement made after the important decision. It is the important decision, and the guarantee is what makes it safe to make badly: a score chosen wrongly produces an interval that is too wide, or the wrong shape, or wide in the wrong places, and never one that lies about how often it covers.
A score can be valid and say nothing
The limiting case makes the point without any measurement at all being needed to interpret it. Score each point by |x| — the covariate, with the response ignored. That is exchangeable with the calibration scores, so the rank argument applies unchanged, and the resulting procedure covers 94.30% of the time.
What it returns is not an interval. If |x| for the test point falls at or below the calibration quantile, every value of y is conforming and the set is the whole real line; if it does not, no value is conforming and the set is empty. The share of draws on which the set is the whole line is 94.30% — the same number as the coverage, because those are exactly the draws on which it covers. The other 5.70% are empty sets, which cover nothing.
A coverage audit of that procedure returns 94.30% and passes. The procedure has no information in it. This is the strongest available statement of what a distribution-free guarantee does not include, and it is worth having in the form of a counted number rather than as a hypothetical, because the hypothetical version invites the reply that no one would do that — while the measured version says that if someone did, nothing in the reporting would show it.
The defect’s symptom is absence: there is no error, no warning, no failed check, and a column of coverages that looks like every other column of coverages. The only instrument that detects it is the one that reports width, where an infinite entry cannot be averaged and has to be counted separately.
The score also decides which end the misses are at
There is a third column, and it is the one the first rung of this ladder found the textbook interval failing on: a two-sided 95% interval promises 2.5% of misses above and 2.5% below, and a coverage column cannot see when it breaks that promise at both ends in opposite directions.
An absolute residual is symmetric in the sign of the residual, so the interval it produces is symmetric about the fit, and on a skewed noise distribution it will miss upward more often than downward — exactly as the textbook interval does, for exactly the same reason, at exactly the same total coverage.
The repair is again a choice of score rather than a repair to the procedure. Scoring by the signed distance outside a fitted quantile band — max(q̂ₗₒ(x) − y, y − q̂ₕᵢ(x)) — gives an interval with its own two ends, each inherited from a fitted quantile, and the calibration then corrects the pair jointly rather than symmetrically. That is what the conformalised quantile band above is, and it is why its two group coverages straddle the promise rather than sitting on one side of it. Which end a rule misses at is a question that has to be asked separately in every field on this site, and here the answer is decided entirely by the score.
What could have made these numbers wrong
Two scores produced identical readings, which looks exactly like a bug. The exchanged normalisation and the plain absolute residual return the same coverage to two decimal places, the same two group coverages, and the same width for the noisy group — 10.043 in both. The first reading of that is that one of the two is not being computed. It is: the quiet group’s widths differ by a factor of nearly three, 29.227 against 10.074. The reason for the agreement is that the pooled calibration quantile is set by the noisy component under both normalisations, so the same order statistic is taken and the noisy group’s interval is identical; only the quiet group’s, which is scaled by the other group’s spread, changes. An agreement that survives a check of what differs is a finding rather than a defect.
The coverage column’s flatness could be a claim the trial count cannot support. It is a null result, and a null result needs its resolution stated. At 2,000 draws the standard error of a difference between two of these rates is 0.69 points, and the observed spread across six of them is 0.60 — so the measurement is capable of resolving a difference of about one and a half points and found nothing above half of that. It is not capable of ruling out a difference of a tenth of a point, and nothing here claims to.
The quantile regression could have been fitted wrongly, quietly. The band the conformalised score is built from comes from a majorisation that iterates to a fixed point, which is the kind of routine that converges to something plausible and wrong. It is checked against a second route sharing none of its arithmetic: a linear quantile regression on two columns has an optimum passing through two of the sample points, so the answer lies in a finite set of C(n,2) lines, and that set can be walked exhaustively. Over eighteen fits the worst relative gap in pinball loss is 0.0049 — half a per cent, at an extreme quantile. That gap costs a little width and no coverage at all, which is this essay’s own thesis applied to its own machinery.
The scores could have been compared on different draws. They are not: one training set, one calibration set and one test point are drawn, and all six scores are computed on them before the next draw. So the coverage column’s flatness is a paired comparison, which makes it a stronger null than an unpaired one would be — six independent sweeps would produce a spread of about a point from sampling alone, and this column has six readings of the same draws differing by 0.60.
And the widths are means over draws whose distribution is skewed. A draw with a large maximum calibration score widens every score’s interval at once, so the comparisons above are paired — one sample, all six scores, then the next sample — and the ratios are ratios of means on shared draws. Reported as means of ratios they would weight the narrow draws more heavily and the factor of 2.361 would shrink.
What none of the six scores could repair
Every score above is a function of a point and a fitted model, and there is a class of defect none of them addresses, because the defect is in the model rather than in how the model is scored.
The population here has the same mean function in both groups. The line is right, and the groups differ only in how far points scatter around it. That was chosen deliberately: it makes the difference between the scores attributable to the scoring alone, since no score is compensating for a misfit that another score would have to compensate for differently.
Under a mean function that is wrong — a line fitted to data that curves — every score’s calibration set fills up with residuals that carry the curvature, and every interval widens to swallow it. The coverage stays exactly where it is, because the rank argument does not read the model; the width absorbs the whole of the misspecification, and it absorbs it uniformly, since the calibration quantile is one number applied everywhere. So a curved truth produces an interval that is far too wide in the region where the line is nearly right and, in the marginal average, exactly correct. That is the same failure this essay has already described twice — a conditional defect concealed by a marginal guarantee — and it arrives here from the model rather than from the population.
The score that would repair it is one that reads where in the covariate space the point sits, which is what the quantile band does and what the group-normalised score does for the one covariate it is told about. The general version is not measured here.
What this leaves for the last rung
The guarantee has now survived a wrong model, a misaimed score, a score that reads nothing, a sample split anywhere in a wide range, and a population with parts it knows nothing about. Every one of those changed the interval and none of them changed how often it covered.
One assumption has done all of the work and has not yet been touched: that the calibration scores and the test score are exchangeable. Everything above rests on it, and nothing above is true without it. The natural expectation is that it fails gracefully — that a small departure costs a small amount of coverage — and the more useful question is which departures a practitioner would notice, since a departure that is caught is a departure that costs nothing. Those two orderings turn out to be opposites.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A detector built for the ordering — both name calibration set, conformal prediction, distribution-free, exchangeability, heteroskedasticity, nonconformity score
- Right for the wrong reason — both name coverage, heteroskedasticity, interval width
- Robust is not free — both name coverage, heteroskedasticity, interval width
- The bias that lands in the slope — both name coverage, interval width, model misspecification
- When the prior is confident and wrong — both name coverage, interval width, model misspecification
- A covariance with no parameter in it — both name least squares, model misspecification
Named objects
A flat tag is an object no other essay names yet.
Calibration setConditional coverageConformal predictionConformalised quantile regressionCoverageDistribution-freeExchangeabilityHeteroskedasticityInterval widthLeast squaresModel misspecificationNonconformity scorePinball lossQuantile regression