Coverage without a distribution

Marginal is not conditional

One exactly valid interval covers 100.00% of a quiet group and 90.66% of a noisy one, and the floor is arithmetic rather than a measurement — a group of share π is guaranteed only 1 − α/π, which is zero when the group is as rare as the miss rate.

Worth reading first: What the 95% refers to · Coverage from exchangeability alone.

Take a population with two equally common groups whose noise scales are 1 and 3, and give it one conformal interval built the ordinary way. Its marginal coverage is exactly what the rank argument promises — 95.0249% at 200 calibration points, counted at 95.28% over six thousand draws. Inside the quiet group it covers 100.00%. Inside the noisy group it covers 90.66%.

Neither of those is a failure of the procedure. The procedure delivered precisely what it guaranteed, and the guarantee is an average over everything that varies, which includes which member of the population is being predicted. The average is kept and no part of it is right, and the gap between the parts is 9.34 points.

The closed form says the same thing without any draws at all. The half-width making the mixture cover 95% is 4.9346, and two normal cdfs then give 99.9999% in the quiet group and 90.0001% in the noisy one, averaging to 95.0000% by construction. Those four decimal places are not decoration: the second number is not 90% by coincidence but because it has already reached a floor, and the floor is arithmetic.

One promise, kept on average and inside neither group. What each calibration scheme covers inside each of two equally common groups whose noise scales are 1 and 3, over 6000 draws with 200 calibration points. One interval for everybody covers 100.00% of the quiet group and 90.66% of the noisy one, averaging to 95.28% — and the closed form for that population says 99.9999% and 90.0001% at a half-width of 4.9346, from two normal cdfs and no simulation. Dividing by an estimated per-group scale gives 95.29% and 95.48%; calibrating separately inside each group gives 95.49% and 96.01% against a closed-form expectation of 95.4645%. Only the last of those is a guarantee rather than a repair, because the rank argument runs inside each group.
Fig. 1 What three calibration schemes cover inside each of two equally common groups whose noise scales are 1 and 3, over 6,000 draws. One interval for everybody covers 100.00% and 90.66%; the two repairs bring both groups to within a point of the promise.

The floor is arithmetic, which is what makes it damning

Suppose a group of share π is covered at rate c and everything else is covered as often as it possibly can be, which is always. Then the marginal coverage is at most πc + (1 − π), and a procedure promising 1 − α needs

πc+(1π)  1αc  1απ.\pi c + (1 - \pi) \ \ge\ 1 - \alpha \qquad \Longrightarrow \qquad c \ \ge\ 1 - \frac{\alpha}{\pi}.

That is the whole derivation. It assumes nothing about the model, nothing about the score, nothing about the noise and nothing about the sample size. Any procedure whose marginal coverage is 95%, however it was built, permits a group of share π to be covered as little as 1 − α/π — and permits it exactly, in the sense that nothing further can be deduced.

At equal shares that floor is 1 − 2α = 90.0%, which is where the noisy half of the population above sits once the scale ratio passes three. At a group share of a tenth it is 50.0%. At a group share of α itself — one in twenty, when the promise is nineteen in twenty — it is 0.0%. A group as rare as the miss rate may be missed every single time inside a procedure whose marginal coverage is exactly what it claimed, and no audit of the marginal rate could detect it.

How far a group can be dropped while the promise is kept. What the smaller group is covered at, in closed form, against how much noisier it is than the rest of the population, for group shares of 0.5, 0.25, 0.1, 0.05. Every curve is a procedure whose marginal coverage is exactly 95.0%. The dashed lines are 1 − α/π, which is arithmetic: if a group of share π is covered at c and everything else is covered at most always, then πc + (1 − π) ≥ 1 − α. At equal shares that floor is 90.0% and the curve reaches it by a scale ratio of three. At a share of 0.5 the floor is 90.0% and the group is covered 90.00% of the time at a ratio of 50. Nothing was estimated, nothing was simulated, and no assumption about the model entered it.
Fig. 2 The closed-form coverage of the smaller group against how much noisier it is, for four group shares, with 1 − α/π drawn for each. At equal shares the curve reaches its floor of 90% by a scale ratio of three and stays there.

The equal-shares curve reaching its floor at a ratio of three is what makes the four decimal places in the opening worth quoting. 90.0001% is not 90% rounded; it is 90% approached from above and effectively arrived at, and the reason the curve then flattens is that the quiet group has run out of room. Once the quiet group is covered 100.000% of the time, every further widening of the interval buys nothing there, so the whole of the marginal requirement has to be met by the noisy group and its coverage cannot fall further. The floor is reached when one group is saturated, which is why a larger scale ratio does not make things worse at equal shares.

How far the arithmetic goes when the group is small

The floor falls as the group gets rarer, and the population reaches it faster too, because a small group contributes little to the marginal average and so exerts little pull on the width. Both effects push the same way.

A tenth of the population, three times noisier than the rest, is covered 60.162% of the time. At a ratio of ten it is covered 50.000% — its floor exactly. A twentieth of the population is covered 53.556% at a ratio of three, 20.167% at a ratio of ten, and 4.815% at a ratio of fifty.

The intermediate cases are the ones that make it a gradient rather than a pathology. A quarter of the population, only twice as noisy as the rest, is covered 82.142% of the time against a floor of 80% — so a mild difference in a moderately sized subgroup has already spent nearly the whole of what the guarantee permits. It does not take an extreme population to reach the floor; it takes a scale ratio of two.

How far a group can be dropped while the promise is kept. What the smaller group is covered at, in closed form, against how much noisier it is than the rest of the population, for group shares of 0.5, 0.25, 0.1, 0.05. Every curve is a procedure whose marginal coverage is exactly 95.0%. The dashed lines are 1 − α/π, which is arithmetic: if a group of share π is covered at c and everything else is covered at most always, then πc + (1 − π) ≥ 1 − α. At equal shares that floor is 90.0% and the curve reaches it by a scale ratio of three. At a share of 0.1 the floor is 50.0% and the group is covered 50.00% of the time at a ratio of 50. Nothing was estimated, nothing was simulated, and no assumption about the model entered it.
Fig. 3 The same table with a tenth of the population in the smaller group. Its floor is 50%, and a scale ratio of ten is enough to reach it.

That last number is worth stating in the form a reader would meet it. A procedure is deployed, audited, and found to cover 95% of the time — exactly, verifiably, with a distribution-free guarantee behind it and no modelling assumption to argue about. One twentieth of the people it is applied to are covered less than one time in twenty. The audit that found the 95% is the correct audit of the quantity that was promised, and it is incapable of noticing.

How far a group can be dropped while the promise is kept. What the smaller group is covered at, in closed form, against how much noisier it is than the rest of the population, for group shares of 0.5, 0.25, 0.1, 0.05. Every curve is a procedure whose marginal coverage is exactly 95.0%. The dashed lines are 1 − α/π, which is arithmetic: if a group of share π is covered at c and everything else is covered at most always, then πc + (1 − π) ≥ 1 − α. At equal shares that floor is 90.0% and the curve reaches it by a scale ratio of three. At a share of 0.05 the floor is 0.0% and the group is covered 4.82% of the time at a ratio of 50. Nothing was estimated, nothing was simulated, and no assumption about the model entered it.
Fig. 4 A twentieth of the population, whose floor is zero. At a scale ratio of fifty it is covered 4.815% of the time, inside a procedure whose marginal coverage is exactly 95%.

Nothing in that paragraph is a criticism of conformal prediction in particular. It is a statement about marginal coverage, and every interval on this site that has ever reported one is subject to it — the eight rules and windows whose coverage was measured as marginal rates, the intervals for a proportion in the first field, the forecast bands. What is unusual here is only that the guarantee is exact, so the gap cannot be blamed on an approximation and has to be read as what the promise actually says.

The structure is also the one Simpson’s reversal is made of, in a different currency. There a marginal association is the wrong sign in every part of the population; here a marginal rate is the wrong number in every part of it. In both cases the aggregate is not a summary of the parts, it is a weighted average that the parts can be arbitrarily far from, and the weights are the group sizes.

Two repairs, and only one of them is a guarantee

The gap is closable, and the two ways of closing it are not the same kind of thing.

Normalise the score. Divide each residual by an estimated scale for its own group, so a point in the noisy group is not called nonconforming for being noisy. The interval then widens and narrows with the group, and the two groups come to 95.29% and 95.48%. This is an approximation. The scale is estimated from the training half, the estimate has error in it, and nothing guarantees the two group rates will land where they landed — what is guaranteed is the marginal rate, exactly as before, because the rank argument does not read the score.

Calibrate inside each group. Split the calibration set by group and take a separate order statistic in each — Mondrian calibration. The rank argument then runs within the group, so the guarantee is conditional on the group and is exact in finite samples, not approximately restored. Counted, the two groups come to 95.49% and 96.01% against a closed-form expectation of 95.4645%, which is a finite sum over the binomial distribution of how many calibration points land in the group.

The distinction between the two is the same one that separates the three estimators for eight hospitals: a scheme that pools everything, a scheme that treats each group as its own problem, and a scheme in between. Mondrian calibration is the no-pooling answer and it is exact for the same reason that answer is always exact — it never uses a point from one group to say anything about another — and it pays for that in the size of the sets it has left. The normalised score is the partial-pooling answer: one line and one shape of interval for the whole population, with a per-group scale doing the adapting, and its error is concentrated in that estimated scale.

Three intervals, one shortfall. What each of three intervals actually covers, at four rules and two block windows, over 300 samples of 120 rows. All three are built from the same resamples on the same draws, so a difference between them is a difference in what is done with the resampled series. Not one of the twenty-four cells reaches the ninety-five per cent it promises. The studentised interval runs from 75.7% to 92.3%, the percentile interval — the earlier field's — from 80.0% to 89.7%, and a normal interval on the same scale from 81.7% to 89.0%. The standard repair for a percentile interval's shortfall does not repair it.
Fig. 5 The same idea in a resampling setting: an interval that carries the scale of the thing it is built around rather than one scale for everything. Dividing by an estimated spread is the device, and the estimate’s own error is what it costs.

The closed-form expectation is above the pooled promise — 95.4645% against 95.0249% — and that difference is the price of the guarantee, paid in overcoverage. Each group’s calibration set is half the size of the pooled one, so each sits further up the sawtooth, and the expectation is a sum over the sizes the binomial can deliver rather than a single tooth. Mondrian calibration buys a conditional guarantee and pays for it by being conservative twice over.

The repair is cheaper than the thing it repairs

The expected trade here is fairness against width: cover the noisy group properly and the noisy group’s interval gets wider, and somebody pays. The measurement says the opposite. The single interval’s mean width is 10.047. The normalised score’s is 8.112 — 19% narrower — and Mondrian calibration’s is 8.313.

The repair is not paid for in width. The width each calibration scheme gives each group, over 6000 draws with 200 calibration points. One interval for everybody is 10.031 and 10.062 — the same interval twice, since it cannot read the group. Dividing the residual by an estimated per-group scale gives 4.120 and 12.025, and calibrating separately inside each group gives 4.217 and 12.328. Averaged over the population the repairs are 8.112 and 8.313 against 10.047, so the mean width falls: the single interval was not a compromise between the two groups, it was the noisy group's interval handed to everybody.
Fig. 6 The width each scheme gives each group. One interval for everybody is 10.031 and 10.062 — the same interval twice, since it cannot read the group. The normalised score gives 4.120 and 12.025, and averaging those is narrower than averaging the two identical ones.

The reason is visible once the two group widths are separated. The single interval gives 10.031 to the quiet group and 10.062 to the noisy one, which is one interval reported twice, since nothing in it reads the group. The normalised score gives 4.120 and 12.025. It is wider where the noise is and much narrower where it is not, and the population average falls because half the population was being handed an interval two and a half times wider than its own noise warranted.

So the single interval was never a compromise between the two groups. It was the noisy group’s interval handed to everybody — near enough, at a scale ratio of three, that the quiet group’s coverage saturates at 100.000%. Conditional validity here is not bought with width; it is what stops width being wasted, and the usual framing of the trade has the sign backwards.

That is a claim about this population and it should not be generalised past it. With groups of nearly equal noise the normalised score has nothing to exploit and its estimated scales add noise for no gain; with a rare and extremely noisy group the pooled interval is set by the common group and normalising genuinely does widen the rare group’s interval. The finding is that the trade is not automatic, not that it always inverts. What the normalised score is doing is what residuals given back their own variance does in a resampling setting: putting each point on a common scale before treating the points as interchangeable. The difference is that there the rescaling is a correction to a known distortion, and here it is the entire mechanism by which an interval learns anything about the point it is being built for.

The group has to be named, and usually is not

Both repairs need the group. Mondrian calibration needs it to partition the calibration set; the normalised score needs it to pick the scale to divide by. In the population here it is an observed binary label, present on every point, which is the easiest case there is and is not the case a reader usually has.

Three things go wrong outside it, in increasing order of how badly.

The group is continuous rather than categorical. Noise that grows with the covariate has no groups in it at all, and Mondrian calibration has nothing to partition on. Binning the covariate recovers a partition, and the guarantee it then delivers is conditional on the bin rather than on the covariate — which is a real guarantee about a coarser object, and every bin boundary is a place where the coverage jumps. The normalised score has no such trouble because a fitted scale function is a continuous thing, and pays for it by no longer being exact anywhere.

The group is known but small. A partition into k groups divides the calibration set k ways, and each part has to clear nineteen points before it has a finite interval at five per cent. Two hundred calibration points support ten groups and not a hundred, and a group with eighteen calibration points gets the real line — the same feasibility edge that made three of nine splits infinite, arriving now once per group rather than once per sweep. A partition fine enough to be interesting is usually a partition too fine to calibrate.

And the group is not recorded. This is the ordinary case and there is nothing to be done about it inside the procedure. A subgroup nobody measured cannot be conditioned on, cannot be audited, and is covered at whatever rate the floor permits. The only defence is that the floor is computable in advance from the group’s share alone: any group making up a twentieth of the population has a guaranteed coverage of zero, whether or not anybody knows the group exists.

What could have made these numbers wrong

A closed form can be wrong in a way a simulation cannot check, and one was. The mixture quantile — the half-width making the two groups’ coverages average to 95% — is found by bisection on a cdf, and the bracket was a fixed upper end of forty. That is right for every scale ratio up to about twenty and returns its own endpoint at fifty, where the true answer is 82.243. The number it produced for the equal-shares population at a scale ratio of fifty was 57.6%, against a floor of 90% for a group that size — which is impossible, and was caught by an assertion comparing the closed form against the arithmetic floor rather than against another measurement. The bracket is grown now. This is the same defect a fixed quantile bracket produced elsewhere on this site, made a second time in a second file, and the lesson it carries is narrow and useful: the assertion that caught it compared a number against a closed form, not against itself. Two routes to a number are only two routes if one of them can refuse the other.

The counted and closed-form group rates could have agreed for the wrong reason. They agree at 100.00% against 99.9999% and 90.66% against 90.0001%, and the second gap is 0.66 points against a binomial standard error of about 0.53 points at three thousand draws in that group — a little over one standard error, which is agreement. But the closed form is computed for a known half-width, and the counted version has to estimate the score quantile from 200 calibration points, so the counted noisy-group rate is expected to sit slightly above the closed form: the calibration size promises 95.0249% rather than 95%, and the sawtooth’s extra 2.5 hundredths of a point of marginal coverage lands almost entirely in the noisy group, because the quiet group has no room for it.

The group could have differed in its mean rather than in its noise. It does not — the line is the same in both groups by construction, and the group enters only the noise scale. That matters because a group whose mean is wrong is a misfit model, and the repair for a misfit model is to fit it better, which would confuse the finding with an ordinary modelling error. Here the model is right for both groups and the interval is still wrong for both, which is the point: this is not a defect that fitting harder cures.

And the two repairs could have been compared at different coverages. They are not; both land within a point of the promise in both groups, and the width comparison is therefore at approximately fixed coverage. Mondrian’s slightly larger width — 8.313 against 8.112 — is consistent with its slightly higher coverage, which is the overcoverage its smaller calibration sets force.

What cannot be done at all, and it is a theorem rather than a measurement

The two repairs answer two different questions and neither answers the question a reader would most like answered.

Coverage conditional on a group of positive probability is attainable exactly. Mondrian calibration is how, and its guarantee is the same rank argument run inside a smaller set. Coverage conditional on the covariate itself — holding for every value of a continuously distributed x, distribution-free, in finite samples — forces an interval of infinite expected length almost everywhere. A group is a set the sample contains many members of; a point on a continuum is a set the sample contains none of, and there is nothing to calibrate against.

That is stated here rather than demonstrated, because it is a theorem about every procedure rather than about this one, and this ladder does not prove theorems. What it means in practice is that the normalised score is an approximation to something unattainable and Mondrian calibration is an exact answer to a weaker question, and that a reader choosing between them is choosing between those two positions rather than between two implementations.

Neither is a way round the floor. The floor is what the marginal guarantee permits, and both repairs work by no longer being the marginal guarantee — by conditioning on something, and therefore by requiring that the something be named in advance. A group nobody thought to name is covered at whatever rate the arithmetic at the top of this essay allows, and the score that decides how wide the interval is for whom is the only remaining lever. It turns out to move everything except the number the guarantee is about — and the last thing left standing, the exchangeability the whole ladder rests on, turns out to be breakable by ordinary things, one of which reproduces this essay’s ninety per cent floor as a marginal loss.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Calibration setClosed formConditional coverageConformal predictionCoverageDistribution-freeExchangeabilityHeteroskedasticityInterval widthMarginal coverageMondrian conformalNonconformity scoreOrder statisticOvercoverage