Concept

Overcoverage — where it appears

Covering more often than promised, which is not free. An interval covering 97% when it claimed 95% is wider than it needed to be, and for a conformal interval the excess is a known function of the calibration set's size rather than an accident of the data in hand.

Named by 6 essays across 3 fields — each of them below, with the objects they name alongside it.

The coverage is exact and it is not the nominal rate. ⌈(m+1)(1−α)⌉/(m+1) against m, the number of calibration points, at α = 0.05. It is a closed form and needs no data. It never falls below 95.0% and never reaches 1−α+1/(m+1), the two bounds the rank argument gives. It equals 95.0% exactly at 10 of the 182 sizes drawn — the sizes where (m+1)α is a whole number, which are 20 apart — and sits above it everywhere else, worst at 38 points where it is 97.4359%, or 2.4359% of coverage nobody asked for. Below 19 points there is no such order statistic and the interval is the whole line, which is where the curve starts.

Coverage from exchangeability alone

A conformal interval's coverage is a fact about the ranks of m+1 numbers, so it can be enumerated before any data arrive — all 40,320 orderings of eight values, agreeing with the closed form to machine precision. What that exactness delivers is not 95%.

conformal · Exchangeability
The split decides the width. The width of the interval against the share of 200 observations spent on fitting rather than on calibrating, over 3000 draws. Spending more on the fit shrinks the residuals; spending more on calibration builds the interval at a less extreme order statistic. The two meet at 0.5, where the width is 4.0416 against 4.1603 at 0.1 and 4.3820 at 0.9. Full conformal, which spends the same 200 points on both jobs, is 3.9865 — so the whole cost of splitting is 1.38%.

What the split costs

Splitting a sample between fitting and calibrating looks like a trade against the guarantee, and it is not: coverage moves 0.63 points across nine splits and every reading sits on its own promise. The whole cost is 1.38% of width — and at sixty observations the width falls, rises and falls again.

conformal · Exchangeability
One promise, kept on average and inside neither group. What each calibration scheme covers inside each of two equally common groups whose noise scales are 1 and 3, over 6000 draws with 200 calibration points. One interval for everybody covers 100.00% of the quiet group and 90.66% of the noisy one, averaging to 95.28% — and the closed form for that population says 99.9999% and 90.0001% at a half-width of 4.9346, from two normal cdfs and no simulation. Dividing by an estimated per-group scale gives 95.29% and 95.48%; calibrating separately inside each group gives 95.49% and 96.01% against a closed-form expectation of 95.4645%. Only the last of those is a guarantee rather than a repair, because the rank argument runs inside each group.

Marginal is not conditional

One exactly valid interval covers 100.00% of a quiet group and 90.66% of a noisy one, and the floor is arithmetic rather than a measurement — a group of share π is guaranteed only 1 − α/π, which is zero when the group is as rare as the miss rate.

conformal · Exchangeability
The interval that over-covers when the instrument fails. Counted coverage of two nominal 95.0% intervals for the same causal effect, read off the same 2000 draws of 200 rows at each first stage. The exact Anderson–Rubin set covers 95.3% at every setting — flat, because the statistic it inverts is built from y − tβ, which contains no π at all, and is therefore the same number on the same draw whatever the instrument is worth. The conventional interval covers 99.1% at π = 0.02 and 95.6% at π = 0.6: it goes wrong at the weak end by covering too MUCH, at a median width of 7.320, because its standard error is computed from residuals taken at an estimate that has itself gone wrong. A weak instrument does not make this interval lie about its coverage; it makes it useless while telling the truth.

What the first stage does not know

A single weak instrument does not make the conventional interval undercover — it makes it cover 99.1% at a width of 7.320. Where the promise actually breaks is many instruments — coverage falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583.

instrument · Exclusion
Too small breaks it and too large does not. Coverage of the weighted interval against the factor the true likelihood ratio is multiplied by, at a test population 80.0% drawn from the noisier group and 200 calibration points. The exact weight is the factor of 1 and covers 95.70%. Overstating it costs nothing: 96.13% at sixteen times too large. Understating it costs, and costs steeply below about a half — 94.93% at half, 88.37% at an eighth and 67.90% at a thirtieth. The question this answers was whether a wrong weight degrades smoothly or falls off a cliff, and the answer is that it does neither symmetrically: the curve is smooth and one-sided.

The weight that has to be estimated

A likelihood ratio sixteen times too large costs 5.5% of interval width and no coverage at all; one a thirtieth of the right size covers 67.90%. The estimate from a batch of five unlabelled covariates covers 95.10% against an exact repair's 95.30%, and the binomial says why.

conformal · Exchangeability
Three promises, and no procedure keeps all three. Average coverage and worst-case coverage for four 95% intervals for a proportion at n = 40, computed exactly. Their expected widths are 0.2418, 0.2417, 0.2472, 0.2641 in the same order. The textbook interval and the score interval have the same expected width to four digits — 0.2418 and 0.2417 — and worst-case coverages of 55.31% and 92.21%. The exact interval never breaks its promise and is 9.3% wider than the score interval to do it. Each of the three columns orders the four procedures differently.

An interval that covers and says nothing

A procedure returning the whole line 95% of the time and the empty set otherwise has coverage exactly 95% at every parameter value. Two real intervals at forty observations have expected widths of 0.2418 and 0.2417 and worst-case coverages of 55.31% and 92.21%.

intervals · Coverage

Named alongside it

The objects these essays reach for when they reach for this one.

Closed formInterval widthCalibration setConformal predictionCoverageExchangeabilityMarginal coverageEmpirical quantileOrder statisticBinomial proportionConditional coverageConservative interval

All concepts