Concept

Distribution-free — where it appears

Holding for every distribution in a very large class rather than for a named one. Such a guarantee is bought by making the claim about ranks or about the design instead of about the data, and what it costs is a statement weaker than the one a correct model would have given.

Named by 6 essays across 2 fields — each of them below, with the objects they name alongside it.

The coverage is exact and it is not the nominal rate. ⌈(m+1)(1−α)⌉/(m+1) against m, the number of calibration points, at α = 0.05. It is a closed form and needs no data. It never falls below 95.0% and never reaches 1−α+1/(m+1), the two bounds the rank argument gives. It equals 95.0% exactly at 10 of the 182 sizes drawn — the sizes where (m+1)α is a whole number, which are 20 apart — and sits above it everywhere else, worst at 38 points where it is 97.4359%, or 2.4359% of coverage nobody asked for. Below 19 points there is no such order statistic and the interval is the whole line, which is where the curve starts.

Coverage from exchangeability alone

A conformal interval's coverage is a fact about the ranks of m+1 numbers, so it can be enumerated before any data arrive — all 40,320 orderings of eight values, agreeing with the closed form to machine precision. What that exactness delivers is not 95%.

conformal · Exchangeability
One promise, kept on average and inside neither group. What each calibration scheme covers inside each of two equally common groups whose noise scales are 1 and 3, over 6000 draws with 200 calibration points. One interval for everybody covers 100.00% of the quiet group and 90.66% of the noisy one, averaging to 95.28% — and the closed form for that population says 99.9999% and 90.0001% at a half-width of 4.9346, from two normal cdfs and no simulation. Dividing by an estimated per-group scale gives 95.29% and 95.48%; calibrating separately inside each group gives 95.49% and 96.01% against a closed-form expectation of 95.4645%. Only the last of those is a guarantee rather than a repair, because the rank argument runs inside each group.

Marginal is not conditional

One exactly valid interval covers 100.00% of a quiet group and 90.66% of a noisy one, and the floor is arithmetic rather than a measurement — a group of share π is guaranteed only 1 − α/π, which is zero when the group is as rare as the miss rate.

conformal · Exchangeability
How many observations the smallest and largest of them need. The interval between the extremes of n draws holds at least 95% of the population with probability 1 - n p^(n-1) + (n-1) p^n, whatever the population is. Reaching 95% confidence takes 93 observations.

Ninety-three observations, and nothing assumed

The interval between the smallest and largest of a sample holds a share of the population whose distribution does not depend on the population — Beta(n − 1, 2), for anything continuous. Buying the 95/95 that normality buys at ten observations costs 93 of them, and that number is the exchange rate between an assumption and data.

estimated · Bands
Six scores, one coverage. What each nonconformity score's interval covers, over 2000 draws with 200 calibration points, against the 95.0249% the rank argument promises. The column runs from 94.30% to 94.90%, a spread of 0.60% against a standard error of a difference of 0.69% — one number, six times. That includes a score aimed five units off the fit and a score that never reads the response at all, because the rank argument does not read the score either: it needs the scores exchangeable and nothing else. Every decision a modeller makes has to show up somewhere else, and the next two readings are where.

The score is the modelling

Six nonconformity scores on the same draws cover within 0.60 points of each other, against a standard error of a difference of 0.69 — one number six times. Their widths run over a factor of 2.361 and their adaptivity over a factor of 8.377.

conformal · Exchangeability
What a scale that grows across the sample costs. What the interval covers when the noise scale grows across the sample, against how far the departure has gone, over 1500 draws at each setting. The coverage runs from 94.47% at no departure to 83.93% at the end of the sweep, a loss of 11.07%. The rank argument needs the 200 calibration scores and the test score to be exchangeable, and this is one of the three ways that fails. A test built for it reaches 80% power at 4.054, where the coverage is 85.13% — so 9.87% of the loss is inside the region such a test would have missed.

When the order matters

Three ways of breaking exchangeability cost 4.93, 11.07 and 1.07 points of coverage, and the ordering by cost is the reverse of the ordering by how soon a test would have caught them. The departure practitioners check for is the cheapest one.

conformal · Exchangeability
Three detectors for one departure, all at 5%. How often each of three checks on the calibration scores fires, against the size of the drift, with every critical value simulated under no drift so that all three sit at 5.0% exactly. The incumbent — a rank comparison of the first half of the scores against the second — reaches four-in-five power at a growth factor of 4.31. Reading each score's rank against its position reaches it at 2.65, and the largest running departure of the scores from their mean at 2.12. The ordering of the three is the ordering by how much of the sample's arrangement each one uses.

A detector built for the ordering

The best of three checks for a drifting scale fires at half the growth factor the standard one needs — 2.12 against 4.31 — and still leaves 6.50 points of coverage gone before it does, against 0.51 for serial correlation. The reversal was not a property of the test.

conformal · Exchangeability

Named alongside it

The objects these essays reach for when they reach for this one.

Calibration setConformal predictionCoverageExchangeabilityNonconformity scoreMarginal coverageHeteroskedasticityOrder statisticAutocorrelationClosed formConditional coverageExchangeability test

All concepts