Coverage without a distribution

Coverage from exchangeability alone

A conformal interval's coverage is a fact about the ranks of m+1 numbers, so it can be enumerated before any data arrive — all 40,320 orderings of eight values, agreeing with the closed form to machine precision. What that exactness delivers is not 95%.

Worth reading first: What the 95% refers to.

The rule this site runs on is that nothing is called 95% until it has been counted, and every field before this one has counted a coverage that a model was supposed to deliver and did not. This field is the construction where the counting can be done in advance. The coverage of a conformal prediction interval is a fact about the ranks of a set of exchangeable numbers, so it can be enumerated: all 40,320 orderings of eight values, each rank taken in exactly 5,040 of them, agreeing with the closed form to machine precision.

That is not a simulation with a large trial count. It is the whole sample space of the question, walked. There is no seed, no standard error and nothing to converge — which is the same distinction the essay that summed a coverage over every possible sample drew for a proportion, arriving here from the other direction: there the finite space was the twenty-one counts a sample of twenty can produce, here it is the orderings of the calibration scores and the future one.

And the first thing the enumeration says is that the exactness is not the nominal rate.

The coverage is exact and it is not the nominal rate⌈(m+1)(1−α)⌉/(m+1) against m, the number of calibration points, at α = 0.05. It is a closed form and needs no data. It never falls below 95.0% and never reaches 1−α+1/(m+1), the two bounds the rank argument gives. It equals 95.0% exactly at 10 of the 182 sizes drawn — the sizes where (m+1)α is a whole number, which are 20 apart — and sits above it everywhere else, worst at 38 points where it is 97.4359%, or 2.4359% of coverage nobody asked for. Below 19 points there is no such order statistic and the interval is the whole line, which is where the curve starts.0.9600.980150100150200calibration pointswhat the interval covers, exactly38 points: 97.4359%1−α + 1/(m+1)1−α = 95.0%closed form, α = 0.05, from 19 calibration pointsexact at 10 sizes, 20 apart
Fig. 1 The exact coverage of a conformal interval against the number of calibration points, at a promised 95%. It equals 95% at ten of the 182 sizes drawn and sits above it at the other 172, worst at 38 points where it is 97.4359%.

The rank of one number decides all of it

The construction has four steps and none of them looks at a distribution. Fit whatever model is wanted on one part of the sample. Score every point of a second part — the calibration set — by how badly the model does on it, which for a line is the absolute residual. Sort those m scores. Build the interval so that it contains every future value whose own score would fall at or below the k-th smallest of them.

The claim is then about ranks. If the m calibration scores and the future score are exchangeable — if every ordering of the m+1 of them is as likely as every other — then the future score’s rank among all m+1 is uniform on the m+1 possibilities. The interval covers exactly when that rank is k or below, so

P(covered)  =  km+1,k=(m+1)(1α).P(\text{covered}) \;=\; \frac{k}{m+1}, \qquad k = \lceil (m+1)(1-\alpha)\rceil .

At 100 calibration points and a promised 95% that index is 96, and the coverage is 96/101 = 95.0495%. At 200 points it is 191/201 = 95.0249%. Neither number depends on the model, the sample size used to fit it, or the shape of the noise. The model can be a straight line through data that curves, a constant, or a random number generator; the coverage does not move, because nothing in the argument above read the model’s output for anything except its rank.

That is a strange kind of guarantee and it is worth being precise about what has been given up to get it. The interval is not promised to be short, and it is not promised to be right for any particular future observation. It is promised to contain the future value in a stated share of repetitions averaged over everything — over the training set, over the calibration set, and over which member of the population arrives next. What that averaging conceals is the whole of the third argument in this field, and it is large.

Forty thousand orderings, and not one of them a draw

The closed form above is the rank argument written down. It is worth having the rank argument counted, because the step a reader has to believe — that each of the m+1 ranks is equally likely — is exactly the step that is asserted rather than shown.

So take m+1 distinct values and deal them out in every possible order: m of them to calibration slots, one to the test slot. Sort the calibration slots, find the k-th, and ask whether the test value is at or below it. Nothing in that walk divides by m+1 and nothing takes a ceiling, so an agreement with the closed form is an agreement between two pieces of arithmetic rather than one piece checked against itself — which is the standard every number here is held to.

Every rank is equally likely, and that is the whole argument. All 40320 orderings of 8 exchangeable values, counted by the rank the future observation takes among them. Each rank occurs in exactly 5040 of them, because fixing the future observation's rank leaves 7! ways to arrange the rest. An interval built at the 6-th smallest of the 7 calibration scores therefore covers whenever that rank is 6 or below, which is 30240 of the 40320 orderings — 0.7500, exactly ⌈(m+1)(1−α)⌉/(m+1) at m = 7 and α = 0.25. Nothing about the values entered the count, which is why the coverage does not depend on where they came from.
Fig. 2 All 40,320 orderings of eight exchangeable values, counted by the rank the future observation takes. Each rank occurs in exactly 5,040 of them, and the six ranks at or below the sixth order statistic are the covered ones — 0.750000 at α = 0.25, exactly ⌈8 × 0.75⌉/8.

The counts are flat, and their flatness is the mechanism rather than a consequence of it. Fixing the future observation’s rank leaves 7! = 5,040 ways to arrange the other seven values, and 5,040 is the same number whichever rank was fixed. The eight numbers are therefore not merely equally likely on average across some distribution; they are equal because a factorial does not know which slot it was counting around.

The choice of α = 0.25 for the picture is not cosmetic either. At five per cent and seven calibration points the index would be eight, which exceeds seven, so every ordering would be covered and the bars would all be one colour — a true count of a rule that is the whole real line. A quarter is the largest miss rate at which eight values still separate the covered ranks from the missed ones, and separating them is what the figure is for.

At m = 7 and α = 0.25 the enumerated coverage is 0.750000 against a closed form of 6/8, and the same agreement holds at every m from three to seven and at each of four miss rates. It stops there for a reason that is not a limitation: eight factorial is 40,320 and nine factorial is ten times that, and the claim being checked is about a rank distribution. A rank distribution that is uniform on eight things is uniform on two hundred, because the counting argument does not mention how many there were.

The exactness is not ninety-five per cent

The coverage is a ratio of two whole numbers, and 0.95 is a ratio of two whole numbers only sometimes. ⌈(m+1)(1−α)⌉/(m+1) equals 1−α exactly when (m+1)α is an integer, which at five per cent means m+1 a multiple of twenty — nineteen, thirty-nine, fifty-nine, and so on. Between 19 and 200 calibration points that is 10 of the 182 sizes. At the other 172 the coverage is above the promise.

The overshoot is not small at the worst of them. At 38 calibration points the index is also 38 — the largest score in the set — so the coverage is 38/39 = 97.4359%, which is 2.4359 points of coverage nobody asked for. Averaged across the whole range the excess is 0.5661 points, which is small in a table and is not nothing: an interval covering 97.4% when 95% was wanted is wider than the question required, and width is what an interval is spent on.

The coverage is exact and it is not the nominal rate. ⌈(m+1)(1−α)⌉/(m+1) against m, the number of calibration points, at α = 0.1. It is a closed form and needs no data. It never falls below 90.0% and never reaches 1−α+1/(m+1), the two bounds the rank argument gives. It equals 90.0% exactly at 20 of the 192 sizes drawn — the sizes where (m+1)α is a whole number, which are 10 apart — and sits above it everywhere else, worst at 18 points where it is 94.7368%, or 4.7368% of coverage nobody asked for. Below 9 points there is no such order statistic and the interval is the whole line, which is where the curve starts.
Fig. 3 The same closed form at a promised 90%. The teeth are ten calibration points apart rather than twenty, and the same shape holds: exact where (m+1)α is a whole number, above it everywhere else.

The teeth are wider at a smaller α and narrower at a larger one, and the reason is the same arithmetic: the exact sizes are the multiples of 1/α less one. A reader who reads the first curve as a convergence — noise that a larger calibration set will grind away — has read it wrong. The envelope narrows, because the excess is bounded by 1/(m+1), but the sawtooth does not smooth out, and the difference between 200 and 201 calibration points is not a difference in precision. It is a different guarantee, exactly stated in both cases.

Below nineteen calibration points there is no interval at all

The index k = ⌈(m+1)(1−α)⌉ can exceed m, and when it does there is no such order statistic. The honest quantile is then infinite and the honest interval is the whole real line. That happens below ⌈1/α⌉ − 1 calibration points — 19 at five per cent — and it is not a convention adopted to keep the theorem tidy. It is the correct answer: with eighteen scores in hand, no finite distance can be claimed to be exceeded by at most one future score in twenty, because there is no eighteen-of-nineteen ordering that would license it.

Counted at a sample of twenty split in half, every draw under every one of six noise shapes returns an infinite interval, and the counted coverage is 100.00% — a true statement about a rule that says nothing. The first feasible size, twenty calibration points, covers 20/21 = 95.2381%.

Coverage against sample size, true proportion 0.15. Coverage does not improve monotonically. n = 19 covers 93.8% while the larger n = 20 covers 81.9%. The sample space is discrete, so the endpoints jump as n changes.
Fig. 4 The same shape from a different construction: the exactly-summed coverage of an interval for a proportion against the sample size. A coverage that is exact and built on a discrete object oscillates rather than converging, and reading the wobble as noise is the error in both cases.

That picture is from a different field and it is here because the two failures are the same failure. The coverage of an interval for a proportion oscillates with the sample size because the sample space is discrete and a sum of binomial terms rarely lands on 0.95; a sample of twenty covers twelve points worse than a sample of nineteen there. The coverage of a conformal interval oscillates with the calibration size because a ratio of small whole numbers rarely lands on 0.95 either. In neither case is more data a repair, and in neither case is the wobble an artefact of measurement — both were computed rather than sampled, which is the only reason the structure is visible at all.

Six noise shapes, and one column that does not move

The interval a reader is actually handed for a future observation is yˉ±ts1+1/n\bar{y} \pm t\,s\sqrt{1 + 1/n}, which is exact at 1−α under a normal parent and derived under nothing else. Measured on the same draws, at 200 observations, under six noise shapes standardised so that only their shape differs, it covers 95.13% under a normal, 94.47% under a t on three degrees of freedom, 96.07% under a lognormal, 94.60% under an exponential, 96.73% under a contaminated normal and 97.80% under a Cauchy — a spread of 3.33 points. The conformal column across the same six spans 0.77 points, which at 3,000 draws is about two standard errors of a binomial rate and is therefore consistent with the one number the closed form says it must be.

One column is a closed form and the other is a measurement. What each interval covers over 3000 draws of 200 observations, under six noise shapes standardised so that only their shape differs — except the Cauchy, which has neither a mean nor a variance to standardise. The conformal interval's coverage is 95.0495% before any data arrive, because it is ⌈(m+1)(1−α)⌉/(m+1) at 100 calibration points, and the counted column stays within 0.77% of itself across all six. The textbook interval runs from 94.47% to 97.80%, a spread of 3.33% — and its widest failure is the one where its own ingredients do not exist: under the Cauchy it is 619.18 wide against 42.08.
Fig. 5 What each interval covers over 3,000 draws of 200 observations under six noise shapes. The conformal column is 95.0495% before any data arrive; the textbook column runs from 94.47% to 97.80%.

The Cauchy row is the one worth stopping at, because it is where the comparison stops being about accuracy. A Cauchy has neither a mean nor a variance, so the two quantities the textbook interval is assembled from do not exist, and every sample estimate of them is a statistic with no limit to converge to. The interval is nevertheless computable, and it covers 97.80% — for a mean width of 619.175 against the conformal interval’s 42.077. Fifteen times the width for two and a half points of coverage that were not wanted. Nothing warns of this in the output; the arithmetic runs, the number prints, and the interval is useless in a way that a coverage column alone reports as a success.

The level is kept and the ends are not

A coverage column can be right while the interval is wrong, and the skewed parents are where that happens. A two-sided 95% interval promises 2.5% of misses above and 2.5% below. Under an exponential at 200 observations the textbook interval misses 5.40% above and 0.00% below; under a lognormal, 3.93% above and 0.00% below. Its overall coverage under those two parents is 94.60% and 96.07%, which is to say that one of them looks fine and the other looks conservative, and both are broken at both ends.

The level is right and the ends are not. Which end of the textbook prediction interval ȳ ± t·s√(1+1/n) the misses are at, over 3000 draws of 200 observations under each of six noise shapes. A two-sided 95.0% interval promises 2.5% at each end. Under an exponential it misses 5.40% above and 0.00% below; under a lognormal, 3.93% above and 0.00% below. Its overall coverage under those two parents is 94.60% and 96.07%, which is why a coverage column cannot show this: the promise is broken at both ends and the two errors cancel in the total.
Fig. 6 Which end of the textbook interval the misses are at, under each of six noise shapes. On a skewed parent every miss is on the same side, and the two tails’ errors cancel in the total.

This is the defect the essay on which tail a threshold sits in is about, arriving in a prediction problem: an interval built symmetrically around a centre cannot be right at both ends of an asymmetric distribution, and averaging the two errors is what hides it. A reader making a one-sided decision — is the next value going to be too high — is being told 2.5% and getting 5.4%, and no amount of the coverage column would have said so.

The conformal interval built from an absolute residual has the same defect and for the same reason: |y − ŷ| is symmetric in the sign of the residual, so the interval it produces is symmetric too. Its coverage is still exactly what the rank argument promises, because the rank argument does not read the score. That is the first sign of where this field is going: the symmetry is a property of the score, the score is a modelling choice, and everything a reader wants from an interval turns out to live there rather than in the guarantee.

What could have made the count wrong

Three ways the numbers above could have come out as they did without the claim holding, and what was done about each.

The enumeration could have been the closed form in disguise. If the walk over orderings had computed k from the same ceiling and divided by the same m+1, an agreement would prove only that the expression was typed twice. It does not: the walk sorts m values, finds the rank of the test value by comparison, and adds one to a counter, and the coverage it reports is a count divided by a count of orderings. The agreement is between a factorial argument and a piece of integer arithmetic that share no line.

The parent columns are counted rather than summed, and 3,000 draws is not many. A binomial rate near 0.95 at 3,000 draws has a standard error of about 0.40 points, so a 0.77-point spread across six conformal readings is what six independent readings of one number look like, and a 3.33-point spread across six textbook readings is not. The three largest textbook departures — the contaminated normal, the Cauchy and the lognormal — sit four to seven standard errors from the normal row, which is the margin the comparison rests on. The tail rates are the sharper evidence anyway, because 5.40% against 0.00% is not a rate that noise produces at any trial count worth arguing about.

The parents could have differed in scale rather than in shape. Five of the six are standardised to mean zero and variance one on purpose, so that an interval reading a variance is handed the same variance under all of them and anything that moves is attributable to shape. The sixth cannot be standardised, which is the point of including it. Without that step the width column would have been a comparison of six different problems.

And the scores could have tied. The rank argument needs the m+1 scores to have a strict ordering; ties split a rank between two slots and the uniformity above is a statement about a permutation with no repeats in it. Nothing here produces one, because every score is a continuous function of a continuous draw and the probability of an exact tie is zero — but that is an argument about the parents used, not about the method, and a score taking finitely many values would need the ties broken at random before any of this applied. The enumeration deals distinct values for the same reason.

What has not been ruled out is that the counted conformal column agrees with its closed form for a reason other than the one claimed — the code path that builds the interval and the code path that computes 96/101 are separate, but they were written by the same hand on the same afternoon, and a shared misunderstanding is not something an agreement can detect. The enumeration is the guard against that, and it is why it is in the field at all.

Where the guarantee stops

Three limits, all of them measured further along this ladder rather than argued here.

The guarantee is marginal. It averages over which member of the population arrives, and an average can be kept exactly while no part of it is right: one interval for two equally common groups whose noise scales differ by three covers 100.00% of one and 90.66% of the other. How far that can be pushed is arithmetic rather than a measurement, and the answer is all the way to zero.

The guarantee says nothing about width, which is the whole of what the choice of score and the choice of split decide. What splitting the sample costs is 1.38% of width and nothing at all in coverage, and the reason the coverage column does not move is the sawtooth above: every split has its own exact promise and keeps it.

And exchangeability is a real assumption, not a technicality. It is weaker than independence and it is broken by ordinary things — a drifting scale, a shift in who is being predicted, serial correlation — and the three of them cost 4.93, 11.07 and 1.07 points respectively, in an order that is the reverse of how easily a test would have found them.

What none of that touches is the claim this essay opened with. The coverage was known before any data arrived, it was counted rather than asserted, and it held at 200 observations under a distribution with no mean. Every other field on this site had to measure what its interval delivered because the promise and the delivery were different quantities. Here they are the same quantity — and the interesting question turns out to be what a promise that exact is worth, which is the question the rest of this ladder answers.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Calibration setClosed formConformal predictionCoverageDistribution-freeEmpirical quantileExact enumerationExchangeabilityForecast intervalMarginal coverageMonte CarloNonconformity scoreOrder statisticOvercoverage