Coverage without a distribution

What the split costs

Splitting a sample between fitting and calibrating looks like a trade against the guarantee, and it is not: coverage moves 0.63 points across nine splits and every reading sits on its own promise. The whole cost is 1.38% of width — and at sixty observations the width falls, rises and falls again.

Worth reading first: What the 95% refers to.

Split conformal spends part of the sample on fitting a model and the rest on scoring it, and the fraction is a dial somebody has to set. The natural reading of that dial is that it trades the quality of the fit against the quality of the guarantee: a bigger calibration set, a better estimate of the quantile, a coverage nearer the promise.

That reading is wrong, and the measurement that says so is the coverage column. Swept across nine training shares at 200 observations, it runs from 95.27% to 95.90% — a spread of 0.63 points, which at 3,000 draws is under two standard errors — and every reading sits on the promise its own calibration size makes rather than on some common target it approaches from below. The dial was never deciding coverage. It decides width, and the whole of what it costs is 1.38%.

The split decides the width. The width of the interval against the share of 200 observations spent on fitting rather than on calibrating, over 3000 draws. Spending more on the fit shrinks the residuals; spending more on calibration builds the interval at a less extreme order statistic. The two meet at 0.5, where the width is 4.0416 against 4.1603 at 0.1 and 4.3820 at 0.9. Full conformal, which spends the same 200 points on both jobs, is 3.9865 — so the whole cost of splitting is 1.38%.
Fig. 1 The interval’s width against the share of 200 observations spent on fitting rather than calibrating, over 3,000 draws. The minimum is at a half, 4.0416, and the horizontal rule is full conformal at 3.9865 — the same 200 points doing both jobs.

The column that was supposed to move

The rank argument says the coverage of an interval built at the k-th smallest of m calibration scores is ⌈(m+1)(1−α)⌉/(m+1), and it says so for every m. So a split leaving 180 calibration points promises 95.03% and a split leaving 20 promises 95.24%, and each of them keeps its own promise exactly. There is no size at which the guarantee is better and no size at which it is approached; there is a different exact number at every size, all of them at least 95%.

The split decides nothing about the coverage. What the interval covers against the share of 200 observations spent on fitting, over 3000 draws, with two standard errors either side and each split's own promise beside it. The column runs from 95.27% to 95.90% — a spread of 0.63% — against promises from 95.03% to 95.24%. It does not move, because the split was never deciding it: the rank argument holds at every calibration size, and what changes with the size is the interval's width and the amount by which it overcovers.
Fig. 2 What each split covers, with two standard errors either side and its own calibration size’s promise beside it. The column spans 0.63 points across nine settings and tracks the promises rather than a common target.

This is the control for everything else in the essay, and it is worth saying why a control that does nothing is the most informative column in the table. If the coverage had drifted with the split, every width comparison below would be a comparison between procedures delivering different things, and choosing the narrowest would be choosing the one that covers least — which is the failure the essay that priced the shortest interval is about. Because the column does not move, a width comparison is a comparison at fixed coverage, and the narrowest is simply the best.

The coverage column also rules out a specific way the sweep could have been broken. Nine splits measured on independent draws would produce a column with a spread of about a point by sampling alone, and a sweep whose calibration sets overlapped incorrectly — reusing a point for both fitting and calibrating — would produce coverage below the promise, because a point the model has seen has a smaller residual than a point it has not. Neither is visible. Every reading is within four standard errors of its own promise and none is systematically below it.

What the dial actually decides

Two forces pull against each other and they act on the same number.

Spending more of the sample on the fit shrinks the residuals. A least-squares line from twenty points is noticeably worse than one from a hundred, and every calibration score is the distance from a point to that line, so a bad line inflates every score and widens the interval that is built from them.

Spending more on calibration moves the order statistic. At m calibration points the interval is built at the ⌈(m+1)(1−α)⌉-th smallest score, which is the 96th of 100 or the 20th of 20; and the more extreme that index is as a fraction of the set, the larger and the noisier the score sitting there. Twenty calibration points means taking the maximum, and a maximum is the statistic that converges slowest and moves most.

Neither force is about coverage and it is worth being explicit about why, because the instinct that a noisy quantile estimate must cost coverage is a good instinct that is wrong here. A conformal interval does not estimate a quantile of the score distribution. It takes an order statistic of the calibration scores, and the claim attached to it is a statement about where a new score’s rank falls among the m+1 — a statement that is true at m = 20 and at m = 180 with no error term anywhere in it. The noisiness of the twentieth of twenty scores is entirely a fact about how wide the resulting interval is on any given draw, which is precisely what a fit does not choose and what the calibration set does.

The two meet at a half. The width is 4.0416 at a training share of 0.5, 4.1603 at 0.1 — where 180 calibration points are paid for with a line fitted on twenty — and 4.3820 at 0.9, where the fit is excellent and the interval is the largest of twenty scores. The interior minimum is what a trade-off looks like when both of its arms are real.

The whole cost of splitting is one and a third per cent

The comparison that prices the split is not against another split. It is against not splitting at all: full conformal, where every one of the 200 points does both jobs by being refitted with each candidate response included.

That sounds like an infinite loop over a continuum and is not, for least squares. With the candidate response z included, every residual is an affine function of z, so the comparison between the candidate’s own residual and each of the others changes its answer at two values of z and nowhere else:

ri(z)=ai+biz,ri(z)  rn+1(z).r_i(z) = a_i + b_i z, \qquad |r_i(z)| \ \ge\ |r_{n+1}(z)| .

Collecting the 2n crossings and counting once inside each cell gives the exact set, with no grid and nothing approximated. It is 3.9865 wide, so the best split costs 1.0138 times that, a penalty of 1.38%.

What splitting the sample is worth. Each split's interval width as a multiple of full conformal's, which fits and calibrates on all 60 points at once and is computed exactly rather than on a grid, over 3000 draws. The best split is 0.3 to the fit at 1.0668 times full conformal's 4.1328 — a penalty of 6.68% — and the worst feasible one is 0.1 at 1.2447. The 3 largest training shares have no finite interval at all: they leave 18, 12, 6 calibration points against the 19 a 95.0% interval needs, so their bars are drawn at a fixed height and mean infinity.
Fig. 3 Each split’s width as a multiple of full conformal’s, at sixty observations. The three largest training shares leave fewer than nineteen calibration points and have no finite interval at all, which is why their bars are drawn at a fixed height.

One and a third per cent is smaller than the expectation that motivated the measurement, and small enough that it changes the recommendation rather than qualifying it: the split is close to free at this sample size, and full conformal’s cost is a model refit per candidate response for every prediction rather than one refit ever. The number that makes the comparison honest is that the conformal set for a fitted line came out as a single interval on every one of the 3,000 draws. The algorithm does not assume it — a conformal set for a fitted model need not be an interval, and reporting the convex hull of a set with holes in it would be reporting something wider than what was computed — so the single piece is a measurement rather than a definition.

The sawtooth arrives as a property of a tuning knob

At sixty observations the dial stops behaving like a dial. The width does not fall to a minimum and rise; it falls, rises and falls again: 5.1442 at a training share of 0.1, 4.6412 at 0.2, 4.4089 at 0.3, then 4.9906 at 0.4 — wider than the two either side of it — 4.8070 at 0.5, and 4.6095 at 0.6.

The split decides the width. The width of the interval against the share of 60 observations spent on fitting rather than on calibrating, over 3000 draws. Spending more on the fit shrinks the residuals; spending more on calibration builds the interval at a less extreme order statistic. The two meet at 0.3, where the width is 4.4089 against 5.1442 at 0.1 and 4.6095 at 0.6. Full conformal, which spends the same 60 points on both jobs, is 4.1328 — so the whole cost of splitting is 6.68%. The 3 largest shares are not on the axis: they leave too few calibration points for a finite interval.
Fig. 4 The same sweep at sixty observations. The width is not monotone on either side of its minimum: the reading at a training share of 0.4 is wider than the readings at 0.3 and at 0.5.

The reason is the sawtooth. A training share of 0.4 leaves 36 calibration points, and ⌈37 × 0.95⌉ = 36 is the largest score in the set — so that split promises 97.30% rather than 95%, and an interval covering 97.3% of the time is wider than one covering 95% of the time by exactly as much as the extra coverage costs. The neighbouring share of 0.3 leaves 42 points, builds at the 41st of 42, and promises 95.35%. Nothing about the fit changed between them; the fit at 0.4 is strictly better, on twenty-four points rather than eighteen. What changed is which order statistic the arithmetic landed on.

The coverage is exact and it is not the nominal rate. ⌈(m+1)(1−α)⌉/(m+1) against m, the number of calibration points, at α = 0.05. It is a closed form and needs no data. It never falls below 95.0% and never reaches 1−α+1/(m+1), the two bounds the rank argument gives. It equals 95.0% exactly at 5 of the 82 sizes drawn — the sizes where (m+1)α is a whole number, which are 20 apart — and sits above it everywhere else, worst at 38 points where it is 97.4359%, or 2.4359% of coverage nobody asked for. Below 19 points there is no such order statistic and the interval is the whole line, which is where the curve starts.
Fig. 5 The closed form that produces it: exact coverage against calibration size over the range a sixty-observation split can reach. The teeth at 36 and at 24 points are the two settings the width sweep jumps at.

So a curve that was drawn as a fact about a promise in the first essay of this ladder is here a fact about a knob, measured on data, in a units a practitioner reads — and the practical form of it is unpleasant. The width is not a smooth function of the split fraction, so it cannot be optimised by a search that assumes it is, and two adjacent settings of a tuning parameter can differ by 13% of the interval for reasons that have nothing to do with the sample in hand. The best split at sixty observations is 0.3, at 1.0668 times full conformal — a penalty of 6.68%, five times the penalty at two hundred.

Three of nine splits have no interval at all

The other edge from the rank argument arrives in the same sweep, and it is more abrupt. Below nineteen calibration points there is no order statistic to build at, so the interval is the whole real line. At sixty observations the three largest training shares — 0.7, 0.8 and 0.9 — leave 18, 12 and 6 calibration points, and every draw at every one of them returns an infinite interval. The largest feasible training share is 0.6833 at sixty observations and 0.9050 at two hundred.

This is a failure whose symptom is not an error. Nothing throws, nothing warns, and the coverage column at those three settings reads a perfect 100.00%. A sweep that reported only coverage would rank them the best three settings of the dial, and a sweep that reported mean width would have to decide what the mean of three thousand infinities is. Reading a coverage column without a width column beside it is exactly the reading that makes a useless interval look good, and here it is not a subtlety — the winning rule is the one that returns the real line.

There is a version of this that bites in practice rather than in a sweep. A pipeline that holds out a fixed fraction — a fifth, a quarter — and applies it to whatever sample arrives will silently cross the feasibility boundary on small batches, and the interval it returns will be the real line without anything in the code being wrong. A held-out fifth reaches nineteen calibration points only at ninety-five observations, and a held-out tenth only at a hundred and ninety; below those the interval is infinite at every sample size, silently, and the coverage it reports is a perfect hundred per cent. The dial’s safe range is a function of the sample size, and it is the one part of the setting that has to be checked rather than chosen.

Expected width against coverage, n = 30, p = 0.15. The Wald interval is the shortest and covers 94.2%. Clopper–Pearson covers 98.3% and is 13% wider. Shortness is not a virtue on its own — an interval of zero width is the shortest of all.
Fig. 6 The same trade in the setting that named it: what four interval rules for a proportion pay in width for the coverage they achieve. Coverage alone never chooses; width is the second number every one of these comparisons is decided on.

What the split is paying for elsewhere

The 1.38% is a price, and a price is only interpretable beside what it buys. What splitting buys here is a model that has not seen the points it is being scored on — and that is the same device, used for the same reason, as everywhere else on this site that a sample is divided.

An estimate computed on the data that chose it is biased upward, by an amount that grows with how hard the choosing looked; the essay that measured what selection costs an estimate prices it, and the interval built after the same choice loses four and a half points of coverage on top of the two that estimation alone costs. The repair in both cases is a part of the sample the choice did not touch. Here the choice is the model fit, the untouched part is the calibration set, and the guarantee that follows is stronger than in either of those settings — exact rather than approximately restored — because the rank argument needs nothing about how the model was arrived at.

That last point is the one worth carrying, and it inverts a familiar worry. A search through models is expensive precisely because the search is invisible in the output, and the usual remedy is to account for it. Conformal prediction does not account for it. The training half may be searched, tuned, overfitted and chosen by whim, and the coverage is unchanged, because the calibration scores are exchangeable with the test score whatever produced the function they are scoring. What the search costs is width — an overfitted model has large residuals out of sample, so its calibration scores are large and its intervals are wide. The bill arrives, in full, in the column the guarantee says nothing about.

What could have made these numbers wrong

The comparisons above are between columns of one table, and the table has four ways of lying that were worth closing.

The splits could have been measured on different draws. They are not: one sample is drawn, and all nine splits and full conformal are computed on it before the next sample is drawn. So a difference between two columns is a paired difference, and the sampling error that would dominate a 1.38% comparison at 3,000 unpaired draws is largely differenced away. The coverage column, by contrast, is not helped by the pairing in the direction that matters — the readings share draws, so its 0.63-point spread is if anything an underestimate of what independent readings would show.

Full conformal could have been computed wrongly, and cheaply so. The crossings algorithm is clever and clever is where errors hide. It is checked against a second route that shares none of its arithmetic: the same question asked at each of twenty thousand grid points, refitting the line each time and counting residuals. The two agree, and the grid route is far too slow to use in a sweep, which is exactly why the fast one is allowed to be clever.

The interior minimum could have been a coincidence of one sample size. It is not the same split at both sizes — 0.5 at two hundred and 0.3 at sixty — which is the honest version of the finding: there is an interior optimum, and where it sits depends on n through which order statistics are reachable. A recommendation of “split in half” is right at two hundred observations and nine per cent wide at sixty.

The two sample sizes could be measuring different populations. They are not: the same draw rule generates both, with the same line and the same noise, and only n changes. So the difference between a 1.38% penalty at two hundred and a 6.68% penalty at sixty is attributable to the sample size, and specifically to the fact that at sixty observations every split lands on a coarse part of the sawtooth. That is a claim the closed form can be checked against rather than a pattern read off two numbers, and the closed form agrees: 37 calibration points promise 97.30% and 43 promise 95.35%, computed with no data at all.

And the 1.38% could be an artefact of measuring a width whose distribution is skewed. The mean width is reported, and the widths across draws are not symmetric — a draw with a large maximum residual widens every split at once. That is another reason the pairing matters, and it is why the penalty is reported as a ratio of means on shared draws rather than as a mean of ratios, which would weight the draws with narrow intervals more heavily.

Where the one-and-a-third per cent does not generalise

The affine-residual trick is a property of a linear fit and of nothing else. For any model whose residuals are not affine in the candidate response — a tree, a smoother, a neural network, a fit with a penalty that moves with the data — the exact conformal set cannot be read off 2n crossings, and full conformal has to be evaluated on a grid of candidate responses at one model refit per grid point. The comparison here therefore prices the split against a competitor that is cheap in the one case where it is cheap.

So 1.38% is not a general number, and it should not be carried out of this setting as one. What is general is the structure: the coverage column does not move, the width column does, and the split is a decision about the second alone. That structure is what the rank argument guarantees, and it holds for a tree exactly as it holds for a line.

The comparison also stops short of the question a reader with a fixed budget actually has, which is not how to split n observations but how many to collect. The width at the best split falls from 4.4089 at sixty observations to 4.0416 at two hundred — an 8% improvement for three and a third times the data — while the guarantee is identical at both, and exact at both. That is the shape of every diminishing return on this site, and it is sharper here than usual because the coverage is not improving at all. More data buys width and nothing else.

The other thing this essay has not measured is the split’s effect on anything but the average width. An interval can be narrow on average and the wrong width in every individual case, and whether it is depends on the score rather than on the split — which is where the modelling actually went, and where the averaging that makes all of this exact stops being harmless. Before that, the average itself has to be taken apart: the coverage column that has been so reassuringly flat across this whole sweep is an average over the population, and the population has parts.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Calibration setClosed formConformal predictionCoverageEmpirical quantileExchangeabilityFull conformalInterval widthLeast squaresOrder statisticOvercoverageSample splittingSplit conformalTuning parameter