What the split costs
Worth reading first: What the 95% refers to.
Split conformal spends part of the sample on fitting a model and the rest on scoring it, and the fraction is a dial somebody has to set. The natural reading of that dial is that it trades the quality of the fit against the quality of the guarantee: a bigger calibration set, a better estimate of the quantile, a coverage nearer the promise.
That reading is wrong, and the measurement that says so is the coverage column. Swept across nine training shares at 200 observations, it runs from 95.27% to 95.90% — a spread of 0.63 points, which at 3,000 draws is under two standard errors — and every reading sits on the promise its own calibration size makes rather than on some common target it approaches from below. The dial was never deciding coverage. It decides width, and the whole of what it costs is 1.38%.
The column that was supposed to move
The rank argument says the coverage of an interval built at the k-th smallest of m calibration scores is ⌈(m+1)(1−α)⌉/(m+1), and it says so for every m. So a split leaving 180 calibration points promises 95.03% and a split leaving 20 promises 95.24%, and each of them keeps its own promise exactly. There is no size at which the guarantee is better and no size at which it is approached; there is a different exact number at every size, all of them at least 95%.
This is the control for everything else in the essay, and it is worth saying why a control that does nothing is the most informative column in the table. If the coverage had drifted with the split, every width comparison below would be a comparison between procedures delivering different things, and choosing the narrowest would be choosing the one that covers least — which is the failure the essay that priced the shortest interval is about. Because the column does not move, a width comparison is a comparison at fixed coverage, and the narrowest is simply the best.
The coverage column also rules out a specific way the sweep could have been broken. Nine splits measured on independent draws would produce a column with a spread of about a point by sampling alone, and a sweep whose calibration sets overlapped incorrectly — reusing a point for both fitting and calibrating — would produce coverage below the promise, because a point the model has seen has a smaller residual than a point it has not. Neither is visible. Every reading is within four standard errors of its own promise and none is systematically below it.
What the dial actually decides
Two forces pull against each other and they act on the same number.
Spending more of the sample on the fit shrinks the residuals. A least-squares line from twenty points is noticeably worse than one from a hundred, and every calibration score is the distance from a point to that line, so a bad line inflates every score and widens the interval that is built from them.
Spending more on calibration moves the order statistic. At m calibration points the interval is built at the ⌈(m+1)(1−α)⌉-th smallest score, which is the 96th of 100 or the 20th of 20; and the more extreme that index is as a fraction of the set, the larger and the noisier the score sitting there. Twenty calibration points means taking the maximum, and a maximum is the statistic that converges slowest and moves most.
Neither force is about coverage and it is worth being explicit about why, because the instinct that a noisy quantile estimate must cost coverage is a good instinct that is wrong here. A conformal interval does not estimate a quantile of the score distribution. It takes an order statistic of the calibration scores, and the claim attached to it is a statement about where a new score’s rank falls among the m+1 — a statement that is true at m = 20 and at m = 180 with no error term anywhere in it. The noisiness of the twentieth of twenty scores is entirely a fact about how wide the resulting interval is on any given draw, which is precisely what a fit does not choose and what the calibration set does.
The two meet at a half. The width is 4.0416 at a training share of 0.5, 4.1603 at 0.1 — where 180 calibration points are paid for with a line fitted on twenty — and 4.3820 at 0.9, where the fit is excellent and the interval is the largest of twenty scores. The interior minimum is what a trade-off looks like when both of its arms are real.
The whole cost of splitting is one and a third per cent
The comparison that prices the split is not against another split. It is against not splitting at all: full conformal, where every one of the 200 points does both jobs by being refitted with each candidate response included.
That sounds like an infinite loop over a continuum and is not, for least squares. With the candidate response z included, every residual is an affine function of z, so the comparison between the candidate’s own residual and each of the others changes its answer at two values of z and nowhere else:
Collecting the 2n crossings and counting once inside each cell gives the exact set, with no grid and nothing approximated. It is 3.9865 wide, so the best split costs 1.0138 times that, a penalty of 1.38%.
One and a third per cent is smaller than the expectation that motivated the measurement, and small enough that it changes the recommendation rather than qualifying it: the split is close to free at this sample size, and full conformal’s cost is a model refit per candidate response for every prediction rather than one refit ever. The number that makes the comparison honest is that the conformal set for a fitted line came out as a single interval on every one of the 3,000 draws. The algorithm does not assume it — a conformal set for a fitted model need not be an interval, and reporting the convex hull of a set with holes in it would be reporting something wider than what was computed — so the single piece is a measurement rather than a definition.
The sawtooth arrives as a property of a tuning knob
At sixty observations the dial stops behaving like a dial. The width does not fall to a minimum and rise; it falls, rises and falls again: 5.1442 at a training share of 0.1, 4.6412 at 0.2, 4.4089 at 0.3, then 4.9906 at 0.4 — wider than the two either side of it — 4.8070 at 0.5, and 4.6095 at 0.6.
The reason is the sawtooth. A training share of 0.4 leaves 36 calibration points, and ⌈37 × 0.95⌉ = 36 is the largest score in the set — so that split promises 97.30% rather than 95%, and an interval covering 97.3% of the time is wider than one covering 95% of the time by exactly as much as the extra coverage costs. The neighbouring share of 0.3 leaves 42 points, builds at the 41st of 42, and promises 95.35%. Nothing about the fit changed between them; the fit at 0.4 is strictly better, on twenty-four points rather than eighteen. What changed is which order statistic the arithmetic landed on.
So a curve that was drawn as a fact about a promise in the first essay of this ladder is here a fact about a knob, measured on data, in a units a practitioner reads — and the practical form of it is unpleasant. The width is not a smooth function of the split fraction, so it cannot be optimised by a search that assumes it is, and two adjacent settings of a tuning parameter can differ by 13% of the interval for reasons that have nothing to do with the sample in hand. The best split at sixty observations is 0.3, at 1.0668 times full conformal — a penalty of 6.68%, five times the penalty at two hundred.
Three of nine splits have no interval at all
The other edge from the rank argument arrives in the same sweep, and it is more abrupt. Below nineteen calibration points there is no order statistic to build at, so the interval is the whole real line. At sixty observations the three largest training shares — 0.7, 0.8 and 0.9 — leave 18, 12 and 6 calibration points, and every draw at every one of them returns an infinite interval. The largest feasible training share is 0.6833 at sixty observations and 0.9050 at two hundred.
This is a failure whose symptom is not an error. Nothing throws, nothing warns, and the coverage column at those three settings reads a perfect 100.00%. A sweep that reported only coverage would rank them the best three settings of the dial, and a sweep that reported mean width would have to decide what the mean of three thousand infinities is. Reading a coverage column without a width column beside it is exactly the reading that makes a useless interval look good, and here it is not a subtlety — the winning rule is the one that returns the real line.
There is a version of this that bites in practice rather than in a sweep. A pipeline that holds out a fixed fraction — a fifth, a quarter — and applies it to whatever sample arrives will silently cross the feasibility boundary on small batches, and the interval it returns will be the real line without anything in the code being wrong. A held-out fifth reaches nineteen calibration points only at ninety-five observations, and a held-out tenth only at a hundred and ninety; below those the interval is infinite at every sample size, silently, and the coverage it reports is a perfect hundred per cent. The dial’s safe range is a function of the sample size, and it is the one part of the setting that has to be checked rather than chosen.
What the split is paying for elsewhere
The 1.38% is a price, and a price is only interpretable beside what it buys. What splitting buys here is a model that has not seen the points it is being scored on — and that is the same device, used for the same reason, as everywhere else on this site that a sample is divided.
An estimate computed on the data that chose it is biased upward, by an amount that grows with how hard the choosing looked; the essay that measured what selection costs an estimate prices it, and the interval built after the same choice loses four and a half points of coverage on top of the two that estimation alone costs. The repair in both cases is a part of the sample the choice did not touch. Here the choice is the model fit, the untouched part is the calibration set, and the guarantee that follows is stronger than in either of those settings — exact rather than approximately restored — because the rank argument needs nothing about how the model was arrived at.
That last point is the one worth carrying, and it inverts a familiar worry. A search through models is expensive precisely because the search is invisible in the output, and the usual remedy is to account for it. Conformal prediction does not account for it. The training half may be searched, tuned, overfitted and chosen by whim, and the coverage is unchanged, because the calibration scores are exchangeable with the test score whatever produced the function they are scoring. What the search costs is width — an overfitted model has large residuals out of sample, so its calibration scores are large and its intervals are wide. The bill arrives, in full, in the column the guarantee says nothing about.
What could have made these numbers wrong
The comparisons above are between columns of one table, and the table has four ways of lying that were worth closing.
The splits could have been measured on different draws. They are not: one sample is drawn, and all nine splits and full conformal are computed on it before the next sample is drawn. So a difference between two columns is a paired difference, and the sampling error that would dominate a 1.38% comparison at 3,000 unpaired draws is largely differenced away. The coverage column, by contrast, is not helped by the pairing in the direction that matters — the readings share draws, so its 0.63-point spread is if anything an underestimate of what independent readings would show.
Full conformal could have been computed wrongly, and cheaply so. The crossings algorithm is clever and clever is where errors hide. It is checked against a second route that shares none of its arithmetic: the same question asked at each of twenty thousand grid points, refitting the line each time and counting residuals. The two agree, and the grid route is far too slow to use in a sweep, which is exactly why the fast one is allowed to be clever.
The interior minimum could have been a coincidence of one sample size. It is not the same split at both sizes — 0.5 at two hundred and 0.3 at sixty — which is the honest version of the finding: there is an interior optimum, and where it sits depends on n through which order statistics are reachable. A recommendation of “split in half” is right at two hundred observations and nine per cent wide at sixty.
The two sample sizes could be measuring different populations. They are not: the same draw rule generates both, with the same line and the same noise, and only n changes. So the difference between a 1.38% penalty at two hundred and a 6.68% penalty at sixty is attributable to the sample size, and specifically to the fact that at sixty observations every split lands on a coarse part of the sawtooth. That is a claim the closed form can be checked against rather than a pattern read off two numbers, and the closed form agrees: 37 calibration points promise 97.30% and 43 promise 95.35%, computed with no data at all.
And the 1.38% could be an artefact of measuring a width whose distribution is skewed. The mean width is reported, and the widths across draws are not symmetric — a draw with a large maximum residual widens every split at once. That is another reason the pairing matters, and it is why the penalty is reported as a ratio of means on shared draws rather than as a mean of ratios, which would weight the draws with narrow intervals more heavily.
Where the one-and-a-third per cent does not generalise
The affine-residual trick is a property of a linear fit and of nothing else. For any model whose residuals are not affine in the candidate response — a tree, a smoother, a neural network, a fit with a penalty that moves with the data — the exact conformal set cannot be read off 2n crossings, and full conformal has to be evaluated on a grid of candidate responses at one model refit per grid point. The comparison here therefore prices the split against a competitor that is cheap in the one case where it is cheap.
So 1.38% is not a general number, and it should not be carried out of this setting as one. What is general is the structure: the coverage column does not move, the width column does, and the split is a decision about the second alone. That structure is what the rank argument guarantees, and it holds for a tree exactly as it holds for a line.
The comparison also stops short of the question a reader with a fixed budget actually has, which is not how to split n observations but how many to collect. The width at the best split falls from 4.4089 at sixty observations to 4.0416 at two hundred — an 8% improvement for three and a third times the data — while the guarantee is identical at both, and exact at both. That is the shape of every diminishing return on this site, and it is sharper here than usual because the coverage is not improving at all. More data buys width and nothing else.
The other thing this essay has not measured is the split’s effect on anything but the average width. An interval can be narrow on average and the wrong width in every individual case, and whether it is depends on the score rather than on the split — which is where the modelling actually went, and where the averaging that makes all of this exact stops being harmless. Before that, the average itself has to be taken apart: the coverage column that has been so reassuringly flat across this whole sweep is an average over the population, and the population has parts.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The weight that has to be estimated — both name calibration set, closed form, conformal prediction, empirical quantile, exchangeability, interval width, overcoverage
- An interval that covers and says nothing — both name closed form, coverage, interval width, overcoverage
- When the order matters — both name calibration set, conformal prediction, coverage, exchangeability
- A standard error that knows about the instruments — both name closed form, coverage, interval width
- Intervals for the findings — both name closed form, coverage, interval width
- The check worth more than the check — both name closed form, coverage, interval width
Named objects
A flat tag is an object no other essay names yet.
Calibration setClosed formConformal predictionCoverageEmpirical quantileExchangeabilityFull conformalInterval widthLeast squaresOrder statisticOvercoverageSample splittingSplit conformalTuning parameter