Series

Exchangeability — the series

7 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. The coverage is exact and it is not the nominal rate. ⌈(m+1)(1−α)⌉/(m+1) against m, the number of calibration points, at α = 0.05. It is a closed form and needs no data. It never falls below 95.0% and never reaches 1−α+1/(m+1), the two bounds the rank argument gives. It equals 95.0% exactly at 10 of the 182 sizes drawn — the sizes where (m+1)α is a whole number, which are 20 apart — and sits above it everywhere else, worst at 38 points where it is 97.4359%, or 2.4359% of coverage nobody asked for. Below 19 points there is no such order statistic and the interval is the whole line, which is where the curve starts.

    Coverage from exchangeability alone

    A conformal interval's coverage is a fact about the ranks of m+1 numbers, so it can be enumerated before any data arrive — all 40,320 orderings of eight values, agreeing with the closed form to machine precision. What that exactness delivers is not 95%.

    part 1 · conformal
  2. The split decides the width. The width of the interval against the share of 200 observations spent on fitting rather than on calibrating, over 3000 draws. Spending more on the fit shrinks the residuals; spending more on calibration builds the interval at a less extreme order statistic. The two meet at 0.5, where the width is 4.0416 against 4.1603 at 0.1 and 4.3820 at 0.9. Full conformal, which spends the same 200 points on both jobs, is 3.9865 — so the whole cost of splitting is 1.38%.

    What the split costs

    Splitting a sample between fitting and calibrating looks like a trade against the guarantee, and it is not: coverage moves 0.63 points across nine splits and every reading sits on its own promise. The whole cost is 1.38% of width — and at sixty observations the width falls, rises and falls again.

    part 2 · conformal
  3. One promise, kept on average and inside neither group. What each calibration scheme covers inside each of two equally common groups whose noise scales are 1 and 3, over 6000 draws with 200 calibration points. One interval for everybody covers 100.00% of the quiet group and 90.66% of the noisy one, averaging to 95.28% — and the closed form for that population says 99.9999% and 90.0001% at a half-width of 4.9346, from two normal cdfs and no simulation. Dividing by an estimated per-group scale gives 95.29% and 95.48%; calibrating separately inside each group gives 95.49% and 96.01% against a closed-form expectation of 95.4645%. Only the last of those is a guarantee rather than a repair, because the rank argument runs inside each group.

    Marginal is not conditional

    One exactly valid interval covers 100.00% of a quiet group and 90.66% of a noisy one, and the floor is arithmetic rather than a measurement — a group of share π is guaranteed only 1 − α/π, which is zero when the group is as rare as the miss rate.

    part 3 · conformal
  4. Six scores, one coverage. What each nonconformity score's interval covers, over 2000 draws with 200 calibration points, against the 95.0249% the rank argument promises. The column runs from 94.30% to 94.90%, a spread of 0.60% against a standard error of a difference of 0.69% — one number, six times. That includes a score aimed five units off the fit and a score that never reads the response at all, because the rank argument does not read the score either: it needs the scores exchangeable and nothing else. Every decision a modeller makes has to show up somewhere else, and the next two readings are where.

    The score is the modelling

    Six nonconformity scores on the same draws cover within 0.60 points of each other, against a standard error of a difference of 0.69 — one number six times. Their widths run over a factor of 2.361 and their adaptivity over a factor of 8.377.

    part 4 · conformal
  5. What a scale that grows across the sample costs. What the interval covers when the noise scale grows across the sample, against how far the departure has gone, over 1500 draws at each setting. The coverage runs from 94.47% at no departure to 83.93% at the end of the sweep, a loss of 11.07%. The rank argument needs the 200 calibration scores and the test score to be exchangeable, and this is one of the three ways that fails. A test built for it reaches 80% power at 4.054, where the coverage is 85.13% — so 9.87% of the loss is inside the region such a test would have missed.

    When the order matters

    Three ways of breaking exchangeability cost 4.93, 11.07 and 1.07 points of coverage, and the ordering by cost is the reverse of the ordering by how soon a test would have caught them. The departure practitioners check for is the cheapest one.

    part 5 · conformal
  6. Too small breaks it and too large does not. Coverage of the weighted interval against the factor the true likelihood ratio is multiplied by, at a test population 80.0% drawn from the noisier group and 200 calibration points. The exact weight is the factor of 1 and covers 95.70%. Overstating it costs nothing: 96.13% at sixteen times too large. Understating it costs, and costs steeply below about a half — 94.93% at half, 88.37% at an eighth and 67.90% at a thirtieth. The question this answers was whether a wrong weight degrades smoothly or falls off a cliff, and the answer is that it does neither symmetrically: the curve is smooth and one-sided.

    The weight that has to be estimated

    A likelihood ratio sixteen times too large costs 5.5% of interval width and no coverage at all; one a thirtieth of the right size covers 67.90%. The estimate from a batch of five unlabelled covariates covers 95.10% against an exact repair's 95.30%, and the binomial says why.

    part 6 · conformal
  7. Three detectors for one departure, all at 5%. How often each of three checks on the calibration scores fires, against the size of the drift, with every critical value simulated under no drift so that all three sit at 5.0% exactly. The incumbent — a rank comparison of the first half of the scores against the second — reaches four-in-five power at a growth factor of 4.31. Reading each score's rank against its position reaches it at 2.65, and the largest running departure of the scores from their mean at 2.12. The ordering of the three is the ordering by how much of the sample's arrangement each one uses.

    A detector built for the ordering

    The best of three checks for a drifting scale fires at half the growth factor the standard one needs — 2.12 against 4.31 — and still leaves 6.50 points of coverage gone before it does, against 0.51 for serial correlation. The reversal was not a property of the test.

    part 7 · conformal

All series