Series

Reference — the series

11 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. The allocations the rule could have made, from these exact patients. One 200-patient trial allocated by response-adaptive randomisation, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the same rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 1.417, 258 of the 999 re-randomisations reach it, and the p-value is (1 + 258)/(1 + 999) = 0.2590. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 2.101.

    The experiments that could have happened

    An adaptive trial's allocation is a function of the outcomes it will later be compared against, so the ordinary analysis rejects a true null 9.2% of the time. Hold the outcomes fixed, re-run the rule that assigned them, and count — the same statistic against a reference distribution the trial could actually have drawn from is back at 4.0%.

    part 1 · exact
  2. The allocations the rule could have made, from these exact patients. One 200-patient trial allocated by response-adaptive randomisation, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs a fair coin rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 1.417, 155 of the 999 re-randomisations reach it, and the p-value is (1 + 155)/(1 + 999) = 0.1560. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 1.946.

    The test that needs the rule

    A randomisation test assumes almost nothing about the data and one thing about the experiment. Tell it a fair coin produced an allocation that an adaptive rule produced — which is what every off-the-shelf permutation routine does — and it rejects 8.0% of true nulls where knowing the rule gives 4.0%.

    part 2 · exact
  3. Where the +1 matters, and why nobody has noticed that it does. The true size of the two rules at every B, computed rather than simulated: under the null the count of re-randomisations reaching the observed statistic is uniform over {0 … B}, so both sizes are integer arithmetic. With the +1 the size is (⌊α(B+1)⌋)/(B+1), which never exceeds 5%. Without it the size is (⌊αB⌋+1)/(B+1), which is larger except at B = 19, 39, 59 — the values with B + 1 a multiple of 1/α, and the values everybody uses. At B = 19 the two rules are the same rule; at B = 20 the uncorrected one is a 9.5% test. The marks are simulated on 500 trials of 120 patients, as the second route to the same numbers.

    The plus one and the round number

    A sampled randomisation test counts the observed allocation as one of its own reference draws, and the correction is invisible at B = 19, 39, 59 and 999 — every value anybody uses. At B = 20 the version without it is an 8.00% test where the corrected one is 3.80%, and the convention protecting everybody is a preference for round numbers minus one.

    part 3 · exact
  4. What the corner costs when the table is full of hopeless candidates. The benchmark holds two predictors, one of which is worth 1; a third predictor, worth the amount on the horizontal axis, is held only by candidates the benchmark does not contain. At the left the null is true and both procedures hold their level. To the right there is a genuinely better candidate, and the uncorrected reality check finds it 2.7% of the time while the same test with the clearly bad columns recentred finds it 51.0% of the time. The columns doing the damage are the ones nobody would have looked at twice: they are so far behind that they cannot win, and calibrating as though they might is what makes the test blind.

    The corner the test is calibrated at

    "No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.

    part 4 · select
  5. A bias against a variance, with the answer in between. How wrong one sample's reference distribution is, split into the two things it is wrong by. Sharing the multiplier over more rows keeps more of the dependence and closes the bias from 1.688 to 0.835; every row it is shared over also removes an independent sign from the 101 the sample started with, and the spread of the resulting quantile rises from 1.307 to 2.172. The distance a practitioner with one sample is actually exposed to is the two together, and it is smallest at ℓ = 5.

    How long a block a multiplier shares

    Sharing a sign over more rows keeps more of the dependence and leaves fewer independent signs to build a distribution from. The bias falls from 1.6885 to 0.8479 and the spread rises from 1.3073 to 2.1716, and the rejection rate walks straight through its nominal level on the way from 11.3% to 1.3%.

    part 5 · proxy
  6. The rate falls geometrically; the count does not fall at all. The share of equal splits of two hundred units that a tolerance of 1 coin-spreads admits, against the number of functions the tolerance is stated for, with the closed form (2Φ(1) − 1)^k drawn beside it. The rate falls by about two thirds with every constraint. The admissible count is that rate times C(200, 100), and it goes from 2^195 to 2^192 — it does not fall in any sense a trial cares about. What the falling rate costs is sampling: 9,878 draws to collect a thousand admissible ones at six constraints, against 1,465 at one.

    What a reference distribution costs to sample

    A randomisation test on a trial too large to enumerate has to sample its reference distribution, at 1/p attempts per draw and a p-value resolved to 1/(B + 1). Six constraints cost 9,878 attempts per thousand draws, and a thousand draws resolve p to 9.99·10⁻⁴ and not one digit finer.

    part 6 · product
  7. The ceiling a multiplier cannot reach past. A wild-type resampling forms e*_t = e_t·w_t with the multiplier independent of the residual, so what comes out has autocovariance γ_resid(k)·γ_w(k) — the residuals' own, multiplied by the multiplier's. Since |γ_w| ≤ 1 the reference distribution's dependence is bounded above by the residuals', and the residuals' is already below the errors'. The two shortfalls compose. For a block of ℓ the multiplier's autocorrelation is exactly the triangle (1 − k/ℓ)⁺, drawn here as the dashed prediction against the realised resamples at ℓ = 5; the bound is attained only at ℓ = n, where the reference distribution is built from one sign.

    What a multiplier cannot keep

    Two reasons were named for the quarter a blocked resampling falls short, and taking either away makes the gap larger. What is left is a bound — a multiplier can only take dependence out, and the residuals' own is already below the errors'.

    part 7 · effective
  8. Where a walk is cheaper than a hunt. Both costs in the same unit. A rejection sampler evaluates 1/p assignments per independent draw and does not care how large the trial is; a walk evaluates one per step and yields an effective draw every τ steps, and τ is a property of the constraint and the statistic together. They cross at a tolerance of 0.194 standard deviations, where about one assignment in 396 is admissible — far tighter than any trial is designed at. And the walk does not remove the acceptance cost; it pays it once, hunting for somewhere to start.

    Draws that repeat each other

    A hunt costs 1/p evaluations per independent draw. A walk costs one per step and yields an effective draw every τ steps. Both are counted in the same unit, and the walk is dearer at every tolerance a trial is designed at.

    part 8 · joint
  9. What each construction carries, against what there was. The autocorrelation of a resampled error series at five lags, averaged over 60 samples of 40 resamples each. Three facts are in the picture. The residuals lie below the errors at every lag, which is the ceiling a multiplier cannot exceed. The blocked multiplier and the fixed-length block lie on top of each other below it — they attenuate identically, because the attenuation is the join — while the stationary bootstrap, whose runs are geometric rather than fixed, sits above them both. And the sieve is the exception in kind rather than in degree: at lag six it carries 0.0638 where the residuals have 0.0300 and the multiplier has -0.0011, because a fitted model extrapolates past the lags it was told about and a truncated sample sequence cannot.

    Errors generated from a fitted model

    The one construction that is not bounded by the residuals, because a model extrapolates past the lags it was told about and a truncated sample sequence cannot. It is nearly exact where the only defect is dependence, and it pays for it where there are two.

    part 9 · banded
  10. How far each reference distribution's 95% point falls short. Seven constructions on rows that repeat each other, at a block length of 5 and 200 draws, against the statistic's own 95% point of 3.0224 computed from three thousand draws of the same world. Reading down: a multiplier on every row keeps no dependence at all and is 44% short; a multiplier shared along a block keeps the triangle; a fixed block keeps the same triangle and is 8% closer, which is the pair that says a taper is not what decides this; the stationary bootstrap; the two tapered blocks, both further short than the untapered one at this block length; and errors generated from a fitted model, which is the only construction here not bounded by what the residuals report.

    A taper and a critical value

    Two constructions whose tapers visibly differ give the same critical value, and two that share a taper exactly do not. Adding a construction whose taper is a decision rather than an accident says which half of that is true.

    part 10 · taper
  11. The one thing a trial always reports is the one thing that survives. How wrong three p-values are when they are computed over the half of the admissible set a single walk can reach, rather than over all of it, at a fourteen-unit trial where the whole set can be enumerated. The two-sided p-value on the difference in arm means — the number a trial publishes — is wrong by exactly nothing, at every row, to machine precision. That is not luck: the two components are complement pairs and the difference in arm means is exactly negated by the complement, so the distribution of its absolute value is the same on both. A one-sided p-value on the same statistic is out by as much as 0.112, and the largest response observed in the treated arm — a safety reading rather than an effect, and the one statistic here that is not odd under the complement — by as much as 0.172. The defect survived because the commonest thing anybody computes is the one quantity it cannot touch.

    Half a reference distribution

    A walk that reaches half its admissible set reports the two-sided p-value exactly right, to the last digit, for ever. A one-sided one it puts on the wrong side of five per cent about once in thirty.

    part 11 · after

All series