The reference the covariates supply
Worth reading first: Randomisation is not balance · The experiments that could have happened.
The previous essay ended with three answers and none of them satisfactory. An unadjusted comparison that assumes nothing and rejects 0.6% of true nulls where it claims 5%. An adjusted comparison that is correct and assumes the outcome model. And two rows of the table that had not been explained.
This essay explains them, and the construction is not new — it is the one the exact field built for a rule that reads outcomes, applied to a rule that reads covariates. What is new is the price, and the price is the finding.
Three lines, and the second one is different here
Fix the outcomes. Under the sharp null the treatment changed nothing for anybody, so each patient’s outcome is the same whichever arm they were sent to. The observed sequence of outcomes stops being a sample and becomes a fixed list.
Re-run the rule. Not a coin: the same minimisation, fed the same patients with the same recorded factors in the same arrival order, with a fresh random stream.
Count. p = (1 + #{as extreme}) / (1 + B), with the +1 the observed allocation counting as one of its own reference draws.
The middle line is where this field differs from the last one, and the difference is everything. In the adaptive field the rule read the outcomes, so re-running it required feeding it the fixed outcomes and the reference distribution depended on what those outcomes were. Here the rule reads only the covariates, which were recorded before anything was given to anybody and are not quantities the trial is trying to estimate.
So the reference distribution can be generated without looking at an outcome at all. It is a property of the design and the enrolled cohort, and in principle it could be computed the day recruitment closes and before a single measurement is taken.
What the two distributions look like
The observed statistic on this trial is 0.4796. Against the rule’s own reference distribution the p-value is 0.3850; against a shuffled one it is 0.6725.
The distributions themselves say why. The rule’s 5% point is 1.0948. A coin’s is 1.9531 — and 1.9531 is, to two decimal places, the 1.96 a t table would have supplied, which is the whole of the previous essay in one number: shuffling the labels and reading a t table are the same mistake.
The slider is worth watching to the end. At p = 0.5 minimisation is a coin, and the two distributions coincide; the field’s own check requires them to agree there rather than to differ, because requiring a difference that must not exist would be a check that passes for the wrong reason.
What the picture is not
It is worth saying plainly what the bars are, because every other histogram on this site is something else and the difference is the whole construction.
A sampling distribution is what a statistic would do over repetitions of the data: draw new patients, new outcomes, new noise, and see where the number lands. Every distribution in the intervals field, the testing field and the forecasting field is one of those.
These bars are what the statistic would do over repetitions of the allocation, with the data held exactly as it came out. No outcome is redrawn anywhere in the figure. The 120 patients are the 120 patients, with the outcomes they had; what varies is only who got what.
That is why the p-value needs no assumption about how outcomes are distributed. It is not that the assumption has been weakened — it is that the outcomes are not random in this calculation at all, and there is nothing left about them to assume.
The size, counted
Seven hundred trials of 120 patients with no treatment effect, both statistics against the rule’s own reference distribution:
- unadjusted: 5.1% at p = 0.8 and 5.6% at p = 1
- adjusted: 4.1% at p = 0.8 and 5.0% at p = 1
All four are 5% tests within their own noise, against 0.6% and 0.0% for the unadjusted statistic read against a t table on the identical trials.
And they are 5% tests for a reason that has nothing to do with the outcome model being right. The construction conditions on the outcomes that happened, so no assumption about how the covariates enter them is used anywhere. That is the model-free virtue the unadjusted analysis was defended for, delivered without the conservatism.
The price, which is the finding
The exact field’s version of this cost nineteen points of power against a competitor held to the same size, and the essay there had to be rewritten around it: the construction it set out to recommend turned out to be expensive, and what it bought had to be found instead.
Here it is nearly free.
Against a real effect of 0.45 standard deviations on the same design, the adjusted statistic against the re-randomisation distribution gets 64.3% and the same statistic against a t table gets 66.7%. Two and a half points, for a test that needs no model at all.
And the unadjusted statistic is where the striking number is. Against a t table it gets 20.3%. Against the rule’s own reference distribution, on identical data with the identical statistic, it gets 55.0%. Thirty-five points of power recovered by changing nothing but what the number is compared with.
The statistic and the reference distribution are two choices
The four-way table exists because the two choices are independent, and the subject usually discusses them as one. It is worth reading the table as a two-by-two rather than as four methods.
Down one axis: which statistic. The plain difference of means, or the same difference after the balanced factors have been put in the model. That choice decides how much of the prognostic variation is taken out of the residual, which decides the power.
Across the other: which reference distribution. A t table, which is a coin’s, or the set of allocations the rule could have produced. That choice decides whether the test holds its level.
Read that way, the numbers arrange themselves without any further explanation. The two right-hand cells are 5% tests and the two left-hand cells are whatever the mismatch makes them. The two bottom cells are powerful and the two top cells are not. The cell everybody reports — plain statistic, t table — is the only one of the four that is wrong on both axes, and it is the default because each of its two halves is what arrives by default when no decision is made.
The independence also says something about what to do when the model is doubtful. Adjustment buys power and costs an assumption; the reference distribution buys validity and costs neither. So the combination in the bottom-right cell — adjust the statistic and re-randomise it — is the one with no downside available: if the model is right it has the model-based test’s power, and if the model is wrong it is still a 5% test, because its level never depended on the model. The counted numbers say that costs two and a half points against the model-based test when the model happens to be right.
The two repairs are substitutes, not complements
Reading the four cells as a two-by-two invites the next question, which is whether the two effects add. They do not, and the extent to which they fail to is the most useful number the table holds.
Take the unadjusted t-table cell at 20.3% as the corner. Moving across to the right reference distribution is worth +34.7 points. Moving down to the adjusted statistic, still against a t table, is worth +46.4. If the two were separate repairs of separate defects, doing both would land near a hundred, which is not a place power can be. It lands at 64.3%.
So the interaction is about −37 points: each repair recovers most of what the other would have recovered, and doing the second one after the first buys single figures. That is the power side of the sentence the null figure already carries — adjusting fixes it, re-randomising fixes it, doing either is enough — and it is worth having in both places, because a reader who accepts substitutes on the level might still expect complements on the power.
Why they are substitutes, in one ratio
The two distributions in this trial differ by a single factor. The rule’s 5% point is 1.0948 and a coin’s is 1.9531, so the t table is applying a critical value 1.784 times too large for the allocations the trial could actually have produced.
Squared, that says the minimisation removed 68.6% of the variance a coin’s allocation would have given the statistic. And that is the same variance the adjusted statistic removes, by a different route: the factors put into the model are exactly the factors the rule balanced, so taking them out of the residual takes out the variation the allocation was already prevented from carrying.
Two constructions, one quantity. The reference distribution accounts for it exactly, by counting the allocations that were available; the adjustment accounts for it approximately, by estimating the coefficients that carry it from a hundred and twenty rows. The two and a half points between the adjusted exact test and the adjusted t table is the whole of the difference between exactly and approximately, which is why it is small and why it is in that direction.
What the re-randomisation count is worth
The p-value is (1 + k)/(1 + B) at B = 399, so it moves in steps of 0.0025 and a nominal 5% decision turns on whether k is above or below 19.
The count is itself a Monte Carlo estimate, and near the boundary its standard error is √(0.05 × 0.95 / 399) = 0.011 — about a fifth of the level being tested. That is not negligible beside the two and a half points the construction costs, and it is the one part of the price that buying more computation removes: at B = 999, the count the figures drawn against the tail use, the same standard error is 0.0069.
For the reading this essay actually makes, 399 is ample. The observed statistic’s p-value is 0.3850 against the rule and 0.6725 against a coin, and neither is anywhere near a boundary where a hundredth decides anything.
Why it is cheap here and expensive there
The two fields ran the same construction and got opposite answers about its cost, so the reason is worth isolating.
In the adaptive field the rule chased outcomes. That made the null distribution of the statistic depend on the success rate — how imbalanced the allocation gets depends on how often anything succeeds — so a critical value simulated once was a 5% test at one success rate and at no other, moving from 1.668 at a rate of 0.05 to 2.718 at 0.8. The randomisation test avoided that by conditioning on the outcomes that happened, and conditioning is what cost the power.
Here the rule depends on the covariates, which are observed. There is no nuisance parameter to condition away, because there was never one in the reference distribution to begin with. So the conditioning costs almost nothing, and what would have been the cheap alternative — simulate the design under its null once and take the critical value — is available too and gives the same answer.
The separating condition is now stated in both fields and it is the same sentence: calibration works where the null distribution depends only on quantities the protocol fixes. In a covariate-adaptive trial it does. In an outcome-adaptive one it does not, and the exact test is the only thing that repairs it.
The refusal, and it is what every off-the-shelf routine does
The construction rests on one assumption and it is the same one the exact field’s does: the analysis has to be told the rule that produced the allocation.
Every general-purpose permutation routine reshuffles labels. That is the reference distribution of a coin, because a coin is the rule nobody thinks to state, and a minimised trial never drew from it. Fed to a minimised trial the shuffled distribution is too wide, and the test rejects 0.4% where it claims 5%, against 6.6% for the same statistic told the right rule.
The direction is the reverse of the exact field’s, where the wrong reference distribution was too narrow and the test over-rejected. Both come from the same place: the wrong reference distribution is a coin’s, and a coin’s allocations are less balanced than an adaptive rule’s — which spreads the statistic out here and concentrated it there, because there the rule was making the arms unequal in size and here it is making them alike.
So the same mistake, made by the same routine, breaks the test in opposite directions in two adjacent fields. Neither direction is detectable from the output.
What this field ends up saying about the one before it
Four essays, and every one of them came out the other way from the adaptive field’s answer to the same question. Putting the pairs side by side is the fairest summary of what reading covariates rather than outcomes actually changes.
The rule’s effect on the error rate. Response-adaptive randomisation rejects 7.8% of true nulls with no time trend at all, because the allocation is a function of the outcomes it is later compared with. A covariate-adaptive rule rejects 0.6% — too few rather than too many, because the allocation is a function of something the comparison does not contain.
What repairs it. There, blocking by arrival time fixes the confounding and not the adaptation, and the randomisation test is the only thing that restores the level. Here, two independent repairs work and either is sufficient.
What the repair costs. There, nineteen points of power against a fair competitor. Here, two and a half — and thirty-five points recovered if the statistic being repaired is the unadjusted one.
And what the critical value depends on. There, a success rate nobody has, moving it from 1.668 to 2.718. Here, nothing but the covariates and the rule, both of which are written down before the trial starts.
The single sentence that generates all four differences is that a covariate is fixed before the experiment and an outcome is produced by it. Everything else follows, and it follows in the direction that makes this family of rules safe to use and the other family difficult.
What is being claimed here, and what is not
This essay claims the randomisation test for a covariate-adaptive allocation: its exactness under two statistics and two settings of the rule, its power against the model-based alternative, and the refusal that shows what a shuffled reference distribution does.
What stays out: the computational cost, which is B re-randomisations of an n-patient cohort and is the same count the exact field reports for its own construction — a count of operations rather than a duration, on the fleet’s usual boundary. Confidence intervals by inverting the test, which is available and is a different amount of work. And rules with more than two arms, where the margin score generalises in more than one way.
The most useful thing this field ends with is not a method, though. It is a reason to record one. Every number in this essay depends on the rule being stated — which arm the score preferred, what p was, how ties were broken, which factors were in the margin. A trial that reports “minimisation” and nothing else has not supplied enough for its own analysis to be reconstructed, and the previous essay’s finding says the default reconstruction is a coin’s and is badly wrong. The rule is part of the data, and here that is not a slogan: it is an input the test will not run without.
The checks
Three claims are gated in this field’s library.
The exact route holds its level under both statistics at p = 1 and p = 0.8, with the unadjusted t-table analysis required to be more than a point below the same statistic’s exact rate — which is what makes the difference attributable to the reference distribution rather than to the statistic.
The exactness is cheap, asserted as three things at once: the adjusted exact test must reach better than 40% power at the effect measured, must be within five points of the model-based test it replaces, and must beat the unadjusted exact test by more than three points. The last of those is the half nobody separates — the choice of statistic and the choice of reference distribution are independent, and a check that only compared the two reference distributions would suggest the statistic did not matter.
And a shuffled reference distribution is rejected, with the right rule required to hold its level and the wrong one required to fall more than a point below it on the identical trials. If the two ever agreed the check throws explicitly, because a refusal that has stopped refusing proves nothing, which is a habit this collection has had repeated cause to write down.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A statistic that is exact twice — both name exact test, nuisance parameter, permutation test, randomisation test, reference distribution, sharp null, statistical power
- The null the exactness is for — both name exact test, nuisance parameter, permutation test, randomisation test, reference distribution, sharp null
- Walking the admissible set — both name exact test, permutation test, randomisation test, reference distribution, sharp null
- A probe nobody chose — both name randomisation test, reference distribution, sharp null, statistical power
- The analysis and the shape — both name covariate adjustment, exact test, randomisation test, reference distribution
- The plus one and the round number — both name permutation test, randomisation test, reference distribution, sharp null
Named objects
A flat tag is an object no other essay names yet.
Covariate-adaptive randomisationCovariate adjustmentExact testMinimisationNuisance parameterPermutation testRandomisation testReference distributionSharp nullStatistical power