Balancing on what was recorded first

The reference the covariates supply

Hold the outcomes fixed, re-run the rule that assigned them, count. The same construction cost nineteen points of power in the adaptive field, because its rule chased outcomes and its critical value depended on a rate nobody has. Here the rule reads only what was recorded before anything happened, and the same unadjusted statistic goes from 20.3% power to 55.0% by being read against the right distribution.

Worth reading first: Randomisation is not balance · The experiments that could have happened.

The previous essay ended with three answers and none of them satisfactory. An unadjusted comparison that assumes nothing and rejects 0.6% of true nulls where it claims 5%. An adjusted comparison that is correct and assumes the outcome model. And two rows of the table that had not been explained.

This essay explains them, and the construction is not new — it is the one the exact field built for a rule that reads outcomes, applied to a rule that reads covariates. What is new is the price, and the price is the finding.

The allocations this trial could have made, and the ones it could not. One 120-patient trial allocated by minimisation at p = 1, re-randomised 399 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the rule again over the same patients in the same order with the same recorded factors, so what is drawn is the set of experiments that could have happened. The bars are that set; the outline is what shuffling the labels gives, which is the reference distribution of a coin and is what every off-the-shelf permutation routine assumes. The coin's is wider — its 5% point is 1.95 against the rule's 1.09 — because a coin's allocations are less balanced and a less balanced allocation gives a larger statistic. Reading this trial against it makes the test conservative rather than anti-conservative, which is the opposite error from the outcome-adaptive case and for the same structural reason.
Fig. 1 One 120-patient trial, allocated by deterministic minimisation, re-randomised 399 times. The bars are the allocations the rule could have produced from these exact patients. The outline is what shuffling the labels gives, which is a coin’s distribution and is what the trial never drew from.

Three lines, and the second one is different here

Fix the outcomes. Under the sharp null the treatment changed nothing for anybody, so each patient’s outcome is the same whichever arm they were sent to. The observed sequence of outcomes stops being a sample and becomes a fixed list.

Re-run the rule. Not a coin: the same minimisation, fed the same patients with the same recorded factors in the same arrival order, with a fresh random stream.

Count. p = (1 + #{as extreme}) / (1 + B), with the +1 the observed allocation counting as one of its own reference draws.

The middle line is where this field differs from the last one, and the difference is everything. In the adaptive field the rule read the outcomes, so re-running it required feeding it the fixed outcomes and the reference distribution depended on what those outcomes were. Here the rule reads only the covariates, which were recorded before anything was given to anybody and are not quantities the trial is trying to estimate.

So the reference distribution can be generated without looking at an outcome at all. It is a property of the design and the enrolled cohort, and in principle it could be computed the day recruitment closes and before a single measurement is taken.

What the two distributions look like

The observed statistic on this trial is 0.4796. Against the rule’s own reference distribution the p-value is 0.3850; against a shuffled one it is 0.6725.

The distributions themselves say why. The rule’s 5% point is 1.0948. A coin’s is 1.9531 — and 1.9531 is, to two decimal places, the 1.96 a t table would have supplied, which is the whole of the previous essay in one number: shuffling the labels and reading a t table are the same mistake.

The allocations this trial could have made, and the ones it could notOne 120-patient trial allocated by minimisation at p = 0.8, re-randomised 399 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the rule again over the same patients in the same order with the same recorded factors, so what is drawn is the set of experiments that could have happened. The bars are that set; the outline is what shuffling the labels gives, which is the reference distribution of a coin and is what every off-the-shelf permutation routine assumes. The coin's is wider — its 5% point is 1.95 against the rule's 1.28 — because a coin's allocations are less balanced and a less balanced allocation gives a larger statistic. Reading this trial against it makes the test conservative rather than anti-conservative, which is the opposite error from the outcome-adaptive case and for the same structural reason.020400123the statistic the trial would have producedre-randomisationsobserved 0.39bars: the rule's own allocations · outline: what shuffling the labels gives399 re-randomisations of one 120-patient trial, outcomes held fixedp = 0.6125 against the rule, 0.7150 against a coin
Fig. 2 The same trial under a rule that keeps a coin one time in five. Its reference distribution is wider — 5% point 1.28 rather than 1.09 — because a less deterministic rule produces less balanced allocations. Drag p to watch the bars spread out towards the outline as the rule becomes a coin.

The slider is worth watching to the end. At p = 0.5 minimisation is a coin, and the two distributions coincide; the field’s own check requires them to agree there rather than to differ, because requiring a difference that must not exist would be a check that passes for the wrong reason.

What the picture is not

It is worth saying plainly what the bars are, because every other histogram on this site is something else and the difference is the whole construction.

A sampling distribution is what a statistic would do over repetitions of the data: draw new patients, new outcomes, new noise, and see where the number lands. Every distribution in the intervals field, the testing field and the forecasting field is one of those.

These bars are what the statistic would do over repetitions of the allocation, with the data held exactly as it came out. No outcome is redrawn anywhere in the figure. The 120 patients are the 120 patients, with the outcomes they had; what varies is only who got what.

That is why the p-value needs no assumption about how outcomes are distributed. It is not that the assumption has been weakened — it is that the outcomes are not random in this calculation at all, and there is nothing left about them to assume.

The reference distribution a fair coin gives, and the curve it matches. One 200-patient trial allocated by a fair coin, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the same rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 0.663, 507 of the 999 re-randomisations reach it, and the p-value is (1 + 507)/(1 + 999) = 0.5080. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 1.997.
Fig. 3 The exact field’s version of the same picture, on a trial allocated by a fair coin, where the re-randomisation distribution and the normal curve agree — which is what makes a t table right in the ordinary case. Everything in this field is what happens when the rule producing the allocation is not a coin and the table has not been told.

The size, counted

Seven hundred trials of 120 patients with no treatment effect, both statistics against the rule’s own reference distribution:

  • unadjusted: 5.1% at p = 0.8 and 5.6% at p = 1
  • adjusted: 4.1% at p = 0.8 and 5.0% at p = 1

All four are 5% tests within their own noise, against 0.6% and 0.0% for the unadjusted statistic read against a t table on the identical trials.

And they are 5% tests for a reason that has nothing to do with the outcome model being right. The construction conditions on the outcomes that happened, so no assumption about how the covariates enter them is used anywhere. That is the model-free virtue the unadjusted analysis was defended for, delivered without the conservatism.

The price, which is the finding

The exact field’s version of this cost nineteen points of power against a competitor held to the same size, and the essay there had to be rewritten around it: the construction it set out to recommend turned out to be expensive, and what it bought had to be found instead.

Here it is nearly free.

Against a real effect of 0.45 standard deviations on the same design, the adjusted statistic against the re-randomisation distribution gets 64.3% and the same statistic against a t table gets 66.7%. Two and a half points, for a test that needs no model at all.

And the unadjusted statistic is where the striking number is. Against a t table it gets 20.3%. Against the rule’s own reference distribution, on identical data with the identical statistic, it gets 55.0%. Thirty-five points of power recovered by changing nothing but what the number is compared with.

The same four, against a real effect of 0.45. 700 trials of 120 patients allocated by minimisation at p = 0.8, with the prognostic factors carrying a real effect on the outcome and a treatment effect of 0.45. Two statistics, the plain difference and the same after adjusting for the balanced factors, each read against two reference distributions: a t table, and the set of allocations the rule could have produced from these covariates. The re-randomised rows give up almost nothing: 64.3% against 66.7%. The exactness is nearly free here, where the same construction cost nineteen points of power against a rule that read outcomes.
Fig. 4 The power side of the four-way table. The two re-randomised rows are in the fifties and sixties; the unadjusted t-table row is at twenty. Nothing about the data differs between them.
The allocations this trial could have made, and the ones it could not. One 240-patient trial allocated by minimisation at p = 1, re-randomised 399 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the rule again over the same patients in the same order with the same recorded factors, so what is drawn is the set of experiments that could have happened. The bars are that set; the outline is what shuffling the labels gives, which is the reference distribution of a coin and is what every off-the-shelf permutation routine assumes. The coin's is wider — its 5% point is 2.01 against the rule's 1.18 — because a coin's allocations are less balanced and a less balanced allocation gives a larger statistic. Reading this trial against it makes the test conservative rather than anti-conservative, which is the opposite error from the outcome-adaptive case and for the same structural reason.
Fig. 5 The same construction on a trial twice the size. The rule’s distribution has not moved towards the coin’s — a balancing rule holds its margins flat as the trial grows — so the mismatch a t table produces does not wash out with more patients.

The statistic and the reference distribution are two choices

The four-way table exists because the two choices are independent, and the subject usually discusses them as one. It is worth reading the table as a two-by-two rather than as four methods.

Down one axis: which statistic. The plain difference of means, or the same difference after the balanced factors have been put in the model. That choice decides how much of the prognostic variation is taken out of the residual, which decides the power.

Across the other: which reference distribution. A t table, which is a coin’s, or the set of allocations the rule could have produced. That choice decides whether the test holds its level.

Read that way, the numbers arrange themselves without any further explanation. The two right-hand cells are 5% tests and the two left-hand cells are whatever the mismatch makes them. The two bottom cells are powerful and the two top cells are not. The cell everybody reports — plain statistic, t table — is the only one of the four that is wrong on both axes, and it is the default because each of its two halves is what arrives by default when no decision is made.

Four analyses of the same trials, with no treatment effect at all. 700 trials of 120 patients allocated by minimisation at p = 1, with the prognostic factors carrying a real effect on the outcome and no treatment effect — every rejection below is a false one. Two statistics, the plain difference and the same after adjusting for the balanced factors, each read against two reference distributions: a t table, and the set of allocations the rule could have produced from these covariates. The unadjusted comparison rejects 0.0% where it claims 5% — conservative, which is a loss of power rather than an error, and nothing on the output says so. Adjusting puts it back at 5.7%. Both re-randomised versions are at their nominal level by construction, whatever statistic goes into them.
Fig. 6 The null side at full determinism, where the two axes separate most cleanly: the unadjusted t-table cell rejects nothing at all in seven hundred trials, and the other three are within a standard error of 5%. Adjusting fixes it. Re-randomising fixes it. Doing either is enough.

The independence also says something about what to do when the model is doubtful. Adjustment buys power and costs an assumption; the reference distribution buys validity and costs neither. So the combination in the bottom-right cell — adjust the statistic and re-randomise it — is the one with no downside available: if the model is right it has the model-based test’s power, and if the model is wrong it is still a 5% test, because its level never depended on the model. The counted numbers say that costs two and a half points against the model-based test when the model happens to be right.

The two repairs are substitutes, not complements

Reading the four cells as a two-by-two invites the next question, which is whether the two effects add. They do not, and the extent to which they fail to is the most useful number the table holds.

Take the unadjusted t-table cell at 20.3% as the corner. Moving across to the right reference distribution is worth +34.7 points. Moving down to the adjusted statistic, still against a t table, is worth +46.4. If the two were separate repairs of separate defects, doing both would land near a hundred, which is not a place power can be. It lands at 64.3%.

So the interaction is about −37 points: each repair recovers most of what the other would have recovered, and doing the second one after the first buys single figures. That is the power side of the sentence the null figure already carries — adjusting fixes it, re-randomising fixes it, doing either is enough — and it is worth having in both places, because a reader who accepts substitutes on the level might still expect complements on the power.

Why they are substitutes, in one ratio

The two distributions in this trial differ by a single factor. The rule’s 5% point is 1.0948 and a coin’s is 1.9531, so the t table is applying a critical value 1.784 times too large for the allocations the trial could actually have produced.

Squared, that says the minimisation removed 68.6% of the variance a coin’s allocation would have given the statistic. And that is the same variance the adjusted statistic removes, by a different route: the factors put into the model are exactly the factors the rule balanced, so taking them out of the residual takes out the variation the allocation was already prevented from carrying.

Two constructions, one quantity. The reference distribution accounts for it exactly, by counting the allocations that were available; the adjustment accounts for it approximately, by estimating the coefficients that carry it from a hundred and twenty rows. The two and a half points between the adjusted exact test and the adjusted t table is the whole of the difference between exactly and approximately, which is why it is small and why it is in that direction.

What the re-randomisation count is worth

The p-value is (1 + k)/(1 + B) at B = 399, so it moves in steps of 0.0025 and a nominal 5% decision turns on whether k is above or below 19.

The count is itself a Monte Carlo estimate, and near the boundary its standard error is √(0.05 × 0.95 / 399) = 0.011 — about a fifth of the level being tested. That is not negligible beside the two and a half points the construction costs, and it is the one part of the price that buying more computation removes: at B = 999, the count the figures drawn against the tail use, the same standard error is 0.0069.

For the reading this essay actually makes, 399 is ample. The observed statistic’s p-value is 0.3850 against the rule and 0.6725 against a coin, and neither is anywhere near a boundary where a hundredth decides anything.

Why it is cheap here and expensive there

The two fields ran the same construction and got opposite answers about its cost, so the reason is worth isolating.

In the adaptive field the rule chased outcomes. That made the null distribution of the statistic depend on the success rate — how imbalanced the allocation gets depends on how often anything succeeds — so a critical value simulated once was a 5% test at one success rate and at no other, moving from 1.668 at a rate of 0.05 to 2.718 at 0.8. The randomisation test avoided that by conditioning on the outcomes that happened, and conditioning is what cost the power.

Here the rule depends on the covariates, which are observed. There is no nuisance parameter to condition away, because there was never one in the reference distribution to begin with. So the conditioning costs almost nothing, and what would have been the cheap alternative — simulate the design under its null once and take the critical value — is available too and gives the same answer.

The separating condition is now stated in both fields and it is the same sentence: calibration works where the null distribution depends only on quantities the protocol fixes. In a covariate-adaptive trial it does. In an outcome-adaptive one it does not, and the exact test is the only thing that repairs it.

What each rule leaves behind, at 120 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.
Fig. 7 The balance table this field opened with, which is the same information the reference distribution carries. A rule that holds a margin flat produces allocations that are alike, and allocations that are alike produce statistics that are small — which is the whole mechanism in one sentence.

The refusal, and it is what every off-the-shelf routine does

The construction rests on one assumption and it is the same one the exact field’s does: the analysis has to be told the rule that produced the allocation.

Every general-purpose permutation routine reshuffles labels. That is the reference distribution of a coin, because a coin is the rule nobody thinks to state, and a minimised trial never drew from it. Fed to a minimised trial the shuffled distribution is too wide, and the test rejects 0.4% where it claims 5%, against 6.6% for the same statistic told the right rule.

The direction is the reverse of the exact field’s, where the wrong reference distribution was too narrow and the test over-rejected. Both come from the same place: the wrong reference distribution is a coin’s, and a coin’s allocations are less balanced than an adaptive rule’s — which spreads the statistic out here and concentrated it there, because there the rule was making the arms unequal in size and here it is making them alike.

So the same mistake, made by the same routine, breaks the test in opposite directions in two adjacent fields. Neither direction is detectable from the output.

The allocations this trial could have made, and the ones it could not. One 120-patient trial allocated by minimisation at p = 0.7, re-randomised 399 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the rule again over the same patients in the same order with the same recorded factors, so what is drawn is the set of experiments that could have happened. The bars are that set; the outline is what shuffling the labels gives, which is the reference distribution of a coin and is what every off-the-shelf permutation routine assumes. The coin's is wider — its 5% point is 1.95 against the rule's 1.45 — because a coin's allocations are less balanced and a less balanced allocation gives a larger statistic. Reading this trial against it makes the test conservative rather than anti-conservative, which is the opposite error from the outcome-adaptive case and for the same structural reason.
Fig. 8 And at p = 0.7, where the rule follows its preference only two times in three. The reference distribution has widened towards the outline without reaching it, which is the slider’s whole range in one static picture.

What this field ends up saying about the one before it

Four essays, and every one of them came out the other way from the adaptive field’s answer to the same question. Putting the pairs side by side is the fairest summary of what reading covariates rather than outcomes actually changes.

The rule’s effect on the error rate. Response-adaptive randomisation rejects 7.8% of true nulls with no time trend at all, because the allocation is a function of the outcomes it is later compared with. A covariate-adaptive rule rejects 0.6% — too few rather than too many, because the allocation is a function of something the comparison does not contain.

What repairs it. There, blocking by arrival time fixes the confounding and not the adaptation, and the randomisation test is the only thing that restores the level. Here, two independent repairs work and either is sufficient.

What the repair costs. There, nineteen points of power against a fair competitor. Here, two and a half — and thirty-five points recovered if the statistic being repaired is the unadjusted one.

And what the critical value depends on. There, a success rate nobody has, moving it from 1.668 to 2.718. Here, nothing but the covariates and the rule, both of which are written down before the trial starts.

The single sentence that generates all four differences is that a covariate is fixed before the experiment and an outcome is produced by it. Everything else follows, and it follows in the direction that makes this family of rules safe to use and the other family difficult.

What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 120 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 49.7% at p = 0.5, which is a coin and cannot be beaten, and 87.6% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.5 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.75 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).
Fig. 9 The one thing that does not come out the other way. Predictability is the price this family pays, it has no analogue in the outcome-adaptive case — an outcome-adaptive rule is unguessable because outcomes are not yet known — and it is the reason the coin kept in reserve is a real design decision rather than a formality.

What is being claimed here, and what is not

This essay claims the randomisation test for a covariate-adaptive allocation: its exactness under two statistics and two settings of the rule, its power against the model-based alternative, and the refusal that shows what a shuffled reference distribution does.

What stays out: the computational cost, which is B re-randomisations of an n-patient cohort and is the same count the exact field reports for its own construction — a count of operations rather than a duration, on the fleet’s usual boundary. Confidence intervals by inverting the test, which is available and is a different amount of work. And rules with more than two arms, where the margin score generalises in more than one way.

The most useful thing this field ends with is not a method, though. It is a reason to record one. Every number in this essay depends on the rule being stated — which arm the score preferred, what p was, how ties were broken, which factors were in the margin. A trial that reports “minimisation” and nothing else has not supplied enough for its own analysis to be reconstructed, and the previous essay’s finding says the default reconstruction is a coin’s and is badly wrong. The rule is part of the data, and here that is not a slogan: it is an input the test will not run without.

The checks

Three claims are gated in this field’s library.

The exact route holds its level under both statistics at p = 1 and p = 0.8, with the unadjusted t-table analysis required to be more than a point below the same statistic’s exact rate — which is what makes the difference attributable to the reference distribution rather than to the statistic.

The exactness is cheap, asserted as three things at once: the adjusted exact test must reach better than 40% power at the effect measured, must be within five points of the model-based test it replaces, and must beat the unadjusted exact test by more than three points. The last of those is the half nobody separates — the choice of statistic and the choice of reference distribution are independent, and a check that only compared the two reference distributions would suggest the statistic did not matter.

And a shuffled reference distribution is rejected, with the right rule required to hold its level and the wrong one required to fall more than a point below it on the identical trials. If the two ever agreed the check throws explicitly, because a refusal that has stopped refusing proves nothing, which is a habit this collection has had repeated cause to write down.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Covariate-adaptive randomisationCovariate adjustmentExact testMinimisationNuisance parameterPermutation testRandomisation testReference distributionSharp nullStatistical power