Concept

Covariate-adaptive randomisation — where it appears

An allocation rule that decides each arrival's arm from the baseline covariates already enrolled, so as to keep the arms alike on them. Because the assignment then depends on the data, the analysis has to condition on the rule that produced it, which is a randomisation test rather than a t test.

Named by 12 essays across 3 fields — each of them below, with the objects they name alongside it.

What a median split can see. A standard normal covariate with its median marked, and the mean of each category as a vertical rule: -0.7979, 0.7979. A rule that balances the categories is balancing those numbers and nothing else, so the part of the covariate it can act on is the variance between them — 0.6366 of the total, which at two categories is exactly 2/π because the two half-normal means are ±√(2/π). The rest, 0.3634, is variation inside the categories that the rule cannot see and does not touch: the assignment within a category is still a coin. Everything the next figure measures is a consequence of this one, and it is available before any unit has arrived.

A covariate with no levels

Every balancing rule on this site reads a level. Age and blood pressure have none, so somebody cuts them into categories — and a median split can see exactly 2/π of a normal covariate, whatever the rule does with the halves.

continuous · Assignment
What each rule leaves behind, at 120 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.

Balancing what is known in advance

Four allocation rules, three definitions of balance, and no rule that holds more than one of them. Minimisation keeps the worst factor margin near three patients whether the trial has forty or six hundred and forty — and lets the imbalance in the cross-classified cells climb to 86% of a coin's, because the cells are not what it is watching.

covadapt · Assignment
Where one rule becomes three. Every arrival in 200 simulated trials is put to all three scores, and the picture is how often they would send that patient to different arms. The range and the pairwise sum are the same rule at two arms and at three — for sorted counts the pairwise sum is twice the range, so the arm that minimises one minimises the other — and they part company at four, where the pairwise sum is 3(d − a) + (c − b) and the range still sees only d − a. The variance disagrees with both from two arms onwards, on 5.4% of arrivals at two and 27.0% at five, because the scores are summed over 3 factors and a sum of squares does not order the candidates the way a sum of absolute values does. All three are called minimisation.

Three arms and three scores

Minimisation balances a trial by keeping the arms' counts even inside every prognostic factor. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms.

multiarm · Assignment
A trial designed 2:1:1, and what two scores deliver. 500 trials of 180 patients, three arms, a target of 2:1:1. The shaded bars are a minimisation score that divides each arm's count by the share that arm is supposed to receive before measuring the spread; it delivers 49.9% : 25.1% : 25.1%. The others are the same rule with the counts left raw, which delivers 33.4% : 33.3% : 33.3% — the balance it enforces inside every factor level is equality, and equality is what it gets. The marks are the shares that were asked for.

Balancing towards unequal targets

A three-arm trial allocating two to one to one is the ordinary case, and a balancing rule built from raw counts does not know it. It balances the arms towards equality inside every factor level, delivers a third to each arm, and reports that it minimised imbalance.

multiarm · Allocation
What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 120 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 49.7% at p = 0.5, which is a coin and cannot be beaten, and 87.6% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.5 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.75 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).

The rule that can be guessed

A balancing rule improves as it becomes more deterministic, and a deterministic rule can be worked out in advance from information the person enrolling the patient already has. At full determinism 87.6% of assignments are guessable, and an investigator who acts on the guess produces a treatment effect of three quarters of a standard deviation where the truth is zero.

covadapt · Assignment
Two rates, not a factor. The standard deviation of the covariate imbalance under three rules, at five trial sizes, 260 trials each, on log axes. The upper line is a coin: its slope is -0.489, against a closed form of exactly −½. The middle line is minimisation on a median split; its slope is -0.519 — the same rate — because inside a category the assignment is still a coin, and what it buys is the constant, 0.654 of a coin's at n = 200. The lower line is the rule that reads x and maximises the information about the treatment effect: slope -0.987, nearly twice as steep. Its advantage is therefore not a number that can be quoted — it is 0.258 of a coin's at n = 50 and 0.065 at n = 800, and it keeps going.

The rule that reads the number

Stop categorising and let the rule read the covariate itself. What it should minimise is not an invented distance but the variance of the effect being estimated — and what comes back is not a better constant but a different rate.

continuous · Assignment
Four analyses of the same 3-arm trials, under a true null. 250 trials of 150 patients, 3 arms, minimisation with p = 0.85, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Only the first cell is wrong, and it is wrong in the direction that costs power rather than the one that manufactures findings: 0.0% where 5% is claimed. Either repair works — adjusting for what the rule balanced, or asking the rule what it would have done.

The analysis after three arms

An unadjusted analysis after a two-arm balancing rule rejects 0.6% of true nulls where it claims 5%. With three arms and a deterministic rule it rejects none at all — and the repair is the same repair, which is a sentence and a column in the model.

multiarm · Assignment
Four analyses of the same trials, with no treatment effect at all. 700 trials of 120 patients allocated by minimisation at p = 0.8, with the prognostic factors carrying a real effect on the outcome and no treatment effect — every rejection below is a false one. Two statistics, the plain difference and the same after adjusting for the balanced factors, each read against two reference distributions: a t table, and the set of allocations the rule could have produced from these covariates. The unadjusted comparison rejects 0.6% where it claims 5% — conservative, which is a loss of power rather than an error, and nothing on the output says so. Adjusting puts it back at 5.4%. Both re-randomised versions are at their nominal level by construction, whatever statistic goes into them.

The analysis has to know the rule

A trial balanced by minimisation and analysed by comparing the two arms' means rejects a true null 0.6% of the time where it claims 5%, and at full determinism 0.0%. That is not an error anybody complains about — it is a test that has stopped working, paid for by a balance the analysis then refused to use.

covadapt · Assignment
Three analyses of the same trials, none of them wrong about the data. 320 trials at n = 60 with no treatment effect at all, so every rejection counted is a false one, and a covariate that drives the outcome with coefficient 1. The unadjusted comparison is at 5.94% after a coin — its level — and at 0.00% after the rule that reads the covariate: the design removed the imbalance and the analysis is still pricing it. Adjusting for the covariate gives 4.06%, and the rule's own reference distribution — hold the outcomes, re-run the rule 199 times, count — gives 3.13% against the 4.5% that 199 draws can deliver. The last of the three has to be told the assignment rule and nothing else, which is the one thing the experimenter certainly knows.

What the balanced trial is worth

A rule that reads the covariate removes three quarters of the imbalance. An analysis that does not know it happened prices the imbalance anyway, rejects one true null in two hundred instead of one in twenty, and finds a real effect less often than a coin-tossed trial does.

continuous · Randomisation
What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 12.7% of a coin's imbalance in it; balancing eight leaves 23.2% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.

Balancing more than one number

The criterion generalises to several covariates without a word changing, which makes the question what it is worth rather than whether it can be done. Each one added takes a share of the assignment's freedom, and the imbalance left in every one of them rises.

continuous · Blocking
What a guesser gets, and what a guesser gets for nothing. 600 trials of 150 patients under a fully deterministic rule, with an investigator who knows the rule, the factors and every assignment so far. The guess rate falls with the number of arms — 87.8% at 2, 86.2% at 3, 81.0% at 4 — which reads like a trial getting safer and is not: what a guesser can trade on is the excess over the 50%, 33%, 25% they would get by naming an arm at random, and that goes the other way, from 1.76× chance at two arms to 3.24× at 4. The gap a guesser manufactures between the best and worst arm under a true null is 0.759, 0.733, 0.711 standard deviations — nearly unchanged.

Guessing one arm in three

A balancing rule is guessable because it is balancing. With three arms the next assignment is worked out less often than with two — and by more, relative to what a guesser gets for nothing, and the damage they can do is almost unchanged.

multiarm · Assignment
The allocations this trial could have made, and the ones it could not. One 120-patient trial allocated by minimisation at p = 1, re-randomised 399 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the rule again over the same patients in the same order with the same recorded factors, so what is drawn is the set of experiments that could have happened. The bars are that set; the outline is what shuffling the labels gives, which is the reference distribution of a coin and is what every off-the-shelf permutation routine assumes. The coin's is wider — its 5% point is 1.95 against the rule's 1.09 — because a coin's allocations are less balanced and a less balanced allocation gives a larger statistic. Reading this trial against it makes the test conservative rather than anti-conservative, which is the opposite error from the outcome-adaptive case and for the same structural reason.

The reference the covariates supply

Hold the outcomes fixed, re-run the rule that assigned them, count. The same construction cost nineteen points of power in the adaptive field, because its rule chased outcomes and its critical value depended on a rate nobody has. Here the rule reads only what was recorded before anything happened, and the same unadjusted statistic goes from 20.3% power to 55.0% by being read against the right distribution.

covadapt · Assignment

Named alongside it

The objects these essays reach for when they reach for this one.

MinimisationMonte CarloRandomisationExperimental designPermuted blocksStudy designContinuous covariateCovariate adjustmentCovariate balanceError rateReference distributionStatistical power

All concepts