Three arms and three scores
Worth reading first: Balancing what is known in advance · Randomisation is not balance.
The covariate-adaptive field measures four allocation rules and finds that each of them holds one kind of imbalance flat while letting the others grow exactly as a coin’s do. The rule at the centre of it is minimisation: as each patient arrives, look at the levels of the prognostic factors they belong to, work out how uneven the arms already are inside those levels, and send the patient wherever makes that unevenness smallest — with probability p, so that the rule is not a function anybody can invert.
Every one of those trials has two arms. With two arms the unevenness inside a factor level is |n₁ − n₂|, and there is nothing to choose. The literature’s original proposal offered three different measures of it and never had to say which, because for two arms they are the same number up to a constant.
They are not the same number for three.
Three ways to say how uneven K counts are
Take a factor level in which the arms currently hold (5, 3, 2) patients. How unbalanced is that?
The range — largest minus smallest — says 3. The variance of the counts about their mean says 1.56. The sum of the pairwise gaps — |5−3| + |5−2| + |3−2| — says 6. Each is a defensible summary and each is what somebody’s minimisation implements, and the rule they define is whichever arm makes their own total smallest, summed across the factors the patient belongs to.
For two arms these collapse. With counts (a, b), the range is |a − b|, the pairwise sum is also |a − b|, and the variance is (a − b)²/2. The first two are identical and the third is a monotone transformation of them level by level.
For three, the first two still collapse — sorted counts a ≤ b ≤ c give a pairwise sum of (c − a) + (c − b) + (b − a) = 2(c − a), exactly twice the range — so those two remain one rule. For four, the pairwise sum is 3(d − a) + (c − b), which is not a function of the range at all, and they part company.
The disagreement rates are 5.4% of arrivals at two arms, 14.4% at three, 22.4% at four and 30.8% at five. Nothing in a trial protocol distinguishes the three, and by four arms nearly a quarter of patients would have been allocated differently under a description that reads identically.
The first thing to ask of a rate like that is whether it is a small-trial artefact, and it is not. Run the same three-arm comparison on twice as many patients — three hundred arrivals rather than a hundred and fifty — and the disagreement rate is 15.0%, against 14.4% before. It did not fall, and there is no sample size at which it would. A disagreement here is not two rules failing to converge on an answer; it is two rules that were never asking the same question, and asking each of them about more patients gives more occasions on which they differ rather than fewer. That is what makes the choice between them something to declare in the protocol rather than something a larger trial outgrows.
What a disagreement rate is counted on
The number above needs one sentence about how it was measured, because there is a much larger and much less meaningful number available for the same words.
Running three separate trials under three scores and comparing the allocations would measure how far three histories diverge, and two histories that diverge once diverge for ever afterwards: the counts differ, so every subsequent decision is being taken in a different trial. That number would be large and would say nothing about the rules.
What is counted here is a decision rather than a history. One trial runs under one named score, and at every arrival all three scores are asked which arm they would prefer given the counts that trial has actually reached. A disagreement is one arrival on which the preferred sets differ, and two scores that are both indifferent between the same two arms have not disagreed about anything.
Why the variance disagrees even at two arms
The two-arm disagreement is the surprising one, since within a single factor level the variance is a monotone function of the range. The answer is that the rule does not act on a single level: it sums its score across the factors the patient belongs to, and a sum of squares does not order candidates the way a sum of absolute values does.
Suppose an arriving patient’s three factor levels currently stand at gaps of (2, 0, 0) under one candidate arm and (1, 1, 1) under another. The summed range says 2 against 3 and prefers the first. The summed variance says 2 against 1.5 and prefers the second. A rule that squares is willing to accept several small imbalances to avoid one large one, and a rule that does not, is not.
That is a statement about what the rule is for, not a numerical detail, and it is the reason the choice cannot be settled by measuring balance without saying which balance.
The score named after the range is not the best at controlling the range
The comparison in that picture is worth reading slowly, because it contains a result that no description of the rules would suggest.
The range score leaves an average range of 0.9658 across factor levels. The pairwise-sum score leaves 0.9658 — the identical number, because at three arms they are the same rule and, given the same coins, they produce the same trial patient for patient. The variance score leaves 0.9322.
The variance score controls the range better than the score named after it. It also leaves a smaller worst margin — 1.724 against 1.950 — and a smaller imbalance in prognosis between the arms, 0.0477 against 0.0513 standard deviations.
The obvious explanation is that a range is a coarse function, that it often cannot separate two candidate arms, and that the rule then settles the tie with a coin. That explanation is testable and it is false.
What separates them is the aggregation. A sum of ranges is indifferent between one large imbalance and several small ones; a sum of squares is not, and the worst margin — the quantity a reader of a baseline table actually looks at — is what a concentrated imbalance produces. The ratio of worst margin to average margin is 2.019 for the range score and 1.849 for the variance score, which is the same fact in the form that shows the mechanism.
The two-arm case was not a special case
It is worth being clear about what the two-arm field measured, since this essay is partly a correction to how its result should be read.
That field ran minimisation with the range score — |n₁ − n₂| summed across factors — and every number it reports is that rule’s. Nothing in it is wrong. What it could not see is that the rule it was measuring is one of a family, because at two arms two of the three members coincide exactly and the third differs on only 5.4% of arrivals.
A five percent disagreement produces trials that are almost the same, and every conclusion of that field survives: each rule holds one imbalance flat and is a coin on the rest, minimisation’s worst factor margin is 3.00 at forty patients against a coin’s 6.32, and the cells are left where a coin leaves them. What changes at three arms is not the conclusion but the specificity: the same sentence now describes three procedures instead of one.
What none of them do
Every minimising rule holds the margins — one factor level at a time — and the cross-classification is not a margin.
The worst cell of the twenty-four-cell cross-classification is 4.438 under the range score and 5.600 under a uniform draw — a rule that balances nothing at all. Blocks inside every cell hold the cells at 1.000 by construction and let the margins go to 4.108, four times the minimising rules’ worst margin.
A rule chosen for balance promises nothing about an interaction, which is the two-arm field’s finding re-measured where there are half as many patients per cell per arm to make good on it. If the treatment effect differs between men in one centre and women in another, minimisation has done nothing to make that comparison fair, and it has not claimed to.
The prognosis, which is what the balance was for
Balance on the factors is a means. What a trial is trying to achieve is that the arms are comparable in the thing the factors predict, and that quantity is measurable directly here because the simulation knows each patient’s prognosis.
Under a uniform draw the arms differ in average prognosis by 0.2398 standard deviations. Under minimisation on the range they differ by 0.0513, and on the variance by 0.0477 — a factor of about five better than a coin, from a rule that never saw the prognosis and only ever saw the three factors it is made of.
That factor of five is the whole case for these rules, and it is worth putting beside the differences between them: the gap between any minimising score and no rule at all is twenty times the gap between the scores.
What to choose, and how to say what was chosen
The measurements support a short answer and a shorter piece of advice.
The variance score is the better rule on every measure of margin balance here, at every number of arms measured, and by margins that are small — three to twelve percent — but consistent. It is also the one whose behaviour is easiest to describe: it treats imbalance as a squared quantity and therefore refuses to concentrate it.
The choice has to be stated. “Patients were allocated by minimisation” describes three different rules that would have allocated up to a quarter of the patients differently, and no reader can reconstruct which was run. Naming the score, the factors, the weights and the probability p is four short clauses and makes the allocation reproducible.
The differences are smaller than the difference from randomising. A uniform draw leaves an average range of 6.7789 where every minimising rule leaves under one. Arguing about which score to use while not using any is the error worth avoiding.
The disagreement grows linearly in the arms
The four disagreement rates — 5.4%, 14.4%, 22.4% and 30.8% at two, three, four and five arms — are worth differencing rather than only reading. The successive steps are 9.0, 8.0 and 8.4 points, which is a straight line to within the noise of two hundred trials: each additional arm puts about eight and a half per cent more of the cohort in dispute.
The linearity is not obvious in advance and it is the useful form of the finding. A rate that grew as the number of pairs would have gone as K(K − 1)/2 and reached three quarters of the cohort by five arms; one that saturated would have flattened. Neither happens over the range measured, and the smaller trial’s rates — 5.9%, 14.0%, 20.9%, 28.8% — have the same slope, 8.1 points an arm, which is what says the slope belongs to the scores rather than to the cohort.
Extrapolating a straight line four arms past the last measurement is an extrapolation and is offered as one: at eight arms it puts the disagreement near 56%, so a platform trial described only as “minimisation” would have more than half its patients allocated differently under two readings of the same word. Nothing here measures that, and nothing here suggests the slope should change before it.
What can be said without extrapolating is the ratio. Going from two arms to five multiplies the disagreement by 5.7 while the trial gains three arms, and the reason is that every extra arm adds both a candidate to prefer and a count to the vector being summarised — so the scores get more to disagree about at both ends of their own construction.
Where the variance score’s advantage actually sits
The variance score wins on all three columns at three arms, and the sizes of the three wins are not alike: 3.5% on the average range, 7.0% on the prognosis gap and 11.6% on the worst margin. Ranked that way they are in exactly the order of how much of a tail each column reads.
The average range is a mean over factor levels, and a mean is the summary least sensitive to where the imbalance is concentrated — so a rule whose whole advantage is refusing to concentrate has least to show there. The worst margin is a maximum, which is nothing but concentration, and it is where the advantage is largest. The prognosis gap sits between them because it is a sum across factors of imbalances weighted by how much each factor predicts, so a single large imbalance contributes more than several small ones without contributing all of it.
That ordering is the mechanism showing up three times rather than three separate results, and the concentration ratios say the same thing in one number: worst margin over average margin is 2.019 for the range score and 1.849 for the variance score. Divide the two columns and the ratio of ratios reproduces the difference between the 3.5% and the 11.6% exactly, because that is arithmetic rather than a second measurement.
The practical consequence is about which column a trial should be judged on. A baseline table in a published report shows the margins, one factor at a time, and a reader’s eye goes to the worst row — which is the quantity the two scores differ on most and the one neither rule’s name mentions. A trial reporting its average imbalance is reporting the column where the choice of score matters least, and the point of balancing at all was never the average.
Why this is the same shape of problem as the criterion field’s
A reader who has been through the design half of this site will recognise the structure, and the recognition is worth making explicit because it recurs whenever a scalar summary is required of a matrix or a vector.
An optimal design has to reduce an information matrix to one number before it can be optimised, and the family of ways to do it — A, D, E and everything between — turns out to matter: the design that wins at one end of the family is nearly worst at the other, and the letter is a choice with a 50.2% efficiency gap behind it.
A balancing rule has to reduce a vector of K counts to one number before it can minimise anything, and the family of ways to do that matters in the same way and for the same reason. Both are cases of the same demand: state what unbalanced means, or state what informative means, and notice that the answer does not follow from the words.
The difference is in how visible the choice is. In optimal design the letter is written down in the paper. In minimisation it is not written down anywhere, and a reader has no way to recover it from the trial report — which is why this essay’s practical recommendation is about disclosure rather than about which score to use.
What is claimed, and what is not
This field takes covariate-adaptive allocation to more than two arms: the score as a choice rather than a definition, where the choices coincide and where they part, what each leaves behind on three measures at once, and the unequal-allocation problem the next essay is about. The covariate-adaptive field named allocation to more than two arms as unclaimed when it was written.
What stays out and is named: continuous covariates in a balancing rule, which need a different score entirely and are a real gap; minimisation with factor weights, where the score is a weighted sum and the weights are another undeclared choice of the same kind; and the sequential-parallel and platform designs where arms enter and leave during the trial, which is a different subject with a different error-rate problem.
The boundary against the allocation field is that it owns how many units each arm should get and this owns which unit goes where; they meet in the next essay, where an unequal allocation ratio has to survive a balancing rule.
The checks, and the refusal
Four claims are gated in this field’s library. The pairwise sum must be a fixed multiple of the range at two and three arms and not at four, checked over three thousand random count vectors rather than argued. The two must therefore prefer the same arm on every arrival at three arms, and part company at four. The variance must disagree with the range from two arms onwards, and by more as the arms increase. And the variance score must leave a smaller average range than the range score, with the concentration ratio as the mechanism and the tie rates as the explanation that was tested and rejected.
The refusal belongs to the next essay: a score that does not know what allocation ratio it was asked for.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A proposal that moves more than two units — both name covariate balance, experimental design, minimisation, monte carlo, randomisation
- Where the gain is, and where the decision is — both name covariate balance, experimental design, minimisation, monte carlo, randomisation
- A count that has to be estimated — both name allocation, covariate balance, monte carlo, randomisation
- A cut is not a polynomial, and it does not have to be — both name covariate balance, experimental design, interaction, randomisation
- A dictionary that is a product — both name allocation, covariate balance, interaction, randomisation
- Stationary is not convergent — both name covariate balance, experimental design, monte carlo, randomisation
Named objects
A flat tag is an object no other essay names yet.
AllocationCovariate-adaptive randomisationCovariate balanceExperimental designImbalance scoreInteractionMinimisationMonte CarloPermuted blocksRandomisationStratified randomisationStudy design