More arms than two

Three arms and three scores

Minimisation balances a trial by keeping the arms' counts even inside every prognostic factor. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms.

Worth reading first: Balancing what is known in advance · Randomisation is not balance.

The covariate-adaptive field measures four allocation rules and finds that each of them holds one kind of imbalance flat while letting the others grow exactly as a coin’s do. The rule at the centre of it is minimisation: as each patient arrives, look at the levels of the prognostic factors they belong to, work out how uneven the arms already are inside those levels, and send the patient wherever makes that unevenness smallest — with probability p, so that the rule is not a function anybody can invert.

Every one of those trials has two arms. With two arms the unevenness inside a factor level is |n₁ − n₂|, and there is nothing to choose. The literature’s original proposal offered three different measures of it and never had to say which, because for two arms they are the same number up to a constant.

They are not the same number for three.

Three ways to say how uneven K counts are

Take a factor level in which the arms currently hold (5, 3, 2) patients. How unbalanced is that?

The range — largest minus smallest — says 3. The variance of the counts about their mean says 1.56. The sum of the pairwise gaps — |5−3| + |5−2| + |3−2| — says 6. Each is a defensible summary and each is what somebody’s minimisation implements, and the rule they define is whichever arm makes their own total smallest, summed across the factors the patient belongs to.

For two arms these collapse. With counts (a, b), the range is |a − b|, the pairwise sum is also |a − b|, and the variance is (a − b)²/2. The first two are identical and the third is a monotone transformation of them level by level.

For three, the first two still collapse — sorted counts a ≤ b ≤ c give a pairwise sum of (c − a) + (c − b) + (b − a) = 2(c − a), exactly twice the range — so those two remain one rule. For four, the pairwise sum is 3(d − a) + (c − b), which is not a function of the range at all, and they part company.

Where one rule becomes threeEvery arrival in 200 simulated trials is put to all three scores, and the picture is how often they would send that patient to different arms. The range and the pairwise sum are the same rule at two arms and at three — for sorted counts the pairwise sum is twice the range, so the arm that minimises one minimises the other — and they part company at four, where the pairwise sum is 3(d − a) + (c − b) and the range still sees only d − a. The variance disagrees with both from two arms onwards, on 5.4% of arrivals at two and 27.0% at five, because the scores are summed over 3 factors and a sum of squares does not order the candidates the way a sum of absolute values does. All three are called minimisation.00.1000.2000.3002345arms in the trialshare of arrivals the two scores would send to different armsrange vs variancerange vs pairsvariance vs pairs30000 arrivals per point, cohorts of 150range and pairwise sum are one rule to three arms
Fig. 1 Every arrival in two hundred simulated trials put to all three scores, and how often they would send that patient to different arms. The range and the pairwise sum agree on every arrival at two and three arms, and disagree at four. The variance disagrees with both from two arms onwards.

The disagreement rates are 5.4% of arrivals at two arms, 14.4% at three, 22.4% at four and 30.8% at five. Nothing in a trial protocol distinguishes the three, and by four arms nearly a quarter of patients would have been allocated differently under a description that reads identically.

The first thing to ask of a rate like that is whether it is a small-trial artefact, and it is not. Run the same three-arm comparison on twice as many patients — three hundred arrivals rather than a hundred and fifty — and the disagreement rate is 15.0%, against 14.4% before. It did not fall, and there is no sample size at which it would. A disagreement here is not two rules failing to converge on an answer; it is two rules that were never asking the same question, and asking each of them about more patients gives more occasions on which they differ rather than fewer. That is what makes the choice between them something to declare in the protocol rather than something a larger trial outgrows.

What a disagreement rate is counted on

The number above needs one sentence about how it was measured, because there is a much larger and much less meaningful number available for the same words.

Running three separate trials under three scores and comparing the allocations would measure how far three histories diverge, and two histories that diverge once diverge for ever afterwards: the counts differ, so every subsequent decision is being taken in a different trial. That number would be large and would say nothing about the rules.

What is counted here is a decision rather than a history. One trial runs under one named score, and at every arrival all three scores are asked which arm they would prefer given the counts that trial has actually reached. A disagreement is one arrival on which the preferred sets differ, and two scores that are both indifferent between the same two arms have not disagreed about anything.

Why the variance disagrees even at two arms

The two-arm disagreement is the surprising one, since within a single factor level the variance is a monotone function of the range. The answer is that the rule does not act on a single level: it sums its score across the factors the patient belongs to, and a sum of squares does not order candidates the way a sum of absolute values does.

Suppose an arriving patient’s three factor levels currently stand at gaps of (2, 0, 0) under one candidate arm and (1, 1, 1) under another. The summed range says 2 against 3 and prefers the first. The summed variance says 2 against 1.5 and prefers the second. A rule that squares is willing to accept several small imbalances to avoid one large one, and a rule that does not, is not.

That is a statement about what the rule is for, not a numerical detail, and it is the reason the choice cannot be settled by measuring balance without saying which balance.

Five rules, three measures of what they left, 3 arms. 300 cohorts of 150 patients through 3 prognostic factors and 24 cells, every rule run on the same patients with the same coins. Each column is a different way of measuring the imbalance the rule left: the average range across factor levels, the average variance across them, and the worst cell of the cross-classification. The three minimising rules are close together and far from a uniform draw — and the ordering between them changes with the column, which is the finding: the variance leaves 0.934 on the range where the range leaves 0.963, so the score named after a quantity is not the one that controls it best. Blocks inside every cell are the opposite trade — the only rule that holds the cells, and the loosest margins of the four.
Fig. 2 Five rules over the same cohorts with the same coins, and three measures of what each left behind. The two rules that balance nothing on purpose — a uniform draw and blocks inside every cell — sit at the extremes; the three minimising rules are close together and not in the same order in every column.

The score named after the range is not the best at controlling the range

The comparison in that picture is worth reading slowly, because it contains a result that no description of the rules would suggest.

The range score leaves an average range of 0.9658 across factor levels. The pairwise-sum score leaves 0.9658 — the identical number, because at three arms they are the same rule and, given the same coins, they produce the same trial patient for patient. The variance score leaves 0.9322.

The variance score controls the range better than the score named after it. It also leaves a smaller worst margin — 1.724 against 1.950 — and a smaller imbalance in prognosis between the arms, 0.0477 against 0.0513 standard deviations.

The obvious explanation is that a range is a coarse function, that it often cannot separate two candidate arms, and that the rule then settles the tie with a coin. That explanation is testable and it is false.

The explanation that is not the explanation. The obvious account of why the variance score controls the range better than the range score does is that a range is coarse, ties often, and settles the tie with a coin. This is that account measured, and it is false: at three arms the variance score is indifferent between candidate arms on 32.9% of arrivals and the range score on 26.2%. What actually separates them is the aggregation across factors — the scores are summed over 3 factors, and a sum of squares refuses to trade one large imbalance for several small ones where a sum of ranges is indifferent between them.
Fig. 3 How often each score is indifferent between the candidate arms. The range ties on 26.2% of arrivals at three arms and the variance on 32.9% — the finer score ties more, not less, so tie-breaking is not what separates them.

What separates them is the aggregation. A sum of ranges is indifferent between one large imbalance and several small ones; a sum of squares is not, and the worst margin — the quantity a reader of a baseline table actually looks at — is what a concentrated imbalance produces. The ratio of worst margin to average margin is 2.019 for the range score and 1.849 for the variance score, which is the same fact in the form that shows the mechanism.

The explanation that is not the explanation. The obvious account of why the variance score controls the range better than the range score does is that a range is coarse, ties often, and settles the tie with a coin. This is that account measured, and it is false: at three arms the variance score is indifferent between candidate arms on 32.1% of arrivals and the range score on 25.1%. What actually separates them is the aggregation across factors — the scores are summed over 3 factors, and a sum of squares refuses to trade one large imbalance for several small ones where a sum of ranges is indifferent between them.
Fig. 4 The same tie rates on twice as many patients. Ties become rarer as a trial grows, because larger counts are less often exactly equal, and the ordering between the scores is unchanged — the variance is indifferent more often at every size measured.

The two-arm case was not a special case

It is worth being clear about what the two-arm field measured, since this essay is partly a correction to how its result should be read.

That field ran minimisation with the range score — |n₁ − n₂| summed across factors — and every number it reports is that rule’s. Nothing in it is wrong. What it could not see is that the rule it was measuring is one of a family, because at two arms two of the three members coincide exactly and the third differs on only 5.4% of arrivals.

A five percent disagreement produces trials that are almost the same, and every conclusion of that field survives: each rule holds one imbalance flat and is a coin on the rest, minimisation’s worst factor margin is 3.00 at forty patients against a coin’s 6.32, and the cells are left where a coin leaves them. What changes at three arms is not the conclusion but the specificity: the same sentence now describes three procedures instead of one.

What none of them do

Every minimising rule holds the margins — one factor level at a time — and the cross-classification is not a margin.

Five rules, three measures of what they left, 2 arms. 300 cohorts of 150 patients through 3 prognostic factors and 24 cells, every rule run on the same patients with the same coins. Each column is a different way of measuring the imbalance the rule left: the average range across factor levels, the average variance across them, and the worst cell of the cross-classification. The three minimising rules are close together and far from a uniform draw — and the ordering between them changes with the column, which is the finding: the variance leaves 0.726 on the range where the range leaves 0.741, so the score named after a quantity is not the one that controls it best. Blocks inside every cell are the opposite trade — the only rule that holds the cells, and the loosest margins of the four.
Fig. 5 The same table at two arms, where the range and the pairwise sum are again identical and the variance is again slightly better on every column. Two arms is not a special case that hides the disagreement; it is a case where the disagreement is small.

The worst cell of the twenty-four-cell cross-classification is 4.438 under the range score and 5.600 under a uniform draw — a rule that balances nothing at all. Blocks inside every cell hold the cells at 1.000 by construction and let the margins go to 4.108, four times the minimising rules’ worst margin.

A rule chosen for balance promises nothing about an interaction, which is the two-arm field’s finding re-measured where there are half as many patients per cell per arm to make good on it. If the treatment effect differs between men in one centre and women in another, minimisation has done nothing to make that comparison fair, and it has not claimed to.

The prognosis, which is what the balance was for

Balance on the factors is a means. What a trial is trying to achieve is that the arms are comparable in the thing the factors predict, and that quantity is measurable directly here because the simulation knows each patient’s prognosis.

Under a uniform draw the arms differ in average prognosis by 0.2398 standard deviations. Under minimisation on the range they differ by 0.0513, and on the variance by 0.0477 — a factor of about five better than a coin, from a rule that never saw the prognosis and only ever saw the three factors it is made of.

That factor of five is the whole case for these rules, and it is worth putting beside the differences between them: the gap between any minimising score and no rule at all is twenty times the gap between the scores.

What to choose, and how to say what was chosen

The measurements support a short answer and a shorter piece of advice.

The variance score is the better rule on every measure of margin balance here, at every number of arms measured, and by margins that are small — three to twelve percent — but consistent. It is also the one whose behaviour is easiest to describe: it treats imbalance as a squared quantity and therefore refuses to concentrate it.

The choice has to be stated. “Patients were allocated by minimisation” describes three different rules that would have allocated up to a quarter of the patients differently, and no reader can reconstruct which was run. Naming the score, the factors, the weights and the probability p is four short clauses and makes the allocation reproducible.

The differences are smaller than the difference from randomising. A uniform draw leaves an average range of 6.7789 where every minimising rule leaves under one. Arguing about which score to use while not using any is the error worth avoiding.

Five rules, three measures of what they left, 4 arms. 300 cohorts of 150 patients through 3 prognostic factors and 24 cells, every rule run on the same patients with the same coins. Each column is a different way of measuring the imbalance the rule left: the average range across factor levels, the average variance across them, and the worst cell of the cross-classification. The three minimising rules are close together and far from a uniform draw — and the ordering between them changes with the column, which is the finding: the variance leaves 1.047 on the range where the range leaves 1.103, so the score named after a quantity is not the one that controls it best. Blocks inside every cell are the opposite trade — the only rule that holds the cells, and the loosest margins of the four.
Fig. 6 And four arms, where the range and the pairwise sum have separated. The two columns they used to share now differ, and the variance score’s advantage on the range is unchanged — which is what says the aggregation argument and the pairwise-sum identity are two different things.

The disagreement grows linearly in the arms

The four disagreement rates — 5.4%, 14.4%, 22.4% and 30.8% at two, three, four and five arms — are worth differencing rather than only reading. The successive steps are 9.0, 8.0 and 8.4 points, which is a straight line to within the noise of two hundred trials: each additional arm puts about eight and a half per cent more of the cohort in dispute.

The linearity is not obvious in advance and it is the useful form of the finding. A rate that grew as the number of pairs would have gone as K(K − 1)/2 and reached three quarters of the cohort by five arms; one that saturated would have flattened. Neither happens over the range measured, and the smaller trial’s rates — 5.9%, 14.0%, 20.9%, 28.8% — have the same slope, 8.1 points an arm, which is what says the slope belongs to the scores rather than to the cohort.

Extrapolating a straight line four arms past the last measurement is an extrapolation and is offered as one: at eight arms it puts the disagreement near 56%, so a platform trial described only as “minimisation” would have more than half its patients allocated differently under two readings of the same word. Nothing here measures that, and nothing here suggests the slope should change before it.

What can be said without extrapolating is the ratio. Going from two arms to five multiplies the disagreement by 5.7 while the trial gains three arms, and the reason is that every extra arm adds both a candidate to prefer and a count to the vector being summarised — so the scores get more to disagree about at both ends of their own construction.

Where the variance score’s advantage actually sits

The variance score wins on all three columns at three arms, and the sizes of the three wins are not alike: 3.5% on the average range, 7.0% on the prognosis gap and 11.6% on the worst margin. Ranked that way they are in exactly the order of how much of a tail each column reads.

The average range is a mean over factor levels, and a mean is the summary least sensitive to where the imbalance is concentrated — so a rule whose whole advantage is refusing to concentrate has least to show there. The worst margin is a maximum, which is nothing but concentration, and it is where the advantage is largest. The prognosis gap sits between them because it is a sum across factors of imbalances weighted by how much each factor predicts, so a single large imbalance contributes more than several small ones without contributing all of it.

That ordering is the mechanism showing up three times rather than three separate results, and the concentration ratios say the same thing in one number: worst margin over average margin is 2.019 for the range score and 1.849 for the variance score. Divide the two columns and the ratio of ratios reproduces the difference between the 3.5% and the 11.6% exactly, because that is arithmetic rather than a second measurement.

The practical consequence is about which column a trial should be judged on. A baseline table in a published report shows the margins, one factor at a time, and a reader’s eye goes to the worst row — which is the quantity the two scores differ on most and the one neither rule’s name mentions. A trial reporting its average imbalance is reporting the column where the choice of score matters least, and the point of balancing at all was never the average.

Why this is the same shape of problem as the criterion field’s

A reader who has been through the design half of this site will recognise the structure, and the recognition is worth making explicit because it recurs whenever a scalar summary is required of a matrix or a vector.

An optimal design has to reduce an information matrix to one number before it can be optimised, and the family of ways to do it — A, D, E and everything between — turns out to matter: the design that wins at one end of the family is nearly worst at the other, and the letter is a choice with a 50.2% efficiency gap behind it.

A balancing rule has to reduce a vector of K counts to one number before it can minimise anything, and the family of ways to do that matters in the same way and for the same reason. Both are cases of the same demand: state what unbalanced means, or state what informative means, and notice that the answer does not follow from the words.

The difference is in how visible the choice is. In optimal design the letter is written down in the paper. In minimisation it is not written down anywhere, and a reader has no way to recover it from the trial report — which is why this essay’s practical recommendation is about disclosure rather than about which score to use.

Where one rule becomes three. Every arrival in 200 simulated trials is put to all three scores, and the picture is how often they would send that patient to different arms. The range and the pairwise sum are the same rule at two arms and at three — for sorted counts the pairwise sum is twice the range, so the arm that minimises one minimises the other — and they part company at four, where the pairwise sum is 3(d − a) + (c − b) and the range still sees only d − a. The variance disagrees with both from two arms onwards, on 5.9% of arrivals at two and 25.2% at five, because the scores are summed over 3 factors and a sum of squares does not order the candidates the way a sum of absolute values does. All three are called minimisation.
Fig. 7 Eighty patients rather than a hundred and fifty: 5.9%, 14.0%, 20.9% and 28.8% at two, three, four and five arms, against 5.4%, 14.4%, 22.4% and 30.8%. The rates barely move with the size of the trial, which is what says they are a property of the scores rather than of how much data they are being asked to sort.
Five rules, three measures of what they left, 3 arms. 300 cohorts of 300 patients through 3 prognostic factors and 24 cells, every rule run on the same patients with the same coins. Each column is a different way of measuring the imbalance the rule left: the average range across factor levels, the average variance across them, and the worst cell of the cross-classification. The three minimising rules are close together and far from a uniform draw — and the ordering between them changes with the column, which is the finding: the variance leaves 0.922 on the range where the range leaves 0.970, so the score named after a quantity is not the one that controls it best. Blocks inside every cell are the opposite trade — the only rule that holds the cells, and the loosest margins of the four.
Fig. 8 Three hundred patients rather than a hundred and fifty. Every imbalance falls, the ordering between the rules does not change, and the variance score keeps its advantage on the range — which is what says the finding is arithmetic rather than a small-sample accident.

What is claimed, and what is not

This field takes covariate-adaptive allocation to more than two arms: the score as a choice rather than a definition, where the choices coincide and where they part, what each leaves behind on three measures at once, and the unequal-allocation problem the next essay is about. The covariate-adaptive field named allocation to more than two arms as unclaimed when it was written.

What stays out and is named: continuous covariates in a balancing rule, which need a different score entirely and are a real gap; minimisation with factor weights, where the score is a weighted sum and the weights are another undeclared choice of the same kind; and the sequential-parallel and platform designs where arms enter and leave during the trial, which is a different subject with a different error-rate problem.

The boundary against the allocation field is that it owns how many units each arm should get and this owns which unit goes where; they meet in the next essay, where an unequal allocation ratio has to survive a balancing rule.

The checks, and the refusal

Four claims are gated in this field’s library. The pairwise sum must be a fixed multiple of the range at two and three arms and not at four, checked over three thousand random count vectors rather than argued. The two must therefore prefer the same arm on every arrival at three arms, and part company at four. The variance must disagree with the range from two arms onwards, and by more as the arms increase. And the variance score must leave a smaller average range than the range score, with the concentration ratio as the mechanism and the tie rates as the explanation that was tested and rejected.

The refusal belongs to the next essay: a score that does not know what allocation ratio it was asked for.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AllocationCovariate-adaptive randomisationCovariate balanceExperimental designImbalance scoreInteractionMinimisationMonte CarloPermuted blocksRandomisationStratified randomisationStudy design