More arms than two

Balancing towards unequal targets

A three-arm trial allocating two to one to one is the ordinary case, and a balancing rule built from raw counts does not know it. It balances the arms towards equality inside every factor level, delivers a third to each arm, and reports that it minimised imbalance.

Worth reading first: Not half and half · Balancing what is known in advance.

The allocation field’s whole subject is that equal allocation is a default rather than an optimum. A shared control against k treatments should be sampled √k times as heavily as each treatment; an arm with more variance should be sampled more; an arm that costs more should be sampled less. Every one of those rules produces a target ratio that is not one to one.

The covariate-adaptive field’s whole subject is a rule that keeps the arms comparable on what was recorded before treatment. Its score, at every arrival, asks which arm makes the counts inside the patient’s factor levels most even.

Put the two together and the second quietly destroys the first.

What a raw-count score delivers

Take a three-arm trial designed to allocate 2:1:1 — a control arm at half the patients and two treatments at a quarter each, which is what one control against several treatments calls for. Run minimisation on the counts as they stand.

A trial designed 2:1:1, and what two scores deliver500 trials of 180 patients, three arms, a target of 2:1:1. The shaded bars are a minimisation score that divides each arm's count by the share that arm is supposed to receive before measuring the spread; it delivers 49.9% : 25.1% : 25.1%. The others are the same rule with the counts left raw, which delivers 33.4% : 33.3% : 33.3% — the balance it enforces inside every factor level is equality, and equality is what it gets. The marks are the shares that were asked for.arm 1 — asked for 50%49.9%arm 1 — raw counts33.4%arm 2 — asked for 25%25.1%arm 2 — raw counts33.3%arm 3 — asked for 25%25.1%arm 3 — raw counts33.3%500 trials of 180, p = 0.85, three factors, one set of seedsraw counts deliver 33.4:33.3:33.3 where 2:1:1 was designed
Fig. 1 Five hundred trials of a hundred and eighty patients, three arms, a design of 2:1:1. The shaded bars are a score that compares each arm’s count against the share it was supposed to receive. The others are the same rule with the counts left raw. The marks are the shares that were asked for.

The weighted score delivers 49.9 : 25.1 : 25.1. The raw-count score delivers 33.4 : 33.3 : 33.3.

It is not approximately wrong. The design has been replaced, arm by arm, with an equal-allocation trial, and every quantity that depended on the ratio — the power of the control-versus-treatment comparisons, the precision of the control mean, whatever cost calculation produced the 2:1:1 in the first place — has been silently recomputed at the wrong numbers.

Why it happens, which is not a bug

Nothing has gone wrong inside the rule. Minimisation was asked to make counts even and it made them even; evenness inside a factor level is exactly what a raw-count score measures, and a 2:1:1 trial is maximally uneven by that measure. The rule spends every arrival correcting the imbalance the design asked for.

The randomisation step does not save it either. Minimisation is usually run with a probability p of following its own preference — 0.85 here — and the remaining probability spread over the other arms. That randomisation is symmetric across arms, so it pulls towards equality as well.

The repair is to compare each arm’s count against its target rather than against the other arms: divide by the share that arm is meant to receive, and then measure the spread of what comes out. An arm at half the patients and an arm at a quarter of them are perfectly balanced in that scale when they hold exactly twice and once of something, which is what the design meant by balanced all along.

That is one division, it costs nothing, and it turns 33.4 : 33.3 : 33.3 into 49.9 : 25.1 : 25.1.

How the failure grows with the ratio

The more unequal the design, the less of it survives. Four target ratios, 400 trials of 180 patients each. The rising curve is a minimisation score built from raw counts: it balances the arms towards equality inside every factor level, which is a different trial from the one designed, and the gap grows with the ratio — 0.0 points at 1:1:1, 16.6 points at 2:1:1, 26.5 points at 3:1:1, 33.1 points at 4:1:1. The flat curve is the same rule comparing each arm's count against its own target, which delivers what was asked for at every ratio. Nothing in a trial's output distinguishes the two: both report a minimisation algorithm, and only one of them ran the design.
Fig. 2 Four target ratios, four hundred trials each. The rising curve is what a raw-count score delivers, measured as the worst departure from the share asked for; the flat one is the weighted score, which is within a point of its target at every ratio.

The worst departures are 0.0 points at 1:1:1, 16.6 at 2:1:1, 26.5 at 3:1:1 and 33.1 at 4:1:1. That is the shape it has to have: an unweighted score always delivers equal allocation, so the error is exactly the distance between the design and equality, and it grows without limit as the design becomes more unequal.

Two consequences follow and both are worth stating in the form a trialist would meet them.

The failure is invisible in a 1:1:1 trial, where the two rules agree exactly, so a procedure validated on an equal-allocation trial and reused on an unequal one carries no warning.

The failure is largest exactly where the ratio was most deliberate. A 4:1:1 design is not an accident; somebody computed it. It is the design whose intention is most completely undone.

A trial designed 4:1:1, and what two scores deliver. 500 trials of 180 patients, three arms, a target of 4:1:1. The shaded bars are a minimisation score that divides each arm's count by the share that arm is supposed to receive before measuring the spread; it delivers 66.0% : 17.0% : 17.0%. The others are the same rule with the counts left raw, which delivers 33.5% : 33.2% : 33.2% — the balance it enforces inside every factor level is equality, and equality is what it gets. The marks are the shares that were asked for.
Fig. 3 The 4:1:1 case in the same form as the first picture. The weighted rule delivers 66.0 : 17.0 : 17.0 against a target of 66.7 : 16.7 : 16.7; the raw-count rule delivers 33.5 : 33.2 : 33.2, and the distance between the marks and the unshaded bars is the entire design decision being overwritten.

The three scores, under an unequal target

The previous essay’s three ways of measuring unevenness all need the same repair and none of them supplies it. The range of standardised counts, their variance and their pairwise sum are three summaries of the same standardised vector, so the weighting sits underneath all three and the choice between them is unchanged by it.

That is worth one sentence rather than a section, and the sentence is that the two choices are independent: which score, and what it is measured on. A trial report that says “minimisation” has left both of them unstated, and the second is the one that can move half the patients.

Where one rule becomes three. Every arrival in 200 simulated trials is put to all three scores, and the picture is how often they would send that patient to different arms. The range and the pairwise sum are the same rule at two arms and at three — for sorted counts the pairwise sum is twice the range, so the arm that minimises one minimises the other — and they part company at four, where the pairwise sum is 3(d − a) + (c − b) and the range still sees only d − a. The variance disagrees with both from two arms onwards, on 5.9% of arrivals at two and 27.6% at five, because the scores are summed over 3 factors and a sum of squares does not order the candidates the way a sum of absolute values does. All three are called minimisation.
Fig. 4 The other undeclared choice, from the previous essay: how often the three scores would send the same patient to different arms. Both choices are invisible in the same sentence of the same protocol, and they compose — a trial can be run under any of six combinations that all read identically.

What it costs, in the currency the design was chosen in

The allocation field measures why unequal ratios exist, and each of its reasons converts the departure above into a loss.

A shared control. With k treatments against one control, the variance-minimising allocation puts √k times as many units on the control as on each treatment. At k = 2 that is 1.41 : 1 : 1, at k = 4 it is 2 : 1 : 1 : 1 : 1. A rule that equalises the arms is running the design that field measures as suboptimal, and the loss it computes there is the loss here.

Unequal variances. The optimal allocation is proportional to the standard deviations, and the same argument applies unchanged.

Cost. The optimal ratio is proportional to σ/√c, and an equalised trial simply spends more money for the same precision.

In every case the arithmetic of what is lost is in the allocation field, and what this essay adds is that the loss can be incurred by a rule that was never asked about ratios at all.

The more unequal the design, the less of it survives. Four target ratios, 400 trials of 360 patients each. The rising curve is a minimisation score built from raw counts: it balances the arms towards equality inside every factor level, which is a different trial from the one designed, and the gap grows with the ratio — 0.0 points at 1:1:1, 16.6 points at 2:1:1, 26.6 points at 3:1:1, 33.2 points at 4:1:1. The flat curve is the same rule comparing each arm's count against its own target, which delivers what was asked for at every ratio. Nothing in a trial's output distinguishes the two: both report a minimisation algorithm, and only one of them ran the design.
Fig. 5 The same four ratios on twice as many patients. Nothing improves: at 2:1:1 on three hundred and sixty patients the raw-count rule delivers 33.4 : 33.3 : 33.3 and the weighted one 50.0 : 25.0 : 25.0, so a larger trial simply converges more exactly on the shares each rule was aiming at.

That last picture is the one worth keeping. A larger trial does not fix this, because the rule is converging on the allocation it was told to produce; it was told the wrong one.

How much of the repair the randomisation step manages on its own

The section below says randomising in the target proportions is “nowhere near sufficient”, and the four departures already measured say exactly how near.

If the raw-count rule delivered perfect equality the departure would be the whole distance from the design to equality: 16.7, 26.7 and 33.4 points at 2:1:1, 3:1:1 and 4:1:1. The measured departures are 16.6, 26.5 and 33.1.

So the 1 − p branch — fifteen per cent of arrivals, spread in the target proportions — recovers 0.1, 0.2 and 0.3 points of the correction needed. As a share of the damage that is 0.6%, 0.8% and 0.9%.

A repair that delivers under one per cent of what is required is not a partial repair. It is a repair that leaves the number in the report indistinguishable from no repair at all, and it grows in absolute terms with the ratio, which is exactly the pattern that makes it look like it is doing something.

Which quantity the equalisation actually damages

“Every quantity that depended on the ratio has been recomputed at the wrong numbers” is true and it is worth saying which ones move and by how much, because they do not all move and one of them moves the other way.

Take a hundred and eighty patients and the variance of a single control-versus-treatment difference, which is proportional to 1/nc+1/nt1/n_c + 1/n_t.

  • At 2:1:1 the arms hold 90, 45 and 45, giving 1/90 + 1/45 = 1/30. Equalised at 60, 60, 60 it is 2/60 = 1/30. Identical, exactly.
  • At 4:1:1 the arms hold 120, 30 and 30, giving 1/120 + 1/30 = 1/24. Equalised it is 1/30, which is 20% better.
  • At the √k optimum for two treatments, 1.41:1:1, the arms hold 74.5, 52.7 and 52.7, and the sum is 2.9% below either of the above.

So on a single pairwise comparison, equalising a 2:1:1 trial costs nothing at all and equalising a 4:1:1 trial is an improvement. The √k rule is what the shared-control argument produces for two treatments, and it is 1.41 rather than 2 — so neither 3:1:1 nor 4:1:1 comes from that argument in the first place.

What the equalisation does destroy is the control mean itself. At 2:1:1 the control’s variance goes from σ²/90 to σ²/60, half as large again; at 4:1:1 from σ²/120 to σ²/60, twice. A design that puts two thirds of its patients on the control is a design estimating the control precisely — because it is a reference for several comparisons at once, or because a cost or variance argument put it there — and that is the quantity the raw-count rule halves the precision of.

So the loss is real and it is not where the ratio’s usual justification points. A trialist checking whether the equalisation mattered by recomputing one pairwise power will find nothing wrong, which is the worst possible outcome for a defect this large.

Where the weighting has to go, exactly

The repair is one division and it has to be in the right place, which is worth spelling out because there are two plausible places and only one of them works.

Inside the score, before the spread is measured. Each arm’s count is divided by its target share and the imbalance is measured on the standardised counts, so an arm holding twice as many patients as another is balanced when its target is twice as large.

Not in the randomisation step alone. Giving the preferred arm probability p and spreading the rest in proportion to the targets — which this implementation also does — is necessary and nowhere near sufficient: with an unweighted score, the preferred arm is whichever arm is behind on raw counts, so the rule spends its p pulling towards equality and its 1 − p pulling towards the target. The result is a compromise nobody designed, and at 2:1:1 with p = 0.85 it is much nearer equality than the target.

The distinction matters because the second repair is the one that looks sufficient. An implementation that randomises in the target proportions and scores on raw counts will produce arm sizes that are wrong in a way that varies with p — which is a parameter chosen for predictability reasons — and nothing in the output will connect the two.

The coin costs the ratio a little even with the score weighted, and the cost grows with how unequal the design is. At 4:1:1 the weighted rule delivers 67.0% of the runs to the first arm at p = 1, 66.0% at p = 0.85 and 62.3% at p = 0.7, because the 1 − p branch spreads over the other arms in proportion and cannot pull as hard as the score does. At 2:1:1 the same sweep runs 50.2%, 49.9%, 49.1%. It is a second-order effect beside the sixteen and thirty-three points a raw-count score gives away, and it is the reason the check on this figure is a relative bound rather than an absolute one.

A design decision is not a preference

One more framing, because it is the reason this essay sits in the allocation field’s ladder rather than the assignment field’s.

The allocation field’s rules are derived: √k for a shared control, proportional to σ for unequal variances, proportional to σ/√c when the arms cost different amounts. Each is the solution of a minimisation with a stated objective, and the ratio it returns is not a taste. It is the answer.

A balancing rule that overwrites it is not making a trade-off between balance and allocation — it is not aware there is anything to trade. The weighted score is what turns the two into a genuine joint problem, at which point they compose without conflict: the ratio decides how many, the score decides which, and neither has to give anything up.

The alternative that does not need the repair

There is one rule in the covariate-adaptive family that handles unequal targets without any weighting, and it is worth naming because it is what many trials use.

Permuted blocks hold the ratio by construction: a block of four containing two controls and one of each treatment delivers 2:1:1 exactly at the end of every block, whatever the covariates do. What blocks do not do is balance any covariate — the two-arm field measures a stratified block scheme failing exactly where the cells are thin, and blocks without strata balance nothing at all.

So the choice is between a rule that gets the ratio right and the covariates wrong, and a rule that gets both right provided somebody remembered to divide by the target. That is not much of a choice, and the reason it is worth an essay is that the second rule fails silently: a blocked trial that loses covariate balance shows it in the baseline table, and a minimised trial that loses the ratio shows it only in the arm sizes, which are usually reported as the number of patients recruited rather than as a design decision that was or was not honoured.

Five rules, three measures of what they left, 3 arms. 300 cohorts of 180 patients through 3 prognostic factors and 24 cells, every rule run on the same patients with the same coins. Each column is a different way of measuring the imbalance the rule left: the average range across factor levels, the average variance across them, and the worst cell of the cross-classification. The three minimising rules are close together and far from a uniform draw — and the ordering between them changes with the column, which is the finding: the variance leaves 0.949 on the range where the range leaves 0.951, so the score named after a quantity is not the one that controls it best. Blocks inside every cell are the opposite trade — the only rule that holds the cells, and the loosest margins of the four.
Fig. 6 The rules from the previous essay, for the comparison: blocks inside every cell are the only rule that holds the cross-classification, and their margins are four times looser than any minimising rule’s. Neither of the two things a trialist wants comes free.

What the weighted score does to balance

A fair question about the repair is whether it costs anything in the balance it was there to buy. It does not.

At 2:1:1 the weighted rule leaves a worst standardised margin of 3.17 against the raw-count rule’s 21.72 — the raw-count rule is worse on the very quantity it was optimising, once that quantity is measured on the scale the design defined. That is not a coincidence: the standardised margin asks how far each arm is from its own target, and a rule driving all three arms to a third is driving two of them a long way from theirs.

Measured on its own scale the raw-count rule wins, and measured on the design’s scale it loses badly, which is the same lesson the previous essay’s three scores taught in a milder form. A comparison between allocation rules has to state what balance means before any of it means anything.

A trial designed 2:1:1, and what two scores deliver. 500 trials of 90 patients, three arms, a target of 2:1:1. The shaded bars are a minimisation score that divides each arm's count by the share that arm is supposed to receive before measuring the spread; it delivers 50.3% : 24.8% : 24.9%. The others are the same rule with the counts left raw, which delivers 33.3% : 33.4% : 33.3% — the balance it enforces inside every factor level is equality, and equality is what it gets. The marks are the shares that were asked for.
Fig. 7 A smaller trial under a fully deterministic rule, where the randomisation step is switched off entirely. The raw-count score still delivers a third to each arm — so the failure is the score’s and not the coin’s, which is the control the previous section’s argument needs.

Why nobody notices

The failure has three properties that between them make it hard to see, and they are worth listing because they are the properties that make any silent defect survive.

It produces a perfectly ordinary trial. Equal allocation is not a pathology. The arms are balanced on the covariates, the analysis runs, the standard errors are honest for the trial that was run, and every diagnostic a trialist looks at is clean.

The number that is wrong is one nobody treats as an output. Arm sizes are reported as a fact about recruitment, not as a design parameter that was met or missed, so the comparison against the target is a comparison nobody makes.

The loss is in power rather than in validity. Nothing here inflates a false-positive rate or biases an estimate; the trial is simply less efficient than the one that was designed, and a trial that misses significance for that reason looks exactly like a trial whose effect was smaller than hoped.

This is the same shape as the two-arm field’s finding that an unadjusted analysis after minimisation is conservative rather than anticonservative: a defect that costs power and breaks no rule is a defect that nothing in the standard apparatus is looking for.

The more unequal the design, the less of it survives. Four target ratios, 400 trials of 180 patients each. The rising curve is a minimisation score built from raw counts: it balances the arms towards equality inside every factor level, which is a different trial from the one designed, and the gap grows with the ratio — 0.0 points at 1:1:1, 16.7 points at 2:1:1, 26.7 points at 3:1:1, 33.3 points at 4:1:1. The flat curve is the same rule comparing each arm's count against its own target, which delivers what was asked for at every ratio. Nothing in a trial's output distinguishes the two: both report a minimisation algorithm, and only one of them ran the design.
Fig. 8 The ratio sweep under a deterministic rule. The picture is the one at the top of this section: the departures are unchanged, so the whole of the effect belongs to the score and none of it to the randomisation probability.

What is claimed, and what is not

The claim is what an unequal allocation ratio does to a covariate-adaptive rule, and what the rule does to it: the delivered shares under both scores, the departure as a function of the ratio, its invariance to the size of the trial, and the standardised margin under each.

What stays out: response-adaptive ratios that change during the trial, which belong to the adaptive field and where the target is a moving object; allocation ratios chosen to maximise power against a specific alternative rather than to minimise a variance; and the analysis of an unequally allocated trial, where the unequal arms change the degrees of freedom and the multiple-comparison structure — that is Dunnett’s problem and it is measured there.

What to check in a report

Three things, all of which a reader can do from the arm sizes alone, and none of which requires the allocation code.

Compare the realised shares against the design. A 2:1:1 trial that recruited 180 patients should report about 90, 45 and 45. A report of 60, 60, 60 under the heading of minimisation is this failure, and it is not a recruitment accident — the arms are too equal, not too unequal.

Look at how the imbalance was defined. If a protocol says the rule minimised the difference in the numbers allocated within each stratum, the score is on raw counts and the ratio was not protected.

Check whether the ratio was ever tested. An implementation validated on an equal-allocation trial has never exercised the code path that matters here, because at 1:1:1 the two scores are the same function.

A trial designed 3:1:1, and what two scores deliver. 500 trials of 180 patients, three arms, a target of 3:1:1. The shaded bars are a minimisation score that divides each arm's count by the share that arm is supposed to receive before measuring the spread; it delivers 59.6% : 20.2% : 20.2%. The others are the same rule with the counts left raw, which delivers 33.5% : 33.2% : 33.3% — the balance it enforces inside every factor level is equality, and equality is what it gets. The marks are the shares that were asked for.
Fig. 9 The 3:1:1 case, between the two already drawn: 59.6 : 20.2 : 20.2 from the weighted score against a target of 60 : 20 : 20, and 33.5 : 33.2 : 33.3 from the raw-count one. The failure is the same failure at every ratio, and its size is exactly the distance from equality.

The checks, and the refusal

Two claims are gated and one refusal carries the field. The weighted score must deliver every arm’s target share to within two points, and the raw-count score must miss the first arm’s by more than five — which is the repair and the failure asserted as bounds rather than described. The departure must grow with how unequal the target is, at every ratio measured.

The refusal is the raw-count score itself, against the standard that a rule asked for 2:1:1 delivers 2:1:1. It delivers 33.4 : 33.3 : 33.3, off by 16.6 points, where the weighted score is off by 0.1. The check requires it to fail, because an assertion that has never rejected anything proves nothing.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AllocationAllocation ratioCovariate-adaptive randomisationCovariate balanceDunnettExperimental designImbalance scoreMinimisationMonte CarloNeyman allocationPermuted blocksStudy design