More arms than two

Guessing one arm in three

A balancing rule is guessable because it is balancing. With three arms the next assignment is worked out less often than with two — and by more, relative to what a guesser gets for nothing, and the damage they can do is almost unchanged.

Worth reading first: Balancing what is known in advance · Randomisation is not balance.

Balance and predictability are one property seen from two sides. A rule that corrects an imbalance is a rule whose next move can be worked out by anybody who can see the imbalance, and the two-arm field measures what that is worth to an investigator who does the working out: a deterministic minimisation is guessable 87.5% of the time, and a guessing investigator extracts 0.75 of a standard deviation of treatment effect where none exists.

With three arms the guesser has more ways to be wrong. Whether that makes the trial safer is a question with three answers, and they do not agree.

Three quantities, three directions

What a guesser gets, and what a guesser gets for nothing600 trials of 150 patients under a fully deterministic rule, with an investigator who knows the rule, the factors and every assignment so far. The guess rate falls with the number of arms — 87.8% at 2, 86.2% at 3, 81.0% at 4 — which reads like a trial getting safer and is not: what a guesser can trade on is the excess over the 50%, 33%, 25% they would get by naming an arm at random, and that goes the other way, from 1.76× chance at two arms to 3.24× at 4. The gap a guesser manufactures between the best and worst arm under a true null is 0.759, 0.733, 0.711 standard deviations — nearly unchanged.2 arms — guessed87.8%2 arms — by chance50.0%3 arms — guessed86.2%3 arms — by chance33.3%4 arms — guessed81.0%4 arms — by chance25.0%600 trials of 150, deterministic minimisation, three factors1.76× chance at two arms and 3.24× at 4
Fig. 1 Six hundred trials of a hundred and fifty patients under a fully deterministic rule, at two, three and four arms. Each pair is what a guesser who knows the rule, the factors and every assignment so far actually achieves, against what they would achieve by naming an arm at random.

The raw guess rate falls. It is 87.8% at two arms, 86.2% at three and 81.0% at four. More arms means more ways for a guess to miss, and the rule’s preferred arm is not always unique.

The advantage over chance rises. A guesser naming an arm at random is right half the time with two arms and a third of the time with three, so the ratio of what they achieve to what they would get for nothing runs 1.76×, 2.59×, 3.24×. By that measure a four-arm trial is more predictable than a two-arm one, not less.

The damage is nearly unchanged. The quantity a three-arm trial actually reports is the gap between its best and worst arm, and a guessing investigator manufactures 0.759, 0.733 and 0.711 standard deviations of it at two, three and four arms — under a true null, from a mechanism that breaks no rule.

The three statements are all true and they answer three different questions. Which one matters depends on what the guesser is trying to do, and the third is the one that reaches the reader of the trial.

Where the two-arm formula stops applying

The two-arm field’s closed form is worth carrying over as far as it goes, because where it fails is informative.

With two arms, a correct guess enrols a patient whose prognosis is δ better than average into the favoured arm, and by symmetry an incorrect guess enrols one δ worse into the other. The arms move in opposite directions, so the difference in means moves by twice δ times the excess of the guess rate over a half: 2δ(2g − 1), and the factor of two is the part that is easy to leave out.

With three arms there is no symmetry to exploit. A correct guess moves one arm up; the compensation is spread over two others, so each of them moves down by half as much, and the effect on the favoured arm against the average of the rest is not the same quantity as the effect on the largest gap between any two arms. Both are counted here rather than derived, and they differ: 0.650 against 0.733 at three arms, from the same trials.

That divergence is why the field reports two summaries. A trial with several arms has no single treatment effect, and a bias that manufactures one contrast does not manufacture all of them equally.

How the damage is done, which involves no cheating

The mechanism is the two-arm field’s, unchanged: an investigator who can predict the next assignment does not tamper with it. They decide who to enrol next, which is the one thing an allocation rule cannot control.

When the guess says the next patient goes to the arm the investigator favours, they enrol somebody whose prognosis is a little better than average. When it says otherwise, they wait, or enrol somebody a little worse. Nothing about the randomisation is violated, no assignment is changed, and the arms end up differing in prognosis by an amount that is entirely a function of the guess rate.

With two arms the arithmetic is clean — a correct guess moves one arm up and the other down, so the effect is 2δ(2g − 1) with g the guess rate — and with three arms it is not, because a correct guess moves one arm up and shares the compensation among the others. What is counted here rather than derived is the effect on the two summaries a trial reports: the favoured arm against the average of the rest, and the largest gap between any two arms.

At three arms and a deterministic rule those are 0.650 and 0.733 standard deviations, against a standard error of 0.0074, so none of it is noise.

What a guesser is assumed to know, and what they cannot

The guess rate above is an upper bound with a specific construction behind it, and both halves matter.

The guesser is given the rule, the score, the factors, and every assignment made so far. From those they compute the same margin score the rule computes and name whichever arm it prefers. That is what an unblinded investigator enrolling into an open trial has, which is why it is the right assumption for a bound.

What they are not given is the rule’s own coin. Where two or more arms are tied on the score the rule chooses between them at random, and the guesser can only name one of them; the earlier version of this measurement handed the guesser the realised tie-break and reported every rule as 100% guessable at p = 1, which is a measurement of a rule nobody runs. The ties are the whole of the shortfall from certainty: at three arms a deterministic rule is guessed 86.2% of the time, and the remaining 13.8% is arrivals on which the score could not separate the candidates.

That has a design consequence worth naming, and it is the rare one that costs nothing. A score that ties more often is less guessable, and the variance score — the one that controls balance better — is indifferent on 32.9% of arrivals against the range score’s 26.2%. Measured directly, it is guessed 82.4% of the time against the range score’s 86.2%, and the bias a guesser extracts falls from 0.650 to 0.598. The score that balances best is also the one that hides most, so the two undeclared choices of this field point the same way.

What the randomisation probability buys

Minimisation is normally run with a probability p of following its own preference — the standard choice is 0.8 — precisely to make this harder.

What a guesser gets, and what a guesser gets for nothing. 600 trials of 300 patients under a fully deterministic rule, with an investigator who knows the rule, the factors and every assignment so far. The guess rate falls with the number of arms — 88.0% at 2, 86.7% at 3, 81.6% at 4 — which reads like a trial getting safer and is not: what a guesser can trade on is the excess over the 50%, 33%, 25% they would get by naming an arm at random, and that goes the other way, from 1.76× chance at two arms to 3.27× at 4. The gap a guesser manufactures between the best and worst arm under a true null is 0.754, 0.703, 0.682 standard deviations — nearly unchanged.
Fig. 2 The same measurement on twice as many patients. The rates barely move: predictability is a property of the rule rather than of how long it has been running, which is what makes it something to design against rather than to outgrow.

At three arms the guess rate is 86.2% at p = 1, 75.4% at p = 0.85 and 55.5% at p = 0.6. The bias to the favoured arm follows: 0.650, 0.507, 0.270. Buying unpredictability with randomisation works, and it works in proportion.

What it costs is balance, and the two-arm field measured the trade with the conclusion that there is no knee in it: the guess rate is close to linear in p and the imbalance close to geometric, so every increment of one buys a proportion of the other and where to sit is a judgement rather than an optimum. Nothing about a third arm changes that shape.

What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 150 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 50.1% at p = 0.5, which is a coin and cannot be beaten, and 87.9% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.5 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.76 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).
Fig. 3 The two-arm field’s own measurement of the same trade, where the guess rate and the bias it buys are drawn against p. The three-arm numbers sit on the same curves with a different vertical scale.
What a guesser gets, and what a guesser gets for nothing. 600 trials of 150 patients under a fully deterministic rule, with an investigator who knows the rule, the factors and every assignment so far. The guess rate falls with the number of arms — 87.8% at 2, 86.2% at 3, 81.0% at 4 — which reads like a trial getting safer and is not: what a guesser can trade on is the excess over the 50%, 33%, 25% they would get by naming an arm at random, and that goes the other way, from 1.76× chance at two arms to 3.24× at 4. The gap a guesser manufactures between the best and worst arm under a true null is 0.759, 0.733, 0.711 standard deviations — nearly unchanged.
Fig. 4 The same six bars, with the trial size held at the value the rest of this field uses. What a guesser achieves and what they would achieve for nothing are drawn as a pair at every width so that neither number can be read without the other — which is the whole difficulty of describing predictability with one figure.

Why more arms are more predictable relative to chance

The ratio going up while the rate goes down is worth explaining rather than reporting, because it is the one number in this field that is genuinely counterintuitive.

A guesser facing K arms has a free floor of 1/K, and the rule’s preference is a ranking rather than a single arm. With two arms, the rule is nearly always indifferent or decisive, and a decisive rule is guessed with certainty — hence 87.8%, with the shortfall coming from arrivals where two arms tie and the rule tosses a coin the guesser cannot see.

With four arms ties are more common — the score is coarser relative to the number of candidates — so the raw rate falls to 81.0%. But the floor has fallen much faster, from 50% to 25%, so what the guesser knows that they did not know before has grown: they have eliminated three arms rather than one.

The quantity that matters for the damage is neither of those, which is why the gap between the best and worst arm barely moves: manufacturing a difference between two named arms requires knowing which of them the next patient is going to, and that is roughly as easy at four arms as at two.

What a guesser gets, and what a guesser gets for nothing. 600 trials of 90 patients under a fully deterministic rule, with an investigator who knows the rule, the factors and every assignment so far. The guess rate falls with the number of arms — 87.5% at 2, 85.8% at 3, 80.5% at 4 — which reads like a trial getting safer and is not: what a guesser can trade on is the excess over the 50%, 33%, 25% they would get by naming an arm at random, and that goes the other way, from 1.75× chance at two arms to 3.22× at 4. The gap a guesser manufactures between the best and worst arm under a true null is 0.756, 0.742, 0.754 standard deviations — nearly unchanged.
Fig. 5 Ninety patients, where the counts are smaller and ties more frequent: every guess rate is a little lower and the pattern across arms is unchanged. Nothing here is a large-trial phenomenon.

What the trial actually loses

A manufactured difference of seven tenths of a standard deviation under a true null is easier to judge with something to compare it against.

The trials in this field are powered to detect an effect of about half a standard deviation with seventy percent probability, so a guessing investigator can generate an effect larger than the one the trial was designed to find, with no data fabricated, no assignment altered and no protocol violated. The p-value that results is a correct p-value for the data that were collected; what is wrong is that the patients in the arms are not comparable, which is the one thing randomisation was supposed to guarantee.

That is why predictability is a design problem rather than an analysis problem. No adjustment recovers it — the covariates that would explain the difference are the ones the investigator chose on and did not record — and the exact test of the previous essay does not touch it either, because the allocation rule really did produce that allocation. The reference distribution is right and the patients are wrong.

What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 150 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 50.1% at p = 0.5, which is a coin and cannot be beaten, and 87.9% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.25 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.38 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).
Fig. 6 The same trade at half the enrolment advantage: a guesser willing to move a patient’s prognosis by a quarter of a standard deviation rather than a half. Everything scales, which is what says δ is a statement about the investigator and g is a statement about the rule.

What a trial can do about it

Four things, of which only the first two are about the allocation rule.

Keep a coin in the rule. p = 0.8 rather than 1 costs a little balance and takes the guess rate from 86.2% to about three quarters, and the bias with it.

Do not publish the score. The guess rate measured here assumes the guesser knows the rule, the factors and every assignment so far. The last of those is what an unblinded investigator has, and concealment of the assignments already made is worth more than any amount of randomisation inside the rule.

Blind the treatment. The bias enters through who is enrolled after the guess, and an investigator who cannot tell which arm a patient went to cannot maintain the count that makes the guess possible.

Report the balance. A trial whose baseline table is suspiciously good on the balanced factors and suspiciously uneven on prognosis outside them is a trial worth asking about, and the standardised prognosis gap in the previous essays is what that comparison would look like.

The fourth summary, which is flat

Three quantities disagree about whether more arms are safer, and there is a fourth that neither the rate nor the ratio is: the share of the guesser’s uncertainty that the rule removed, which is (g − 1/K)/(1 − 1/K).

At two, three and four arms that is 0.756, 0.793 and 0.747. Flat to within five points across a doubling of the arms, where the raw rate falls seven points and the ratio to chance nearly doubles.

The reason it is the right scale is that it is the only one of the four with both ends fixed: a guesser who knows nothing scores 0 and one who knows everything scores 1, at every K. The raw rate’s floor moves with K and its ceiling does not; the ratio’s ceiling moves with K and its floor does not. Each of them therefore reports a change in the number of arms as though it were a change in the rule, which is exactly the disagreement this field ends on.

Read on the fixed scale the three trials say one thing: a deterministic minimisation gives away about three quarters of what there was to give away, whatever the number of arms. That is the finding, and the two summaries the field reports are it seen through two moving floors.

It also explains why the manufactured gap barely moves — 0.759, 0.733, 0.711 — while the other two summaries move a great deal. The damage tracks the quantity that is flat, because what an investigator can extract depends on how much of the rule’s uncertainty they have removed and not on how many candidates were on the list.

The closed form the three-arm case does have

The two-arm expression 2δ(2g − 1) is described above as having no three-arm counterpart, and the counterpart is in this field’s own numbers, three settings of p and one change of score away.

Take the excess of the guess rate over chance, g − 1/3, and divide the manufactured bias by it. At p = 1 that is 0.650/0.5287 = 1.229. At p = 0.85, 0.507/0.4207 = 1.205. At p = 0.6, 0.270/0.2217 = 1.218. And the variance score, which is a different rule rather than a different setting of this one, gives 0.598/0.4907 = 1.219.

Four points, from two rules and three randomisation probabilities, on a straight line through the origin with a slope of 1.22 standard deviations per unit of excess — agreeing to within two per cent of themselves, against a standard error on the bias of 0.0074.

bias ≈ 1.22·(g − 1/3)

The two-arm formula is the same shape with a different slope: 4δ(g − ½) is 2.00(g − ½) at the enrolment advantage of half a standard deviation used here, and it gives 0.756 at g = 0.878 against the counted 0.759. So both cases are linear in the excess over chance and differ only in the constant, which is what the compensation being shared among K − 1 arms rather than one does to it.

That makes the design question quantitative in the currency the trial reports. Moving p from 1 to 0.85 removes 0.108 of guess rate and therefore about 0.13 standard deviations of manufactured effect; moving it to 0.6 removes 0.31 and about 0.38. An experimenter deciding where to sit on the trade between balance and predictability can now price one side of it without running anything.

The three numbers, and which to report

The field ends where it started: with three quantities that describe one property and do not agree about whether more arms are safer.

The guess rate is what an investigator achieves and it falls with the number of arms. It is the number usually quoted and it is the least informative of the three, because it cannot be read without the floor beneath it.

The ratio to chance is what the rule has given away, and it rises: 1.76×, 2.59×, 3.24×. It is the right measure of how much information the rule leaks and the wrong measure of how much harm that does.

The manufactured gap is the harm and it barely moves: 0.759, 0.733, 0.711. It is the one to report, because it is denominated in the units the trial’s own result is denominated in.

A field that reported only the first would conclude that adding arms protects a trial. A field that reported only the second would conclude the opposite. Both would be describing the same six numbers.

Five rules, three measures of what they left, 3 arms. 300 cohorts of 150 patients through 3 prognostic factors and 24 cells, every rule run on the same patients with the same coins. Each column is a different way of measuring the imbalance the rule left: the average range across factor levels, the average variance across them, and the worst cell of the cross-classification. The three minimising rules are close together and far from a uniform draw — and the ordering between them changes with the column, which is the finding: the variance leaves 0.934 on the range where the range leaves 0.963, so the score named after a quantity is not the one that controls it best. Blocks inside every cell are the opposite trade — the only rule that holds the cells, and the loosest margins of the four.
Fig. 7 And the balance those guesses were bought with, from the field’s first essay. Every rule in that table is guessable in proportion to how well it balances, and the two properties cannot be separated by any choice of score — only traded, by keeping a coin in the rule.
The conservatism does not go away with more arms. 800 trials of 150 patients at each number of arms, all under a true null. The lower curve is the unadjusted F test — an analysis of variance on the arms alone, which is what a trial report normally shows: 0.0% at 2 arms, 0.0% at 3 arms, 0.0% at 4 arms, 0.0% at 6 arms, where 5% is claimed throughout. The rule has already removed the variation the balanced factors carry, and the unadjusted denominator does not know that, so the statistic is too small. The upper curve is the same trials analysed with the factors in the model, which puts the level back at every number of arms. Nothing about the arithmetic gets better or worse with K; the loss is in what the analysis was not told.
Fig. 8 The analysis-side consequence of the same determinism, across widths: an unadjusted test on the floor at every number of arms. A rule tuned to be less guessable is a rule that balances less, and a rule that balances less leaves less for the analysis to be ignorant of — the two problems are one dial.

What is claimed, and what is not

The claim is the predictability of a covariate-adaptive rule with more than two arms: the guess rate, its ratio to chance, the bias a guessing investigator manufactures on the two summaries a multi-arm trial reports, and how all three move with the randomisation probability.

What stays out: guessers with partial information — knowing the factors but not the previous assignments, which is the realistic case and needs a model of what an investigator sees; deliberate manipulation of the covariates recorded, which is a different failure with a different repair; and the whole question of what predictability does to a trial’s credibility rather than to its estimate.

The boundary against the two-arm field is that it owns the mechanism and the closed form 2δ(2g − 1), and this owns what happens to both when there are more than two arms — where the closed form stops applying and the two summaries a trial reports stop agreeing with each other.

Four analyses of the same 3-arm trials, under a true null. 250 trials of 150 patients, 3 arms, minimisation with p = 1, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Only the first cell is wrong, and it is wrong in the direction that costs power rather than the one that manufactures findings: 0.0% where 5% is claimed. Either repair works — adjusting for what the rule balanced, or asking the rule what it would have done.
Fig. 9 And the other side of the same rule, from the previous essay: the analysis after a deterministic minimisation, where the unadjusted test rejects none of the true nulls it claims to reject 5% of. Determinism is what makes the rule both maximally balancing and maximally guessable.

The checks, and the refusal

Three claims are gated in this field’s library. The raw guess rate must fall as arms are added, and its ratio to chance must rise — the two statements that sound contradictory and are both true. The gap a guesser manufactures between the best and worst arm must be within a tenth of itself at four arms and at two, which is the finding that the damage does not fall. And the bias must be more than four standard errors from zero, so that none of it can be noise.

The refusal for this field is the unequal-target score two essays back. This essay adds none: its findings are measurements of a rule behaving exactly as designed, and the thing worth refusing here — an investigator’s enrolment decisions — is not a piece of arithmetic that can be fed bad input.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AllocationBlindingCovariate-adaptive randomisationError rateExperimental designMinimisationMonte CarloPermuted blocksPredictabilityRandomisationSelection biasStudy design