Guessing one arm in three
Worth reading first: Balancing what is known in advance · Randomisation is not balance.
Balance and predictability are one property seen from two sides. A rule that corrects an imbalance is a rule whose next move can be worked out by anybody who can see the imbalance, and the two-arm field measures what that is worth to an investigator who does the working out: a deterministic minimisation is guessable 87.5% of the time, and a guessing investigator extracts 0.75 of a standard deviation of treatment effect where none exists.
With three arms the guesser has more ways to be wrong. Whether that makes the trial safer is a question with three answers, and they do not agree.
Three quantities, three directions
The raw guess rate falls. It is 87.8% at two arms, 86.2% at three and 81.0% at four. More arms means more ways for a guess to miss, and the rule’s preferred arm is not always unique.
The advantage over chance rises. A guesser naming an arm at random is right half the time with two arms and a third of the time with three, so the ratio of what they achieve to what they would get for nothing runs 1.76×, 2.59×, 3.24×. By that measure a four-arm trial is more predictable than a two-arm one, not less.
The damage is nearly unchanged. The quantity a three-arm trial actually reports is the gap between its best and worst arm, and a guessing investigator manufactures 0.759, 0.733 and 0.711 standard deviations of it at two, three and four arms — under a true null, from a mechanism that breaks no rule.
The three statements are all true and they answer three different questions. Which one matters depends on what the guesser is trying to do, and the third is the one that reaches the reader of the trial.
Where the two-arm formula stops applying
The two-arm field’s closed form is worth carrying over as far as it goes, because where it fails is informative.
With two arms, a correct guess enrols a patient whose prognosis is δ better than average into the favoured arm, and by symmetry an incorrect guess enrols one δ worse into the other. The arms move in opposite directions, so the difference in means moves by twice δ times the excess of the guess rate over a half: 2δ(2g − 1), and the factor of two is the part that is easy to leave out.
With three arms there is no symmetry to exploit. A correct guess moves one arm up; the compensation is spread over two others, so each of them moves down by half as much, and the effect on the favoured arm against the average of the rest is not the same quantity as the effect on the largest gap between any two arms. Both are counted here rather than derived, and they differ: 0.650 against 0.733 at three arms, from the same trials.
That divergence is why the field reports two summaries. A trial with several arms has no single treatment effect, and a bias that manufactures one contrast does not manufacture all of them equally.
How the damage is done, which involves no cheating
The mechanism is the two-arm field’s, unchanged: an investigator who can predict the next assignment does not tamper with it. They decide who to enrol next, which is the one thing an allocation rule cannot control.
When the guess says the next patient goes to the arm the investigator favours, they enrol somebody whose prognosis is a little better than average. When it says otherwise, they wait, or enrol somebody a little worse. Nothing about the randomisation is violated, no assignment is changed, and the arms end up differing in prognosis by an amount that is entirely a function of the guess rate.
With two arms the arithmetic is clean — a correct guess moves one arm up and the other down, so the effect is 2δ(2g − 1) with g the guess rate — and with three arms it is not, because a correct guess moves one arm up and shares the compensation among the others. What is counted here rather than derived is the effect on the two summaries a trial reports: the favoured arm against the average of the rest, and the largest gap between any two arms.
At three arms and a deterministic rule those are 0.650 and 0.733 standard deviations, against a standard error of 0.0074, so none of it is noise.
What a guesser is assumed to know, and what they cannot
The guess rate above is an upper bound with a specific construction behind it, and both halves matter.
The guesser is given the rule, the score, the factors, and every assignment made so far. From those they compute the same margin score the rule computes and name whichever arm it prefers. That is what an unblinded investigator enrolling into an open trial has, which is why it is the right assumption for a bound.
What they are not given is the rule’s own coin. Where two or more arms are tied on the score the rule chooses between them at random, and the guesser can only name one of them; the earlier version of this measurement handed the guesser the realised tie-break and reported every rule as 100% guessable at p = 1, which is a measurement of a rule nobody runs. The ties are the whole of the shortfall from certainty: at three arms a deterministic rule is guessed 86.2% of the time, and the remaining 13.8% is arrivals on which the score could not separate the candidates.
That has a design consequence worth naming, and it is the rare one that costs nothing. A score that ties more often is less guessable, and the variance score — the one that controls balance better — is indifferent on 32.9% of arrivals against the range score’s 26.2%. Measured directly, it is guessed 82.4% of the time against the range score’s 86.2%, and the bias a guesser extracts falls from 0.650 to 0.598. The score that balances best is also the one that hides most, so the two undeclared choices of this field point the same way.
What the randomisation probability buys
Minimisation is normally run with a probability p of following its own preference — the standard choice is 0.8 — precisely to make this harder.
At three arms the guess rate is 86.2% at p = 1, 75.4% at p = 0.85 and 55.5% at p = 0.6. The bias to the favoured arm follows: 0.650, 0.507, 0.270. Buying unpredictability with randomisation works, and it works in proportion.
What it costs is balance, and the two-arm field measured the trade with the conclusion that there is no knee in it: the guess rate is close to linear in p and the imbalance close to geometric, so every increment of one buys a proportion of the other and where to sit is a judgement rather than an optimum. Nothing about a third arm changes that shape.
Why more arms are more predictable relative to chance
The ratio going up while the rate goes down is worth explaining rather than reporting, because it is the one number in this field that is genuinely counterintuitive.
A guesser facing K arms has a free floor of 1/K, and the rule’s preference is a ranking rather than a single arm. With two arms, the rule is nearly always indifferent or decisive, and a decisive rule is guessed with certainty — hence 87.8%, with the shortfall coming from arrivals where two arms tie and the rule tosses a coin the guesser cannot see.
With four arms ties are more common — the score is coarser relative to the number of candidates — so the raw rate falls to 81.0%. But the floor has fallen much faster, from 50% to 25%, so what the guesser knows that they did not know before has grown: they have eliminated three arms rather than one.
The quantity that matters for the damage is neither of those, which is why the gap between the best and worst arm barely moves: manufacturing a difference between two named arms requires knowing which of them the next patient is going to, and that is roughly as easy at four arms as at two.
What the trial actually loses
A manufactured difference of seven tenths of a standard deviation under a true null is easier to judge with something to compare it against.
The trials in this field are powered to detect an effect of about half a standard deviation with seventy percent probability, so a guessing investigator can generate an effect larger than the one the trial was designed to find, with no data fabricated, no assignment altered and no protocol violated. The p-value that results is a correct p-value for the data that were collected; what is wrong is that the patients in the arms are not comparable, which is the one thing randomisation was supposed to guarantee.
That is why predictability is a design problem rather than an analysis problem. No adjustment recovers it — the covariates that would explain the difference are the ones the investigator chose on and did not record — and the exact test of the previous essay does not touch it either, because the allocation rule really did produce that allocation. The reference distribution is right and the patients are wrong.
What a trial can do about it
Four things, of which only the first two are about the allocation rule.
Keep a coin in the rule. p = 0.8 rather than 1 costs a little balance and takes the guess rate from 86.2% to about three quarters, and the bias with it.
Do not publish the score. The guess rate measured here assumes the guesser knows the rule, the factors and every assignment so far. The last of those is what an unblinded investigator has, and concealment of the assignments already made is worth more than any amount of randomisation inside the rule.
Blind the treatment. The bias enters through who is enrolled after the guess, and an investigator who cannot tell which arm a patient went to cannot maintain the count that makes the guess possible.
Report the balance. A trial whose baseline table is suspiciously good on the balanced factors and suspiciously uneven on prognosis outside them is a trial worth asking about, and the standardised prognosis gap in the previous essays is what that comparison would look like.
The fourth summary, which is flat
Three quantities disagree about whether more arms are safer, and there is a fourth that neither the rate nor the ratio is: the share of the guesser’s uncertainty that the rule removed, which is (g − 1/K)/(1 − 1/K).
At two, three and four arms that is 0.756, 0.793 and 0.747. Flat to within five points across a doubling of the arms, where the raw rate falls seven points and the ratio to chance nearly doubles.
The reason it is the right scale is that it is the only one of the four with both ends fixed: a guesser who knows nothing scores 0 and one who knows everything scores 1, at every K. The raw rate’s floor moves with K and its ceiling does not; the ratio’s ceiling moves with K and its floor does not. Each of them therefore reports a change in the number of arms as though it were a change in the rule, which is exactly the disagreement this field ends on.
Read on the fixed scale the three trials say one thing: a deterministic minimisation gives away about three quarters of what there was to give away, whatever the number of arms. That is the finding, and the two summaries the field reports are it seen through two moving floors.
It also explains why the manufactured gap barely moves — 0.759, 0.733, 0.711 — while the other two summaries move a great deal. The damage tracks the quantity that is flat, because what an investigator can extract depends on how much of the rule’s uncertainty they have removed and not on how many candidates were on the list.
The closed form the three-arm case does have
The two-arm expression 2δ(2g − 1) is described above as having no three-arm counterpart, and the counterpart is in this field’s own numbers, three settings of p and one change of score away.
Take the excess of the guess rate over chance, g − 1/3, and divide the manufactured bias by it. At p = 1 that is 0.650/0.5287 = 1.229. At p = 0.85, 0.507/0.4207 = 1.205. At p = 0.6, 0.270/0.2217 = 1.218. And the variance score, which is a different rule rather than a different setting of this one, gives 0.598/0.4907 = 1.219.
Four points, from two rules and three randomisation probabilities, on a straight line through the origin with a slope of 1.22 standard deviations per unit of excess — agreeing to within two per cent of themselves, against a standard error on the bias of 0.0074.
bias ≈ 1.22·(g − 1/3)
The two-arm formula is the same shape with a different slope: 4δ(g − ½) is 2.00(g − ½) at the enrolment advantage of half a standard deviation used here, and it gives 0.756 at g = 0.878 against the counted 0.759. So both cases are linear in the excess over chance and differ only in the constant, which is what the compensation being shared among K − 1 arms rather than one does to it.
That makes the design question quantitative in the currency the trial reports. Moving p from 1 to 0.85 removes 0.108 of guess rate and therefore about 0.13 standard deviations of manufactured effect; moving it to 0.6 removes 0.31 and about 0.38. An experimenter deciding where to sit on the trade between balance and predictability can now price one side of it without running anything.
The three numbers, and which to report
The field ends where it started: with three quantities that describe one property and do not agree about whether more arms are safer.
The guess rate is what an investigator achieves and it falls with the number of arms. It is the number usually quoted and it is the least informative of the three, because it cannot be read without the floor beneath it.
The ratio to chance is what the rule has given away, and it rises: 1.76×, 2.59×, 3.24×. It is the right measure of how much information the rule leaks and the wrong measure of how much harm that does.
The manufactured gap is the harm and it barely moves: 0.759, 0.733, 0.711. It is the one to report, because it is denominated in the units the trial’s own result is denominated in.
A field that reported only the first would conclude that adding arms protects a trial. A field that reported only the second would conclude the opposite. Both would be describing the same six numbers.
What is claimed, and what is not
The claim is the predictability of a covariate-adaptive rule with more than two arms: the guess rate, its ratio to chance, the bias a guessing investigator manufactures on the two summaries a multi-arm trial reports, and how all three move with the randomisation probability.
What stays out: guessers with partial information — knowing the factors but not the previous assignments, which is the realistic case and needs a model of what an investigator sees; deliberate manipulation of the covariates recorded, which is a different failure with a different repair; and the whole question of what predictability does to a trial’s credibility rather than to its estimate.
The boundary against the two-arm field is that it owns the mechanism and the closed form 2δ(2g − 1), and this owns what happens to both when there are more than two arms — where the closed form stops applying and the two summaries a trial reports stop agreeing with each other.
The checks, and the refusal
Three claims are gated in this field’s library. The raw guess rate must fall as arms are added, and its ratio to chance must rise — the two statements that sound contradictory and are both true. The gap a guesser manufactures between the best and worst arm must be within a tenth of itself at four arms and at two, which is the finding that the damage does not fall. And the bias must be more than four standard errors from zero, so that none of it can be noise.
The refusal for this field is the unequal-target score two essays back. This essay adds none: its findings are measurements of a rule behaving exactly as designed, and the thing worth refusing here — an investigator’s enrolment decisions — is not a piece of arithmetic that can be fed bad input.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A covariate with no levels — both name covariate-adaptive randomisation, experimental design, minimisation, monte carlo, randomisation
- The rule that reads the number — both name covariate-adaptive randomisation, experimental design, minimisation, monte carlo, randomisation
- What the balanced trial is worth — both name covariate-adaptive randomisation, error rate, experimental design, monte carlo, randomisation
- A proposal that moves more than two units — both name experimental design, minimisation, monte carlo, randomisation
- Balancing more than one number — both name covariate-adaptive randomisation, experimental design, monte carlo, randomisation
- The analysis has to know the rule — both name covariate-adaptive randomisation, error rate, minimisation, permuted blocks
Named objects
A flat tag is an object no other essay names yet.
AllocationBlindingCovariate-adaptive randomisationError rateExperimental designMinimisationMonte CarloPermuted blocksPredictabilityRandomisationSelection biasStudy design