More arms than two

The analysis after three arms

An unadjusted analysis after a two-arm balancing rule rejects 0.6% of true nulls where it claims 5%. With three arms and a deterministic rule it rejects none at all — and the repair is the same repair, which is a sentence and a column in the model.

Worth reading first: Balancing what is known in advance · The experiments that could have happened.

The two-arm covariate-adaptive field ends with a finding that sounds harmless and is not. A trial allocated by minimisation and then analysed without mentioning the factors that were balanced rejects a true null 0.6% of the time where it claims 5%, and at full determinism 0.0%. The test is conservative, the direction is the safe-sounding one, and what it costs is power that nothing on the output reports.

With three arms the analysis is an F rather than a t, the comparison has two degrees of freedom rather than one, and the question is whether anything about that changes the picture.

The measurement

Four analyses of the same 3-arm trials, under a true null250 trials of 150 patients, 3 arms, minimisation with p = 0.85, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Only the first cell is wrong, and it is wrong in the direction that costs power rather than the one that manufactures findings: 0.0% where 5% is claimed. Either repair works — adjusting for what the rule balanced, or asking the rule what it would have done.the arms alone, against a table0.0%the factors in the model, against a table5.6%the arms alone, against the rule's own5.2%the factors in the model, against the rule's own6.8%5%, which all four claim250 trials of 150, 3 arms, 99 re-randomisations each0.0% · 5.6% · 5.2% · 6.8%
Fig. 1 Two hundred and fifty three-arm trials of a hundred and fifty patients, minimised on three prognostic factors, under a true null. Two statistics — an F on the arms alone, and an F on the arms after the balanced factors — against two reference distributions: the table, and the distribution the allocation rule itself generates.

The unadjusted F against its own table rejects 0.0% of true nulls. Not a reduced rate — none, in two hundred and fifty trials, and none in two thousand at a deterministic rule.

The other three cells are all at their nominal level: adjusting for the balanced factors gives 5.6%, and the rule’s own reference distribution gives 5.2% for the unadjusted statistic and 6.8% for the adjusted one.

The gap between the first cell and the other three is the whole essay, and its size is worth pausing on: a test that claims a one-in-twenty false-positive rate delivered none at all in two hundred and fifty trials, and the three repairs available are a column in a model, a re-run of the allocation rule, or both.

Why it happens, in one sentence and then in three

The rule has removed the variation the balanced factors carry, and the unadjusted denominator does not know it.

Spelled out: the F statistic divides the variation between arms by the variation within them. Under minimisation, the between-arm variation has been deliberately stripped of everything the factors could contribute, so the numerator is smaller than a random allocation would produce. The denominator is unchanged — the within-arm variation still contains all of the factors’ contribution, since nothing about the analysis knows they exist. A statistic with a shrunken numerator and a full denominator is too small, and a test that compares it against a table computed for randomly allocated data rejects too rarely.

Everything in that argument turns on how much of the outcome the balanced factors explain, and that is measurable.

The conservatism does not go away with more arms. 800 trials of 150 patients at each number of arms, all under a true null. The lower curve is the unadjusted F test — an analysis of variance on the arms alone, which is what a trial report normally shows: 4.4% at 2 arms, 7.0% at 3 arms, 5.6% at 4 arms, 5.6% at 6 arms, where 5% is claimed throughout. The rule has already removed the variation the balanced factors carry, and the unadjusted denominator does not know that, so the statistic is too small. The upper curve is the same trials analysed with the factors in the model, which puts the level back at every number of arms. Nothing about the arithmetic gets better or worse with K; the loss is in what the analysis was not told.
Fig. 2 The same measurement when the balanced factors have nothing to do with the outcome. Both curves sit at their nominal level at every number of arms: with nothing for the rule to remove, there is nothing for the analysis to be ignorant of.

At a deterministic rule and two thousand trials, the unadjusted test rejects 4.85% when the factors are unrelated to the outcome, 3.30% when they carry a quarter of the effect, 1.25% at half, and 0.00% at one. The adjusted test rejects 4.80% throughout. The conservatism tracks how much the balancing was worth, which is what says the mechanism is the one above rather than something about the rule’s arithmetic.

More arms does not change it

The conservatism does not go away with more arms. 800 trials of 150 patients at each number of arms, all under a true null. The lower curve is the unadjusted F test — an analysis of variance on the arms alone, which is what a trial report normally shows: 0.0% at 2 arms, 0.0% at 3 arms, 0.0% at 4 arms, 0.0% at 6 arms, where 5% is claimed throughout. The rule has already removed the variation the balanced factors carry, and the unadjusted denominator does not know that, so the statistic is too small. The upper curve is the same trials analysed with the factors in the model, which puts the level back at every number of arms. Nothing about the arithmetic gets better or worse with K; the loss is in what the analysis was not told.
Fig. 3 Two, three, four and six arms, all under a true null with strongly prognostic factors. The unadjusted curve is on the floor at every width; the adjusted one is at its nominal level at every width.

At a deterministic rule the unadjusted F rejects 0.10% at two arms and 0.00% at three, four and six. The adjusted F rejects 5.40%, 4.80%, 5.25% and 5.60%.

Nothing about the number of arms matters, which is worth stating because it would have been reasonable to expect otherwise: with K arms there are K − 1 contrasts to balance and the rule is spread thinner, so the balancing might have been expected to bite less. It does not, because the mechanism is about the denominator — the within-arm variation containing the factors’ effect — and that is the same whatever the numerator’s degrees of freedom.

What the rule has actually done to the data

It is worth looking at the mechanism from the data’s side rather than the statistic’s, because the statement the rule removed variation the analysis does not know about can be checked directly.

The three prognostic factors here explain a substantial share of the outcome by construction, and the allocation rule has made the arms nearly identical on them: the standardised difference in average prognosis between the best and worst arm is 0.0513 under minimisation against 0.2398 under a uniform draw. So the between-arm variation that a random allocation would have produced by chance — one arm happening to hold more of the good-prognosis patients — has been removed on purpose.

What has not been removed is any of the within-arm variation. Every arm still contains a mixture of prognoses spanning the full range, and the F statistic’s denominator is estimated from exactly that mixture. The rule has shrunk the numerator’s null distribution and left the denominator alone.

An adjusted analysis takes the factors out of both, which is why it is the repair and not a refinement.

What the conservatism costs

A test that never rejects a true null sounds like a test that has bought safety. It has bought nothing: the level was already correct in the analysis that adjusts, and what has been given up is the ability to detect an effect that is there.

Four analyses of the same 3-arm trials, against a real effect. 250 trials of 150 patients, 3 arms, minimisation with p = 0.85, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Against an effect of 0.5 the comparison is power rather than level, and the ordering is the same one the level implies: 70.4% adjusted against 10.4% unadjusted.
Fig. 4 The same four analyses against a real effect of half a standard deviation. The unadjusted table analysis finds it 10.4% of the time. The adjusted one finds it 70.4% of the time. The two exact analyses are at 61.2% and 67.2%.

Sixty points of power, on the same trials, from putting the balanced factors into the model. That is not a refinement — it is the difference between a trial that answers its question and a trial that does not, and the only thing separating the two analyses is which columns are in the design matrix.

The exact routes are worth reading too. The unadjusted statistic against the rule’s own reference distribution has 61.2% power: the statistic is still the weak one, but it is being compared against the distribution that statistic actually has under this rule rather than against a table for a different design, and that alone recovers most of the loss. The adjusted statistic against the same reference distribution has 67.2%.

Four analyses of the same 3-arm trials, against a real effect. 250 trials of 90 patients, 3 arms, minimisation with p = 0.85, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Against an effect of 0.5 the comparison is power rather than level, and the ordering is the same one the level implies: 48.8% adjusted against 5.6% unadjusted.
Fig. 5 The same power comparison on ninety patients rather than a hundred and fifty. Every rate falls and the gap between the adjusted and unadjusted analyses does not close — a smaller trial is exactly the one that cannot afford to give away sixty points.

Why the exact route is cheap here

The randomisation test in this essay costs a few hundred re-runs of the allocation rule and no assumptions. That is worth one paragraph of comparison, because the same construction in the outcome-adaptive field was expensive enough to change what was recommended there.

A reference distribution built by re-running the rule needs the rule to be re-runnable without knowing anything the trial is trying to estimate. Here it is: the rule reads covariates and the assignments already made, both of which are fixed when the outcomes are held fixed, so the whole distribution is generated from quantities the trial already has. In the outcome-adaptive field the rule reads the responses, so re-running it requires either the responses to move — which changes the null — or a nuisance parameter to be supplied.

That is the whole reason the exactness costs two and a half points of power in the covariate case and nineteen in the outcome case, and the K-arm generalisation changes none of it.

The two repairs, and when each is available

Two things fix the level and they fix it in different ways.

Adjust. Put the balanced factors in the model. The denominator then excludes what the rule removed, the numerator and denominator are talking about the same thing again, and the level is restored. It costs the degrees of freedom the factors use — nine here, of a hundred and fifty — and it requires that the factors are recorded, which they are, since the rule used them.

Condition on the rule. Hold the outcomes fixed, re-run the allocation rule many times, and compare the observed statistic against the distribution of statistics it produces. That is the exact route the covariate-adaptive field built, applied unchanged: nothing about it cares how many arms there are, and the reference distribution depends on no quantity the trial is estimating.

The second is more expensive and more general. It works when the model is wrong, when the factors enter non-linearly, and when nobody can say what adjusting would even mean. What it does not do is recover the power: the unadjusted statistic under its own reference distribution is at 61.2% against the adjusted statistic’s 70.4%, because a conditioning argument can fix a level and cannot manufacture information the statistic is not using.

How much randomisation the rule keeps does not rescue it either

The measurements above are at p = 0.85 and at p = 1, and the difference between them is smaller than it looks.

At full determinism the rule is a function of the covariates alone and the conservatism is total: 0.00%. At p = 0.85 the rule follows its own preference most of the time and randomises the rest, which puts a little of the factors’ variation back into the between-arm comparison — and the unadjusted test still rejects 0.0% of true nulls in two hundred and fifty trials.

The reason the coin does not help is that it is a coin on which arm, not on how balanced. An arrival sent to a non-preferred arm makes that factor level slightly less balanced, and the next few arrivals are then steered to repair it. The rule’s randomisation buys unpredictability, which is what the next essay is about, and it does not buy back the between-arm variation the analysis is missing.

Four analyses of the same 3-arm trials, under a true null. 250 trials of 150 patients, 3 arms, minimisation with p = 0.6, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Only the first cell is wrong, and it is wrong in the direction that costs power rather than the one that manufactures findings: 0.8% where 5% is claimed. Either repair works — adjusting for what the rule balanced, or asking the rule what it would have done.
Fig. 6 The four analyses under a rule that follows its own preference only three times in five: the unadjusted test rejects 0.8% where it claims 5%, and the other three cells are at 4.4%, 4.0% and 3.2%. Loosening the rule recovers less than a point of the level, so p is not a lever on this problem at any setting a trial would use.

The two repairs are very nearly the same repair

The four power figures at three arms — 10.4%, 70.4%, 61.2% and 67.2% — are a complete two-by-two, since one factor is whether the balanced factors are in the model and the other is whether the statistic is read against a table or against the rule’s own distribution. Read as a factorial they say something neither pair says alone.

Adjusting is worth 60.0 points when the statistic is read against the table, and 6.0 points when it is read against the rule’s own distribution. Conditioning is worth 50.8 points on the unadjusted statistic and −3.2 on the adjusted one. Either repair alone recovers most of what the unadjusted table analysis gave away — 100% and 85% of it respectively — and applying both recovers 94.7%, which is less than adjusting alone.

The two repairs overlap almost completely, and the small amount they do not overlap on is a cost rather than a gain. That is worth saying plainly because the instinct on being offered two fixes for one defect is to take both. Here taking both is worse than taking the better one, by the three points that exactness costs, and the reason is that they are addressing the identical deficiency from two sides: the analysis is comparing a statistic whose numerator the rule shrank against a distribution that does not know it was shrunk. Adjusting unshrinks the numerator; conditioning replaces the distribution. Once either is done there is nothing left for the other to correct.

The practical ordering follows from the sizes rather than from principle. Adjust when the model can be written down, because it is the cheaper repair and the better one. Condition when it cannot, and accept three points. Doing both is defensible only as a check that the two agree — which, at 70.4% and 67.2%, they do.

The level is flat in the arms and the power is not

Two of this essay’s sweeps point in different directions and putting them together is the useful part.

The level does not depend on the number of arms: the unadjusted F rejects 0.10% at two arms and 0.00% at three, four and six, and the adjusted one sits between 4.80% and 5.60% throughout. The mechanism is in the denominator, and the denominator does not know how many arms there are.

The power is not flat at all. Going from three arms to four on the same hundred and fifty patients takes the adjusted analysis from 70.4% to 57.2% — a fall of 13.2 points, or a fifth of what it had — and the unadjusted analysis from 10.4% to 3.2%, which is a fall to less than a third. The weak analysis degrades more than three times as fast as the strong one as arms are added.

So the cost of not adjusting is not a fixed penalty; it compounds with the width of the trial. A six-arm trial analysed without its balanced factors is spreading the same patients over more comparisons and handing back most of what the balancing bought, and the two losses multiply rather than add. The rule of thumb worth carrying is that the more arms a trial has, the less it can afford the analysis it is most likely to be given.

What a three-arm trial should report

The list is short and each item is a consequence of a number above.

Say which factors the rule balanced, because the analysis needs them and a reader needs to know whether they are in it.

Put them in the model. Sixty points of power at the setting measured here, and no cost beyond nine degrees of freedom.

If the model is doubtful, condition on the rule instead. It costs a few hundred re-runs of the allocation and holds the level exactly, and it is the only option that needs no modelling assumption at all.

Do not report the unadjusted test as conservative and therefore safe. It is conservative and therefore weak, and the two words describe the same number from the point of view of two different people.

Four analyses of the same 4-arm trials, against a real effect. 250 trials of 150 patients, 4 arms, minimisation with p = 0.85, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Against an effect of 0.5 the comparison is power rather than level, and the ordering is the same one the level implies: 57.2% adjusted against 3.2% unadjusted.
Fig. 7 And the same four analyses at four arms against the same effect: 3.2%, 57.2%, 47.2%, 58.0%. Every rate is lower than at three arms, because the same patients are spread over more arms and the effect is being detected against two more degrees of freedom — and the ordering and the gaps between the four analyses are unchanged.

The same shape, three fields along

This is the third time a rule that reads something before deciding has changed what the analysis afterwards means, and the three results together are more informative than any of them alone.

Reading outcomes to allocate breaks the error rate upwards: a response-adaptive rule inflates the rejection rate and biases the estimate of the arm it favoured, and repairing it costs nineteen points of power.

Reading covariates to allocate breaks the error rate downwards: the level falls to zero and the cost is power, and repairing it costs nine degrees of freedom.

Reading responses to place experimental runs breaks almost nothing, because the design’s objective is the analysis’s objective.

The pattern is not about how much the rule knows. It is about what the rule’s objective has to do with the analysis’s: opposed, orthogonal, or identical, and the three cases have three different signs.

What a guesser gets, and what a guesser gets for nothing. 600 trials of 150 patients under a fully deterministic rule, with an investigator who knows the rule, the factors and every assignment so far. The guess rate falls with the number of arms — 87.8% at 2, 86.2% at 3, 81.0% at 4 — which reads like a trial getting safer and is not: what a guesser can trade on is the excess over the 50%, 33%, 25% they would get by naming an arm at random, and that goes the other way, from 1.76× chance at two arms to 3.24× at 4. The gap a guesser manufactures between the best and worst arm under a true null is 0.759, 0.733, 0.711 standard deviations — nearly unchanged.
Fig. 8 The other consequence of a rule that reads covariates, which the next essay takes up: how much of the allocation a guesser can work out. Balance and predictability are the same property seen from two sides, and the analysis problem in this essay is the price of the first while the next essay is the price of the second.

What is claimed, and what is not

The claim is the analysis after a covariate-adaptive rule with more than two arms: the unadjusted F’s conservatism, its dependence on how much the balanced factors explain, its independence of the number of arms, the power the adjustment recovers, and the rule’s own reference distribution as the assumption-free alternative.

What stays out and is named: pairwise comparisons after the F, where the multiplicity structure of several arms against one control sits on top of everything here; covariate adjustment for factors the rule did not balance, which is a different and older question; and non-linear outcome models, where the conservatism has the same source and the arithmetic of the adjustment is not a partial F.

The boundary against the two-arm field is that it owns the finding and this owns its behaviour as the number of arms and the strength of the factors vary. Nothing here contradicts it; the numbers are its numbers with two degrees of freedom instead of one.

What this does to a trial’s stated sample size

The practical consequence is a number nobody computes, and it is worth computing once.

A trial powered at 80% under the assumption that its analysis behaves as advertised, and then analysed without the balanced factors, is not running at 80%. At the setting measured here — a hundred and fifty patients, three arms, an effect of half a standard deviation — the adjusted analysis has 70.4% power and the unadjusted 10.4%. To reach the adjusted analysis’s power with the unadjusted one would take several times the patients, and no protocol amendment would be required to notice, because nothing about the trial has visibly changed.

That is the strongest version of the two-arm field’s point about conservatism being expensive. A level that is too low is not a safety margin. It is a sample size calculation quietly multiplied.

The checks, and the refusal

Three claims are gated in this field’s library. The unadjusted F after a deterministic rule must reject far below its nominal level, at strongly prognostic factors. Adjusting must restore the level to within two points of 5%. And the conservatism must vanish when the balanced factors are unrelated to the outcome, which is the mechanism asserted rather than described — if the rate stayed low there, the explanation in this essay would be wrong.

The exact route is gated separately: both statistics against the rule’s own reference distribution must hold the level to within two and a half points, over two hundred re-randomisation tests of a hundred and ninety-nine draws each.

The field’s refusal is the unequal-target score of the previous essay.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Conservative intervalCovariate-adaptive randomisationCovariate adjustmentDegrees of freedomError rateExact testMinimisationMonte CarloNull hypothesisRandomisation testReference distributionStatistical power