The analysis after three arms
Worth reading first: Balancing what is known in advance · The experiments that could have happened.
The two-arm covariate-adaptive field ends with a finding that sounds harmless and is not. A trial allocated by minimisation and then analysed without mentioning the factors that were balanced rejects a true null 0.6% of the time where it claims 5%, and at full determinism 0.0%. The test is conservative, the direction is the safe-sounding one, and what it costs is power that nothing on the output reports.
With three arms the analysis is an F rather than a t, the comparison has two degrees of freedom rather than one, and the question is whether anything about that changes the picture.
The measurement
The unadjusted F against its own table rejects 0.0% of true nulls. Not a reduced rate — none, in two hundred and fifty trials, and none in two thousand at a deterministic rule.
The other three cells are all at their nominal level: adjusting for the balanced factors gives 5.6%, and the rule’s own reference distribution gives 5.2% for the unadjusted statistic and 6.8% for the adjusted one.
The gap between the first cell and the other three is the whole essay, and its size is worth pausing on: a test that claims a one-in-twenty false-positive rate delivered none at all in two hundred and fifty trials, and the three repairs available are a column in a model, a re-run of the allocation rule, or both.
Why it happens, in one sentence and then in three
The rule has removed the variation the balanced factors carry, and the unadjusted denominator does not know it.
Spelled out: the F statistic divides the variation between arms by the variation within them. Under minimisation, the between-arm variation has been deliberately stripped of everything the factors could contribute, so the numerator is smaller than a random allocation would produce. The denominator is unchanged — the within-arm variation still contains all of the factors’ contribution, since nothing about the analysis knows they exist. A statistic with a shrunken numerator and a full denominator is too small, and a test that compares it against a table computed for randomly allocated data rejects too rarely.
Everything in that argument turns on how much of the outcome the balanced factors explain, and that is measurable.
At a deterministic rule and two thousand trials, the unadjusted test rejects 4.85% when the factors are unrelated to the outcome, 3.30% when they carry a quarter of the effect, 1.25% at half, and 0.00% at one. The adjusted test rejects 4.80% throughout. The conservatism tracks how much the balancing was worth, which is what says the mechanism is the one above rather than something about the rule’s arithmetic.
More arms does not change it
At a deterministic rule the unadjusted F rejects 0.10% at two arms and 0.00% at three, four and six. The adjusted F rejects 5.40%, 4.80%, 5.25% and 5.60%.
Nothing about the number of arms matters, which is worth stating because it would have been reasonable to expect otherwise: with K arms there are K − 1 contrasts to balance and the rule is spread thinner, so the balancing might have been expected to bite less. It does not, because the mechanism is about the denominator — the within-arm variation containing the factors’ effect — and that is the same whatever the numerator’s degrees of freedom.
What the rule has actually done to the data
It is worth looking at the mechanism from the data’s side rather than the statistic’s, because the statement the rule removed variation the analysis does not know about can be checked directly.
The three prognostic factors here explain a substantial share of the outcome by construction, and the allocation rule has made the arms nearly identical on them: the standardised difference in average prognosis between the best and worst arm is 0.0513 under minimisation against 0.2398 under a uniform draw. So the between-arm variation that a random allocation would have produced by chance — one arm happening to hold more of the good-prognosis patients — has been removed on purpose.
What has not been removed is any of the within-arm variation. Every arm still contains a mixture of prognoses spanning the full range, and the F statistic’s denominator is estimated from exactly that mixture. The rule has shrunk the numerator’s null distribution and left the denominator alone.
An adjusted analysis takes the factors out of both, which is why it is the repair and not a refinement.
What the conservatism costs
A test that never rejects a true null sounds like a test that has bought safety. It has bought nothing: the level was already correct in the analysis that adjusts, and what has been given up is the ability to detect an effect that is there.
Sixty points of power, on the same trials, from putting the balanced factors into the model. That is not a refinement — it is the difference between a trial that answers its question and a trial that does not, and the only thing separating the two analyses is which columns are in the design matrix.
The exact routes are worth reading too. The unadjusted statistic against the rule’s own reference distribution has 61.2% power: the statistic is still the weak one, but it is being compared against the distribution that statistic actually has under this rule rather than against a table for a different design, and that alone recovers most of the loss. The adjusted statistic against the same reference distribution has 67.2%.
Why the exact route is cheap here
The randomisation test in this essay costs a few hundred re-runs of the allocation rule and no assumptions. That is worth one paragraph of comparison, because the same construction in the outcome-adaptive field was expensive enough to change what was recommended there.
A reference distribution built by re-running the rule needs the rule to be re-runnable without knowing anything the trial is trying to estimate. Here it is: the rule reads covariates and the assignments already made, both of which are fixed when the outcomes are held fixed, so the whole distribution is generated from quantities the trial already has. In the outcome-adaptive field the rule reads the responses, so re-running it requires either the responses to move — which changes the null — or a nuisance parameter to be supplied.
That is the whole reason the exactness costs two and a half points of power in the covariate case and nineteen in the outcome case, and the K-arm generalisation changes none of it.
The two repairs, and when each is available
Two things fix the level and they fix it in different ways.
Adjust. Put the balanced factors in the model. The denominator then excludes what the rule removed, the numerator and denominator are talking about the same thing again, and the level is restored. It costs the degrees of freedom the factors use — nine here, of a hundred and fifty — and it requires that the factors are recorded, which they are, since the rule used them.
Condition on the rule. Hold the outcomes fixed, re-run the allocation rule many times, and compare the observed statistic against the distribution of statistics it produces. That is the exact route the covariate-adaptive field built, applied unchanged: nothing about it cares how many arms there are, and the reference distribution depends on no quantity the trial is estimating.
The second is more expensive and more general. It works when the model is wrong, when the factors enter non-linearly, and when nobody can say what adjusting would even mean. What it does not do is recover the power: the unadjusted statistic under its own reference distribution is at 61.2% against the adjusted statistic’s 70.4%, because a conditioning argument can fix a level and cannot manufacture information the statistic is not using.
How much randomisation the rule keeps does not rescue it either
The measurements above are at p = 0.85 and at p = 1, and the difference between them is smaller than it looks.
At full determinism the rule is a function of the covariates alone and the conservatism is total: 0.00%. At p = 0.85 the rule follows its own preference most of the time and randomises the rest, which puts a little of the factors’ variation back into the between-arm comparison — and the unadjusted test still rejects 0.0% of true nulls in two hundred and fifty trials.
The reason the coin does not help is that it is a coin on which arm, not on how balanced. An arrival sent to a non-preferred arm makes that factor level slightly less balanced, and the next few arrivals are then steered to repair it. The rule’s randomisation buys unpredictability, which is what the next essay is about, and it does not buy back the between-arm variation the analysis is missing.
The two repairs are very nearly the same repair
The four power figures at three arms — 10.4%, 70.4%, 61.2% and 67.2% — are a complete two-by-two, since one factor is whether the balanced factors are in the model and the other is whether the statistic is read against a table or against the rule’s own distribution. Read as a factorial they say something neither pair says alone.
Adjusting is worth 60.0 points when the statistic is read against the table, and 6.0 points when it is read against the rule’s own distribution. Conditioning is worth 50.8 points on the unadjusted statistic and −3.2 on the adjusted one. Either repair alone recovers most of what the unadjusted table analysis gave away — 100% and 85% of it respectively — and applying both recovers 94.7%, which is less than adjusting alone.
The two repairs overlap almost completely, and the small amount they do not overlap on is a cost rather than a gain. That is worth saying plainly because the instinct on being offered two fixes for one defect is to take both. Here taking both is worse than taking the better one, by the three points that exactness costs, and the reason is that they are addressing the identical deficiency from two sides: the analysis is comparing a statistic whose numerator the rule shrank against a distribution that does not know it was shrunk. Adjusting unshrinks the numerator; conditioning replaces the distribution. Once either is done there is nothing left for the other to correct.
The practical ordering follows from the sizes rather than from principle. Adjust when the model can be written down, because it is the cheaper repair and the better one. Condition when it cannot, and accept three points. Doing both is defensible only as a check that the two agree — which, at 70.4% and 67.2%, they do.
The level is flat in the arms and the power is not
Two of this essay’s sweeps point in different directions and putting them together is the useful part.
The level does not depend on the number of arms: the unadjusted F rejects 0.10% at two arms and 0.00% at three, four and six, and the adjusted one sits between 4.80% and 5.60% throughout. The mechanism is in the denominator, and the denominator does not know how many arms there are.
The power is not flat at all. Going from three arms to four on the same hundred and fifty patients takes the adjusted analysis from 70.4% to 57.2% — a fall of 13.2 points, or a fifth of what it had — and the unadjusted analysis from 10.4% to 3.2%, which is a fall to less than a third. The weak analysis degrades more than three times as fast as the strong one as arms are added.
So the cost of not adjusting is not a fixed penalty; it compounds with the width of the trial. A six-arm trial analysed without its balanced factors is spreading the same patients over more comparisons and handing back most of what the balancing bought, and the two losses multiply rather than add. The rule of thumb worth carrying is that the more arms a trial has, the less it can afford the analysis it is most likely to be given.
What a three-arm trial should report
The list is short and each item is a consequence of a number above.
Say which factors the rule balanced, because the analysis needs them and a reader needs to know whether they are in it.
Put them in the model. Sixty points of power at the setting measured here, and no cost beyond nine degrees of freedom.
If the model is doubtful, condition on the rule instead. It costs a few hundred re-runs of the allocation and holds the level exactly, and it is the only option that needs no modelling assumption at all.
Do not report the unadjusted test as conservative and therefore safe. It is conservative and therefore weak, and the two words describe the same number from the point of view of two different people.
The same shape, three fields along
This is the third time a rule that reads something before deciding has changed what the analysis afterwards means, and the three results together are more informative than any of them alone.
Reading outcomes to allocate breaks the error rate upwards: a response-adaptive rule inflates the rejection rate and biases the estimate of the arm it favoured, and repairing it costs nineteen points of power.
Reading covariates to allocate breaks the error rate downwards: the level falls to zero and the cost is power, and repairing it costs nine degrees of freedom.
Reading responses to place experimental runs breaks almost nothing, because the design’s objective is the analysis’s objective.
The pattern is not about how much the rule knows. It is about what the rule’s objective has to do with the analysis’s: opposed, orthogonal, or identical, and the three cases have three different signs.
What is claimed, and what is not
The claim is the analysis after a covariate-adaptive rule with more than two arms: the unadjusted F’s conservatism, its dependence on how much the balanced factors explain, its independence of the number of arms, the power the adjustment recovers, and the rule’s own reference distribution as the assumption-free alternative.
What stays out and is named: pairwise comparisons after the F, where the multiplicity structure of several arms against one control sits on top of everything here; covariate adjustment for factors the rule did not balance, which is a different and older question; and non-linear outcome models, where the conservatism has the same source and the arithmetic of the adjustment is not a partial F.
The boundary against the two-arm field is that it owns the finding and this owns its behaviour as the number of arms and the strength of the factors vary. Nothing here contradicts it; the numbers are its numbers with two degrees of freedom instead of one.
What this does to a trial’s stated sample size
The practical consequence is a number nobody computes, and it is worth computing once.
A trial powered at 80% under the assumption that its analysis behaves as advertised, and then analysed without the balanced factors, is not running at 80%. At the setting measured here — a hundred and fifty patients, three arms, an effect of half a standard deviation — the adjusted analysis has 70.4% power and the unadjusted 10.4%. To reach the adjusted analysis’s power with the unadjusted one would take several times the patients, and no protocol amendment would be required to notice, because nothing about the trial has visibly changed.
That is the strongest version of the two-arm field’s point about conservatism being expensive. A level that is too low is not a safety margin. It is a sample size calculation quietly multiplied.
The checks, and the refusal
Three claims are gated in this field’s library. The unadjusted F after a deterministic rule must reject far below its nominal level, at strongly prognostic factors. Adjusting must restore the level to within two points of 5%. And the conservatism must vanish when the balanced factors are unrelated to the outcome, which is the mechanism asserted rather than described — if the rate stayed low there, the explanation in this essay would be wrong.
The exact route is gated separately: both statistics against the rule’s own reference distribution must hold the level to within two and a half points, over two hundred re-randomisation tests of a hundred and ninety-nine draws each.
The field’s refusal is the unequal-target score of the previous essay.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What the balanced trial is worth — both name covariate-adaptive randomisation, covariate adjustment, error rate, monte carlo, reference distribution, statistical power
- A statistic that is exact twice — both name exact test, monte carlo, randomisation test, reference distribution, statistical power
- The analysis and the shape — both name covariate adjustment, error rate, exact test, randomisation test, reference distribution
- The corner the test is calibrated at — both name error rate, monte carlo, null hypothesis, reference distribution, statistical power
- The null the exactness is for — both name error rate, exact test, monte carlo, randomisation test, reference distribution
- A distribution drawn from the null — both name monte carlo, null hypothesis, reference distribution, statistical power
Named objects
A flat tag is an object no other essay names yet.
Conservative intervalCovariate-adaptive randomisationCovariate adjustmentDegrees of freedomError rateExact testMinimisationMonte CarloNull hypothesisRandomisation testReference distributionStatistical power