What the balanced trial is worth
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
The design half of this field ends well. A rule that reads the covariate leaves an eighth of the imbalance complete randomisation leaves, at a rate that improves with the size of the trial, and it costs almost nothing to run.
What happens next is the whole point of the field, and it is not the reassuring part. A trial is analysed by somebody, using a model, and the model does not know what the design did.
Three ways to read the same trial
Sixty units, a covariate that drives the outcome, no treatment effect at all — so every rejection counted is a false one — and three analyses of the identical data.
The unadjusted comparison is a two-sample test of the outcome: the difference in means, its standard error, a t. It is what gets reported when the covariate has been “handled by the design”.
The adjusted comparison puts the covariate in the model: y on arm and x, and the arm coefficient’s t.
The rerandomisation test holds the outcomes fixed, re-runs the assignment rule two hundred times, and counts how often it produces a statistic at least as large as the observed one. It has to be told the rule and nothing else.
After a coin the unadjusted test rejects 3.92% — its level, to within the noise of twelve hundred trials. After the rule that reads the covariate it rejects 0.50%.
One true null in two hundred, against a claim of one in twenty. The design removed the imbalance and the analysis is still pricing it.
Why, one level down
The mechanism is a ratio and it is worth computing rather than describing.
After a coin the ratio is 1.010: the standard error is right, because the imbalance it is pricing is really there. After the rule that reads the covariate it is 1.384.
The reason is that the unadjusted standard error comes from the residual spread of the outcome, and that spread still contains everything the covariate contributes. The estimate’s actual spread does not, because the design has removed the part of it that came from the arms differing on x: the imbalance falls from 0.249 standard deviations under a coin to 0.061 under the rule, and the estimate’s own standard deviation falls with it while the reported standard error does not move at all.
A test whose standard error is 38% too large rejects far less often than it claims, and there is nothing in the data that tells it so. The information that would fix it is the design, and the design is not in the model.
The whole table, read across
The four rules make the mechanism a gradient rather than a pair of cases, and reading down the column is more useful than any one number in it.
Under a coin the unadjusted test rejects 3.92% and its standard error is 1.010 times the estimate’s real spread. Under blocking inside two categories it rejects 1.75% at a ratio of 1.211. Under minimisation on the same two categories, 2.50% at 1.188. Under the rule that reads the number, 0.50% at 1.384.
The conservatism tracks the design’s quality exactly. It is not a property of one rule or of one kind of rule; it is a property of the gap between what the design removed and what the analysis knows about, and every improvement to the first without a corresponding change to the second makes it larger.
The adjusted analysis, meanwhile, reads 5.42%, 5.08%, 6.42% and 4.58% across the same four rules — at its level under all of them, because it is pricing the imbalance that is actually left rather than the imbalance a coin would have left. And the rerandomisation test reads 3.92%, 4.58%, 4.92% and 4.08%, at the 4.5% a two-hundred-draw reference distribution can deliver.
Two of the three analyses are indifferent to the design and one of them is wrecked by it.
What 0.50% says about the standard error
The rejection rate can be run backwards into the quantity that produced it, and doing so turns a striking number into a measurement of the analysis’s error.
A two-sided test at 5% rejects when the statistic exceeds about 1.96. If the reported standard error is a factor r larger than the estimate’s actual spread, the statistic is the true one divided by r, so the test rejects when the true one exceeds 1.96r. Setting the rate to 0.50% gives 1.96r = 2.807, so
The unadjusted analysis is reporting a standard error 43% larger than the spread of the thing it is attached to, which is a variance overstated by a factor of 2.05. The trial is being priced at twice the uncertainty it has.
That is a more useful statement of the defect than the rejection rate, because it is the number an interval is built from. A 95% interval on this trial is about 43% too wide, and its coverage is correspondingly near 99.5% rather than 95% — the same fact the rejection rate reports, said about the quantity anybody actually publishes.
And what that says about the covariate
One more step backwards prices the design rather than the analysis.
Let the estimate’s variance under a coin be a residual part plus a covariate-driven part, and let the balancing rule remove seven eighths of the second. The unadjusted standard error is estimating the coin’s total; the true spread is the balanced one. Setting their ratio to the measured 2.05 and solving, the covariate-driven part must be about 1.4 times the residual part — that is, the covariate accounts for roughly 59% of the estimate’s variance under complete randomisation.
Two readings follow, and they are the two halves of the field.
The design is working, and working hard. It removes seven eighths of the majority component of the variance, which is why the true spread is only 70% of the coin’s.
And that is exactly why the analysis fails so visibly. A rule that removed seven eighths of something contributing five per cent of the variance would leave a t test rejecting 4.7% instead of 5%, and nobody would ever have looked. The size of the conservatism is not a measure of how badly the analysis is wrong; it is a measure of how well the design worked, read through an analysis that was never told.
The conservatism is not free
A test that is too cautious under the null is sometimes waved through as safe. It is not safe, and the comparison that says so is against a coin.
The unadjusted comparison finds the effect 36.1% of the time after the rule that reads the covariate and 43.0% of the time after a coin.
The better-balanced trial finds less. Every point the design bought has been spent by the analysis and seven more with it, and the trial would have been better off assigning its units at random.
That is the strongest form the field’s argument takes. A design improvement that is invisible to the analysis is not merely wasted — it is negative, because the analysis’ standard error is calibrated to a variability the design has removed.
Two repairs, and what each needs to know
Both repairs recover it, and they need different things.
Adjusting for the covariate gives 4.58% under the null and 70.5% power — the best power in the table, and it needs the covariate’s values and a model that is right about how it enters.
The rerandomisation test gives 4.08% under the null and 67.0% power, and it needs the assignment rule. Nothing else: not the covariate’s relationship to the outcome, not linearity, not normality. Hold the outcomes fixed, re-run the rule, count.
The exactness of the second is worth stating carefully. With two hundred draws the smallest achievable p-value is 1/200 and a test that rejects when p < 0.05 is a 4.5% test rather than a 5% test, so 4.08% is the level it can deliver. It is exact for any rule by construction, because its reference distribution is the distribution the rule actually induces, and this field’s contribution is that the construction survives a rule that reads a continuous number rather than a level.
What the rerandomisation test is worth against a coin
One number in the power table is easy to miss and says what the design is actually for.
The rerandomisation test after a coin finds the effect 41.0% of the time. After the rule that reads the covariate it finds it 67.0%. Same test, same units, same effect, and twenty-six points of difference — all of it the design.
That is the balancing rule paying off. Under a coin the reference distribution is wide, because the assignments the rule could have produced include badly imbalanced ones; under the balancing rule every assignment in the reference distribution is well balanced, so the observed statistic is being compared against a much tighter set of alternatives.
The design is worth twenty-six points of power to an analysis that knows about it and minus seven to one that does not. Both numbers are about the same trial.
The two error rates move in opposite directions
There is a way of describing the unadjusted analysis that makes it sound acceptable, and it is worth dismantling because it is what allows the practice to continue.
The description is: the test is conservative, conservatism protects against false positives, and a trial that reports fewer false positives than it claims is erring on the safe side. Every clause of that is true and the conclusion does not follow.
A test has two error rates and they trade against each other. Pushing the false-positive rate from 5% to 0.5% is not free protection; it is the same movement along the same trade that takes the power from 43% to 36%. The trial has not become more careful, it has become less able to detect anything, and it has done so by an amount nobody chose and nobody reported.
The multiplicity field makes the same point about corrections that control the wrong rate: a procedure has to be judged by both of its numbers, and one of them alone is never evidence of anything. What is unusual here is that the movement was not a choice at all. Nobody decided to run a 0.5% test. It is what happens when a design and an analysis are chosen by different people, or by the same person on different days.
Where this differs from the field it inherits
The covariate field measured the same three analyses after rules that balance levels, and reported the unadjusted test at 0.6% and the rerandomisation test as costing almost nothing. The numbers here are recognisably the same family and two things are different.
The conservatism is worse: 0.50% here against 0.6% there, from a rule that leaves a fifth of the imbalance rather than a half. That is the same mechanism scaled — a better design makes an unadjusted analysis more wrong, not less.
And the rerandomisation test’s advantage over the coin is much larger, for the same reason. A rule that balances categories produces a reference distribution that is tighter than a coin’s; a rule that reads the number produces one that is much tighter than that.
Why adjusting is not simply better
Adjustment wins on power in every row of the table, and it would be easy to end there. Two things stop it.
The first is that the adjusted analysis assumes a model. It puts x in linearly because that is how the design’s criterion put it in, and if the outcome depends on x in some other way — a threshold, a curve, an interaction with the arm — then the adjustment removes the linear part and leaves the rest, while the standard error behaves as though everything had been removed. Nothing in this essay measures that case, and it is the obvious place for the result above to stop holding.
The second is that adjustment has to be pre-specified, and pre-specification is a promise about a decision taken before the data exists. An adjusted analysis chosen after seeing which covariates came out imbalanced is a different procedure with a different error rate, and it is the forking-paths problem wearing a covariate’s clothes.
The rerandomisation test has neither difficulty. It assumes no model at all, so nothing about how x enters the outcome can invalidate it, and there is nothing to pre-specify beyond the rule, which was fixed before the trial started by definition. What it costs is a few points of power against a correctly specified adjustment, and a record of the assignment process detailed enough to replay.
What an experimenter should do
The three findings compose into one instruction and it is short.
If the design read the covariate, the analysis has to. Either put it in the model, or use the rule’s own reference distribution, or the trial is worse than one that assigned at random. Which of the two repairs to prefer is a genuine choice: adjustment is more powerful and assumes a model, rerandomisation assumes nothing and needs the rule to be recorded exactly enough to re-run.
The second of those is not a small requirement. A rerandomisation test needs the rule, its probability, the order the units arrived in and the covariate values — all of which exist during the trial and none of which is usually kept in a form that can be replayed years later. The covariate field records the same condition and it applies here in a stronger form, because a rule that reads a continuous number has more to record.
What gets reported, and what it means
The last thing worth measuring is what a reader of the trial actually sees, because none of the above is visible in a published result.
A trial that used a balancing rule and an unadjusted analysis reports a p-value, a confidence interval and a statement that the arms were well balanced on the covariates. All three are true. The p-value is larger than it should be, the interval is wider than it should be, and the balance statement is the reason for both — and nothing in the report connects them.
The most likely visible consequence is a trial that reports a null result with a wide interval and concludes that the effect, if any, is small. That conclusion is drawn from an interval inflated by 38% relative to the estimate’s real spread, and the correct version of it would have been narrower and might have excluded values the published one includes.
So the failure is not that somebody gets a false positive. It is that a well-designed trial reports a weaker conclusion than its own data support, in a way that looks like caution and is invisible to everybody including its authors.
What is claimed here, and what is not
This essay claims the analysis after a rule that balances a continuous covariate: that the unadjusted comparison is enormously conservative, that its conservatism costs more power than the balancing bought, and that both an adjusted analysis and the rule’s own reference distribution recover it.
What stays out and is named as a decision: model misspecification in the adjusted analysis, where the covariate enters the truth non-linearly and the adjustment is partly wrong — which matters more here than usual, because the design balanced only the linear part; the rerandomisation test’s power against alternatives other than a constant shift; and the case where the covariate is measured after assignment, which is a different problem with a different name.
The checks, and the refusal that makes them mean something
Three claims are gated in this field’s library. The unadjusted comparison after the rule that reads the covariate is required to reject under 2% of true nulls while the same test after a coin is required to be at its level — the pair, because either alone would be a statement about the test rather than about the design. The adjusted analysis and the rerandomisation test are both required to be at their levels. And against a real effect, the unadjusted comparison after the rule is required to find it less often than after a coin, which is the finding that makes the conservatism a cost rather than a caution.
The refusal is the unadjusted analysis itself, on the standard that a stated error rate has to be the error rate of the whole procedure with its design included. A test that states 5% and delivers half a per cent, while the same data analysed with the covariate in the model delivers 4.58%, is refused.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The analysis after three arms — both name covariate-adaptive randomisation, covariate adjustment, error rate, monte carlo, reference distribution, statistical power
- Guessing one arm in three — both name covariate-adaptive randomisation, error rate, experimental design, monte carlo, randomisation
- The corner the test is calibrated at — both name conservative test, error rate, monte carlo, reference distribution, statistical power
- A proposal that moves more than two units — both name experimental design, monte carlo, randomisation, reference distribution
- Stationary is not convergent — both name experimental design, monte carlo, randomisation, reference distribution
- The null the exactness is for — both name error rate, experimental design, monte carlo, reference distribution
Named objects
A flat tag is an object no other essay names yet.
ANCOVAConservative testContinuous covariateCovariate-adaptive randomisationCovariate adjustmentError rateExperimental designMonte CarloRandomisationReference distributionRerandomisation testStandard errorStatistical power