The analysis has to know the rule
Worth reading first: Randomisation is not balance · The observations that repeat each other.
The two essays before this one measured what the rules do to the assignment. This one measures what they do to the test, and the answer is the reason the whole family is more than an administrative preference.
An allocation rule changes the joint distribution of the covariates and the arms. A two-sample comparison of means assumes that distribution is the one a coin produces. Those two facts do not sit together, and what happens when they are made to is not subtle.
Counted
Seven hundred trials of 120 patients, allocated by minimisation at p = 0.8, with the prognostic factors carrying a real effect on the outcome and no treatment effect at all. Every rejection below is a false positive, and all four claim 5%.
- the plain difference of means against a t table: 0.6%
- the same difference adjusted for the balanced factors: 5.4%
- the plain difference against the trial’s own re-randomisation distribution: 5.1%
- the adjusted difference against the same: 4.1%
The first number is the finding. A test that rejects six true nulls in a thousand when it claims fifty has stopped being a 5% test in any useful sense, and it is the analysis that a very large share of trials report.
At p = 1 it is worse: 0.0%. Not one rejection in seven hundred.
Why it goes that way
The direction is the opposite of everything the adaptive field measured, and the mechanism is straightforward once it is stated.
The t statistic divides the difference between the arms by an estimate of its standard error, and that estimate is computed from the total variance of the outcomes — which contains the prognostic variation the covariates carry. Under a coin, the difference between the arms also contains that variation, at exactly the rate the standard error assumes, and the two match.
A balancing rule removes the covariate variation from the numerator and leaves it in the denominator. The difference between the arms is smaller than a coin would have produced, and the standard error it is divided by has not been told. So the statistic is systematically too small, and the test rejects too rarely.
The size of the effect is therefore governed by exactly one thing: how much of the outcome the balanced covariates explain.
Sweeping that at four hundred trials: with the factors at no effect the unadjusted test is at 4.8%; at half strength 2.3%; at full strength 0.8%; at double 0.0%. The adjusted test is at 5.5% at every one of the four, and the re-randomised rows do not move either. The conservatism is proportional to the balance being worth something, which means a rule that works better makes the naive analysis worse.
That is an uncomfortable sentence and it is the honest reading. The better the covariates predict, the more worth balancing they are, and the more badly the analysis that ignores them behaves.
What conservatism costs
A test that rejects too rarely under the null is not “safe”. It is a test whose power has been spent, and the spend is visible the moment there is something to detect.
Against a real treatment effect of 0.45 standard deviations, on the same design:
- unadjusted against a t table: 20.3%
- adjusted against a t table: 66.7%
Forty-six points of power, from an analysis choice, on data that had already been collected. And the comparison against a coin-allocated trial is the one that makes it stark: under complete randomisation the same unadjusted analysis gets 28.6% — so the balanced trial’s unadjusted analysis is less powerful than the unbalanced trial’s. The rule improved the design and, analysed that way, made the result worse.
Adjusting is the repair and it is not free of assumptions
Putting the factors into the model restores the level: 5.4% under the null, within a standard error of nominal, and 66.7% power. That is the whole repair and it is one line of code.
What it costs is a model. The adjusted analysis assumes the covariates enter the outcome the way the model says they do, and here they do, because the outcomes were simulated from that model. In a real trial they do not exactly — a factor may act non-linearly, or interact with something, or the recorded band may be a coarsening of an underlying continuous quantity.
That is not a reason to prefer the unadjusted analysis, which is wrong in a way that is worse and does not depend on any model being right. It is a reason to notice that the two available answers so far are a test that is badly conservative and assumes nothing, and a test that is correct and assumes a model. The third and fourth rows of the table are a third answer, and they are the subject of the next essay: both statistics read against the set of allocations the rule could have produced, which needs no model and holds its level whichever statistic goes into it.
The same arithmetic, from the other end of the site
There is a field on this site that measured the mirror image of this and did not know it, and putting the two together is the clearest way to see what is going on.
The dependence field is about observations that repeat each other. Fifty observations at a lag-one correlation of 0.8 carry about six observations’ worth of information, the standard error divides by √50 as though they were independent, and the 95% interval covers 47.0%. A numerator that is too large for the denominator it is divided by.
This is the same fraction with the error in the other place. Here the numerator has been made smaller — the arms are more alike than chance would make them — and the denominator has not been told. Positive dependence between observations makes an interval too narrow; negative dependence between the assignment and the covariates makes a test too weak. One statistic, two ways for its two halves to stop matching, and opposite symptoms.
The pairing also says which of the two is more dangerous, and the answer is the one that is not here. An interval that covers 47% when it claims 95% publishes false certainty; a test that rejects 0.6% when it claims 5% publishes nothing at all, and the trial that produced it is simply wasted. The first is a wrong answer and the second is a missing one, and only the first survives into the literature to mislead anybody.
Is it a defect or a design choice
There is a defence of the unadjusted analysis and it deserves to be stated properly rather than dismissed, because it is not silly.
The defence: an unadjusted comparison of two randomised arms is valid without any model at all. It does not assume the covariates enter linearly, it does not assume the bands are the right bands, and it cannot be wrong about a functional form because it never states one. A conservative test is a test whose stated level is an upper bound, and an upper bound is what a regulator asks for.
The reply has two parts and the second is the one that matters.
First, the level is not what an upper bound is usually taken to protect. A test that rejects 0.6% of true nulls will also, at any given alternative, reject far fewer real effects — the level and the power move together, and there is no sense in which the conservatism is bought and the power kept.
Second, and this is the part the table settles: the model-free virtue is available without the conservatism. The third row — the unadjusted difference read against the trial’s own re-randomisation distribution — assumes nothing about the outcomes either, and it is a 5% test at 5.1% with 55.0% power against the same alternative. So the choice is not between assuming a model and losing power. It is between assuming a model, losing power, and doing the arithmetic that the design has already made available.
What the conservatism is, as a scale factor
The rejection rates are the symptom and the mechanism is a mismatch of scales, so it is worth converting the rates back into the scale they imply. A statistic whose true spread is s times what its table assumes rejects 2(1 − Φ(1.96/s)) of the time, and inverting that at each rate gives s directly.
At the factors’ full strength, 0.8% implies s = 0.71: the statistic is running at about seven-tenths of the spread the t table was written for, which is to say the rule removed roughly half the variance of the numerator. At half strength, 2.3% implies s = 0.86 and about a quarter removed. At no strength, 4.8% implies s = 0.99 and 1.7% removed — which is nothing, and is the check that says the conversion is measuring the right thing: with no prognostic variation there is nothing for a balancing rule to take out, and the arithmetic finds nothing.
Read that way the sweep is one quantity rather than four rejection rates: the share of the numerator’s variance the rule has removed and the denominator still contains, running 2%, 26%, 49% and effectively all of it as the factors strengthen. The rejection rate is a very steep function of that share — halving the numerator’s variance takes the size from 5% to under 1% — which is why the failure looks catastrophic while the underlying removal is only a factor of two.
That steepness is also the reason the effect is invisible in a trial report. Nothing about a design that removed half the variance of a difference looks alarming; a balance table showing well-matched arms is exactly what removing it produces, and is what a reader would call a success.
Three quarters of the loss, with no model
The power figures contain a comparison the essay makes qualitatively and it is worth having as arithmetic, because it prices the third row against the second.
Against an effect of 0.45 standard deviations, the unadjusted t-table analysis reaches 20.3% and the adjusted one 66.7% — a gap of 46.4 points. The unadjusted statistic read against the trial’s own re-randomisation distribution reaches 55.0%, which recovers 34.7 of those points, or 75% of the loss, while assuming nothing at all about how the covariates enter the outcome.
So the model buys the last quarter. That is a much smaller claim for adjustment than the pair of t-table rows suggests, and it changes what the choice is about: not between a valid weak test and a powerful assumed one, but between two valid tests differing by twelve points of power, one of which needs a correctly specified outcome model and one of which needs the allocation rule to be re-runnable.
The remaining comparison is the one that puts a floor under all of it. A coin-allocated trial analysed the same plain way reaches 28.6%, above the balanced trial’s 20.3% and far below every other row. Balancing and then adjusting is worth 38 points over not balancing at all; balancing and then not adjusting is worth minus eight. The rule is not the thing that decides whether the design was worth running — the sentence in the analysis plan is.
Which rules need it
Not all of them, and the pattern says exactly what the conservatism is caused by.
- A coin: unadjusted 5.1%. Nothing to repair.
- Permuted blocks: 4.9%. Nothing to repair — blocking the totals removes no covariate variation, which is the previous essay’s point arriving where it matters.
- Stratified blocks: 0.1%.
- Minimisation at p = 0.8: 0.6%.
So the rules that balance covariates need the adjustment and the rules that do not, do not. That is not a coincidence and it is the cleanest confirmation available that the mechanism above is the right one: the conservatism tracks the balancing, rule by rule, exactly as it tracks the covariates’ prognostic strength.
Stratified blocks at 0.1% is worth its own note, because stratification is often described as the conservative-in-the-safe-sense option. It is the most conservative rule in the table, and for the same reason as the rest.
The general rule this is a case of
There is a statement this site has arrived at three times now from different directions, and this is the fourth.
A procedure’s advertised property is a claim about the whole procedure. A p-value is defined relative to a sampling plan, so a stopping rule changes it without changing a number in the dataset. An estimate after a selection is not the estimate it looks like. An interval after a model selection does not cover what it says. And a test after an allocation rule is not the test the table was written for.
In every one of the four the omitted step happens outside what anybody thinks of as the analysis, and in every one the fix is the same in shape: either put the omitted step into the model, or build the reference distribution that contains it.
What the fourth row is doing at 4.1%
One number in the table has not been explained and it should not be left standing. The adjusted statistic against the re-randomisation distribution comes out at 4.1%, which is below nominal by about a standard error and a half.
Two things are worth saying about it. It is not the conservatism the essay is about — that one is at 0.6%, an order of magnitude further down, and moves with the covariates’ strength while this one does not. And a randomisation test’s attainable sizes are a lattice: with 99 re-randomisations the p-value can only take the values 1/100, 2/100 and so on, so the largest attainable size at or below 5% is exactly 5% and small departures from it are the discreteness rather than a defect. The exact field measured that lattice directly and found the correction for it worth exactly zero or exactly one step, with the meeting points at every B anybody uses.
What would be worth chasing is a rate that drifted with the design, and it does not: the fourth row sits at 4.1% at every value of the covariates’ strength in the sweep above, which is what a test whose validity does not depend on the outcome model looks like.
What is being claimed here, and what is not
This essay claims the size and power of two analyses after four allocation rules, at one sample size, one set of factors and one outcome model, with the conservatism’s dependence on the covariates’ strength measured rather than argued.
What stays out: the asymptotic theory that says which analyses are valid under which rules, which is where these results come from in the literature and is not derived here; misspecified adjustment, where the model put into the analysis is not the model the outcomes came from — named as the honest worry about the repair and not measured; and covariates that were recorded but not balanced, which are a different case with a different answer.
One more, because it is the question a reader will have and the answer is genuinely not here. Nothing above says how much of this survives at a trial of two thousand rather than a hundred and twenty. The mechanism does not weaken with n — a balancing rule holds the margins flat while a coin’s grow, so the relative removal of covariate variation from the numerator gets larger rather than smaller — but the outcome is a race between that and the shrinking standard error, and a race is a measurement rather than an argument. The numbers here are for one size and are labelled as such throughout.
The checks
Two claims are gated in this field’s library.
The unadjusted analysis is conservative and the adjusted one is not, with the unadjusted rate required to be more than two standard errors below 5%, the adjusted rate within two and a half points of it, and — the part that makes it attributable — the same unadjusted analysis after a coin required to be correct. Without that last clause the check would pass for a simulation whose t test was simply broken.
And the exact route holds its level under both statistics and at two values of p, which is the next essay’s claim and is gated here because it is the control that says the conservatism belongs to the reference distribution rather than to the statistic. The unadjusted statistic against a t table and the identical statistic against the rule’s own reference distribution differ by more than a percentage point on the same trials, and only one of them is a 5% test.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Guessing one arm in three — both name covariate-adaptive randomisation, error rate, minimisation, permuted blocks
- Three arms and three scores — both name covariate-adaptive randomisation, minimisation, permuted blocks, stratified randomisation
- What the exactness buys — both name error rate, randomisation test, reference distribution, statistical power
- A probe nobody chose — both name randomisation test, reference distribution, statistical power
- A proposal that moves more than two units — both name minimisation, randomisation test, reference distribution
- A statistic that is exact twice — both name randomisation test, reference distribution, statistical power
Named objects
A flat tag is an object no other essay names yet.
Conservative intervalCovariate-adaptive randomisationCovariate adjustmentError rateMinimisationPermuted blocksRandomisation testReference distributionStatistical powerStratified randomisation