The repair, with a covariate in it
Worth reading first: The experiments that could have happened.
A statistic that is exact twice repaired a permutation test by changing what it compares. Shuffling treatment labels and recomputing a difference in means builds a reference distribution with the wrong variance whenever the treated units vary more than the controls and the split is uneven, and it rejected 20.47% of true weak nulls at a quarter treated. Dividing the difference by its own separate-variance standard error inside every shuffle took that to 6.07% and kept the test exact under the sharp null. At an even split the two tests were the same test, draw for draw.
It ended on the statistic trials actually report. A trial that measured a baseline covariate adjusts for it, and its estimate is a regression coefficient rather than a difference in means. The essay suspected the repair would not carry over unchanged: an adjusted estimator’s variance involves the covariate’s design as well as the arms’ spreads, and the identity that made the two tests coincide at an even split used the fact that the only thing being subtracted was a mean. Both suspicions are right, and the second one has a large consequence the essay did not guess.
Four statistics, one construction
Each unit carries a covariate , measured before randomisation, and potential outcomes untreated and treated, so the covariate explains about half the outcome’s variance. The effect varies between units in one of two ways. In the first it varies with the outcome a unit would have had anyway — the earlier essay’s arrangement, in which the units who would do well untreated are the ones the treatment helps most. In the second it varies with the covariate alone: a treatment that works better for patients with a higher baseline, which is the heterogeneity a subgroup analysis looks for. Either way its average over the 150 units is exactly zero, so the weak null holds in the sample.
The trial regresses the observed outcome on treatment and the covariate. Two regressions are used: the ordinary adjustment, on , and the one that also includes the product of treatment and the centred covariate, , whose treatment coefficient is the average effect even when the effect varies with and the split is uneven. Each coefficient is permuted as it stands and over its HC2 sandwich standard error — each squared residual divided by one minus its leverage — which is exactly the separate-variance standard error when there is no covariate. Every test fixes the observed outcomes, re-randomises the assignment 199 times, and is therefore exact under the sharp null whatever its statistic; under the sharp null all four reject between 5.1% and 5.9% of a thousand experiments, within their counting error of 5%.
Where the effect follows the outcome
At a quarter of the units treated, with the effect varying with the untreated outcome by a standard deviation of 2, the adjusted coefficient’s permutation test rejects 14.8% of true weak nulls: the failure the earlier essay measured for the difference in means, somewhat smaller because the covariate absorbs part of the outcome variance the shuffle mis-assigns. The coefficient from the interacted model fails the same way unstudentised, 14.0%.
Studentised, both come down, and not to the same place. The adjusted coefficient over its sandwich rejects 3.0%, below the nominal rate by two points — conservative, where the difference in means’ studentised test was a point liberal. The interacted coefficient over its sandwich rejects 3.9%, and across the whole sweep of effect spreads it stays between 3.9% and 5.7%, where the adjusted one drifts from 5.7% down to 3.0%.
A conservative test is valid, so the adjusted statistic’s repair is not a failure the way the unstudentised one is. But it is a sign that the sandwich is estimating the wrong variance. When the effect varies with an outcome that the covariate partly predicts, the adjusted model leaves part of that variation in its residuals, and the sandwich reads the residual variance of the treated arm as larger than the estimator’s actual variance under re-randomisation. The interacted model fits a separate slope in each arm and leaves less of the heterogeneity in the residuals, so its sandwich is nearer the right variance.
Where the effect follows the covariate
The second kind of heterogeneity separates the statistics much further, and the hero figure’s slider shows it.
When the effect varies with the covariate by a spread of 2, the adjusted coefficient’s permutation test at a quarter treated rejects 5.5% of true weak nulls, which looks fine. Studentised, it rejects 1.0%. The interacted coefficient permuted as it stands rejects 0.3%; over its sandwich, 4.6%. Across the sweep the studentised interacted test stays between 4.6% and 5.7%. Two of the others drift towards nothing as the spread grows, and the unstudentised adjusted test holds near 5.5% throughout — the one case in which it looks right, and the arrangement in which the next section finds it rejecting nothing at an even split.
The reason is that an effect that varies with the covariate is a treatment-by-covariate interaction, and a statistic that does not model it treats it as noise. The adjusted model’s residuals contain the interaction, which inflates the sandwich; the permutation distribution of the interacted model’s coefficient, re-randomised without the interaction’s variance being accounted for in the statistic, is too wide. Only the statistic that fits the interaction and scales by the variance that remains is compared against a reference distribution of the right width.
The even split, which protected the difference, protects nothing here
For a difference in means an even split made the unstudentised and studentised tests identical and both nearly right. With a covariate in the model it does neither. With the effect following the untreated outcome, at an even split the adjusted tests reject 1.6% of true weak nulls, studentised or not, and the interacted test over its sandwich 3.1%. With the effect following the covariate, the adjusted test rejects 0.0% of a thousand experiments, studentised or not, the interacted test as it stands 0.1% — and the interacted test over its sandwich 5.0%.
The identity behind the even-split protection was an identity about subtracting two means, and with a covariate there is more than a mean in the statistic. The explanation that fits the numbers is that a shuffle which scrambles which units are treated also scrambles which units’ slopes enter the estimate, so the reference distribution includes the between-unit variation of the effect as if it were part of the estimator’s noise, and a test compared against a distribution that wide almost never rejects. Four statistics and two arrangements are the evidence for it here.
What a strict test gives up
A test that almost never rejects a true null is valid. It is also nearly powerless when the null is false in the same heterogeneous way, and the size of that is the practical content of the result.
At an even split, with the effect varying with the covariate by a spread of 2, an average effect of 0.2 standard deviations is detected by the adjusted permutation test in 5.2% of experiments and by the interacted test over its sandwich in 40.7%. At 0.3, 24.2% against 71.8%; at 0.4, 52.0% against 91.2%; at 0.5, 78.8% against 99.2%. The unstudentised interacted test and the studentised adjusted one track the adjusted test within a point or two throughout. A trial analysed with the ordinary covariate-adjusted permutation test would need roughly twice the average effect to see it, whenever that effect depends on the covariate it adjusted for.
When the effect does not vary, there is nothing to choose between them. At an average effect of 0.2 the four detect it in 38.0%, 38.3%, 38.0% and 38.0% of experiments; at 0.4, in 90.3%, 89.7%, 90.3% and 90.0%. The interacted, studentised test costs nothing where its advantages are not needed — the same shape as the earlier essay’s finding that studentising a difference in means cost 0.8 points of power against a constant effect — and gains an effect size’s worth of power where they are.
Why the interaction is the adjustment that composes
The earlier essay’s repair worked because it made the statistic a pivot: dividing by a standard error computed inside every shuffle removed the nuisance parameter — the arms’ different spreads — from the statistic’s distribution. For an adjusted estimate the nuisance parameters are the spreads and the way the effect varies with the covariate, and a standard error can only remove what the model has already separated from noise.
The ordinary adjustment assumes one slope for both arms. When the effect varies with the covariate, the arms have different slopes, and the model’s residuals carry the difference; its sandwich standard error is then a correct estimate of a quantity that is not the estimator’s variance under re-randomisation. The interacted model estimates a slope in each arm, and the heterogeneity it cannot explain is heterogeneity unrelated to the covariate — which is the kind the earlier essay’s studentisation already handled. It is the same move which weights are the inverse variances made for blocked designs: when the arms differ in more than their means, the estimator has to carry the difference rather than average over it. Adjusting with the interaction and studentising is the difference in means’ repair, carried over with the covariate moved from the noise into the model.
It also explains why the even split’s protection vanished. For the difference in means, the split decided how the two arms’ variances entered the reference distribution, and an even split balanced them. With a covariate, what enters is the covariate-by-arm structure, and no split balances a slope that differs between arms; only modelling the slope does.
What the interaction costs a trial of this size
The interacted model has one more coefficient than the ordinary adjustment, and a reader used to hearing that interactions need four times the sample might expect it to cost precision the ordinary model does not pay. It does not, and a term built from the others says why: the treatment coefficient in the interacted model is estimated from the centred covariate’s own variation within each arm, and a treatment’s interaction with a normal covariate rests on the same third of the units as the covariate’s own coefficient. The interaction’s coefficient is the noisy one; the average effect, estimated alongside it with the covariate centred, loses almost nothing, which is why the constant-effect powers above are identical to within the counting error.
The cost appears only with a skewed covariate or a small trial. A covariate with a long tail makes the interaction rest on a handful of units, and its sandwich standard error, which leans on those units’ residuals and leverages, becomes noisy in turn; at thirty units the HC2 leverage correction, which is small at 150, becomes large. Neither is measured here, and both would be the first things to check before recommending the interacted model for a small trial with a skewed baseline. What the measurements here support is narrower and still useful: at 150 units with a normal covariate, a trial that pre-specifies the interacted model and permutes its studentised coefficient gives up nothing when the effect is constant and gains an effect size’s worth of power when it is not, and the choice between the two models does not have to be made after looking at whether the effect varies.
Where this sits among the field’s constructions
The field began with the experiments that could have run: fix the outcomes, re-run the assignment, count, and the test is exact for the sharp null whatever the statistic. The null the exactness is for found that exactness does not reach the weak null a trial usually cares about, and the essay before this one found that studentising the statistic carries it most of the way there. This essay adds a condition: the statistic being studentised has to model everything that makes the arms differ, or its standard error estimates the variance of a misspecified model rather than the estimator’s variance under re-randomisation.
That condition has relatives elsewhere in this collection. The reference the covariates supply met it from the design side, where the covariates entered through the rule that assigned treatment and the re-randomisation had to re-run that rule to stay exact; and adjusting for everything met its opposite in observational data, where what is put into a model can create the bias it was meant to remove. Here every estimator is consistent for the average effect under randomisation, nothing is biased, and what is left out of the model miscalibrates the test instead — which is the quieter failure, since an estimate that is right and a p-value that is wrong look like a result.
What a trial adjusting for a covariate should permute
The treatment coefficient from the model with a treatment-by-covariate interaction, over its HC2 sandwich standard error. Under the sharp null it is exact, as every permutation test is. Under every weak null measured here — effects varying with the untreated outcome or with the covariate, at a quarter or half treated — it rejected between 3.1% and 5.7% of true nulls, and against a constant effect it costs no power.
Not the ordinary adjusted coefficient, studentised or not. Unstudentised it fails under an uneven split, as the difference in means did; studentised it is conservative, to 1.0% at a quarter treated and to nothing at an even split when the effect varies with the covariate, and it gives up most of its power exactly then.
And not the even split as a substitute for the right statistic. It protected the unadjusted difference and does not protect an adjusted one; the interaction term does.
The rates are over a thousand simulated experiments each at the sizes and 600 at the powers, with 199 re-randomisations per experiment, so a rate near 5% carries a standard error of about 0.7 points; the conclusions rest on differences of several points or on rates of nothing where the nominal is five. The sandwich is HC2 throughout, checked to equal the separate-variance standard error exactly for a coefficient on a lone indicator. Not measured: more than one covariate, where the interacted model has a slope per covariate per arm and the degrees of freedom begin to matter at 150 units; a binary outcome, where the adjustment is a logistic or risk-difference model and the sandwich behaves differently; and the other sandwich forms, HC0 and HC3, which differ from HC2 by leverage factors that are small here and would not be at thirty units.
Still open: the covariate chosen after randomisation
Every covariate here was chosen before the trial: the analysis adjusts for whatever the data show. Trials often choose which baseline covariates to adjust for after seeing them — the ones that turned out imbalanced, or the ones that most reduce the residual variance — and a cut that fitted best found that choosing a covariate’s form from the trial’s own data flatters its precision.
For a permutation test the question has a clean form. If the choice of covariate is re-made inside every re-randomisation — the analysis rule applied to each shuffled assignment as it was applied to the real one — the test stays exact under the sharp null by construction; if it is made once, from the observed data, and held fixed across the shuffles, it is not. Whether the difference is small, as the results on admitted families of analyses might suggest, or large, and whether the interacted, studentised statistic keeps its weak-null behaviour when the covariate it interacts with was itself chosen from the data, is a calculation on the same experiments and has not been made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The analysis after three arms — both name covariate adjustment, exact test, randomisation test, statistical power
- Walking the admissible set — both name exact test, permutation test, randomisation test, sharp null
- A probe nobody chose — both name randomisation test, sharp null, statistical power
- The analysis and the shape — both name covariate adjustment, exact test, randomisation test
- The analysis has to know the rule — both name covariate adjustment, randomisation test, statistical power
- The plus one and the round number — both name permutation test, randomisation test, sharp null
Named objects
A flat tag is an object no other essay names yet.
Analysis of covarianceCovariate adjustmentExact testHeteroskedasticity-consistentInteractionPermutation testRandomisation testSharp nullStatistical powerTreatment effect heterogeneity