The reference distribution the design supplies

The covariate chosen afterwards

A trial that chooses which baseline covariate to adjust for after randomisation has two ways to run its permutation test: hold the chosen covariate fixed across the shuffles, which is what an analysis that reports adjusting for it actually does, or re-make the choice inside every shuffle, which is exact by construction. Holding it fixed, a choice made on significance rejects 10.4% of true sharp nulls, a choice made on imbalance only 3.0% — and pays seven and a half points of power for it — while a choice that never reads the arm labels is the same test either way. The interacted, studentised statistic keeps its weak-null behaviour only when the covariate chosen is the one the effect varies with.

Worth reading first: The experiments that could have happened · Randomisation is not balance.

The repair, with a covariate in it found that a permutation test of a covariate-adjusted treatment effect holds a weak null — an average effect of zero with units differing — only when the statistic is the coefficient from the model with a treatment-by-covariate interaction, divided by its sandwich standard error. Every covariate in that essay was named before the trial. Trials often name it afterwards: the analysis adjusts for whichever baseline covariate turned out imbalanced between the arms, or whichever most reduced the residual variance, and a cut that fitted best found that choosing a covariate’s form from the trial’s own data flatters the precision it reports.

For a permutation test the question has a clean form, and it is the one the test that needs the rule answered for the assignment itself. If the choice of covariate is re-made inside every re-randomisation — the analysis rule applied to each shuffled assignment as it was to the real one — the statistic is one fixed function of the assignment and the test is exact under the sharp null by construction. If the choice is made once, from the observed data, and the chosen covariate held fixed across the shuffles, nothing guarantees it. How far the second strays, in which direction, and why, depends entirely on what the choice reads.

How often an adjusted permutation test rejects a true sharp null when the covariate is chosen after randomisation, held fixed or re-chosen inside every shuffle150 units, half treated, six baseline covariates of which three are prognostic; 2,000 experiments and 99 re-randomisations each, read on the adjusted coefficient over its sandwich standard error. Choosing the covariate most correlated with the outcome: 5.0% held fixed, 5.0% re-chosen. Choosing the covariate with the smallest residual variance: 5.5% held fixed, 5.3% re-chosen. Choosing the covariate most imbalanced between the arms: 3.0% held fixed, 4.8% re-chosen. Choosing the covariate with the smallest adjusted p-value: 10.4% held fixed, 5.1% re-chosen. Naming the most prognostic covariate in advance: 5.5%.0%2.5%5%7.5%10%12.5%most correlated with the outcome5.0%5.0%smallest residual variance5.5%5.3%most imbalanced between arms3.0%4.8%smallest adjusted p-value10.4%5.1%named in advance: 5.5%chosen once, held fixedre-chosen inside every shuffle2,000 experiments, 99 re-randomisations eachthe target line is 5%
Fig. 1 How often the permutation test of the adjusted treatment coefficient, over its sandwich standard error, rejects a true sharp null when the covariate is chosen after randomisation by each of four rules — held fixed across the re-randomisations, and re-chosen inside each — with the test that names the most prognostic covariate in advance as the dashed line. The slider switches to power at an effect of 0.3, and to a weak null.

Four ways to choose, two ways to test

The experiments have 150 units, half treated, and six baseline covariates measured before randomisation. Three are prognostic — the outcome untreated is 0.5x1+0.4x2+0.3x30.5x_1 + 0.4x_2 + 0.3x_3 plus noise, with the noise’s variance chosen so the outcome has unit variance — and three are not. Each experiment is one randomisation and 99 re-randomisations of the same outcomes, and every rate below is over 2,000 experiments unless it says otherwise, which puts the counting error of a 5% rate near half a point.

Four rules choose one covariate to adjust for, each a rule a real analysis plan contains:

  • most correlated with the outcome, pooled across both arms — a choice made without ever reading which unit was treated;
  • smallest residual variance in the regression of the outcome on treatment and the covariate — “the one that explains most”;
  • most imbalanced between the arms — “adjust for what randomisation failed to balance”;
  • smallest adjusted p-value — the rule nobody writes down, and the one twenty analyses of nothing is about.

The hero figure sets out the result. The rule that never reads the labels gives one test, not two: on every shuffle it picks the covariate it picked on the data, because its choice does not depend on the assignment at all, and it rejects 5.0% of true sharp nulls held fixed or re-chosen. The residual-variance rule reads the labels a little, switches covariate on 2.1% of shuffles, and rejects 5.5% fixed and 5.3% re-chosen, both within counting error of the covariate named in advance. The other two rules come apart.

Chosen on significance and held fixed

Choosing the covariate that gives the smallest adjusted p-value, then permuting with that covariate held fixed, rejects 10.4% of true sharp nulls — more than twice the nominal rate. That is no surprise: it is six analyses with the best one reported, and its p-value is computed as though only one had been run. The surprise, if there is one, is how cheaply it is repaired. Re-making the choice inside every shuffle — taking, on each re-randomised assignment, the smallest p-value over all six covariates — rejects 5.1%, and the test is exact because the minimum over six analyses is simply the statistic.

What holding a chosen covariate fixed does to the permutation p-value under the sharp null. The distribution function of the studentised adjusted statistic's permutation p-value over 2,000 experiments with no effect. A valid p-value lies on the diagonal. Chosen on significance and held fixed, 10.4% of p-values are at or below 0.05 and 76.5% at or below one half. Chosen on imbalance and held fixed, 3.0% and 46.4%. Re-chosen inside every shuffle, 4.8% and 51.5%. A choice blind to the arms — the covariate most correlated with the outcome — gives 5.0% and 51.0%, held fixed or not.
Fig. 2 The distribution of the permutation p-value under the sharp null over 2,000 experiments, for the covariate chosen on significance and held fixed, chosen on imbalance and held fixed, and chosen on imbalance and re-chosen inside every shuffle. A valid p-value lies on the diagonal.

The p-value figure shows where the excess comes from. Held fixed, the significance-chosen p-value is piled towards zero across the whole range, not just below 5%: the selection pushes every p-value down, so a reader who discounts only the ones near the threshold has not discounted enough. Re-chosen, the same rule’s p-value lies on the diagonal, because its reference distribution now contains the same search the observed statistic went through.

Held fixed, 76.5% of the significance-chosen p-values fall at or below one half, where a valid p-value puts half of them. The chosen-on-imbalance p-value, held fixed, errs the other way, with 46.4% at or below one half; re-chosen, the imbalance rule’s p-values put 51.5% there. The blind rule, which is one test however it is run, puts 51.0% there, which is the diagonal to within the counting error of two thousand experiments.

What each choice costs when there is an effect

An error rate is half of what a test is for. The other half is how often it finds an effect that is there, and the four rules differ there too.

How often an adjusted permutation test detects an effect of 0.3 when the covariate is chosen after randomisation, held fixed or re-chosen inside every shuffle. 150 units, half treated, six baseline covariates of which three are prognostic; 1,000 experiments and 99 re-randomisations each, read on the adjusted coefficient over its sandwich standard error. Choosing the covariate most correlated with the outcome: 52.1% held fixed, 52.1% re-chosen. Choosing the covariate with the smallest residual variance: 53.2% held fixed, 52.5% re-chosen. Choosing the covariate most imbalanced between the arms: 44.6% held fixed, 53.9% re-chosen. Choosing the covariate with the smallest adjusted p-value: 64.8% held fixed, 51.7% re-chosen. Naming the most prognostic covariate in advance: 52.8%.
Fig. 3 How often each test detects an average effect of 0.3 standard deviations, over a thousand experiments with 99 re-randomisations each, with the covariate chosen by each rule and held fixed or re-chosen, against the covariate named in advance.

The price of honesty for the significance rule is the price of the search, and it is small. Against an effect of 0.3 standard deviations, the re-chosen significance rule detects it in 51.7% of a thousand experiments, against 52.0% for the covariate named in advance. Searching six covariates and paying for the search costs almost nothing when one of the six is genuinely prognostic and the search usually finds a good one. The fixed version detects it in 64.8%, and every point of that difference was bought with excess error: a test that rejects 10.4% of true nulls will reject more of the false ones too, and nobody should count that as power.

The blind rule detects the effect in 52.1%, and the residual-variance rule in 53.2% fixed and 52.5% re-chosen. Both find the strongly prognostic covariate most of the time — in 77.1% and 82.9% of experiments under the null — and when they miss it they find the second, which removes almost as much of the noise. A rule that reads the outcome is reading the one thing that decides which covariate is worth adjusting for, which is why the rules that read it lose nothing to the covariate named in advance.

Chosen on imbalance and held fixed

The imbalance rule goes the other way. Choosing the covariate whose arm means differ most and holding it fixed across the shuffles rejects 3.0% of true sharp nulls, and its p-values pile towards one: the test is conservative. Re-chosen, it rejects 4.8%.

The reason is what the choice certifies. To be the most imbalanced of six, the chosen covariate’s imbalance must exceed every other covariate’s, so the observed assignment is one in which the other five are less imbalanced than the chosen one. Five times in six the chosen covariate is not the strongly prognostic one — it was chosen 16.7% of the time — and adjusting for a weak or useless covariate leaves the prognostic covariates’ imbalance in the statistic. On the observed assignment that leftover imbalance is smaller than usual, because the selection guaranteed it. On a re-randomised assignment, with the chosen covariate held fixed, nothing guarantees anything, and the prognostic covariates are as imbalanced as randomisation makes them. The reference distribution is wider than the distribution the observed statistic was drawn from, and it rejects too rarely.

That account makes a prediction, and it holds. With all six covariates unrelated to the outcome, the choice certifies nothing about the outcome, and over a thousand experiments the fixed imbalance test rejects 3.7%, the re-chosen one 4.0% and the test with a named covariate 3.9% — the same to within counting error. The conservativeness is a property of choosing among covariates some of which matter.

Conservative sounds safe and is not free. Against the effect of 0.3, as the power figure shows, the fixed imbalance test detects it in 44.6% of experiments; re-chosen, 53.9%; named in advance, 52.1%. Adjusting for whatever randomisation failed to balance, the most respectable-sounding of the four rules, costs 7.5 points of power against the covariate named in advance, and 9.3 against the same rule re-chosen, when it is analysed as though the covariate had been named — power that is recovered entirely by putting the rule inside the re-randomisation. Randomisation is not balance, and the remedy for that is not to chase the imbalance after the fact, but to account for the chase.

At a quarter treated

An even split hides some of what a choice can do, because at an even split many statistics coincide; the null the exactness is for found the studentised and unstudentised permutation tests returning the same p-value in every draw there. With a quarter of the 150 units treated the four rules, run over a thousand experiments each, keep their order and change their sizes. The significance rule held fixed rejects 12.6% of true sharp nulls, more than at the even split, because with an uneven split the six adjusted statistics are noisier and their minimum p-value is pulled further down. Re-chosen, it rejects 6.0%, which is a count, not a defect: the test is exact by construction and a thousand experiments put a 5% rate anywhere between about 3.6% and 6.4% one time in twenty.

The imbalance rule held fixed rejects 4.1%, nearer the nominal rate than at the even split, and re-chosen 5.6%. The blind rule rejects 5.7% and the residual-variance rule 6.0% fixed and 5.5% re-chosen. Nothing at a quarter treated changes the conclusion at a half; it moves the liberal rule further from 5% and the conservative one nearer.

What a rule must not read

The four rules sort themselves by one property: how much the choice depends on the assignment. A rule that reads only the outcomes and the covariates, pooled, cannot depend on it, and for such a rule fixing and re-choosing are literally the same test. The residual-variance rule reads the assignment, but only through a regression that the treatment barely moves under the null, and it switches covariate on 2.1% of shuffles; its fixed test is within counting error of exact. The imbalance and significance rules depend on the assignment by design — one reads the covariates’ arm means, the other the outcome’s — and each switches covariate on most shuffles, 83.2% and 76.0%.

Why the residual-variance rule reads the assignment so little is worth a sentence. The residual variance after adjusting for a covariate is the outcome’s variance less the part the covariate explains, and under the null the treatment explains almost none of it, so the rule’s ranking of the six covariates is the ranking by how much of the outcome each explains — very nearly the blind rule’s ranking. Only when two covariates explain almost the same amount can a shuffle reorder them, and with one covariate strongly prognostic that happened on 2.1% of shuffles. On a list of covariates that are all weakly and similarly prognostic it would happen more, and the fixed test would stray further; the measurement here is for a list with a clear leader.

That is a usable rule for a protocol. A covariate chosen by a criterion blind to the arms needs no re-randomisation of the choice and no correction, which is the same reason a sample size re-estimated from the blinded variance needs none of either. A covariate chosen by anything that reads the arms belongs inside the permutation, and once it is there every one of these rules is exact.

Under a weak null

The interacted, studentised statistic earned its place because it holds a weak null when the effect varies with the covariate. With the covariate chosen from the data, it can only do that if the chosen covariate is the one the effect varies with.

How often an adjusted permutation test rejects a true weak null when the covariate is chosen after randomisation, held fixed or re-chosen inside every shuffle. 150 units, half treated, six baseline covariates of which three are prognostic; 1,000 experiments and 99 re-randomisations each, read on the interacted coefficient over its sandwich standard error. Choosing the covariate most correlated with the outcome: 4.8% held fixed, 4.8% re-chosen. Choosing the covariate with the smallest residual variance: 5.1% held fixed, 5.1% re-chosen. Choosing the covariate most imbalanced between the arms: 1.3% held fixed, 2.4% re-chosen. Choosing the covariate with the smallest adjusted p-value: 9.4% held fixed, 4.0% re-chosen. Naming the most prognostic covariate in advance: 4.8%.
Fig. 4 How often the interacted coefficient over its sandwich standard error rejects a true weak null — an average effect of zero, with each unit’s effect equal to its value of the first covariate — when the covariate is chosen by each rule, held fixed or re-chosen, over a thousand experiments.

Let the effect vary with the first covariate — a unit’s effect is x1x_1 times one, averaging zero — and read the interacted statistic over its sandwich standard error. Named in advance, the first covariate is also the most prognostic, and it rejects 4.8% of true weak nulls. The rules that read the outcome pick that covariate in every experiment, because a covariate the effect varies with is one the treated outcomes depend on strongly, and they reject 4.8% and 5.1%. The imbalance rule picks it in only one experiment in six, interacts the treatment with an irrelevant covariate the rest of the time, and loses the repair: 1.3% held fixed, 2.4% re-chosen. The significance rule, held fixed, rejects 9.4%; re-chosen, 4.0%.

The unstudentised statistics, for comparison, reject 2.2% of the same weak nulls under every rule that reads the outcome and under the named covariate, which is the even-split conservativeness the repair with a covariate in it found: they are conservative rather than wrong. Under the two rules that read the arms they move with the rule — 1.0% held fixed on imbalance, 7.1% held fixed on significance — for the reasons the sharp null already showed. The studentised interacted statistic is the one with something to lose, and the imbalance rule takes it away.

Re-choosing inside the shuffles is exact only for the sharp null, and the weak null is not the sharp null. What re-choosing cannot do is make the interaction model right when the covariate in it is the wrong one; the studentised interacted statistic’s protection is a protection for the covariate it is given, and a rule that chooses on imbalance usually gives it a different one.

What an analysis plan can say

Name the covariate in advance when it can be named. Every rule here was at best as good as naming the most prognostic covariate, and that is the comparison the rules lose least on only because one of the six was strongly prognostic and easy to find.

Choose blind to the arms when it cannot. A choice made from the outcomes and covariates pooled — the covariate most correlated with the outcome — is the same test whether or not it is re-made, and it rejected 5.0% of sharp nulls here.

Write the rule down, because the rule is the test. A re-chosen test is exact only for the rule that was re-run, and an analysis plan that says “adjusted for imbalanced covariates” without saying how imbalance is measured, among which covariates, and with what tie-break has not specified a test. The test that needs the rule made the same point about the assignment mechanism; here it applies to the analysis. Re-running a stated rule costs nothing worth mentioning — six regressions per shuffle instead of one.

Otherwise put the rule inside the re-randomisation. Every rule re-chosen on every shuffle was exact: 5.3%, 4.8% and 5.1% for the three that read the arms. Holding the choice fixed is liberal for a rule that reads the outcome’s significance and conservative for one that reads the covariates’ imbalance, and the conservative case costs seven and a half points of power at the effect measured.

The rates are counted over 2,000 experiments for the sharp null and 1,000 for the weak null and the effect, each with 99 re-randomisations. A rate of 5% carries a counting error of about half a point at 2,000 experiments and seven tenths at 1,000; differences smaller than that are not claimed. The choice rules are applied exactly as stated on every shuffle, the adjusted fits are ordinary least squares in closed form on centred sums, and the sandwich is HC2. The fixed significance rule is refused as a valid test: it rejects 10.4% of true sharp nulls.

Still open: more covariates than units can bear

Six candidates is a short list. Modern trials record dozens of baseline variables, and the rules analysts use there — a lasso on the outcome, a stepwise search, a propensity model for the imbalance — choose sets rather than single covariates. Re-choosing inside every shuffle remains exact for any of them, at the cost of running the whole selection a few hundred times, and a lasso run blind to the arms remains the same test fixed or re-chosen.

Two things are not known from these sums. One is how much power a re-chosen set-selection rule gives up against the named oracle when no single covariate is strongly prognostic and the selection is noisy, which is where the search’s price could stop being small. The other is the fixed imbalance rule’s conservativeness as the list grows: the certificate a selection issues — every other covariate less imbalanced than this one — gets stronger with more candidates, and how fast the power it costs grows with the length of the list has not been measured.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Analysis of covarianceCovariate adjustmentExact testHeteroskedasticity-consistentInteractionPermutation testRandomisation testSharp nullStatistical powerTreatment effect heterogeneity