The reference distribution the design supplies

The null the exactness is for

A permutation test is exact under the hypothesis that the treatment changed nothing for anybody. Under the hypothesis it changed nothing on average, with a quarter of the units treated and the effect varying between them, it rejects a true null 22.93% of the time.

Worth reading first: The experiments that could have happened.

This field’s construction is the strongest thing on the site. Fix the outcomes, re-run the assignment rule, count how many of the experiments that could have happened are at least as extreme as the one that did. It needs no population, no distributional assumption and no large sample; it needs to know the rule, and feeding it the wrong rule is the one way to break it.

There is a second way, and it is not a way of breaking the test. It is a way of asking it a question it was never answering.

Fixing the outcomes is an operational statement of a hypothesis: patient i’s outcome is what it is, whichever arm they were sent to. That is the sharp null — the treatment changed nothing for anybody. The hypothesis a trial is normally run against is that the treatment changed nothing on average, which is implied by the sharp null and does not imply it, and the gap between the two is a quantity no experiment can measure, because no unit is ever observed under both arms.

An exact test rejecting a true hypothesis a fifth of the time. How often each analysis reports an effect when the average treatment effect is exactly zero and the effect varies between units, at 150 units with 25% treated. The permutation test on the difference in means reads 4.20% where the effect is constant — where the two nulls coincide and its exactness applies — and 22.93% where the effect varies with a standard deviation of 3. The same test on the studentised difference reads 6.27% there, and the ordinary large-sample t, which makes no exactness claim at all, reads 6.60%.
Fig. 1 How often each statistic’s permutation test rejects when the average treatment effect is exactly zero and the effect varies between units, at 150 units with a quarter treated. The permutation test on the difference in means reads 4.20% where the effect is constant and 22.93% where its standard deviation is 3.

An exact test, run correctly, on a true hypothesis, rejecting it nearly a quarter of the time.

What varies, and why nothing can estimate it

Every unit here has two potential outcomes: what it would do under treatment and what it would do under control. The difference between them is that unit’s effect. The average of those differences is the average treatment effect, and the weak null says that average is zero.

It can be zero while the individual effects are large. Half the units helped by one unit and half harmed by one unit averages to nothing, and so does an arrangement in which the units who would have done well anyway are helped and the rest are harmed. Every experiment in the sweep above is of the second kind, with the average held at exactly zero and only the spread moving.

The spread is a nuisance parameter in the strictest available sense. It is not merely unknown; it is unidentifiable from any experiment, because seeing a unit under one arm tells nothing about what it would have done under the other. So it cannot be estimated, cannot be conditioned on, and cannot be eliminated by any of the devices this field has used on the variance ratios it has met before.

The construction, and the sentence it makes operational

The reference distribution a fair coin gives, and the curve it matches. One 200-patient trial allocated by a fair coin, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the same rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 0.663, 507 of the 999 re-randomisations reach it, and the p-value is (1 + 507)/(1 + 999) = 0.5080. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 1.997.
Fig. 2 The reference distribution this field is built on: the statistic over the experiments the same design could have produced, with the outcomes held exactly as they came out. Holding them is what makes the test exact, and it is also the sentence being assumed.

Every re-randomisation in that picture reuses the observed outcome list. A unit that recorded a 7 records a 7 in every one of them, whichever arm the shuffle sends it to. That is not an approximation or a convenience — it is the hypothesis, written as an instruction, and it is the reason the test needs nothing else.

It is also the whole of what is at issue. If a unit’s outcome would have been different in the other arm, the re-randomised experiments are not experiments the design could have produced under the null being entertained; they are experiments under a stricter null that happens to imply it. The construction cannot tell, because the counterfactual outcome is exactly the thing no experiment contains.

The mechanism, which is a variance and not a mean

The failure has a cause that can be watched directly, and it is not about the average of anything.

Where the failure comes from. The ratio of the two arms' sample variances against how much the treatment effect varies between units, with the average effect held at zero throughout. At no variation the arms are alike and the ratio is 1.01. At a spread of 3 the treated arm has 16.3 times the variance of the control arm, on an experiment where the treatment did nothing on average. The permutation distribution of a difference in means is built by shuffling labels, which makes both arms look alike — so it is a reference distribution for a world in which the arms have the same spread, and that world is not this one.
Fig. 3 The ratio of the two arms’ sample variances against how much the effect varies, with the average effect zero throughout. At no variation the arms are alike; at a spread of 3 the treated arm has 16.3 times the variance of the control arm.

Adding a varying effect to the treated units adds variance to the treated arm and none to the control arm. At a spread of 1 the treated arm’s variance is 4.07 times the control arm’s; at 3 it is 16.26 times. The two arms are not exchangeable, and they are not exchangeable in a way the average effect knows nothing about.

The permutation distribution is built by shuffling the labels. Shuffling makes both arms look alike — every re-randomisation draws units from one common pool — so the reference distribution it produces is the distribution of a difference in means between two arms with the same spread. That is a reference distribution for a world in which the effect is constant, and the statistic is being read against it in a world where it is not.

That is exactly the same defect found in a standard error computed for a model that is wrong and in a t statistic compared against a table its null does not fit. What is unusual here is that the test is exact, and the exactness is not in doubt: it is an exact test of a hypothesis the trial is not asking about.

Why an even split escapes it

The size of the failure depends on something the experimenter chooses, which makes it more than a curiosity.

The split decides whether the wrong null matters. How often each permutation test reports an effect against a true weak null — no average effect, with the effect varying between units at a spread of 2 — against the share assigned to treatment, at 150 units. On the difference in means the rate runs from 4.33% at an even split to 28.67% at 15% treated. On the studentised difference it stays between 4.33% and 6.00% across the whole range. An even split is the one arrangement in which the wrong reference distribution does no damage.
Fig. 4 How often each statistic’s permutation test rejects a true weak null against the share assigned to treatment, at an effect spread of 2. On the difference in means the rate runs from 4.33% at an even split to 28.67% at fifteen per cent treated.

At an even split the rate is 4.33%, which is the level. At 40% treated it is 8.87%, at 30% 15.73%, at 25% 20.60%, at 20% 26.27% and at 15% 28.67%. The failure is entirely a property of the imbalance.

The reason is the arithmetic of a difference in means. Its true variance is v1/n1+v0/n0v_1/n_1 + v_0/n_0 and the permutation distribution’s variance is built from a pooled spread applied to both arms. With n1=n0n_1 = n_0 those two expressions coincide whatever v1v_1 and v0v_0 are — the weights on the two variances are equal, so pooling them changes nothing. With n1<n0n_1 < n_0 the true variance puts more weight on the smaller group’s variance, which here is the treated group, which here is the larger variance. The pooled version under-weights it, the reference distribution comes out too narrow, and the pooled test rejects too often.

The same arithmetic is why an even split is the only protected one. Nothing about it makes the arms exchangeable — the treated arm still has sixteen times the variance at a spread of 3 — so a reader who explains the result by saying the split “balances” the two worlds has the wrong account. What an even split does is make the pooled variance and the true variance two names for the same number, which is a coincidence of the weights rather than a property of the data.

Reversing the imbalance reverses the direction: with the larger group carrying the larger variance the test would be conservative instead. The failure is not a bias towards rejection; it is a mismatch between two weightings, and which way it points depends on which arm is both smaller and noisier.

The two nulls coincide at one end of the sweep

The left-hand end of the first figure is the calibration and it is worth reading as more than that, because it says precisely where the sharp null is enough.

At a spread of zero every unit has the same effect, so “no effect on average” and “no effect for anybody” are the same statement, and the test reads 4.20% — its level. The exactness applies because there is nothing to distinguish the two hypotheses. As the spread grows the two hypotheses come apart and the readings climb: 10.27% at a spread of half a standard deviation, 14.07% at one, 16.67% at one and a half.

So the size of the failure is a monotone function of a quantity nobody can measure, starting from exactly nominal at a value nobody can rule out. There is no reading, no diagnostic and no sample size that places a trial on the curve. What an experiment can do is bound the damage by choosing the split, which is the subject of the figure two below this one, and that is the whole of the control available.

The guarantee that is not in doubt

None of this is an argument that the construction is unsound, and the column that says so is the one worth reading first.

Under the sharp null, both tests are exact at every split. How often each permutation test reports an effect when the treatment changed nothing for anybody, over 1,500 experiments of 150 units at each allocation split, with 249 re-randomisations apiece. Every reading is within a standard error or two of the 5% promised, at splits from even to fifteen per cent treated. This is the guarantee the construction actually makes, and it does not weaken as the split gets uneven — which is what makes the next figure's failure a failure of aim rather than of arithmetic.
Fig. 5 How often each statistic’s permutation test rejects when the treatment changed nothing for anybody, at four allocation splits. Every reading is within a standard error or two of the 5% promised, from an even split to fifteen per cent treated.

Under the sharp null the test reads 4.27%, 4.67%, 3.93% and 4.40% at splits from even to fifteen per cent treated. It does not weaken as the allocation becomes uneven, it does not need a large sample, and nothing about the outcome distribution enters it. The guarantee is intact and it is a guarantee about a different statement.

That distinction matters for how the failure should be described. Nothing has gone wrong with the arithmetic, no assumption has been violated, and no diagnostic could catch it — because there is nothing to catch. The test answers the question it was built to answer, and the question a reader takes the answer to be about is a different one.

It also means the failure is invisible in every direction a reader might look. The p-value is a valid p-value, computed correctly, for a hypothesis that is stated in every textbook account of the method. The construction’s assumptions are met. The arms are balanced on everything that was measured, because they were randomised. The sample is adequate. Nothing in the output, and nothing a referee could ask for, separates a trial in which the two hypotheses coincide from one in which they do not — which is the shape a defect that is an omission rather than an error always has.

Which hypothesis a trial actually means

It is worth being honest that the two hypotheses are not always different in practice, and that the sharp null is sometimes exactly right.

Where the sharp null is the question. A trial asking whether a treatment does anything at all — a first test of a new compound, an experiment where any effect at all would be a finding — is asking the sharp question. Rejecting it means “something happened to somebody”, which is what was wanted. Here the exactness is exactly aimed.

Where the weak null is the question. A trial asking whether a policy is worth adopting is asking about an average, because the average is what a population-level decision is made on. A treatment that helps half the units and harms the other half equally is not worth adopting, and the sharp null is false while the weak null is true. That is the case in which a fifth of experiments produce a significant result.

And a subgroup finding is the sharp null’s territory. A trial whose headline is null and whose subgroup analysis is not is, read charitably, a trial whose weak null held and whose sharp null did not. The permutation test on the whole trial is the right instrument for the second question, and reporting it as the answer to the first is what the figures above price.

And the gap is exactly the case that matters most. Treatments with no effect at all are rare and uninteresting; treatments that help some and harm others are the normal state of the world and the reason subgroup analysis exists. So the region where the two hypotheses differ is not a technical corner — it is where most real questions sit.

What a defensible use looks like

Say which null is being tested. “The randomisation test gives p = 0.03” is a complete sentence with an incomplete meaning. The sharp version is a claim about every unit and the weak version is a claim about an average, and a reader has no way to tell which was computed.

Split evenly when the average is the question. An even split is free of this failure by arithmetic rather than approximately, and there are other reasons to prefer one — it is also the split that minimises the variance when the arms have equal spreads. Where an uneven split is forced, the failure is a function of the imbalance and this sweep prices it.

Report the two arms’ spreads. They are computed on the way to any analysis and they are the one visible symptom of the condition that makes the two hypotheses differ. A treated arm with four times the control arm’s variance is not proof that the effect varies — the arms could differ for other reasons — but it is the reading that should stop a permutation test on a difference in means from being reported as assumption-free. That is the same discipline a robust standard error’s diagnostics ask for, arriving in a design where there was supposed to be nothing to diagnose.

And do not read an uneven-split permutation test as conservative. The folk description of a permutation test is that it is the safe, assumption-free option. At fifteen per cent treated with a varying effect it rejects a true weak null 28.67% of the time, which is nearly six times its stated level — a worse error rate than the large-sample t it was chosen over, which reads 6.60% at the same setting because it estimates the two variances separately and never pooled them in the first place.

What is claimed here and what is not

The effect variation is arranged, not incidental. The units the treatment helps are the ones that would have done well anyway, which is what makes the treated arm’s variance larger rather than merely different. An arrangement in which the effect varies independently of the baseline adds variance too but less of it, and an arrangement anti-correlated with the baseline can reduce the treated arm’s variance and make the test conservative instead. The sweep is at the arrangement that produces the failure, which is stated rather than hidden, and the direction of the arrangement is a quantity nothing can observe.

One hundred and fifty units. The failure is not a small-sample effect and does not go away with more data: the mismatch is between two variance formulas rather than between an approximation and a limit, so a larger experiment estimates both more precisely and rejects a true weak null just as often. What a larger sample would change is the precision of the readings here, not the readings.

The comparison to the ordinary t is not an endorsement of it. The separate-variance t reads between 4.73% and 6.60% across the sweep, which is better than the permutation test on the difference and is not exact anywhere. It makes an asymptotic argument and it happens to make the right one; at a smaller sample it would drift, and it carries none of the sharp null’s guarantee.

The re-randomisation count is 249 and the level is read at 5%. With 249 re-randomisations the p-value takes 250 values and 5% is exactly attainable, so no part of the readings above is the discreteness this field has priced elsewhere. A smaller count would put the achievable levels on a coarser lattice and would move every number by an amount that has nothing to do with the hypothesis being tested, which is why it is stated.

And “exact” is used here in its technical sense throughout. The permutation test’s size under the sharp null is exactly the nominal level up to the discreteness of the p-value, which is what the +1 in the construction delivers and is why the measured readings sit a little below 5% rather than scattered around it. Nothing in this essay weakens that.

Still open: whether the same construction can be aimed at the other null

The failure is a mismatch between a statistic’s true variance and the variance the permutation distribution gives it. That is a statement about the statistic, not about the construction — the re-randomisations, the fixed outcomes and the counting are all doing exactly what they should.

Which suggests the repair is a change of statistic inside the same construction rather than a different one. A statistic that already carries its own standard error, so that its scale does not depend on which arm has the larger spread, would be shuffled the same way and counted the same way and might be read against a reference distribution that fits it. Whether one statistic can be exact under the sharp null and right under the weak one at the same time, and what it costs, is the next thing this field has to measure.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Allocation ratioAverage treatment effectComposite nullError rateExact testExperimental designHeteroskedasticityMonte CarloNuisance parameterPermutation testRandomisation testReference distributionSharp nullTreatment effect heterogeneity