The reference distribution the design supplies

A statistic that is exact twice

Dividing the difference in means by its own separate-variance standard error before permuting takes the rejection rate under a true weak null from 20.47% to 6.07%, keeps the exactness under the sharp null at 4.07%, and costs 0.8 points of power against a real effect. At an even split it changes nothing at all, in every draw.

Worth reading first: The experiments that could have happened.

The essay before this one left the failure with a precise address. The permutation construction — fix the outcomes, re-run the assignment, count — is doing exactly what it should. What fails is the statistic: a difference in means whose true variance is v1/n1+v0/n0v_1/n_1 + v_0/n_0, read against a reference distribution built by shuffling labels, which gives both arms a common spread and therefore the wrong variance whenever the arms differ and the split is uneven.

That is a complaint about a scale, and a statistic that carries its own scale does not have it. It is the same repair a fixed-width interval buys with a variance ratio, arriving where there is no ratio to be had and the scale has to be estimated inside every draw. Divide the difference by its own separate-variance standard error before permuting — the same shuffles, the same fixed outcomes, a different number compared across them — and the reference distribution is a distribution of a quantity whose scale has already been removed.

One statistic that is right under both hypotheses. Rejection rates for both statistics under both nulls, at 25% of 150 units treated, with the weak-null readings taken at an effect spread of 3. The difference in means is exact under the sharp null and rejects 22.93% of true weak nulls. The studentised difference is exact under the sharp null — 4.07% — and reads 6.27% under the weak one. The repair is a change of statistic inside the same construction: the same re-randomisations, the same fixed outcomes, a different number compared across them.
Fig. 1 Rejection rates for both statistics under both nulls, at a quarter of 150 units treated. The difference in means is exact under the sharp null and rejects 20.47% of true weak nulls. The studentised difference is exact under the sharp null — 4.07% — and reads 6.07% under the weak one.

One statistic, inside one construction, that is exact under the hypothesis the construction was built for and close to right under the hypothesis the trial is usually asking about.

Why the scale is what mattered

The difference in means and the studentised difference are the same quantity up to a divisor, so it is worth being precise about why swapping one for the other repairs anything.

A permutation test compares the observed statistic against the same statistic computed on shuffled labels. Shuffling destroys the treatment assignment, so every shuffled experiment has both arms drawn from one common pool — and for a difference in means, the spread of that pool is the only spread there is. The reference distribution’s variance is therefore a pooled variance, and the observed statistic’s variance is not.

For the studentised difference the divisor is computed inside each shuffle. A shuffled experiment with an unusually noisy treated group gets an unusually large standard error and a correspondingly smaller statistic. The scale is removed from each draw separately rather than assumed common across them, which is what makes the reference distribution a distribution of something that no longer depends on which arm carries the variance.

This is the same device that makes a t interval work where a z interval does not, and it is the same one the inverse-variance weights of this field turn on: a quantity whose distribution does not depend on the nuisance parameter can be compared against a fixed reference, and a quantity whose distribution does cannot. What is different here is that the reference is generated rather than tabulated, so the property is bought by construction instead of by a theorem.

The repair, swept

The single comparison above is one setting, and the shape across the sweep is what says the repair is a repair rather than a coincidence.

An exact test rejecting a true hypothesis a fifth of the time. How often each analysis reports an effect when the average treatment effect is exactly zero and the effect varies between units, at 150 units with 25% treated. The permutation test on the difference in means reads 4.20% where the effect is constant — where the two nulls coincide and its exactness applies — and 22.93% where the effect varies with a standard deviation of 3. The same test on the studentised difference reads 6.27% there, and the ordinary large-sample t, which makes no exactness claim at all, reads 6.60%.
Fig. 2 How often each statistic’s permutation test rejects a true weak null, against how much the effect varies between units. The unstudentised column rises from 4.20% to 22.93%. The studentised column reads 4.33%, 4.47%, 4.93%, 5.47%, 6.07% and 6.27% across the same range.

The studentised column is not flat at exactly 5% and is not claimed to be. It drifts upward — 4.33% at no variation, 6.27% at a spread of 3 — and that drift is the honest limit of the repair: the studentised statistic’s permutation distribution is not exactly the right one under the weak null, only asymptotically right, so a finite sample leaves a residue. A residue of one and a quarter points against a failure of eighteen is the trade being offered.

The split decides whether the wrong null matters. How often each permutation test reports an effect against a true weak null — no average effect, with the effect varying between units at a spread of 2 — against the share assigned to treatment, at 150 units. On the difference in means the rate runs from 4.33% at an even split to 28.67% at 15% treated. On the studentised difference it stays between 4.33% and 6.00% across the whole range. An even split is the one arrangement in which the wrong reference distribution does no damage.
Fig. 3 The same comparison against the allocation split, at an effect spread of 2. The unstudentised rate runs from 4.33% to 28.67% as the split becomes uneven; the studentised rate stays between 5.20% and 6.00% across the whole range.

Across the split it is flatter still. The unstudentised test degrades monotonically with the imbalance and the studentised one does not degrade at all — 5.33%, 5.27%, 6.00%, 5.53% and 5.20% at 40%, 30%, 25%, 20% and 15% treated. Whatever residue the repair leaves is a function of the effect variation rather than of the design, which is the better of the two dependencies to have: the design is chosen and the variation is not.

The exactness that is not given up

The claim has two halves and the second is the one that makes the first worth having. A statistic can be made valid under the weak null by abandoning the construction entirely — read a separate-variance t against a normal table and the rate is between 4.73% and 6.60% across the whole sweep, which is fine and makes no exactness claim anywhere.

Under the sharp null, both tests are exact at every split. How often each permutation test reports an effect when the treatment changed nothing for anybody, over 1,500 experiments of 150 units at each allocation split, with 249 re-randomisations apiece. Every reading is within a standard error or two of the 5% promised, at splits from even to fifteen per cent treated. This is the guarantee the construction actually makes, and it does not weaken as the split gets uneven — which is what makes the next figure's failure a failure of aim rather than of arithmetic.
Fig. 4 Both statistics under the sharp null at four allocation splits. The studentised version reads 4.27%, 4.67%, 4.07% and 4.80% from an even split to fifteen per cent treated — the same exactness the unstudentised version has, unweakened.

What the studentised permutation test keeps is the guarantee the t does not have: under the sharp null it is exact at every split, at every sample size, whatever the outcomes look like. So the trade is not between exactness and validity. It is a statistic that holds one guarantee outright and buys most of a second one, against a statistic that holds the first and has nothing to say about the second.

That is why the repair belongs inside the construction rather than beside it. Abandoning the re-randomisations would give up the property this whole field exists for — a p-value that needs no population — in exchange for a fix that can be had without giving anything up.

What it costs

A repair that is free under the conditions it repairs is worth nothing if it is expensive otherwise, and the case to check is the one the unstudentised test is exactly right for: a real average effect, the same for every unit.

The repair is free where nothing needed repairing. How often each permutation test finds a real average effect, with the effect the same for every unit — the case the unstudentised test is exactly right for. The two curves are within 0.8 points of each other at every setting: 33.33% against 32.53% at an effect of 0.3, and 88.73% against 88.47% at 0.6. Whatever the studentised version gives up, it is not power here.
Fig. 5 How often each statistic’s permutation test rejects against a real average effect, with the effect constant across units. The two curves are within 0.8 points of each other everywhere: 33.33% against 32.53% at an effect of 0.3, and 88.73% against 88.47% at 0.6.

Nothing measurable. At an effect of 0.15 standard deviations the two read 12.67% and 12.73%; at 0.3, 33.33% and 32.53%; at 0.45, 65.87% and 65.27%; at 0.6, 88.73% and 88.47%; at 0.8 they are identical at 98.27%. The largest gap anywhere is under a point and it is not consistently in one direction.

That is a stronger result than it looks. Studentising normally costs something — a statistic divided by an estimated quantity is noisier than one divided by a constant, and the loss is what a t distribution’s heavier tails are. Here it costs nothing, because the divisor is applied to the observed statistic and to every re-randomised one, so the extra noise is in the reference distribution as well and cancels.

What the reference distribution now is

It is worth saying plainly what the studentised permutation test’s reference distribution is a distribution of, because the phrase “the permutation distribution” is used for both and they are different objects.

For the difference in means it is the distribution, over label shuffles, of a quantity measured in the outcome’s units. Its spread is a property of the pooled data and has nothing to do with which arm is noisy, which is why the observed statistic can sit in its tail for a reason the distribution knows nothing about.

For the studentised difference it is the distribution, over the same shuffles, of a dimensionless quantity. Each shuffle computes its own divisor, so a shuffle that happens to put the noisy units together produces a large standard error and a statistic no larger than one that did not. The spread of the reference distribution is therefore close to one regardless of how the arms differ, and “close to one regardless” is the whole of what a pivotal quantity is.

The construction never changed. What changed is that the thing being counted is a quantity whose scale was removed before the counting, which is the same move a t interval makes on a sample whose spread is unknown — except that here no distribution has to be assumed for the divisor, because the shuffles supply its distribution too.

The split where the repair is not a repair

One reading of the sweep needs an explanation and it turns out to be an exact fact rather than a near-coincidence: at an even split the two tests give the same answer.

At an even split the two tests are one test. How often the two permutation tests return exactly the same p-value on the same experiment, over 600 experiments of 150 units at each split. At an even split they agree in every single draw and the largest disagreement is 0. The reason is arithmetic rather than approximate: with equal group sizes the within-group sum of squares is the total minus a constant times the squared difference in means, so the studentised statistic is a strictly increasing function of the difference — and a permutation test compares ranks, which a monotone function does not change. Away from an even split the relation breaks and so does the agreement.
Fig. 6 How often the two tests return exactly the same p-value on the same experiment. At an even split they agree in all 600 draws with a largest disagreement of 0. At a 40% split they agree 3.67% of the time and disagree by as much as 0.128.

The count is worth being literal about. The two p-values are equal — not close, not equal in distribution, but the same number on the same experiment — in every one of six hundred draws at an even split, and the largest absolute disagreement across all of them is 0. That is a check on the algebra rather than a measurement of anything: an identity that held approximately would not produce a column of exact zeros, and a single non-zero entry would falsify the derivation below.

Six hundred of six hundred, with a largest gap of exactly zero. The reason is arithmetic. With equal group sizes the within-group sum of squares equals the total sum of squares minus a constant times the squared difference in means, so the standard error is a strictly decreasing function of that difference and the studentised statistic is a strictly increasing function of it. A permutation test compares ranks, and a strictly increasing function does not change a rank — so the two tests order the same experiments the same way and their p-values are equal draw by draw.

Away from an even split the identity breaks immediately and completely: at 40% treated the two agree on 3.67% of experiments, which is about what two different statistics agree on by accident given a discrete p-value.

That is worth stating as a practical rule rather than as an observation. An even split is already studentised. A trial that allocates evenly gets the repair for free and cannot tell the difference between the two analyses; a trial that does not needs to choose, and at fifteen per cent treated the choice is between 5.20% and 28.67%.

What a defensible procedure looks like

Report both statistics where the split is uneven. They are computed on the same shuffles, so the second costs one extra number per re-randomisation, and the gap between them is the clearest available signal that the effect varies — a large gap means the arms’ spreads differ enough to matter, which is a finding about the treatment rather than a technicality. That makes it the same kind of instrument as a diagnostic run before the standard error rather than after it.

Prefer the studentised version even where the split is even. It changes nothing there, so the cost is zero and the benefit is that the analysis plan does not have to be revised if the realised split turns out uneven — which it does whenever units drop out, and dropout is not randomised. An analysis specified on the difference in means is an analysis whose validity depends on a balance nobody controls after the fact.

Studentise by default. It is exact where the unstudentised test is exact, it is nearly right where the unstudentised test is badly wrong, it costs no measurable power, and at an even split it is literally the same test. There is no setting in the sweep where the unstudentised version is preferable.

Say which null the result is about anyway. The studentised test is exact under the sharp null and approximately valid under the weak one, which is two different strengths of claim about two different hypotheses. A reader who wants the sharp statement gets an exact answer; a reader who wants the weak one gets an answer with a residue in it, and the residue is a function of an unmeasurable quantity.

Split evenly where the design allows it, and then the choice does not arise. An uneven split is chosen for reasons — a scarce treatment, an expensive arm, a desire for more precision about the control — and each of those is a real reason. What the identity says is that an even split buys, for free, immunity to the whole problem this pair of essays is about, and that is a consideration to put beside the others rather than a rule.

And do not expect the residue to vanish with the sample. It is an asymptotic property being used at a finite size, so it shrinks — but the sweep above is at 150 units and reads 6.07%, and whether that is small enough is a judgement about the study rather than a fact about the method. Reporting the split and the two arms’ variances lets a reader make it.

The pattern this shares with the rest of the field

The repair is a change of statistic inside an unchanged construction, and that is the shape several of this field’s repairs take — which is worth collecting, because it says where to look when the next one fails.

The randomisation test’s guarantee comes from the procedure: fix the outcomes, re-run the rule, count. Every failure this field has found has been a failure of what is counted rather than of the counting. Feeding the test the wrong rule changes the reference set and breaks it; dropping the plus one changes the count and breaks it; and here, comparing the wrong number across the shuffles points it at the wrong hypothesis.

In each case the repair is local: use the right rule, keep the plus one, divide by a standard error. None of them touches the construction, and none of them needs a distributional assumption to be introduced. That is an unusual property for a family of repairs and it is the reason the construction is worth the arithmetic it costs — which is B re-randomisations for one p-value.

What is claimed here and what is not

One hundred and fifty units, and the residue is a finite-sample residue. Nothing above argues that the studentised permutation test is exact under the weak null, because it is not: its validity there is asymptotic and the 6.07% is what an asymptotic argument looks like at 150 units. A larger experiment would read closer to 5% and a smaller one further from it, and the sweep is at a size where the residue is visible rather than at one flattering to the method.

The standard error is the separate-variance one. v1/n1+v0/n0v_1/n_1 + v_0/n_0 with each variance computed within its own arm, not a pooled version. A pooled standard error is itself a function of the difference in means and would reintroduce the problem: the whole of the repair is that the divisor tracks whichever arm is noisy, and a pooled divisor by definition does not.

Two hundred and forty-nine re-randomisations. The p-value therefore takes 250 values and 5% is exactly attainable, so nothing above is discreteness. A larger count would tighten every reading and change none of the comparisons, since both statistics are computed on the same shuffles.

The residue is measured at one arrangement of the effect variation. The effects here are arranged so that the treated arm is the noisy one, which is the arrangement that breaks the unstudentised test. An arrangement in which the control arm ends up noisier would make the unstudentised test conservative and would leave the studentised one where it is, so the sweep is at the setting where the comparison is hardest on the repair — and the repair’s residue is a property of how much variation there is rather than of which way it points.

“Costs nothing” is a statement about 1,500 experiments per setting. The standard error of a measured power near 33% at that count is about 1.2 points, so a difference of 0.8 points is inside the noise and the honest claim is that no cost is visible at this resolution rather than that none exists. The direction is not consistent across the sweep either, which is what a difference of zero looks like when it is measured.

And the identity at an even split is exact and is about this design. It rests on equal group sizes and on the statistic being a difference in means divided by a separate-variance standard error. A design with unequal group sizes, a covariate-adjusted statistic, or a stratified analysis with different splits in different strata breaks the algebra, and the 600-of-600 count is the check that the algebra as stated is right rather than a claim about permutation tests in general.

Still open: what the same repair does to a covariate-adjusted statistic

The repair is a change of divisor and it leaves everything else in the construction alone, which suggests it should compose with whatever else the statistic is doing. A trial that adjusts for a baseline covariate is comparing a regression coefficient rather than a difference in means, and the same question arises in the same form: the coefficient’s true variance depends on how the effect varies with the covariate, and the permutation distribution gives it a common one.

There is reason to think the answer is not identical. An adjusted estimator’s variance involves the covariate’s design matrix as well as the two arms’ spreads, so the even-split identity above has no obvious counterpart — the algebra that made the two statistics monotone functions of each other used the fact that the only thing being subtracted was a mean. Whether the studentised version is still exact under the sharp null with an adjustment in it, what its residue is under the weak null, and whether an even split still protects, is a measurement this field has not made and the one that decides whether the rule above generalises past the simplest design.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Allocation ratioClosed formExact testHeteroskedasticityMonte CarloNuisance parameterPermutation testPivotal quantityRandomisation testReference distributionSharp nullStatistical powerStudentised residualTreatment effect heterogeneity