Randomisation test — where it appears
Named by 35 essays across 13 fields — each of them below, with the objects they name alongside it.
A proposal that moves more than two units
The walk's autocorrelation is a fact about its step size and not about its acceptance rate. Exchanging three units from each arm mixes nearly twice as fast as exchanging one, and is refused a third more often.
A test rather than a survey
A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever. That was found by enumerating fourteen units, and enumeration stops at twenty-four.
The experiments that could have happened
An adaptive trial's allocation is a function of the outcomes it will later be compared against, so the ordinary analysis rejects a true null 9.2% of the time. Hold the outcomes fixed, re-run the rule that assigned them, and count — the same statistic against a reference distribution the trial could actually have drawn from is back at 4.0%.
The part the rule already took
A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units.
The statistic the p-value is about
The test for whether a balanced-assignment walk reaches its whole set is run on a covariate function chosen before the trial. Run on the difference in arm means it is the same test, and it is about the number the trial publishes.
What the rule blocks
A balancing rule breaks the admissible set into pieces by refusing exchanges. Which exchanges it refuses is computable from the design and the tolerance alone, before any assignment exists — and it makes a probe.
A model and a count
The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.
A probe chosen from the design
The design's own leverage aligns with the separating direction four times better than a random direction in the same subspace. The concentrated direction the argument invites is worse than random.
A probe nobody chose
On a set that is definitively in two pieces, seven of twenty-four outcomes report nothing at all. Every covariate probe reports it. What separates them is not accuracy — it is that one of them can be chosen and the other is what happened.
Stationary is not convergent
A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.
The statistic that changes sign
A test for an unreachable half needs a quantity that tells one half from the other. Every symmetric reading of a mirror pair is identical, and a magnitude is the natural thing to reach for.
The test that needs the rule
A randomisation test assumes almost nothing about the data and one thing about the experiment. Tell it a fair coin produced an allocation that an adaptive rule produced — which is what every off-the-shelf permutation routine does — and it rejects 8.0% of true nulls where knowing the rule gives 4.0%.
A defect that is about size
The admitted share of a rerandomisation barely moves with the number of units. The number of admissible neighbours grows like the square of it, and that is what decides whether the walk can go everywhere.
A quantity that loses to a heuristic
Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.
Half a reference distribution
A walk that reaches half its admissible set reports the two-sided p-value exactly right, to the last digit, for ever. A one-sided one it puts on the wrong side of five per cent about once in thirty.
The analysis after three arms
An unadjusted analysis after a two-arm balancing rule rejects 0.6% of true nulls where it claims 5%. With three arms and a deterministic rule it rejects none at all — and the repair is the same repair, which is a sentence and a column in the model.
The analysis has to know the rule
A trial balanced by minimisation and analysed by comparing the two arms' means rejects a true null 0.6% of the time where it claims 5%, and at full determinism 0.0%. That is not an error anybody complains about — it is a test that has stopped working, paid for by a balance the analysis then refused to use.
The plus one and the round number
A sampled randomisation test counts the observed allocation as one of its own reference draws, and the correction is invisible at B = 19, 39, 59 and 999 — every value anybody uses. At B = 20 the version without it is an 8.00% test where the corrected one is 3.80%, and the convention protecting everybody is a preference for round numbers minus one.
The set a dictionary leaves
A rule constrained on six functions at a loose tolerance leaves a set as thin as one constrained on three at a tight one. Both sampling methods cross over at the same thinness, and the tolerance where that happens moves by a factor of three.
Walking the admissible set
A rerandomisation test hunts for admissible assignments and throws away the rest. A walk visits them instead — and it is exactly uniform only because it stands still when a proposal fails, which is the step that looks like waste.
What a chosen probe finds
On a chain of eight hundred draws the probe the earlier fields use misses 44% of the sets that are split. Its own residual off the rule's span misses 12%, for one least-squares fit.
Where the gain is, and where the decision is
A bigger proposal is worth a factor of six at a loose tolerance and nothing at a tight one. The tolerances where it helps are the ones where a hunt costs two evaluations a draw, and the crossing barely moves.
Before the trial and after
The same diagnostic run at two moments answers two different questions. Before, a positive verdict changes the design. After, it changes which number gets reported — and only for the numbers the defect can reach.
Counting it exactly does not help
If a modelled active set lost because the model was crude, the exact one would win. It is computed at a cost no trial can pay, and it is worse — so the approximation was never what was costing the probe.
Draws that repeat each other
A hunt costs 1/p evaluations per independent draw. A walk costs one per step and yields an effective draw every τ steps. Both are counted in the same unit, and the walk is dearer at every tolerance a trial is designed at.
The analysis and the shape
An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.
The diagnostic at two hundred
Pointed at a trial size no enumeration reaches, the test gives three answers rather than one — and past a certain thinness it stops agreeing with itself, which is the honest reading and the one nothing could give before.
The reference the covariates supply
Hold the outcomes fixed, re-run the rule that assigned them, count. The same construction cost nineteen points of power in the adaptive field, because its rule chased outcomes and its critical value depended on a rate nobody has. Here the rule reads only what was recorded before anything happened, and the same unadjusted statistic goes from 20.3% power to 55.0% by being read against the right distribution.
The walk that cannot cross
A thin enough admissible set is not one set. It splits into an assignment and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples it is uniform on half the reference distribution for ever.
What a reference distribution costs to sample
A randomisation test on a trial too large to enumerate has to sample its reference distribution, at 1/p attempts per draw and a p-value resolved to 1/(B + 1). Six constraints cost 9,878 attempts per thousand draws, and a thousand draws resolve p to 9.99·10⁻⁴ and not one digit finer.
What the exactness buys
Against a z test calibrated to reject exactly 5% of true nulls on this design, the randomisation test loses nineteen points of power. What it buys is that the calibration needs the success rate — which moves the critical value from 1.668 to 2.718 and is the quantity the trial was run to find out.
When the constraints run out
Every function added to a basis is a constraint the assignment has to satisfy with the same units. At sixteen units and a stated tolerance the admissible assignments run 3,874, then 1,006, then 314, then none — and the count is exact, because the assignment space is finite.
A set of pairs, not a vector
The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.
The null the exactness is for
A permutation test is exact under the hypothesis that the treatment changed nothing for anybody. Under the hypothesis it changed nothing on average, with a quarter of the units treated and the effect varying between them, it rejects a true null 22.93% of the time.
A statistic that is exact twice
Dividing the difference in means by its own separate-variance standard error before permuting takes the rejection rate under a true weak null from 20.47% to 6.07%, keeps the exactness under the sharp null at 4.07%, and costs 0.8 points of power against a real effect. At an even split it changes nothing at all, in every draw.
Named alongside it
The objects these essays reach for when they reach for this one.
Reference distributionCovariate balanceRerandomisationMarkov chain Monte CarloAssignment mechanismImbalanceExact enumerationExperimental designConnected componentMonte CarloEffective sample sizeSharp null