Concept

Exact test — where it appears

A test whose stated error rate is its true rate at every sample size, rather than in the limit of a large one. Exactness is paid for in width or in observations, and a procedure that appears to pay neither is usually one that does not cover.

Named by 9 essays across 8 fields — each of them below, with the objects they name alongside it.

Every way of splitting 16 units into two halves. All 12,870 assignments, enumerated. The spread of the standardised imbalance is exactly 2/√n = 0.500, whatever the covariate's own distribution, and 33.3% of assignments differ by more than 0.5 standard deviations. Randomisation does not deliver balance; it delivers a known distribution of imbalance.

Randomisation is not balance

A third of all ways to split sixteen units leave the two halves more than half a standard deviation apart on a covariate. What randomisation delivers is not balance but a known reference distribution — and it makes a test exact with no assumption about the data's shape at all.

design · Randomisation
Worst and average coverage of six 95% intervals for a proportion, 30 trials. The worst coverage over every proportion beside the average over a uniform one, with the average expected width. Clopper–Pearson: worst 95.05%, average 97.34%, width 0.299. Blaker: worst 95.00%, average 96.31%, width 0.283. Wilson: worst 83.71%, average 95.24%, width 0.271.

What a guaranteed minimum costs

Clopper–Pearson's interval never covers less than 95%, and at thirty trials it averages 97.34% and is 10.4% wider than Wilson's. Blaker's interval keeps the same guarantee, averages 96.31% and is 4.6% wider. The difference is not waste: Clopper–Pearson guarantees each side separately, holding both below 2.5%, and Blaker guarantees only their sum — so at ten trials and a proportion of 0.15 it misses on one side 5.00% of the time.

discrete · Oscillation
Four analyses of the same 3-arm trials, under a true null. 250 trials of 150 patients, 3 arms, minimisation with p = 0.85, 99 re-randomisations for each exact test. Two statistics — an F on the arms alone and an F on the arms after the balanced factors — against two reference distributions: the table the statistic is named for, and the distribution the allocation rule itself generates when the outcomes are held fixed and the rule is re-run. Only the first cell is wrong, and it is wrong in the direction that costs power rather than the one that manufactures findings: 0.0% where 5% is claimed. Either repair works — adjusting for what the rule balanced, or asking the rule what it would have done.

The analysis after three arms

An unadjusted analysis after a two-arm balancing rule rejects 0.6% of true nulls where it claims 5%. With three arms and a deterministic rule it rejects none at all — and the repair is the same repair, which is a sentence and a column in the model.

multiarm · Assignment
How long a walk has to be given. Every equal split of twelve units is enumerated, the 410 admissible ones are found, the transition matrix is built, and the distance from uniform is computed exactly at each step — no simulation anywhere. The walk is started at the least balanced admissible assignment, which is the state a rejection sampler is least likely to have handed it and the one a burn-in has to cover. It is 0.0849 away after twenty steps and 0.00008 after ninety. A real cost, and a small one, and naming it is what stops it being assumed to be zero.

Walking the admissible set

A rerandomisation test hunts for admissible assignments and throws away the rest. A walk visits them instead — and it is exactly uniform only because it stands still when a proposal fails, which is the step that looks like waste.

joint · Randomisation
The interval that over-covers when the instrument fails. Counted coverage of two nominal 95.0% intervals for the same causal effect, read off the same 2000 draws of 200 rows at each first stage. The exact Anderson–Rubin set covers 95.3% at every setting — flat, because the statistic it inverts is built from y − tβ, which contains no π at all, and is therefore the same number on the same draw whatever the instrument is worth. The conventional interval covers 99.1% at π = 0.02 and 95.6% at π = 0.6: it goes wrong at the weak end by covering too MUCH, at a median width of 7.320, because its standard error is computed from residuals taken at an estimate that has itself gone wrong. A weak instrument does not make this interval lie about its coverage; it makes it useless while telling the truth.

What the first stage does not know

A single weak instrument does not make the conventional interval undercover — it makes it cover 99.1% at a width of 7.320. Where the promise actually breaks is many instruments — coverage falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583.

instrument · Exclusion
What each analysis does at a true null, by shape. Four analyses of the same trials — 500 of them at each shape, 120 units, assigned by the rule that reads the covariate. Every rejection is false. The unadjusted analysis is the one that moves: 1.60% against a linear outcome, where the design removed a great deal that the standard error still prices, and 5.20% against a quadratic, where it removed nothing and the standard error is right. Adjusting holds the level in all three columns, and so does the design's own reference distribution, which needs to be told the rule and nothing else.

The analysis and the shape

An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.

shape · Randomisation
The allocations this trial could have made, and the ones it could not. One 120-patient trial allocated by minimisation at p = 1, re-randomised 399 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the rule again over the same patients in the same order with the same recorded factors, so what is drawn is the set of experiments that could have happened. The bars are that set; the outline is what shuffling the labels gives, which is the reference distribution of a coin and is what every off-the-shelf permutation routine assumes. The coin's is wider — its 5% point is 1.95 against the rule's 1.09 — because a coin's allocations are less balanced and a less balanced allocation gives a larger statistic. Reading this trial against it makes the test conservative rather than anti-conservative, which is the opposite error from the outcome-adaptive case and for the same structural reason.

The reference the covariates supply

Hold the outcomes fixed, re-run the rule that assigned them, count. The same construction cost nineteen points of power in the adaptive field, because its rule chased outcomes and its critical value depended on a rate nobody has. Here the rule reads only what was recorded before anything happened, and the same unadjusted statistic goes from 20.3% power to 55.0% by being read against the right distribution.

covadapt · Assignment
An exact test rejecting a true hypothesis a fifth of the time. How often each analysis reports an effect when the average treatment effect is exactly zero and the effect varies between units, at 150 units with 25% treated. The permutation test on the difference in means reads 4.20% where the effect is constant — where the two nulls coincide and its exactness applies — and 22.93% where the effect varies with a standard deviation of 3. The same test on the studentised difference reads 6.27% there, and the ordinary large-sample t, which makes no exactness claim at all, reads 6.60%.

The null the exactness is for

A permutation test is exact under the hypothesis that the treatment changed nothing for anybody. Under the hypothesis it changed nothing on average, with a quarter of the units treated and the effect varying between them, it rejects a true null 22.93% of the time.

exact · Nuisance
One statistic that is right under both hypotheses. Rejection rates for both statistics under both nulls, at 25% of 150 units treated, with the weak-null readings taken at an effect spread of 3. The difference in means is exact under the sharp null and rejects 22.93% of true weak nulls. The studentised difference is exact under the sharp null — 4.07% — and reads 6.27% under the weak one. The repair is a change of statistic inside the same construction: the same re-randomisations, the same fixed outcomes, a different number compared across them.

A statistic that is exact twice

Dividing the difference in means by its own separate-variance standard error before permuting takes the rejection rate under a true weak null from 20.47% to 6.07%, keeps the exactness under the sharp null at 4.07%, and costs 0.8 points of power against a real effect. At an even split it changes nothing at all, in every draw.

exact · Nuisance

Named alongside it

The objects these essays reach for when they reach for this one.

Randomisation testReference distributionPermutation testError rateMonte CarloSharp nullConservative intervalCovariate adjustmentCovariate balanceNuisance parameterStatistical powerAllocation ratio

All concepts