Half a reference distribution
Worth reading first: Randomisation is not balance · The experiments that could have happened.
A walk that samples half its admissible set is sampling half a reference distribution, and a randomisation test’s p-value is a share of a reference distribution. So the p-value is wrong.
It is not. The commonest one is exactly right, to machine precision, on every admissible assignment of every set measured — and the reason is a symmetry rather than an approximation.
The measurement
At fourteen units the admissible set can be enumerated. At the tolerance where it splits it holds 116 assignments in two components of 58, and every assignment’s complement is in the other one.
Take each of those 116 in turn as the observed assignment — which is exactly the population a rejection sampler draws from — and compute three p-values twice: once over the whole set, once over the component the observed assignment sits in, which is all a walk can ever visit.
Two-sided, on the difference in arm means: wrong by exactly nothing. Not close to nothing. The worst difference over 116 observed assignments and five explained shares is 0, at machine precision.
One-sided, on the same statistic: wrong by as much as 0.112. On one worked case the whole set gives 0.526 and the reachable half gives 0.603.
The largest response observed in the treated arm: wrong by as much as 0.172. On the same case, 0.897 against 0.845.
Why the two-sided one survives
The components are complement pairs and the difference in arm means is exactly negated by the complement — that is the fact the whole diagnostic rests on. So if the statistic takes value somewhere in one component, it takes at the mirror assignment in the other.
The multiset of on one component is therefore the multiset of on the other, and a two-sided p-value is a share of assignments with at least as large as the observed one. Two identical multisets give the same share.
It is worth being precise about the strength of that. It is not that the halves are similar, or that the error is small at this sample size, or that it averages out. The two halves carry identical distributions of the absolute statistic, so the two-sided p-value from half the set is the two-sided p-value from all of it, at every observed assignment, at every tolerance, at every trial size, for ever.
Which is why nobody found it
That identity is the answer to the question the whole enumeration was run to ask: if a walk has been sampling half its reference distribution, how did nothing go wrong?
Nothing went wrong because the quantity everybody computes is the one quantity a mirror split cannot touch. Every acceptance rate was in order, every stationary distribution was uniform on what the chain could reach, every diagnostic a chain can run on itself passed — and the number that came out at the end was right.
This collection has a name for the shape and it applies exactly: the symptom is absence. A defect whose only consequence is on quantities nobody computes leaves nothing to see.
The error is an integer over 116
Every error in the measurement above is a count difference divided by the size of the set, which is worth working out because it turns three decimals into a structure.
Write for the number of assignments in the observed component whose statistic is at least the observed , and for the number whose statistic is at most . The other component is the exact negation of this one, so the full set’s one-sided tail is out of 116 and the reachable half’s is out of 58. Subtracting,
The error is an integer over 116, in steps of 0.00862. The worked case’s 0.603 against 0.526 is ; the largest error found, 0.112, is ; and the largest-response statistic’s 0.172 is with its worked case’s 0.052 at .
Four numbers reported to three decimals, all four exactly integers over the size of the enumerated set.
What the bound is, and how far from it the measurement sits
The same expression gives the worst case the construction allows, which is a more useful thing to know than the worst case observed.
Since and are counts within a component of 58, , so the one-sided p-value can be wrong by at most , which is 0.5. The largest error measured is 13 of that possible 58 — 22% of the bound.
So the one-sided statistic is not protected and is not badly damaged either. Its error is bounded at a half by the mirror structure, and on this set it uses a fifth of that.
The largest-response statistic has no such bound. It is not negated by the complement — the largest response among the treated is a different quantity when the arms are swapped — so nothing constrains against and the error is bounded only by 1.
Measured, it is 20/116 against the odd statistic’s 13/116, a factor of 1.54. Which is the right ordering and a much smaller factor than the difference in their bounds suggests: the mirror structure buys the odd statistic a guarantee, and on this particular set it buys it about half a point of actual error.
What is not protected
The protection is a consequence of oddness, and the boundary is sharp.
A statistic that is exactly negated by the complement has a two-sided p-value that survives and a one-sided p-value that does not: a one-sided share counts assignments with at least the observed value, and one component holds the positive tail while the other holds the negative one. The reachable half is the wrong half half the time.
A statistic that is not odd has neither protection. The largest response observed in the treated arm — a safety reading rather than an effect — becomes the largest in the control arm under the complement, which is unrelated to it, and its two-sided p-value is out by up to 0.172.
A third of what looks like a list of statistics is not one. The treated arm’s own mean is the difference in means over two plus a constant that does not depend on the assignment, so its p-values are the difference’s to the last digit — the same statistic in different clothes. The field is smaller than it looks, and the boundary running through it is exactly oddness.
What it costs where it counts
An error of 0.08 in a p-value of 0.60 changes nothing. The reading that decides something is how often the reachable half puts a p-value on the other side of five per cent from where the whole set puts it.
For the one-sided p-value that is 2.6% to 4.3% of observed assignments — about one trial in thirty — and it does not depend much on how much of the outcome the balancing rule explains. For the two-sided one it is zero at every row, as it must be. For the largest-in-arm statistic it is zero as well, but for an uninteresting reason: those p-values are large on this construction and never near five per cent at all, which is a property of the sweep rather than a protection.
One trial in thirty is a rate worth knowing and it is not an emergency. What makes it worth the essay is that it is a rate a practitioner currently has no way to discover, since the diagnostic that would tell them is quiet a quarter of the time on the probe they would run it with.
The direction is not random
One more thing the enumeration gives that a simulation could not: the error has a sign that depends on which component the observed assignment is in.
An observed assignment in the component holding the positive tail sees a reachable distribution shifted towards it, so its one-sided p-value comes out too large. An assignment in the other component sees the opposite. Averaged over the 116 the errors cancel — the mean signed error is essentially zero — which is why nothing about the average would ever reveal this.
That cancellation is itself worth carrying, because it is the reason a simulation study of the whole procedure would find nothing. Draw many trials, compute the one-sided p-value under each, check that the distribution is uniform: it is, because the two directions of error are equally likely and equally large. The defect is invisible to the calibration check that would normally catch a broken reference distribution.
One worked case
The averages hide the shape, so one observed assignment in full.
On the fourteen-unit set at a tolerance of 0.8, with an outcome 36% explained by the balanced basis, take the first admissible assignment a rejection sampler would return. Over the whole set of 116 its two-sided p-value is 0.9655 and over its own component of 58 it is 0.9655 — the same number, not a rounding of two nearby ones. Its one-sided p-value is 0.5259 over the whole set and 0.6034 over the component.
Nothing about the second pair looks wrong. Both are large, both are the kind of number a trial reports as no evidence, and a reader comparing them without the enumeration in front of them would have no reason to prefer either.
That is what makes the enumeration the load-bearing part of this field rather than a nicety. The two p-values differ by 0.08 and there is no internal signal — no diagnostic, no convergence statistic, no disagreement between chains — that says which is the right one.
What a practitioner does
Three sentences.
Report the two-sided p-value on an odd statistic and the defect cannot reach it. That is the usual practice, it is what most trials do, and it is exactly right for a reason nobody was relying on.
A one-sided test, a non-inferiority margin, or any statistic that treats the two arms differently needs the walk checked. Those are not exotic: a one-sided hypothesis is standard in several fields, and a safety reading on the treated arm is standard in all of them.
And check the walk with a probe that can see, not with the one the trial is about. The loudest reading over several probes is the verdict; a quiet reading from one is not.
Why an odd statistic is the usual case
The protection covers a class that is worth describing, because it is larger than it first appears and its boundary is not where intuition puts it.
A statistic is odd under the complement when swapping every unit’s arm negates it. The difference in arm means is; so is a difference in medians, a difference of any location measure computed the same way in both arms, a two-sample t statistic with a pooled variance, a Wilcoxon rank sum measured from the centre, and a covariate-adjusted treatment effect. Almost every effect estimate a trial reports is odd, because an effect is a contrast and a contrast reverses when the labels do.
What is not odd is anything that treats one arm as special: a maximum, a count of events in the treated arm, a proportion above a threshold in one arm, a safety signal. Those are not exotic either — they are the second table of most trial reports.
So the honest summary is that the primary endpoint is protected and the secondary table is not, which is an unusually clean division and not one anybody designed.
The same question one field along
This is not the only place in this collection where a reference distribution is sampled rather than enumerated, and it is worth checking whether the same protection is available elsewhere.
A bootstrap’s reference distribution is drawn rather than walked, so there is no reachability question at all — every resample is available at every draw. The failure mode there is a different one: the distribution being sampled is the wrong distribution, not half the right one.
A specification search’s null is simulated from a fitted model, and there the analogous question is whether the fit is right rather than whether the sampler reaches everything.
So the defect this field is about is specific to a walk over a constrained set, which is the one construction in this collection whose sampler can be stuck. That narrows where the finding applies and it also says where to look for it: any procedure that draws from a restricted randomisation by proposing local moves has the same exposure, and the diagnostic transfers unchanged.
Where the protection ends
The identity holds because the components are complement pairs. That is measured on every set here — the mirror share is 1.000 — and it is not a theorem.
A set that split into components that were not complement pairs would have no such protection, and nothing in this collection has produced one: every split found by enumeration, at every tolerance and every trial size up to twenty-two units, is the mirror pairing. The field that enumerates them reports it as the only structure that has ever come out, and reports it as a measurement rather than an argument.
So the honest statement of the protection is conditional: while the split is the mirror pairing, the two-sided p-value on an odd statistic is exact. If a rule ever produced a different split — a proposal moving more than one pair, a constraint that is not symmetric, an unequal allocation — the protection would have to be re-derived and there is no reason to expect it to survive.
What the enumeration cost
The whole field’s certainty comes from one thing: at fourteen units the 3,432 equal splits can be listed, the admissible ones picked out, the components found by walking single swaps, and every p-value computed as a share rather than estimated as a rate.
That is a small computation and it is the only reason any of the statements above can use the word exactly. A simulation study of the same question would have reported that the two-sided p-value agrees to within simulation error, which is a weaker and much less interesting claim — and it would have missed the cancellation of the one-sided errors entirely, because a cancellation looks like correctness from the outside.
Where the enumeration stops is about twenty-four units, and everything past it is the two-chain diagnostic’s territory. The division of labour is the same one the whole area runs on: enumerate where it is possible so that the test can be checked, and then use the test where enumeration is not.
What is claimed here, and what is not
This essay takes what a walk that reaches half its admissible set does to the number a trial publishes. The claims are that over every one of the 116 admissible assignments of a fourteen-unit set, taken in turn as the observed one, the two-sided p-value on the difference in arm means computed over the reachable component equals the one computed over the whole set to machine precision, at every explained share measured; that a one-sided p-value on the same statistic differs by as much as 0.112 and crosses five per cent on 2.6% to 4.3% of observed assignments; that a statistic which is not odd under the complement — the largest response in the treated arm — differs by as much as 0.172; that the treated arm’s mean is the difference in means over two plus a constant and so is not a separate statistic; and that the signed one-sided errors cancel across the set, so a calibration check on the whole procedure would find nothing.
What stays out, and is named as a decision: the same enumeration at a size a trial would run. Everything here is at fourteen units because everything here is exact, and the arithmetic behind the two-sided identity does not depend on the size — it depends on the split being the mirror pairing, which is measured up to twenty-two units and asserted nowhere. The one-sided error rates are a different matter and are properties of this set at this tolerance; a larger trial’s would have to be measured.
Also out: an estimate rather than a test. Everything here is about a p-value under the sharp null. What a half-reachable walk does to a randomisation-based interval — which inverts a family of tests, not one — is a question with more moving parts and is not answered.
The boundary against the field that found the split is that it establishes the walk reaches half the set and this one asks what that costs the number at the end. The answer is nothing, for the number most trials report, which is why the defect could last.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A defect that is about size — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
- A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test
- A probe chosen from the design — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
- A quantity that loses to a heuristic — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test
- Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test
- The part the rule already took — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
Named objects
A flat tag is an object no other essay names yet.
Assignment mechanismConnected componentCovariate balanceExact enumerationImbalanceMarkov chain Monte CarloOne-sided testp-valuePermutationRandomisation testReference distributionSharp nullSymmetryTreatment effect