Before the trial and after
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
The field has one construction and two moments, and the closing question is what each moment is worth.
Both probes reach the same verdict at every tolerance where the truth is known: quiet at 1.2, 1.0, 0.9 and 0.85 where the admissible set is one piece of 886, 304, 158 and 126 assignments, firing at 0.8, 0.75 and 0.7 where it is two. Neither is more accurate than the other. What differs is when they can be run and what can be done about the answer.
Where the question arises
It arises at a tolerance, and the tolerance is a design choice nobody thinks of as dangerous.
Tightening a balancing rule shrinks the admissible set — 886, 304, 158, 126, 116, 102, 84 assignments of the 3,432 equal splits of fourteen units, as the tolerance goes from 1.2 to 0.7 — and somewhere in that shrinking the set stops being connected under single swaps. On this design it happens between 0.85 and 0.8, between a set of 126 and one of 116.
Nothing about the rule changes there. The constraint is the same shape, the acceptance rate moves smoothly, and every diagnostic the chain can run on itself passes on both sides of it. A practitioner tightening a tolerance from 0.85 to 0.8 for perfectly good reasons — better balance, tighter intervals — crosses a line with no marker on it.
The share admitted, and what it does not mark
Against the 3,432 equal splits of fourteen units, the seven tolerances admit 25.8%, 8.9%, 4.6%, 3.7%, 3.4%, 3.0% and 2.4%.
Read as a share, the crossing is a fall from 3.7% to 3.4%. There is nothing in that step to distinguish it from the ones on either side, and it is the smallest proportional fall anywhere in the sequence — which is the section’s point in numbers: the quantity that changes is connectivity, and connectivity is not a size.
There is a subtler signal in the same counts, and it is worth finding precisely so that its uselessness can be stated. Reading each step as a local power of the tolerance, the set shrinks like from 1.2 to 1.0, like from 1.0 to 0.9, like from 0.9 to 0.85 — and then like across the step where it disconnects, before recovering to 2.0 and 2.8.
So the exponent does dip, and it dips at exactly the right step. It is not a marker anybody can use. It is computed from enumerated counts, which is precisely what a practitioner running a walk does not have; the chain’s own acceptance rate estimates the same shares to sampling error, and 3.7% against 3.4% is a difference no run distinguishes from noise. A signal visible only through the calculation the walk exists to avoid is not a warning.
Why enumeration stops where it does
“Below about twenty-four units” is a threshold with a growth rate behind it, and the rate is what makes the boundary sharp rather than a matter of patience.
Fourteen units give C(14,7) = 3,432 equal splits. Every two units added multiply that by , which is just under four and rising towards it, so the count quadruples for each pair of subjects:
Twenty-four units give 2,704,156. Twenty-six give 10.4 million. Thirty give 155 million. Thirty-four give 2.3 billion.
Which is why the boundary is where it is and why arguing about it is pointless. A tenfold increase in patience buys about three and a half more units, and a trial is not usually within three units of the size somebody wanted. The walk is not a convenience at thirty units; it is the only construction available, and everything this field measures about walks follows from that.
It also fixes the scale of the diagnostic question. The fourteen-unit design in every figure here is enumerable because it was chosen to be — the whole field rests on knowing the truth it is checking a probe against. Nothing here is measured at a size where the probe would actually be needed, and that is a limitation of the evidence rather than of the argument: the two-chain construction’s validity is an argument about complements and sign flips, and it does not depend on fourteen.
Before: the verdict changes the design
Run before the outcomes exist, a positive verdict is actionable in three ways and all three are free.
Loosen the tolerance. The set is connected at 0.85 and split at 0.8, and the cost of the first over the second is a slightly worse expected balance — which the field that prices balancing rules can quantify against whatever the trial is protecting.
Change the proposal. A walk that moves two pairs at a time crosses between components that single swaps cannot, and what that costs in acceptance rate is measured in its own field.
Or stop walking. Below about twenty-four units the admissible set can be enumerated outright, and where that stops is the only reason a walk is used at all.
The diagnostic that supports those decisions is the covariate one, because it is the only one that exists yet. That is the earlier field’s whole justification and it stands.
After: the verdict changes what is reported
Run after the trial, none of those is available. The assignment happened, the outcomes are in, and the design is a matter of record.
What is available is the reading. A trial that knows its walk sampled half the reference distribution knows which p-values are affected, and the answer is unusually clean: a two-sided p-value on any statistic that is negated by swapping the arms is exactly right, and everything else is not.
So the after-the-fact verdict does not invalidate a trial. It partitions the trial’s own results into a part that stands and a part that needs the other half of the reference distribution — which, at a size where enumeration is possible, can simply be computed.
That is a smaller claim than the trial is wrong and a more useful one. It is also a claim a reader can act on rather than only an author: a published trial that reports its balancing rule and its covariates gives a reader everything the diagnostic needs.
What a positive verdict does not mean
Two misreadings are available and both are worth heading off, because the verdict is a blunt word.
It does not mean the assignment was bad. Every assignment the walk produced is admissible: it satisfies the balancing rule exactly as required, and its covariate balance is whatever the rule promised. What the verdict says is that the reference distribution the walk sampled was half of the one the p-value is defined against. The trial’s balance is fine; its null is halved.
And it does not mean the walk was broken. The chain is doing precisely what it was written to do — proposing single swaps, accepting admissible ones, and converging to the uniform distribution on everything it can reach. It converges. It is stationary. It satisfies detailed balance. The set is simply not connected under the move it was given, and nothing the chain can be asked about itself will ever say so.
That is what makes the two-chain construction necessary rather than merely convenient: the fact being tested for is a property of the set and the move together, and no property of one run of the chain carries it.
Which probe at which moment
The two diagnostics are not simply the same test at two times, because one of them can be chosen and the other cannot.
Before, run several covariate probes and take the loudest. They cost nothing, they can be computed from the design alone, and a quiet reading from one of them is not evidence — a probe whose enumerated separation is 0.001 will be quiet on a split set for ever.
After, run the outcome probe and the covariate probes. The outcome probe is the one the p-value is about, which is what makes it worth running; the covariate probes are the ones that can be checked and replaced, which is what makes them worth running beside it. Three in ten outcomes on the set measured here see nothing at all.
The recommendation is the same at both moments — several probes, loudest wins — and only the reason differs. Before, because a better probe may exist. After, because the probe the trial handed over may be one of the blind ones.
The two moments are not equally well served
It is worth being plain that the before-the-trial diagnostic is the better one, and that this field’s contribution is mostly to the worse one.
Before the trial everything is available: the design is known, the rule is known, several probes can be tried, the enumeration can be run if the trial is small, and every repair is free. A practitioner who runs the check then has a decision procedure with no residual uncertainty in it at small sizes and a good one at large.
After the trial the situation is worse in every respect. One outcome, no repairs, and a probe that is blind three times in ten. What the after-the-fact check buys is a partition of the results rather than a fix, and it buys it in a case where the fix is unavailable.
So the honest ordering is: run it before, and if that was not done, run it after. The second is not a substitute for the first and this field does not claim it is. What it claims is that the second is possible at all, valid under the null the p-value is computed under, and worth doing — including by somebody who did not run the trial.
What the field does not cover
Three limits, and the first of them is structural rather than a gap in the measurement.
An unequal allocation breaks the construction. Everything here splits the units evenly, which is what makes an assignment’s complement an assignment of the same shape, and the complement is where the second chain starts. Under a two-to-one allocation the complement is not admissible — it is not even a valid assignment — so there is no second chain to run and no symmetry to test with. That is not a limitation of this essay; it is a limitation of the whole two-chain idea, and no repair is offered.
A constraint that is not symmetric breaks the protection. The two-sided p-value’s exactness rests on the components being complement pairs, which rests on the admissible set being closed under complementation, which rests on the rule admitting a vector exactly when it admits its negative. Every rule in this collection does. A rule that penalised imbalance in one direction more than the other would not, and the whole of the previous essay’s protection would have to be re-derived.
And an estimate is not a test. Everything here is about a p-value under the sharp null. A randomisation-based interval inverts a family of tests rather than one, so what a half-reachable walk does to it is a question with more moving parts and is not answered anywhere in this field.
What this field found that the earlier one did not have
Three things, and they are worth separating from the earlier field’s own findings.
The outcome is a valid probe, under the sharp null, with no new theory — because the sharp null fixes the outcome column and the difference in arm means is then the signed imbalance of it. That is what makes the after-the-fact diagnostic exist at all.
The probe is what decides whether the diagnostic works, and the range is enormous: separations from 0.001 to 4.938 across two dozen outcomes on one set, with seven of them reporting nothing. The earlier field chose its probe on a reasonable-sounding rule — a function the balancing rule was not handed — and the probe it uses sits 98.9% inside the span it was supposed to avoid, so whatever produced the working probe, it was not that rule.
And the defect costs the published number nothing, exactly, for ever, whenever the statistic is odd. That is why nobody found it and it is the most useful sentence in the field: a practitioner reading a two-sided p-value on an effect estimate has nothing to worry about, and everybody else does.
What a trial should actually do
Four lines, in the order they arise.
Before the design is fixed: compute the admissible set if the trial is small enough to enumerate, and if it is not, run several covariate probes on the design and take the loudest. If the verdict is positive, loosen the tolerance or change the proposal — both are cheap and neither costs anything a trial cares about.
When choosing the analysis: prefer a two-sided test on an odd statistic. This is what most trials do anyway, and it happens to be immune.
After the outcomes are in: run the diagnostic again, on the outcome and on the covariate probes together. A quiet reading from one probe is not a verdict.
And if the verdict is positive after the fact: report the two-sided p-value on the primary endpoint as it stands, and recompute anything one-sided or arm-specific over the whole admissible set if the size allows it. What is not available is a repair by re-running the walk, because a longer walk does not reach a component it cannot enter.
Where the four essays leave it
The first establishes that the outcome is a probe: under the sharp null the outcome column is fixed, the difference in arm means is its signed imbalance, and the two-chain test applies unchanged. Checked against an enumeration at seven tolerances.
The second finds that a probe’s power is a property of the set rather than of the run — separations from 0.001 to 4.938 — and that seven of twenty-four outcomes see nothing on a set that is definitively split.
The third enumerates what the defect costs the published number and finds it is exactly nothing for a two-sided p-value on an odd statistic, as much as 0.112 for a one-sided one, and 0.172 for a statistic that is not odd.
This one asks what to do about it, and the answer divides on when the question is asked.
One thing is worth doing at every stage and costs nothing: record the tolerance and the proposal. The verdict depends on both, neither is usually reported, and a reader who has them can run the whole check from outside. A trial that reports “balanced by rerandomisation” has told a reader nothing they can act on; one that reports the rule, the tolerance and the move has told them everything.
One consequence for anybody writing a protocol is worth stating plainly. The diagnostic is cheap, it needs no data, and its inputs — the units’ covariates, the balancing rule and the tolerance — all exist before randomisation. There is no reason a trial should reach its analysis without knowing whether its sampler could reach the whole admissible set, and the only thing standing between most trials and that answer is that nobody has asked.
What is claimed here, and what is not
This essay takes what each of the two diagnostics is for. The claims are that both probes reach the enumerated verdict at all seven tolerances measured, quiet where the set is one piece of 886, 304, 158 and 126 assignments and firing where it is two of 58; that the set stops being connected between tolerances of 0.85 and 0.8 with nothing else about the rule changing; that a positive verdict before the trial admits three repairs and a positive verdict afterwards admits none of them; that after the fact the verdict partitions the reported numbers rather than invalidating them, because a two-sided p-value on an odd statistic is exact on half a reference distribution; and that the recommendation at both moments is several probes with the loudest taken, for two different reasons.
What stays out, and is named as a decision: an unequal allocation. The complement of an assignment under a two-to-one split is not an assignment, so the second chain has nothing to start from and the whole construction has no analogue. Naming a repair would mean inventing one, and none has been tested.
Also out: a recommendation about tolerance. The set splits between 0.85 and 0.8 on this design, with these fourteen units and this basis, and there is no reason to think that number transports. What transports is the method: compute where it splits for the design in hand, which is an enumeration at small sizes and the diagnostic at large ones.
The boundary against the field that found the split is that it establishes the defect and this field asks what it costs and who can see it. Both answers turned out to be smaller than expected and neither is nothing.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A defect that is about size — both name assignment mechanism, connected component, covariate balance, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
- A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- A quantity that loses to a heuristic — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- The diagnostic at two hundred — both name connected component, covariate balance, ergodicity, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
- The statistic that changes sign — both name connected component, covariate balance, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
Named objects
A flat tag is an object no other essay names yet.
Assignment mechanismConnected componentCovariate balanceErgodicityExact enumerationImbalanceMarkov chain Monte Carlop-valueRandomisation testReference distributionRerandomisationSharp nullStudy designTreatment effect