The statistic the p-value is about
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
The two-chain test asks whether the walk that draws a balanced assignment can reach the whole admissible set, and it asks it on a covariate function the balancing rule was not handed. That choice is deliberate, it is stated as a limitation in the field that makes it, and the reason given is good: the diagnostic runs before the outcomes exist, which is the moment a practitioner can still loosen the rule or change the sampler.
What the same test says on the difference in arm means is a different question, and it is the one the p-value is actually about.
The test carries over unchanged
A randomisation test conditions on the sharp null: every unit’s outcome is what it is whatever arm it lands in. Under that null the outcome is a fixed vector — not a random one, not a model’s prediction, a column of numbers that does not move when the assignment does.
The difference in arm means is then
the signed imbalance of that column, which is exactly the object the covariate probe is with a different column in it. It is negated exactly by the complement, which is what the two-chain construction rests on. Nothing has to be re-derived.
The identity is checked rather than asserted: on twenty admissible assignments of a fourteen-unit trial, the signed probe on the outcome column and the difference in arm means agree to twelve decimal places, up to the standardisation the probe applies. That check exists because the two claims about p-values in this field would both be satisfied by an enumeration of the wrong object.
Both probes agree with the truth
At fourteen units the admissible set can be enumerated, so the answer is known and the diagnostic can be scored rather than trusted.
At tolerances of 1.2, 1.0, 0.9 and 0.85 the set is connected — 886, 304, 158 and 126 assignments in one piece — and both probes are quiet, at statistics between −1.8 and 2.0 against a threshold of 4. At 0.8, 0.75 and 0.7 the set is in two mirror components, and both fire: the covariate probe at 9.5, 13.3 and 14.6, the outcome probe at 9.0, 10.5 and 12.2.
Seven tolerances, two probes, fourteen verdicts, all correct.
That is what makes the outcome probe the same test rather than a resemblance of it. Whatever else follows about how well it works — and the next essay is entirely about that — it is not a different diagnostic with its own theory.
What the split is
The defect being tested for is worth restating, because everything in this field is about what it does to a number rather than about what it is.
Every balancing rule in this collection is symmetric under the joint sign flip: the imbalance of is minus the imbalance of , and both the box rule and the ellipsoid admit a vector exactly when they admit its negative. So the admissible set is closed under complementation, whether or not the walk can get there.
A walk that moves by single swaps cannot always get there. At a tight enough tolerance the 116 admissible assignments of a fourteen-unit trial fall into two components of 58, and every assignment’s complement is in the other one — so a chain started anywhere is uniform on its own half for ever, and every diagnostic a chain can run on itself passes.
Ten assignments is what disconnects it
The seven tolerances are close enough together to say how abruptly the defect arrives, and the answer is that it arrives all at once.
At a tolerance of 0.85 the admissible set has 126 members and is one piece. At 0.8 it has 116 and is two pieces of 58. Ten assignments removed — eight per cent of the set — is the whole distance between a walk that mixes and a walk that is confined to half of what it is sampling.
There is no gradual stage in between. Connectivity is a property of the graph rather than of its size, so nothing in the counts warns that the boundary is near, and the four connected tolerances give a sequence — 886, 304, 158, 126 — whose only visible feature is that it is shrinking.
That is why the diagnostic has to be a test rather than a look at the numbers. The quantity a practitioner can see moves smoothly through the point where the quantity that matters changes discontinuously, which is the same shape as every other absence this collection tests for.
How fast the set shrinks, and what that says about its dimension
The counts also say how many constraints the rule is really imposing, and the answer is what the rule was written to impose.
Between the two tightest connected tolerances the set falls from 158 to 126 as the tolerance goes from 0.9 to 0.85, an exponent of log(1.254) / log(1.0588) = 4.0. A tolerance that is a half-width in each of d balanced dimensions admits a count proportional to its d-th power once the region is small enough for the density inside it to be flat, so an exponent of four is the four balanced factors, counted from outside.
At the loose end the exponent is larger — about 5.9 between 1.2 and 1.0, and 6.2 between 1.0 and 0.9 — because there the region is wide enough that the imbalance distribution’s own falling density adds to the volume effect. The exponent settling towards the dimension as the tolerance tightens is the signature that says the count is being driven by geometry rather than by the shape of the law.
The outcome probe is quieter, and only slightly
On the three split tolerances the two probes can be compared directly. The outcome probe reads 9.0, 10.5 and 12.2 where the covariate probe reads 9.5, 13.3 and 14.6 — ratios of 0.95, 0.79 and 0.84, averaging 0.86.
So the outcome the trial happened to produce is about six sevenths as loud as a probe chosen to be loud, on the sets where both fire. Against a threshold of 4 that leaves margins of 5.0, 6.5 and 8.2 rather than 5.5, 9.3 and 10.6 — smaller, and nowhere near the threshold.
Which is worth holding beside the claim two sections down that three outcomes in ten report nothing at all on a set that is definitively split. Both are true and they are about different features of the same distribution: the typical outcome is nearly as informative as a chosen probe, and the distribution over outcomes has a long lower tail that a single reading cannot show. A practitioner gets one draw from it and has no way to tell which part of it they got.
Why the outcome is a probe nobody chose
The covariate probe is chosen. It is a function of the recorded covariates, it can be picked to be anything the rule was not handed, and if it comes back quiet a second one can be tried.
An outcome is what the trial produced. There is one of it, it arrives after every design decision has been made, and a practitioner who does not like what it says about the walk has no second one.
That difference has no consequence for validity — the section above settles that — and it has a large one for power, which is the next essay’s whole subject. The short form is that on a set which is genuinely split, three in ten outcomes report nothing at all, and every covariate probe tried reports the split.
What it can and cannot repair
The two diagnostics run at two moments and what they enable is different at each.
Before the trial, a positive verdict is actionable in the strongest sense: loosen the tolerance until the set is connected, change the proposal so it moves more than one pair at a time, or abandon the walk and enumerate. All three are design changes and all three are free, because nothing has happened yet. A proposal that moves more than two units is the cheapest of them and is measured in its own field.
After the trial, none of those is available. What is available is the reference distribution: a practitioner who knows the walk reached half the set knows which half, and can either enumerate the other half where that is possible or report the statistic whose p-value is unaffected. Which statistics those are is a whole essay’s worth of answer and it is a stranger one than it sounds.
The outcome is built to order
Everything measured in this field uses an outcome constructed with a stated share of its variance inside the span of what the balancing rule was told to hold. That is a modelling choice and it is worth being explicit about why it is the right one.
A real outcome is a function of the covariates plus noise, and how much of it the covariates explain is exactly the quantity a practitioner has an opinion about. Constructing outcomes across that range — from a tenth of the variance explained to nineteen twentieths — is what makes the sweep a sweep over a stated quantity rather than over a knob.
The realised share is measured rather than assumed, because on fourteen units a random vector already has about a third of its variance inside any four-dimensional span by dimension counting alone. So an outcome built with none of its signal in the span reports 25% explained, and the axis in every figure here is the measured share rather than the target.
What the sharp null is doing
The whole construction rests on one assumption and it is worth being honest about how strong it is.
Under the sharp null the outcome column is fixed, which is what makes it a probe. Under an alternative it is not: a unit’s outcome depends on which arm it got, so the column changes with the assignment and the signed imbalance is no longer a fixed function of the walk’s state.
That sounds like a restriction and it is nearly the opposite. The sharp null is precisely the hypothesis a randomisation test conditions on — it is what the p-value is computed under — so the diagnostic is valid exactly where the p-value it is diagnosing is defined. The test with no table is the field that establishes what that conditioning buys, and what it buys here is that no extra assumption is needed at all.
Where it does bite is in interpretation. A large statistic says the walk did not reach the whole set under the sharp null’s reference distribution, which is the reference distribution the p-value uses. It says nothing about a walk’s behaviour under an alternative, and nothing about estimation.
Where the two probes come apart, and where they do not
It is worth being precise about the sense in which the outcome probe is “the same test”, because two different things could be meant and only one of them is true.
The construction is the same. Two chains, one from an admissible assignment and one from its complement, compared on a signed statistic at effective standard errors, with a replicate chain from the same starting point to say which of two explanations a large statistic has. Every word of that is the covariate diagnostic’s with a column swapped, and the column is handed in rather than named — which is why the shared file takes a column at all.
The power is not the same, and it is not systematically worse either. Both probes are testing the same property of the same set, and how loudly each one reports it depends on how far apart the two components are on that probe. That is an enumerable quantity at fourteen units and it is what the next essay measures, and the range across outcomes is enormous.
The threshold, and why it is four
The verdict rule compares the standardised difference between the two chains against a threshold of four, and four is a convention rather than a calibration.
The statistic is a difference of two means over their own standard errors, so under a connected set it is approximately standard normal and four is about a one-in-sixteen-thousand event. That margin is deliberately wide, because the standard errors are computed from effective draws rather than from the number taken, and the effective sample size is itself estimated — which the covariate field found is the least reliable quantity in the whole construction, disagreeing between two chains from the same starting point by more than three standard errors.
A wide threshold is the right response to an unreliable denominator, and it is why the field reports a verdict rather than a p-value. The numbers in this essay clear it by factors of two to three, and the ones in the next clear it by factors of twenty-four or not at all.
Why this had to be asked at all
The covariate diagnostic is complete and it works, so a reader is entitled to ask what the outcome version adds.
Three things, and the third is the one that made the field.
It answers the question a reader of the earlier field will have asked. A diagnostic about a reference distribution that is demonstrated on a quantity the reference distribution is not about invites exactly one question, and leaving it unanswered is the kind of gap that becomes a claim by default.
It is the only version available to somebody reading a published trial. The covariates of a finished trial are often reported and the balancing rule sometimes is; a reader who wants to know whether the reported p-value came from half a reference distribution has the outcome and the rule, and nothing else. Whether the diagnostic can be run from the outside decides whether it is a design tool or a reviewing tool, and it is both.
And it turns out to have a failure mode the chosen probe does not. That is not what the field was looking for and it is what it found: an outcome is a probe nobody selected, and a quarter of them cannot see the defect at all.
And what a positive verdict is worth
The last thing worth saying before the field turns to power is what the verdict buys, because a diagnostic that cannot change anything is a diagnostic nobody runs.
After the trial, a positive verdict says the reported p-value came from a walk that sampled half its reference distribution. Whether that matters depends entirely on which p-value was reported, and the answer is not the one anybody expects: the commonest one is exactly right, to the last digit, for a reason that is a symmetry rather than an approximation.
A last observation about scope. Every set measured here is small enough to enumerate, which is what makes the verdicts checkable, and no trial is that size. What carries to a real trial is the construction rather than the numbers: two chains, one from an assignment and one from its complement, compared on a signed statistic at effective standard errors, with a replicate chain to say which of the two explanations of a large statistic applies. The verdicts here are what say the construction reaches the right answer where the right answer is known.
What is claimed here, and what is not
This essay takes whether the walk’s reachability diagnostic can be run on the statistic a trial reports. The claims are that under the sharp null the outcome column is fixed, so the difference in arm means is the signed imbalance of a column and is exactly negated by the complement; that the signed probe on the outcome and the difference in arm means agree to twelve decimal places on twenty admissible assignments, up to the probe’s standardisation; that at seven tolerances of a fourteen-unit rule both probes reach the enumerated verdict, quiet at 1.2, 1.0, 0.9 and 0.85 where the set is one piece and firing at 0.8, 0.75 and 0.7 where it is two; and that the outcome probe’s statistics there are 9.0, 10.5 and 12.2 against the covariate probe’s 9.5, 13.3 and 14.6.
What stays out, and is named as a decision: an unequal allocation. Every rule here splits the units evenly, which is what makes an assignment’s complement an assignment of the same shape and the whole symmetry argument available. Under a two-to-one allocation the complement is not a valid assignment at all, the admissible set is not closed under it, and the two-chain construction has nothing to start its second chain from. That is a real limitation of the diagnostic rather than of this essay, and no repair for it is offered.
Also out: an outcome with a treatment effect in it. The construction is valid under the sharp null and the sharp null is what a randomisation test conditions on, so nothing here needs an alternative. What a large statistic would mean on a trial with a real effect is a different question and is not answered.
The boundary against the covariate diagnostic is that it is run before the outcomes exist and this one after, and the two reach the same verdict at every tolerance where the truth is known.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A defect that is about size — both name assignment mechanism, connected component, covariate balance, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
- The diagnostic at two hundred — both name connected component, covariate balance, effective sample size, ergodicity, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
- The statistic that changes sign — both name connected component, covariate balance, effective sample size, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
- A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- A quantity that loses to a heuristic — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
Named objects
A flat tag is an object no other essay names yet.
Assignment mechanismConnected componentCovariate balanceEffective sample sizeErgodicityExact enumerationImbalanceMarkov chain Monte CarloPermutationRandomisation testReference distributionRerandomisationSharp nullTreatment effect