The diagnostic after the trial

The statistic the p-value is about

The test for whether a balanced-assignment walk reaches its whole set is run on a covariate function chosen before the trial. Run on the difference in arm means it is the same test, and it is about the number the trial publishes.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

The two-chain test asks whether the walk that draws a balanced assignment can reach the whole admissible set, and it asks it on a covariate function the balancing rule was not handed. That choice is deliberate, it is stated as a limitation in the field that makes it, and the reason given is good: the diagnostic runs before the outcomes exist, which is the moment a practitioner can still loosen the rule or change the sampler.

What the same test says on the difference in arm means is a different question, and it is the one the p-value is actually about.

The test carries over unchanged

A randomisation test conditions on the sharp null: every unit’s outcome is what it is whatever arm it lands in. Under that null the outcome is a fixed vector — not a random one, not a model’s prediction, a column of numbers that does not move when the assignment does.

The difference in arm means is then

2ni±Yi,\frac{2}{n}\sum_i \pm Y_i,

the signed imbalance of that column, which is exactly the object the covariate probe is with a different column in it. It is negated exactly by the complement, which is what the two-chain construction rests on. Nothing has to be re-derived.

The identity is checked rather than asserted: on twenty admissible assignments of a fourteen-unit trial, the signed probe on the outcome column and the difference in arm means agree to twelve decimal places, up to the standardisation the probe applies. That check exists because the two claims about p-values in this field would both be satisfied by an enumeration of the wrong object.

Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.
Fig. 1 Both probes at seven tolerances of a fourteen-unit rule, against the enumerated truth.

Both probes agree with the truth

At fourteen units the admissible set can be enumerated, so the answer is known and the diagnostic can be scored rather than trusted.

At tolerances of 1.2, 1.0, 0.9 and 0.85 the set is connected — 886, 304, 158 and 126 assignments in one piece — and both probes are quiet, at statistics between −1.8 and 2.0 against a threshold of 4. At 0.8, 0.75 and 0.7 the set is in two mirror components, and both fire: the covariate probe at 9.5, 13.3 and 14.6, the outcome probe at 9.0, 10.5 and 12.2.

Seven tolerances, two probes, fourteen verdicts, all correct.

Where the set stops being one set. How many of the 3,432 equal splits of fourteen units a balancing rule admits, as the tolerance tightens, with the number of components single swaps leave it in. The set falls from 886 to 84 assignments, and somewhere in that fall it stops being connected: at 0.8 it is in 2 pieces and every assignment's complement is in the other one. Nothing about the rule changes at that point and nothing a chain reports changes either, which is the whole difficulty — the acceptance rate, the stationary distribution and the detailed balance are all in order on both sides of it.
Fig. 2 How the admissible set shrinks as the rule tightens, and where it stops being one set.

That is what makes the outcome probe the same test rather than a resemblance of it. Whatever else follows about how well it works — and the next essay is entirely about that — it is not a different diagnostic with its own theory.

What a probe can see, and what a chain says about it. Each point is one probe on one fourteen-unit admissible set: horizontally what the enumeration says the two components differ by on that probe, in within-component spreads; vertically what a pair of chains 20,000 steps long reports. The two agree, which is the check — a chain that reports a large statistic where the enumeration says the components are nearly coincident would be reporting its own failure to mix. The bars through the points are the range over four chain seeds. What the picture is for is the horizontal axis: the covariate probes sit at 0.431 and 0.419, and the outcomes run from 0.001 to 4.938. A probe is chosen; an outcome is what happened.
Fig. 3 What each probe can see against what a chain reports, which is the check on the whole construction.

What the split is

The defect being tested for is worth restating, because everything in this field is about what it does to a number rather than about what it is.

Every balancing rule in this collection is symmetric under the joint sign flip: the imbalance of z-z is minus the imbalance of zz, and both the box rule and the ellipsoid admit a vector exactly when they admit its negative. So the admissible set is closed under complementation, whether or not the walk can get there.

A walk that moves by single swaps cannot always get there. At a tight enough tolerance the 116 admissible assignments of a fourteen-unit trial fall into two components of 58, and every assignment’s complement is in the other one — so a chain started anywhere is uniform on its own half for ever, and every diagnostic a chain can run on itself passes.

A thin enough set is not one set. Every admissible set of 14 units this table can enumerate, by how much of the assignment space it admits and how many pieces it falls into under single swaps. A walk is uniform on the piece it starts in and never leaves it. The pieces are not fragments: at 522 admissible assignments the set splits into 3 halves of exactly 520 each, and every assignment's complement is in the other half — no sequence of admissible single swaps takes an assignment to its own mirror image. Two-swap proposals reconnect four of the six disconnected sets here — the two they do not are the thinnest, where a two-unit move rarely lands anywhere admissible either — which makes a bigger proposal a correctness repair rather than the speed dial it was measured as.
Fig. 4 The set in two mirror components, enumerated, which is the defect both probes are looking for.

Ten assignments is what disconnects it

The seven tolerances are close enough together to say how abruptly the defect arrives, and the answer is that it arrives all at once.

At a tolerance of 0.85 the admissible set has 126 members and is one piece. At 0.8 it has 116 and is two pieces of 58. Ten assignments removed — eight per cent of the set — is the whole distance between a walk that mixes and a walk that is confined to half of what it is sampling.

There is no gradual stage in between. Connectivity is a property of the graph rather than of its size, so nothing in the counts warns that the boundary is near, and the four connected tolerances give a sequence — 886, 304, 158, 126 — whose only visible feature is that it is shrinking.

That is why the diagnostic has to be a test rather than a look at the numbers. The quantity a practitioner can see moves smoothly through the point where the quantity that matters changes discontinuously, which is the same shape as every other absence this collection tests for.

How fast the set shrinks, and what that says about its dimension

The counts also say how many constraints the rule is really imposing, and the answer is what the rule was written to impose.

Between the two tightest connected tolerances the set falls from 158 to 126 as the tolerance goes from 0.9 to 0.85, an exponent of log(1.254) / log(1.0588) = 4.0. A tolerance that is a half-width in each of d balanced dimensions admits a count proportional to its d-th power once the region is small enough for the density inside it to be flat, so an exponent of four is the four balanced factors, counted from outside.

At the loose end the exponent is larger — about 5.9 between 1.2 and 1.0, and 6.2 between 1.0 and 0.9 — because there the region is wide enough that the imbalance distribution’s own falling density adds to the volume effect. The exponent settling towards the dimension as the tolerance tightens is the signature that says the count is being driven by geometry rather than by the shape of the law.

The outcome probe is quieter, and only slightly

On the three split tolerances the two probes can be compared directly. The outcome probe reads 9.0, 10.5 and 12.2 where the covariate probe reads 9.5, 13.3 and 14.6 — ratios of 0.95, 0.79 and 0.84, averaging 0.86.

So the outcome the trial happened to produce is about six sevenths as loud as a probe chosen to be loud, on the sets where both fire. Against a threshold of 4 that leaves margins of 5.0, 6.5 and 8.2 rather than 5.5, 9.3 and 10.6 — smaller, and nowhere near the threshold.

Which is worth holding beside the claim two sections down that three outcomes in ten report nothing at all on a set that is definitively split. Both are true and they are about different features of the same distribution: the typical outcome is nearly as informative as a chosen probe, and the distribution over outcomes has a long lower tail that a single reading cannot show. A practitioner gets one draw from it and has no way to tell which part of it they got.

Why the outcome is a probe nobody chose

The covariate probe is chosen. It is a function of the recorded covariates, it can be picked to be anything the rule was not handed, and if it comes back quiet a second one can be tried.

An outcome is what the trial produced. There is one of it, it arrives after every design decision has been made, and a practitioner who does not like what it says about the walk has no second one.

That difference has no consequence for validity — the section above settles that — and it has a large one for power, which is the next essay’s whole subject. The short form is that on a set which is genuinely split, three in ten outcomes report nothing at all, and every covariate probe tried reports the split.

The probe a trial has is the probe a trial got. What the two-chain test says when it is run on the trial's own difference in arm means, over 24 outcomes on one fourteen-unit set. The set is in 2 mirror components — that is enumerated, not inferred — so every quiet reading is a miss. 29% of them are quiet. The reason is in the enumerated set rather than in the run: how far the two components are apart on a given probe ranges from 0.001 to 4.938 of a within-component spread across these outcomes, a factor of several thousand. Both covariate probes — chosen before any outcome existed, and replaceable if they had been quiet — report the split. An outcome cannot be chosen and cannot be replaced.
Fig. 5 What the two-chain test says when it is run on the trial’s own statistic, over twenty-four outcomes on one set that is definitively split.
How often half a reference distribution changes the answer. The share of observed assignments on which the one-sided randomisation p-value computed from the component the walk can reach falls on the other side of five per cent from the one computed over the whole admissible set. Every assignment of the 116-strong set is taken as the observed one in turn, which is exactly the population a rejection sampler draws from. It runs between 2.6% and 4.3% — about one trial in thirty — and it does not depend much on how much of the outcome the balancing rule already explains. The same figure for the two-sided p-value is zero at every row, and not approximately zero: the components are complement pairs and the statistic is exactly negated by the complement, so the two halves carry identical distributions of its absolute value.
Fig. 6 How often the verdict at five per cent changes, which is what a repair after the fact is worth.

What it can and cannot repair

The two diagnostics run at two moments and what they enable is different at each.

Before the trial, a positive verdict is actionable in the strongest sense: loosen the tolerance until the set is connected, change the proposal so it moves more than one pair at a time, or abandon the walk and enumerate. All three are design changes and all three are free, because nothing has happened yet. A proposal that moves more than two units is the cheapest of them and is measured in its own field.

After the trial, none of those is available. What is available is the reference distribution: a practitioner who knows the walk reached half the set knows which half, and can either enumerate the other half where that is possible or report the statistic whose p-value is unaffected. Which statistics those are is a whole essay’s worth of answer and it is a stranger one than it sounds.

A proposal that moves more, refused more often. The two halves of the trade, both exact, on the 410 admissible assignments of twelve units. The integrated autocorrelation time of an imbalance the rule was never handed falls from 7.30 at one swap to 3.97 at three, and the acceptance rate falls with it, from 58.8% to 40.8%. A rejected proposal costs one evaluation and leaves the chain where it was, so acceptance is not the price of anything and the ranking by acceptance is the reverse of the ranking by cost. Past three the family folds: exchanging k of six from each arm is the complement of exchanging six − k, so k = 5 has the same 36 proposals as k = 1 and k = 6 has 1.
Fig. 7 One of the repairs available before the trial and not after: a proposal that can cross between the components.

The outcome is built to order

Everything measured in this field uses an outcome constructed with a stated share of its variance inside the span of what the balancing rule was told to hold. That is a modelling choice and it is worth being explicit about why it is the right one.

A real outcome is a function of the covariates plus noise, and how much of it the covariates explain is exactly the quantity a practitioner has an opinion about. Constructing outcomes across that range — from a tenth of the variance explained to nineteen twentieths — is what makes the sweep a sweep over a stated quantity rather than over a knob.

The realised share is measured rather than assumed, because on fourteen units a random vector already has about a third of its variance inside any four-dimensional span by dimension counting alone. So an outcome built with none of its signal in the span reports 25% explained, and the axis in every figure here is the measured share rather than the target.

What the sharp null is doing

The whole construction rests on one assumption and it is worth being honest about how strong it is.

Under the sharp null the outcome column is fixed, which is what makes it a probe. Under an alternative it is not: a unit’s outcome depends on which arm it got, so the column changes with the assignment and the signed imbalance is no longer a fixed function of the walk’s state.

That sounds like a restriction and it is nearly the opposite. The sharp null is precisely the hypothesis a randomisation test conditions on — it is what the p-value is computed under — so the diagnostic is valid exactly where the p-value it is diagnosing is defined. The test with no table is the field that establishes what that conditioning buys, and what it buys here is that no extra assumption is needed at all.

Where it does bite is in interpretation. A large statistic says the walk did not reach the whole set under the sharp null’s reference distribution, which is the reference distribution the p-value uses. It says nothing about a walk’s behaviour under an alternative, and nothing about estimation.

Where the two probes come apart, and where they do not

It is worth being precise about the sense in which the outcome probe is “the same test”, because two different things could be meant and only one of them is true.

The construction is the same. Two chains, one from an admissible assignment and one from its complement, compared on a signed statistic at effective standard errors, with a replicate chain from the same starting point to say which of two explanations a large statistic has. Every word of that is the covariate diagnostic’s with a column swapped, and the column is handed in rather than named — which is why the shared file takes a column at all.

The power is not the same, and it is not systematically worse either. Both probes are testing the same property of the same set, and how loudly each one reports it depends on how far apart the two components are on that probe. That is an enumerable quantity at fourteen units and it is what the next essay measures, and the range across outcomes is enormous.

The threshold, and why it is four

The verdict rule compares the standardised difference between the two chains against a threshold of four, and four is a convention rather than a calibration.

The statistic is a difference of two means over their own standard errors, so under a connected set it is approximately standard normal and four is about a one-in-sixteen-thousand event. That margin is deliberately wide, because the standard errors are computed from effective draws rather than from the number taken, and the effective sample size is itself estimated — which the covariate field found is the least reliable quantity in the whole construction, disagreeing between two chains from the same starting point by more than three standard errors.

A wide threshold is the right response to an unreliable denominator, and it is why the field reports a verdict rather than a p-value. The numbers in this essay clear it by factors of two to three, and the ones in the next clear it by factors of twenty-four or not at all.

Why this had to be asked at all

The covariate diagnostic is complete and it works, so a reader is entitled to ask what the outcome version adds.

Three things, and the third is the one that made the field.

It answers the question a reader of the earlier field will have asked. A diagnostic about a reference distribution that is demonstrated on a quantity the reference distribution is not about invites exactly one question, and leaving it unanswered is the kind of gap that becomes a claim by default.

It is the only version available to somebody reading a published trial. The covariates of a finished trial are often reported and the balancing rule sometimes is; a reader who wants to know whether the reported p-value came from half a reference distribution has the outcome and the rule, and nothing else. Whether the diagnostic can be run from the outside decides whether it is a design tool or a reviewing tool, and it is both.

And it turns out to have a failure mode the chosen probe does not. That is not what the field was looking for and it is what it found: an outcome is a probe nobody selected, and a quarter of them cannot see the defect at all.

And what a positive verdict is worth

The last thing worth saying before the field turns to power is what the verdict buys, because a diagnostic that cannot change anything is a diagnostic nobody runs.

After the trial, a positive verdict says the reported p-value came from a walk that sampled half its reference distribution. Whether that matters depends entirely on which p-value was reported, and the answer is not the one anybody expects: the commonest one is exactly right, to the last digit, for a reason that is a symmetry rather than an approximation.

The one thing a trial always reports is the one thing that survives. How wrong three p-values are when they are computed over the half of the admissible set a single walk can reach, rather than over all of it, at a fourteen-unit trial where the whole set can be enumerated. The two-sided p-value on the difference in arm means — the number a trial publishes — is wrong by exactly nothing, at every row, to machine precision. That is not luck: the two components are complement pairs and the difference in arm means is exactly negated by the complement, so the distribution of its absolute value is the same on both. A one-sided p-value on the same statistic is out by as much as 0.112, and the largest response observed in the treated arm — a safety reading rather than an effect, and the one statistic here that is not odd under the complement — by as much as 0.172. The defect survived because the commonest thing anybody computes is the one quantity it cannot touch.
Fig. 8 What each kind of p-value costs when it is computed on the reachable half, which is what a positive verdict after the trial is worth knowing.

A last observation about scope. Every set measured here is small enough to enumerate, which is what makes the verdicts checkable, and no trial is that size. What carries to a real trial is the construction rather than the numbers: two chains, one from an assignment and one from its complement, compared on a signed statistic at effective standard errors, with a replicate chain to say which of the two explanations of a large statistic applies. The verdicts here are what say the construction reaches the right answer where the right answer is known.

What is claimed here, and what is not

This essay takes whether the walk’s reachability diagnostic can be run on the statistic a trial reports. The claims are that under the sharp null the outcome column is fixed, so the difference in arm means is the signed imbalance of a column and is exactly negated by the complement; that the signed probe on the outcome and the difference in arm means agree to twelve decimal places on twenty admissible assignments, up to the probe’s standardisation; that at seven tolerances of a fourteen-unit rule both probes reach the enumerated verdict, quiet at 1.2, 1.0, 0.9 and 0.85 where the set is one piece and firing at 0.8, 0.75 and 0.7 where it is two; and that the outcome probe’s statistics there are 9.0, 10.5 and 12.2 against the covariate probe’s 9.5, 13.3 and 14.6.

What stays out, and is named as a decision: an unequal allocation. Every rule here splits the units evenly, which is what makes an assignment’s complement an assignment of the same shape and the whole symmetry argument available. Under a two-to-one allocation the complement is not a valid assignment at all, the admissible set is not closed under it, and the two-chain construction has nothing to start its second chain from. That is a real limitation of the diagnostic rather than of this essay, and no repair for it is offered.

Also out: an outcome with a treatment effect in it. The construction is valid under the sharp null and the sharp null is what a randomisation test conditions on, so nothing here needs an alternative. What a large statistic would mean on a trial with a real effect is a different question and is not answered.

The boundary against the covariate diagnostic is that it is run before the outcomes exist and this one after, and the two reach the same verdict at every tolerance where the truth is known.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A defect that is about size — both name assignment mechanism, connected component, covariate balance, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
  • The diagnostic at two hundred — both name connected component, covariate balance, effective sample size, ergodicity, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
  • The statistic that changes sign — both name connected component, covariate balance, effective sample size, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
  • A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • A quantity that loses to a heuristic — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismConnected componentCovariate balanceEffective sample sizeErgodicityExact enumerationImbalanceMarkov chain Monte CarloPermutationRandomisation testReference distributionRerandomisationSharp nullTreatment effect