The diagnostic after the trial

Before the trial and after

The same diagnostic run at two moments answers two different questions. Before, a positive verdict changes the design. After, it changes which number gets reported — and only for the numbers the defect can reach.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

The field has one construction and two moments, and the closing question is what each moment is worth.

Both probes reach the same verdict at every tolerance where the truth is known: quiet at 1.2, 1.0, 0.9 and 0.85 where the admissible set is one piece of 886, 304, 158 and 126 assignments, firing at 0.8, 0.75 and 0.7 where it is two. Neither is more accurate than the other. What differs is when they can be run and what can be done about the answer.

Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.
Fig. 1 Both probes against the enumerated truth, at seven tolerances of a fourteen-unit rule.

Where the question arises

It arises at a tolerance, and the tolerance is a design choice nobody thinks of as dangerous.

Tightening a balancing rule shrinks the admissible set — 886, 304, 158, 126, 116, 102, 84 assignments of the 3,432 equal splits of fourteen units, as the tolerance goes from 1.2 to 0.7 — and somewhere in that shrinking the set stops being connected under single swaps. On this design it happens between 0.85 and 0.8, between a set of 126 and one of 116.

Where the set stops being one set. How many of the 3,432 equal splits of fourteen units a balancing rule admits, as the tolerance tightens, with the number of components single swaps leave it in. The set falls from 886 to 84 assignments, and somewhere in that fall it stops being connected: at 0.8 it is in 2 pieces and every assignment's complement is in the other one. Nothing about the rule changes at that point and nothing a chain reports changes either, which is the whole difficulty — the acceptance rate, the stationary distribution and the detailed balance are all in order on both sides of it.
Fig. 2 How the admissible set shrinks as the rule tightens, and where it stops being one set.

Nothing about the rule changes there. The constraint is the same shape, the acceptance rate moves smoothly, and every diagnostic the chain can run on itself passes on both sides of it. A practitioner tightening a tolerance from 0.85 to 0.8 for perfectly good reasons — better balance, tighter intervals — crosses a line with no marker on it.

The share admitted, and what it does not mark

Against the 3,432 equal splits of fourteen units, the seven tolerances admit 25.8%, 8.9%, 4.6%, 3.7%, 3.4%, 3.0% and 2.4%.

Read as a share, the crossing is a fall from 3.7% to 3.4%. There is nothing in that step to distinguish it from the ones on either side, and it is the smallest proportional fall anywhere in the sequence — which is the section’s point in numbers: the quantity that changes is connectivity, and connectivity is not a size.

There is a subtler signal in the same counts, and it is worth finding precisely so that its uselessness can be stated. Reading each step as a local power of the tolerance, the set shrinks like t5.9t^{5.9} from 1.2 to 1.0, like t6.2t^{6.2} from 1.0 to 0.9, like t4.0t^{4.0} from 0.9 to 0.85 — and then like t1.4t^{1.4} across the step where it disconnects, before recovering to 2.0 and 2.8.

So the exponent does dip, and it dips at exactly the right step. It is not a marker anybody can use. It is computed from enumerated counts, which is precisely what a practitioner running a walk does not have; the chain’s own acceptance rate estimates the same shares to sampling error, and 3.7% against 3.4% is a difference no run distinguishes from noise. A signal visible only through the calculation the walk exists to avoid is not a warning.

Why enumeration stops where it does

“Below about twenty-four units” is a threshold with a growth rate behind it, and the rate is what makes the boundary sharp rather than a matter of patience.

Fourteen units give C(14,7) = 3,432 equal splits. Every two units added multiply that by 2(2m+1)/(m+1)2(2m+1)/(m+1), which is just under four and rising towards it, so the count quadruples for each pair of subjects:

Twenty-four units give 2,704,156. Twenty-six give 10.4 million. Thirty give 155 million. Thirty-four give 2.3 billion.

Which is why the boundary is where it is and why arguing about it is pointless. A tenfold increase in patience buys about three and a half more units, and a trial is not usually within three units of the size somebody wanted. The walk is not a convenience at thirty units; it is the only construction available, and everything this field measures about walks follows from that.

It also fixes the scale of the diagnostic question. The fourteen-unit design in every figure here is enumerable because it was chosen to be — the whole field rests on knowing the truth it is checking a probe against. Nothing here is measured at a size where the probe would actually be needed, and that is a limitation of the evidence rather than of the argument: the two-chain construction’s validity is an argument about complements and sign flips, and it does not depend on fourteen.

Before: the verdict changes the design

Run before the outcomes exist, a positive verdict is actionable in three ways and all three are free.

Loosen the tolerance. The set is connected at 0.85 and split at 0.8, and the cost of the first over the second is a slightly worse expected balance — which the field that prices balancing rules can quantify against whatever the trial is protecting.

Change the proposal. A walk that moves two pairs at a time crosses between components that single swaps cannot, and what that costs in acceptance rate is measured in its own field.

Or stop walking. Below about twenty-four units the admissible set can be enumerated outright, and where that stops is the only reason a walk is used at all.

A proposal that moves more, refused more often. The two halves of the trade, both exact, on the 410 admissible assignments of twelve units. The integrated autocorrelation time of an imbalance the rule was never handed falls from 7.30 at one swap to 3.97 at three, and the acceptance rate falls with it, from 58.8% to 40.8%. A rejected proposal costs one evaluation and leaves the chain where it was, so acceptance is not the price of anything and the ranking by acceptance is the reverse of the ranking by cost. Past three the family folds: exchanging k of six from each arm is the complement of exchanging six − k, so k = 5 has the same 36 proposals as k = 1 and k = 6 has 1.
Fig. 3 The first repair, priced: a proposal that can cross between components, and what it costs in refusals.

The diagnostic that supports those decisions is the covariate one, because it is the only one that exists yet. That is the earlier field’s whole justification and it stands.

The probe a trial has is the probe a trial got. What the two-chain test says when it is run on the trial's own difference in arm means, over 24 outcomes on one fourteen-unit set. The set is in 2 mirror components — that is enumerated, not inferred — so every quiet reading is a miss. 29% of them are quiet. The reason is in the enumerated set rather than in the run: how far the two components are apart on a given probe ranges from 0.001 to 4.938 of a within-component spread across these outcomes, a factor of several thousand. Both covariate probes — chosen before any outcome existed, and replaceable if they had been quiet — report the split. An outcome cannot be chosen and cannot be replaced.
Fig. 4 The share of outcomes that report nothing, which is why the after-the-fact check is not the outcome probe alone.

After: the verdict changes what is reported

Run after the trial, none of those is available. The assignment happened, the outcomes are in, and the design is a matter of record.

What is available is the reading. A trial that knows its walk sampled half the reference distribution knows which p-values are affected, and the answer is unusually clean: a two-sided p-value on any statistic that is negated by swapping the arms is exactly right, and everything else is not.

So the after-the-fact verdict does not invalidate a trial. It partitions the trial’s own results into a part that stands and a part that needs the other half of the reference distribution — which, at a size where enumeration is possible, can simply be computed.

The one thing a trial always reports is the one thing that survives. How wrong three p-values are when they are computed over the half of the admissible set a single walk can reach, rather than over all of it, at a fourteen-unit trial where the whole set can be enumerated. The two-sided p-value on the difference in arm means — the number a trial publishes — is wrong by exactly nothing, at every row, to machine precision. That is not luck: the two components are complement pairs and the difference in arm means is exactly negated by the complement, so the distribution of its absolute value is the same on both. A one-sided p-value on the same statistic is out by as much as 0.112, and the largest response observed in the treated arm — a safety reading rather than an effect, and the one statistic here that is not odd under the complement — by as much as 0.172. The defect survived because the commonest thing anybody computes is the one quantity it cannot touch.
Fig. 5 Which reported numbers the defect can reach, which is what an after-the-fact verdict is for.

That is a smaller claim than the trial is wrong and a more useful one. It is also a claim a reader can act on rather than only an author: a published trial that reports its balancing rule and its covariates gives a reader everything the diagnostic needs.

What a positive verdict does not mean

Two misreadings are available and both are worth heading off, because the verdict is a blunt word.

It does not mean the assignment was bad. Every assignment the walk produced is admissible: it satisfies the balancing rule exactly as required, and its covariate balance is whatever the rule promised. What the verdict says is that the reference distribution the walk sampled was half of the one the p-value is defined against. The trial’s balance is fine; its null is halved.

And it does not mean the walk was broken. The chain is doing precisely what it was written to do — proposing single swaps, accepting admissible ones, and converging to the uniform distribution on everything it can reach. It converges. It is stationary. It satisfies detailed balance. The set is simply not connected under the move it was given, and nothing the chain can be asked about itself will ever say so.

That is what makes the two-chain construction necessary rather than merely convenient: the fact being tested for is a property of the set and the move together, and no property of one run of the chain carries it.

Everything passes, and half the set is unreachable. The three checks a practitioner would run on a chain over the 116 admissible assignments of a fourteen-unit trial, computed on the chain's exact transition matrix rather than simulated. The acceptance rate is 16%, which is ordinary. The stationary distribution is uniform to 5.2e-18, which is what the construction promises. The transition matrix is symmetric to 0.0e+0, so detailed balance holds exactly. And the set is in 2 pieces, so the chain is uniform on half of it for ever. Every one of the three is computed from where the chain goes, which is why none of them can report on where it does not.
Fig. 6 The diagnostics a chain can run on itself, all in order on a set it reaches half of.
What a probe can see, and what a chain says about it. Each point is one probe on one fourteen-unit admissible set: horizontally what the enumeration says the two components differ by on that probe, in within-component spreads; vertically what a pair of chains 20,000 steps long reports. The two agree, which is the check — a chain that reports a large statistic where the enumeration says the components are nearly coincident would be reporting its own failure to mix. The bars through the points are the range over four chain seeds. What the picture is for is the horizontal axis: the covariate probes sit at 0.431 and 0.419, and the outcomes run from 0.001 to 4.938. A probe is chosen; an outcome is what happened.
Fig. 7 What each probe can see against what a chain reports, which is the quantity the recommendation is about.

Which probe at which moment

The two diagnostics are not simply the same test at two times, because one of them can be chosen and the other cannot.

Before, run several covariate probes and take the loudest. They cost nothing, they can be computed from the design alone, and a quiet reading from one of them is not evidence — a probe whose enumerated separation is 0.001 will be quiet on a split set for ever.

After, run the outcome probe and the covariate probes. The outcome probe is the one the p-value is about, which is what makes it worth running; the covariate probes are the ones that can be checked and replaced, which is what makes them worth running beside it. Three in ten outcomes on the set measured here see nothing at all.

The recommendation is the same at both moments — several probes, loudest wins — and only the reason differs. Before, because a better probe may exist. After, because the probe the trial handed over may be one of the blind ones.

The two moments are not equally well served

It is worth being plain that the before-the-trial diagnostic is the better one, and that this field’s contribution is mostly to the worse one.

Before the trial everything is available: the design is known, the rule is known, several probes can be tried, the enumeration can be run if the trial is small, and every repair is free. A practitioner who runs the check then has a decision procedure with no residual uncertainty in it at small sizes and a good one at large.

After the trial the situation is worse in every respect. One outcome, no repairs, and a probe that is blind three times in ten. What the after-the-fact check buys is a partition of the results rather than a fix, and it buys it in a case where the fix is unavailable.

So the honest ordering is: run it before, and if that was not done, run it after. The second is not a substitute for the first and this field does not claim it is. What it claims is that the second is possible at all, valid under the null the p-value is computed under, and worth doing — including by somebody who did not run the trial.

How often half a reference distribution changes the answer. The share of observed assignments on which the one-sided randomisation p-value computed from the component the walk can reach falls on the other side of five per cent from the one computed over the whole admissible set. Every assignment of the 116-strong set is taken as the observed one in turn, which is exactly the population a rejection sampler draws from. It runs between 2.6% and 4.3% — about one trial in thirty — and it does not depend much on how much of the outcome the balancing rule already explains. The same figure for the two-sided p-value is zero at every row, and not approximately zero: the components are complement pairs and the statistic is exactly negated by the complement, so the two halves carry identical distributions of its absolute value.
Fig. 8 The rate at which a one-sided verdict changes, which is what the protection does not extend to.

What the field does not cover

Three limits, and the first of them is structural rather than a gap in the measurement.

An unequal allocation breaks the construction. Everything here splits the units evenly, which is what makes an assignment’s complement an assignment of the same shape, and the complement is where the second chain starts. Under a two-to-one allocation the complement is not admissible — it is not even a valid assignment — so there is no second chain to run and no symmetry to test with. That is not a limitation of this essay; it is a limitation of the whole two-chain idea, and no repair is offered.

A constraint that is not symmetric breaks the protection. The two-sided p-value’s exactness rests on the components being complement pairs, which rests on the admissible set being closed under complementation, which rests on the rule admitting a vector exactly when it admits its negative. Every rule in this collection does. A rule that penalised imbalance in one direction more than the other would not, and the whole of the previous essay’s protection would have to be re-derived.

And an estimate is not a test. Everything here is about a p-value under the sharp null. A randomisation-based interval inverts a family of tests rather than one, so what a half-reachable walk does to it is a question with more moving parts and is not answered anywhere in this field.

What this field found that the earlier one did not have

Three things, and they are worth separating from the earlier field’s own findings.

The outcome is a valid probe, under the sharp null, with no new theory — because the sharp null fixes the outcome column and the difference in arm means is then the signed imbalance of it. That is what makes the after-the-fact diagnostic exist at all.

The probe is what decides whether the diagnostic works, and the range is enormous: separations from 0.001 to 4.938 across two dozen outcomes on one set, with seven of them reporting nothing. The earlier field chose its probe on a reasonable-sounding rule — a function the balancing rule was not handed — and the probe it uses sits 98.9% inside the span it was supposed to avoid, so whatever produced the working probe, it was not that rule.

And the defect costs the published number nothing, exactly, for ever, whenever the statistic is odd. That is why nobody found it and it is the most useful sentence in the field: a practitioner reading a two-sided p-value on an effect estimate has nothing to worry about, and everybody else does.

What a trial should actually do

Four lines, in the order they arise.

Before the design is fixed: compute the admissible set if the trial is small enough to enumerate, and if it is not, run several covariate probes on the design and take the loudest. If the verdict is positive, loosen the tolerance or change the proposal — both are cheap and neither costs anything a trial cares about.

When choosing the analysis: prefer a two-sided test on an odd statistic. This is what most trials do anyway, and it happens to be immune.

After the outcomes are in: run the diagnostic again, on the outcome and on the covariate probes together. A quiet reading from one probe is not a verdict.

And if the verdict is positive after the fact: report the two-sided p-value on the primary endpoint as it stands, and recompute anything one-sided or arm-specific over the whole admissible set if the size allows it. What is not available is a repair by re-running the walk, because a longer walk does not reach a component it cannot enter.

Where the four essays leave it

The first establishes that the outcome is a probe: under the sharp null the outcome column is fixed, the difference in arm means is its signed imbalance, and the two-chain test applies unchanged. Checked against an enumeration at seven tolerances.

The second finds that a probe’s power is a property of the set rather than of the run — separations from 0.001 to 4.938 — and that seven of twenty-four outcomes see nothing on a set that is definitively split.

The third enumerates what the defect costs the published number and finds it is exactly nothing for a two-sided p-value on an odd statistic, as much as 0.112 for a one-sided one, and 0.172 for a statistic that is not odd.

This one asks what to do about it, and the answer divides on when the question is asked.

Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.
Fig. 9 The same seven tolerances at a much higher explained share, where the verdicts are unchanged — which is what says the division above is about timing rather than about the outcome.

One thing is worth doing at every stage and costs nothing: record the tolerance and the proposal. The verdict depends on both, neither is usually reported, and a reader who has them can run the whole check from outside. A trial that reports “balanced by rerandomisation” has told a reader nothing they can act on; one that reports the rule, the tolerance and the move has told them everything.

One consequence for anybody writing a protocol is worth stating plainly. The diagnostic is cheap, it needs no data, and its inputs — the units’ covariates, the balancing rule and the tolerance — all exist before randomisation. There is no reason a trial should reach its analysis without knowing whether its sampler could reach the whole admissible set, and the only thing standing between most trials and that answer is that nobody has asked.

What is claimed here, and what is not

This essay takes what each of the two diagnostics is for. The claims are that both probes reach the enumerated verdict at all seven tolerances measured, quiet where the set is one piece of 886, 304, 158 and 126 assignments and firing where it is two of 58; that the set stops being connected between tolerances of 0.85 and 0.8 with nothing else about the rule changing; that a positive verdict before the trial admits three repairs and a positive verdict afterwards admits none of them; that after the fact the verdict partitions the reported numbers rather than invalidating them, because a two-sided p-value on an odd statistic is exact on half a reference distribution; and that the recommendation at both moments is several probes with the loudest taken, for two different reasons.

What stays out, and is named as a decision: an unequal allocation. The complement of an assignment under a two-to-one split is not an assignment, so the second chain has nothing to start from and the whole construction has no analogue. Naming a repair would mean inventing one, and none has been tested.

Also out: a recommendation about tolerance. The set splits between 0.85 and 0.8 on this design, with these fourteen units and this basis, and there is no reason to think that number transports. What transports is the method: compute where it splits for the design in hand, which is an enumeration at small sizes and the diagnostic at large ones.

The boundary against the field that found the split is that it establishes the defect and this field asks what it costs and who can see it. Both answers turned out to be smaller than expected and neither is nothing.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A defect that is about size — both name assignment mechanism, connected component, covariate balance, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
  • A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • A quantity that loses to a heuristic — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • The diagnostic at two hundred — both name connected component, covariate balance, ergodicity, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
  • The statistic that changes sign — both name connected component, covariate balance, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismConnected componentCovariate balanceErgodicityExact enumerationImbalanceMarkov chain Monte Carlop-valueRandomisation testReference distributionRerandomisationSharp nullStudy designTreatment effect