The diagnostic after the trial

Half a reference distribution

A walk that reaches half its admissible set reports the two-sided p-value exactly right, to the last digit, for ever. A one-sided one it puts on the wrong side of five per cent about once in thirty.

Worth reading first: Randomisation is not balance · The experiments that could have happened.

A walk that samples half its admissible set is sampling half a reference distribution, and a randomisation test’s p-value is a share of a reference distribution. So the p-value is wrong.

It is not. The commonest one is exactly right, to machine precision, on every admissible assignment of every set measured — and the reason is a symmetry rather than an approximation.

The measurement

At fourteen units the admissible set can be enumerated. At the tolerance where it splits it holds 116 assignments in two components of 58, and every assignment’s complement is in the other one.

Take each of those 116 in turn as the observed assignment — which is exactly the population a rejection sampler draws from — and compute three p-values twice: once over the whole set, once over the component the observed assignment sits in, which is all a walk can ever visit.

The one thing a trial always reports is the one thing that survives. How wrong three p-values are when they are computed over the half of the admissible set a single walk can reach, rather than over all of it, at a fourteen-unit trial where the whole set can be enumerated. The two-sided p-value on the difference in arm means — the number a trial publishes — is wrong by exactly nothing, at every row, to machine precision. That is not luck: the two components are complement pairs and the difference in arm means is exactly negated by the complement, so the distribution of its absolute value is the same on both. A one-sided p-value on the same statistic is out by as much as 0.112, and the largest response observed in the treated arm — a safety reading rather than an effect, and the one statistic here that is not odd under the complement — by as much as 0.172. The defect survived because the commonest thing anybody computes is the one quantity it cannot touch.
Fig. 1 How wrong each kind of p-value is on the reachable half, against how much of the outcome the balancing rule explains.

Two-sided, on the difference in arm means: wrong by exactly nothing. Not close to nothing. The worst difference over 116 observed assignments and five explained shares is 0, at machine precision.

One-sided, on the same statistic: wrong by as much as 0.112. On one worked case the whole set gives 0.526 and the reachable half gives 0.603.

The largest response observed in the treated arm: wrong by as much as 0.172. On the same case, 0.897 against 0.845.

Why the two-sided one survives

The components are complement pairs and the difference in arm means is exactly negated by the complement — that is the fact the whole diagnostic rests on. So if the statistic takes value tt somewhere in one component, it takes t-t at the mirror assignment in the other.

The multiset of t|t| on one component is therefore the multiset of t|t| on the other, and a two-sided p-value is a share of assignments with t|t| at least as large as the observed one. Two identical multisets give the same share.

It is worth being precise about the strength of that. It is not that the halves are similar, or that the error is small at this sample size, or that it averages out. The two halves carry identical distributions of the absolute statistic, so the two-sided p-value from half the set is the two-sided p-value from all of it, at every observed assignment, at every tolerance, at every trial size, for ever.

What half a reference set costs, and what it does not. The 95% point of a difference in means over the 116 admissible assignments of 14 units, computed over the whole set and over each of the two components a single-swap walk can reach. The statistic changes sign under the complement and the components are exactly the complement pairs, so the two halves carry mirror images of one distribution: the upper point is 0.9344 on one and 0.4170 on the other, against 0.7749 for the set. A one-sided test run on the wrong half uses a critical value 46% too small. A two-sided test reads the same number either way — 0.9344 against the whole set's 0.9434 — because a mirror image has the same absolute values.
Fig. 2 The two halves of the reference distribution, which are mirror images — so their absolute values are the same distribution.

Which is why nobody found it

That identity is the answer to the question the whole enumeration was run to ask: if a walk has been sampling half its reference distribution, how did nothing go wrong?

Nothing went wrong because the quantity everybody computes is the one quantity a mirror split cannot touch. Every acceptance rate was in order, every stationary distribution was uniform on what the chain could reach, every diagnostic a chain can run on itself passed — and the number that came out at the end was right.

This collection has a name for the shape and it applies exactly: the symptom is absence. A defect whose only consequence is on quantities nobody computes leaves nothing to see.

What a probe can see, and what a chain says about it. Each point is one probe on one fourteen-unit admissible set: horizontally what the enumeration says the two components differ by on that probe, in within-component spreads; vertically what a pair of chains 20,000 steps long reports. The two agree, which is the check — a chain that reports a large statistic where the enumeration says the components are nearly coincident would be reporting its own failure to mix. The bars through the points are the range over four chain seeds. What the picture is for is the horizontal axis: the covariate probes sit at 0.431 and 0.419, and the outcomes run from 0.001 to 4.938. A probe is chosen; an outcome is what happened.
Fig. 3 What each probe can see, which is what decides whether a trial would ever learn its walk was split.

The error is an integer over 116

Every error in the measurement above is a count difference divided by the size of the set, which is worth working out because it turns three decimals into a structure.

Write aa for the number of assignments in the observed component whose statistic is at least the observed tt, and bb for the number whose statistic is at most t-t. The other component is the exact negation of this one, so the full set’s one-sided tail is a+ba + b out of 116 and the reachable half’s is aa out of 58. Subtracting,

phalfpfull=a58a+b116=ab116.p_{\text{half}} - p_{\text{full}} = \frac{a}{58} - \frac{a+b}{116} = \frac{a-b}{116}.

The error is an integer over 116, in steps of 0.00862. The worked case’s 0.603 against 0.526 is 9/1169/116; the largest error found, 0.112, is 13/11613/116; and the largest-response statistic’s 0.172 is 20/11620/116 with its worked case’s 0.052 at 6/1166/116.

Four numbers reported to three decimals, all four exactly integers over the size of the enumerated set.

What the bound is, and how far from it the measurement sits

The same expression gives the worst case the construction allows, which is a more useful thing to know than the worst case observed.

Since aa and bb are counts within a component of 58, ab58|a-b| \le 58, so the one-sided p-value can be wrong by at most 58/11658/116, which is 0.5. The largest error measured is 13 of that possible 58 — 22% of the bound.

So the one-sided statistic is not protected and is not badly damaged either. Its error is bounded at a half by the mirror structure, and on this set it uses a fifth of that.

The largest-response statistic has no such bound. It is not negated by the complement — the largest response among the treated is a different quantity when the arms are swapped — so nothing constrains aa against bb and the error is bounded only by 1.

Measured, it is 20/116 against the odd statistic’s 13/116, a factor of 1.54. Which is the right ordering and a much smaller factor than the difference in their bounds suggests: the mirror structure buys the odd statistic a guarantee, and on this particular set it buys it about half a point of actual error.

What is not protected

The protection is a consequence of oddness, and the boundary is sharp.

A statistic that is exactly negated by the complement has a two-sided p-value that survives and a one-sided p-value that does not: a one-sided share counts assignments with tt at least the observed value, and one component holds the positive tail while the other holds the negative one. The reachable half is the wrong half half the time.

A statistic that is not odd has neither protection. The largest response observed in the treated arm — a safety reading rather than an effect — becomes the largest in the control arm under the complement, which is unrelated to it, and its two-sided p-value is out by up to 0.172.

How often half a reference distribution changes the answer. The share of observed assignments on which the one-sided randomisation p-value computed from the component the walk can reach falls on the other side of five per cent from the one computed over the whole admissible set. Every assignment of the 116-strong set is taken as the observed one in turn, which is exactly the population a rejection sampler draws from. It runs between 2.6% and 4.3% — about one trial in thirty — and it does not depend much on how much of the outcome the balancing rule already explains. The same figure for the two-sided p-value is zero at every row, and not approximately zero: the components are complement pairs and the statistic is exactly negated by the complement, so the two halves carry identical distributions of its absolute value.
Fig. 4 How often the one-sided verdict at five per cent changes on the reachable half, by explained share.

A third of what looks like a list of statistics is not one. The treated arm’s own mean is the difference in means over two plus a constant that does not depend on the assignment, so its p-values are the difference’s to the last digit — the same statistic in different clothes. The field is smaller than it looks, and the boundary running through it is exactly oddness.

What it costs where it counts

An error of 0.08 in a p-value of 0.60 changes nothing. The reading that decides something is how often the reachable half puts a p-value on the other side of five per cent from where the whole set puts it.

For the one-sided p-value that is 2.6% to 4.3% of observed assignments — about one trial in thirty — and it does not depend much on how much of the outcome the balancing rule explains. For the two-sided one it is zero at every row, as it must be. For the largest-in-arm statistic it is zero as well, but for an uninteresting reason: those p-values are large on this construction and never near five per cent at all, which is a property of the sweep rather than a protection.

The probe a trial has is the probe a trial got. What the two-chain test says when it is run on the trial's own difference in arm means, over 24 outcomes on one fourteen-unit set. The set is in 2 mirror components — that is enumerated, not inferred — so every quiet reading is a miss. 29% of them are quiet. The reason is in the enumerated set rather than in the run: how far the two components are apart on a given probe ranges from 0.001 to 4.938 of a within-component spread across these outcomes, a factor of several thousand. Both covariate probes — chosen before any outcome existed, and replaceable if they had been quiet — report the split. An outcome cannot be chosen and cannot be replaced.
Fig. 5 Whether the trial would even know: the outcome probe’s verdicts on this set, a quarter of them quiet.

One trial in thirty is a rate worth knowing and it is not an emergency. What makes it worth the essay is that it is a rate a practitioner currently has no way to discover, since the diagnostic that would tell them is quiet a quarter of the time on the probe they would run it with.

Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.
Fig. 6 The verdicts at seven tolerances, which is where the enumeration says the split is real.

The direction is not random

One more thing the enumeration gives that a simulation could not: the error has a sign that depends on which component the observed assignment is in.

An observed assignment in the component holding the positive tail sees a reachable distribution shifted towards it, so its one-sided p-value comes out too large. An assignment in the other component sees the opposite. Averaged over the 116 the errors cancel — the mean signed error is essentially zero — which is why nothing about the average would ever reveal this.

That cancellation is itself worth carrying, because it is the reason a simulation study of the whole procedure would find nothing. Draw many trials, compute the one-sided p-value under each, check that the distribution is uniform: it is, because the two directions of error are equally likely and equally large. The defect is invisible to the calibration check that would normally catch a broken reference distribution.

20,000 p-values from a true null, n = 30Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0228 (p = 0.01). That flatness is the check that catches an error a single rejection rate would miss.05001e+30.2000.4000.6000.800p-valuecountflatKS distance 0.0228a p-value that is not uniform is not a p-value
Fig. 7 The check that would normally catch a broken reference distribution, and which this defect passes because its two directions of error cancel.

One worked case

The averages hide the shape, so one observed assignment in full.

On the fourteen-unit set at a tolerance of 0.8, with an outcome 36% explained by the balanced basis, take the first admissible assignment a rejection sampler would return. Over the whole set of 116 its two-sided p-value is 0.9655 and over its own component of 58 it is 0.9655 — the same number, not a rounding of two nearby ones. Its one-sided p-value is 0.5259 over the whole set and 0.6034 over the component.

Nothing about the second pair looks wrong. Both are large, both are the kind of number a trial reports as no evidence, and a reader comparing them without the enumeration in front of them would have no reason to prefer either.

That is what makes the enumeration the load-bearing part of this field rather than a nicety. The two p-values differ by 0.08 and there is no internal signal — no diagnostic, no convergence statistic, no disagreement between chains — that says which is the right one.

What a practitioner does

Three sentences.

Report the two-sided p-value on an odd statistic and the defect cannot reach it. That is the usual practice, it is what most trials do, and it is exactly right for a reason nobody was relying on.

A one-sided test, a non-inferiority margin, or any statistic that treats the two arms differently needs the walk checked. Those are not exotic: a one-sided hypothesis is standard in several fields, and a safety reading on the treated arm is standard in all of them.

And check the walk with a probe that can see, not with the one the trial is about. The loudest reading over several probes is the verdict; a quiet reading from one is not.

Why an odd statistic is the usual case

The protection covers a class that is worth describing, because it is larger than it first appears and its boundary is not where intuition puts it.

A statistic is odd under the complement when swapping every unit’s arm negates it. The difference in arm means is; so is a difference in medians, a difference of any location measure computed the same way in both arms, a two-sample t statistic with a pooled variance, a Wilcoxon rank sum measured from the centre, and a covariate-adjusted treatment effect. Almost every effect estimate a trial reports is odd, because an effect is a contrast and a contrast reverses when the labels do.

What is not odd is anything that treats one arm as special: a maximum, a count of events in the treated arm, a proportion above a threshold in one arm, a safety signal. Those are not exotic either — they are the second table of most trial reports.

So the honest summary is that the primary endpoint is protected and the secondary table is not, which is an unusually clean division and not one anybody designed.

The same question one field along

This is not the only place in this collection where a reference distribution is sampled rather than enumerated, and it is worth checking whether the same protection is available elsewhere.

A bootstrap’s reference distribution is drawn rather than walked, so there is no reachability question at all — every resample is available at every draw. The failure mode there is a different one: the distribution being sampled is the wrong distribution, not half the right one.

A specification search’s null is simulated from a fitted model, and there the analogous question is whether the fit is right rather than whether the sampler reaches everything.

So the defect this field is about is specific to a walk over a constrained set, which is the one construction in this collection whose sampler can be stuck. That narrows where the finding applies and it also says where to look for it: any procedure that draws from a restricted randomisation by proposing local moves has the same exposure, and the diagnostic transfers unchanged.

Where the set stops being one set. How many of the 3,432 equal splits of fourteen units a balancing rule admits, as the tolerance tightens, with the number of components single swaps leave it in. The set falls from 886 to 84 assignments, and somewhere in that fall it stops being connected: at 0.8 it is in 2 pieces and every assignment's complement is in the other one. Nothing about the rule changes at that point and nothing a chain reports changes either, which is the whole difficulty — the acceptance rate, the stationary distribution and the detailed balance are all in order on both sides of it.
Fig. 8 The tolerances at which the set splits at all, which is the condition the protection is stated under.

Where the protection ends

The identity holds because the components are complement pairs. That is measured on every set here — the mirror share is 1.000 — and it is not a theorem.

A set that split into components that were not complement pairs would have no such protection, and nothing in this collection has produced one: every split found by enumeration, at every tolerance and every trial size up to twenty-two units, is the mirror pairing. The field that enumerates them reports it as the only structure that has ever come out, and reports it as a measurement rather than an argument.

So the honest statement of the protection is conditional: while the split is the mirror pairing, the two-sided p-value on an odd statistic is exact. If a rule ever produced a different split — a proposal moving more than one pair, a constraint that is not symmetric, an unequal allocation — the protection would have to be re-derived and there is no reason to expect it to survive.

What the enumeration cost

The whole field’s certainty comes from one thing: at fourteen units the 3,432 equal splits can be listed, the admissible ones picked out, the components found by walking single swaps, and every p-value computed as a share rather than estimated as a rate.

That is a small computation and it is the only reason any of the statements above can use the word exactly. A simulation study of the same question would have reported that the two-sided p-value agrees to within simulation error, which is a weaker and much less interesting claim — and it would have missed the cancellation of the one-sided errors entirely, because a cancellation looks like correctness from the outside.

Where the enumeration stops is about twenty-four units, and everything past it is the two-chain diagnostic’s territory. The division of labour is the same one the whole area runs on: enumerate where it is possible so that the test can be checked, and then use the test where enumeration is not.

What is claimed here, and what is not

This essay takes what a walk that reaches half its admissible set does to the number a trial publishes. The claims are that over every one of the 116 admissible assignments of a fourteen-unit set, taken in turn as the observed one, the two-sided p-value on the difference in arm means computed over the reachable component equals the one computed over the whole set to machine precision, at every explained share measured; that a one-sided p-value on the same statistic differs by as much as 0.112 and crosses five per cent on 2.6% to 4.3% of observed assignments; that a statistic which is not odd under the complement — the largest response in the treated arm — differs by as much as 0.172; that the treated arm’s mean is the difference in means over two plus a constant and so is not a separate statistic; and that the signed one-sided errors cancel across the set, so a calibration check on the whole procedure would find nothing.

What stays out, and is named as a decision: the same enumeration at a size a trial would run. Everything here is at fourteen units because everything here is exact, and the arithmetic behind the two-sided identity does not depend on the size — it depends on the split being the mirror pairing, which is measured up to twenty-two units and asserted nowhere. The one-sided error rates are a different matter and are properties of this set at this tolerance; a larger trial’s would have to be measured.

Also out: an estimate rather than a test. Everything here is about a p-value under the sharp null. What a half-reachable walk does to a randomisation-based interval — which inverts a family of tests, not one — is a question with more moving parts and is not answered.

The boundary against the field that found the split is that it establishes the walk reaches half the set and this one asks what that costs the number at the end. The answer is nothing, for the number most trials report, which is why the defect could last.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A defect that is about size — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
  • A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test
  • A probe chosen from the design — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
  • A quantity that loses to a heuristic — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test
  • Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test
  • The part the rule already took — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismConnected componentCovariate balanceExact enumerationImbalanceMarkov chain Monte CarloOne-sided testp-valuePermutationRandomisation testReference distributionSharp nullSymmetryTreatment effect