The diagnostic after the trial

A probe nobody chose

On a set that is definitively in two pieces, seven of twenty-four outcomes report nothing at all. Every covariate probe reports it. What separates them is not accuracy — it is that one of them can be chosen and the other is what happened.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

Take one fourteen-unit trial whose admissible set is enumerated and known to be in two mirror components. Run the two-chain diagnostic on twenty-four outcomes drawn across the range of how much of them the balancing rule explains, and on two covariate probes.

Both covariate probes report the split. Seven of the twenty-four outcomes report nothing.

The probe a trial has is the probe a trial got. What the two-chain test says when it is run on the trial's own difference in arm means, over 24 outcomes on one fourteen-unit set. The set is in 2 mirror components — that is enumerated, not inferred — so every quiet reading is a miss. 29% of them are quiet. The reason is in the enumerated set rather than in the run: how far the two components are apart on a given probe ranges from 0.001 to 4.938 of a within-component spread across these outcomes, a factor of several thousand. Both covariate probes — chosen before any outcome existed, and replaceable if they had been quiet — report the split. An outcome cannot be chosen and cannot be replaced.
Fig. 1 Where the twenty-four outcomes land against a threshold of four, on a set that is split in every case.

Twelve of the twenty-four report it at between four and twenty, and five report it above twenty — one of them at 97.2. So the same defect, on the same set, at the same run length, is reported at anything from 0.7 to 97.2 depending on which outcome the trial happened to produce.

The reason is in the set, not in the run

A statistic that varies over a factor of a hundred invites the obvious explanation — the chains have not mixed, the effective sample size is being mis-estimated, one run is an anecdote — and that explanation is available and is wrong here.

At fourteen units the components can be enumerated, so what each probe can see is computable exactly: the difference between the two components’ mean probe values, over the spread inside a component. Call it the separation. It runs from 0.001 to 4.938 across these twenty-four outcomes — five thousand-fold — and the two covariate probes sit at 0.431 and 0.419.

What a probe can see, and what a chain says about it. Each point is one probe on one fourteen-unit admissible set: horizontally what the enumeration says the two components differ by on that probe, in within-component spreads; vertically what a pair of chains 20,000 steps long reports. The two agree, which is the check — a chain that reports a large statistic where the enumeration says the components are nearly coincident would be reporting its own failure to mix. The bars through the points are the range over four chain seeds. What the picture is for is the horizontal axis: the covariate probes sit at 0.431 and 0.419, and the outcomes run from 0.001 to 4.938. A probe is chosen; an outcome is what happened.
Fig. 2 Each probe’s enumerated separation against what a pair of chains reports, with the range over four chain seeds drawn through each point.

The chains agree with the enumeration. That agreement is the check: a chain reporting a large statistic where the enumeration says the two components nearly coincide would be reporting its own failure to mix, and none of them does.

So the spread is not noise. An outcome with a separation of 0.001 cannot see the split however long the chains are run, because on that probe the two halves of the reference distribution have the same mean.

Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.
Fig. 3 Both probes against the enumerated truth, which is what says the spread below is about power rather than validity.

What separation is

The quantity is worth a sentence of its own because it is the only thing in this field that predicts anything.

The two components are complement pairs, so on a probe that is exactly negated by the complement they have means +m+m and m-m. The chains estimate those two means, and the test’s statistic is their difference over its standard error. Whether that difference is detectable depends on 2m2m against the spread inside a component — how much the probe varies among the assignments the chain can actually reach.

A probe with a large mm and a small within-component spread is a probe on which the two halves are two obviously different distributions. A probe with mm near zero is one on which they are the same distribution, and the walk’s failure is invisible on it for ever.

What half a reference set costs, and what it does not. The 95% point of a difference in means over the 116 admissible assignments of 14 units, computed over the whole set and over each of the two components a single-swap walk can reach. The statistic changes sign under the complement and the components are exactly the complement pairs, so the two halves carry mirror images of one distribution: the upper point is 0.9344 on one and 0.4170 on the other, against 0.7749 for the set. A one-sided test run on the wrong half uses a critical value 46% too small. A two-sided test reads the same number either way — 0.9344 against the whole set's 0.9434 — because a mirror image has the same absolute values.
Fig. 4 The two halves of a reference distribution, which is what a probe with a small separation cannot tell apart.
The one thing a trial always reports is the one thing that survives. How wrong three p-values are when they are computed over the half of the admissible set a single walk can reach, rather than over all of it, at a fourteen-unit trial where the whole set can be enumerated. The two-sided p-value on the difference in arm means — the number a trial publishes — is wrong by exactly nothing, at every row, to machine precision. That is not luck: the two components are complement pairs and the difference in arm means is exactly negated by the complement, so the distribution of its absolute value is the same on both. A one-sided p-value on the same statistic is out by as much as 0.112, and the largest response observed in the treated arm — a safety reading rather than an effect, and the one statistic here that is not odd under the complement — by as much as 0.172. The defect survived because the commonest thing anybody computes is the one quantity it cannot touch.
Fig. 5 What the defect costs each kind of p-value, which is what a blind probe fails to warn about.

Not the explained share

The obvious candidate for what decides separation is how much of the outcome the balancing rule already explains, and it is not.

Across the twenty-four outcomes the explained share runs from 9% to 95%, and the separations at each end run over most of the whole range: at around 10% explained the separations are 0.985, 0.611 and 0.176; at around 92% they are 0.385, 1.698, 2.554, 0.451, 0.633 and 4.938. There is no trend to find.

That is worth being explicit about because the natural theory says there should be one. The rule holds the covariates near balance on every admissible assignment, so an outcome the rule explains ought to be held near balance too and ought to have less to separate the components with. The theory is wrong because the components are the constrained directions split by sign: an admissible assignment has a small imbalance in the basis and a definite one, its complement has the negative, and a probe correlated with the basis is exactly the probe that tells one half from the other.

Both mechanisms are real and they pull opposite ways, which is why the answer is neither and the measured separation is what has to be read.

Two diagnostics, one answer, two different moments. The two-chain statistic on a covariate probe and on the trial's own difference in arm means, at seven tolerances of a fourteen-unit rule, against the enumerated truth. Both are quiet wherever the set is one set and both fire wherever it is not, at every tolerance — which is what says the outcome probe is the same test rather than a resemblance of it. The difference between them is not accuracy and it is not power. It is when: the covariate probe can be run before a single outcome exists, when a practitioner can still loosen the rule or change the sampler, and it can be run again on a different function if it comes back quiet. The outcome probe runs after the trial, on the one column the trial produced, and what it can do with a positive verdict is repair the p-value rather than the design.
Fig. 6 The two probes at a much higher explained share, where the verdicts are the same as at a low one.

What the twenty-four outcomes are

The sweep is over outcomes constructed with a stated share of their variance inside the span of what the rule was told to balance — four shares, six draws each — and the realised share is measured rather than assumed, because on fourteen units a random vector already sits a third inside any four-dimensional span by dimension counting.

That construction is deliberately simple: an outcome is a linear combination of the balanced basis plus noise. A real outcome has more in it, and every direction it has that the basis does not is a direction the components are not split along, so a richer outcome should have a smaller separation on average rather than a larger one. The sweep is therefore optimistic about the outcome probe, and the finding — a quarter of them blind — is a lower bound on how often the situation arises.

The zero was a fact about independenceWhat a balancing rule handed every main effect of both covariates removes of a pure interaction, as the covariates are allowed to move together. At ρ = 0 it is exactly nothing — at machine precision, at any number of main effects — which is the independent-covariate result and is correct. It is not small anywhere else: the product of the two covariates loses 64.0% of itself by ρ = 0.5, because h₁h₁ = h₀ + √2·h₂ and Mehler pairs h₂ with h₂ at ρ². Four interactions are drawn and none of them keeps the zero.00.2500.5000.750100.2000.4000.6000.800correlation between the two covariatesshare of the interaction the eight main effects removeexactly zero, and only hereρ = 0.5the product of the two covariatesone covariate times the other's squarethe product of the two squaresone covariate times the other's cubeclosed geometry at every correlation, no simulation64.0% gone by ρ = 0.5
Fig. 7 The kind of structure a real outcome has that a linear combination of the basis does not, which is what makes the sweep optimistic.

There is a second respect in which it is optimistic. Every outcome here is generated once and used for both chains, which is exactly right under the sharp null — the outcome column is fixed — and it means the sweep has no outcome noise in it at all beyond the twenty-four draws.

Chosen against given

The practical difference between the two kinds of probe is not power. It is control.

A covariate probe is chosen before any outcome exists, from a dictionary of functions, and if it comes back quiet another can be tried — and should be, because a quiet reading from one probe is not evidence of connectivity when a probe with a separation of 0.001 exists. The loudest reading over several probes is what a practitioner should take, and over the twenty-six probes here the loudest is 97.2 against a threshold of 4.

An outcome is one column, arriving after every design decision has been made. If it is one of the seven that see nothing, there is no second outcome.

This is a recommendation with a clean shape: run the diagnostic on several probes and take the loudest, and treat a quiet verdict from one probe as no information. It costs nothing before the trial and it is the only defence available afterwards.

Why the covariate probes are not blind

Neither of the two covariate probes here is quiet, and it would be wrong to read that as a property of covariate probes in general.

They are polynomial and threshold functions of the same covariate the rule was told to balance, so they sit largely inside the span the components are split along — the fourth power has 98.9% of its variance in that span on these fourteen units, and the threshold 95.3%. Their separations, 0.431 and 0.419, are middling rather than large: five of the outcomes beat both of them.

What makes them reliable is not that they are strong. It is that they can be checked and replaced: their separation is a property of the design and the rule, both of which exist before the trial, so a practitioner can compute what a probe will be able to see before committing to it.

What this says about the earlier field’s choice

The covariate diagnostic picks a function the balancing rule was not handed, and gives a reason: a randomisation test’s reference distribution is about the quantities the rule did not constrain, so those are where a defect in it shows.

The measurement here does not support that reasoning and does not contradict the choice. The probe it uses sits 98.9% inside the span of what the rule was handed, so whatever the rule of thumb was, it is not what produced the probe that gets used — and the probe that gets used works.

What the measurement replaces the reasoning with is a computable quantity. Instead of “choose a function the rule was not handed”, the rule is “choose a function with a large separation, and separation is a property of the design and the rule that can be computed before the trial”. That is a better rule in the sense that it can be checked, and a worse one in the sense that checking it needs the components, which is the enumeration the diagnostic exists to avoid.

The honest position is that the covariate probes here are known to work by having been checked against an enumeration, at the sizes where an enumeration is possible, and that is all that is known.

What a longer run does not fix

The two failure modes have to be kept apart because they have different repairs and look identical from a single number.

An unmixed chain reports a large statistic because it has not settled, and the field’s own replicate control catches it: a third chain from the same starting point, which nothing structural can separate from the first, disagreeing by as much as the mirror pair does. That failure is repaired by a longer run, and at two hundred units a longer run repairs it.

A blind probe reports a small statistic because there is nothing on that probe to report. No run length repairs it, the replicate control says nothing about it, and the verdict is “quiet” — which is the same word the diagnostic uses for a connected set.

That collision is the field’s own instance of the pattern it keeps finding: the same output for “nothing is wrong” and “this instrument cannot see what is wrong”. The only defence is a second instrument, which is the recommendation above.

The number a practitioner would report

One consequence is worth spelling out because it is what a trial’s reader faces.

Suppose a trial runs a balanced-assignment sampler, gets its assignment, collects its outcomes, and runs the diagnostic on the difference in arm means. The verdict comes back quiet.

On the set measured here that verdict is wrong seven times in twenty-four, and the trial has no way to know which case it is in. What it can do is run the diagnostic again on a covariate function, or on several — which costs nothing, needs no new data, and is exactly the probe the earlier field established works.

So the practical conclusion of this essay is not that the outcome probe is a poor instrument. It is that a quiet reading from any single probe is not evidence, and the outcome is the one probe a practitioner is most likely to run alone because it is the one the trial is about.

Where the set stops being one set. How many of the 3,432 equal splits of fourteen units a balancing rule admits, as the tolerance tightens, with the number of components single swaps leave it in. The set falls from 886 to 84 assignments, and somewhere in that fall it stops being connected: at 0.8 it is in 2 pieces and every assignment's complement is in the other one. Nothing about the rule changes at that point and nothing a chain reports changes either, which is the whole difficulty — the acceptance rate, the stationary distribution and the detailed balance are all in order on both sides of it.
Fig. 8 The tolerances at which the question arises at all, which is where a practitioner has to know whether their sampler is in trouble.

How small a separation is too small

The threshold is four on the statistic and the two covariate probes clear it at separations of 0.431 and 0.419, which locates the detectable separation at these settings: about 0.4 at twenty thousand draws.

That converts into a statement about run length, because the statistic is a difference over a standard error and the standard error falls as 1/√B. Quadrupling the run halves the separation a probe needs. So a probe at 0.176 would need about four times the draws and a probe at 0.001 would need (0.4/0.001)² = a hundred and sixty thousand times them — three billion draws to see a defect that is definitively there.

“Cannot see it however long the chains are run” is not a figure of speech and it is not a limit either. It is a number, and the number says the blindest of the twenty-four outcomes is unreachable by any run anybody would perform while being perfectly reachable in principle. A practitioner cannot tell the two situations apart from the verdict, which is the whole argument for a second probe.

Not “no trend” — a trend the sweep cannot resolve

The explained share is reported as having no relation to the separation, and the two groups of draws say something slightly weaker and more useful.

At around a tenth explained the three separations average 0.59; at around nine tenths the six average 1.78 — a factor of three in the mean. On a logarithmic scale the difference is 0.89 with a standard error of 0.65, which is 1.4 standard errors.

So the sweep is consistent with no relation and equally consistent with the factor of three it measured. Three draws and six draws cannot separate a threefold difference in a quantity whose within-group spread is five- and thirteen-fold, and the honest statement is that the sweep does not resolve it rather than that there is nothing there.

That matters for the mechanism the essay proposes. Two effects pulling opposite ways would produce exactly this — a small net trend the sweep cannot see — and so would one effect dominating mildly. The argument for the first is structural rather than measured, and it is worth marking as such.

The headline rate deserves the same treatment. Seven of twenty-four is 29%, with a standard error of 9.3 points on twenty-four draws, so the interval runs from about a tenth to about a half. “A quarter of outcomes are blind” is the point estimate of a quantity this sweep pins to within a factor of four, which is enough to establish that the failure is common and not enough to say how common.

Where else a chosen instrument beats a given one

The shape here is not about randomisation and it recurs across this collection, so it is worth naming.

A quantity chosen to be sensitive beats a quantity that happens to be reported, and the second is usually the one available. A criterion read on a benchmark that is one candidate among many is the same trade; so is a diagnostic run on residuals that are not the errors. In each case the instrument a practitioner has is the one the procedure produced, and its sensitivity is a property of the procedure rather than a choice.

What separates this instance is that both instruments are available at once and cost nothing. There is no reason to run only the outcome probe, and the only thing that would make somebody do it is the thought that a diagnostic about a p-value ought to be run on the statistic the p-value is about.

That thought is right about validity, which is the first essay’s subject, and wrong about what to run. Both being true at once is the least intuitive thing in this field.

There is one further asymmetry worth naming. A covariate probe can be run again, on a second function, at no cost and with no new data; an outcome cannot, because a trial has one. So the recommendation that follows is not symmetric between the two moments either: before the trial, several probes are a cheap insurance, and afterwards they are the only insurance there is.

What is claimed here, and what is not

This essay takes how much a probe decides what the diagnostic can see. The claims are that on one fourteen-unit admissible set enumerated to be in two mirror components, seven of twenty-four outcomes give a statistic below the threshold of four and both covariate probes give one above it; that the enumerated separation between the two components runs from 0.001 to 4.938 across those outcomes, so the spread is a property of the set rather than of the runs; that the chains agree with the enumeration, which is what says so; that the separation does not track how much of the outcome the balancing rule explains, because two mechanisms pull opposite ways; and that the loudest reading over the twenty-six probes here is 97.2 against a threshold of 4.

What stays out, and is named as a decision: a rule for choosing probes. The recommendation is to run several and take the loudest, and that is weaker than it could be: since the separation is computable from the design and the rule before any outcome exists, a probe could be chosen to maximise it. Doing that properly needs the components, which is an enumeration, which is exactly what the diagnostic exists to avoid — so the honest form is a heuristic and it is offered as one.

Also out: the same measurement at two hundred units. Everything here is at fourteen because everything here is checked against an enumeration. Whether the range of separations is as wide at a scale nothing can enumerate is unmeasured, and the field that runs the diagnostic there reports a different difficulty — an effective sample size two chains disagree about — which would have to be separated from this one first.

The boundary against the first essay of the field is that it establishes the outcome probe is the same test and this one measures how often it works.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A defect that is about size — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
  • A quantity that loses to a heuristic — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, projection, randomisation test, rerandomisation
  • Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, projection, randomisation test, rerandomisation
  • A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • A set of pairs, not a vector — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, projection, randomisation test, rerandomisation
  • The statistic that changes sign — both name connected component, covariate balance, effective sample size, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismConnected componentCovariate balanceEffective sample sizeExact enumerationImbalanceMarkov chain Monte CarloProjectionRandomisation testReference distributionRerandomisationSharp nullStatistical powerTreatment effect