A probe nobody chose
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
Take one fourteen-unit trial whose admissible set is enumerated and known to be in two mirror components. Run the two-chain diagnostic on twenty-four outcomes drawn across the range of how much of them the balancing rule explains, and on two covariate probes.
Both covariate probes report the split. Seven of the twenty-four outcomes report nothing.
Twelve of the twenty-four report it at between four and twenty, and five report it above twenty — one of them at 97.2. So the same defect, on the same set, at the same run length, is reported at anything from 0.7 to 97.2 depending on which outcome the trial happened to produce.
The reason is in the set, not in the run
A statistic that varies over a factor of a hundred invites the obvious explanation — the chains have not mixed, the effective sample size is being mis-estimated, one run is an anecdote — and that explanation is available and is wrong here.
At fourteen units the components can be enumerated, so what each probe can see is computable exactly: the difference between the two components’ mean probe values, over the spread inside a component. Call it the separation. It runs from 0.001 to 4.938 across these twenty-four outcomes — five thousand-fold — and the two covariate probes sit at 0.431 and 0.419.
The chains agree with the enumeration. That agreement is the check: a chain reporting a large statistic where the enumeration says the two components nearly coincide would be reporting its own failure to mix, and none of them does.
So the spread is not noise. An outcome with a separation of 0.001 cannot see the split however long the chains are run, because on that probe the two halves of the reference distribution have the same mean.
What separation is
The quantity is worth a sentence of its own because it is the only thing in this field that predicts anything.
The two components are complement pairs, so on a probe that is exactly negated by the complement they have means and . The chains estimate those two means, and the test’s statistic is their difference over its standard error. Whether that difference is detectable depends on against the spread inside a component — how much the probe varies among the assignments the chain can actually reach.
A probe with a large and a small within-component spread is a probe on which the two halves are two obviously different distributions. A probe with near zero is one on which they are the same distribution, and the walk’s failure is invisible on it for ever.
Not the explained share
The obvious candidate for what decides separation is how much of the outcome the balancing rule already explains, and it is not.
Across the twenty-four outcomes the explained share runs from 9% to 95%, and the separations at each end run over most of the whole range: at around 10% explained the separations are 0.985, 0.611 and 0.176; at around 92% they are 0.385, 1.698, 2.554, 0.451, 0.633 and 4.938. There is no trend to find.
That is worth being explicit about because the natural theory says there should be one. The rule holds the covariates near balance on every admissible assignment, so an outcome the rule explains ought to be held near balance too and ought to have less to separate the components with. The theory is wrong because the components are the constrained directions split by sign: an admissible assignment has a small imbalance in the basis and a definite one, its complement has the negative, and a probe correlated with the basis is exactly the probe that tells one half from the other.
Both mechanisms are real and they pull opposite ways, which is why the answer is neither and the measured separation is what has to be read.
What the twenty-four outcomes are
The sweep is over outcomes constructed with a stated share of their variance inside the span of what the rule was told to balance — four shares, six draws each — and the realised share is measured rather than assumed, because on fourteen units a random vector already sits a third inside any four-dimensional span by dimension counting.
That construction is deliberately simple: an outcome is a linear combination of the balanced basis plus noise. A real outcome has more in it, and every direction it has that the basis does not is a direction the components are not split along, so a richer outcome should have a smaller separation on average rather than a larger one. The sweep is therefore optimistic about the outcome probe, and the finding — a quarter of them blind — is a lower bound on how often the situation arises.
There is a second respect in which it is optimistic. Every outcome here is generated once and used for both chains, which is exactly right under the sharp null — the outcome column is fixed — and it means the sweep has no outcome noise in it at all beyond the twenty-four draws.
Chosen against given
The practical difference between the two kinds of probe is not power. It is control.
A covariate probe is chosen before any outcome exists, from a dictionary of functions, and if it comes back quiet another can be tried — and should be, because a quiet reading from one probe is not evidence of connectivity when a probe with a separation of 0.001 exists. The loudest reading over several probes is what a practitioner should take, and over the twenty-six probes here the loudest is 97.2 against a threshold of 4.
An outcome is one column, arriving after every design decision has been made. If it is one of the seven that see nothing, there is no second outcome.
This is a recommendation with a clean shape: run the diagnostic on several probes and take the loudest, and treat a quiet verdict from one probe as no information. It costs nothing before the trial and it is the only defence available afterwards.
Why the covariate probes are not blind
Neither of the two covariate probes here is quiet, and it would be wrong to read that as a property of covariate probes in general.
They are polynomial and threshold functions of the same covariate the rule was told to balance, so they sit largely inside the span the components are split along — the fourth power has 98.9% of its variance in that span on these fourteen units, and the threshold 95.3%. Their separations, 0.431 and 0.419, are middling rather than large: five of the outcomes beat both of them.
What makes them reliable is not that they are strong. It is that they can be checked and replaced: their separation is a property of the design and the rule, both of which exist before the trial, so a practitioner can compute what a probe will be able to see before committing to it.
What this says about the earlier field’s choice
The covariate diagnostic picks a function the balancing rule was not handed, and gives a reason: a randomisation test’s reference distribution is about the quantities the rule did not constrain, so those are where a defect in it shows.
The measurement here does not support that reasoning and does not contradict the choice. The probe it uses sits 98.9% inside the span of what the rule was handed, so whatever the rule of thumb was, it is not what produced the probe that gets used — and the probe that gets used works.
What the measurement replaces the reasoning with is a computable quantity. Instead of “choose a function the rule was not handed”, the rule is “choose a function with a large separation, and separation is a property of the design and the rule that can be computed before the trial”. That is a better rule in the sense that it can be checked, and a worse one in the sense that checking it needs the components, which is the enumeration the diagnostic exists to avoid.
The honest position is that the covariate probes here are known to work by having been checked against an enumeration, at the sizes where an enumeration is possible, and that is all that is known.
What a longer run does not fix
The two failure modes have to be kept apart because they have different repairs and look identical from a single number.
An unmixed chain reports a large statistic because it has not settled, and the field’s own replicate control catches it: a third chain from the same starting point, which nothing structural can separate from the first, disagreeing by as much as the mirror pair does. That failure is repaired by a longer run, and at two hundred units a longer run repairs it.
A blind probe reports a small statistic because there is nothing on that probe to report. No run length repairs it, the replicate control says nothing about it, and the verdict is “quiet” — which is the same word the diagnostic uses for a connected set.
That collision is the field’s own instance of the pattern it keeps finding: the same output for “nothing is wrong” and “this instrument cannot see what is wrong”. The only defence is a second instrument, which is the recommendation above.
The number a practitioner would report
One consequence is worth spelling out because it is what a trial’s reader faces.
Suppose a trial runs a balanced-assignment sampler, gets its assignment, collects its outcomes, and runs the diagnostic on the difference in arm means. The verdict comes back quiet.
On the set measured here that verdict is wrong seven times in twenty-four, and the trial has no way to know which case it is in. What it can do is run the diagnostic again on a covariate function, or on several — which costs nothing, needs no new data, and is exactly the probe the earlier field established works.
So the practical conclusion of this essay is not that the outcome probe is a poor instrument. It is that a quiet reading from any single probe is not evidence, and the outcome is the one probe a practitioner is most likely to run alone because it is the one the trial is about.
How small a separation is too small
The threshold is four on the statistic and the two covariate probes clear it at separations of 0.431 and 0.419, which locates the detectable separation at these settings: about 0.4 at twenty thousand draws.
That converts into a statement about run length, because the statistic is a difference over a standard error and the standard error falls as 1/√B. Quadrupling the run halves the separation a probe needs. So a probe at 0.176 would need about four times the draws and a probe at 0.001 would need (0.4/0.001)² = a hundred and sixty thousand times them — three billion draws to see a defect that is definitively there.
“Cannot see it however long the chains are run” is not a figure of speech and it is not a limit either. It is a number, and the number says the blindest of the twenty-four outcomes is unreachable by any run anybody would perform while being perfectly reachable in principle. A practitioner cannot tell the two situations apart from the verdict, which is the whole argument for a second probe.
Not “no trend” — a trend the sweep cannot resolve
The explained share is reported as having no relation to the separation, and the two groups of draws say something slightly weaker and more useful.
At around a tenth explained the three separations average 0.59; at around nine tenths the six average 1.78 — a factor of three in the mean. On a logarithmic scale the difference is 0.89 with a standard error of 0.65, which is 1.4 standard errors.
So the sweep is consistent with no relation and equally consistent with the factor of three it measured. Three draws and six draws cannot separate a threefold difference in a quantity whose within-group spread is five- and thirteen-fold, and the honest statement is that the sweep does not resolve it rather than that there is nothing there.
That matters for the mechanism the essay proposes. Two effects pulling opposite ways would produce exactly this — a small net trend the sweep cannot see — and so would one effect dominating mildly. The argument for the first is structural rather than measured, and it is worth marking as such.
The headline rate deserves the same treatment. Seven of twenty-four is 29%, with a standard error of 9.3 points on twenty-four draws, so the interval runs from about a tenth to about a half. “A quarter of outcomes are blind” is the point estimate of a quantity this sweep pins to within a factor of four, which is enough to establish that the failure is common and not enough to say how common.
Where else a chosen instrument beats a given one
The shape here is not about randomisation and it recurs across this collection, so it is worth naming.
A quantity chosen to be sensitive beats a quantity that happens to be reported, and the second is usually the one available. A criterion read on a benchmark that is one candidate among many is the same trade; so is a diagnostic run on residuals that are not the errors. In each case the instrument a practitioner has is the one the procedure produced, and its sensitivity is a property of the procedure rather than a choice.
What separates this instance is that both instruments are available at once and cost nothing. There is no reason to run only the outcome probe, and the only thing that would make somebody do it is the thought that a diagnostic about a p-value ought to be run on the statistic the p-value is about.
That thought is right about validity, which is the first essay’s subject, and wrong about what to run. Both being true at once is the least intuitive thing in this field.
There is one further asymmetry worth naming. A covariate probe can be run again, on a second function, at no cost and with no new data; an outcome cannot, because a trial has one. So the recommendation that follows is not symmetric between the two moments either: before the trial, several probes are a cheap insurance, and afterwards they are the only insurance there is.
What is claimed here, and what is not
This essay takes how much a probe decides what the diagnostic can see. The claims are that on one fourteen-unit admissible set enumerated to be in two mirror components, seven of twenty-four outcomes give a statistic below the threshold of four and both covariate probes give one above it; that the enumerated separation between the two components runs from 0.001 to 4.938 across those outcomes, so the spread is a property of the set rather than of the runs; that the chains agree with the enumeration, which is what says so; that the separation does not track how much of the outcome the balancing rule explains, because two mechanisms pull opposite ways; and that the loudest reading over the twenty-six probes here is 97.2 against a threshold of 4.
What stays out, and is named as a decision: a rule for choosing probes. The recommendation is to run several and take the loudest, and that is weaker than it could be: since the separation is computable from the design and the rule before any outcome exists, a probe could be chosen to maximise it. Doing that properly needs the components, which is an enumeration, which is exactly what the diagnostic exists to avoid — so the honest form is a heuristic and it is offered as one.
Also out: the same measurement at two hundred units. Everything here is at fourteen because everything here is checked against an enumeration. Whether the range of separations is as wide at a scale nothing can enumerate is unmeasured, and the field that runs the diagnostic there reports a different difficulty — an effective sample size two chains disagree about — which would have to be separated from this one first.
The boundary against the first essay of the field is that it establishes the outcome probe is the same test and this one measures how often it works.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A defect that is about size — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution, rerandomisation
- A quantity that loses to a heuristic — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, projection, randomisation test, rerandomisation
- Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, projection, randomisation test, rerandomisation
- A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- A set of pairs, not a vector — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, projection, randomisation test, rerandomisation
- The statistic that changes sign — both name connected component, covariate balance, effective sample size, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
Named objects
A flat tag is an object no other essay names yet.
Assignment mechanismConnected componentCovariate balanceEffective sample sizeExact enumerationImbalanceMarkov chain Monte CarloProjectionRandomisation testReference distributionRerandomisationSharp nullStatistical powerTreatment effect