The statistic that changes sign
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
The test that does not enumerate compares two chains, one started at an assignment and one at its complement, on the signed imbalance in a function the balancing rule was not handed. Two of those choices look like details and neither is.
Signed is why the test can see anything at all. And two chains is not enough — a third is needed, and what it rules out is the commoner of the two explanations for a large reading.
Every symmetric reading is blind
The admissible set is closed under complementation, and when it splits it splits into an arrangement and its mirror image. So the two components carry the same distribution with the sign reversed.
Anything computed from a distribution that does not change when the whole distribution is reflected is the same number on both components. The absolute imbalance is one. So is the variance, the interquartile range, the two-sided critical value, and the acceptance rate.
The enumeration makes it exact rather than approximate. The two components’ means of the signed statistic differ by 0.1679; the two components’ means of its absolute value differ by 0.0000. Not a small number — zero, because the multisets are reflections of each other and reflection does not move a magnitude.
A magnitude is the natural statistic to reach for, because an imbalance is usually reported as a size. Reaching for it here produces a diagnostic that is exactly blind to the thing it is pointed at.
Where the zero comes from
The exact zero above is worth one more sentence, because “the multisets are reflections of each other” is a claim about the enumerated set and the reason it holds is a claim about the construction.
Every assignment in a component has its complement in the other, and the complement map is a bijection between the two — that is what mirrored means and it is measured rather than assumed, as the share of assignments whose complement lies in a different component, which comes out at 1.00. Under a bijection that negates the statistic, the second component’s multiset of values is exactly the first’s negated. Taking absolute values then gives the identical multiset, so every quantile, every moment and every function of the magnitudes matches to the last bit rather than to a tolerance.
That is why the number reported is 0.0000 rather than 0.0003. It is not a small difference that could grow on a different design; it is an identity, and the only way for it to fail is for the components not to be a mirror pair — which is a case the enumeration also produces, at tolerances thin enough to leave four pieces rather than two.
The blindness is in the summary, not in the covariate
It is easy to read the parity requirement as a constraint on which covariate function to probe with, and it is not. Every imbalance is already the right parity.
An imbalance is with — linear in the assignment, whatever is. So complementing the assignment negates it, exactly, for a covariate, for its square, for its fourth power, for a random direction, for anything at all.
What is even is the summary. An absolute value, a variance, an interquartile range, a two-sided critical value, an acceptance rate: each of those is a function of the imbalance that does not change when the imbalance changes sign, and each is therefore identical on the two components.
So the failure has one step in it and it is not the step anybody would examine. Choosing the covariate function well is the part this line of fields spends its effort on, and no choice there can rescue a diagnostic that takes a magnitude at the end. Choosing to keep the sign costs nothing and is the whole of the repair.
Which makes the decomposition exhaustive
Because every statistic splits into an even part and an odd part, and the two components’ distributions are exact reflections, the accounting has no residue.
The odd part has means that are exact negatives on the two components — here +0.0840 and −0.0840, differing by 0.1679, which is 2 × 0.0840 to the last digit printed.
The even part has means that are exactly equal — differing by 0.0000, not by something small.
There is no third part. A statistic’s usable separation between the components is entirely the separation of its odd part, and its even part contributes nothing but noise to the comparison. A probe that is ninety per cent even and ten per cent odd is not a slightly worse probe; it is a probe carrying a tenth of a signal underneath nine tenths of a common component that cancels in the difference and does not cancel in the variance.
That is why the exact zero is worth reporting as a zero. It is not a measurement that happened to come out small on this enumeration; it is what the bijection between the components forces, and the measured share of assignments whose complement lies in the other component — 1.00 — is the statement that the bijection is there.
And so is the test most trials run
The same blindness has a consequence that is not about diagnostics at all, and it is why the defect matters unevenly.
A randomisation test’s critical value is a quantile of the reference distribution. Read one-sided, the whole set’s 95% point is 0.7749 and the two components give 0.9344 and 0.4170 — so a chain confined to one of them reports a critical value nearly half the right one, or a fifth too large, depending where it started. Read two-sided, the whole set gives 0.9434 and each component gives 0.9344: a gap of 0.0090 against the one-sided gap of 0.3579, a factor of forty.
So a two-sided randomisation test on a set the chain reaches half of is essentially correct, and a one-sided one is not. The property that makes the first true — every symmetric reading is unaffected — is exactly the property that makes a symmetric diagnostic useless.
The parity, stated
The blindness has a one-line statement and it is the same statement the essay on the zero that rests on a symmetry makes about a balancing dictionary, which is worth noticing because the two fields are about different things.
The complement map takes an assignment to its negative. The signed imbalance in any function is odd under that map: negate every arm and the sum negates. Any function of the signed imbalance that is even under negation — a magnitude, a square, a two-sided quantile — is therefore constant across the two components, because they are each other’s images.
So which statistics can see the split is decided by parity, and by nothing about the trial, the units, the tolerance or the basis. That is the same shape as the parity result about interactions in what a dictionary buys: a symmetry of the law decides which inner products are exactly zero, and the answer does not depend on the sizes of anything.
Three readings, one verdict
A large signed difference between two chains has two explanations and only one of them is about the set.
They are in mirror components and cannot reach each other. Then their absolute-value distributions are identical and their signed means are opposite.
They simply have not mixed. Then they are in two arbitrary regions of one connected set, and they differ on everything — the signed statistic, its magnitude, and any other reading anybody takes.
Two comparisons separate the two, and the second is the one that does not depend on any estimate being right.
The replicate is a third chain, started at the same assignment as the first, on a different stream. Nothing structural about the set can separate two chains started at the same point: whatever they disagree about is what a run of this length disagrees with itself about. So a replicate reading of the same size as the complement reading says the run is too short, whatever the mirror pair reports.
The magnitude comparison is the other control, and it is the one the symmetry predicts: quiet under a mirror split, loud under a failure to mix.
At fourteen units both controls stay inside 2.7 standard errors at every tolerance, split or not, while the signed comparison runs to −14.58 on the split sets. That combination — one loud, two quiet — has exactly one explanation.
Why the replicate is the load-bearing one
The magnitude control is elegant and the replicate control is the one that works, and the difference between them is worth stating.
The magnitude comparison is quiet under a mirror split only once both chains have explored their own components well enough for the magnitude to have settled. Under a chain that has barely moved, the two magnitudes are two arbitrary numbers and the comparison is loud — so the magnitude control cannot distinguish a mirror split from a stalled chain when the chain is stalled, which is exactly the case it is needed for.
The replicate has no such condition. Two chains from one point are exchangeable by construction, at any run length, on any statistic. Their difference has mean zero whatever the mixing is, and its standard error is estimated from the same autocorrelated sequences the main comparison uses — so if the standard error is wrong, it is wrong in both places and the ratio is right.
A control that does not depend on the estimate it is controlling is the useful shape, and it is why the test has three chains rather than two.
The standard error is the thing being controlled
That last point deserves its own paragraph, because it is what the third chain is really for.
The comparison divides a difference of means by a standard error, and the standard error is computed from an estimated integrated autocorrelation time. At fourteen units that estimate is about thirty and the sequences are twenty thousand long, so it has plenty of data and is believable. At two hundred units it is about three hundred and fifty, and the sequences are not long enough for it — the estimate truncates early, reports too little dependence, and produces a standard error several times too small.
An instrument whose standard error is wrong reports significance everywhere. The replicate comparison uses the same wrong standard error on a difference that is genuinely zero, so it reports the same spurious significance — and the test’s verdict, which requires one reading loud and the other quiet, correctly refuses to say anything.
The control’s job is not to be right about the set. It is to be wrong in the same way, so that the comparison of the two is right.
What the three readings cost
Running three chains rather than one is three times the work, and it is worth being clear that nothing cheaper does the job.
One chain and its own diagnostics is what the field already had, and all of them pass on a set the chain reaches half of.
Two chains from independently hunted starting points would detect a split only when the two happened to land in different components, which is half the time, and would give no way of telling that from a failure to mix. The complement is what makes the two starting points guaranteed to be in different components when there are two.
Two chains from an assignment and its complement, without a replicate, is the test as first built, and it is a coin at two hundred units: it reports a split at some run lengths and not at others on the same set, because the reading it depends on is the standard error rather than the structure.
So three chains is the cheapest arrangement that answers the question, and the third one is not redundancy — it is the only part that does not assume the thing being estimated.
Choosing the probe
One more choice is a choice: which function to measure the imbalance in.
It has to be a function the balancing rule was not handed, because a function the rule holds has its imbalance constrained to a box and carries almost no information about which component an assignment is in. It also has to vary enough on the units drawn to have a distribution at all — the first version of this test used a threshold at one standard deviation on fourteen units, where exactly one unit exceeded the threshold, so the signed statistic took two values and its magnitude took exactly one. Every magnitude comparison in that version returned zero over zero.
The probe used throughout is the fourth power of the covariate, standardised, on a rule handed the covariate, its square, its cube and its median split. It is continuous, it is not in the span, and it varies on every draw.
That is a small mechanical detail and it is recorded because it produced a diagnostic that returned a number rather than an error, which is the failure mode this collection watches for: an instrument that reports something when it should report nothing at all.
A statistic that could see it and does not
There is one more reading worth putting in the table, because it is the one a practitioner would defend.
A randomisation test’s own statistic — the difference in outcome means between the arms — is signed and is odd under the complement, so in principle it can tell the components apart. The reason the test does not use it is that it is a statistic about the outcomes, and the outcomes are fixed once the trial has run: every hypothetical assignment is scored on the same vector of observed values, which is what makes a randomisation test a randomisation test.
That makes it a perfectly good probe, and it makes it the wrong one for a diagnostic. A diagnostic run before the outcomes exist — which is when a practitioner would want to know whether their sampler works — has no outcome vector to score. The imbalance in a covariate function is available from the design alone, which is available before anybody is treated.
So the choice of probe is not only about parity and variability. It is about when the question is asked, and the answer a practitioner needs is one they can get in time to change the sampler.
What is claimed here, and what is not
This essay takes why the test is built the way it is. The claims are that the two components of a split admissible set carry mirror images of one distribution, so every symmetric reading of them is identical — measured as a gap of exactly 0.0000 in the mean absolute imbalance against 0.1679 in the signed mean; that the consequence for a randomisation test is a one-sided critical value of 0.4170 against a whole-set 0.7749 and a two-sided gap forty times smaller; that a large signed difference between two chains has two explanations, and a replicate chain from the same starting point separates them without depending on the autocorrelation estimate being right; and that at fourteen units both controls stay inside 2.7 standard errors at every tolerance while the signed comparison reaches 14.58.
What stays out, and is named as a decision: a general theory of which probes are best. The probe here is the fourth power because it is outside the span and varies; a systematic comparison of probes — which separates the components fastest, which needs the shortest run — is a measurement this field does not make. The one thing established is that a probe inside the span or degenerate on the units drawn is useless, and both were found by trying them.
The boundary against the essay that builds the test is that it establishes the test works and this one establishes why it has the shape it has. Neither is a claim that this is the only shape that would.
The checks, and the refusals that make them mean something
Two claims are gated and both are exact rather than statistical. The signed probe of a complement is required to be exactly minus the signed probe, and its absolute value exactly equal — checked to twelve decimal places on twenty random assignments, because the whole design of the test and of its control rest on that one line of arithmetic. And the admissible set is required to be closed under complementation, on the enumerated set, because a rule with an asymmetric constraint would break the closure quietly.
The refusal is the natural alternative. The same test run on a statistic that does not change sign is refused, with the signed comparison’s standard errors printed beside the magnitude comparison’s on the same draws: 9.5 against 1.3, on a set the enumeration says is in two pieces.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The statistic the p-value is about — both name connected component, covariate balance, effective sample size, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
- A probe nobody chose — both name connected component, covariate balance, effective sample size, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
- Before the trial and after — both name connected component, covariate balance, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
- What a chosen probe finds — both name connected component, covariate balance, effective sample size, exact enumeration, markov chain monte carlo, mixing time, randomisation test, reference distribution
- Draws that repeat each other — both name burn-in, critical value, detailed balance, effective sample size, markov chain monte carlo, randomisation test, reference distribution
- Half a reference distribution — both name connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
Named objects
A flat tag is an object no other essay names yet.
Burn-inConnected componentCovariate balanceCritical valueDetailed balanceEffective sample sizeErgodicityExact enumerationImbalanceMarkov chain Monte CarloMixing timeParityRandomisation testReference distribution