What a chain cannot report

The statistic that changes sign

A test for an unreachable half needs a quantity that tells one half from the other. Every symmetric reading of a mirror pair is identical, and a magnitude is the natural thing to reach for.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

The test that does not enumerate compares two chains, one started at an assignment and one at its complement, on the signed imbalance in a function the balancing rule was not handed. Two of those choices look like details and neither is.

Signed is why the test can see anything at all. And two chains is not enough — a third is needed, and what it rules out is the commoner of the two explanations for a large reading.

Every symmetric reading is blind

The admissible set is closed under complementation, and when it splits it splits into an arrangement and its mirror image. So the two components carry the same distribution with the sign reversed.

Two halves of one reference distribution. The reference distribution of a randomisation test on the 116 admissible assignments of a fourteen-unit trial, drawn separately for each of the two components the walk cannot cross between. They are mirror images: the means are 0.084 and -0.084, and the spreads are 0.389 and 0.389. A chain reports one of them, chosen by where it happened to start. Every symmetric reading of the two is identical — the absolute values differ by 0.000 against 0.168 for the signed means — which is why a two-sided test is untouched by this and a one-sided one is read off the wrong half.
Fig. 1 The two components of a fourteen-unit admissible set, enumerated. Means +0.0840 and −0.0840, spreads identical.

Anything computed from a distribution that does not change when the whole distribution is reflected is the same number on both components. The absolute imbalance is one. So is the variance, the interquartile range, the two-sided critical value, and the acceptance rate.

The enumeration makes it exact rather than approximate. The two components’ means of the signed statistic differ by 0.1679; the two components’ means of its absolute value differ by 0.0000. Not a small number — zero, because the multisets are reflections of each other and reflection does not move a magnitude.

A magnitude is the natural statistic to reach for, because an imbalance is usually reported as a size. Reaching for it here produces a diagnostic that is exactly blind to the thing it is pointed at.

Where the zero comes from

The exact zero above is worth one more sentence, because “the multisets are reflections of each other” is a claim about the enumerated set and the reason it holds is a claim about the construction.

Every assignment in a component has its complement in the other, and the complement map is a bijection between the two — that is what mirrored means and it is measured rather than assumed, as the share of assignments whose complement lies in a different component, which comes out at 1.00. Under a bijection that negates the statistic, the second component’s multiset of values is exactly the first’s negated. Taking absolute values then gives the identical multiset, so every quantile, every moment and every function of the magnitudes matches to the last bit rather than to a tolerance.

That is why the number reported is 0.0000 rather than 0.0003. It is not a small difference that could grow on a different design; it is an identity, and the only way for it to fail is for the components not to be a mirror pair — which is a case the enumeration also produces, at tolerances thin enough to leave four pieces rather than two.

The blindness is in the summary, not in the covariate

It is easy to read the parity requirement as a constraint on which covariate function to probe with, and it is not. Every imbalance is already the right parity.

An imbalance is izif(xi)\sum_i z_i f(x_i) with zi=±1z_i = \pm1linear in the assignment, whatever ff is. So complementing the assignment negates it, exactly, for a covariate, for its square, for its fourth power, for a random direction, for anything at all.

What is even is the summary. An absolute value, a variance, an interquartile range, a two-sided critical value, an acceptance rate: each of those is a function of the imbalance that does not change when the imbalance changes sign, and each is therefore identical on the two components.

So the failure has one step in it and it is not the step anybody would examine. Choosing the covariate function well is the part this line of fields spends its effort on, and no choice there can rescue a diagnostic that takes a magnitude at the end. Choosing to keep the sign costs nothing and is the whole of the repair.

Thinness stays put and reachability does not. The same balancing rule and the same tolerance at five trial sizes. The admitted share barely moves — one admissible assignment in 27, 30, 30, 20, 18 — because the acceptance rate of a rerandomisation is a fact about the basis rather than about the number of units. What does move is the number of single swaps available: 36, 49, 64, 81, 100, growing like a quarter of the square of the trial size. Filled marks are sizes whose admissible set falls into more than one piece. The set is in 4 pieces at 12 units, 2 at 14, and one piece from 16 upwards. The fourteen-unit result is a statement about fourteen units.
Fig. 2 The same rule and tolerance at five trial sizes. The admitted share barely moves — one assignment in 27, 30, 30, 20, 18 — while the number of single swaps grows 36, 49, 64, 81, 100. Filled marks are the sizes whose admissible set falls into more than one piece, which is what a symmetric summary cannot see.

Which makes the decomposition exhaustive

Because every statistic splits into an even part and an odd part, and the two components’ distributions are exact reflections, the accounting has no residue.

The odd part has means that are exact negatives on the two components — here +0.0840 and −0.0840, differing by 0.1679, which is 2 × 0.0840 to the last digit printed.

The even part has means that are exactly equal — differing by 0.0000, not by something small.

There is no third part. A statistic’s usable separation between the components is entirely the separation of its odd part, and its even part contributes nothing but noise to the comparison. A probe that is ninety per cent even and ten per cent odd is not a slightly worse probe; it is a probe carrying a tenth of a signal underneath nine tenths of a common component that cancels in the difference and does not cancel in the variance.

That is why the exact zero is worth reporting as a zero. It is not a measurement that happened to come out small on this enumeration; it is what the bijection between the components forces, and the measured share of assignments whose complement lies in the other component — 1.00 — is the statement that the bijection is there.

And so is the test most trials run

The same blindness has a consequence that is not about diagnostics at all, and it is why the defect matters unevenly.

A randomisation test’s critical value is a quantile of the reference distribution. Read one-sided, the whole set’s 95% point is 0.7749 and the two components give 0.9344 and 0.4170 — so a chain confined to one of them reports a critical value nearly half the right one, or a fifth too large, depending where it started. Read two-sided, the whole set gives 0.9434 and each component gives 0.9344: a gap of 0.0090 against the one-sided gap of 0.3579, a factor of forty.

So a two-sided randomisation test on a set the chain reaches half of is essentially correct, and a one-sided one is not. The property that makes the first true — every symmetric reading is unaffected — is exactly the property that makes a symmetric diagnostic useless.

The parity, stated

The blindness has a one-line statement and it is the same statement the essay on the zero that rests on a symmetry makes about a balancing dictionary, which is worth noticing because the two fields are about different things.

The complement map takes an assignment to its negative. The signed imbalance in any function is odd under that map: negate every arm and the sum negates. Any function of the signed imbalance that is even under negation — a magnitude, a square, a two-sided quantile — is therefore constant across the two components, because they are each other’s images.

So which statistics can see the split is decided by parity, and by nothing about the trial, the units, the tolerance or the basis. That is the same shape as the parity result about interactions in what a dictionary buys: a symmetry of the law decides which inner products are exactly zero, and the answer does not depend on the sizes of anything.

Three readings, one verdict

Three readings, one verdict. Every tolerance of a fourteen-unit trial, with three comparisons on each. The first is between a chain started at an assignment and a chain started at its complement, which is what a mirror split separates. The second is between two chains started at the same assignment on different streams, which nothing about the set can separate — so a large reading there says the run is too short and not that the set is in pieces. The third is the same comparison on the statistic's absolute value, which is symmetric under the complement and therefore blind to the split by construction. The verdict is the pattern rather than any one line: the split is called only where the first fires and the other two do not, which happens at exactly the tolerances the enumeration calls disconnected — 0.8, 0.75, 0.7.
Fig. 3 The three comparisons, at every tolerance of a fourteen-unit trial.

A large signed difference between two chains has two explanations and only one of them is about the set.

They are in mirror components and cannot reach each other. Then their absolute-value distributions are identical and their signed means are opposite.

They simply have not mixed. Then they are in two arbitrary regions of one connected set, and they differ on everything — the signed statistic, its magnitude, and any other reading anybody takes.

Two comparisons separate the two, and the second is the one that does not depend on any estimate being right.

The replicate is a third chain, started at the same assignment as the first, on a different stream. Nothing structural about the set can separate two chains started at the same point: whatever they disagree about is what a run of this length disagrees with itself about. So a replicate reading of the same size as the complement reading says the run is too short, whatever the mirror pair reports.

The magnitude comparison is the other control, and it is the one the symmetry predicts: quiet under a mirror split, loud under a failure to mix.

At fourteen units both controls stay inside 2.7 standard errors at every tolerance, split or not, while the signed comparison runs to −14.58 on the split sets. That combination — one loud, two quiet — has exactly one explanation.

Why the replicate is the load-bearing one

The magnitude control is elegant and the replicate control is the one that works, and the difference between them is worth stating.

The magnitude comparison is quiet under a mirror split only once both chains have explored their own components well enough for the magnitude to have settled. Under a chain that has barely moved, the two magnitudes are two arbitrary numbers and the comparison is loud — so the magnitude control cannot distinguish a mirror split from a stalled chain when the chain is stalled, which is exactly the case it is needed for.

The replicate has no such condition. Two chains from one point are exchangeable by construction, at any run length, on any statistic. Their difference has mean zero whatever the mixing is, and its standard error is estimated from the same autocorrelated sequences the main comparison uses — so if the standard error is wrong, it is wrong in both places and the ratio is right.

A control that does not depend on the estimate it is controlling is the useful shape, and it is why the test has three chains rather than two.

The standard error is the thing being controlled

That last point deserves its own paragraph, because it is what the third chain is really for.

The comparison divides a difference of means by a standard error, and the standard error is computed from an estimated integrated autocorrelation time. At fourteen units that estimate is about thirty and the sequences are twenty thousand long, so it has plenty of data and is believable. At two hundred units it is about three hundred and fifty, and the sequences are not long enough for it — the estimate truncates early, reports too little dependence, and produces a standard error several times too small.

An instrument whose standard error is wrong reports significance everywhere. The replicate comparison uses the same wrong standard error on a difference that is genuinely zero, so it reports the same spurious significance — and the test’s verdict, which requires one reading loud and the other quiet, correctly refuses to say anything.

The control’s job is not to be right about the set. It is to be wrong in the same way, so that the comparison of the two is right.

What the three readings cost

Running three chains rather than one is three times the work, and it is worth being clear that nothing cheaper does the job.

One chain and its own diagnostics is what the field already had, and all of them pass on a set the chain reaches half of.

Two chains from independently hunted starting points would detect a split only when the two happened to land in different components, which is half the time, and would give no way of telling that from a failure to mix. The complement is what makes the two starting points guaranteed to be in different components when there are two.

Two chains from an assignment and its complement, without a replicate, is the test as first built, and it is a coin at two hundred units: it reports a split at some run lengths and not at others on the same set, because the reading it depends on is the standard error rather than the structure.

So three chains is the cheapest arrangement that answers the question, and the third one is not redundancy — it is the only part that does not assume the thing being estimated.

Choosing the probe

One more choice is a choice: which function to measure the imbalance in.

It has to be a function the balancing rule was not handed, because a function the rule holds has its imbalance constrained to a box and carries almost no information about which component an assignment is in. It also has to vary enough on the units drawn to have a distribution at all — the first version of this test used a threshold at one standard deviation on fourteen units, where exactly one unit exceeded the threshold, so the signed statistic took two values and its magnitude took exactly one. Every magnitude comparison in that version returned zero over zero.

The probe used throughout is the fourth power of the covariate, standardised, on a rule handed the covariate, its square, its cube and its median split. It is continuous, it is not in the span, and it varies on every draw.

That is a small mechanical detail and it is recorded because it produced a diagnostic that returned a number rather than an error, which is the failure mode this collection watches for: an instrument that reports something when it should report nothing at all.

What the diagnostic says at two hundred units. The same test run 8 times on independent streams, at five tolerances of a two-hundred-unit trial, 40,000 steps each. At the loosest tolerance every run says the same thing — the walk reaches the whole set — and it keeps saying it as the set is thinned. Past a point the runs stop agreeing with each other: at the tightest tolerance here 6 of 8 report that the chains have not mixed and 2 report a split, which is a diagnostic disagreeing with itself rather than a property of the set. That disagreement is the honest answer at this size, and it is one nothing in this collection could give before: an enumeration stops at about twenty-four units.
Fig. 4 The same test on eight independent streams at five tolerances of a two-hundred-unit trial. At the loosest tolerance every run agrees; at the tightest, six of eight report that the chains have not mixed and two report a split — which is the diagnostic disagreeing with itself rather than a property of the set.

A statistic that could see it and does not

There is one more reading worth putting in the table, because it is the one a practitioner would defend.

A randomisation test’s own statistic — the difference in outcome means between the arms — is signed and is odd under the complement, so in principle it can tell the components apart. The reason the test does not use it is that it is a statistic about the outcomes, and the outcomes are fixed once the trial has run: every hypothetical assignment is scored on the same vector of observed values, which is what makes a randomisation test a randomisation test.

That makes it a perfectly good probe, and it makes it the wrong one for a diagnostic. A diagnostic run before the outcomes exist — which is when a practitioner would want to know whether their sampler works — has no outcome vector to score. The imbalance in a covariate function is available from the design alone, which is available before anybody is treated.

So the choice of probe is not only about parity and variability. It is about when the question is asked, and the answer a practitioner needs is one they can get in time to change the sampler.

What is claimed here, and what is not

This essay takes why the test is built the way it is. The claims are that the two components of a split admissible set carry mirror images of one distribution, so every symmetric reading of them is identical — measured as a gap of exactly 0.0000 in the mean absolute imbalance against 0.1679 in the signed mean; that the consequence for a randomisation test is a one-sided critical value of 0.4170 against a whole-set 0.7749 and a two-sided gap forty times smaller; that a large signed difference between two chains has two explanations, and a replicate chain from the same starting point separates them without depending on the autocorrelation estimate being right; and that at fourteen units both controls stay inside 2.7 standard errors at every tolerance while the signed comparison reaches 14.58.

What stays out, and is named as a decision: a general theory of which probes are best. The probe here is the fourth power because it is outside the span and varies; a systematic comparison of probes — which separates the components fastest, which needs the shortest run — is a measurement this field does not make. The one thing established is that a probe inside the span or degenerate on the units drawn is useless, and both were found by trying them.

The boundary against the essay that builds the test is that it establishes the test works and this one establishes why it has the shape it has. Neither is a claim that this is the only shape that would.

The checks, and the refusals that make them mean something

Two claims are gated and both are exact rather than statistical. The signed probe of a complement is required to be exactly minus the signed probe, and its absolute value exactly equal — checked to twelve decimal places on twenty random assignments, because the whole design of the test and of its control rest on that one line of arithmetic. And the admissible set is required to be closed under complementation, on the enumerated set, because a rule with an asymmetric constraint would break the closure quietly.

The refusal is the natural alternative. The same test run on a statistic that does not change sign is refused, with the signed comparison’s standard errors printed beside the magnitude comparison’s on the same draws: 9.5 against 1.3, on a set the enumeration says is in two pieces.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • The statistic the p-value is about — both name connected component, covariate balance, effective sample size, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
  • A probe nobody chose — both name connected component, covariate balance, effective sample size, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
  • Before the trial and after — both name connected component, covariate balance, ergodicity, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution
  • What a chosen probe finds — both name connected component, covariate balance, effective sample size, exact enumeration, markov chain monte carlo, mixing time, randomisation test, reference distribution
  • Draws that repeat each other — both name burn-in, critical value, detailed balance, effective sample size, markov chain monte carlo, randomisation test, reference distribution
  • Half a reference distribution — both name connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, reference distribution

Named objects

A flat tag is an object no other essay names yet.

Burn-inConnected componentCovariate balanceCritical valueDetailed balanceEffective sample sizeErgodicityExact enumerationImbalanceMarkov chain Monte CarloMixing timeParityRandomisation testReference distribution