What the rule blocks
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
The field that asked which direction a diagnostic should be run in ends by naming what its best feasible probe is not. The design’s own leverage carries 3.836 of separation against a random direction’s 0.942 and an oracle’s 10.565, and it works — but leverage is a heuristic about which units the rule has most to say about, and the rule does not say anything about units.
What a balancing rule does is refuse exchanges. An assignment’s admissible neighbours are the ones a single swap can reach without leaving the tolerance box; the admissible set falls into pieces exactly because some swaps are refused; and the quantity the whole diagnostic is about is therefore which swaps are refused.
Nothing in that field uses it. This one does.
What a swap moves
The arithmetic is short and it is the whole construction.
An assignment is a vector of signs, half of them positive. Its standardised imbalance in basis column k is the signed column sum divided by the coin’s own standard deviation of the same quantity, which for an equal split of standardised units is (n/2)(2/√(n − 1)). Swapping unit i from one arm to the other and unit j back changes the signed sum by −2(c_i − c_j), so
That is a fact about the design’s columns and about nothing else. No outcomes, no assignment, no chain.
What a box refuses
The tolerance rule admits an assignment when every |z_k| is at most a. So a swap from an admissible point is refused when the displacement takes some coordinate outside the box.
Under a uniform position inside the box, a displacement of d in one coordinate keeps the point inside with probability 1 − |d|/2a, and the coordinates multiply. So
is a per-unit number computable from the design and the tolerance alone.
It is a model of the active set rather than the active set, and the model is the uniform-position one. The next essay is what that assumption costs.
Why the model, and what it assumes
The step above is exact. What follows it is not, and the approximation is worth stating precisely because it is the thing the next essay measures.
The exact statement is: from a given admissible assignment z, the swap of i and j is refused when z + Δ leaves the box. That depends on where z is, and z is an admissible assignment rather than an arbitrary point — so to turn it into a number per unit, something has to be assumed about how admissible assignments are distributed inside the box.
The assumption made here is uniform, which is the simplest thing that is not obviously wrong and is obviously wrong. An imbalance is a signed sum of n standardised terms, so it is nearly Gaussian rather than uniform, and conditioning on it being inside the box concentrates it towards the middle rather than spreading it evenly. A point nearer the middle survives more displacements than a uniform one, so the model should over-state how much is blocked.
It does the opposite, which is one of the small surprises of the next essay: the counted share is 0.8442 against the modelled 0.8170.
The other assumption is quieter. The average over j runs over every other unit, not over the units in the other arm — because the probe is built before any assignment exists and there is no other arm yet. That is not an approximation to be repaired; it is what “computable before any outcome” costs.
The check that matters
There is one place this could be silently wrong, and it is the index in the swap step.
An off-by-one there would produce a probe that is finite, is a direction, standardises fine, projects fine, and is simply not the quantity it claims to be. Every downstream number would still be reportable and none of them would look odd.
So the closed form is checked against the thing it is a closed form for: apply the exchange, recompute the imbalance from scratch with the same function the admissibility rule uses, and require the difference to match to a part in 10¹⁰. Twenty designs, three swaps each, every basis column.
A quantity that only appears in its own downstream products has to be checked against something outside them.
The same discipline applies to the tolerance. A box of 0.8 on four orthonormalised columns admits 116 of 3,432 equal splits on the design the figures are drawn from — about one in thirty — and that share is what makes the enumeration worth doing and the walk necessary. A tolerance loose enough to admit most splits would leave the set connected and nothing here would have anything to measure; a tolerance tight enough to admit a handful would leave it in pieces so obvious that no diagnostic would be needed.
The setting, which is the enumerable one
Everything here is the earlier field’s setting unchanged: fourteen units, a basis of the covariate, its square, its cube and a median cut, a tolerance of 0.8, and components taken under single swaps.
Fourteen units is 3,432 equal splits, so the admissible set can be enumerated and the answer is available. Of two hundred designs drawn, 50.0% have an admissible set that falls into two or more pieces, and those are the designs every measurement here is made on — a design whose set is connected has no separating direction and nothing to be aligned with.
At the design the figures are drawn from, the admissible set holds 116 assignments of the 3,432, and the single swaps out of them number 5,684, of which the box refuses 4,748.
Four fifths of the exchanges the rule could make, it refuses. That is not a small perturbation of a random walk; it is a walk on a graph most of whose edges are missing, which is why the set falls apart at all.
Eight neighbours out of forty-nine, and still connected on paper
The refusal count converts into a statement about the graph the walk is on, and the statement is sharper than “most of whose edges are missing”.
From any equal split of fourteen units there are 7 × 7 = 49 single swaps. The admissible set holds 116 assignments and admits 936 of the 5,684 swaps out of them, so each admissible assignment has on average 8.07 admissible neighbours — a sixth of the forty-nine available.
Eight is not a sparse graph. A random graph on a hundred and sixteen nodes is connected with overwhelming probability once its average degree passes ln 116 = 4.75, and this one is at 8.07 and in two pieces.
So the set does not fall apart because there are too few edges. It falls apart because the missing edges are missing in a structured way: the box is a product of intervals in a basis, the two components are complement pairs, and a single swap moves two units where crossing between an assignment and its complement takes seven. Adding admissible assignments — loosening the tolerance a little — would raise the degree and not touch the obstruction, which is why the diagnostic is about connectivity rather than about density.
That also settles what the fifty per cent of split designs means. It is not a marginal population of awkward designs; it is half of all designs at a tolerance a trial would use, on a graph that would be connected if its refusals were arbitrary.
The closed acceptance rate over-states this design by three
One number in the setting is worth checking against the closed form the neighbouring fields use, because the discrepancy is the field’s own warning arriving here.
The asymptotic acceptance rate for an orthonormal basis of k columns at a tolerance of a is (2Φ(a) − 1)^k, which at a = 0.8 and four columns is 11.0%. The design admits 116 of 3,432, which is 3.38%.
A factor of 3.3 low, on a design of fourteen units. The closed rate is a limit and it is quoted in this collection as arriving late — fifteen per cent high at the tightest tolerance measured on two hundred units — and at fourteen units it is out by a factor rather than by a percentage.
That matters for reading the enumeration rather than for the probe. A practitioner choosing a tolerance from the closed rate, expecting one assignment in nine to be admissible, would find one in thirty — and would be choosing a much tighter rule than intended, on a set correspondingly more likely to be in pieces.
What the probe is
A per-unit number is not yet a direction. Two more steps make it one, and both are inherited.
Standardise. Centre and scale to unit variance, so the column is comparable with the others in the table.
And project off the rule’s span. The earlier field’s first finding is that a probe inside the span the rule balanced is a probe the rule has already answered: the fourth power of the covariate is 92% inside that span at fourteen units, and projecting it off raises what it can see by a factor of three. Every feasible probe in that table is projected, so this one is too.
The projection matters more here than it looks. The blocked share is built out of the design’s columns, so a great deal of it lies inside the span those columns define — and what is left after the projection is the part that carries any information the rule has not already used.
What the diagnostic is for
One paragraph of context, because a probe is only interesting if something reads it.
A rerandomisation trial draws its assignment by walking the admissible set with single swaps, and the reference distribution its p-value is read against is what that walk reaches. If the set is in pieces the walk reaches one piece, and the reference distribution is half a reference distribution.
The two-chain test is the diagnostic: run two chains from independent starts, compute a probe’s signed imbalance along each, and compare the two means. If the set is connected both chains sample the same distribution and the two means agree; if it is in pieces and the probe distinguishes the pieces, the two means are of opposite sign.
Which is why the probe matters. A probe orthogonal to the direction that separates the pieces reports the same distribution in both, however long the chains run, and the diagnostic stays quiet on a set that is broken. So choosing the probe is choosing whether the test can see anything at all, and the whole earlier field is about that choice.
Where it lands
The table it joins has six rows in the earlier field and eight here. Over a hundred split designs, the alignment with the separating direction is:
the separating direction itself 1.0000, the fourth power projected off the span 0.6583, the design’s own leverage 0.5395, the modelled active set 0.4272, the counted active set 0.3854, a random direction in the same subspace 0.2622, the most concentrated direction a projection pursuit finds 0.2001, and the raw fourth power 0.1720.
It is a probe. It beats a random direction at 4.94 paired standard errors and it beats both of the earlier field’s own dictionary probes.
And it loses. It is below leverage at 4.43 paired standard errors and below the projected fourth power at 7.16.
The third essay of this field is that result and what to make of it. This essay’s job is the construction, and the construction is sound: the step is checked against the imbalance it moves, the share is a probability at every unit by arithmetic, and the count it approximates is checked against the chain machinery’s own neighbour count.
The second route to the count
The modelled share has a closed form and the counted one does not, so the counted one is checked against a second implementation rather than against a formula.
The tally here walks every admissible assignment, tries every single exchange out of it, and records for each unit whether the result is still admissible. The chain machinery in the field that walks these sets already has a function that returns an assignment’s admissible neighbours, written for a different purpose and sharing no loop with this one.
The two have to agree, and the arithmetic that connects them is small: every blocked exchange is tallied against both of its units, so the per-unit totals sum to twice the count of blocked exchanges, and the count of blocked exchanges is the total number of exchanges less the neighbour count summed over the set. At the design the figures are drawn from that is 9,496 against twice 4,748, exactly.
Two implementations of “which exchanges are blocked” is exactly the pair worth disagreeing, and this is the check that would catch a set built with the wrong tolerance, a neighbour function with a different notion of a swap, or an enumeration that had quietly lost half its members.
Why a per-unit quantity at all
One design choice deserves a defence, because the active set is a set of exchanges and what has been built is a number per unit.
A probe has to be a column: the two-chain diagnostic reads a vector of length n and computes its signed imbalance under an assignment. So whatever is extracted from the active set has to be summarised to one number per unit, and the natural summary is the share of the exchanges that unit takes part in that are refused.
That throws information away. The active set is a structure on pairs — which particular exchanges are blocked, and therefore how the set is connected — and averaging over j collapses it to a marginal. Two designs with very different blocking patterns can give the same per-unit shares.
What is not thrown away is what the diagnostic can use, and that is the honest defence: a probe is a direction, a direction is a column, and a column cannot carry a structure on pairs. A diagnostic that read the pair structure directly would not be this diagnostic.
What the shares look like
It is worth reading the fourteen numbers once, because a per-unit share is an unfamiliar object and the shape of it is informative.
At the design the figures are drawn from, the counted shares are 0.803, 0.813, 1.000, 1.000, 0.961, 0.867, 0.709, 0.704, 0.729, 0.729, 0.867, 0.803, 1.000 and 0.709.
Three units are at exactly one: every exchange they take part in is refused. Those are units whose columns are extreme enough that moving them across the arms breaks the tolerance whoever they are swapped with, so they sit at whichever arm the admissible assignments put them in and never move. They are, in a precise sense, the units the rule has decided about.
The rest run from 0.704 to 0.961, so even the most mobile unit in this design has seven exchanges in ten refused.
The modelled shares on the same design are 0.739, 0.758, 0.993, 0.987, 0.786, 0.847, 0.771, 0.756, 0.752, 0.751, 0.869, 0.741, 0.831 and 0.827. The model gets the three pinned units nearly right — 0.993, 0.987 and 0.831 against three ones — and compresses the rest towards the middle, which is what an assumption of uniform position inside the box would do.
A unit that never moves is what breaks an admissible set into pieces, and a probe built on this quantity is therefore aimed at something the separating direction ought to know about. Whether it is aimed well is the third essay.
What the box being a box does
One structural point, because it changes what the model above can be.
The rule here is a box: every column’s imbalance separately inside a tolerance. An ellipsoid rule — a Mahalanobis distance inside a cutoff — would give the same admissible set a different shape, and the swap-step model would be a different product.
A box is not basis-invariant, and that is the reason this matters: a box drawn in one basis is a different constraint from the same box drawn in an orthogonalised one, so the active set is a property of the rule as written rather than of the span it constrains. The blocked share inherits that: it is computed from the orthonormalised columns this collection’s designs carry, and it would be a different column on the same units under a different orthogonalisation.
That is a limitation on transporting the number and not on the argument, and it is the same limitation the acceptance rate carries.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Counting it exactly does not help — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, hat matrix, imbalance, leverage, markov chain monte carlo, randomisation test, rerandomisation
- A set of pairs, not a vector — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, imbalance, leverage, randomisation test, rerandomisation
- A defect that is about size — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- A probe nobody chose — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- Before the trial and after — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- The statistic the p-value is about — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
Named objects
A flat tag is an object no other essay names yet.
Assignment mechanismBasisClosed formConnected componentCovariate balanceDetailed balanceExact enumerationExperimental designHat matrixImbalanceLeverageMarkov chain Monte CarloRandomisation testRerandomisationStandardised difference