Counting it exactly does not help
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
The previous essay leaves an escape route, and a reader is entitled to take it. The active-set probe is built on a model — the assumption that an admissible assignment sits uniformly inside its tolerance box — and that model is measurably wrong: it is low on the level by 0.0272 and it agrees with the count about the ordering of the units at a correlation of 0.8141 rather than at one.
So of course it loses. Approximate a quantity badly and the approximation will not carry what the quantity carries.
This essay closes that route, because the exact quantity is available at fourteen units and it is on the table.
The exact column
Over every admissible assignment and every single exchange out of it, tally per unit whether the exchange leaves the set. That is the active set, not a model of it. At the design the figures are drawn from it walks 116 admissible assignments and 5,684 exchanges, and finds 4,748 of them refused.
Nothing about it is estimated. It needs the enumeration, so no trial has it, and it is on the table for the same reason the separating direction is: as the thing a feasible probe is short of.
It is also checked against a second implementation rather than trusted. The chain machinery in the field that walks these sets already returns an assignment’s admissible neighbours, written for a different purpose; every blocked exchange is tallied against both of its units, so the per-unit totals must sum to twice the count of blocked exchanges the neighbour function implies. At the design the figures are drawn from that is 9,496 against twice 4,748, exactly.
What was expected
It is worth writing down the prediction before the measurement, because the prediction was confident and it was made inside this field rather than by a reader.
The counted column is the active set. The modelled column is a uniform-position approximation to it that is demonstrably low by 0.0272 on the level and that gets the ordering of the units backwards on half a per cent of designs. A probe built on the exact quantity cannot carry less of what the exact quantity carries than a probe built on a noisy version of it.
That reasoning has one hole and it is the whole essay: it assumes the exact quantity carries something. If the quantity is only weakly related to the answer, then a noisy version of it is a mixture of the quantity and noise, and the noise can be pointed anywhere — including, by accident, somewhere better.
Which is what happens here. The model’s noise is leverage. Its departures from the count are correlated with the design’s own leverage, because the product of clamped survival factors is a smooth function of how large a unit’s columns are, and that is what leverage measures. So approximating the active set badly, in this particular way, moved the column towards a better probe.
And it is worse
The counted active set aligns with the separating direction at 0.3854 against the modelled column’s 0.4272. Paired on the design, the difference is −0.0315 at −0.83 standard errors — so on the alignment the two are indistinguishable, with the exact one nominally behind.
On the verdict they are not indistinguishable. A chain of eight hundred draws misses 39.4% of split sets with the counted probe against 29.4% with the modelled one, and 20.6% with leverage.
And against leverage the counted column is behind by 0.1511 at −3.62 paired standard errors — further behind than the model is.
The exact version of the quantity is not better than an approximation to it. Whatever cost the probe, it was not the uniform-position assumption.
What the enumeration says about the set as well as the units
The tally was built to price a probe and it measures the admissible set on the way, in a form worth reading beside what the neighbouring fields say about these sets.
Fourteen units means seven per arm, so every assignment has exactly 7 × 7 = 49 single exchanges available, and 116 × 49 = 5,684 is the whole enumeration with nothing left over. Of those, 4,748 are refused, so 936 are admitted — an average of 8.07 admissible neighbours per assignment.
Two things follow and they pull in opposite directions.
The set is dense, not sparse. Eight neighbours apiece is four times the expected degree at which a set of this shape stops falling into pieces, and this set is in two pieces anyway. Whatever separates its two components is not a shortage of edges.
It cannot be. The two components are an assignment and its complement, and a single exchange moves two units where reaching the complement moves all fourteen — seven exchanges away, at minimum, through states that need not be admissible. A distance obstruction is invisible to a degree count, and a walk with eight neighbours at every step will search half a set for ever just as surely as one with two.
And admissibility is strongly clustered. The share of proposals admitted from an admissible state is 936 / 5,684 = 16.5%, where the share of all 3,432 equal splits that are admissible at all is about 3.3%. A neighbour of an admissible assignment is five times more likely to be admissible than a stranger is, which is what a walk trades on and what a rejection sampler cannot use.
The check that the two implementations agree
The cross-check is worth one line more than it is usually given, because of what it does and does not establish.
Every blocked exchange involves two units and is tallied against both, so the per-unit totals must sum to 9,496 — twice 4,748 — and they do, exactly.
That is an identity rather than an agreement: it would hold even if both implementations shared the same wrong notion of admissibility. What it rules out is the failure mode that actually threatens a per-unit tally, which is a unit being credited once for an exchange it is half of, or the same exchange being counted from both ends into one total. Those are off-by-a-factor-of-two errors, they produce a column that is perfectly correlated with the right one, and no downstream comparison in this field would notice.
What the exact column costs to have
It is worth pricing the thing that turned out not to be worth having, because the price is the reason nobody would run it and the reason it is worth running once.
At fourteen units there are 3,432 equal splits. Screening each against the box is four sums and four comparisons; the admissible set that survives holds around a hundred. Walking every single exchange out of every member is another 5,684 admissibility checks. All told it is a few milliseconds per design, which is why two hundred designs is affordable here.
At twenty units there are 184,756 splits, which is where the field that enumerates these sets stops. At forty there are about 1.4 × 10¹¹, and the exchange walk is a further factor of four hundred on top of whatever survives the screen.
A trial has forty units or two hundred, and the diagnostic exists precisely because the admissible set cannot be enumerated at those sizes. So the counted column is available exactly where nobody needs it — the same shape as the separating direction, which is also computed here and is also unavailable to a trial.
That is the standing arrangement in this line of fields: everything is measured at a size where the answer is known, so that the feasible things can be scored against it, and the scores are then a statement about what a trial gives up. Here the answer is that it gives up nothing.
Which is the result the field turns on
It is worth being explicit about the logic, because a null comparison is doing the work.
Had the counted column beaten the modelled one, the reading would have been: the active set is the right quantity, the model is a poor approximation, build a better model. That is a normal outcome and it would have left the deferral’s argument intact.
Had it beaten leverage, the reading would have been sharper still: the active set is the right quantity, it needs the enumeration, and here is what a trial is giving up by not having it.
Neither. The exact quantity is no better than its approximation and both are behind a heuristic, so the deferral’s argument fails at its premise rather than at its execution. What the rule blocks is not what the diagnostic needs to know.
Why the two columns differ so much and land so close
The two active-set columns are not near-copies of each other, which makes their near-identical performance more interesting than it looks.
Projected off the rule’s span and standardised, the modelled column agrees with the design’s own leverage at |r| = 0.8359 ± 0.0114. The counted column agrees with leverage at 0.4239 ± 0.0216 — half as much.
So they are two genuinely different directions. One is largely leverage with a sixth of something else; the other is about as far from leverage as it is from the model. And they align with the separating direction at 0.4272 and 0.3854 — within a standard error of each other.
There is a way to check that reading rather than assert it. If the shared part were doing the work, then a probe made of only the shared part — leverage itself — should carry at least as much as either of them. It carries 0.5395 against 0.4272 and 0.3854, which is more than either. Two directions that share less than half of themselves land in the same place, and the thing they share carries more than either of them does. That is what it looks like when the quantity they are both about is only weakly related to the answer: the part they have in common — leverage — carries most of what either of them carries, and their departures from it are worth about the same as each other, which is nothing much.
What “no better” means on 97 designs
The null comparison is doing the work, so it deserves the treatment a positive one would get.
The counted probe is computable on 97 of the hundred split designs rather than all hundred, because on three of them the enumerated shares come out constant across units and standardising a constant column returns nothing. The paired comparison is made on the designs where both exist.
The paired difference in alignment is −0.0315 with a t of −0.83. That is not a demonstration that the two probes are identical; it is a failure to separate them at this sweep length, and a difference of a twentieth in alignment would need several times the designs to resolve.
What the sweep does separate is everything the field’s conclusion rests on. The counted column against leverage is −0.1511 at −3.62, and the modelled column against leverage is −0.1124 at −4.43. Both are decisive at this length, and both point the same way.
So the honest statement is: the exact column is not measurably better than its model, and both are measurably worse than a heuristic. The first half is a null result and the second half is not, and only the second half is load-bearing.
The one place the exact column is ahead
There is a column where the counted probe reads higher, and it should be reported rather than left out.
On the enumerated separation — the distance between the two components’ mean probe values over the spread inside a component — the counted column reads 2.0170 against the modelled column’s 1.9529, a paired difference of +0.9779 at 0.80 standard errors.
That is a non-difference, and its sign is the opposite of the alignment’s. Two readings of the same pair of probes, neither separating, pointing different ways.
The verdict column is what breaks the tie: 39.4% missed against 29.4%, over 34 split designs. A ten-point gap on that few designs is itself not decisive — a difference of three or four designs — but it is the reading a practitioner would experience, and it points the same way as the alignment.
The honest summary is that the two probes are within noise of each other on two readings and the exact one is behind on the third. What is not within noise is that both are behind leverage.
The escape routes that are left
Two remain, and both are named rather than closed, because closing them is not this field’s work.
A different summary of the active set. Everything here collapses a structure on pairs to one number per unit, by averaging over the partner. A diagnostic that read the pair structure — which particular exchanges are refused, and therefore how the graph of admissible assignments is actually connected — would be using more of the quantity. It would also not be this diagnostic, which reads a column of length n and computes its signed imbalance, so it would be a different test rather than a better probe.
And a different constraint. The rule here is a box: every column’s imbalance separately inside a tolerance. An ellipsoid rule gives the same span a differently shaped admissible set, and its active set would be a different quantity — plausibly one more aligned with the span, since an ellipsoid is a function of the whole imbalance vector rather than of its coordinates one at a time. Whether the argument fares better there is untested.
What is closed is the version that was proposed. The active set of a box, summarised per unit, is worse than the design’s own leverage as a probe, whether it is modelled or counted, on the alignment and on the separation and on the verdict a short chain returns.
What is actually wrong with the argument
The deferral reasoned from a true premise to a false conclusion, and it is worth locating the step.
True: a balancing rule breaks the admissible set into pieces by refusing exchanges, so the pieces are a consequence of the active set.
Also true: leverage is a heuristic about which units the rule has most to say about, and it is not the active set.
False: therefore the active set is a better probe.
The gap is that the diagnostic does not need to know which exchanges are refused. It needs a direction that distinguishes the two components — a column whose signed imbalance differs between them. The active set is a fact about the graph’s edges; the separating direction is a fact about which side of the split each unit tends to sit on. Those are related and they are not the same, and the relation is weak enough that a per-unit summary of the first carries less of the second than a simpler summary of the design does.
A unit can have every exchange refused because its own column is extreme in a coordinate that has nothing to do with how the set falls in two. The blocked share counts it as maximally important; the separating direction does not care about it at all.
What the field leaves
Four essays, two new probes, and a table whose best feasible row did not change. Three things are worth carrying.
The construction is sound and reusable. The swap step is checked against the imbalance it produces, the counted share against the chain machinery’s own neighbour count, and the modelled share is a probability at every unit by arithmetic. A negative finding from broken machinery is indistinguishable from a negative finding, so the checks matter more here than in a field that found something.
A quantity being the right one to think about does not make it the right one to measure. The active set is what the rule does; leverage is a summary of the design. The diagnostic wants the second, and no amount of getting the first exactly right changes that.
And an exact version is the way to test an approximation’s excuse. Had only the model been built, the field would have ended with a probe that lost and a plausible reason it lost, and the plausible reason would have been wrong. Building the infeasible version cost one enumeration per design and settled it.
That last is the transferable half. Whenever a feasible approximation to a quantity underperforms, there are two explanations — the approximation is bad, or the quantity is the wrong one — and they recommend opposite next steps. If the exact version is computable anywhere, computing it once is what tells them apart, and it is usually cheaper than building the better approximation the first explanation asks for.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a chosen probe finds — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, leverage, markov chain monte carlo, model diagnostics, orthogonality, projection, randomisation test
- A probe nobody chose — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, projection, randomisation test, rerandomisation
- A test rather than a survey — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- Before the trial and after — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- The statistic the p-value is about — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
- Half a reference distribution — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test
Named objects
A flat tag is an object no other essay names yet.
Approximation errorAssignment mechanismConnected componentCovariate balanceExact enumerationExperimental designHat matrixImbalanceLeverageMarkov chain Monte CarloModel diagnosticsOrthogonalityProjectionRandomisation testRerandomisation