A set of pairs, not a vector
Worth reading first: Randomisation is not balance · Balancing what is known in advance.
The essay that computed the active set exactly ends by naming one thing it did not try. Every probe built from what a balancing rule blocks — modelled from the design and the tolerance, or counted over the whole enumerated admissible set — collapses the quantity to one number per unit by averaging over the partner. A blocked exchange is a pair of units. The averaging throws the pairing away.
That is a specific complaint with a specific test behind it, and this essay runs it.
The vector that was a degree
Write the quantity out as it actually is. For units and let be the share of the exchanges between them that the tolerance box refuses — modelled under the uniform-position assumption, or counted over every admissible assignment where that exchange was available. is symmetric, because an exchange has no direction, and its diagonal is empty, because a unit is never exchanged with itself.
The per-unit share every earlier probe is built from is
which is the degree of unit in the weighted graph . That is not an analogy. It is checked as an identity, at every unit of every design in the sweep, against the two loops that produced the per-unit shares in the first place — one arithmetic route through the survival product and one through the enumeration’s tally, neither of which knows the matrix exists.
So the question is not whether the pairing was ignored. It was, exactly, and by the amount a degree sequence differs from the graph it came from.
That identity is the one check this essay would be worthless without, and it is worth saying why it is checked twice rather than asserted once. The modelled matrix and the modelled per-unit share are built from the same product of clamped survival factors, so their agreement is arithmetic and a failure would mean a transcription slip. The counted pair tally and the counted per-unit share are not: one increments an entry of a matrix and the other increments two entries of a vector, over the same double loop, and the row mean of the first has to be the availability-weighted mean of the second. A tally that credited a unit once for an exchange it is half of would produce a matrix perfectly correlated with the right one, a degree sequence off by a factor of two, and no downstream number that looked wrong. The whole of this essay is a comparison between a matrix and its own row means, so an error that moved one and not the other is the error it is least equipped to notice.
How much of it a per-unit summary discards
The degree is a summary, and a summary can be complete. If were well approximated by up to a constant, then a probe built from the pairing would be a probe built from the degrees and there would be nothing here.
The measurement is the share of the off-diagonal variation the additive degree fit leaves. Over the designs the earlier essays are measured on it is 0.6902 ± 0.0066 for the modelled matrix and 0.7540 ± 0.0099 for the counted one. More than two-thirds of the object, and on the exact version more than three-quarters, is outside the summary of it every probe so far has used.
A second reading says the same thing in a form a degree sequence cannot express at all. Weight each pair by how blocked it is and correlate the two endpoints’ degrees: the exchange graph is disassortative, at −0.1484 ± 0.0042 modelled and −0.1701 ± 0.0049 counted. A refused exchange joins a heavily-blocked unit to a lightly-blocked one more often than it joins two heavily-blocked ones. Both numbers are more than thirty standard errors from zero, so this is not a rounding artefact of a nearly-saturated matrix; it is an arrangement.
That is worth pausing on, because it is the first thing in this line of essays that could not have been said before. Every quantity the earlier probes report is a function of a column of length fourteen. Assortativity is a function of ninety-one numbers and is invariant to nothing the column can change. The pairing is real structure, it is most of the object, and until now none of it had been read.
Two probes a graph has and a vector does not
A probe here is a column of length whose signed imbalance the diagnostic reads, so the pairing has to be turned back into a column before it can compete. There are two obvious ways and they are not the same.
The direction. Subtract the additive degree fit, leaving off the diagonal and zero on it, and take the dominant eigenvector of . That is the leading direction of exactly what the degree throws away, and it is orthogonal to the degrees by construction rather than by luck.
The cut. Take the signs of that eigenvector. This is coarser — it keeps which side of a partition each unit is on and discards how far — and it is the natural summary of a graph rather than of a matrix.
Both are computed twice, once from the modelled blocking matrix, which needs only the design and the tolerance, and once from the counted one, which needs the enumeration and so is available to nobody running a real trial. Both are then standardised and projected off the rule’s own span, exactly as every feasible probe in this line is, because a probe inside the span is a probe the rule has already answered. Four new rows, on the same designs and the same seeds as the eight already on the table, through the same table-building function — so the eight old rows come out identical to the last bit, and a difference between the two tables would mean the construction had moved rather than that the probes had.
Counted, the pairing recovers most of what was lost
The counted per-unit share aligns with the separating direction at 0.3854. The counted pairing’s own direction reads 0.5266 and the cut it induces 0.5258.
Paired on the design, so that no design-to-design variation is carried twice, that is a gain of 0.1326 ± 0.0388 at 3.42 standard errors for the direction and 0.1404 at 3.49 for the cut. On the separation each probe carries — how far apart the two components of the admissible set sit on it, over the spread inside a component — the same change runs from 2.0170 to 3.5949 and 4.2786.
So the complaint was right about the object. Reading the active set as a graph rather than as a vector recovers something real, and it recovers most of the gap between the per-unit share and the heuristic that beat it.
Most, and not all. The design’s own leverage reads 0.5395. Against it the counted pairing is −0.0129 ± 0.0354, at −0.37 paired standard errors, and the cut −0.0101 at −0.24. Neither of those separates anything. The honest statement is that the pairing draws level with leverage — not that it ties it, which would be a demonstration, and not that it loses, which the sweep cannot show either.
And it does not pass the probe a trial can already build
Against the projected fourth power — the best feasible probe the field that opened this question ever found — the counted pairing is −0.1317 ± 0.0330, at −3.99 paired standard errors. That is a loss and it is decisive.
Which settles the practical question the whole line of essays exists to answer. Whatever a trial should run its diagnostic in, it is not a summary of the constraint’s active set: not the per-unit share, modelled or counted, and not the pairing either. The best feasible probe on a table of twelve is the one it was on a table of five, before any of this was proposed.
The reason this is worth saying carefully is that the earlier finding was stronger than the evidence now supports. The essay that first priced the active set against leverage concluded that what the rule blocks is not what the diagnostic needs to know. That reading survives at the level of the recommendation and does not survive as a statement about the quantity: a probe that recovers 0.1326 of alignment when the pairing is kept is a probe whose quantity was carrying something the summary destroyed.
Modelled, the direction recovers nothing at all
The counted matrix needs an enumeration, so the reading above is a statement about what a trial gives up rather than about what it could do. The modelled matrix needs only the design and the tolerance. What does the pairing buy there?
The modelled per-unit share aligns at 0.4272. The modelled pairing’s own direction aligns at 0.4274 — a paired gain of 0.0003 at 0.01 standard errors, which is not a small effect but the absence of one, on a matrix two-thirds of which lies outside the degree.
The cut is different. It reads 0.4954, a gain of 0.0793 at 1.91 standard errors: nominal rather than decisive, and pointing the right way. Against leverage the modelled cut is −0.0387 at −0.86, which again separates nothing.
So the four new rows split, and they split along the wrong axis. It is not that counted beats modelled everywhere: the counted cut beats the modelled cut by 0.0354 at 0.78 paired standard errors, which is nothing, while the counted direction beats the modelled direction by 0.0992 at 3.41, which is everything. The coarse reading of the pairing survives the approximation. The fine one does not.
The obvious explanation, checked and refused
There is a ready account of why a modelled pairing should be worth less than a counted one, and it was checked before anything was concluded from it, because it is the account the previous essay in this field would have reached for. The model is measurably wrong — it assumes an admissible assignment sits uniformly inside its tolerance box, which it does not — so presumably it gets the pairing wrong.
It does not. The two blocking matrices agree entry for entry at 0.8041 ± 0.0162. Their degrees agree at 0.7742 ± 0.0223. And their residuals past the degree — the pair structure itself, with everything the two already agree about removed — agree at 0.8219 ± 0.0149, which is higher than the agreement on the whole matrix rather than lower.
The uniform-position model is not missing the pairing. Whatever separates the two directions is finer than the pair structure is.
What it is can be measured. The two residual eigenvectors put 0.8274 ± 0.0109 of the units on the same side of the cut — eleven and a half of fourteen, read up to a global flip, since a cut and its complement are one cut. They agree about the partition and disagree about the magnitudes, and an eigenvector is a fine functional of a matrix while a sign pattern is a coarse one. Once the magnitudes are discarded the two probes are nearly the same object, which is why the cut survives the approximation and the direction does not.
That is the mechanism as far as this essay can establish it, and it is worth being clear that it is a description rather than a derivation. Read as columns, after the projection and the standardisation, the two directions agree at 0.7129 ± 0.0193 and the two cuts at 0.5778 ± 0.0255 — which is the wrong way round if column agreement were what decided performance. The alignment a probe delivers is not a smooth function of the column, and nothing here says by how much it fails to be.
What a chain of eight hundred draws makes of it
The alignment and the separation are population quantities computed on an enumerated set. A practitioner has a chain, and a chain of finite length is what decides whether a split set is reported as split.
On that reading the counted pairing does slightly better than the enumeration suggests. It misses 17.6% of the split sets against the counted per-unit share’s 39.4% and the design’s own leverage at 20.6%, with its cut at 18.2%. The projected fourth power still misses least of anything feasible, at 11.8%, and the separating direction itself — which no trial has, because having it means already knowing the answer — misses 8.8%.
The gap between the pairing and leverage there is three points on thirty-four designs, which is one design. It is reported because leaving out the reading that happens to favour the new probe would be worse than reporting a number too small to lean on, and it is not leaned on.
What the chain does confirm is the ordering that matters. The pairing is a large improvement on the vector it replaces — twenty-two points of missed sets — and it is not an improvement on the probe that was already available.
What this changes about the field’s conclusion
Three sentences, and the third is the one that moved.
The recommendation is unchanged. A trial running a randomisation diagnostic should run it in the projected fourth power, not in anything built from what its balancing rule refuses. That was true against a table of five probes, against a table of eight, and it is true against a table of twelve.
The exactness is still not what was costing the probe. Counting the active set instead of modelling it gains nothing on the per-unit reading and gains on the pair reading only because the model’s magnitudes are wrong; both facts sit inside a comparison that leverage wins either way.
But the quantity was carrying something, and the summary was destroying it. The earlier reading — that what the rule blocks is simply not related to where the set splits — is too strong. What the rule blocks, read as the set of pairs it is, is about as informative as the design’s own leverage. It is a summary of it, taken by averaging over the partner, that reads as barely better than noise.
That distinction has a use beyond this field, and it is the transferable half. A quantity that underperforms as a probe has three explanations rather than two: the quantity is wrong, the approximation to it is wrong, or the summary of it is wrong. The first two are the ones anybody tests, because they are the ones a modelling argument suggests. The third is invisible, because the summary usually arrives with the quantity and is never separately named — here it arrived as an average over inside a loop, and it cost a factor of a seventh in alignment and twenty-two points of detection.
What this opens and does not measure
Three things, each named where it could have been run.
A cut is a partition, and a partition of the units is not the only one available. The cut used here comes from the signs of one eigenvector of one residual matrix. The admissible set falls into two connected components, and those are a partition of assignments rather than of units; whether the pairing’s cut and the set’s own split are related as partitions is a question about two objects of different types and is untouched. It would need a way of pushing a partition of assignments down onto the units, and the obvious ones are exactly the per-unit averages this essay is about.
The second eigenvector is not read. The residual matrix has fourteen of them and only the dominant one is used, on the reasoning that a probe is one column. A diagnostic that read two columns is a different test — the field that walks these sets already has statistics that take more than one — and nothing here says whether the second direction carries anything.
The blocking matrix is a weighted graph and is read here as an unweighted one, twice over. Both the eigenvector and the cut treat every entry of the residual on the same footing, where the entries themselves run from nought to one and on one design seventy-three of ninety-one sit at the ceiling. A saturated pair carries no information about the design and a great deal of weight in the eigenvector, and a probe that dropped the saturated pairs before taking the direction is a different probe. Nothing here says whether it is a better one, and it is the cheapest of the three things left.
And the whole thing is measured at fourteen units. That is where the admissible set can be enumerated, and the enumeration stops not far past it. The pair matrix is numbers against a degree sequence’s , so the share of the object a per-unit summary discards should grow with the trial size — and every number above is measured at the one size where the counted version exists at all. Whether the pairing’s advantage over the vector grows, holds or fades at forty units is untested, and the modelled matrix, which needs no enumeration, is the one instrument that could answer it.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What the rule blocks — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, imbalance, leverage, randomisation test, rerandomisation
- A probe nobody chose — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, projection, randomisation test, rerandomisation
- A test rather than a survey — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, randomisation test, rerandomisation
- Before the trial and after — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, randomisation test, rerandomisation
- The statistic the p-value is about — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, randomisation test, rerandomisation
- When the constraints run out — both name assignment mechanism, covariate balance, exact enumeration, experimental design, imbalance, randomisation test, rerandomisation
Named objects
A flat tag is an object no other essay names yet.
Assignment mechanismAssortativityConnected componentCovariate balanceEigenvectorExact enumerationExchange graphExperimental designGraph cutImbalanceLeverageModel diagnosticsProjectionRandomisation testRerandomisation