A probe from what the rule blocks

A set of pairs, not a vector

The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

The essay that computed the active set exactly ends by naming one thing it did not try. Every probe built from what a balancing rule blocks — modelled from the design and the tolerance, or counted over the whole enumerated admissible set — collapses the quantity to one number per unit by averaging over the partner. A blocked exchange is a pair of units. The averaging throws the pairing away.

That is a specific complaint with a specific test behind it, and this essay runs it.

The vector that was a degree

Write the quantity out as it actually is. For units ii and jj let BijB_{ij} be the share of the exchanges between them that the tolerance box refuses — modelled under the uniform-position assumption, or counted over every admissible assignment where that exchange was available. BB is symmetric, because an exchange has no direction, and its diagonal is empty, because a unit is never exchanged with itself.

The per-unit share every earlier probe is built from is

di=1n1jiBij,d_i = \frac{1}{n-1}\sum_{j \neq i} B_{ij},

which is the degree of unit ii in the weighted graph BB. That is not an analogy. It is checked as an identity, at every unit of every design in the sweep, against the two loops that produced the per-unit shares in the first place — one arithmetic route through the survival product and one through the enumeration’s tally, neither of which knows the matrix exists.

So the question is not whether the pairing was ignored. It was, exactly, and by the amount a degree sequence differs from the graph it came from.

That identity is the one check this essay would be worthless without, and it is worth saying why it is checked twice rather than asserted once. The modelled matrix and the modelled per-unit share are built from the same product of clamped survival factors, so their agreement is arithmetic and a failure would mean a transcription slip. The counted pair tally and the counted per-unit share are not: one increments an entry of a matrix and the other increments two entries of a vector, over the same double loop, and the row mean of the first has to be the availability-weighted mean of the second. A tally that credited a unit once for an exchange it is half of would produce a matrix perfectly correlated with the right one, a degree sequence off by a factor of two, and no downstream number that looked wrong. The whole of this essay is a comparison between a matrix and its own row means, so an error that moved one and not the other is the error it is least equipped to notice.

The active set is a graph on the units. Every pair of the 14 units of one design, at the share of the exchanges between them that the tolerance box refuses — counted over all 116 admissible assignments. The darkest cells are at 1.0000 — pairs the box refuses every time they are available — and the diagonal is empty, because a unit is never exchanged with itself. The bars on the right are each unit's row mean, which is the one number per unit every earlier reading of this quantity uses; the 91 cells behind them are what that summary averages away. On this matrix 83.4% of the variation between pairs is not accounted for by the two units' own row means.
Fig. 1 The counted blocking matrix of one design: every pair of the fourteen units, at the share of the exchanges between them the box refuses. Seventy-three of the ninety-one pairs are refused every time they come up.

How much of it a per-unit summary discards

The degree is a summary, and a summary can be complete. If BijB_{ij} were well approximated by di+djd_i + d_j up to a constant, then a probe built from the pairing would be a probe built from the degrees and there would be nothing here.

The measurement is the share of the off-diagonal variation the additive degree fit leaves. Over the designs the earlier essays are measured on it is 0.6902 ± 0.0066 for the modelled matrix and 0.7540 ± 0.0099 for the counted one. More than two-thirds of the object, and on the exact version more than three-quarters, is outside the summary of it every probe so far has used.

A second reading says the same thing in a form a degree sequence cannot express at all. Weight each pair by how blocked it is and correlate the two endpoints’ degrees: the exchange graph is disassortative, at −0.1484 ± 0.0042 modelled and −0.1701 ± 0.0049 counted. A refused exchange joins a heavily-blocked unit to a lightly-blocked one more often than it joins two heavily-blocked ones. Both numbers are more than thirty standard errors from zero, so this is not a rounding artefact of a nearly-saturated matrix; it is an arrangement.

That is worth pausing on, because it is the first thing in this line of essays that could not have been said before. Every quantity the earlier probes report is a function of a column of length fourteen. Assortativity is a function of ninety-one numbers and is invariant to nothing the column can change. The pairing is real structure, it is most of the object, and until now none of it had been read.

What a number per unit throws away. Two readings of the blocking matrix that a vector of per-unit shares cannot carry, over 192 designs of 14 units. The first pair is the share of the variation between pairs that the two units' own degrees do not account for: 0.6902 ± 0.0066 modelled and 0.7540 ± 0.0099 counted, so more than two-thirds of the object is outside the summary of it the earlier readings use. The second pair is whether a refused exchange joins two units of like degree: -0.1484 ± 0.0042 and -0.1701 ± 0.0049, both negative, so the refusals spread across the units rather than concentrating among the units that are refused most.
Fig. 2 Two readings a vector cannot carry: how much of the pair structure the units’ own degrees fail to account for, and whether a refused exchange joins two units of like degree.

Two probes a graph has and a vector does not

A probe here is a column of length nn whose signed imbalance the diagnostic reads, so the pairing has to be turned back into a column before it can compete. There are two obvious ways and they are not the same.

The direction. Subtract the additive degree fit, leaving Rij=Bijdidj+BˉR_{ij} = B_{ij} - d_i - d_j + \bar{B} off the diagonal and zero on it, and take the dominant eigenvector of RR. That is the leading direction of exactly what the degree throws away, and it is orthogonal to the degrees by construction rather than by luck.

The cut. Take the signs of that eigenvector. This is coarser — it keeps which side of a partition each unit is on and discards how far — and it is the natural summary of a graph rather than of a matrix.

Both are computed twice, once from the modelled blocking matrix, which needs only the design and the tolerance, and once from the counted one, which needs the enumeration and so is available to nobody running a real trial. Both are then standardised and projected off the rule’s own span, exactly as every feasible probe in this line is, because a probe inside the span is a probe the rule has already answered. Four new rows, on the same designs and the same seeds as the eight already on the table, through the same table-building function — so the eight old rows come out identical to the last bit, and a difference between the two tables would mean the construction had moved rather than that the probes had.

Counted, the pairing recovers most of what was lost

The counted per-unit share aligns with the separating direction at 0.3854. The counted pairing’s own direction reads 0.5266 and the cut it induces 0.5258.

Paired on the design, so that no design-to-design variation is carried twice, that is a gain of 0.1326 ± 0.0388 at 3.42 standard errors for the direction and 0.1404 at 3.49 for the cut. On the separation each probe carries — how far apart the two components of the admissible set sit on it, over the spread inside a component — the same change runs from 2.0170 to 3.5949 and 4.2786.

So the complaint was right about the object. Reading the active set as a graph rather than as a vector recovers something real, and it recovers most of the gap between the per-unit share and the heuristic that beat it.

Most, and not all. The design’s own leverage reads 0.5395. Against it the counted pairing is −0.0129 ± 0.0354, at −0.37 paired standard errors, and the cut −0.0101 at −0.24. Neither of those separates anything. The honest statement is that the pairing draws level with leverage — not that it ties it, which would be a demonstration, and not that it loses, which the sweep cannot show either.

The pairing recovers most of it and passes nothing. How much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split in two. The four rows the pairing adds are the dominant direction of what each blocking matrix keeps past its degrees, and the cut that direction's signs induce. Counted, they read 0.5266 and 0.5258 against the counted per-unit share's 0.3854 — most of the gap between that share and the design's own leverage at 0.5395, closed. Modelled, they read 0.4274 and 0.4954 against 0.4272. Nothing built from the active set passes leverage, and the projected fourth power is still ahead of all of them at 0.6583.
Fig. 3 The twelve-row probe table. The four rows the pairing adds sit between the per-unit shares they are built from and the design’s own leverage, and the projected fourth power is still ahead of all of them.

And it does not pass the probe a trial can already build

Against the projected fourth power — the best feasible probe the field that opened this question ever found — the counted pairing is −0.1317 ± 0.0330, at −3.99 paired standard errors. That is a loss and it is decisive.

Which settles the practical question the whole line of essays exists to answer. Whatever a trial should run its diagnostic in, it is not a summary of the constraint’s active set: not the per-unit share, modelled or counted, and not the pairing either. The best feasible probe on a table of twelve is the one it was on a table of five, before any of this was proposed.

The reason this is worth saying carefully is that the earlier finding was stronger than the evidence now supports. The essay that first priced the active set against leverage concluded that what the rule blocks is not what the diagnostic needs to know. That reading survives at the level of the recommendation and does not survive as a statement about the quantity: a probe that recovers 0.1326 of alignment when the pairing is kept is a probe whose quantity was carrying something the summary destroyed.

What each probe can see, with the pairing on the table. How far apart the two components of the admissible set are on each probe, over the spread inside a component, on 100 designs whose set is enumerated and split. The four rows the pairing adds run 3.1025 and 3.7210 modelled and 3.5949 and 4.2786 counted, against the per-unit shares they are built from at 1.9529 and 2.0170. So every one of them roughly doubles what the vector could see. The design's own leverage reads 3.8362, the projected fourth power 5.0800 and the separating direction itself 10.5646.
Fig. 4 The same twelve probes on the separation between the admissible set’s two components, which is the population quantity a chain is trying to report rather than a correlation with an answer nobody has.

Modelled, the direction recovers nothing at all

The counted matrix needs an enumeration, so the reading above is a statement about what a trial gives up rather than about what it could do. The modelled matrix needs only the design and the tolerance. What does the pairing buy there?

The modelled per-unit share aligns at 0.4272. The modelled pairing’s own direction aligns at 0.4274 — a paired gain of 0.0003 at 0.01 standard errors, which is not a small effect but the absence of one, on a matrix two-thirds of which lies outside the degree.

The cut is different. It reads 0.4954, a gain of 0.0793 at 1.91 standard errors: nominal rather than decisive, and pointing the right way. Against leverage the modelled cut is −0.0387 at −0.86, which again separates nothing.

So the four new rows split, and they split along the wrong axis. It is not that counted beats modelled everywhere: the counted cut beats the modelled cut by 0.0354 at 0.78 paired standard errors, which is nothing, while the counted direction beats the modelled direction by 0.0992 at 3.41, which is everything. The coarse reading of the pairing survives the approximation. The fine one does not.

What the pairing buys, and where it stops. Paired differences in alignment with the separating direction, design by design over the 100 designs whose admissible set is split — so each bar carries no design-to-design variation. Reading the counted active set as a graph rather than as a vector is worth 0.1326 at 3.42 standard errors; reading the modelled one that way is worth 0.0003 at 0.01, which is nothing. Against the design's own leverage the counted pairing is -0.0129 at -0.37 — a failure to separate rather than a win — and against the projected fourth power it is -0.1317 at -3.99, which is a loss.
Fig. 5 The paired differences the argument turns on, design by design. The pairing gains on the vector it replaces, gains nothing on the modelled side unless it is read as a cut, and loses to the projected column.

The obvious explanation, checked and refused

There is a ready account of why a modelled pairing should be worth less than a counted one, and it was checked before anything was concluded from it, because it is the account the previous essay in this field would have reached for. The model is measurably wrong — it assumes an admissible assignment sits uniformly inside its tolerance box, which it does not — so presumably it gets the pairing wrong.

It does not. The two blocking matrices agree entry for entry at 0.8041 ± 0.0162. Their degrees agree at 0.7742 ± 0.0223. And their residuals past the degree — the pair structure itself, with everything the two already agree about removed — agree at 0.8219 ± 0.0149, which is higher than the agreement on the whole matrix rather than lower.

The uniform-position model is not missing the pairing. Whatever separates the two directions is finer than the pair structure is.

What it is can be measured. The two residual eigenvectors put 0.8274 ± 0.0109 of the units on the same side of the cut — eleven and a half of fourteen, read up to a global flip, since a cut and its complement are one cut. They agree about the partition and disagree about the magnitudes, and an eigenvector is a fine functional of a matrix while a sign pattern is a coarse one. Once the magnitudes are discarded the two probes are nearly the same object, which is why the cut survives the approximation and the direction does not.

That is the mechanism as far as this essay can establish it, and it is worth being clear that it is a description rather than a derivation. Read as columns, after the projection and the standardisation, the two directions agree at 0.7129 ± 0.0193 and the two cuts at 0.5778 ± 0.0255 — which is the wrong way round if column agreement were what decided performance. The alignment a probe delivers is not a smooth function of the column, and nothing here says by how much it fails to be.

The model is not missing the pairing. How much of the counted blocking matrix its uniform-position model recovers, read six ways over 192 designs. Entry for entry the two agree at 0.8041 ± 0.0162; strip out the degrees they both agree about at 0.7742 and the agreement between what is left is 0.8219 ± 0.0149, which is higher rather than lower. So the obvious explanation of why a modelled pairing is a worse probe than a counted one — that the model gets the pairing wrong — is not the explanation. What the two disagree about is finer: as probes their leading directions agree at 0.7129 and the cuts those directions induce at 0.5778, while 82.7% of units are put on the same side of the cut by both.
Fig. 6 Six readings of how much of the counted pairing its uniform-position model recovers. The agreement rises rather than falls when the degrees both matrices agree about are stripped out.

What a chain of eight hundred draws makes of it

The alignment and the separation are population quantities computed on an enumerated set. A practitioner has a chain, and a chain of finite length is what decides whether a split set is reported as split.

On that reading the counted pairing does slightly better than the enumeration suggests. It misses 17.6% of the split sets against the counted per-unit share’s 39.4% and the design’s own leverage at 20.6%, with its cut at 18.2%. The projected fourth power still misses least of anything feasible, at 11.8%, and the separating direction itself — which no trial has, because having it means already knowing the answer — misses 8.8%.

The gap between the pairing and leverage there is three points on thirty-four designs, which is one design. It is reported because leaving out the reading that happens to favour the new probe would be worse than reporting a number too small to lean on, and it is not leaned on.

What the chain does confirm is the ordering that matters. The pairing is a large improvement on the vector it replaces — twenty-two points of missed sets — and it is not an improvement on the probe that was already available.

What a chain of eight hundred draws misses. The share of admissible sets that are known to be split on which a two-chain diagnostic of 800 draws stays quiet, by probe, over 34 designs. The counted pairing misses 17.6% and the cut it induces 18.2%, against the counted per-unit share's 39.4% and the design's own leverage at 20.6% — so on the reading a practitioner would experience the pairing is ahead of the vector it replaces and nominally ahead of leverage, on a difference of one design. The projected fourth power still misses least of anything feasible, at 11.8%, against the separating direction itself at 8.8%.
Fig. 7 What a two-chain diagnostic of eight hundred draws misses with each probe, which is the reading a practitioner would experience rather than the one the enumeration supports.

What this changes about the field’s conclusion

Three sentences, and the third is the one that moved.

The recommendation is unchanged. A trial running a randomisation diagnostic should run it in the projected fourth power, not in anything built from what its balancing rule refuses. That was true against a table of five probes, against a table of eight, and it is true against a table of twelve.

The exactness is still not what was costing the probe. Counting the active set instead of modelling it gains nothing on the per-unit reading and gains on the pair reading only because the model’s magnitudes are wrong; both facts sit inside a comparison that leverage wins either way.

But the quantity was carrying something, and the summary was destroying it. The earlier reading — that what the rule blocks is simply not related to where the set splits — is too strong. What the rule blocks, read as the set of pairs it is, is about as informative as the design’s own leverage. It is a summary of it, taken by averaging over the partner, that reads as barely better than noise.

That distinction has a use beyond this field, and it is the transferable half. A quantity that underperforms as a probe has three explanations rather than two: the quantity is wrong, the approximation to it is wrong, or the summary of it is wrong. The first two are the ones anybody tests, because they are the ones a modelling argument suggests. The third is invisible, because the summary usually arrives with the quantity and is never separately named — here it arrived as an average over jj inside a loop, and it cost a factor of a seventh in alignment and twenty-two points of detection.

What this opens and does not measure

Three things, each named where it could have been run.

A cut is a partition, and a partition of the units is not the only one available. The cut used here comes from the signs of one eigenvector of one residual matrix. The admissible set falls into two connected components, and those are a partition of assignments rather than of units; whether the pairing’s cut and the set’s own split are related as partitions is a question about two objects of different types and is untouched. It would need a way of pushing a partition of assignments down onto the units, and the obvious ones are exactly the per-unit averages this essay is about.

The second eigenvector is not read. The residual matrix has fourteen of them and only the dominant one is used, on the reasoning that a probe is one column. A diagnostic that read two columns is a different test — the field that walks these sets already has statistics that take more than one — and nothing here says whether the second direction carries anything.

The blocking matrix is a weighted graph and is read here as an unweighted one, twice over. Both the eigenvector and the cut treat every entry of the residual on the same footing, where the entries themselves run from nought to one and on one design seventy-three of ninety-one sit at the ceiling. A saturated pair carries no information about the design and a great deal of weight in the eigenvector, and a probe that dropped the saturated pairs before taking the direction is a different probe. Nothing here says whether it is a better one, and it is the cheapest of the three things left.

And the whole thing is measured at fourteen units. That is where the admissible set can be enumerated, and the enumeration stops not far past it. The pair matrix is n(n1)/2n(n-1)/2 numbers against a degree sequence’s nn, so the share of the object a per-unit summary discards should grow with the trial size — and every number above is measured at the one size where the counted version exists at all. Whether the pairing’s advantage over the vector grows, holds or fades at forty units is untested, and the modelled matrix, which needs no enumeration, is the one instrument that could answer it.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • What the rule blocks — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, imbalance, leverage, randomisation test, rerandomisation
  • A probe nobody chose — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, projection, randomisation test, rerandomisation
  • A test rather than a survey — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, randomisation test, rerandomisation
  • Before the trial and after — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, randomisation test, rerandomisation
  • The statistic the p-value is about — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, randomisation test, rerandomisation
  • When the constraints run out — both name assignment mechanism, covariate balance, exact enumeration, experimental design, imbalance, randomisation test, rerandomisation

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismAssortativityConnected componentCovariate balanceEigenvectorExact enumerationExchange graphExperimental designGraph cutImbalanceLeverageModel diagnosticsProjectionRandomisation testRerandomisation