A probe from what the rule blocks

How far, not whether

The blocking matrix records whether a tolerance box refuses an exchange. Dropping the pairs it refuses for certain changes nothing. Recording instead how far each exchange moves the balance vector turns the pairing into its degree plus a part that lies inside the rule's own span — and the degree is the design's leverage, exactly. The heuristic that beat every reading of what the rule blocks was the unclamped pairing's degree all along.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

A set of pairs, not a vector read the tolerance box’s blocking matrix as the graph it is, rather than as the per-unit share every earlier probe had used, and found that the share was the graph’s degree and discarded two thirds of it. The graph’s own direction recovered part of what the degree lost and drew level with the design’s leverage rather than passing it. It named three things it had not run, and called one of them the cheapest.

That one was about saturation. The modelled blocking matrix is one minus a product of clamped survival factors, so any exchange that moves a single coordinate of the balance vector past twice the tolerance is refused with certainty and its entry is exactly one. A saturated pair, the essay argued, carries the matrix’s largest value and no information about how far outside the box the exchange lands; a probe that dropped those pairs before taking the direction is a different probe, and nothing said whether it was better.

It is not. But asking why leads to the reading the field has been missing since the essay where leverage first won, and it changes what that win meant.

Dropping the certain refusals

The setting is the one every essay on this question has used: designs of fourteen units, four balance columns, a tolerance of 0.8, and the hundred designs out of two hundred whose admissible set is in two pieces, so that there is a separating direction to find. Every probe is projected off the rule’s span and scored two ways — its alignment with the separating direction, and the median separation it carries between the two pieces.

How many pairs saturate depends on which matrix is meant.

How many refusals are certain. Over 192 designs of fourteen units, the share of the ninety-one pairs whose entry in the blocking matrix is at the ceiling. In the modelled matrix — one minus a product of clamped survival factors, which reaches one whenever a single coordinate of the swap step exceeds twice the tolerance — 9.3% of pairs are saturated, between 2.2% and 22.0% on a single design. In the counted matrix, where an entry is one when the exchange is refused on every admissible assignment that offers it, 57.3% are. Of the pairs the model puts at the ceiling, 96.4% are refused every time in the count as well.
Fig. 1 The share of the ninety-one pairs at the ceiling in the modelled and the counted blocking matrix, and how often the count agrees with the model’s certain refusals.

In the modelled matrix 9.3% of pairs sit at the ceiling on average, between 2.2% and 22.0% on a single design. In the counted matrix — where an entry is one when the exchange is refused on every admissible assignment that offers it — 57.3% do. Of the pairs the model puts at the ceiling, the count refuses 96.4% every time as well, so the model’s certain refusals are nearly always real ones; it simply finds fewer of them than the enumeration does.

Dropping the modelled matrix’s saturated pairs — setting their entries in the residual past the degree to zero, and taking the leading direction of what is left — leaves the probe almost where it was. Its alignment with the separating direction is 0.4302 against the undropped pairing’s 0.4274, a difference of 0.0028 at 0.41 paired standard errors, and its median separation is 2.917 against 3.103. Neither is a change. The saturated pairs were a small share of the modelled matrix, and the direction did not depend on them.

How far, not whether

The objection to a saturated entry was that it records whether an exchange is refused and discards how far outside the box it lands. The direct repair is to record the second thing instead. For every pair, take the swap step — the change in the balance vector that exchanging the two units would cause — and measure its squared length in units of twice the tolerance. Call that the pair’s depth. It has no ceiling: an exchange that lands just outside the box and one that lands three widths outside are refused alike, and their depths differ by a factor of nine.

Separation, how far against whether. The median separation each probe carries between the pieces of a split admissible set, over the 100 designs whose set is split. The share refused per unit carries 1.953; the refusal pairing's direction 3.103, and with the saturated pairs dropped 2.917. The depth pairing's direction carries 2.345, and each unit's mean depth 3.836 — exactly the leverage's. A random direction carries 0.942 and the projected fourth power 5.080.
Fig. 2 The median separation each probe carries between the pieces of a split admissible set: the per-unit refusal share, the refusal pairing’s direction with and without its saturated pairs, the depth pairing’s direction and each unit’s mean depth, beside a random direction and the projected fourth power.

The depth pairing’s direction, read the way the refusal pairing’s was, carries a median separation of 2.345 — less than the refusal pairing’s. But the depth matrix’s degree, each unit’s mean depth over its partners, carries 3.836. That is more than any feasible probe built from the box’s refusals carries on this table, and it is not a near miss of the leverage’s number. It is the leverage’s number. The alignment tells the same story to every digit: the mean depth aligns 0.5395 and so does the leverage.

Alignment with the split, how far against whether. The mean alignment of each probe with the direction that separates the admissible set, over the 100 designs whose set is split. The refusal pairing's direction aligns 0.4274, and with its saturated pairs dropped 0.4302. The depth pairing's direction aligns 0.4293; each unit's mean depth 0.5395, which is the leverage's 0.5395 to every digit. The share refused per unit aligns 0.4272, a random direction 0.2622, and the projected fourth power 0.6583.
Fig. 3 The same probes by their mean alignment with the separating direction.

Why the mean depth is the leverage

That is an identity, and it is short. The balance columns are centred and orthogonal, each with squared length nn. The swap step of units ii and jj is a constant times the difference of their rows, xi−xjx_i - x_j, so a pair’s depth is a constant times

∥xi−xj∥2=hi+hj−2 xi⋅xj,\lVert x_i - x_j \rVert^2 = h_i + h_j - 2\,x_i \cdot x_j,

where hi=∥xi∥2h_i = \lVert x_i \rVert^2 is unit ii’s leverage on the rule’s columns. Averaging over the partner jj, the centring makes the mean of xjx_j over the other units −xi/(n−1)-x_i/(n-1), and the mean of hjh_j over them (np−hi)/(n−1)(np - h_i)/(n-1) with pp the number of columns. Put together, a unit’s mean depth is

nn−1 hi+npn−1\frac{n}{n-1}\,h_i + \frac{np}{n-1}

times the step’s constant — an affine function of the leverage, with nothing else in it. Every probe here is projected off the span and standardised, which removes the constant and the scale, so the mean-depth probe and the leverage probe are the same column.

A unit's mean depth is its leverage. For each of the 14 units of one design, the mean over its partners of how far an exchange with that partner moves the balance vector — the squared length of the swap step, in units of twice the tolerance — against the unit's leverage, the diagonal of the hat matrix of the rule's columns. Every point lies on the line the algebra predicts, with slope n/(n − 1) times the step's scale: with centred orthogonal columns the squared distance between two units is the sum of their leverages minus twice their inner product, and averaging over the partner removes the inner product up to a term that is itself proportional to the leverage. The largest gap between a point and the line is 1.1e-16.
Fig. 4 For each unit of one design, its mean depth against its leverage, with the line the algebra predicts.

On a single design every unit lies on the predicted line, and the largest gap between a point and the line is 1.1e-16, which is the arithmetic’s rounding. The leverages on that design run from 1.913 to 11.867 and sum to 56, which is fourteen units times four columns — as they must, since each column has squared length equal to the number of units and the leverages are the columns’ squares summed across. The minimum correlation between mean depth and leverage across all two hundred designs is one to fifteen decimal places.

What leverage was all along

That reframes four essays. The essay where leverage first won called it a heuristic about which units a balancing rule has most to say about, and set it against the constraint’s active set, the thing the rule actually does. Counting the active set exactly made it worse. Reading it as a graph drew level. In every round the heuristic was the benchmark the quantity derived from the rule could not beat.

The identity says the two were never different kinds of thing. Leverage is the degree of a pairing derived from the rule — how far each exchange moves the balance the rule constrains. The refusal share is the degree of the same swap steps passed through the box: one minus a product of clamped survival factors, a function of each coordinate of the step that is zero inside one tolerance-width, ramps, and is certain beyond two. So the comparison every round made was between one pairing’s degree read linearly and the same pairing’s degree read through a clamp. The linear reading won every time, by 0.1124 in alignment at 4.43 paired standard errors over the refusal share, and the clamp is the whole of the difference.

That is a sharper statement than “the heuristic wins”, and a more useful one — but it is worth being exact about what the clamp loses, because the obvious guess is wrong.

The clamp keeps the order and loses the spacing

The obvious guess is that a clamp throws away the ordering: a unit whose exchanges all land far outside the box and one whose exchanges land just outside it are refused alike. On these designs that hardly happens. The rank correlation between each unit’s refusal share and its leverage averages 0.909 across the two hundred designs, and the two put the same unit first on 98% of them. The clamp leaves the order of the units nearly intact.

What it changes is the spacing. Leverage on one of these designs runs from 1.913 to 11.867 — the most influential unit carries six times the least — and the refusal share is a probability, so it maps that range into the unit interval, through clamps set by the tolerance rather than by the design. The order survives; the distances between units do not.

And after projection off the span, distances are all a probe has. Every probe here is a column of fourteen numbers with the rule’s own columns regressed out, and the regression is linear: it removes whatever part of the column the balance columns can reproduce, and the part that survives depends on the column’s values, not its ranks. A monotone reshaping of leverage — which is what the refusal share nearly is — leaves a different residual from leverage itself, and on these designs a worse one. The difference in alignment, 0.1124, is the cost of reshaping a quantity that was already in the right units.

That is the general lesson, and it applies to more than this box. A probe that has to survive a linear projection should be built on a scale on which differences mean something. The depth is such a scale: it is a squared length in the rule’s own coordinates. The refusal share is not: it is a probability under a model of where the assignment sits, and its spacing reflects the model’s clamps rather than the design.

What only the count can see

One kind of probe built from the box does carry more than leverage, and it is worth placing. The counted pairing’s cut — the partition the leading direction of the enumerated refusal matrix induces — carries a median separation of 4.279, and the counted pairing’s direction 3.595, against the leverage’s 3.836. Their alignments, 0.5258 and 0.5266, sit just below the leverage’s 0.5395.

Both are infeasible. They are built by enumerating the admissible set, which is the thing the diagnostic exists because nobody can do at a realistic trial size. And the identity says why they can carry something leverage cannot: a counted refusal is a fact about the admissible set itself — which assignments the box actually admits — and not a function of the design’s geometry. Leverage, the depth and the modelled refusal share are all functions of the balance columns alone. Every feasible probe derived from the box is therefore a reading of one object, the rows of the design in the rule’s coordinates, and its best reading is the linear one. To beat it a probe has to know something the design’s rows do not say, and the only thing on the table that does is the enumeration.

That bounds the programme the last five essays have pursued. Every feasible probe derived from what the rule blocks is a transformation of the design’s rows; the best of them is the leverage, exactly; and the projected fourth power, which beats it, is not a transformation of the rule’s blocking at all but of the covariate’s own shape.

Where the depth matrix keeps what it knows

The depth pairing’s own direction carried less than its degree, which is the opposite of what the refusal pairing did. The anatomy of the two matrices says why.

Where each matrix keeps what it knows. How the off-diagonal variation of two matrices on the same pairs divides, averaged over 200 designs. The depth matrix — how far each exchange moves the balance vector — keeps 71.6% of it in its degree, which is the leverage exactly, and of the 28.4% past the degree, 98.7% is the Gram matrix of the units inside the rule's own span, which every probe is projected off. Almost nothing is left for a direction past the span to find. The refusal matrix, the same pairs clamped at the ceiling, keeps only 31.2% in its degree and 68.8% outside it.
Fig. 5 How the off-diagonal variation of the depth matrix and of the refusal matrix divides between the degree, the rule’s own span and the rest.

The depth matrix keeps 71.6% of its off-diagonal variation in its degree. Of the 28.4% past the degree, 98.7% is the Gram matrix of the units inside the rule’s own span — the −2 xi⋅xj-2\,x_i \cdot x_j term of the identity, centred — and every probe is projected off that span, because a probe inside it is one the rule has already answered. What is left for a direction past the span to find is a sliver, and the depth pairing’s direction, once projected, is whatever that sliver happens to point at.

The refusal matrix is the other way round: 31.2% in its degree and 68.8% outside it. The clamp has moved most of the structure out of the degree and out of the span. That is why its direction was worth reading where the depth matrix’s is not — and also why it could only draw level with leverage. What the clamp moves out of the degree is the ordering the degree needed.

Reading the earlier essays again

The identity is a lens on the whole line of essays, and several of its findings read differently through it.

What the rule blocks proposed the refusal share as a probe on the ground that which exchanges a balancing rule refuses is computable from the design and the tolerance before any assignment exists. That is true, and the identity adds the qualification: the share is computable from the design because it is a clamped function of the same swap steps whose unclamped mean is the leverage. It was never a second source of information about the design; it was a second scale for the first.

A model and a count found that the modelled and counted refusal shares order the units alike, at a correlation of 0.81, and disagree about the level. Through the identity that is expected: the modelled share is a reshaping of leverage, the counted share carries what the enumeration adds, and the agreement between them is the part both inherit from the design’s rows. The disagreement is where the enumeration knows something the rows do not, and the counted cut above is that knowledge put to use.

And the essay that chose a probe from the design ranked leverage third of six, behind the projected fourth power and the separating direction itself, on the reasoning that leverage measures how much of the rule’s span a unit carries. The identity gives the same quantity a second meaning that is closer to what the diagnostic needs: leverage is how far, on average, exchanging that unit with another moves the balance the rule constrains. A unit whose exchanges move the balance a long way is a unit the rule cannot let move freely, and the separations measured here say the admissible set tends to split along such units. That is the mechanism the leverage heuristic was assumed to have, and it is now an identity rather than an assumption.

None of this changes a number those essays reported. It changes what the numbers are about: four essays of comparison between a derived quantity and a heuristic turn out to have been comparisons between two scales of one derived quantity, and the scale the rule itself uses — the box’s clamps — was the worse one for a diagnostic every time.

What this changes about the field’s conclusion

The recommendation stands. The projected fourth power carries 5.080 and no feasible probe built from the rule’s blocking carries more than 3.836. A trial running the diagnostic should still run it in the projected fourth power.

The heuristic is now a derivation. Leverage is the mean squared exchange step in the rule’s own coordinates, with an exact identity behind it, so it no longer needs defending as a guess about which units matter. It is the pairing the rule induces, summarised the one way that summary loses nothing past the span.

The summary lesson has a boundary. The essay on pairs found that a per-unit summary of the refusal matrix destroyed information its graph still carried. For the depth matrix the summary destroys almost nothing that survives the projection: the degree is the leverage, the rest is inside the span. Whether a summary loses depends on where the matrix keeps what it knows, and the clamp is what moved it.

Still open: a clamp chosen for the probe

The box’s clamp is a fixed function of the depth: zero below one tolerance-width, a ramp, certainty above two. It was chosen by the rule, not by the diagnostic. A probe could pass the depth through a different function — a smooth one that keeps the ordering within the refused pairs, or a steeper one that sharpens it near the edge of the box — and read the degree of the result. At one extreme it is leverage and at the other it is the refusal share; the question is whether some function between them carries more separation than either. The enumeration stops not far past fourteen units, so the measurement has to be made here, on these designs, and then carried to the larger trials where the modelled matrices — which need no enumeration — are the only instrument; the identity above says what the linear end of that family already delivers.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A probe nobody chose — both name assignment mechanism, covariate balance, exact enumeration, imbalance, projection, randomisation test, rerandomisation
  • The part the rule already took — both name assignment mechanism, covariate balance, exact enumeration, experimental design, leverage, projection, randomisation test
  • What a chosen probe finds — both name assignment mechanism, covariate balance, exact enumeration, experimental design, leverage, projection, randomisation test
  • When the constraints run out — both name assignment mechanism, covariate balance, exact enumeration, experimental design, imbalance, randomisation test, rerandomisation
  • A defect that is about size — both name assignment mechanism, covariate balance, exact enumeration, imbalance, randomisation test, rerandomisation
  • A test rather than a survey — both name assignment mechanism, covariate balance, exact enumeration, imbalance, randomisation test, rerandomisation

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismClosed formCovariate balanceEigenvectorExact enumerationExchange graphExperimental designImbalanceLeverageProjectionRandomisation testRerandomisation