A probe from what the rule blocks

A quantity that loses to a heuristic

Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

The argument that opened this field is a good one, and it is worth stating at its strongest before it fails.

A balancing rule breaks an admissible set into pieces by refusing exchanges. The pieces are what a diagnostic exists to detect. So the quantity a diagnostic should be aimed at is which exchanges the rule refuses — the constraint’s active set — and not the design’s leverage, which is a summary of how much of the rule’s span each unit carries and is related to the active set only by intuition.

The earlier field says exactly that about its own best feasible probe: leverage is a mechanism rather than a theorem, and the measurement is what decides whether it is right.

The measurement decides against the active set.

The table

Over a hundred designs of fourteen units whose admissible set is enumerated and split, the alignment each probe carries with the separating direction:

probe alignment separation
the separating direction itself 1.0000 10.5646
the fourth power, projected off the span 0.6583 5.0800
the design’s own leverage 0.5395 3.8362
the active set, modelled 0.4272 1.9529
the active set, counted 0.3854 2.0170
a random direction in the same subspace 0.2622 0.9422
the most concentrated direction 0.2001 0.5431
the fourth power, raw 0.1720 1.5427

The active set beats a random direction by 0.1650 of alignment at 4.94 paired standard errors, and it loses to leverage by 0.1124 at 4.43.

What the rule blocks is not where it splitsHow much of the separating direction each probe carries, over 100 designs of 14 units whose admissible set is enumerated and split into two pieces. The deferral this field answers proposed the constraint's active set — which exchanges the tolerance box actually blocks — as a better probe than the design's own leverage, on the ground that leverage is a heuristic and the active set is the quantity. Modelled from the design and the tolerance, it reads 0.4272 against leverage's 0.5395, at 4.43 paired standard errors the wrong way. Counted exactly over the enumerated set — at a cost no trial can pay — it reads 0.3854, worse again. Both beat a random direction at 0.2622, so they are probes; neither beats the two the earlier field already had.the fourth power, raw0.1720± 0.0154the fourth power, projected0.6583± 0.0279the design's own leverage0.5395± 0.0298the most concentrated direction0.2001± 0.0150a random direction0.2622± 0.0174the active set, modelled0.4272± 0.0280the active set, counted0.3854± 0.0272the separating direction1.0000± 0.0000100 split designs of 14 units, 50.0% of those drawnone is the answer itself
Fig. 1 How much of the separating direction each probe carries. The slider changes how many designs are read.

What a paired standard error is measuring here

The comparisons above are all paired, and the pairing is what makes a difference of 0.11 in a quantity between 0 and 1 readable.

Each design contributes one alignment per probe. Designs vary enormously: the separating direction is at 1 by construction on every one, but the projected fourth power ranges over most of the unit interval across designs, with a mean of 0.6583 and a median of 0.7579. An unpaired comparison of two probes’ means would be comparing two averages each carrying that whole spread.

Paired, the comparison is design by design: on this design, does the active-set column align better than the leverage column? The differences are much less variable than the levels, because a design on which one probe does well is a design on which the other tends to as well.

That is why 0.4272 against 0.5395 comes out at 4.43 standard errors rather than at one or two, and it is why the one comparison that does not separate — the counted active set against the modelled one, at −0.83 — can be reported as a genuine non-difference rather than as a sweep that was too short.

It is a probe

That is worth establishing before anything else, because a column that lost to a random direction would make the rest of the field uninteresting.

It does not. Both active-set columns beat a random direction in the same subspace on the alignment and on the separation, at 4.94 and 3.64 paired standard errors on the first and 4.92 and 2.97 on the second. Both also beat the raw fourth power — the column the earlier fields used before any of this — by a wide margin.

So the field’s finding is not that a column of nothing lost to a heuristic. It is that a real probe, built from the right quantity, is beaten by a heuristic about that quantity.

And it loses

By 4.43 paired standard errors to leverage, and by 7.16 to the projected fourth power.

Both comparisons are paired on the design, which matters here more than usual: designs differ enormously in how sharply their admissible sets split, so an unpaired comparison of two probes would carry the design-to-design variation twice and would separate almost nothing. Paired, the differences are small and clean.

The separation column says the same thing in the units a chain reports: 1.9529 for the modelled active set against 3.8362 for leverage and 5.0800 for the projected fourth power, on an oracle of 10.5646.

Two readings of the same table, and why they differ

The alignment column and the separation column are not the same measurement, and the gap between them for one probe is worth reading.

Alignment is the absolute correlation between a probe and the separating direction, so it is a property of the two columns and of nothing else. It runs from 0 to 1 and it is what predicts whether a chain can see anything.

Separation is the distance between the two components’ mean probe values over the spread inside a component, enumerated. It is the population quantity a chain is estimating, it has no upper bound, and it is reported as a median over designs because its mean is dominated by a tail — the most concentrated direction reads a median separation of 0.5431 and a mean of 65.7323, which is one design with a probe that is nearly constant inside a component.

The two mostly agree. Where they part is instructive: the raw fourth power aligns at 0.1720, below a random direction’s 0.2622, and carries a separation of 1.5427, above the random direction’s 0.9422. A column can be poorly aligned and still carry separation if what alignment it has is concentrated where the spread inside a component is small.

The active-set columns sit the same way round on both readings, so nothing in this field turns on the choice. It is reported because a table with two columns that usually agree and occasionally do not is more honest than a table with one.

What a practitioner should take from the table

Short, because the table’s practical content is one line.

Run the projected fourth power. It is the best feasible probe here, it was the best in the earlier field, and adding two probes built from better arguments has not displaced it. It costs one Gram–Schmidt pass against the rule’s own columns.

Leverage is a reasonable second, and it has the advantage of needing no dictionary — it is the diagonal of the hat matrix of the basis the rule already balanced, so it exists on any design whatever functions were chosen.

And the active set is not worth building. It costs n² products, it needs the tolerance as well as the design, and it carries less than either. That is the field’s recommendation and it is a negative one.

What a short chain says

The separation is what a probe could see given an unlimited chain. It is worth checking that a chain a trial would actually run reports the same ordering, because a finding that held on the enumerated quantity and reversed on the verdict would be an artefact of the enumeration.

Over 34 split designs, with chains of eight hundred draws, the share of split sets each probe misses:

the separating direction 8.8%, the projected fourth power 11.8%, leverage 20.6%, the modelled active set 29.4%, the counted active set 39.4%, and the raw fourth power 44.1%.

Same order. A trial using the active-set probe would miss three of every ten broken sets where one using leverage misses two.

A short chain agrees with the enumeration. The share of admissible sets that are known to be split on which a two-chain diagnostic of 800 draws stays quiet, by probe, over 34 designs. The separation is what a probe could see given an unlimited chain; this is what a chain a trial would actually run returns, and it ranks the probes the same way. The projected fourth power misses 11.8%, the design's own leverage 20.6%, the modelled active set 29.4% and the counted one 39.4%, against the separating direction itself at 8.8% and the raw column the earlier fields used at 44.1%. So the two probes built from what the constraint actually blocks are better than doing nothing and worse than the two the earlier field already had.
Fig. 2 What a chain of eight hundred draws misses with each probe, over the designs known to be split.

Why, as far as this can say

The mechanism is available and it is not flattering to the active set.

Projected off the rule’s span and standardised, the modelled active-set column agrees with the leverage column at |r| = 0.8359 ± 0.0114. So it is largely leverage under another name: a unit whose exchanges are refused is a unit whose basis columns are large, and a unit whose basis columns are large is a unit with high leverage.

What it is not is the sixth of itself that is something else, and that sixth points away from the separating direction. A column that is leverage plus a small orthogonal component, and lands below leverage, has an orthogonal component that costs rather than pays.

That is as far as this can be established here. The next essay is what happens when the same question is asked of a column that is not mostly leverage.

One of them is mostly leverage. How much of the design's own leverage direction each active-set probe carries, once both are standardised and projected off the rule's span — which is what a probe is, so it is the comparison that matters. Over 189 designs the modelled active set agrees with leverage at |r| = 0.8359 ± 0.0114 and the counted one at 0.4239 ± 0.0216. So the modelled probe is largely leverage under another name and the counted one is genuinely a different direction — and the counted one is the worse probe, at 0.3854 of alignment against 0.4272. What the active set contains beyond leverage points away from where the set splits.
Fig. 3 How much of the leverage direction each active-set column carries, which is the mechanism.

The designs this is measured on

Half the designs drawn are thrown away, and which half matters.

Of two hundred designs of fourteen units at a tolerance of 0.8, 50.0% have an admissible set that falls into two or more components. The other half have a connected set, on which there is no separating direction, no separation to measure, and nothing for a probe to be aligned with. Those are dropped rather than counted as zeroes, and the share that split is reported so the drop is visible.

That is a selection, and it is the right one for this question: the probes are being compared on their ability to detect a split, so the comparison has to be made where there is a split to detect. A probe’s behaviour on a connected set is a separate question — the earlier field’s own control is that the two-chain test stays quiet on connected sets, and it does.

It also means every number in this field is conditional on a design whose set is broken. A practitioner does not know which half they are in, which is exactly why they run the diagnostic, and the miss rates above are the numbers that matter to them: on the designs where something is wrong, this is how often each probe fails to say so.

Which column predicts the verdict

The alignment and separation columns are reported together on the grounds that they mostly agree, and the miss rates are a third measurement that decides between them where they do not.

Rank the six probes by alignment — 1.0000, 0.6583, 0.5395, 0.4272, 0.3854, 0.1720 — and the miss rates follow in exactly that order: 8.8%, 11.8%, 20.6%, 29.4%, 39.4%, 44.1%. Six probes, six ranks, no inversions.

Rank them by separation instead and there is one. The counted active set carries 2.0170 of separation against the modelled one’s 1.9529, so separation puts it ahead — and it misses 39.4% of split sets where the modelled column misses 29.4%.

So alignment is the column that predicts what a chain does and separation is not, which is a stronger statement than the essay’s “the two mostly agree”. The one place they disagree is the one place a practitioner would have been misled, and it is between two probes that differ in nothing except whether the active set was modelled or counted — the pair the field reports as a genuine non-difference at −0.83 standard errors on alignment.

That also says which of the two to report elsewhere. Separation is the population quantity and it has a tail that makes its mean useless — a median of 0.5431 against a mean of 65.7 for one column — while alignment is bounded, paired cleanly, and orders the verdicts. The unbounded quantity is the one with the interpretation and the bounded one is the one that works.

The mechanism is dilution, not misdirection

The active set correlates with leverage at 0.8359, which is 70% of its variance shared and 30% orthogonal, and the essay reads the orthogonal part as pointing away from the split. The arithmetic says something milder and more useful.

If the orthogonal 30% contributed nothing at all, the active set’s alignment would be 0.8359 × 0.5395 = 0.4510. It is 0.4272. So the orthogonal component’s own alignment with the separating direction is −0.043 — near enough zero, and nowhere near enough to be called pointing away.

The active set is leverage with three tenths of its length aimed at nothing in particular, and a probe with three tenths of its length wasted retains 0.792 of leverage’s alignment, against the 0.836 the correlation alone predicts. The whole gap between the two columns is dilution.

That is a better finding than misdirection would have been, because it says what a repair would have to do. There is nothing to subtract off — no direction the active set is wrongly leaning into — only a component carrying no information about how the set splits. Removing it would recover 0.024 of alignment and leave the column at 0.451, still well short of leverage’s 0.5395 and further short of the projected fourth power’s 0.6583. The active set is not a probe that has been spoiled; it is a probe that was never going to be better than the quantity it is three quarters of.

What being blocked and being decisive are

There is a plainer way to say why the active set is the wrong quantity, and it is worth having because the original argument is so nearly right.

The separating direction distinguishes the two components of the admissible set: it is the difference between the two components’ average sign vectors, so a unit contributes to it in proportion to how differently the two components treat it. A unit that is treated with one sign in one component and the other sign in the other is a unit the direction is about.

The blocked share measures something adjacent and different: how often a unit’s exchanges fail. A unit whose exchanges fail often is a unit whose columns are large in some direction, and the failure can be in any coordinate of the box.

Those coincide when the constraint is tight in the direction that separates the components and part company when it is tight elsewhere. A unit can be immobile because its own column is extreme in a coordinate that has nothing to do with how the set falls in two, and the blocked share counts that unit as important while the separating direction does not.

Leverage is a better summary of “the rule has decided about this unit” than the active set is, and the active set is a better summary of “this unit cannot move”. The diagnostic needs the first.

The two come apart most where the box’s coordinates are least alike. A tolerance box constrains each column separately, so a unit can be pinned by the fourth column and free in the first three, and the blocked share cannot tell which coordinate did the pinning. Leverage sums the squared entries across columns, which is also blind to direction — but it weights the columns the way the span does rather than the way the box’s corners do, and the separating direction is a fact about the span.

That is a difference of weighting rather than of kind, and it is the size of the difference between 0.5395 and 0.4272.

What the winner is, and why it is odd that it wins

The best feasible probe on this table is not leverage. It is the fourth power of the covariate projected off the rule’s span, at 0.6583 of alignment and 5.0800 of separation — ahead of leverage by a wide margin and ahead of the active set by more.

That is worth dwelling on, because it is a column with no argument behind it at all.

The fourth power is a dictionary function the earliest field in this line chose precisely for being outside what the rule was handed: the rule balances the covariate, its square, its cube and a median cut, so its fourth power is the next thing along and is the obvious diagnostic. That choice was made before anybody knew the set could fall apart, and it was made on the reasoning that a probe the rule has already balanced cannot see anything.

Raw, it is nearly useless: 92% of it lies inside the span the rule balanced, and it aligns at 0.1720. Projecting it off that span triples what it can see and produces the best feasible probe anybody here has found.

So the table’s ordering is: an arbitrary dictionary function, cleaned up, beats a designed summary of the rule’s own geometry, which beats the rule’s own active set. Two probes built from arguments about what the rule does lose to one built from a function the rule was not handed, and that is not what any of the three arguments predicted.

What would have been the other answer

It is worth saying what would have happened if the field had come out the other way, because that is what makes this a measurement rather than a demonstration.

Had the active set beaten leverage, the reading would have been the deferral’s own: a heuristic was standing in for a quantity, the quantity is computable, use the quantity. The probe would have gone into the table as the best feasible one, and the field would have been three essays long.

Had it beaten the projected fourth power as well, it would have changed what a practitioner should run — the projected column is the earlier field’s recommendation, and it survives.

Neither happened, and what is left is a specific negative claim with a mechanism attached: the part of the active set that is not leverage points away from where the set splits. That claim is checkable, it is the kind of thing another setting could refute, and it is more useful than a fourth probe would have been.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A model and a count — both name assignment mechanism, basis, connected component, covariate balance, exact enumeration, experimental design, imbalance, leverage, markov chain monte carlo, randomisation test, rerandomisation
  • What a chosen probe finds — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, leverage, markov chain monte carlo, model diagnostics, orthogonality, projection, randomisation test
  • A probe nobody chose — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, projection, randomisation test, rerandomisation
  • A defect that is about size — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • Before the trial and after — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation
  • The statistic the p-value is about — both name assignment mechanism, connected component, covariate balance, exact enumeration, imbalance, markov chain monte carlo, randomisation test, rerandomisation

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismBasisConnected componentCovariate balanceExact enumerationExperimental designHat matrixImbalanceLeverageMarkov chain Monte CarloModel diagnosticsOrthogonalityProjectionRandomisation testRerandomisation