A probe chosen rather than picked

A probe chosen from the design

The design's own leverage aligns with the separating direction four times better than a random direction in the same subspace. The concentrated direction the argument invites is worse than random.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

Projecting a probe off the span the rule balanced takes it from a separation of 1.543 to 5.080, against an oracle of 10.565. Half the gap to the enumeration closes for one least-squares fit and no choice at all.

The other half needs a choice: which direction in the orthogonal complement. On fourteen units with four balanced functions that complement has nine dimensions, and the separating direction is one particular ray in it. The enumeration would say which; nothing else does.

This essay is what can be done without it, and it has two halves that point in opposite directions.

The subspace, and what is in it

A rule balancing four functions on fourteen units, with the probe centred, leaves a complement of nine dimensions. That is where every candidate lives once the projection has been made, and the first thing to establish is how much room a choice has in it.

The complement is where every candidate probe now lives, so it is worth knowing how much room a choice has.

A random direction in it aligns with the separating direction at 0.2622 ± 0.0174. That is the floor: what a probe gets for being in the right subspace and nothing else. The projected fourth power gets 0.6583 ± 0.0279 — so the fourth power, once its pinned part is removed, is not a generic direction in the complement, it is a much better one.

But the fourth power is a dictionary function, and choosing it is the thing this field was asked to replace. Picked from a dictionary and hoped for is exactly what the deferral names. What is wanted is a direction computed from the design and the rule.

What a chosen probe cannot do

Before the choosing, one thing the projection has already fixed and no direction in the complement can change.

The two-chain test compares two chains started at complementary assignments on a statistic that changes sign under the complement. Every probe here is such a statistic, so every probe here is testing the same hypothesis and differs only in how loudly. A probe that is symmetric under the complement — the absolute imbalance, say — is blind to the mirror split by construction, whatever subspace it is in, and the field that built the test carries it as the control precisely because it must stay quiet.

So the choosing here is about power and not about validity. A badly chosen probe misses splits; it does not invent them. That is worth knowing before reading a ladder of separations, because it says which end of the ladder is dangerous: the bottom, where a diagnostic reports nothing and a reader concludes the walk is fine.

The floor is what nine dimensions predicts

The random-direction figure is worth checking against what a uniformly random direction in a nine-dimensional space has to give, because the check confirms the geometry the whole essay rests on.

For a uniform unit vector in dd dimensions, the squared alignment with any fixed direction follows a Beta(½, (d−1)/2), so at d = 9 its mean is exactly 1/9, and the mean of the alignment itself is

Γ(1)Γ(4.5)Γ(0.5)Γ(5)=0.2734.\frac{\Gamma(1)\,\Gamma(4.5)}{\Gamma(0.5)\,\Gamma(5)} = 0.2734 .

Measured: 0.2622 ± 0.0174 — six tenths of a standard error away.

So the complement really is nine-dimensional and the random directions really are uniform in it. That is not a formality: had the measured floor come in at 0.33 or at 0.20, the complement’s effective dimension would have been something other than nine and every ratio in the ladder would be against the wrong baseline.

Where the projected fourth power sits among random directions

With the floor’s whole distribution known rather than just its mean, the fourth power’s 0.6583 can be placed as a percentile rather than as a multiple.

In squared terms that is 0.4334 against a random direction’s average of 0.1111 — 3.9 times the expected alignment, or 5.8 times in the variance a probe actually gets to use.

And integrating the Beta(½, 4) density above 0.4334 puts the chance that a random direction does at least that well at about 4%.

So the projected fourth power is at roughly the 96th percentile of directions in the complement. That is a real advantage and it is a bounded one. A reader told that a dictionary function beats a random direction by a factor of four might imagine it is nearly the best direction available; it is beaten by one random draw in twenty-five.

Which sets the size of the prize this essay is chasing. The oracle direction is at 1.000 by construction, the fourth power is at 0.658, and the remaining third of the way is not a matter of refining a good choice — it is the difference between the 96th percentile and the top, in a space where a random guess already reaches 0.27.

How much of the separating direction each probe is. The alignment between each probe and the direction that separates the two components of the admissible set, averaged over the 100 of 200 designs whose set is enumerated and found split. One is the separating direction itself, which needs the enumeration. The fourth power as the earlier fields use it reads 0.1720 ± 0.0154; the same column projected off the span the rule balances reads 0.6583 ± 0.0279, a paired gain of 0.4863 at 16.7 standard errors. The design's own leverage, which needs no dictionary at all, reads 0.5395. A random direction in the same subspace reads 0.2622 — so the projection is most of the gain and the choice of direction is the rest.
Fig. 1 Every probe against the direction that separates the two components, averaged over the hundred designs whose set is enumerated and found split. The fourth power as the earlier fields use it reads 0.1720, the same column projected off the rule’s span 0.6583, the design’s own leverage 0.5395, and a random direction in the same subspace 0.2622.

Leverage

One quantity is available and has a mechanism behind it.

The leverage of a unit is how much of the rule’s span it carries: the diagonal of the basis’s hat matrix, hi=b(xi)G1b(xi)h_i = b(x_i)^{\top}G^{-1}b(x_i). It is a function of the fourteen covariate values and the four balanced functions, and of nothing else — no outcome, no enumeration, no dictionary.

The mechanism is about which assignments the constraint can move between. A unit with high leverage is one whose sign the constraint has most to say about: flipping it moves the imbalance a long way, so a swap involving it is the most likely to leave the admissible box. The units that pin the set into pieces should be those units, and the separating direction — which is exactly the units whose sign is nearly determined within a component — should therefore be concentrated on them.

How far apart the two components are, on each probe. The median separation between the two components of the admissible set — the difference in their mean probe values, over the spread inside a component — over the 100 of 200 designs whose set is enumerated and found split. The separating direction carries 10.565 and needs the enumeration. The fourth power as the earlier fields use it carries 1.543; projected off the span the rule balances, 5.080. The design's own leverage, which uses no dictionary and no outcome, carries 3.836. A random direction in the same subspace carries 0.942, and a direction chosen by looking for concentrated structure carries 0.543 — below random, and the one heuristic here that is worse than not choosing at all.
Fig. 2 The median separation each probe carries between the two components, over the hundred split designs.

Leverage, projected off the span, aligns at 0.5395 ± 0.0298 and carries a separation of 3.836. Against a random direction in the same subspace — 0.2622 and 0.942 — that is a paired gain of 0.2774 ± 0.0332 at 8.4 standard errors, and four times the separation.

So the mechanism is right and the choice is worth having. It is not worth as much as the projected fourth power, which reaches 0.6583 and 5.080; the paired gap between them is 0.1188 ± 0.0262 at 4.5 standard errors, in the dictionary’s favour. Leverage is the honest form — it needs nothing but the design and the rule — and it buys about two thirds of what a well-chosen dictionary function buys.

What leverage is and is not, here

The word does duty for two different things in this collection and it is worth keeping them apart.

In a regression leverage is how much one observation can move its own fitted value, and a high-leverage point is a hazard: it can own a slope. Here the same arithmetic on the same hat matrix is being read as a resource — the units the constraint has most to say about are the units whose signs distinguish the components, so the same diagonal that flags a dangerous observation flags an informative probe direction.

Nothing about the two readings conflicts. They are the same number answering two questions: how much of the fit does this unit determine, and how much does the constraint determine this unit. A rule that balances a span constrains hardest where the span is largest, which is where the leverage is.

That is also why the leverage vector has to be projected off the span before it is used. It is built out of the basis columns and is a quadratic form in them, so a large part of it is inside their span by construction, and handing it to a diagnostic un-projected would be the same mistake the first essay of this field is about, made with a different column.

And a heuristic that is worse than nothing

The obvious next move is to look for the direction directly, and it fails in a way worth reporting.

The separating direction is concentrated: it is large on the few units whose signs are pinned and small elsewhere. So a direction in the complement chosen for being concentrated — maximising the fourth moment, which is projection pursuit’s oldest index, by fixed-point iteration from eight random starts and kept by the index itself — ought to find something like it.

It aligns at 0.2001 ± 0.0150 and carries a separation of 0.543.

That is below a random direction on both readings, and the paired comparison says so: random beats pursuit by 0.0621 ± 0.0233, at 2.7 standard errors.

The failure is legible once measured. The most concentrated direction in a nine-dimensional subspace of a fourteen-unit space is very nearly a single unit’s indicator, residualised — it puts almost all its weight on one row. The separating direction is concentrated on several units and has structure among them, and a single-unit spike is nearly orthogonal to it. The index found the wrong kind of concentration, and the search made it worse than not searching: a random direction spreads its weight and picks up some of the separating direction by chance, where the optimised one commits to a spike and misses.

A heuristic can be worse than nothing. This is the one a reader would propose after the first essay, and it is reported as a rung rather than dropped for that reason.

How often each probe finds a split that is there. The share of 34 designs — every one of them enumerated to be in two components — on which a two-chain test of 800 draws declares the split, by probe. The fourth power as the earlier fields use it finds it on 55.9%, so it misses 44.1% of the sets that have one. The same column projected off the rule's span finds it on 88.2%, and the separating direction itself on 91.2%. The design's own leverage, chosen without any dictionary, gets 79.4%. A random direction in the same subspace gets 44.1%, and the direction chosen for being concentrated gets 38.2% — worse than random, which is what a heuristic that finds the wrong structure looks like from the outside.
Fig. 3 The same six probes on the verdict rather than on the separation: the share of thirty-four enumerated split sets a two-chain test of eight hundred draws actually declares. The concentrated direction gets 38.2% against a random direction’s 44.1% — worse than random, which is what a heuristic that has found the wrong structure looks like from outside.

What a chosen direction is being chosen against

It is worth being precise about the benchmark, because “a random direction in the complement” is doing a lot of work in the comparison and it is not an obvious baseline.

The alternative baseline would be zero: a probe that sees nothing. That is the wrong comparison, because being in the complement at all is already most of the way — the projected fourth power gets 5.080 and a random member of the same subspace gets 0.942, so the subspace supplies a sixth of what the best feasible probe has and the direction supplies the rest. Measuring a chosen direction against zero would credit the choice with the projection’s work.

Against a random direction, the comparison is clean: same subspace, same standardisation, same designs, drawn from a stream that touches nothing else. What is being priced is the choice and only the choice.

And the random baseline is exactly what it should be. A uniformly random direction in dd dimensions has an expected absolute cosine with any fixed direction of 2/πd\sqrt{2/\pi d}. The complement here has nine dimensions — fourteen units, less the mean, less the four functions the rule balances — so the prediction is 0.2660, and the measurement is 0.2622 ± 0.0174.

That is a closed form and a Monte Carlo agreeing to a quarter of a standard error, and it is the check that matters most in this essay: it says the orthogonal complement is not secretly aligned with the separating direction, so everything a chosen probe beats the baseline by is the choice rather than the subspace.

The ladder

Five feasible probes and one that is not, on the same hundred designs, by the separation they carry:

0.543 the most concentrated direction off the span · 0.942 a random direction off the span · 1.543 the fourth power as it is used · 3.836 the design’s own leverage, off the span · 5.080 the fourth power, off the span · 10.565 the separating direction itself.

Two things are worth reading off that list rather than one.

The projection is the large effect and the choice is the small one. From 1.543 to 5.080 is the projection applied to one column; from 0.942 to 3.836 is choosing well inside the complement. Both are worth having and the first is free.

And the ordering is not the ordering of how clever the probe is. The most computed of the six — an optimised projection-pursuit direction — is last. The least computed that works — leverage, one hat-matrix diagonal — is third from the top. What separates them is whether the quantity being maximised is the quantity that matters, and the pursuit index is not.

What the ladder does not settle

Two limits, and both are about what a hundred designs of fourteen units can support.

The ordering between leverage and the projected fourth power is not large. 3.836 against 5.080 on medians, 0.1188 ± 0.0262 on the paired alignment. It says the dictionary function is better on this basis, at this size, with this covariate law; it does not say a dictionary is better than a design-based rule in general, and the obvious way to find out — vary the basis — is not run.

And every design here shares one basis. All hundred split designs balance xx, x2x^2, x3x^3 and a median cut. What differs between them is the fourteen covariate values. So the field varies the sample and holds the rule fixed, which is the right design for this question and leaves untouched whether a rule with two functions or six behaves the same way. A rule with fewer functions leaves a larger complement and a harder choice; a rule with more leaves a smaller one and less to choose from.

What this changes about running the diagnostic

The recommendation that comes out of two essays is short and it is a change to a line of code rather than to a procedure.

Project the probe. Whatever column is being used — a dictionary function, an outcome, anything — regress it on the rule’s own columns and keep the residual. It costs one Gram–Schmidt pass, it needs nothing the implementation does not already have, and it triples the separation on the setting measured here.

And if there is no natural column, use the leverage. It needs no dictionary, no outcome and no choice, it is four times a random direction, and it is two thirds of the best dictionary function this field found.

What is not recommended is searching for a good direction. The one search tried here made things worse than not searching, and the reason it did — the index maximised is not the quantity that matters — is not specific to the fourth-moment index. Any index computable without the enumeration is a guess about what the separating direction looks like, and the one guess with a mechanism behind it, leverage, is available without any searching at all.

One number that is not in the ladder

The separations above are medians, and the means are much larger and much less useful — which is worth one paragraph because it says something about the quantity rather than about the summary.

A probe’s separation is a ratio: the distance between the two components’ mean probe values, over the spread inside a component. The denominator can be very small. A probe that is nearly constant within a component — which happens when it is close to a function of the constraint itself — has almost no within-component spread and a ratio that runs away, and the field that defined it reports those as degenerate rather than dividing by zero and shipping fifteen digits.

So a handful of designs produce enormous ratios for reasons that have nothing to do with how good the probe is, and the mean of a hundred of them is dominated by three. The medians are the ladder and the means are not reported. That is an ordinary robustness choice and it is stated because the alternative summary would reverse one of the rungs: the concentrated direction, which is last on medians, has the largest mean of the six.

Why the oracle is still ahead

The gap from 5.080 to 10.565 is a factor of two and it does not close, so it is worth saying what is in it.

Part of it is not in the complement at all. The separating direction has a pinned share of 0.1266: an eighth of it lies inside the span the rule balanced, so no probe built by projecting can reach it. That is a ceiling on the whole approach and it is a low one only in the sense that an eighth of a direction is not most of one.

The rest is that the complement has nine dimensions and the separating direction is one ray in it. The projected fourth power aligns at 0.66 — it is two thirds of the way round — and the remaining third is exactly the part that depends on which fourteen covariate values were drawn, which is what the enumeration knows and the design alone does not.

So a heuristic gets half the separation and there is no argument here that a better one gets much more. What would decide it is a probe built from something the design and the rule determine that the two constructions here do not use — the constraint’s active set, say, or which swaps are blocked — and neither is tried.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A defect that is about size — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
  • Before the trial and after — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
  • Half a reference distribution — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
  • The set a dictionary leaves — both name assignment mechanism, basis functions, covariate balance, experimental design, markov chain monte carlo, randomisation test, reference distribution
  • The statistic the p-value is about — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
  • The walk that cannot cross — both name assignment mechanism, basis functions, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismBasis functionsConnected componentCovariate balanceExact enumerationExperimental designHat matrixLeverageMarkov chain Monte CarloNumerical methodsOrthogonalityProjectionRandomisation testReference distributionRobustness