A probe chosen rather than picked

The part the rule already took

A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units.

Worth reading first: Randomisation is not balance · Balancing what is known in advance.

A two-chain diagnostic asks whether a randomisation walk can reach the whole admissible set, and it asks through a probe: a column the balancing rule was not handed, whose signed imbalance changes sign under the complement. The field that ran it on the trial’s own outcome found that seven of twenty-four outcomes report nothing at all on a set enumerated to be in two components, and that how far apart the components are on a probe ranges over five thousand-fold across those outcomes.

That leaves an obvious question and the field says so: the separation is a property of the design and the rule, with no outcome in it, so a probe could be chosen rather than picked from a dictionary and hoped for.

Choosing it properly needs the components, which needs the enumeration, which is what the diagnostic exists to avoid. But there is something available before that, it costs one least-squares fit, and it turns out to be most of what there is.

The setting, stated once

Everything below is measured in the one place the answer is knowable. Fourteen units, a covariate drawn standard normal, a rule balancing xx, x2x^2, x3x^3 and a cut at the median inside a box of tolerance 0.8, and assignments split equally between two arms.

At that size the admissible set can be enumerated: all 3,432 equal splits are checked, the admissible ones kept, and the components under single swaps found exactly. That is what supplies the separating direction — the difference between the two components’ mean assignment vectors — against which every probe here is measured. It is also what no trial has: the enumeration stops at about twenty-four units and trials are not that size.

So the design of this field is the standard one for a benchmark. Measure the feasible things where the infeasible thing is available, and report what each recovers, knowing that the setting is small and that smallness is exactly what the first finding is about.

What the rule has already taken

A balancing rule constrains the imbalance in a span. If it balances xx, x2x^2, x3x^3 and a cut at the median, then every admissible assignment has all four of those imbalances inside a tolerance. A probe that overlaps that span is measuring a quantity the rule has already forced towards zero, and the part of it that carries any information about which component the assignment is in is the part orthogonal to the span.

That much is obvious once stated, and it produces a check nobody had run: how much of each probe is inside the span?

What is left of a probe after the rule has had it. The share of each dictionary function a rule balancing x, x2, x3, cut0 has already taken, on trials of 14 units, averaged over 100 designs. Four of the eight functions are the basis, so their share is exactly one: a randomisation test run on one of them is asking about a quantity the rule forced to zero, and one of them is the default probe of the field this measurement comes from. The four that are not still read 0.919, 0.873, 0.903, 0.832 — between 0.832 and 0.919 of them is inside the span — against closed-form removed shares of 0.000, 0.692, 0.590, 0.692. At 14 units a rule with four functions in it takes most of anything it is shown.
Fig. 1 The share of each dictionary function a rule balancing four of them has already taken, on fourteen units, over a hundred designs.

Four of the eight functions in the dictionary are the basis, so their share is exactly one. A randomisation test run through one of them is asking about a quantity the rule set to zero by construction, and it will report whatever the tolerance allows and nothing about the set.

The four that are not in the basis read 0.9192, 0.8727, 0.9034 and 0.8322. The fourth power of the covariate — the one the earlier fields reach for when the basis contains the first three — is 91.92% inside the span the rule balanced.

Which is a fact about fourteen units, not about the fourth power

The number that makes this worth an essay is beside it in the same figure.

The exact share a rule balancing xx, x2x^2, x3x^3 and the median cut removes from x4x^4 is 0.0000. Not small: zero. The Hermite polynomials are orthogonal under the covariate’s own law, and the median cut is orthogonal to x4x^4 as well, so in the population the fourth power is exactly the direction the rule cannot touch. It is the right probe, chosen for the right reason.

Orthogonal in the population, absorbed in the sample. How much of two probes a balancing rule has already taken, against the size of the trial. The fourth power of the covariate is exactly orthogonal to the rule's basis in the population — its removed share is 0.0000 in closed form, because the Hermite polynomials are orthogonal and the median cut is orthogonal to it too. On fourteen units the rule has taken 0.8989 of it. The share falls to 0.5501 at a hundred units and 0.0957 at sixteen hundred, which is the rate a sample correlation converges at and is not fast. A cut at one is genuinely correlated with the basis and sits at its own closed form of 0.6924 throughout. A probe chosen for being orthogonal to what the rule reads is orthogonal to it in the population and inside it in the sample, and a fourteen-unit trial is where the difference is nearly total.
Fig. 2 The same probe’s pinned share against the size of the trial, with its exact removed share marked. Orthogonality arrives slowly.

On fourteen units the rule has 0.8989 of it. At fifty units, 0.7233. At a hundred, 0.5501. At four hundred, 0.2295. At sixteen hundred, 0.0957.

Orthogonality in the population is not orthogonality in the sample, and a trial the size an enumeration allows is where the difference is nearly total. Four functions plus an intercept span five of a fourteen-unit design’s thirteen centred dimensions, and a column with the variance x4x^4 has is absorbed into them almost entirely.

A cut at one behaves differently and is the control for the reading: it is genuinely correlated with the basis, its exact removed share is 0.6924, and its sample share is 0.8842 at fourteen units and 0.7008 at sixteen hundred — converging to its own closed form rather than to zero. So the measurement is not simply reporting that small samples absorb everything; it separates a probe that is orthogonal and looks absorbed from one that is genuinely half-absorbed.

Why a rule’s own basis is in the probe dictionary at all

There is a design decision underneath the first figure that deserves stating, because it is the sort of thing that looks like a bug and is not.

The eight functions the probes are drawn from and the four the rule balances come from one dictionary. That is deliberate and it is what makes the collection’s fields comparable: the same eight shapes are what a balancing rule may be asked to balance, what an outcome may depend on, and what a diagnostic may probe with. The field that built the dictionary treats it as the vocabulary of the whole subject rather than as three separate lists.

The consequence is that half the probes available on any given trial are the rule’s own basis, and a reader choosing one at random has an even chance of probing a quantity that was set to zero on purpose. That is not a hazard the earlier fields fell into — they choose the fourth power precisely to avoid it — but it is the reason the check is worth running rather than assumed: a probe’s independence from the rule is a property of the pair, and nothing in the code makes the pair explicit.

Half absorbed at a hundred and twenty-five units

“Orthogonality arrives slowly” is the essay’s summary of the five pinned shares, and the five have a shape with a scale in it.

Taking 1/pinned − 1 at n = 14, 50, 100 and 400 gives 0.113, 0.383, 0.818 and 3.36 — very nearly proportional to n, at about n/125. So

pinned ≈ 125 / (125 + n)

which gives 0.899, 0.714, 0.556 and 0.238 against counted values of 0.8989, 0.7233, 0.5501 and 0.2295. Four of the five to within a hundredth; the sixteen-hundred-unit reading is the one that departs, at 0.0957 against a predicted 0.0725.

The number worth carrying is the scale. The fourth power is half absorbed at about a hundred and twenty-five units, a quarter absorbed at three hundred and seventy-five, and at the two hundred units a real trial might have it is still 38% taken. A probe chosen for being exactly orthogonal in the population is a third gone at the sizes trials come in, and the projection that gives it back is not a refinement for small designs — it is the ordinary case.

Twenty-five dimensions where five were counted

The absorption is not dimension counting, and comparing it with what dimension counting would predict says by how much.

A random column of a fourteen-unit design has an expected share of p/(n − 1) = 5/13 = 38.5% inside a five-dimensional span. The fourth power has 89.9% — 2.3 times as much. At sixteen hundred units the two are 0.31% and 9.57%, a factor of 31.

Both fall like 1/n, so the ratio settles: with pinned ≈ 125/n and a random column at 5/n, the fourth power is absorbed twenty-five times more than a random direction of the same nominal dimension.

The span behaves as though it had a hundred and twenty-five dimensions rather than five against this particular column. That is the honest description of what is happening: x⁴ is not merely one more direction among many, it is nearly collinear with x² and x³ in any sample of a normal covariate, because a normal’s high moments are dominated by a handful of extreme units and those units drive all three columns together.

The cut at one is the control and it behaves as a control should. Its excess over its own exact share falls from 0.192 at fourteen units to 0.008 at sixteen hundred — a factor of twenty-three — converging on 0.6924 rather than on zero. A genuinely correlated probe converges to its correlation and an orthogonal one converges to zero, and both do it slowly enough that neither is at its limit at any size a trial reaches.

Giving it back

The repair is a projection. Regress the probe on the rule’s own columns, keep the residual, standardise it. It needs no enumeration, no outcome, and no assumption about the design; it is available to anybody who knows what the rule balanced, which is everybody who ran it.

How much of the separating direction each probe is. The alignment between each probe and the direction that separates the two components of the admissible set, averaged over the 100 of 200 designs whose set is enumerated and found split. One is the separating direction itself, which needs the enumeration. The fourth power as the earlier fields use it reads 0.1720 ± 0.0154; the same column projected off the span the rule balances reads 0.6583 ± 0.0279, a paired gain of 0.4863 at 16.7 standard errors. The design's own leverage, which needs no dictionary at all, reads 0.5395. A random direction in the same subspace reads 0.2622 — so the projection is most of the gain and the choice of direction is the rest.
Fig. 3 How much of the separating direction each probe is, over the hundred designs of two hundred whose admissible set is enumerated and found split.

Over a hundred split designs the raw column aligns with the separating direction at 0.1720 ± 0.0154. Its own residual off the rule’s span aligns at 0.6583 ± 0.0279 — a paired gain of 0.4863 ± 0.0291, at 16.7 standard errors.

A factor of nearly four, from a column that was already there. Nothing was added to the probe; a part of it was removed, and the part removed was the part the rule had pinned.

Where the designs that split are

One number in the last figure is a fact about the setting rather than about probes, and every measurement in this field is conditioned on it.

Of two hundred designs of fourteen units drawn from the same law, at the same tolerance, a hundred have an admissible set that splits into two components and a hundred do not. Which of those a given trial is in is decided by the fourteen covariate values it happened to draw, and nothing about the rule, the tolerance or the sample size says which.

Everything measured here is conditional on the split, because on a connected set there is no separating direction and no separation to carry. That is the right conditioning for the question — how well does a probe see a split that is there — and it is the wrong conditioning for a different question a reader may have, which is how often a probe fires when it should not. That one is measured in the field that built the test, against the enumerated truth in both directions, and nothing here touches it.

A better probe is a probe that finds more splits, not one that finds more things. The two-chain statistic’s control — the same comparison on a statistic that is symmetric under the complement, and therefore blind to the split by construction — sits below one for every probe in this field, which is what says the extra firing is signal rather than noise.

What it is worth in the quantity that matters

Alignment is a direction cosine, and what the diagnostic is actually trying to see is the separation: how far apart the two components’ mean probe values are, over the spread inside a component.

How far apart the two components are, on each probe. The median separation between the two components of the admissible set — the difference in their mean probe values, over the spread inside a component — over the 100 of 200 designs whose set is enumerated and found split. The separating direction carries 10.565 and needs the enumeration. The fourth power as the earlier fields use it carries 1.543; projected off the span the rule balances, 5.080. The design's own leverage, which uses no dictionary and no outcome, carries 3.836. A random direction in the same subspace carries 0.942, and a direction chosen by looking for concentrated structure carries 0.543 — below random, and the one heuristic here that is worse than not choosing at all.
Fig. 4 The median separation each probe carries between the two components, over the same hundred designs.

The raw column carries a median separation of 1.543. Its residual carries 5.080. The separating direction itself — which needs the enumeration, and is what everything here is short of — carries 10.565.

So the projection recovers about half of what the enumeration would give, and it triples what the field was working with. The remaining half is the next essay’s subject, and it turns out to be much harder to get.

What one least-squares fit costs

It is worth being blunt about the price, since the whole recommendation rests on it being nothing.

The rule’s columns are already computed — they are what the constraint is checked against, so any implementation has them. Projecting a probe onto their orthogonal complement is one Gram–Schmidt pass: four inner products, four scaled subtractions, and a standardisation. On fourteen units that is under a hundred multiplications, once per trial, before any chain is run.

Against that, the chain the diagnostic runs takes twenty thousand steps and each step re-evaluates the constraint. The projection is free by any measure that matters, and the reason it was not being done is that nothing said to do it — the population argument for choosing an orthogonal function is correct, complete on its own terms, and silent about the realised design.

The thing this is not

It is worth separating this from a different and better-known problem, because the two look alike and the remedies are opposite.

A randomisation test’s reference distribution is about the quantities the rule did not constrain — that is the design principle the probe exists to serve, and it says to choose a function outside the balanced span. The earlier fields follow it: they pick the fourth power precisely because it is orthogonal to the basis.

The principle is right and the implementation of it is a population statement applied to a sample. What a rule balances on a given trial is the span of its columns as realised on those fourteen units, and that span is not the population span. A function orthogonal to the second is, on a small trial, mostly inside the first.

So the remedy is not to choose a different function. It is to take the function already chosen and remove the part of it the realised design has put inside the span — which is a sample operation for a sample problem, and which the population argument gives no reason to do.

Two routes to the same share

The pinned share is measured twice here and by arithmetic that shares nothing, which is what makes the finite-sample reading believable rather than a suspicion about a projection.

The first route is the projection itself: regress the probe on the realised columns, compare the residual sum of squares to the total. It knows about fourteen particular covariate values and nothing else.

The second is the closed form, R²(g | span B) computed as an integral against the covariate’s law — a′G⁻¹a with a the covariances of the probe with the basis and G the basis’s own Gram matrix, all evaluated exactly. It knows about the law and nothing about any sample.

They agree at large n on the three cut functions, to within 0.02 at sixteen hundred units. They disagree by 0.899 at fourteen on the fourth power, and the disagreement is the finding. A check that only ever agreed would have said nothing; this one agrees where it should and the size of where it does not is the essay.

What the number would be on a real trial

A reader whose trial has two hundred units rather than fourteen will want to know whether any of this reaches them, and the honest answer is a partial one.

The pinned share at two hundred units is between the 0.5501 measured at a hundred and the 0.2295 at four hundred — call it a third of the probe, still inside the span, still removable for nothing. That is a smaller effect than the 0.92 at fourteen and it is not negligible: a third of a probe is a third of its variance spent on a quantity the rule set to zero.

What cannot be checked at two hundred units is whether removing it helps, because there is no enumeration and therefore no separating direction to align with. So the recommendation transports on the mechanism and the verification does not, which is the ordinary position for anything measured at a size where the truth is computable.

What is left in the probe

One number in the alignment figure is worth reading against the others, because it says how much room the projection leaves.

The separating direction itself has a pinned share of 0.1266. It is not orthogonal to the rule’s span; about an eighth of it lies inside. So projecting is not simply moving towards the answer — the answer has a component the projection removes, and the projected probe is therefore not merely a worse version of the oracle but a slightly different object.

That is a small correction and it matters for how the second half of this field is set up. The best a projection can do is find the best direction in the orthogonal complement, and the best direction overall is not quite in there. What the enumeration knows that a projection cannot is worth about the eighth it keeps inside the span, plus whatever choosing well inside the complement is worth.

What the alignment comes to in the verdict

Alignment and separation are properties of a direction. What an experimenter gets is a verdict, and the two are not the same thing: a probe can carry a real separation and still be missed by a chain that was not run long enough to resolve it. So the last measurement here is the one the whole field is for — on designs enumerated to be split, so that a quiet verdict is a miss and never a correct silence, how often does the two-chain test actually fire?

On a short chain of eight hundred draws, the raw probe misses 44% of the sets that are genuinely split, and its own residual misses 12%. Those are the same two directions the alignment figure separated by 0.17 against 0.66, and this is what that gap is worth once it has been put through a test of finite length.

The size of it is the point. Removing the component the rule already took is not a refinement that buys a decimal place; it turns a diagnostic that is wrong on nearly half the cases it is shown into one that is wrong on one in eight, without collecting a single additional draw, changing the threshold, or knowing anything the experimenter did not already write down. The whole of the improvement is subtraction — the probe is the same column it always was, minus the part that could never have told the two components apart.

And it is a floor rather than a ceiling for the same reason the pinned share of 0.1266 was a correction rather than a rounding error. The residual is the best direction in the complement, and about an eighth of the separating direction is outside the complement altogether. The 12% is what a projection can reach; what remains is what enumeration knows and projection cannot.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • What a chosen probe finds — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, leverage, markov chain monte carlo, orthogonality, projection, randomisation test, reference distribution
  • A model and a count — both name assignment mechanism, connected component, covariate balance, exact enumeration, experimental design, leverage, markov chain monte carlo, randomisation test
  • A defect that is about size — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
  • A dictionary that is neither — both name basis functions, covariate balance, experimental design, hermite polynomials, orthogonality, projection, variance explained
  • Before the trial and after — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution
  • Half a reference distribution — both name assignment mechanism, connected component, covariate balance, exact enumeration, markov chain monte carlo, randomisation test, reference distribution

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismBasis functionsConnected componentCovariate balanceExact enumerationExperimental designHermite polynomialsLeast squaresLeverageMarkov chain Monte CarloOrthogonalityProjectionRandomisation testReference distributionVariance explained