Choosing what the rule reads

A basis is a subspace

A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.

Worth reading first: A design is a number · The variance removed before the data.

A rule that balances a column of numbers is balancing every multiple of that column, and a rule that balances two columns is balancing every combination of the two. It is not reading functions; it is reading a subspace, and it cannot tell one basis of that subspace from another.

That sounds like a technicality. It is the thing that turns a question with no answer — which functions should the rule read? — into a question with an exact one, because a subspace has a geometry and a list does not.

A rule cannot tell one basis from another with the same span. The first column is a rule handed the covariate and its square. The second is a rule handed x + 2x² and 3x − x², which are two different functions spanning the same plane. Every number in the two columns agrees to 0e+0 — the criterion a balancing rule maximises is a function of the span alone, so choosing what to hand it is choosing a subspace of the space of functions rather than choosing a list. The third column is a genuinely different plane, the covariate and an indicator at 1, and there the numbers move.
Fig. 1 A rule handed the covariate and its square, and a rule handed x + 2x² and 3x − x². They are the same rule, and every number they produce agrees to machine precision.

The identity

Suppose the outcome depends on some unit-variance function g of the covariate, and the rule balances the span of a basis B perfectly. Then the imbalance it leaves in g is the part of g the span cannot reach:

var(imbalance in g)=(1R2(gB))×var(imbalance under a coin),\operatorname{var}(\text{imbalance in } g) = \big(1 - R^2(g \mid B)\big) \times \operatorname{var}(\text{imbalance under a coin}),

with R² the squared multiple correlation of g on the span. The rule removes the projection and leaves the residual, and nothing else about g matters.

The one-function case of this is measured in the field that found the exposure, where it is ρ² and reads as a fact about a threshold: balancing the covariate’s mean removes exactly 2/π of the imbalance in a median split of it. It is not a fact about thresholds. It is a fact about projections, and the constant is one entry of a Gram matrix.

Every number here is a closed form

The inner products are exact rather than quadrature, which is worth a paragraph because it is what makes the rest of this field a piece of arithmetic instead of a simulation study.

Take the orthonormal Hermite functions h_j = He_j/√(j!), which are uncorrelated under a standard normal covariate by construction. Then an indicator’s covariance with one of them is

1{x>c},hj=φ(c)Hej1(c)/j!,\langle \mathbf{1}\{x > c\},\, h_j \rangle = \varphi(c)\, He_{j-1}(c) / \sqrt{j!},

and two indicators’ covariance is Φ̄(max(c, c′)) − Φ̄©Φ̄(c′). So a dictionary of polynomials and cut points has a Gram matrix in closed form, and R²(g | B) = a′G⁻¹a is a solve of a matrix no larger than four by four.

The check that the machinery is right is the constant the neighbouring field derived by a different route. Computed here as a projection, the share of a median split that balancing the covariate removes is 0.6366197724, and 2/π is 0.6366197724. Read the other way — the share of the covariate that balancing a median split removes — it is the same number, which it must be, since a squared correlation is symmetric and the two statements are one integral read in opposite directions.

The six best bases of one function, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 2.3% against every shape in the list, and the worst of the six guarantees 0.0%: the difference between them is entirely which subspace was picked, at the same cost per arrival.
Fig. 2 Single functions, scored against six shapes. Every cell is a closed form; the column of zeros is orthogonality rather than a small number.

What the table says that a list of rules could not

With the geometry in hand, the twelve cells the earlier field measured become part of a much larger table, and three things in it are visible only at that size.

Orthogonality is exact and it is common. A rule reading the covariate removes 0.0000 of a quadratic outcome and 0.0000 of a cubic one, because odd and even Hermite functions are uncorrelated under a symmetric covariate. That is not “very little”; it is nothing, and no amount of data changes it.

A threshold’s protection depends entirely on where it sits. A rule reading the covariate removes 0.4386 of a threshold at 1 and 0.1311 of a threshold at 2. The further into the tail, the less a mean-balancing rule can see of it.

And an indicator is a surprisingly good single function. The best single function to hand the rule, judged by its worst shape, is an indicator at 2 — the one function whose worst cell is not zero, at 0.0233. Every polynomial has a shape it is exactly orthogonal to; an indicator has none, because an indicator is not an even or an odd function and so it has a component on every Hermite order.

The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.
Fig. 3 Pairs of functions. The best row here guarantees 26.8% against every shape in the list, and the worst of the six shown guarantees 15.1%, at the same cost per arrival.

Adding a function does not add a fixed amount

The table also settles a question the twelve-cell version could not, which is whether “a bigger basis” is a meaningful thing to say.

It is not, in general. A rule reading the covariate and its square removes 0.6579 of a threshold at 1, where the covariate alone removes 0.4386 — a square is not orthogonal to an indicator even though it is orthogonal to the covariate, so protecting against curvature protects partly against a cut. But a rule reading the covariate, its square and its cube removes 0.7427 of a median split where the covariate alone removes 0.6366, an improvement of a tenth for two extra functions.

The increments depend on the angles, and the angles depend on which shape is being protected. There is no ordering of bases by size that survives contact with the table, which is why choosing one by its worst case is a separate problem rather than a matter of adding functions until satisfied.

Why the span, and not the functions, is what a rule can see

The invariance is easy to state and easy to doubt, so it is worth showing where it comes from rather than only that it holds.

A balancing rule maximises the information about the treatment effect given the columns it has been handed: S_aa − S_ax′S_xx⁻¹S_ax, with x the columns. Replace those columns by any invertible linear recombination — x → xA — and S_ax becomes S_ax A, S_xx becomes A′S_xx A, and the correction term becomes

SaxA(ASxxA)1ASax=SaxSxx1Sax,S_{ax}A\,(A' S_{xx} A)^{-1} A' S_{ax}' = S_{ax} S_{xx}^{-1} S_{ax}',

with the A’s cancelling exactly. The criterion is unchanged, so the rule makes the same decision at every arrival, so the trial is the same trial.

That is why the first figure is not a coincidence: a rule handed x and x², and a rule handed x + 2x² and 3x − x², are not two rules that happen to agree. They are one rule described twice.

The practical consequence is worth stating because it cuts both ways. Rescaling, centring or recombining the functions handed to a balancing rule does nothing at all — which is reassuring, and means an experimenter cannot improve a design by standardising its columns. And the only way to change what a rule protects is to change the span, which means adding a function that is already a combination of the others is precisely zero work, however different it looks written down.

A rule cannot tell one basis from another with the same span. The first column is a rule handed the covariate and its square. The second is a rule handed x + 2x² and 3x − x², which are two different functions spanning the same plane. Every number in the two columns agrees to 0e+0 — the criterion a balancing rule maximises is a function of the span alone, so choosing what to hand it is choosing a subspace of the space of functions rather than choosing a list. The third column is a genuinely different plane, the covariate and an indicator at 1, and there the numbers move.
Fig. 4 The invariance, once more, with the third column as the control: a genuinely different plane, and there every number moves.

What a rule on four hundred units actually does

The identity is a statement about a rule that balances its span perfectly, and no rule does. So the last thing this essay owes is the gap between the geometry and a rule running on real units, and it turns out to have a shape.

The counted leftover is not 1 − R². It is affine in R², running from some α where the basis removes nothing to some β where it removes everything, and both departures are the finite sample:

  • α is above one, because a constraint that removes nothing about the outcome still competes for the same assignments. A rule reading the covariate leaves 1.0960 of a coin’s imbalance in a quadratic outcome — it is worse than a coin, by nine points, against a shape it cannot see at all.
  • β is above zero, because the rule balances its own columns approximately. A rule reading the covariate leaves 0.0074 of a coin’s imbalance in a linear outcome, and one reading three functions leaves 0.0262.
What the projection predicts and what a rule on real units leaves. Each line is one basis: four shapes scored against it, with the closed-form R² on the horizontal axis and the imbalance a rule on 400 units actually leaves on the vertical. The identity says the points should sit on 1 − R², and they sit on a line from α at the left to β at the right instead. Both departures are the finite sample: α is above one because a constraint that removes nothing still competes for the assignments (1.05, 1.15, 1.08, 1.14 as the basis grows), and β is above zero because the rule balances its own columns approximately (-0.032, 0.020, 0.024, 0.039). The slope is the geometry and it is the same in every row.
Fig. 5 Four bases, five shapes each, with the closed-form projection on one axis and what a rule on four hundred units leaves on the other. Two fitted numbers per basis; the slope is the geometry.

Both fitted numbers move with the size of the basis in the direction the competition between constraints predicts. Reading three functions rather than one takes β from −0.0320 to 0.0388 — the balance the rule achieves in the shape it was built for gets worse — and α from 1.0539 to 1.1361 — the cost against a shape it cannot see gets larger. Every one of the twenty cells sits on its own basis’s line to within 0.0898.

That is the two-route check for this field, and the two routes share nothing at all. One is a four-by-four solve on a Gram matrix of Hermite functions and indicators; the other is a rule assigning four hundred units one at a time, three hundred times over, and the standard deviation of an imbalance.

Why the affine form and not the identity

It is worth being clear that the gap is not an error term to be shrunk away, because one half of it does not shrink.

The β end does: a rule with more units balances its columns better, and at any fixed basis the counted leftover at R² = 1 goes to zero as the trial grows. That is estimation noise and it behaves like estimation noise.

The α end does not. A rule choosing among assignments to make k quantities small is choosing from a restricted set, and the restriction has a cost in every direction — including directions the constraints say nothing about. That cost is a property of the constraint count relative to the number of units, and it is the subject of the last essay of this field, where it is measured by counting rather than by simulating.

So the honest summary is that the projection gives the slope — the exchange rate between what the basis can reach and what the rule removes — and the finite trial gives the two intercepts, one of which is a nuisance and one of which is a real cost that scales with how much is being asked.

What the projection predicts and what a rule on real units leaves. Each line is one basis: four shapes scored against it, with the closed-form R² on the horizontal axis and the imbalance a rule on 200 units actually leaves on the vertical. The identity says the points should sit on 1 − R², and they sit on a line from α at the left to β at the right instead. Both departures are the finite sample: α is above one because a constraint that removes nothing still competes for the assignments (1.11, 1.21, 0.98, 1.22 as the basis grows), and β is above zero because the rule balances its own columns approximately (0.027, 0.034, 0.049, 0.072). The slope is the geometry and it is the same in every row.
Fig. 6 The same twenty cells in a trial half the size. The lines rotate and the ordering does not move, because the slope is geometry and the intercepts are the sample.
What the projection predicts and what a rule on real units leavesEach line is one basis: four shapes scored against it, with the closed-form R² on the horizontal axis and the imbalance a rule on 400 units actually leaves on the vertical. The identity says the points should sit on 1 − R², and they sit on a line from α at the left to β at the right instead. Both departures are the finite sample: α is above one because a constraint that removes nothing still competes for the assignments (1.05, 1.15, 1.08, 1.14 as the basis grows), and β is above zero because the rule balances its own columns approximately (-0.032, 0.020, 0.024, 0.039). The slope is the geometry and it is the same in every row.00.500100.2000.4000.6000.8001R²(g | span B), a closed formimbalance left in g, relative to a coin'sdashed: 1 − R², the population identity300 trials of 400 units, four bases, five shapestwo fitted numbers per basis and the slope given by the geometry
Fig. 7 Drag the trial size. The point at R² = 1 walks towards zero and the point at R² = 0 does not walk towards one.

What a constraint costs, per constraint

The two intercepts are reported as moving with the size of the basis, and dividing by the two functions that were added says at what rate.

α — the leftover against a shape the basis cannot reach at all — goes from 1.0539 at one function to 1.1361 at three, which is 4.1 points per constraint. β — the leftover in the shape the basis reaches perfectly — goes from −0.0320 to 0.0388, which is 3.5 points per constraint.

The two move at nearly the same rate, and that is the useful form of the field’s competition argument. A constraint costs about four points of a coin’s imbalance in every direction, whether the direction is one the constraint speaks to or one it is silent about. There is no sense in which the cost falls on the shapes the rule was not built for; it falls everywhere, evenly, and only the benefit is selective.

A function pays for itself at four points of R²

Putting the slope and the cost together gives a threshold for the decision the table cannot make.

On the affine line, leftover = α − (α − β)·R², and α − β is about 1.05. So a function that raises a shape’s R² by ΔR² buys 1.05·ΔR² of leftover and costs about 0.04 in α. A function breaks even if it raises R² by four hundredths, and everything above that is profit.

That is a very low bar, and the table’s own increments clear it comfortably. Adding the square to the covariate takes a threshold at 1 from 0.4386 to 0.6579 — a ΔR² of 0.219, five times the break-even. Adding the cube to the pair takes a median split from 0.6366 to 0.7427 — 0.106, two and a half times it.

So the essay’s finding that increments are irregular has a floor under it: they are irregular and they are all large. On this dictionary, every function anybody would think of adding is worth adding, and the reason to stop is not that the next one costs too much but that the assignment space runs out — which is a constraint on the count rather than a judgement about any particular column, and is what the field’s last essay measures.

The one case the threshold rules out is the one it should. A function already in the span raises R² by exactly zero and costs its four points, so adding a recombination of the existing columns is not merely useless — it is the one addition that is strictly harmful, which is the span invariance stated in the currency of a trial rather than as an identity.

Six shapes, and why these six

The dictionary and the shape list are both choices, and a field built on exact arithmetic can still be built on the wrong six numbers, so it is worth saying what was picked and why.

The shapes are the covariate itself, its square, its cube, an indicator at the median, one at 1 and one at 2. Each is standardised to unit variance, so how much the covariate matters is held fixed and only the shape differs — without that, a comparison between shapes is partly a comparison between effect sizes, and every number in the table would move if a shape were rescaled.

The polynomials are there because they are the smooth alternatives to a straight line and because they are mutually orthogonal, which makes the arithmetic legible: a basis of polynomials protects exactly the orders it contains and nothing else. The three indicators are there because a threshold is the shape a mean-balancing rule is least able to see and because it is a shape experimenters actually believe in — a dose that matters above a level, an age at which eligibility changes.

The three cut points are not interchangeable. At the median a threshold is 64% explained by the covariate; at 1 it is 44%; at 2 it is 13%. So they are three quite different problems wearing the same name, and a field that used only one of them would reach a different conclusion about how much a basis has to contain depending on which it picked.

What is deliberately absent is any shape that is not a function of one covariate — an interaction with the treatment, most obviously, which is a different object because it changes what the estimand is rather than how precisely it is estimated.

The six best bases of 4 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 69.8% against every shape in the list, and the worst of the six guarantees 59.0%: the difference between them is entirely which subspace was picked, at the same cost per arrival.
Fig. 8 Quadruples, where the guarantee reaches 69.8% and the best row has picked up an indicator in the tail. Nothing about that choice is available from the size of the basis.

What this buys, in one sentence

Everything in this field after this essay is arithmetic on the table above: choosing a row by its worst cell, asking what happens when the class of shapes is a whole subspace instead of a list, and asking how many columns a rule can be given before the assignment space runs out. None of it needs a trial, because none of it is about a trial.

The one thing the geometry cannot supply is which shapes belong in the columns. That is a statement about the world, it is not derivable from anything here, and which shapes are worth protecting is what happens when somebody declines to make it.

The six best bases of 3 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 59.0% against every shape in the list, and the worst of the six guarantees 35.5%: the difference between them is entirely which subspace was picked, at the same cost per arrival.
Fig. 9 Triples, where the best row is the covariate, its square and its cube — the polynomial basis arriving from a list of shapes rather than from a smoothness argument.

What is claimed here, and what is not

This essay takes the projection identity and the span invariance, and the claims are that two bases with the same span produce identical numbers, that the closed-form projection reproduces the 2/π a median split leaves exactly, and that a rule on real units sits on an affine function of R² whose two intercepts move with the size of the basis.

What stays out and is named as a decision: covariates that are not normal, for which the Hermite functions are no longer orthonormal and the whole Gram matrix has to be recomputed — the identity survives and the closed forms do not; more than one covariate, where the span is a subspace of a product space and the dictionary grows combinatorially; and bases that are not polynomials or indicators, splines in particular, which have closed-form inner products against the normal and were left out because they add a knot-placement decision to a field that already has one decision too many.

The boundary against the shape field is the level. That the rule’s criterion is one over the variance of the treatment estimate in a stated model, that reading three functions costs about two points, and what a rule balanced on the wrong function gives up, are established there — measured by simulation on twelve cells, where every cell is one of the closed forms above computed instead. What is new here is that the criterion depends on the span, which makes the choice a geometric one.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. Two bases spanning the same plane are required to produce identical projections for every shape, to ten decimal places — the identity the whole field rests on, checked rather than asserted, and it would fail immediately if the Gram matrix were being inverted in the wrong basis. And the projection is required to reproduce 2/π at the median to twelve decimal places, in both directions, which ties this field’s machinery to a constant derived independently elsewhere on this site.

The refusal that bears on this essay is a basis reported at the shape it happens to be best against. Every row of the table has a cell at 1.0000 somewhere and a much smaller cell elsewhere, and the number quoted for a design is decided entirely by which cell is chosen — so a rule reported by the shape it was built for is a rule reported by its own assumption.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Allocation ruleAssignment mechanismBasis functionsCovariate balanceDesign criterionHermite polynomialsImbalanceInformation matrixModel misspecificationMonte CarloOrthogonalityProjectionThresholdTreatment effectVariance reduction