A basis is a subspace
Worth reading first: A design is a number · The variance removed before the data.
A rule that balances a column of numbers is balancing every multiple of that column, and a rule that balances two columns is balancing every combination of the two. It is not reading functions; it is reading a subspace, and it cannot tell one basis of that subspace from another.
That sounds like a technicality. It is the thing that turns a question with no answer — which functions should the rule read? — into a question with an exact one, because a subspace has a geometry and a list does not.
The identity
Suppose the outcome depends on some unit-variance function g of the covariate, and the rule balances the span of a basis B perfectly. Then the imbalance it leaves in g is the part of g the span cannot reach:
with R² the squared multiple correlation of g on the span. The rule removes the projection and leaves the residual, and nothing else about g matters.
The one-function case of this is measured in the field that found the exposure, where it is ρ² and reads as a fact about a threshold: balancing the covariate’s mean removes exactly 2/π of the imbalance in a median split of it. It is not a fact about thresholds. It is a fact about projections, and the constant is one entry of a Gram matrix.
Every number here is a closed form
The inner products are exact rather than quadrature, which is worth a paragraph because it is what makes the rest of this field a piece of arithmetic instead of a simulation study.
Take the orthonormal Hermite functions h_j = He_j/√(j!), which are uncorrelated under a standard normal covariate by construction. Then an indicator’s covariance with one of them is
and two indicators’ covariance is Φ̄(max(c, c′)) − Φ̄©Φ̄(c′). So a dictionary of polynomials and cut points has a Gram matrix in closed form, and R²(g | B) = a′G⁻¹a is a solve of a matrix no larger than four by four.
The check that the machinery is right is the constant the neighbouring field derived by a different route. Computed here as a projection, the share of a median split that balancing the covariate removes is 0.6366197724, and 2/π is 0.6366197724. Read the other way — the share of the covariate that balancing a median split removes — it is the same number, which it must be, since a squared correlation is symmetric and the two statements are one integral read in opposite directions.
What the table says that a list of rules could not
With the geometry in hand, the twelve cells the earlier field measured become part of a much larger table, and three things in it are visible only at that size.
Orthogonality is exact and it is common. A rule reading the covariate removes 0.0000 of a quadratic outcome and 0.0000 of a cubic one, because odd and even Hermite functions are uncorrelated under a symmetric covariate. That is not “very little”; it is nothing, and no amount of data changes it.
A threshold’s protection depends entirely on where it sits. A rule reading the covariate removes 0.4386 of a threshold at 1 and 0.1311 of a threshold at 2. The further into the tail, the less a mean-balancing rule can see of it.
And an indicator is a surprisingly good single function. The best single function to hand the rule, judged by its worst shape, is an indicator at 2 — the one function whose worst cell is not zero, at 0.0233. Every polynomial has a shape it is exactly orthogonal to; an indicator has none, because an indicator is not an even or an odd function and so it has a component on every Hermite order.
Adding a function does not add a fixed amount
The table also settles a question the twelve-cell version could not, which is whether “a bigger basis” is a meaningful thing to say.
It is not, in general. A rule reading the covariate and its square removes 0.6579 of a threshold at 1, where the covariate alone removes 0.4386 — a square is not orthogonal to an indicator even though it is orthogonal to the covariate, so protecting against curvature protects partly against a cut. But a rule reading the covariate, its square and its cube removes 0.7427 of a median split where the covariate alone removes 0.6366, an improvement of a tenth for two extra functions.
The increments depend on the angles, and the angles depend on which shape is being protected. There is no ordering of bases by size that survives contact with the table, which is why choosing one by its worst case is a separate problem rather than a matter of adding functions until satisfied.
Why the span, and not the functions, is what a rule can see
The invariance is easy to state and easy to doubt, so it is worth showing where it comes from rather than only that it holds.
A balancing rule maximises the information about the treatment effect given the columns it has been handed: S_aa − S_ax′S_xx⁻¹S_ax, with x the columns. Replace those columns by any invertible linear recombination — x → xA — and S_ax becomes S_ax A, S_xx becomes A′S_xx A, and the correction term becomes
with the A’s cancelling exactly. The criterion is unchanged, so the rule makes the same decision at every arrival, so the trial is the same trial.
That is why the first figure is not a coincidence: a rule handed x and x², and a rule handed x + 2x² and 3x − x², are not two rules that happen to agree. They are one rule described twice.
The practical consequence is worth stating because it cuts both ways. Rescaling, centring or recombining the functions handed to a balancing rule does nothing at all — which is reassuring, and means an experimenter cannot improve a design by standardising its columns. And the only way to change what a rule protects is to change the span, which means adding a function that is already a combination of the others is precisely zero work, however different it looks written down.
What a rule on four hundred units actually does
The identity is a statement about a rule that balances its span perfectly, and no rule does. So the last thing this essay owes is the gap between the geometry and a rule running on real units, and it turns out to have a shape.
The counted leftover is not 1 − R². It is affine in R², running from some α where the basis removes nothing to some β where it removes everything, and both departures are the finite sample:
- α is above one, because a constraint that removes nothing about the outcome still competes for the same assignments. A rule reading the covariate leaves 1.0960 of a coin’s imbalance in a quadratic outcome — it is worse than a coin, by nine points, against a shape it cannot see at all.
- β is above zero, because the rule balances its own columns approximately. A rule reading the covariate leaves 0.0074 of a coin’s imbalance in a linear outcome, and one reading three functions leaves 0.0262.
Both fitted numbers move with the size of the basis in the direction the competition between constraints predicts. Reading three functions rather than one takes β from −0.0320 to 0.0388 — the balance the rule achieves in the shape it was built for gets worse — and α from 1.0539 to 1.1361 — the cost against a shape it cannot see gets larger. Every one of the twenty cells sits on its own basis’s line to within 0.0898.
That is the two-route check for this field, and the two routes share nothing at all. One is a four-by-four solve on a Gram matrix of Hermite functions and indicators; the other is a rule assigning four hundred units one at a time, three hundred times over, and the standard deviation of an imbalance.
Why the affine form and not the identity
It is worth being clear that the gap is not an error term to be shrunk away, because one half of it does not shrink.
The β end does: a rule with more units balances its columns better, and at any fixed basis the counted leftover at R² = 1 goes to zero as the trial grows. That is estimation noise and it behaves like estimation noise.
The α end does not. A rule choosing among assignments to make k quantities small is choosing from a restricted set, and the restriction has a cost in every direction — including directions the constraints say nothing about. That cost is a property of the constraint count relative to the number of units, and it is the subject of the last essay of this field, where it is measured by counting rather than by simulating.
So the honest summary is that the projection gives the slope — the exchange rate between what the basis can reach and what the rule removes — and the finite trial gives the two intercepts, one of which is a nuisance and one of which is a real cost that scales with how much is being asked.
What a constraint costs, per constraint
The two intercepts are reported as moving with the size of the basis, and dividing by the two functions that were added says at what rate.
α — the leftover against a shape the basis cannot reach at all — goes from 1.0539 at one function to 1.1361 at three, which is 4.1 points per constraint. β — the leftover in the shape the basis reaches perfectly — goes from −0.0320 to 0.0388, which is 3.5 points per constraint.
The two move at nearly the same rate, and that is the useful form of the field’s competition argument. A constraint costs about four points of a coin’s imbalance in every direction, whether the direction is one the constraint speaks to or one it is silent about. There is no sense in which the cost falls on the shapes the rule was not built for; it falls everywhere, evenly, and only the benefit is selective.
A function pays for itself at four points of R²
Putting the slope and the cost together gives a threshold for the decision the table cannot make.
On the affine line, leftover = α − (α − β)·R², and α − β is about 1.05. So a function that raises a shape’s R² by ΔR² buys 1.05·ΔR² of leftover and costs about 0.04 in α. A function breaks even if it raises R² by four hundredths, and everything above that is profit.
That is a very low bar, and the table’s own increments clear it comfortably. Adding the square to the covariate takes a threshold at 1 from 0.4386 to 0.6579 — a ΔR² of 0.219, five times the break-even. Adding the cube to the pair takes a median split from 0.6366 to 0.7427 — 0.106, two and a half times it.
So the essay’s finding that increments are irregular has a floor under it: they are irregular and they are all large. On this dictionary, every function anybody would think of adding is worth adding, and the reason to stop is not that the next one costs too much but that the assignment space runs out — which is a constraint on the count rather than a judgement about any particular column, and is what the field’s last essay measures.
The one case the threshold rules out is the one it should. A function already in the span raises R² by exactly zero and costs its four points, so adding a recombination of the existing columns is not merely useless — it is the one addition that is strictly harmful, which is the span invariance stated in the currency of a trial rather than as an identity.
Six shapes, and why these six
The dictionary and the shape list are both choices, and a field built on exact arithmetic can still be built on the wrong six numbers, so it is worth saying what was picked and why.
The shapes are the covariate itself, its square, its cube, an indicator at the median, one at 1 and one at 2. Each is standardised to unit variance, so how much the covariate matters is held fixed and only the shape differs — without that, a comparison between shapes is partly a comparison between effect sizes, and every number in the table would move if a shape were rescaled.
The polynomials are there because they are the smooth alternatives to a straight line and because they are mutually orthogonal, which makes the arithmetic legible: a basis of polynomials protects exactly the orders it contains and nothing else. The three indicators are there because a threshold is the shape a mean-balancing rule is least able to see and because it is a shape experimenters actually believe in — a dose that matters above a level, an age at which eligibility changes.
The three cut points are not interchangeable. At the median a threshold is 64% explained by the covariate; at 1 it is 44%; at 2 it is 13%. So they are three quite different problems wearing the same name, and a field that used only one of them would reach a different conclusion about how much a basis has to contain depending on which it picked.
What is deliberately absent is any shape that is not a function of one covariate — an interaction with the treatment, most obviously, which is a different object because it changes what the estimand is rather than how precisely it is estimated.
What this buys, in one sentence
Everything in this field after this essay is arithmetic on the table above: choosing a row by its worst cell, asking what happens when the class of shapes is a whole subspace instead of a list, and asking how many columns a rule can be given before the assignment space runs out. None of it needs a trial, because none of it is about a trial.
The one thing the geometry cannot supply is which shapes belong in the columns. That is a statement about the world, it is not derivable from anything here, and which shapes are worth protecting is what happens when somebody declines to make it.
What is claimed here, and what is not
This essay takes the projection identity and the span invariance, and the claims are that two bases with the same span produce identical numbers, that the closed-form projection reproduces the 2/π a median split leaves exactly, and that a rule on real units sits on an affine function of R² whose two intercepts move with the size of the basis.
What stays out and is named as a decision: covariates that are not normal, for which the Hermite functions are no longer orthonormal and the whole Gram matrix has to be recomputed — the identity survives and the closed forms do not; more than one covariate, where the span is a subspace of a product space and the dictionary grows combinatorially; and bases that are not polynomials or indicators, splines in particular, which have closed-form inner products against the normal and were left out because they add a knot-placement decision to a field that already has one decision too many.
The boundary against the shape field is the level. That the rule’s criterion is one over the variance of the treatment estimate in a stated model, that reading three functions costs about two points, and what a rule balanced on the wrong function gives up, are established there — measured by simulation on twelve cells, where every cell is one of the closed forms above computed instead. What is new here is that the criterion depends on the span, which makes the choice a geometric one.
The checks, and the refusals that make them mean something
Two claims are gated in this field’s library. Two bases spanning the same plane are required to produce identical projections for every shape, to ten decimal places — the identity the whole field rests on, checked rather than asserted, and it would fail immediately if the Gram matrix were being inverted in the wrong basis. And the projection is required to reproduce 2/π at the median to twelve decimal places, in both directions, which ties this field’s machinery to a constant derived independently elsewhere on this site.
The refusal that bears on this essay is a basis reported at the shape it happens to be best against. Every row of the table has a cell at 1.0000 somewhere and a much smaller cell elsewhere, and the number quoted for a design is decided entirely by which cell is chosen — so a rule reported by the shape it was built for is a rule reported by its own assumption.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A zero that was an assumption — both name basis functions, covariate balance, design criterion, hermite polynomials, imbalance, orthogonality, projection
- The arcsine that closes it, and the error that was overstated — both name basis functions, covariate balance, hermite polynomials, monte carlo, orthogonality, projection, threshold
- A dictionary that is neither — both name basis functions, covariate balance, hermite polynomials, model misspecification, orthogonality, projection
- The zero that survives a cut — both name basis functions, covariate balance, hermite polynomials, orthogonality, projection, threshold
- A probe chosen from the design — both name assignment mechanism, basis functions, covariate balance, orthogonality, projection
- A probe nobody chose — both name assignment mechanism, covariate balance, imbalance, projection, treatment effect
Named objects
A flat tag is an object no other essay names yet.
Allocation ruleAssignment mechanismBasis functionsCovariate balanceDesign criterionHermite polynomialsImbalanceInformation matrixModel misspecificationMonte CarloOrthogonalityProjectionThresholdTreatment effectVariance reduction