When the set is too large to walk

A dictionary that is a product

Two covariates make what a balancing rule may read an outer product — eight main effects and sixteen interactions — and every inner product in it is still closed form. What a rule holding all eight main effects removes of a pure interaction is not small. It is zero.

Worth reading first: The variance removed before the data · A design is a number.

Choosing what a balancing rule reads is exact geometry, and every projection in it lives on one covariate. The Hermite functions are orthonormal in L²(φ), the inner product between any two entries of an eight-function dictionary is a closed form, and the best basis of size k is found by walking all C(8, k) subsets of it. It is a satisfying field precisely because nothing in it is estimated.

Real trials record more than one covariate. Age and baseline severity; income and household size; temperature and humidity. The obvious question is whether any of that survives a second covariate, and the answer has two halves that point in opposite directions: the geometry survives completely intact, and something that the one-covariate case could not exhibit appears immediately.

The dictionary is an outer product

Let the two covariates be independent standard normals. Every term a rule may be handed is f(x₁)·g(x₂), where each factor is either an entry of the one-dimensional dictionary or absent. A term with one factor absent is a main effect; a term with neither absent is an interaction. There is no third case.

Two covariates make the dictionary an outer product. Four functions of each covariate, and everything a balancing rule may be handed. The margins are the 8 main effects and the block between them is the 16 interactions, which are 66.7% of the dictionary. Every inner product in it is closed form — ⟨f₁g₁, f₂g₂⟩ = ⟨f₁,f₂⟩⟨g₁,g₂⟩ when the covariates are independent — so nothing about the geometry gets harder. What gets harder is the counting: choosing k of 24 is C(24, k), which is 10,626 at four and 735,471 at eight.
Fig. 1 Four functions of each covariate, and everything a balancing rule may be handed. The margins are the main effects; the block between them is every product of one from each.

Four functions per covariate gives 8 main effects and 16 interactions — 24 terms, of which the interactions are 66.7%. With d functions per covariate it is 2d and d², so the interactions overtake the main effects at d = 3 and the whole dictionary grows quadratically where the number of covariates grows by one.

And the arithmetic does not get harder at all. For independent covariates,

⟨f₁g₁, f₂g₂⟩ = ⟨f₁, f₂⟩ · ⟨g₁, g₂⟩

exactly, because the expectation of a product of independent things factors. Every entry of a twenty-four by twenty-four Gram matrix is therefore a product of two one-dimensional inner products that were already closed forms, and what a rule removes of a shape is the same expression it always was — the shape’s squared multiple correlation on the span of what the rule reads.

That identity is checked rather than trusted. Three hundred thousand draws of a pair of covariates, every one of the three hundred distinct pairs of terms evaluated directly, and the largest disagreement with the closed form is 2.33 standard errors. Each pair is measured in its own standard error rather than against one tolerance, because the product of two fourth-order Hermite functions has a spread the linear terms do not and a single tolerance would be either vacuous or a false alarm.

Exactly nothing

Now the half that is new, and it is not a matter of degree.

A rule holding every main effect removes none of an interaction. What a balancing rule handed all 8 main effects of two covariates removes of each shape, as a share of that shape's variance. The three shapes that are functions of one covariate are removed entirely. The two that are products of a centred function of each are removed exactly — not approximately — nothing, because E[f(x₁)·f′(x₁)g′(x₂)] factors into E[ff′]E[g′] and the second factor is zero. And the mixed shape sits at two thirds, which is the share of its variance that is not in its interaction term: expanding 1{x₁>0}1{x₂>0} gives one interaction and two main effects at equal weight.
Fig. 2 What a rule handed all eight main effects of both covariates removes of each shape. Three of the answers are one, two are zero, and one is two thirds.

Hand a rule every main effect of both covariates — all eight of them — and ask what it removes of an outcome shaped like x₁·x₂, or like x₁ above the second covariate’s median. The answer is 0.000000. Not small, not asymptotically negligible, not a few per cent: zero, at machine precision, and it stays zero however many main effects the rule is given.

The reason is one line. For a centred f′ and a centred g′,

E[ f(x₁) · f′(x₁) g′(x₂) ] = E[ f f′ ] · E[ g′ ] = 0

because the second factor is the mean of a centred function. Every main effect is orthogonal to every pure interaction, by independence, and a projection onto a span of things all orthogonal to a target is the zero projection.

This is the two-covariate version of the exactly-zero guarantee the one-covariate field arrives at, and it arrives for an entirely different reason. There the zero was a dimension count — a basis with fewer functions than the class has dimensions has a principal angle of ninety degrees somewhere. Here it is independence, and it does not care how many functions the rule has.

The shape that is two thirds

One row of that figure is neither zero nor one, and it is the one worth knowing, because it is what a real interaction usually looks like.

Take the outcome to depend on both covariates being above their medians — a corner, a joint eligibility, a compound condition. That sounds like a pure interaction and is not. Writing each uncentred indicator as its centred version plus a half and multiplying out gives an interaction term plus one half of each main effect, and once every part is scaled to unit variance the three parts arrive with equal weight. So the shape has three orthogonal pieces, two of which are main effects.

A rule holding both median splits removes exactly two thirds of it. Not approximately: the number is 2/3 to machine precision, and it falls out of the expansion rather than being derived.

That is the useful case to carry, because it says what happens between the two extremes. A shape that is partly an interaction is protected exactly in the proportion that it is not one, and the proportion is arithmetic on the shape rather than a property of the rule. An experimenter who believes their outcome is interactive is not thereby unprotected; they are unprotected in the share of the outcome’s variance that lies in the interaction terms, and that share is computable from the shape they believe in.

What the zero is not

Two readings of the exact zero are wrong and both are tempting, so it is worth ruling them out before anything is built on it.

It is not a statement that balancing is useless against interactive outcomes. It is a statement about the pure interaction component. Every outcome anybody writes down has main-effect components too, and those are removed entirely; the corner shape is the worked case, at two thirds. What the zero says is that the interaction component is untouched, and how much of the outcome that is depends on the outcome.

And it is not a small-sample phenomenon. There is no n in the argument. E[f f′]·E[g′] = 0 because g′ is centred, and a rule handed a thousand main effects on ten thousand units removes exactly as much of x₁x₂ as one handed two on sixteen — which is nothing. This is worth saying because most of the disappointing results in the surrounding fields are asymptotic ones that improve with the trial size, and this one does not improve at all.

The practical form is a question an experimenter can answer before randomising: what share of the outcome’s variance is believed to lie in terms that are products across covariates? If the answer is none, main effects are the whole story and this essay changes nothing. If the answer is a third — the corner shape — then a rule reading only main effects leaves a third of the imbalance in place, and no amount of adding more main effects will move it.

What a third covariate does to the same counts

Two covariates make the dictionary a grid. Three make it a lattice, and the arithmetic is the same expression with one more exponent in it.

With p covariates and d functions apiece — counting “absent” as a fifth choice on each axis — the whole dictionary is (d+1)p1(d+1)^p - 1 terms. At d = 4 that is 24 at two covariates, 124 at three and 624 at four.

The main effects are always pd: 8, 12 and 16. So their share of the dictionary runs

  • one covariate: 100%
  • two: 33.3%
  • three: 9.7%
  • four: 2.6%

By three covariates, nine terms in ten a balancing rule might be handed are interactions, and by four it is thirty-eight in thirty-nine. The one-covariate case is not a simplification of the general one; it is the single point at which the phenomenon this essay is about does not exist.

The geometry stays free and the search does not

The closed form for an inner product survives all of that untouched, so the Gram matrix costs (d+1)2p(d+1)^{2p} products of numbers already computed: 576 entries at two covariates, 15,376 at three, 389,376 at four. All of them trivial.

What does not survive is the part of the one-covariate field that made it satisfying — walking every subset.

Choosing the best four terms out of eight is (84)=70\binom{8}{4} = 70 subsets, which is a loop. Out of twenty-four it is 10,626, which is still a loop. Out of a hundred and twenty-four it is 9.4 million, which is a few seconds. Out of six hundred and twenty-four it is 4.2 billion, and choosing six out of a hundred and twenty-four is already 4.5 billion.

So the boundary is not where the geometry becomes approximate — it never does — but where the exhaustive search stops fitting inside a session. Two covariates and four functions each is the last configuration where the best basis of any size can be found by looking at all of them, and beyond it the field’s central move has to be replaced by a search that can be wrong.

That is worth naming as the real cost of a second covariate, because the essay’s own headline is that nothing about the arithmetic gets harder. Nothing does. What gets harder is proving that the basis chosen is the best one, and that is the claim the one-covariate field could make and this one cannot.

What correlated covariates give away

The construction above assumes the two covariates are independent, which is the assumption most likely to be objected to, so it is worth pricing.

For a standard bivariate normal of correlation ρ, Mehler’s formula gives the exact answer: E[h_i(x₁) h_j(x₂)] = ρ^i when i = j and zero otherwise. The orthonormal Hermite functions stay uncorrelated across orders and pick up ρ^i within an order.

What correlated covariates give a rule for nothingA rule that balances the jth Hermite function of one covariate removes ρ²ʲ of the same function of the other, exactly, by Mehler's formula — the orthonormal Hermite functions stay uncorrelated across orders and pick up ρ^j within one. At a correlation of 0.5 that is 25.0% of the second covariate's linear shape and 0.39% of its fourth power. The gift is real and it is only ever about the smoothest term there is, so a rule that reads one covariate is not quietly reading the other.Hermite function 125.00%Hermite function 26.25%Hermite function 31.56%Hermite function 40.39%removed of the second covariate, at a correlation of 0.5ρ²ʲ, from Mehler's formula25.0% down to 0.39%
Fig. 3 What a rule balancing one covariate removes of the same function of the other, at a stated correlation, with the correlation on a slider. The gift dies geometrically in the order.

So a rule that balances the jth Hermite function of one covariate removes ρ²ʲ of the same function of the other, for nothing. At ρ = 0.5 that is 25.000% of the second covariate’s linear shape, 6.250% of its square, 1.563% of its cube and 0.391% of its fourth power. Checked against four hundred thousand draws, which is the second route.

Two things follow. The free protection is real and it is only ever about the smoothest term there is — so a rule that reads one covariate is not quietly reading the other except in the one place where it hardly needed to. And the correlation buys nothing at all for the interactions, because the argument above concerned functions of the same order in two coordinates, and an interaction is a product rather than a coordinate. The zero survives correlation.

What the interactions cost to compute, which is nothing

One small consequence of the product identity is worth drawing out, because it is what makes the rest of this field possible at all.

A balancing rule needs, for every basis it might use, the Gram matrix of that basis and the covariances of each candidate outcome shape with it. In one covariate those were eight-by-eight matrices of closed forms. In two they are twenty-four-by-twenty-four, and every entry is a product of two numbers already in the one-dimensional table. Nothing has to be integrated, nothing has to be simulated, and adding a third covariate would multiply again rather than requiring anything new.

That is the difference between this construction and the obvious alternative, which is to estimate the Gram matrix from the units in hand. An estimated Gram matrix on twenty-four terms and two hundred units is a twenty-four-dimensional covariance estimated on two hundred rows, which is noisy enough that the projection it produces is a fact about the sample; and worse, it makes what a rule removes a random quantity, so the guarantee attached to a basis stops being a guarantee.

The whole field turns on the covariates having a known distribution. That is a strong assumption and it is stated rather than buried: standard normal covariates, independent unless a correlation is named. In a trial the covariates’ distribution is not known, and the honest reading of everything here is a statement about a stated population rather than about a realised sample. The one place that assumption is checked against reality is the trials, where a rule on four hundred actual units is shown sitting on the line the projection draws — and that check belongs to the one-covariate field, where it was run.

Why this is not an argument for reading everything

The obvious response to a zero is to hand the rule the interactions as well, and it is worth saying why that is a decision rather than an obvious improvement.

Every function a rule is told to balance costs it randomisation. The constrained set of assignments gets smaller, the reference distribution for any randomisation test gets coarser, and past some number of constraints the rule stops being able to randomise at all — which the one-covariate field measures by walking every split of sixteen units and finding the admissible count falling to nothing.

So balance the interactions too is not free, and the size of the bill is what the third essay of this field is about. It turns out to be much smaller than the sixteen-unit measurement suggests, for a reason that reverses the extrapolation entirely — but the decision still has to be taken, and it has to be taken against a statement about which shapes are worth protecting.

Which is the other cost. A dictionary of twenty-four terms has 16,777,216 subsets, and choosing among them by any criterion means saying what the criterion is protecting against. The next essay is about what happens to that choice when the subsets can no longer be walked.

The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.
Fig. 4 The one-covariate version of the same projection, where a rule balancing the mean of a covariate removes exactly 2/π of a median split of it. Every number in it survives the second covariate unchanged.
The first purchase is an interaction, and it is not optional. What the best basis of each size removes of the shape it is worst against. One function and two guarantee exactly zero, because two of the seven shapes are pure interactions and no combination of main effects touches them at all — so the answer to what should the rule read is settled by the list of shapes before any optimisation happens. The smallest basis with a guarantee is a.x + a.x2 + x*cut0, at 42.44%, and the interaction in it is doing the work no main effect can. Hollow points are sizes the walk can still confirm; filled ones are the exchange algorithm alone.
Fig. 5 What each basis size guarantees against a list of shapes containing two pure interactions: exactly nothing until an interaction term is bought.

The counting, which is the only thing that gets worse

It is worth being precise about what a second covariate costs, since the answer so far has been nothing.

The geometry costs nothing: same expression, same closed forms, one multiplication.

The counting costs a great deal. Choosing k functions from an eight-entry dictionary is C(8, k), which is at most 70. Choosing k from twenty-four is C(24, k), which is 2,024 at three, 10,626 at four, 134,596 at six and 735,471 at eight. The whole power set is 16,777,216.

And the counting is not incidental — it is how the one-covariate field answers its central question. The best basis of a given size against a stated list of shapes is found there by scoring every subset, which is an exact optimisation with an exact answer and no algorithm in it. That method does not survive here, and what replaces it has a warrant of a completely different kind.

There is a third cost that is easy to miss and is the reason the interactions cannot simply be added to everything. Each function a rule balances is a constraint on the assignment, and constraints are what the randomisation is spent on. Sixteen interactions are sixteen constraints, and a rule with twenty-four constraints on a small trial has, as the one-covariate field shows by direct enumeration, no assignments left at all. Whether that survives a real trial size is a separate question with a surprising answer, and it is the second half of this field.

What is claimed here, and what is not

This essay takes the balancing geometry over two covariates, and the claims are the product identity for the inner products, the exact zero against a pure interaction, the exact two thirds against a shape that is one third interaction, and the ρ²ʲ that correlated covariates hand a rule for nothing.

What stays out and is named as a decision: choosing a basis from the dictionary, which is the next essay; what any of it costs in randomisation, which is the third; and the whole of the interaction geometry for correlated covariates, which needs a fourth-order expectation Mehler’s formula does not supply and is not attempted here.

The boundary against the one-covariate field is the second coordinate. That a balancing rule cannot tell one basis from another with the same span, that what it removes is a squared multiple correlation, and the dictionary of polynomials and cut points are all established there and used here without being re-derived.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. Every one of the three hundred pairs of dictionary terms is required to match ⟨f₁,f₂⟩⟨g₁,g₂⟩ on a sample that computes the covariance directly, each within four of its own standard errors — which fails if the product identity were being assumed rather than used. And the projection onto every main effect is required to be exactly zero against each pure interaction, exactly two thirds against the corner, and above 0.999 against each shape that is a function of one covariate, so that the zero is a statement about the shape rather than about the basis being too small.

The refusals for this field belong to its later essays, and the reason is worth naming: nothing in this one is a procedure. It is a Gram matrix and a projection, both exact, and the only way to be wrong about them is to compute them wrongly — which is what the second route is for.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AllocationBasisBlockingClosed formCorrelationCovariateCovariate balanceGram matrixHermite polynomialsInner productInteractionOrthogonalityProjectionRandomisationVariance explained