A zero that was an assumption
Worth reading first: A design is a number · The variance removed before the data.
The result reads like a limitation of balancing rules and is quoted as one. Hand a rerandomisation every main effect of both covariates — four functions of the first, four of the second, eight constraints — and ask what it removes of an outcome that depends on the product of them. The answer is 0.000000, at machine precision, at any number of main effects.
It is correct. The construction is a projection, the projection is onto the span of the main effects, and a pure interaction is orthogonal to every one of them — because for independent covariates E[f(X)g(Y)·h(X)] factorises into E[f(X)h(X)]·E[g(Y)], and the second factor is the mean of a centred function, which is zero.
Every step of that uses independence. Take it away and there is nothing left of the argument.
What the interaction becomes
Write the leading case out. Let the interaction be h₁(x)h₁(y), the product of the two covariates, each standardised. The linearisation says h₁h₁ = h₀ + √2·h₂, so the square of a covariate is what a product of two copies of it contains.
Now Mehler pairs h_i(X) with h_i(Y) at strength ρ^i, so:
The interaction has a component along the square of the first covariate, of size √2·ρ, and by symmetry the same along the square of the second. It has none along the covariates themselves — E[h₁(X)²h₁(Y)] is zero for the same reason it always was — so the component is second-order, not first.
Three inner products finish it. The two squares are not orthogonal to each other either, since E[h₂(X)h₂(Y)] = ρ², so the two-by-two solve gives a quadratic form of 4ρ²/(1 + ρ²). What divides it is the interaction’s own variance, and that is where the care is needed: its second moment is 1 + 2ρ² and its mean is ρ, because a product of two centred functions is not centred once the covariates are dependent. The variance is 1 + ρ², and
is the share of the interaction that a rule holding both squares removes. The six other main effects contribute nothing, which is a statement the closed derivation makes and the general routine — which builds an eight-by-eight Gram matrix and solves it without being told — has to agree with.
They agree to four parts in a thousand at every correlation measured.
The shape of the curve
The expression has two features that are worth more than the expression.
It is quadratic at the origin. At small ρ the removed share is about 4ρ², so at ρ = 0.1 it is 3.92% and at ρ = 0.2 it is 14.79%. That is what makes the zero look robust: a mildly correlated pair of covariates really does leave a rule almost unable to touch their interaction, and somebody checking the result at a plausible correlation would find it nearly holding.
It is monotone, and it reaches one. At ρ = 0.9 the rule removes 98.90% of the interaction, and the endpoint is worth deriving because it is the sanity check on everything else: at ρ = 1 the two covariates are the same variable, so the interaction h₁(x)h₁(y) is h₁(x)², which is h₀ + √2·h₂ — a constant plus something the span of the main effects contains. A constant is not variance, so there is nothing outside the span at all and the removed share is exactly 1.
That last sentence is the correction this expression carries. Until 2026-08-28 the denominator here was the interaction’s second moment rather than its variance, and this paragraph read that the rule “removes the √2·h₂ part and leaves the constant — two units of three in squared norm, which is two thirds”. The derivation was right and its conclusion was not: a constant left behind is not a share of anything left unexplained, because a constant is the same in both arms of the trial and contributes no imbalance whatever. The expression that follows from the corrected denominator rises to one at ρ = 1 rather than falling back to two thirds, and its values are between one and forty per cent larger than the ones this collection quoted for two days.
So the worst case for an experimenter is the most correlated design, and it is a limiting case rather than an interior one: the two covariates become the same variable and the interaction becomes a function the rule already holds.
There is a third feature, and it belongs to the derivation rather than to the curve. Every number above comes from three inner products — the interaction against a square, the two squares against each other, and the interaction against itself — and the six remaining main effects contribute exactly nothing. That is not an approximation made to keep the expression short: the general routine solves an eight-by-eight system without being told, and agrees. So the closed form is a statement that six of the eight constraints are doing no work at all against this shape, which is a sharper claim than the number it produces.
Which interactions lose most
The product of the two covariates is the mildest case. Three others measured on the same axis:
| interaction | ρ = 0.2 | ρ = 0.4 | ρ = 0.5 | ρ = 0.8 |
|---|---|---|---|---|
| one covariate times the other | 14.79% | 47.56% | 64.00% | 95.18% |
| one times the other’s square | 17.62% | 52.10% | 67.86% | 95.95% |
| the two squares | 3.30% | 27.10% | 45.46% | 91.61% |
| one times the other’s cube | 22.60% | 57.73% | 71.41% | 95.94% |
The pattern is that asymmetric interactions lose most. A term mixing orders — one covariate times the other’s square or cube — has a component along a third or fourth power of a single covariate, and the dictionary contains those, so a rule holding all eight main effects reaches most of it by ρ = 0.8. The symmetric ones are protected for longer because the orders they linearise into are lower and fewer.
And the two squares is the slowest of the four at small correlation — 3.30% at ρ = 0.2 against 14.79% for the plain product — because h₂h₂ contains h₀, h₂ and h₄, and its overlap with the main effects starts at ρ² through the h₂ term with a smaller coefficient. It catches up by ρ = 0.8.
None of them keeps the zero.
Why the second order and not the first
There is a detail in the derivation that decides the whole shape of the result, and it is worth isolating: the interaction acquires a component along the covariates’ squares and none along the covariates themselves.
The reason is parity. The linearisation of h₁h₁ contains only even orders — h₀ and h₂ — because a product of two functions of orders i and k contains orders between |i − k| and i + k with the same parity as i + k. So a product of two first-order terms has no first-order part to be paired with anything, and Mehler pairs like with like.
That has two consequences an experimenter can act on.
A rule reading only the covariates themselves removes nothing of their product, at any correlation. The zero survives for a linear-only rule. It is the squares that break it, and a rule handed the first four powers of each covariate contains them.
And the leading term is ρ² rather than ρ. That is why the curve is flat near the origin and why the result looks robust at the correlations most people would test it at. At ρ = 0.1 a rule holding both squares removes 3.92% of the product; at ρ = 0.3 it removes 30.3%. The crossing from negligible to most of it happens over a range of correlation that no design report would think worth distinguishing.
Half the interaction goes at ρ = √2 − 1
The curve’s two landmarks are worth solving for, because both land on ordinary numbers.
Setting 4ρ²/(1 + ρ²)² = ½ gives, with u = ρ², the quadratic u² − 6u + 1 = 0, whose root inside the unit interval is u = 3 − √8. So the correlation at which a rule holding both squares removes half of the plain interaction is
ρ = √2 − 1 = 0.4142
exactly — the same constant that governs the weights of a subset-optimal design elsewhere on this site, arriving here from a different problem.
And the curve is steepest at ρ ≈ 0.36, where its slope is 1.74 per unit of correlation. So around there a change of a tenth in the correlation between two covariates moves the removed share by seventeen points.
Both numbers say the same uncomfortable thing about where the transition sits. It is not at a correlation somebody would describe as strong. It is between about 0.25 and 0.55, which is the range covering most pairs of covariates anybody records on the same units — age and a baseline score, two laboratory values, height and weight — and it is the range in which a design report would say “the covariates are mildly correlated” and stop.
The four interactions agree exactly where it matters
The table’s ordering is a finding at one end of the range and nothing at the other.
At ρ = 0.2 the four removed shares are 3.30%, 14.79%, 17.62% and 22.60% — a factor of 6.8 between the smallest and the largest. At ρ = 0.8 they are 91.61%, 95.18%, 95.94% and 95.95% — a factor of 1.05.
So which interaction an experimenter is worried about matters a great deal at correlations where the answer is the rule removes almost none of it, and does not matter at all at correlations where the answer is the rule removes almost all of it.
There is no regime in which the choice is both consequential and unresolved. At a low correlation the guarantee is nearly intact whichever interaction is feared, so the ordering among the four is a distinction between “3% removed” and “23% removed” and both are close enough to the zero the field started from. At a high one the ordering has collapsed and every interaction is reached.
That is worth stating because the table invites a design decision — which interaction to protect against — and the arithmetic says there is not one to make. What there is to decide is whether the covariates are correlated enough to be in the second regime, and the previous section says that question has to be answered to within about a tenth.
What this is a claim about
It is worth being exact about which sentence is wrong, because the underlying result is not.
Correct: a balancing rule handed a set of functions removes, from any shape, that shape’s squared multiple correlation on the span of what it was handed. This is a projection identity and holds at every correlation.
Correct: when the two covariates are independent, a pure interaction is orthogonal to every main effect of either, so the projection is zero and the rule removes none of it.
Wrong when carried across: a balancing rule cannot protect against an interaction. That sentence drops the premise. On a trial whose covariates are correlated at a half, a rule holding both squares removes nearly two thirds of the plain product and over two thirds of the mixed term — and it does so without having been handed any interaction at all.
The direction of the error is the one that matters. Quoting the zero on a correlated design understates what the rule does, so it understates the protection and overstates what is left for the outcome to depend on. That is the safe direction for a warning and the wrong direction for a design calculation: an experimenter budgeting for residual imbalance in an interaction, and using the zero, is budgeting for something that has already been half removed.
What it does to a basis
The practical consequence is not about interactions at all. It is about which functions to hand the rule.
Choosing a basis is a finite problem: score every subset of a given size by its worst guarantee over a list of shapes worth protecting, and take the best. Every score in that search is one of the projections above, and every one of them moves with ρ. So the argmin moves too, and a basis chosen by solving the problem at ρ = 0 is solving a different problem from the one a correlated trial poses.
That is not a violation of anything. The theorem that a rule cannot tell one basis from another with the same span still holds at every correlation — it is a statement about projections and subspaces, and correlation does not touch it. What moves is every number the theorem’s machinery produces.
A stable theorem with unstable outputs is the most dangerous combination there is, because the result keeps being quotable while the numbers behind it go stale, and nothing in the quoting says which correlation the numbers came from.
The free protection is the mirror image and is worth having beside it. Mehler pairs the jth function of one covariate with the jth of the other, so balancing one removes ρ²ʲ of the other’s — 25.000% of the linear term at ρ = 0.5, 6.250% of the quadratic, 1.563% of the cubic, 0.391% of the fourth. A rule reading only the first covariate is doing a quarter of the second covariate’s linear work by accident, and almost none of its curvature.
So a correlation gives with one hand and takes with the other, and both are geometric: what a rule gets free on the second covariate falls as ρ²ʲ, and what it inadvertently removes of an interaction rises as ρ². Neither is a design decision. Both are properties of the covariates that exist before anybody chooses a rule.
Two sentences that are not the same
The essay ends on a distinction that is easy to lose and is the whole practical content.
A rule cannot protect against an interaction it was not handed — false at any correlation, by the numbers above.
A rule cannot protect against an interaction it was not handed, on independent covariates — true, and the projection is exactly zero.
Neither sentence is about the rule’s cleverness. Both are about geometry, and the second one has a premise in it that a real trial does not satisfy. The covariates a trial balances on are age and baseline severity, height and weight, income and education; they are correlated by construction, because they were chosen for being prognostic and the things that predict an outcome tend to predict each other.
So the honest form for a design report is a number rather than either sentence: at the correlation these covariates actually have, a rule holding both squares removes this share of their product, and this share of the mixed terms. All of those are closed expressions in ρ, none of them needs a simulation, and the correlation is one of the few quantities an experimenter genuinely knows before randomising, because it is a property of the units that have already been recruited.
What is claimed here, and what is not
This essay takes what a balancing rule removes of an interaction once the covariates are correlated, and the claims are that the share is 4ρ²/(1 + ρ²)² for the leading case by two independent routes, that it is quadratic at the origin and rises monotonically to exactly one at perfect correlation, that asymmetric interactions lose most and reach 95.95% by ρ = 0.8, and that the free protection on the second covariate dies as ρ²ʲ.
What stays out and is named as a decision: what any of it does to the variance of a treatment effect in a real trial, since that requires a stated outcome model and this is geometry; thresholds, whose expansions do not terminate; and whether a basis ought to be chosen at an estimated correlation, which is a design question with the same shape as every other estimated-nuisance problem in this collection and is not settled here.
The boundary against the field that built the dictionary is that its zero is exact and this essay is about the premise it is exact under. The boundary against the essay on the shape a covariate enters by is that it varies the outcome’s shape with one covariate and this one varies the covariates’ relationship with the shape held.
The checks, and the refusals that make them mean something
Two claims are gated in this field’s library. The general routine is required to reproduce the closed expression at four correlations, which is two routes to one number sharing no arithmetic. And the removed share is required to be exactly zero at ρ = 0 and above a third at ρ = 0.5 — the first because it is the result being qualified, the second because a qualification that amounted to a rounding error would not be one.
The refusal for this essay is the zero quoted for a trial with correlated covariates. It is not a wrong result being repeated; it is a right result carried past its premise, and it says a rule protects nothing where in fact it removes a third by a correlation of a half.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A basis is a subspace — both name basis functions, covariate balance, design criterion, hermite polynomials, imbalance, orthogonality, projection
- A zero that rests on a symmetry — both name basis functions, covariate adjustment, covariate balance, hermite polynomials, interaction, orthogonality, projection
- The arcsine that closes it, and the error that was overstated — both name basis functions, closed form, correlation, covariate balance, hermite polynomials, orthogonality, projection
- A model and a count — both name basis, closed form, correlation, covariate balance, imbalance, rerandomisation
- A quantity that loses to a heuristic — both name basis, covariate balance, imbalance, orthogonality, projection, rerandomisation
- A split survives what a mean does not — both name basis functions, covariate adjustment, covariate balance, interaction, orthogonality, rerandomisation
Named objects
A flat tag is an object no other essay names yet.
BasisBasis functionsClosed formCorrelationCovariate adjustmentCovariate balanceDesign criterionHermite polynomialsImbalanceInteractionMaximin designMehler's formulaOrthogonalityProjectionRerandomisation