A dictionary that is neither
Worth reading first: A design is a number · The variance removed before the data.
Two results in this collection say the same thing about different dictionaries, and each has its own explanation.
A rule handed both covariates removes exactly none of their product when the covariates are independent, and once they are correlated it removes a great deal — because the product has a component along the squares, and the dictionary contains those.
A rule handed two median splits removes exactly none of their product at every correlation — because a two-valued function squares to a constant, so the product of two of them is orthogonal to each of its own factors whatever ρ is.
Two arguments, each with a hole the other fills. The polynomial one fails for cuts, because a dictionary of cuts does not contain the products a correlation generates. The constant one fails for powers, because x² is not constant. And the dictionary a trial actually holds — the mean of a covariate and the count above a threshold — is neither case: its span contains the cut, so the constant argument is unavailable, and it contains no higher power of the cut, so the polynomial argument is unavailable too.
The rule is parity
Consider the joint sign flip (X, Y) → (−X, −Y). A bivariate normal is invariant under it at every correlation, so the flip is a symmetry of the whole problem rather than of one case.
A function that changes sign under the flip is orthogonal to one that does not. A product of two odd functions is even. A main effect is a function of one covariate, and every dictionary entry a trial normally holds — the mean, the median split — is odd.
So a rule holding odd functions removes exactly nothing of an interaction between two odd functions, at every correlation, and neither of the two arguments above is needed. Both are special cases of it.
The second row of that picture is what makes it a rule rather than a restatement. A dictionary of squares removes 64.00% of the product of the two covariates and exactly nothing of one covariate times the other’s square — because that mixed term is odd, and the squares are even. The zero appears twice, on opposite sides, and neither the constant argument nor the polynomial one predicts the second occurrence.
And the dictionary a trial holds is odd
A rule holding the mean and the median split of each covariate — four functions, both of the things anybody balances — removes exactly none of the product of the two median splits, and exactly none of the product of the two covariates, at every correlation. Adding the second odd function to the first bought nothing at all against either.
That is the answer to the question this field was opened for, and it is not the answer the two existing arguments suggested. The span does contain the cut, so the constant argument really is unavailable; it does contain no higher power, so the polynomial argument really is unavailable; and the removed share is zero anyway, for a reason neither of them names.
A cut away from the median is the case that separates parity from shape. The median split is exactly odd; a cut at one standard deviation is 59.4% odd and the rest even, and a rule holding two of them removes 22.49% of their interaction at ρ = 0.5. It is the even part doing all of it, and the leak grows as the cut moves out because the even part does: at two standard deviations the cut is 48.8% even.
That also settles what adding the covariate to a dictionary of cuts is worth, which is the question the field was opened for. At a cut of one standard deviation, holding the two cuts alone removes 22.49% of their interaction and holding the covariates as well removes 23.22% — seven tenths of a point, from four extra constraints, because the covariates are odd and only the cuts’ own even parts were ever available.
What parity covers that the other two arguments do not
The three explanations are worth counting against the zeros they account for, because the case for replacing two arguments with one is a case about coverage.
There are four combinations of a dictionary’s parity and a target’s, and parity settles two of them outright. An odd dictionary against an even target is exactly zero. An even dictionary against an odd target is exactly zero. The other two diagonal cells are where every number in this collection lives — 64.00% for squares against a product, and everything the mean rule removes of anything.
Set the earlier arguments against that. The polynomial argument explains one cell and only when the dictionary happens to contain the products a correlation generates; it says nothing about cuts, because a dictionary of cuts contains no such products. The constant argument explains one cell and only for two-valued functions; it says nothing about powers, because a square is not a constant. Neither of them says anything at all about the second row of the figure — a dictionary of squares against a covariate times the other’s square — which is zero for the same reason and is a case both arguments are silent on.
So parity accounts for three zeros where each of the earlier arguments accounts for one, and it does it without asking what is in the dictionary beyond whether every entry changes sign.
What the zero is actually resting on, and what breaks it
Parity is a property of the law, not of the rule, and saying which is what makes the guarantee portable or not.
The flip (X, Y) → (−X, −Y) is a symmetry of a bivariate normal at every correlation, which is why the zeros in the figure are flat lines rather than curves through zero at one value of ρ. That flatness is the signature: a zero that holds because of a symmetry holds identically in every parameter the symmetry does not touch, and a zero that holds because two quantities cancel crosses through zero at one point and is a coordinate.
And the symmetry is exactly what a skewed covariate removes. Once the marginal is not symmetric, does not have the law of , the flip is no longer a symmetry of the problem, and the argument has nothing left. The size of what is lost is measured elsewhere in this collection: a rule holding the mean of each covariate removes exactly none of their product under a symmetric marginal and 11.63% under one skewed at 0.75.
So the guarantee this essay establishes is worth 11.63 points and is lost at a skew of 0.75. That is not an argument against stating it — it is the correct statement of what an experimenter has, which is a zero conditional on a symmetry they can check rather than a zero they can rely on.
Every inner product it needs, in one dimension
The geometry the rule needs is closed, and the reduction is one line. For any one-variable f, g and u,
because the conditional expectation of g(Y) given X is Σ_j ⟨g, h_j⟩ρ^j h_j(X). Every term on the right is a one-dimensional expectation of a piecewise smooth function against the normal density, which is a quadrature that a computer answers to the last bit a double carries, and the sum converges geometrically in ρ whatever the functions are.
So a dictionary containing cuts, powers and any combination of them has an exact Gram matrix and exact covariances at any correlation, with nothing truncated and nothing simulated. That is the same move the cut-point geometry makes for a pair of functions, one order up.
Two things check it. The general form is required to reproduce, to twelve places, the hand-derived closed form for two cuts — 22.494% against 22.494% at ρ = 0.5 — which is a comparison between an expression written for one case and a routine that knows nothing about that case. And every entry is checked against draws, which share no arithmetic with either.
One of the two checks nearly did not fire. The zeroth Hermite coefficient of a centred function is zero by definition, and the routine that computes ⟨f, h_j⟩ does not return zero at j = 0: it evaluates the inner product of an indicator with h₀ through a formula containing He₋₁, which is not a Hermite polynomial. Summed from j = 0 the covariance of two cuts at one standard deviation with their own interaction came out at 1.185 where its value is 0.5227 — the spurious term was larger than the answer. It appeared only for cuts away from the median, because for a polynomial that term is zero for a second reason, and it is the kind of defect that a check against a different case catches and a check against draws of the same case would eventually have caught too.
The number underneath both results was wrong
Asking the question this way exposed something in the established answer. The share a rule holding both squares removes of the product was quoted as
which is 53.33% at ρ = 0.5. It divides the projection by the interaction’s second moment, 1 + 2ρ², and every quantity above the line is a covariance. A product of two centred functions is not centred once the covariates are dependent: E[h₁(X)h₁(Y)] is ρ, not zero. What an R² divides by is a variance, which is 1 + ρ².
The correction is settled without measuring anything, at ρ = 1. There the two covariates are the same variable, so the interaction h₁(X)h₁(Y) is h₁(X)², which is a constant plus √2·h₂(X) — and the span of the two squares contains h₂ exactly. A rule holding the squares removes all of it, so the share must be 1. The old expression gives two thirds.
The essay that derived it says so in its own prose: “the rule removes the √2·h₂ part and leaves the constant — two units of three in squared norm, which is two thirds.” The derivation is right and the conclusion does not follow. A constant is not variance. A shape’s constant part is the same in both arms of a trial and creates no imbalance whatever, so it cannot be a share of anything left unprotected.
Corrected, the share is 4ρ²/(1 + ρ²)²: 64.00% at ρ = 0.5, 95.18% at ρ = 0.8, and monotone to 1. Four hundred thousand draws give 64.47% at a half against the corrected 64.00% and the old 53.33%. The affected numbers move by between one and forty per cent of themselves, and they move in the direction that matters — the earlier values understated how much a balancing rule protects against an interaction, which is the safe direction for a warning and the wrong one for a design calculation.
Why it survived
Nothing in this collection was in a position to catch it. Every gate on the geometry checks a closed form against draws, and both routes shared the error: the drawn check regressed the same uncentred product on the same centred dictionary. Two routes to a number are only two routes if they disagree about something, and these two disagreed about nothing except arithmetic.
What did catch it was extending the machinery. The general form written here handles any pair of one-variable functions, and applied to two cuts at one standard deviation it returned 151% — a share above one, which is impossible, and which is what an uncentred denominator does when the interaction’s mean is large. The median split is the case where the mean of the product is smallest relative to its spread, so the error was invisible exactly where the collection had been looking.
The defect is one line in each of two files, and the correction is now checked by a case with no arithmetic in it: as ρ → 1 the removed share must go to one, because the interaction becomes an element of the span.
It is worth being precise about which numbers moved, because most of them did not. An interaction whose mean is zero is unaffected: one covariate times the other’s square is odd, so its mean is zero at every correlation, and its removed share is 17.62% at ρ = 0.2 and 67.86% at ρ = 0.5 exactly as before. What moved is every even interaction — the product of the covariates, from 53.33% to 64.00% at a correlation of a half, and the product of the two squares from 44.62% to 45.46%. The size of the move is the size of the interaction’s own mean relative to its spread, which is why the correction is invisible at small correlations and large at big ones.
And the direction is uniform: the corrected share is always the larger, because a variance is always smaller than a second moment. Every statement this collection has made about what a balancing rule fails to protect against was therefore conservative rather than optimistic, which is the better way round to have been wrong and is not a reason to leave it.
What an experimenter could add that would help
The rule is negative and its practical form is not. If a dictionary of odd functions is worth nothing against an even shape, the repair is to put an even function in it, and the cheapest even function is the one a trial already has the data for.
The square of a standardised covariate is even, and holding it turns the worst case from exactly zero into something. A variance or an absolute deviation is even for the same reason. A cut away from the median is neither, so it contributes its even part: at one standard deviation that is 40.6% of it, at two standard deviations 48.8%, and the further out the cut the more even it is and the fewer units are above it.
What does not help is the thing a trial is most likely to do, which is add another odd function. A second median split at a different quantile is still mostly odd; a rank, a sign, a standardised score are all odd; and each of them is a real constraint that buys real protection against the shapes it spans, and buys exactly nothing against an even interaction.
This is worth stating as a design instruction rather than as geometry, because the instruction is short: hold at least one even function of each covariate, and check that it is even. Parity is a property anybody can compute from a function in one line — average it against its own reflection — and it decides an exact zero rather than a small number.
What is claimed here, and what is not
This essay takes which functions have to be in a dictionary for anything to be removed of an interaction. The claims are that the joint sign flip is a symmetry of the bivariate normal at every correlation, so an odd dictionary removes exactly none of an even interaction and an even dictionary exactly none of an odd one, both verified to machine precision and against draws; that a dictionary containing a covariate and its own median split is odd and therefore removes nothing of either product, which neither of the two existing explanations predicts; that a cut at one standard deviation is 59.4% odd and leaks 22.49% at ρ = 0.5; and that the established share for the polynomial case divided by a second moment rather than a variance, so 53.33% at ρ = 0.5 is 64.00% and the limit at perfect correlation is 1 rather than two thirds.
What stays out and is named as a decision: parity is a property of this family of designs. It holds because the covariates are jointly normal and centred at their own means, and a trial whose covariates are skewed, or recorded on a scale with a meaningful zero somewhere else, has no such symmetry. Nothing here measures what happens then, and the honest expectation is that the exact zeros become small numbers rather than staying zero.
The other omission is what any of it costs a treatment effect. The share removed is geometry; turning it into a variance needs a stated outcome model, and that is the other half of the field — after which what a fixed dictionary leaves of the assignment space is a question of its own.
The checks, and the refusal
Three claims are gated. The parity rule is required to hold in both directions — an odd dictionary against an even shape and an even dictionary against an odd one, both exactly zero — because a rule that only ever produces one kind of zero is indistinguishable from a coincidence. Every closed inner product is required to agree with a sample of hundreds of thousands of draws, which is what says the reduction to one-dimensional quadratures is a statement about the world. And the removed share is required to go to one as the correlation goes to one, which is the check that would have caught the old denominator on the day it was written.
The refusal is the denominator. A projection of covariances divided by a raw second moment is rejected as a share of anything: it is not bounded by one, and it reports 151% for two cuts at one standard deviation.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A dictionary that is a product — both name closed form, correlation, covariate balance, hermite polynomials, interaction, orthogonality, projection, variance explained
- A split survives what a mean does not — both name basis functions, covariate adjustment, covariate balance, interaction, orthogonality, rerandomisation, standardisation
- The part the rule already took — both name basis functions, covariate balance, experimental design, hermite polynomials, orthogonality, projection, variance explained
- A basis is a subspace — both name basis functions, covariate balance, hermite polynomials, model misspecification, orthogonality, projection
- A copula that halves a marginal — both name closed form, correlation, covariate adjustment, covariate balance, experimental design, interaction
- A zero that is arithmetic — both name basis functions, covariate adjustment, covariate balance, interaction, orthogonality, projection
Named objects
A flat tag is an object no other essay names yet.
Basis functionsClosed formCorrelationCovariate adjustmentCovariate balanceExperimental designHermite polynomialsInteractionModel misspecificationOrthogonalityProjectionRerandomisationStandardisationVariance explained