A cut point, at a correlation

The zero that survives a cut

A rule holding both main effects removes half of a pure interaction between correlated powers and exactly none between correlated median splits. The guarantee that a correlation destroyed was never about interactions.

Worth reading first: A design is a number · The variance removed before the data.

The essay that destroyed the interaction guarantee is one of the sharpest results in this collection. A balancing rule handed every main effect of both covariates removes exactly none of a pure interaction when the covariates are independent — that is a theorem — and once they are dependent it removes

4ρ2(1+ρ2)2\frac{4\rho^2}{(1 + \rho^2)^2}

of it, which is 64.00% by ρ = 0.5 and rises to one at perfect correlation. Every number the machinery produces moves, and the guarantee that a rule cannot tell one basis from another with the same span survives untouched.

That result is about a product of powers, XY, against a dictionary of powers. The dictionary a trial is actually handed is full of median splits, and until the geometry of a cut at a correlation was closed exactly the same question could not be asked for them.

Asked, it has the opposite answer. For median splits the removed share is exactly zero at every correlation, and the reason is one line.

sign(x)² = 1

Let c(x) = sign(x), which is the median split centred and scaled to unit variance. The pure interaction is g = c(X)c(Y), and the rule holds the span of {c(X), c(Y)}. The two inner products it needs are

g,c(X)=E[c(X)2c(Y)]=E[c(Y)]=0,\langle g,\, c(X)\rangle = \operatorname{E}\big[c(X)^2 c(Y)\big] = \operatorname{E}\big[c(Y)\big] = 0,

and the same with the roles swapped. The square of a two-valued function is a constant, so the interaction is orthogonal to each of its own factors whatever the correlation between them is, and the projection has nothing to project onto.

Compare what happens for powers. There h₁(x)h₁(y) is not orthogonal to h₂(x): the product h₁h₁ = h₀ + √2·h₂ is a polynomial of higher degree that the dictionary already contains, so the interaction has a component along a main effect the moment ρ is not zero, and the component is √2·ρ. The dictionary is closed under the multiplication that produced the interaction, and that is exactly the problem.

A dictionary of median splits is not. Multiplying c(X) by itself leaves the constant, which no balancing rule constrains and no outcome depends on.

The guarantee that survives a correlation, and the one that does not. What a balancing rule handed both main effects removes of the pure interaction between them, as the covariates become dependent. For median splits it is exactly zero at every correlation, because sign(x)² = 1: the interaction sign(X)sign(Y) is orthogonal to sign(X) and to sign(Y) whatever ρ is. For the product of the raw covariates it is 4ρ²/(1+ρ²)² — 64.00% by ρ = 0.5, rising to all of it at perfect correlation. A cut away from the median sits between them and is not small: 23.01% at a cut of one. The zero is not a fact about interactions. It is a fact about a dictionary whose functions square to a constant, which a polynomial one does not.
Fig. 1 What a rule holding both main effects removes of the interaction between them, as the covariates become dependent. For median splits it is zero at every correlation, to 10⁻¹⁶ rather than to a tolerance. For the product of the raw covariates it is 64.00% by ρ = 0.5.

It is the median, not the cut

The result is narrower than cuts are safe and the boundary is exact, which makes it more useful rather than less.

Away from the median the square of a cut is not constant. For c(x) = (1{x > a} − p)/σ with p = Φ̄(a),

c(x)2=A1{x>a}+B,A=12pσ2,c(x)^2 = A\,\mathbf{1}\{x > a\} + B, \qquad A = \frac{1 - 2p}{\sigma^2},

because an indicator is its own square. A is zero exactly when p = ½, and everything follows: at the median the square carries no information about the covariate at all, and away from it the square is itself a rescaled indicator, so the interaction has a component along the other cut and the projection removes it.

The leak is closed form and it is not small. At ρ = 0.5 a rule holding two cuts at half a standard deviation removes 9.51% of their interaction, at one standard deviation 22.49%, at one and a half 26.84%. At ρ = 0.8 the cut at one removes 51.29%, against 95.18% for the raw product.

The leak is also not monotone in the cut. At ρ = 0.5 it rises to 26.84% at a cut of 1.5 and falls to 23.87% at a cut of 2, because two competing things are happening: A grows without limit as the cut moves out, and the correlation between the two indicators shrinks towards zero. The product turns over somewhere around one and a half standard deviations.

The guarantee that survives a correlation, and the one that does not. What a balancing rule handed both main effects removes of the pure interaction between them, as the covariates become dependent. For median splits it is exactly zero at every correlation, because sign(x)² = 1: the interaction sign(X)sign(Y) is orthogonal to sign(X) and to sign(Y) whatever ρ is. For the product of the raw covariates it is 4ρ²/(1+ρ²)² — 64.00% by ρ = 0.5, rising to all of it at perfect correlation. A cut away from the median sits between them and is not small: 23.01% at a cut of one. The zero is not a fact about interactions. It is a fact about a dictionary whose functions square to a constant, which a polynomial one does not.
Fig. 2 The same curves read for the boundary rather than for the zero: the median is flat on the axis and every other cut is a rising curve between it and the polynomial answer.

The Gram matrix that makes the projection trivial

There is a second reason the median case is clean, and it is worth having because it explains why the answer is zero rather than merely small.

The projection of g onto a span is a′G⁻¹a, where a holds the inner products of g with the basis and G is the basis’s Gram matrix. For two median splits, G is

(1rr1),r=2πarcsinρ,\begin{pmatrix} 1 & r \\ r & 1 \end{pmatrix}, \qquad r = \frac{2}{\pi}\arcsin\rho,

which is not the identity — the two main effects are correlated, and at ρ = 0.8 they are correlated at 0.5903. G⁻¹ is therefore a real matrix with real off-diagonal entries, and none of that matters, because a is the zero vector. The correlation between the main effects is irrelevant to the result, and a great deal of the polynomial answer’s machinery is about exactly that correlation.

That is what makes the zero robust rather than delicate. It does not depend on the basis being orthogonal, on the correlation being anything in particular, or on the two cuts being at the same place: c(X)² is constant for the first cut whatever the second one is, so ⟨g, c(X)⟩ = E[c(Y)] = 0 as long as the first is at its median. A dictionary with one median split and one quartile split has one of the two inner products zero and the other not.

Counted

The closed expressions are algebra about a bivariate normal, and the whole of this collection’s method is that algebra about a distribution is checked against draws from it.

Four hundred thousand correlated pairs, the interaction formed pointwise, the two main effects formed pointwise, the interaction regressed on them and the removed share read off. Nothing in that route touches an orthant probability or a Hermite coefficient.

The zero, and the leak, both counted. What a rule holding both cuts removes of the interaction between them, by two routes: a closed expression in orthant probabilities, and 400,000 draws of a correlated pair with the interaction regressed on the two main effects. At the median it is zero in the closed form and 2.3e-5 in the draws, which is the sampling noise of a quantity that is exactly nothing. Away from the median it is not small — 22.49% at a cut of one and a correlation of a half — and the two routes agree to 4.0e-3. Nothing in the counted route knows what an orthant probability is.
Fig. 3 The closed answer against the counted one, at two cuts and two correlations. At the median the closed value is zero and the counted value is 2.0·10⁻⁵, which is the sampling noise of a quantity that is exactly nothing.

At a cut of one and a correlation of a half the closed answer is 0.224937 and four hundred thousand draws report 0.226704. At ρ = 0.8 they are 0.512881 and 0.508920. The worst departure anywhere in the table is inside three Monte Carlo standard errors.

The median entries are the ones worth dwelling on. A regression of the interaction on the two main effects, computed from real draws, returns an R² of two hundredths of a per mille — which is what an exactly-zero population quantity looks like when it is estimated from a finite sample. Had the closed answer been small rather than zero, this table could not tell the difference; what makes it evidence is that the closed route returns 0 and not 10⁻⁵.

The most exposed cut is met by one unit in sixteen

The leak’s turnover is reported as “somewhere around one and a half standard deviations”, and the four readings at a correlation of a half locate it.

Fitting a parabola through 22.49% at a cut of 1, 26.84% at 1.5 and 23.87% at 2 puts the maximum at a = 1.547, at a leak of 26.9%. That cut is met by Φ̄(1.547) = 6.1% of units — one in sixteen.

Which is worth setting beside the main-effect result from the same dictionary. A rule balancing a covariate’s mean protects a threshold worse the further out it sits, falling monotonically to almost nothing past two standard deviations. The interaction between two thresholds is exposed most at a cut in the near tail and less on either side of it.

So the two failures do not sit at the same cut point. A design worried about a rare-event main effect should worry about the far tail; a design worried about an interaction between two cuts should worry about the near one, at about one and a half standard deviations, where a sixteenth of the units are. Neither warning generalises to the other, and a single sentence about “thresholds in the tail” covers only the first.

The polynomial comparison narrows as the correlation rises. At ρ = 0.5 the powers dictionary removes 64.00% of its interaction against the worst cut’s 26.9% — a factor of 2.4. At ρ = 0.8 it is 95.18% against 51.29%, a factor of 1.9. Both go to one at perfect correlation, so the two dictionaries converge exactly where neither guarantee is worth anything.

What an exact zero looks like when it is counted

The median cell’s counted value of 2.0 × 10⁻⁵ is offered as sampling noise around nothing, and it is worth checking against what a null regression should give.

An R² from regressing on two predictors that explain nothing has an expectation of p/(n − 1), which on four hundred thousand draws is 5.0 × 10⁻⁶, with a standard deviation of about the same size. The counted 2.0 × 10⁻⁵ is a few times that — the right order, and the right order is the whole of what can be asked of it.

That is the strongest form the check takes. A closed answer of 10⁻⁵ and a closed answer of zero would both be consistent with a counted 2.0 × 10⁻⁵; what separates them is that one route returns exactly nothing and the other returns the noise floor of a regression, and the two agreeing at the noise floor is the only evidence a finite simulation can offer about an identity.

What this changes about balancing on a split

The design consequence runs against the usual instinct, and it is worth stating plainly because the usual instinct has a good reason behind it.

Dichotomising a covariate throws information away. The previous essay measures how much: a pair correlated at 0.8 has median splits correlated at 0.5903, so a rule constrained on splits is constrained on a weaker version of the same information. For a main effect that is a straight loss, and a rule handed the covariates protects an outcome linear in them better than a rule handed their splits.

For an interaction it is the other way round. A rule handed both covariates and their powers is, at a correlation, quietly removing half of any pure interaction — which sounds like a benefit and is not: what a rule removes is what its guarantee covers, and what it appears to remove without covering is what the essay on a rule reported at its own best shape is about. A rule handed two median splits removes none of it, at every correlation, which is a statement that can be made in advance and does not move when the covariates turn out to be dependent.

So the two dictionaries are not ordered. Cuts protect main effects worse and behave predictably about interactions; powers protect main effects better and have an interaction guarantee that is exactly zero at independence and nowhere else. Which is preferable depends on what the outcome is thought to depend on, and that is the same conclusion this field reached about which shapes are worth protecting, arriving from the other side.

A note on what “removes” means here

One clarification, because the word does two jobs in this collection and the result reads differently under each.

A balancing rule constrains the imbalance in the functions it was handed. The projection identity says that the imbalance left in an outcome g is the imbalance in the part of g orthogonal to that span, so the removed share is R²(g | span) — the fraction of g’s variance the rule has taken responsibility for. A removed share of zero means the rule has done nothing about g and its guarantee says nothing about g.

So the median split’s zero is not good news about the rule. It is a statement that a trial balanced on two median splits has no protection at all against an outcome that depends on their product, at any correlation, and knows it in advance. The polynomial dictionary’s 64.00% at ρ = 0.5 is protection the rule really does supply — and the reason that result was reported as a loss is that the same field had proved the share is zero under independence, so a rule designed against the independent case was being credited with a guarantee it had not been given.

Both numbers are useful and neither is a verdict. What each says is how much of an interaction the rule’s guarantee covers, and the answer for cuts happens to be a constant.

It is the median that does it, and not the cut

The zero here is exact at every correlation, and exactness of that kind almost always means an algebraic identity rather than a cancellation that happens to hold. It does, and naming it says immediately how far it travels.

A median split takes two values, and the one fact that matters is that its centred version squares to a constant: sign(x)² = 1. So the product of two median splits, regressed on either of its own factors, has a coefficient proportional to E[sign(X)²sign(Y)] = E[sign(Y)] = 0. The interaction is orthogonal to each of its factors, and it is orthogonal to them whatever the correlation between X and Y is, because ρ never enters that expectation at all. Nothing is being cancelled. There was never a term.

Move the cut and the identity goes with it. An indicator away from the median squares to a rescaled copy of itself rather than to a constant, so the product of two of them has a genuine component along each factor, and the size of that component is a function of ρ and of how far the cut sits from centre. It reaches 22.49% at a cut of one standard deviation and a correlation of a half, and turns over near one and a half — because a cut far out in the tail is nearly constant, and a nearly constant function has little of anything to remove.

So the right statement is about the dictionary, not about interactions. A rule handed both main effects removes none of their interaction exactly when the span of those main effects is closed under the multiplication that produces it — when multiplying two of the dictionary’s functions lands outside the dictionary’s own span entirely. Median splits have that property. Powers do not: x·y lies partly along x and partly along y once X and Y are correlated, which is where 64.00% at ρ = 0.5 comes from, and that number is a fact about polynomials rather than about balancing.

This matters more than a special case because a real trial balances on cuts far more often than on powers — above or below median age, above or below a threshold on a baseline score, in one stratum or the other. The dictionary a design actually uses is the one where the zero holds, and it held all along; what it was missing was the closed geometry to prove it in, which the arcsine now supplies.

What is claimed here, and what is not

This essay takes the interaction guarantee for a dictionary of cuts. The claims are that a rule holding both median splits removes exactly none of the interaction between them at every correlation, to 10⁻¹⁶, because a two-valued function squares to a constant; that the same is not true away from the median, where the leak is closed form and reaches 22.49% at a cut of one and a correlation of a half against 64.00% for the raw product; that the leak is not monotone in the cut, turning over near one and a half standard deviations; and that the closed answers agree with four hundred thousand draws to within three standard errors.

What stays out and is named as a decision: a mixed dictionary, containing both a covariate and its own median split. That is the dictionary a cautious trial would actually specify, its Gram matrix is available from the exact representation, and the interaction result for it is neither of the two answers here — the span contains the cut, so the constant argument fails, and it contains no higher power of the cut, so the polynomial argument fails too. Nothing here computes it.

The boundary against the essay that destroyed the polynomial zero is that its result is about a dictionary closed under the multiplication that makes the interaction, and this one is about a dictionary that is not.

The checks, and the refusals that make them mean something

Two claims are gated. The median split’s zero is required to be zero to machine precision at eleven correlations from 0 to 0.99 — an identity, not a small number — and the leak away from the median is required to be positive and to grow both with the cut and with the correlation over the range where it is monotone, which is what makes the zero a fact about the median rather than about cuts.

Two refusals bite. The polynomial dictionary’s 4ρ²/(1+ρ²)² carried to a dictionary of cuts is rejected: it says a rule removes half of an interaction it removes none of, so it overstates the rule and understates what is left for the outcome to depend on. And a cut correlation read as the covariates’ own is rejected in the essay before this one, for the same reason it matters here — every entry of the Gram matrix this result is computed from is a cut correlation and not a covariate one.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A dictionary that is a product — both name basis, closed form, correlation, covariate balance, hermite polynomials, interaction, orthogonality, projection, randomisation, variance explained
  • The part the rule already took — both name basis functions, covariate balance, experimental design, hermite polynomials, orthogonality, projection, variance explained
  • The zero that survives both — both name closed form, continuous covariate, covariate balance, experimental design, interaction, orthogonality, threshold
  • Where the guarantee is exactly zero — both name basis functions, closed form, covariate balance, hermite polynomials, orthogonality, projection, threshold
  • A basis is a subspace — both name basis functions, covariate balance, hermite polynomials, orthogonality, projection, threshold
  • A copula that halves a marginal — both name closed form, continuous covariate, correlation, covariate balance, experimental design, interaction

Named objects

A flat tag is an object no other essay names yet.

BasisBasis functionsClosed formContinuous covariateCorrelationCovariate balanceDependenceExperimental designHermite polynomialsInteractionOrthogonalityProjectionRandomisationThresholdVariance explained