A cut is not a polynomial, and it does not have to be
Worth reading first: The variance removed before the data · Balancing what is known in advance.
The essay that closed a balancing dictionary’s geometry at a correlation did it for polynomials, and said why in its own first paragraph. A product of two Hermite functions of one variable is a finite combination of Hermite functions of that variable, so a dictionary of powers can be multiplied out exactly and Mehler’s formula applied to the products. Sixteen orders is not adequate for such a dictionary, it is exact for it.
A threshold is a different object. Its Hermite coefficients decay like and never stop, so no truncation is exact and the tail past order J falls like 1/√J. At sixty orders 6.57% of a median split’s variance is still outside the sum, and at a hundred and sixty it is still 4.02%. So the dictionary in that field replaced its cut-point entries with polynomial relatives, the closed results were stated for powers, and the cut cases were taken to draws.
A median split is what trials actually balance on. Age above or below the median, a biomarker above a threshold, a site classified as large or small: the dictionary a real balancing rule is handed is full of these. The boundary was therefore in exactly the wrong place, and the question left open was whether a different representation makes the cut cases exact.
One does, and it is not a better truncation. It is a different decomposition, and it takes one line.
Three inner products, and only one is two-dimensional
Everything a dictionary of Hermite functions and indicators can produce is one of three shapes when the two covariates are correlated. Take (X, Y) standard bivariate normal with correlation ρ.
Hermite against Hermite. Mehler’s formula gives
so one term survives and the rest are zero. This is the case the polynomial dictionary is built from.
Hermite against indicator. Because — the conditional expectation of a Hermite function is the same function scaled — the whole two-dimensional expectation collapses:
and the right-hand factor is the one-variable inner product this collection has had since the essay about which shapes a basis protects. No expansion, no truncation, no order to choose.
Indicator against indicator. This one does not collapse:
where Φ₂ is the orthant probability. It is the only genuinely two-dimensional object in the whole construction, and the next essay is about the two routes to it.
That is the entire representation. The conditional expectation does the work a truncation was being asked to do, and it does it for a threshold exactly as readily as for a power, because it never needs the threshold’s expansion at all.
Why the expansion is the wrong object here
It is worth looking at what was being avoided, because the coefficients are a real obstacle to a real calculation — just not to this one.
The decay is genuinely slow. To capture 99% of a median split’s variance takes several thousand orders, and the linearisation coefficients that multiply two Hermite series together grow factorially, so nobody computes them that far. A field that needs a cut’s variance decomposition against a polynomial basis really is stuck with an approximation.
But an inner product between a function of X and a function of Y is not a variance decomposition. It is an expectation over a joint distribution, and the joint distribution has a structure — the conditional expectation identity above — that the expansion route throws away and then tries to rebuild term by term.
The truncation was a consequence of the representation rather than of the problem.
The distinction is worth holding on to because it recurs. A quantity is hard to compute in the basis somebody happened to write it down in, and the reflex is to compute more terms; the alternative is to ask which structural fact about the distribution the basis is failing to use. Here the fact is that a conditional expectation of a Hermite function is that function again — the single property that makes Mehler’s formula true in the first place — and the expansion route rebuilds a consequence of it term by term instead of applying it once.
The orthant probability, and why one integral is enough
The one shape that does not collapse is worth naming precisely, because not elementary and not closed are different claims and only the first is true of it.
Φ₂(a, b; ρ) has no expression in elementary functions for general a and b. It does have a one-dimensional integral representation over the correlation — the derivative of the orthant probability with respect to ρ is the bivariate density at (a, b), so integrating that from 0 gives the whole thing — and after the substitution t = sin θ the integrand is smooth on a finite interval with no singularity at either end. Sixty-four Gauss–Legendre nodes then give it to a part in 10¹⁴.
That is not a simulation. It has no seed, it returns the same number on every machine, and its error is bounded by the smoothness of a function that is analytic on the interval. For the purpose every other formula in this collection is held to — a closed form beside every simulation, so that neither route can confirm itself — it is a closed form. At the median it also has an elementary answer, which is what the next essay uses to check it.
What a cut costs, which is a separate question
None of this says that balancing on a cut is a good idea, and the figures say fairly plainly that it is not, at least not when the covariate itself is available.
A median split of a covariate correlated at 0.8 with another has a correlation of 0.5903 with the other’s median split. A cut at one standard deviation has 0.5429, and a cut at two has 0.4186. Every one of these is below the covariates’ own correlation, and the shortfall is not small: dichotomising at the median throws away about a quarter of the association, and dichotomising at two standard deviations throws away nearly half.
So a rule handed cuts is protecting less than a rule handed the covariates, and the geometry now says exactly how much less. That is a design statement rather than a technical one, and it is the kind of statement the exact representation exists to make: the price of a cut is now a number rather than a caveat.
The whole dictionary, checked against draws
An exact form that nobody has checked against the world is algebra. Every entry of the mixed dictionary is therefore computed both ways: from the closed expressions above, and from four hundred thousand draws of a correlated pair with the covariances taken directly.
Six shapes — three Hermite orders and three cut points, at the median and at ±1 — give thirty-six ordered pairs, including every mixed entry, which is the case the representation exists for. At a correlation of 0.6 the worst departure over all thirty-six is 3.8 standard errors of the Monte Carlo error, and there is no pattern in which pairs are furthest.
Nothing in the counted route knows what an orthant probability is, and nothing in the closed route draws a number.
What this makes available
Three things become computable that were not, and they are the ones a design question actually asks.
A dictionary that mixes powers and cuts. The one a real trial is handed — age above the median, age itself, its square, a biomarker above a threshold — has both kinds of entry, and its Gram matrix at a correlation now has a closed form entry by entry.
A guarantee for a cut-shaped outcome. The projection identity says a balancing rule removes exactly the part of an outcome that lies in the span of what it was handed, and applying it to a threshold outcome needs the inner products between that threshold and the dictionary. Those are now the second shape above, and they are ρ^j times a one-variable number.
The interaction question, for the dictionary a trial has. The essay that destroyed the interaction zero did it for a product of powers. Whether the same thing happens to a product of cuts is a different question with a different answer, and it can only be asked because the inner products are exact.
The one-variable case was always exact, and that is the clue
There is a detail in the older essays that reads differently once this is in place.
The inner products between a threshold and a Hermite function of the same variable have been closed forms in this collection from the start, and so have those between two thresholds of the same variable: the second is Φ̄(max(a, b)) − Φ̄(a)Φ̄(b), which is elementary. A field working with one covariate has never needed a Hermite expansion of a cut for anything except a smoothness argument, where the coefficients are the object.
So the expansion was introduced when a second covariate arrived, to make the product of two functions of one variable computable — and it was the product that needed it, not the cut. Conditioning on Y removes the product from the problem. The one-variable answers were exact the whole time and the two-variable ones are the same answers multiplied by ρ^j.
The tail law, and what it would take
The two truncation figures are enough to confirm the law they are quoted against and then to say what obeying it would cost.
If the variance outside order J falls like 1/√J, then going from sixty orders to a hundred and sixty should divide it by √(160/60) = 1.633. The counted figures are 6.57% and 4.02%, whose ratio is 1.634. The law is not approximately right over that range; it is right to three figures.
Run it forwards and the size of the obstacle is unambiguous. Reaching 1% of the variance outside the sum takes 60 × 6.57² ≈ 2,600 orders, which is the “several thousand” the essay quotes. Reaching a tenth of a per cent takes a hundred times that — a quarter of a million orders — and the linearisation coefficients needed to multiply two such series together grow factorially long before any of it.
So the expansion route is not merely inconvenient at the orders anybody computes; it is unreachable at the accuracy the rest of this collection works to. Every other closed form here is checked to ten decimal places or better, and a representation whose error falls as the inverse square root of the work can never get there. The conditioning identity is not a shortcut to the same answer — it is the only route to an answer of the accuracy the field’s other results are stated at.
Two over pi again, and the floor it sets
The correlation between two median splits is quoted at 0.5903 for covariates correlated at 0.8, and the arcsine the next essay supplies reproduces it: (2/π)·arcsin(0.8) = 0.59027. That agreement is worth using rather than only noting, because dividing through gives a statement about every correlation at once.
The share of the association a pair of median splits retains is
(2/π)·arcsin(ρ) / ρ
which is 0.738 at ρ = 0.8, 0.667 at ρ = 0.5 and 0.641 at ρ = 0.2. As ρ → 1 it rises to one; as ρ → 0 it falls to 2/π = 0.6366 and stops. Dichotomising two covariates never throws away more than 36.3% of their association, however weakly they are related, and throws away less the more strongly they are.
That floor is the third appearance of the same constant in this collection, and the three are one integral. A median split sees 2/π of its own covariate’s variance; a balanced mean removes 2/π of the imbalance in a median split; and two median splits retain 2/π of the correlation between the covariates they cut. The first two are the same statement read in opposite directions and the third is what happens when the operation is applied at both ends — which is why it is the square-root-like quantity rather than 2/π squared: a correlation is a ratio of a covariance to two standard deviations, and the cut costs the numerator and the denominators at compensating rates.
The cuts away from the median have no such floor, and that is the design reading. At ρ = 0.8 a cut at one standard deviation retains 0.679 of the association and a cut at two retains 0.523 — so a threshold in the tail loses nearly half where a median split loses a quarter, and the loss keeps growing as the cut moves out. A dictionary of median splits is the least damaging dictionary of cuts available, and it is still bounded below by a constant nobody chose.
What conditioning bought, and why it is not a trick
The step that closes this is one line, and one line is exactly the sort of thing that reads as a manipulation rather than as a reason. It is worth saying what it actually is.
Everything hard about the two-variable problem came from treating it as two-dimensional: a threshold in X and a threshold in Y, joint over a correlated Gaussian, with no obvious factorisation. But a bivariate normal has a property that removes the second dimension wherever one of the two functions is a polynomial: conditional on Y = y, X is normal with mean ρy and variance 1 − ρ². Taking the expectation of a Hermite polynomial under that conditional gives ρ^j h_j(y) exactly — the Hermite polynomials are the eigenfunctions of the conditioning operator, and ρ^j are its eigenvalues.
That is the whole of it. ⟨h_j(X), h_k(Y)⟩ = δⱼₖρ^j is not a coincidence of the Gaussian integral; it is the statement that a self-adjoint operator has orthogonal eigenfunctions. And ⟨h_j(X), 1{Y > c}⟩ = ρ^j times the one-variable answer is the same statement applied to a function that happens not to be a polynomial: the operator does not care, it just multiplies by ρ^j on the way through. Mehler’s formula, which is usually introduced as a series identity to be verified term by term, is the spectral decomposition of that operator written out. Read that way it is not a trick and it does not need truncating, because nothing is being expanded.
What survives as genuinely two-dimensional is the one object with no polynomial in it at all: the orthant probability, both variables thresholded. There is no eigenfunction argument available there, and it does not reduce. It is a single smooth quadrature over a bounded interval, which is why sixty-four Gauss–Legendre nodes settle it, and at the median it is (2/π)arcsin ρ in closed form.
So the field’s difficulty was in the wrong place. The expansion was never the obstacle; it was an artefact of attacking a two-variable inner product with a two-variable expansion when only one of the two variables ever needed expanding. Once that is seen, the boundary that sent cut points to draws stops looking like a prudent retreat from a hard integral and starts looking like what it was — a boundary drawn around a representation rather than around a problem, and drawn before anything was measured.
What is claimed here, and what is not
This essay takes the representation that makes a cut point’s geometry exact. The claims are that the three inner products a mixed dictionary produces are δⱼₖρ^j, ρ^j times a one-variable indicator–Hermite covariance, and an orthant probability; that only the third is genuinely two-dimensional; that this makes a dictionary containing thresholds closed at any correlation with no order to choose; and that all thirty-six entries agree with four hundred thousand draws to within 3.8 standard errors.
What stays out and is named as a decision: more than two covariates. The conditional expectation identity generalises, but the object at the bottom becomes a three-dimensional orthant probability, which has no one-dimensional integral representation. The construction here is exactly two-variable, and a trial balancing on three correlated covariates would need something else.
The boundary against the essay that closed the polynomial case is that it is about multiplying two expansions together and this one is about not needing an expansion.
The checks, and the refusals that make them mean something
Two claims are gated. Every inner product in the mixed dictionary is required to agree with four hundred thousand draws — thirty-six pairs, both cuts and powers and every mixed entry, worst departure 3.8 standard errors. And the truncated route is required to converge to the exact one as the order rises, which is what makes the exact route a generalisation of the old one rather than a different quantity: for polynomials the two agree to a million-millionth at any order past the degree, and for a cut the truncated route’s error falls monotonically.
The refusal this field opens with is the boundary it removes. A cut’s geometry taken to a truncation and described as asymptotic is rejected, on the ground that the tail the caveat was written against is the ρ = 1 statement and the error at any correlation a trial has is smaller by orders of magnitude.
What links here
Computed from the collection, not written here: the essays that point at this one.
- The zero that survives a cut
- The arcsine that closes it, and the error that was overstated
- The fourth moment that was missing
- A dictionary that is neither
- A zero that was an assumption
- Where the guarantee is exactly zero
- The cut that is not a quantile
- A split survives what a mean does not
- and 1 more
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What the extra function buys — both name basis functions, closed form, correlation, covariate balance, experimental design, interaction, orthogonality, projection
- The zero that survives both — both name closed form, continuous covariate, covariate balance, experimental design, interaction, orthogonality, threshold
- A copula that halves a marginal — both name closed form, continuous covariate, correlation, covariate balance, experimental design, interaction
- A zero that rests on a symmetry — both name basis functions, covariate balance, hermite polynomials, interaction, orthogonality, projection
- The part the rule already took — both name basis functions, covariate balance, experimental design, hermite polynomials, orthogonality, projection
- A count that has to be estimated — both name basis, closed form, covariate balance, orthogonality, randomisation
Named objects
A flat tag is an object no other essay names yet.
BasisBasis functionsClosed formContinuous covariateCorrelationCovariate balanceDependenceDiscretenessExperimental designHermite polynomialsInteractionOrthogonalityProjectionRandomisationThreshold