The arcsine that closes it, and the error that was overstated
Worth reading first: Two routes to every number · The variance removed before the data.
Everything in a mixed balancing dictionary reduces, at a correlation, to one two-dimensional object: the probability that two correlated normals are both above their cut points. The essay that set that out leaves it standing there. This one is about the two routes to it, and about a caveat that turns out to have been quoted against the wrong quantity by six orders of magnitude.
The elementary case
For cuts at the median there is an answer with nothing in it but an arcsine. Two standard normals correlated at ρ are both positive with probability
so their signs agree with probability ½ + arcsin(ρ)/π and
This is Sheppard’s theorem, it is a century and a half old, and it is exact at every correlation there is. At ρ = ½ it gives exactly ⅓, because arcsin ½ is π/6. At 0.8 it gives 0.590334 and at 0.9 0.712867.
The shape of the curve is the first thing worth reading off it. It is above the diagonal at small ρ and below it at large: a pair of covariates correlated at 0.2 have median splits correlated at 0.1282, but a pair correlated at 0.9 have splits correlated at 0.7129. Dichotomising costs more where there is more to lose.
And the general one
Away from the median nothing elementary survives, and something exact does. The derivative of the orthant probability with respect to the correlation is the bivariate normal density at the cut point, so integrating from independence gives
where the substitution t = sin θ has removed the 1/√(1 − t²) that makes the usual statement awkward — the singular factor is the Jacobian. What is left is an analytic integrand on a finite interval, which sixty-four Gauss–Legendre nodes evaluate to a part in 10¹⁴.
At a = b = 0 the integrand is identically 1 and the integral is arcsin(ρ)/2π, which is Sheppard’s formula arriving as a special case rather than as a separate fact. That is what the agreement above is checking: the general construction is required to reproduce the elementary one, at eleven correlations from −0.95 to 0.99 and then again on an eighty-one-point grid that includes both endpoints.
The two ends are taken exactly rather than approached, and it is worth saying why. At ρ = 1 the pair is one variable and the event is X > max(a, b); at ρ = −1 it is a < X < −b or nothing. Clamping to 1 − 10⁻¹² instead — which is what the first version of this did — moves the upper limit of the integral by 1.4·10⁻⁶ of a radian and the answer by 9·10⁻⁷. That is six orders of magnitude worse than the quadrature is everywhere else, and it is the kind of error that hides forever unless the two routes are compared on a grid that reaches the endpoint.
The caveat that was about a different quantity
Now the second half, which changes what the older essays should have said.
The reason a cut’s geometry was taken to draws is that its Hermite coefficients decay like , so the tail past J orders falls like 1/√J: 12.76% of a median split’s variance is outside a sixteen-order truncation, 6.57% outside sixty, 4.02% outside a hundred and sixty. Doubling the work halves nothing.
Every one of those numbers is about a cut against itself. An inner product between a function of X and a function of Y truncated at J leaves
which Cauchy–Schwarz bounds by ρᴶ⁺¹√(tail_a · tail_b). The tail sets the constant and the correlation sets the rate.
Every measured error is inside its own bound — the worst ratio over the whole table is 0.620 — and the consequence for a trial is stark. A balancing dictionary over two covariates correlated at a half, truncated at sixteen orders, has inner products right to seven decimal places while the caveat that was written against it says twelve and three quarters per cent.
So the old boundary was real and was in the wrong place. The 1/√J statement is the ρ = 1 statement, and ρ = 1 is the one correlation at which two covariates are the same covariate and a balancing dictionary has a redundant column.
The rate, checked across thirteen orders of magnitude
The three measured errors at sixteen orders — 9.8 × 10⁻¹⁵, 7.1 × 10⁻⁸ and 1.5 × 10⁻² at correlations of 0.2, 0.5 and 0.95 — are a test of the claimed rate, and they pass it across a span no single measurement could.
If the error goes as ρ^(J+1) = ρ¹⁷, then raising the correlation from 0.2 to 0.5 should multiply it by 2.5¹⁷ = 5.8 × 10⁶, and the counted ratio is 7.2 × 10⁶. From 0.5 to 0.95 the prediction is 1.9¹⁷ = 5.5 × 10⁴ against a counted 2.1 × 10⁵.
Within a factor of four across a range of thirteen orders of magnitude, from a prediction with one exponent in it and no fitting anywhere. The rate is geometric in ρ and the exponent is the truncation order, and the two readings that could have distinguished a geometric rate from an algebraic one are the ones furthest apart.
The bound’s tightness moves the same way: measured error over Cauchy–Schwarz bound rises with the correlation, so the bound is loosest where the error is negligible and tightest where it is not — which is the direction a bound should be wrong in.
How many orders each correlation needs
The rate converts into the one number a practitioner assembling a dictionary needs: how deep a truncation has to go before the inner products are safe.
Keeping the error below a millionth means ρ^(J+1) below about 7.8 × 10⁻⁶, so J + 1 > 11.76/log(1/ρ). At ρ = 0.5 that is J = 16 — exactly the truncation the collection already used, which is why it was accurate to seven decimals and looked like luck. At ρ = 0.8 it is J = 52. At ρ = 0.9, J = 111. At ρ = 0.95, J = 228.
Those last two are where the argument stops being about diminishing returns and becomes about feasibility. The linearisation coefficients that multiply two Hermite series together involve factorials of the orders, so a truncation past about a hundred is not a matter of patience.
So the exact route is not an elegance and it is not insurance. It is the only route available past a correlation of about 0.85, and the truncation is not merely slower there — it is uncomputable. That locates the boundary the field was groping for: not at 1/√J, and not at ρ = 1, but at the correlation where the orders a millionth demands outrun the arithmetic that could supply them.
Which does not make the exact route unnecessary
Two things keep the previous essay’s construction worth having.
The rate is geometric in ρ, so it degrades where the correlation is high, which is exactly where a balancing rule is under most strain. At ρ = 0.95 the truncated route at sixteen orders is wrong in the third decimal place, and a Gram matrix assembled from entries that are wrong in the third decimal place is being inverted.
And the bound is a bound. It says the error is small; it does not say what it is. The exact route says what it is, which is what makes it possible to state — as the figure above does — that the truncated route is accurate, rather than to assume it.
There is also a third reason, and it is the ordinary one. A quantity computed two ways is a quantity that can be wrong in only one place at a time. This collection’s whole method is a closed form beside every simulation, and this is a case where the two routes are both closed and still disagree usefully: they disagree by an amount that is itself predicted.
What the bound is doing, and what it is not
The Cauchy–Schwarz step is worth unpacking once, because a bound that is never attained is a different object from one that is.
Truncating at J leaves , and every ρ^m in that sum is at most ρᴶ⁺¹. Pull it out and what remains is , which is at most — the two tails. So the bound is attained only when two things happen at once: the coefficients past J are proportional to each other, and ρ is so close to one that every term carries nearly the same factor.
For a cut against itself the first holds exactly, so the bound is tight in the limit ρ → 1 and loose everywhere else. The measured worst ratio of 0.620 is what that looseness looks like: the error is well inside its own bound, and the bound is still the right thing to quote, because it is the statement that does not depend on which two functions the dictionary happens to contain.
The practical version is a rule of thumb with an argument behind it. At sixteen orders the truncated route is accurate to roughly ρ^17 times a tenth — which is 10⁻⁸ at ρ = 0.5, 10⁻⁴ at ρ = 0.8, and 10⁻² at ρ = 0.95. A dictionary over covariates correlated below about 0.8 can be assembled by truncation and checked against the exact route; above that the exact route is not a luxury.
What a cut correlation is not
One consequence worth stating on its own, because it is the sort of thing that gets read the wrong way round in practice.
A trial that balances on median splits and reports the balance achieved is reporting a correlation between indicators. Two covariates correlated at 0.8 have splits correlated at 0.5903, so a report of the two balancing variables were correlated at about 0.59 describes covariates correlated at 0.8 — and reading the 0.59 as the covariates’ own association understates it by more than a third.
The direction matters for design. A rule constrained on two nearly-independent-looking cuts may be constrained on two covariates that are far from independent, and the number of admissible assignments — which is what the acceptance rate is about — depends on the covariates rather than on the report.
The same reading error runs the other way in the reporting of a trial. A table of baseline characteristics that shows two dichotomised covariates balanced to within a percentage point is a statement about the indicators; the covariates behind them may be imbalanced in a direction the split cannot see, which is what balancing on the wrong function costs. The projection identity is exact about this: a rule removes what lies in the span of what it was handed, and a span of two indicators is a four-cell table.
Why an arcsine, and not something else
It is worth having the reason the answer is an arcsine, because it explains the shape of every curve in this essay and it is one line.
The probability that two correlated normals are both positive is the probability that a point drawn from a centred bivariate normal falls in the first quadrant. That density is rotationally symmetric once the pair is whitened, so the probability of a wedge is proportional to its angle — and the angle the first quadrant subtends after whitening is π/2 + arcsin ρ. Divide by 2π and the quarter and the arcsine both appear, with nothing else in between.
Two consequences fall straight out. The answer depends on the correlation only through an angle, which is why it is concave and why it saturates: an extra tenth of correlation near ρ = 0.9 buys less angle than one near zero. And it depends on the cut points not at all when they are at the median, because the median is where the wedge argument works — move either cut and the region stops being a quadrant, the symmetry argument fails, and there is no elementary answer left. That is the whole reason the general case needs an integral and the median case does not, and it is a fact about geometry rather than about analysis.
The same picture says what happens at ρ = 1 and ρ = −1 without any limit-taking: the wedge becomes a half-plane and a line respectively, which is exactly the endpoint arithmetic the quadrature is given directly.
Why an arcsine, and why it is the same arcsine everywhere
An arcsine turning up in an answer about thresholds is not a coincidence of this particular integral, and knowing why it appears is what makes the closed form feel inevitable rather than lucky.
Ask the question geometrically. Two standard normals with correlation ρ are, jointly, a spherically symmetric cloud viewed in coordinates whose axes are not orthogonal — or equivalently, two unit vectors separated by an angle θ with cos θ = ρ, each variable being the projection of a rotationally symmetric point onto its own vector. The event both above their medians is the event that the point falls on the positive side of both vectors: the intersection of two half-planes through the origin, which is a wedge. And the probability that a rotationally symmetric point lands in a wedge is the wedge’s angle divided by 2π, because there is nothing else it could be.
The wedge’s angle is π − θ, so the probability is (π − arccos ρ)/2π = 1/4 + arcsin(ρ)/2π, and the correlation between the two signs — four times that, minus one — is (2/π)arcsin ρ. Sheppard’s theorem is that arithmetic and nothing more. There is no series, no special function, and no approximation anywhere in it.
That is why it is exact and why it is exact only at the median. Every step used the origin: the half-planes pass through it, the symmetry is about it, the angle is measured at it. Move a threshold off centre and the two half-planes no longer share a point on the symmetry axis, the region stops being a wedge, and the answer stops being an angle. What replaces it is a general orthant probability, which has no elementary closed form and is computed here by quadrature — Drezner’s form, sixty-four Gauss–Legendre nodes, with the endpoints ρ = ±1 handled exactly rather than by taking a limit into a singular integrand.
The same arcsine appears in the arcsine law for the fraction of time a random walk spends positive, and its close relative the arctangent supplies the Cauchy law for the quotient of two independent normals — where the mechanism is identical and one step shorter, since the angle of an isotropic pair is uniform and the quotient is its tangent. The recurrence is not an analogy. It is one fact about rotational symmetry, asked for an angle in several places.
What is claimed here, and what is not
This essay takes the two routes to a cut point’s inner product. The claims are that the orthant probability’s quadrature and Sheppard’s arcsine agree to 3.3·10⁻¹⁶ over an eighty-one-point grid including both endpoints; that the truncated Mehler sum’s error is bounded by ρᴶ⁺¹√(tail·tail), with every measured error inside its own bound and the worst ratio 0.620; and that the 1/√J tail the older caveat quotes overstates the error in an inner product at ρ = 0.5 and sixteen orders by a factor of about two million.
What stays out and is named as a decision: the variance decomposition of a cut against a polynomial basis, where the tail really is the object and really does fall like 1/√J. Nothing here makes that computation easier, and the smoothness argument that depends on it is unchanged.
The boundary against the essay before it is that it is about which three shapes an inner product can take and this one is about computing the third of them twice.
The checks, and the refusals that make them mean something
Two claims are gated. The quadrature is required to reproduce the arcsine at eleven correlations from −0.95 to 0.99, to a ten-trillionth — the check that licenses using the quadrature at cut points where no elementary form exists. And every truncation error in the table is required to sit inside ρᴶ⁺¹√(tail·tail), with the four entries whose bound has fallen below what a double can represent excluded and counted rather than silently included: at ρ = 0.2 and thirty-two orders the bound is 8·10⁻²⁵ and the measured error is 10⁻¹⁶, which is a statement about floating-point arithmetic rather than about Mehler.
Two refusals bite. A cut’s geometry quoted at its tail and described as asymptotic is rejected, with the factor of two million named. And a cut correlation read as the covariates’ own is rejected: 0.8 becomes 0.5903, and reading the second as the first understates the association by more than a third.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The fourth moment that was missing — both name basis functions, closed form, continuous covariate, correlation, covariate balance, hermite polynomials, monte carlo, orthogonality, projection, threshold
- What the extra function buys — both name basis functions, closed form, correlation, covariate balance, experimental design, orthogonality, projection, variance explained
- A basis is a subspace — both name basis functions, covariate balance, hermite polynomials, monte carlo, orthogonality, projection, threshold
- A dictionary that is a product — both name closed form, correlation, covariate balance, hermite polynomials, orthogonality, projection, variance explained
- A zero that was an assumption — both name basis functions, closed form, correlation, covariate balance, hermite polynomials, orthogonality, projection
- The part the rule already took — both name basis functions, covariate balance, experimental design, hermite polynomials, orthogonality, projection, variance explained
Named objects
A flat tag is an object no other essay names yet.
Approximation errorBasis functionsClosed formContinuous covariateCorrelationCovariate balanceDependenceDiscretenessExperimental designHermite polynomialsMonte CarloOrthogonalityProjectionThresholdVariance explained