The zero that survives both
Worth reading first: Which series does the moving · Balancing what is known in advance.
Three essays of this field are about guarantees coming apart. This one is about the one that does not, and the reason it survives is worth more than the fact.
Thirty combinations, nothing above 10⁻¹⁶
Five copulas, six marginals, and a rule balancing a median split of each covariate. The largest leak anywhere in the table is 1.736 × 10⁻²⁰, at a normal covariate under a Clayton copula, which is the quadrature’s own noise rather than a leak.
That includes the cell where the covariate is exponential and the copula concentrates its dependence in the upper tail — the cell where a mean split leaves 40.288% of its interaction — and the cell where the two asymmetries cancel and a mean split leaves 0.002%. The median split’s zero does not notice any of it.
What 10⁻²⁰ means here
The number is a computed quantity rather than a measured one, and the distinction decides how much weight it carries.
Every leak in this field is a squared multiple correlation evaluated by quadrature on a grid of the copula — sums of products of smooth functions against a joint density, with no draws in them. So a quantity that is mathematically zero comes out as the grid’s rounding error rather than as a small sample estimate, and 1.736 × 10⁻²⁰ against leaks of 0.4 is not a small number, it is nothing.
The grid was checked against two earlier constructions — it reproduces 64.00% and 22.50% at the Gaussian copula, where two other fields compute the same quantities by a nested quadrature and by draws — so it is a third route to numbers two others already produce.
That matters for this particular claim more than for any other in the field. A guarantee is a claim about being exactly zero, and a measurement that could not distinguish zero from a thousandth would be unable to report one.
Why it is arithmetic
The argument is three lines and does not mention either half of the dependence.
A median split of a covariate, centred, takes exactly two values: +½ above the median and −½ below it, for every unit, on every draw, whatever the covariate’s scale is. So its square is a quarter — identically, not on average — and the interaction of two such indicators is a product of two things whose squares are constants.
The rule holds both main effects. The interaction is orthogonal to each of them because the product of an indicator with itself is a constant and the product with the other indicator is the target, and neither projection has anything left in it. The orthogonality is a consequence of the indicator taking two values symmetric about zero, and it holds for any joint law of the two covariates whatever.
Compare that with the mean split. A mean split is the indicator that a covariate exceeds its own mean, and the mean is not the median once the marginal is skewed — so the indicator takes the values and for some , its centred square is not a constant, and the orthogonality has to be recovered from somewhere else. Under a symmetric marginal and a radially symmetric copula it is recovered from parity; when either symmetry goes it is not recovered at all.
So one rule’s guarantee is a fact about the numbers ±½ and the other’s is a fact about two symmetries of the law. That is the whole difference, and it is why one survives a table the other does not.
And what it does not protect
The zero is about the interaction of two median splits and about nothing else, which is narrower than a median split is safe.
It says the rule removes the whole of a target that is the product of the two centred median-split indicators. It says nothing about a target that is the product of the two covariates themselves, or of one covariate and one split, or of a threshold and a split — and those are shapes an outcome may well take. The worst case below is over six such shapes and it is not zero anywhere the mean split’s own guarantee has failed.
So the guarantee is a guarantee about one shape, and the reason it matters is that the shape is the one the rule is built to remove. A rule that cannot remove the thing it is defined by has nothing left, and the mean split’s failure across this table is exactly that failure.
What the third rule does
The dictionary has three rules and the third one never had a zero, which makes it the useful contrast.
A cut at a fixed value on the covariate’s own scale is binary and unbalanced, so its centred square is not a constant and there was never anything to protect. What is drawn is a size, and it moves the opposite way from the mean split’s: 16.417% under a Gaussian copula with a normal covariate against 3.579% with an exponential one, and the same fall under every one of the five copulas.
The mechanism is that a cut at a fixed value on a skewed scale is a cut at an extreme quantile. A right-skewed covariate stretches its upper tail, so a threshold at one standard normal unit falls far into the tail of the transformed variable and the resulting split is heavily unbalanced — and an unbalanced split has less interaction to leak in the first place.
So a skewed covariate makes the mean split worse and the threshold better, on the same table, at the same cells. That is not a paradox and it is a warning about summaries: skewness costs a balancing rule is true of one rule and false of another.
The two routes to the same zero
The arithmetic argument above is a proof, and this collection does not ship proofs without a measurement beside them. The measurement is the table, and the two agree in a way worth stating precisely.
The proof says the leak is exactly zero for any joint law. The table checks thirty joint laws chosen to be as unlike each other as the dictionary allows — five copulas including two that break every symmetry the other guarantees rest on, six marginals from a normal to an exponential — and finds nothing above the grid’s noise.
Neither route could have found what the other found. The proof cannot say whether the code implements the centred median split it describes; the table cannot say whether the zero would survive a thirty-first combination. Together they say the implemented rule has the property the argument gives it, on every law tried, which is the strongest form this collection’s assertions take.
And it is worth noticing which of the two would have caught a mistake. A miscentred split — the indicator taking the values 1 and 0 rather than ±½ — satisfies every word of the proof’s setup and breaks its conclusion, and what would report it is the table.
What a trial is exposed to
The leaks above are on one target. A trial does not know the shape of its outcome’s interaction, so the number to act on is the worst case over the shapes it might take.
The rule every trial actually runs is a mean and a median split of each covariate, and its worst case over six outcome shapes runs from 100.0000% removed with both halves symmetric down to 95.4795% with an exponential covariate under an upper-tail copula.
So the exposure a trial carries, over the shapes it does not know, is 4.5205% in the worst cell of the table and exactly nothing in the best. And it is not monotone in either half: at a covariate skewed at 2.26 the lower-tail copula removes 99.9998% — better than the Gaussian’s 98.5823% — because the two asymmetries cancel.
The median split is what stops the worst cell being much worse. A rule holding the mean split alone leaves between 0.536% and 5.340% of its worst shape depending on the cell; adding the median split takes the best cells to zero and holds the worst to four and a half per cent, because the second indicator’s guarantee does not depend on anything the first one’s does.
What the field’s four essays add up to
The table has thirty cells, three rules and two ways of reading it, and it is worth setting the four findings side by side because only one of them is a guarantee.
The two leaks do not add. Eleven of twenty cells with both halves failing fall below what adding them gives and nine rise above; the extremes are 6.162% against 33.405% predicted and 40.288% against 30.489%. The guess the field was written to test has the wrong sign on more than half the table.
A symmetric marginal is not enough. The heavy-tailed symmetric covariate leaks exactly nothing on its own and doubles what an asymmetric copula leaks, from 7.707% to 14.229%, while satisfying every condition the parity argument states.
And a copula that breaks nothing still halves a marginal’s leak. Three radially symmetric copulas, matched on rank correlation and all harmless alone, put a factor of 1.95 between the same skewed covariate’s leaks.
Against which: one zero holds everywhere. Not because the median split is a better estimator or removes more — it removes the same thing when everything is symmetric — but because its guarantee is a statement about the numbers ±½ rather than about a symmetry of the law, and there is no law to get wrong.
The recommendation, and where it came from twice
The reading is short and it is the same one two fields reached independently.
Balance on a median split rather than a mean, and hold both if the budget allows. The median split removes the same thing when everything is symmetric, and its guarantee is the only one in this collection that survives being wrong about the joint law — not approximately, and not in most cells, but at every one of thirty combinations to twenty decimal places.
The parity field reaches it from the marginals, which it varies with the copula Gaussian. The copula field reaches it from the copulas, which it varies with the marginal normal. This field reaches it from both at once, and the value of doing so is not that the recommendation changed — it is that the recommendation had two independent supports and neither could see the other’s.
Why a rule with no guarantee is not the worst rule
One reading of all this would be that the threshold rule, which never had a zero, is the one to avoid. The table says something more careful.
A threshold at a value leaks between 2.411% and 33.360% across the thirty cells — always something, never a guarantee. A mean split leaks between 0.000% and 40.288% — sometimes nothing, sometimes more than the threshold ever does.
So the rule with a guarantee has the worse worst case. What the guarantee buys is not a smaller maximum; it is a known value in the cases where it applies, and the cases where it applies are exactly the ones a practitioner cannot verify. A threshold rule’s 16.417% under a Gaussian copula with a normal covariate is a number a trial can be told in advance and can size against. A mean split’s 0.000% in the same cell is a smaller number that becomes 21.539% if the covariate turns out to be skewed and 36.213% if the dependence turns out to be asymmetric as well.
A guarantee that holds under conditions nobody checks is not obviously better than a bound that holds always. That is the reading the whole field points at, and the median split is the only rule here that escapes it.
What it costs to hold both
The recommendation is to balance a median split as well as a mean, and a balancing rule’s budget is finite, so the second function is not free.
A rule holds a fixed number of linear constraints on the assignment, and each one it spends on a covariate is one it does not spend elsewhere — on a second covariate, on a higher-order term, on the stratification a trial also wants. The field that priced an extra function measures what each additional balanced function removes of the shapes worth protecting, and finds the returns falling steeply after the first two.
What this field adds is that the two functions are not interchangeable. A rule holding the mean split alone leaves between 0.536% and 5.340% of its worst shape across the cells measured; holding both takes the symmetric cells to exactly zero and the worst to 4.5205%. The second constraint is buying the guarantee rather than buying more removal, and a guarantee is worth a constraint in a way that a marginal improvement is not.
So the arithmetic is: one constraint for a rule whose exposure depends on two properties of the joint law nobody measures, or two for a rule whose exposure is bounded whatever they are.
The median is the only quantile with the property
The argument above turns on the split taking values symmetric about zero, and that invites the obvious question: what happens at a cut somewhere else. The answer is exact and it is worth writing out, because it prices every neighbouring rule at once.
Let be the indicator that a covariate exceeds its cut and let be the share above it. Because is an indicator, , so the centred square is
A constant plus times the centred indicator itself. So the covariance between the interaction and one main effect is
with no joint law entering except through that covariance.
Three readings come straight off it.
At the median the factor is exactly zero, whatever the two splits’ covariance is, which is the guarantee this essay is about — recovered here as the vanishing of a coefficient rather than as a property of ±½.
And nowhere else. is zero at and at no other cut, so the median is not one convenient quantile among several with the property; it is the unique one. There is no near-median family of rules sharing the guarantee approximately.
The departure is first order in the cut. Move the cut to the 55th percentile and the factor is −0.1 against a centred square of about a quarter — a forty per cent relative variation in the quantity that had to be constant, for a five-point move. At the 51st percentile it is already eight per cent. The identity is exact and it is not robust, and those are compatible: the constancy fails linearly in the distance from the median, not quadratically.
It also explains the other rule in one line. A mean split of a skewed covariate is a split at a cut whose exceedance share is not a half, so its leak is times the splits’ covariance, and is exactly what a skewed marginal produces. The two rules on this table are not two mechanisms. They are one formula read at two cuts, and the symmetric marginals leak nothing under it because for them the mean is the median.
What would break it
Naming what would is more useful than the guarantee, because the guarantee is exact and its scope is narrow.
A cut anywhere but the median. The zero is the ±½, which is what a balanced binary split gives. A split at the eightieth percentile is binary and unbalanced, its centred square is not a constant, and the field that checked it found the protection gone — under every copula, including the symmetric ones.
More than two covariates. Every measurement here is a pair. A three-way interaction of median splits is a product of three things whose squares are constants, and whether the same argument closes is not established; nothing in this collection has run it.
And a covariate that is not continuous. The median of a discrete covariate is not a value half the units fall above, so the indicator is not balanced and the ±½ is not ±½. That is a common case in trials and it is outside every measurement in this line of fields.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A split survives what a mean does not — both name covariate adjustment, covariate balance, gaussian copula, interaction, marginal distribution, median split, orthogonality, skewness, threshold
- A zero that rests on a symmetry — both name covariate adjustment, covariate balance, gaussian copula, interaction, marginal distribution, median split, numerical methods, orthogonality, skewness
- The cut that is not a quantile — both name continuous covariate, covariate balance, gaussian copula, interaction, marginal distribution, median split, orthogonality, skewness, threshold
- A margin that turns over — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, skewness, symmetry
- An answer that changes — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, skewness, symmetry
- The other dial — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, skewness, symmetry
Named objects
A flat tag is an object no other essay names yet.
Closed formContinuous covariateCovariate adjustmentCovariate balanceExperimental designGaussian copulaInteractionMarginal distributionMedian splitNumerical methodsOrthogonalitySkewnessSymmetryThresholdTreatment effect