A zero that is arithmetic
Worth reading first: Balancing what is known in advance · A design is a number.
The parity field holds the dependence fixed and changes the marginals, and finds two exact zeros behaving differently. A rule balancing the means of two covariates removes exactly none of their product’s interaction under a symmetric marginal and 11.63% under a mildly skewed one. A rule balancing a median split of each removes exactly none of the product of the splits under every marginal it tries — six of them, to twenty decimal places — and the explanation offered is that both sides are functions of the sign of the latent normal, and a monotone map does not move signs.
The field says in its own last paragraph what it did not do: everything in it holds the copula fixed. A marginal is a relabelling of each axis; a copula is the joint law of the ranks. Changing the second is a different experiment, and the prediction was that the split’s zero would go with it.
It does not, and the reason is better than the prediction.
The split’s zero is arithmetic
A median split centred is a Bernoulli(½) less a half, so it takes the values and — and its square is identically, at every unit, on every draw, whatever the joint law is.
The covariance between the interaction and the main effect is therefore
for any dependence at all. The interaction is orthogonal to both main effects because one factor’s centred square is a constant, and a constant times a centred variable has mean zero.
No symmetry is involved. The parity field’s explanation is true — both sides are functions of the sign of the latent normal — and it is not the reason.
Measured on five copulas
Arithmetic is worth checking, and the check is what makes this a field rather than a footnote.
Five copulas at a Spearman rank correlation of 0.4, with a standard normal covariate throughout so nothing here is about the marginal: a Gaussian, a t on four degrees of freedom, a Frank, a Clayton whose density piles up in the lower tail, and the same Clayton turned over so it piles up in the upper one.
The median split rule removes , , , and of the interaction it is aimed at. Every one of those is zero to the quadrature’s own precision, and the last two are on copulas that are not symmetric under reflection at all.
The mean’s zero is a symmetry, and a different one
The other rule behaves as the prediction said the first one would.
A rule balancing the mean of each covariate removes , and of their product under the Gaussian, t and Frank copulas — exact zeros — and 7.707% under the Clayton and under its reflection.
The three copulas with an exact zero are the three that are radially symmetric: has the same law as . That is measured rather than looked up — the total absolute difference between the density and its own reflection is below for those three and 0.136 for the other two — and the correspondence is perfect: the zero is present exactly where the symmetry is.
So the mean’s zero needs two symmetries, one from each half of the dependence. The parity field shows it needs an odd marginal; this one shows it also needs a radially symmetric copula, with a normal covariate throughout so the marginal cannot be doing any of the work.
Why the Clayton and its reflection read the same number
The two asymmetric copulas leak 7.707% each, to four figures, and the split’s residual on both is — the same digits twice. That is not the quadrature returning a coincidence; it is forced, and the reason says exactly when it will stop being true.
Reflecting a copula sends to . With a standard normal marginal on each axis that is sending to , and a standard normal is symmetric, so the reflected pair has the same marginals as the original. The interaction the mean rule is aimed at is the product , which the double sign flip leaves unchanged, and each main effect changes sign. A covariance between an unchanged quantity and two sign-flipped ones has the same magnitude, so the share removed is identical.
So with a symmetric marginal, a copula and its reflection are the same experiment. The two readings had to agree and their agreeing is a check on the grid rather than a result.
It also predicts where the agreement breaks. Give the covariate a skewed marginal and no longer has the law of , the reflection is no longer undone, and the two copulas become genuinely different pairs. That is what the field that swept the correlation finds: the same covariate under a lower-tail copula and its reflection reads 6.1619% and 33.2638%, a factor of five apart, where here they read the same number to four figures.
The radial symmetry of the copula and the symmetry of the marginal are not two independent conditions. Either one alone is enough to make a copula and its reflection interchangeable, and the mean’s zero needs both to hold at once for a different reason.
One consequence for the figure that plots the leak against the radial gap: the five copulas supply only two distinct gaps, below and 0.136. The correspondence it establishes is therefore presence against absence rather than a slope, and nothing here measures how the leak grows with the gap.
Two zeros, two kinds of thing
Set out plainly, the two guarantees are not the same kind of object at all.
The split’s zero is an identity. It holds for every copula, every marginal, every correlation and every sample size, and it needs nothing to be assumed about anything. A practitioner balancing median splits and protecting against an interaction between the splits has a guarantee that cannot be lost.
The mean’s zero is a coincidence of two symmetries, either of which real data can break. A skewed covariate breaks one and lower-tail dependence breaks the other, and each on its own is enough.
That is a more useful division than the earlier field could reach, because a guarantee that rests on an identity and a guarantee that rests on assumptions are different things to hand somebody, and both were being reported as exactly zero.
What a copula is, and why it is the other half
A joint distribution of two covariates splits into two pieces and this collection had only varied one of them.
The marginals say what each covariate’s own distribution is. Changing a marginal is applying a monotone map to one axis, which moves every quantile and moves no rank at all: the unit at the seventieth percentile stays at the seventieth percentile.
The copula is what is left — the joint distribution of the two ranks. Changing it moves which units are extreme together, and it does so without touching either marginal.
The parity field varies the first, with a Gaussian copula throughout. This one varies the second, with a normal marginal throughout. Between them they cover the two halves, and the reason the split has to be made in exactly that way is that the two act on different things: a marginal decides what a value means and a copula decides what a pair means, and an interaction between two covariates is a statement about pairs.
The five copulas here differ in where their dependence lives: two of them concentrate it in a tail — the Clayton in the lower one and its reflection in the upper — the t copula in both, and the Gaussian and Frank in neither. Concentrating dependence in one tail is exactly what breaks the reflection symmetry, and it is the commonest departure real data shows.
Three routes to one geometry
The quadrature here is a weighted sum over a grid of each copula’s own density, and it shares no arithmetic with either of the two constructions this collection already had for the same quantity.
At the Gaussian copula it must reproduce them, and it does. The mean-and-square dictionary removes 64.00% of the product and a threshold at a value removes 22.50% — the same two numbers the dictionary field computes by a Hermite series and the parity field computes by nested quadrature over the bivariate normal.
Three independent routes to four significant figures is worth more than any one of them, and it is what makes the two exact zeros above statements rather than artefacts of a grid.
The grid, and the one thing it got wrong
One detail of the arithmetic is worth recording because the wrong version looked entirely plausible.
The natural grid for a copula is a uniform one on the unit square, since a copula lives there. On that grid the outermost node is at , which is 3.3 at 240 points — and the functions this field integrates include a square, so the tails being truncated are exactly the region carrying the answer.
The uniform-rank grid reports 62.55% where the earlier fields compute 64.00%. Nothing about 62.55% looks wrong. It is the right order, it is stable under refinement of the grid, and it would have been reported as agreement if the comparison had been to two significant figures.
Laid out on the latent normal scale instead — the copula density times the two normal ordinates, which decays at both ends whatever the copula does at the corners — it reports 63.99996%. That is the number in the sections above, and the check that caught the first version is a comparison against a construction that already existed.
What a practitioner should take
Two rules, and the second is the one worth carrying.
A median split’s interaction zero is unconditional. Balance a median split of each covariate, and the interaction between those splits is removed exactly, whatever the covariates’ joint distribution is. There is nothing to check and nothing that can break it.
A mean’s interaction zero holds under two conditions and both are checkable. The covariate has to be symmetric — which the parity field measures and which a sample can be examined for — and the dependence has to be symmetric under joint reflection, which is a property of the copula and is much less often examined. Lower-tail dependence, which is the commonest departure in practice, breaks it by 7.7% at a rank correlation of 0.4.
What the split’s zero does not cover
It is worth being exact about the scope, because “unconditional” is a strong word and this one has a boundary.
The zero is between the product of the two splits and each split’s own main effect. It says nothing about an outcome that depends on the covariates in some other way: a rule holding median splits removes exactly nothing of a linear term, exactly nothing of a product of the raw covariates, and — what a trial actually faces — its worst case over a list of six outcome shapes is exactly zero under every copula measured.
So the unconditional zero is a guarantee about one shape, and the guarantee a trial needs is about the shape it does not know. The dictionary field’s answer is that a rule holding more functions has a better worst case, and this field’s addition is that the worst case under a Clayton copula is 1.74% against the Gaussian’s exact zero — the asymmetry that broke one zero also removed the shape the zero was protecting.
That inversion is worth its own essay and gets one.
Why an identity is worth finding
The finding is negative in form — a symmetry that turned out not to be needed — and it is worth saying what a negative result of that shape buys.
It widens the guarantee. The split’s zero holds for six marginals is a statement about six marginals. It holds because a centred Bernoulli’s square is a constant is a statement about everything, and it costs nothing to check because there is nothing to check.
It narrows the other one. Once the split’s zero is known to be arithmetic, the mean’s zero can be attributed correctly, and the attribution is two conditions rather than one. The parity field could only see one of them because it never varied the other.
And it removes a wrong mechanism from circulation. Both sides are functions of the sign of the latent normal is a correct sentence that suggests a false generalisation — that any pair of sign-based rules inherits the protection — and it does not: what inherits the protection is any rule whose centred functions square to constants, which is binary functions at the median and nothing else. A split at the eightieth percentile is binary and is not balanced, so its centred square is not a constant, and its zero is gone.
That last consequence is the practical one, and it follows from the identity rather than from any measurement.
Where the identity does and does not reach
Two more rules are worth testing against the same argument, and they come out on opposite sides.
A rule holding a mean and a median split of each covariate — the two things every trial balances — inherits the split’s protection for the splits’ own interaction and not for anything else. Its worst case over six outcome shapes is exactly zero under every radially symmetric copula here and 1.74% under a Clayton, so the pair is governed by the mean’s condition rather than the split’s identity.
A rule holding a threshold at a value on the covariate’s own scale — a dose, a clinical cut-off — has no zero at all and never did. The parity field establishes that across marginals and it is true across copulas too: the removed share runs from 13.20% to 33.36% over the five, with the two extremes being the same copula and its reflection.
The pattern is that the identity covers exactly one construction — a balanced binary split — and every neighbouring rule falls back on the symmetries. That is worth knowing precisely, because the neighbouring rules are the ones a protocol usually specifies.
What is claimed here, and what is not
This essay takes what a median split’s interaction zero actually rests on. The claims are that a centred median split’s square is a quarter identically, so the interaction between two such splits is orthogonal to both main effects for any joint law whatsoever; that this is measured on five copulas at a matched Spearman rank correlation of 0.4 with a normal covariate throughout, giving removed shares between and , including on two copulas that are not radially symmetric; that a rule balancing the means gives exact zeros on the three radially symmetric copulas and 7.707% on the two that are not, with the radial gap measured at below and 0.136 respectively; and that the grid reproduces the two constructions this collection already had at the Gaussian copula, 64.00% and 22.50% against their 22.49%, by arithmetic that shares nothing with either.
What stays out, and is named as a decision: a copula fitted to data. Every copula here is a one-parameter family calibrated to a stated rank correlation, which is what makes the five comparable. What a real dataset’s dependence looks like, and whether it is close enough to radially symmetric for the mean’s zero to be nearly true, is an empirical question this collection has no data for.
Also out: more than two covariates. Every zero here is between two covariates and their product, because that is what the parity field measures and what makes the comparison exact. A trial balancing five covariates has ten pairwise interactions and a much larger dictionary, and whether the arithmetic identity survives — it should, since it is about one factor’s centred square — is not measured.
The boundary against the parity field is that it changes the marginals and this one changes the joint law of the ranks. Between them they show the mean’s zero needs one condition from each half, and the split’s zero needs neither.
Why a thousandth of the probability was worth a per cent of the answer
The truncated grid’s error deserves one line of arithmetic, because “the tails carry the answer” is the kind of sentence that is either obvious or unhelpful depending on whether the size is attached.
The uniform-rank grid at 240 points reaches out to 3.3 standard deviations. The normal mass beyond is 0.097% — a thousandth, which is why the truncation looks harmless.
But the dictionary’s entries are squares, and the share of the second moment beyond that point is
at , twelve times the probability. The products of two squares reach the fourth moment, whose share beyond the same point is 5.4%, fifty-five times the probability.
The measured shortfall — 62.55% against 63.99996%, or 2.3% in relative terms — sits between those two, which is where a quantity built from second and fourth moments should sit. So the discrepancy is not a mystery about grids; it is the exact amount of a fourth-moment integral that a grid stopping at 3.3 does not contain.
The general form is worth carrying past this field. A grid chosen to be uniform in the natural coordinate of the object — ranks, for a copula — is not uniform in the coordinate the integrand lives in, and the mismatch is largest exactly for the moments a dictionary of squares is made of. Refining the grid does not fix it: doubling to 480 points moves the outermost node only to about 3.5, whose fourth-moment tail is still 3.2%. The fix is the change of variable, not the resolution, which is why the wrong answer was stable under refinement and looked like a converged one.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two failures that cancel — both name covariate adjustment, covariate balance, interaction, marginal distribution, median split, numerical methods, parity, symmetry
- A symmetry that was not enough — both name covariate adjustment, covariate balance, interaction, marginal distribution, median split, parity, symmetry
- Balancing a skewed covariate — both name basis functions, covariate balance, interaction, marginal distribution, median split, parity, projection
- A copula that halves a marginal — both name covariate adjustment, covariate balance, interaction, marginal distribution, median split, symmetry
- A dictionary that is neither — both name basis functions, covariate adjustment, covariate balance, interaction, orthogonality, projection
- A zero that was an assumption — both name basis functions, covariate adjustment, covariate balance, interaction, orthogonality, projection
Named objects
A flat tag is an object no other essay names yet.
Basis functionsCopulaCovariate adjustmentCovariate balanceInteractionMarginal distributionMedian splitNumerical methodsOrthogonalityParityProjectionQuadratureRank correlationSymmetry