A margin that turns over
Worth reading first: Which series does the moving · Randomisation is not balance.
The field that varied a copula and a marginal together builds a table of thirty cells and reports twenty of them: eleven where two failures of a dependence cancel and nine where they compound. Every one of those numbers is measured at a Spearman rank correlation of 0.40.
Its own last paragraph says why that matters. The leak is known to grow with the correlation as well as with the skew, so whether a cell’s answer is a property of the pair or of the pair at one strength of dependence is a guess.
This field sweeps the correlation from 0.1 to 0.7 in steps of a tenth, recalibrating every copula to each target. Before any cell can be read, the two margins of the table have to be, and one of them does something the sweep was not expecting.
The two margins
A marginal’s own leak is the cell at that marginal under the Gaussian copula, where the dependence is as harmless as this collection’s copulas get. Across the sweep, for a covariate skewed at 0.75:
0.728%, 2.727%, 5.588%, 8.872%, 12.208%, 15.364%, 18.213%.
Rising at every step. The same is true of the three other skewed marginals, and the two symmetric ones leak exactly nothing at every correlation, which is the parity result the earlier fields established.
A copula’s own leak is the cell at that copula with a normal covariate. Across the sweep, for the lower-tail copula:
0.958%, 3.215%, 5.706%, 7.707%, 8.875%, 9.064%, 8.219%.
It rises to a maximum at a Spearman of 0.6 and falls.
Why the copula’s leak has to turn over
The maximum is not an accident of this copula, and the argument for it is short.
The leak is an interaction: the share of a product of two split indicators that a rule balancing the two main effects fails to remove. At a rank correlation of zero the two variables are independent, the product’s expectation factorises, and there is no interaction to leak.
At a rank correlation of one the copula is comonotone — each variable is a deterministic increasing function of the other — so the product of the two split indicators is a function of either one alone, and the main effects span it exactly. There is no interaction to leak there either.
A quantity that is zero at both ends of an interval and positive inside it has a maximum inside it. The only question is where, and the sweep puts it at 0.6 for the copulas that leak at all.
What the rule being balanced is
A word about what “leaks” means, because everything above is a share of one specific thing.
A trial balances a covariate and an outcome-relevant function of it. The functions on this table are splits: an indicator for the covariate being above a threshold, and the same for a second variable. Balancing their two main effects is easy; what a rule cannot balance without being told about it is their product, and the interaction the product carries is what a rule that balanced only the main effects leaves behind.
The leak is the share of that interaction the main effects fail to remove. It is computed as an integral against the joint law of the two ranks, on a grid, in closed form rather than by simulation — which is what makes a table of two hundred and ten cells affordable and what makes a reading of 10⁻¹⁶ meaningful.
Three rules are on the table, and this field reads the first two. A mean split cuts at the mean and leaks nothing when both the copula and the marginal are symmetric. A median split cuts at the median and leaks nothing ever. A threshold at a fixed value has no guarantee at all. The sweep is reported for the mean split, because it is the rule whose zero is conditional and therefore the rule whose answer can change.
What that does to a reading
The consequence for the table is not subtle, and it is the reason this essay comes first.
Every cell of the table is compared against a baseline built from the two margins: the copula’s own leak plus the marginal’s own leak, less the corner where both are at their base. If one margin rises and the other turns over, then the baseline itself turns over, and a cell’s excess against it is a difference between two curves of different shapes.
So a cell whose excess changes sign across the sweep may be doing so because the cell moved or because the baseline did. Distinguishing those is the field’s third essay, and it can only be done cell by cell.
It also means the correlation of 0.40 is not a neutral place to have measured. It is on the rising part of the copula margin, about two thirds of the way to its peak, and the peak is inside the range a practitioner would work in.
Both margins must turn over, and only one does so in range
The comonotone argument is applied to the copula margin and it applies to the other one just as completely, which changes what the asymmetry between them is.
At a rank correlation of one a Gaussian copula is comonotone too. Then the second variable is a deterministic increasing function of the first, so the second split indicator is the first at a shifted threshold, and their product is whichever of the two indicators has the higher cut — a function already in the span of the main effects. The leak is exactly zero.
So the marginal margin is zero at both ends of the interval as well, and it must have an interior maximum for the same reason the copula margin does. The two margins are not one rising curve and one turning curve. They are two turning curves, and the sweep stops before the second one turns.
The deceleration is already visible. The marginal margin’s increments across the sweep are 2.00, 2.86, 3.28, 3.34, 3.16 and 2.85 — rising to a maximum at a Spearman of about 0.45 and falling after it. The curve’s second derivative has changed sign inside the range measured; only its first derivative has not.
Where it turns is outside the sweep and the shape says it must turn sharply. Extending the increments by the quadratic the last three imply gives about 2.4, 1.9 and 1.2 over the remaining three tenths, so the margin would reach roughly 23.7% at a rank correlation of one rather than the zero the comonotone argument requires. The gentle deceleration in the sweep is nowhere near enough, so the fall has to happen abruptly somewhere above 0.7 — in exactly the range the field declines to enter because the copulas’ calibration strains there.
That is worth stating because it changes how the field’s own framing should be read. The margin does not grow without limit; it grows over the range a trial’s covariates plausibly occupy and collapses somewhere beyond it. And the cells built on it inherit the same shape, so the four that change sides inside the sweep are the leading edge of a turnover the whole table has to undergo rather than a special property of the two most skewed covariates.
Which correlations a reader is in
Seven points from 0.1 to 0.7 covers what a covariate pair in a trial plausibly looks like, and it is worth saying why the range stops where it does at each end.
Below 0.1 every leak is a fraction of a per cent and the differences between cells are smaller than the third decimal place. The table is uninformative there because there is almost nothing to fail to balance.
Above 0.7 the calibration starts to strain. The copulas are parameterised differently — a correlation-like parameter for two of them, a shape parameter for the lower-tail one, a scalar for the one with no tail dependence — and pushing all five to a rank correlation of 0.9 puts three of them near the edge of the range where a grid of the resolution used here resolves them. The sweep stops before that rather than reporting numbers whose calibration is worse than their differences.
The correlation the earlier field measured at, 0.40, sits in the middle. That is the right place to have chosen one point, and it is also the place where the copula margin is still rising steeply and two thirds of the way to its peak.
The reflection, which is a relabelling
One control has to hold across the whole sweep, and it is worth stating because it is the cleanest thing in the field.
A copula and its reflection — the lower-tail one and the upper-tail one — have the same rank correlation, the same Kendall tau and the same radial gap. With a symmetric marginal the reflection is a relabelling: reversing the sign of both variables maps one to the other and leaves a symmetric distribution where it was.
So their leaks with a normal covariate must be identical, not close. They are: 0.958%, 3.215%, 5.706%, 7.707%, 8.875%, 9.064%, 8.219% for both, agreeing to a part in 10¹².
That is checked at every correlation because it is the one way the sweep could be broken without any number looking odd. A calibration that drifted, a grid built at the wrong resolution, a target missed by a hundredth — any of those would show up first as two columns that ought to be identical and are not.
The calibration, checked at every target
The sweep is over a rank correlation, so each copula’s parameter has to be solved for at each target, and a sweep whose copulas are not actually at those correlations is a sweep over a label.
Each copula’s parameter is found by search, its grid is built at that parameter, and the grid’s own Spearman correlation is then computed from the grid rather than from the parameter. The two routes share no arithmetic. At every one of the seven targets and every one of the five copulas the computed correlation is within 0.002 of the target.
At a Spearman of 0.4 the parameters come out at 0.4158 for the Gaussian, 0.4283 for the heavy-tailed one, 2.6099 for the one with no tail dependence, and 0.7587 for the lower-tail copula and its reflection — five different numbers for one rank correlation, which is what makes the calibration necessary in the first place.
What the five copulas are
The sweep recalibrates five copulas at every target and it is worth knowing what they are, because three of them contribute nothing to this essay’s margin and everything to the next two.
The Gaussian copula is the base: radially symmetric, no dependence left in the tails, and the corner every cell’s additivity baseline is measured against. The heavy-tailed one is radially symmetric too, with dependence surviving into both tails. The one with no tail dependence is radially symmetric and has the least tail dependence of the three. All three leak exactly nothing with a symmetric covariate, at every correlation, because a radially symmetric copula and a symmetric marginal give a parity zero.
The lower-tail copula and its reflection are the two that are not radially symmetric, and they are the only two that appear in this essay’s copula margin. Their leak with a normal covariate is the same number, and it is the number that turns over.
So the copula margin is really one curve, drawn by two copulas that must agree, and the other three sit on zero. That is a thin margin to hang a turning point on — one shape of dependence, seen twice — and it is the honest limit of this essay. What the argument for the turning point rests on is the comonotone limit, which applies to any copula, and what the measurement rests on is one family.
The economy that makes seven correlations affordable
Thirty cells at seven correlations is two hundred and ten evaluations, and the reason that is a few minutes rather than an afternoon is inherited.
The earlier field’s construction builds each copula’s grid once and evaluates it at every marginal, because the grid is a joint law of the ranks and the marginal is a relabelling applied on top of it. Rebuilding the grid per marginal — which is what calling the two earlier fields’ own tables in a loop would do — costs six times as much for identical numbers.
This field inherits that and adds one loop outside it. Each correlation needs five calibrations and five grids; each grid is then read at six marginals and three rules. So the cost is linear in the number of correlations and the factor-of-six economy is preserved at every one.
That is why the sweep is seven points rather than three, and why the second essay can afford a second sweep four times finer over a narrow range.
Where the copula margin’s maximum actually is
The sweep is on a grid of tenths, so “the maximum is at 0.6” is a statement about seven points.
The three readings around it are 8.875% at 0.5, 9.064% at 0.6 and 8.219% at 0.7. The rise from 0.5 to 0.6 is a fifth of a point and the fall from 0.6 to 0.7 is eight tenths, so the curve is asymmetric about its peak — flat on the way up and steeper on the way down, which is what a quantity forced to zero at 1 and rising smoothly from zero at 0 should look like.
That means the true maximum is somewhere between 0.5 and 0.65 rather than at 0.6 exactly, and the grid cannot say where. It is reported as an interior maximum rather than as a location, because the location is not measured and the interiority is what the argument needs.
The same shape is worth noticing on the other side of the table: the marginal margin is still rising at 0.7 — 15.364% to 18.213% — with no sign of turning. It has no comonotone limit forcing it down, because a marginal’s leak is not an interaction between two variables at all; it is what a mean split of a skewed variable fails to be, and skewness does not go away as a dependence strengthens.
One margin has a reason to turn over and the other does not, and both behave accordingly.
What a table of thirty cells is for
It is worth restating the object once, because everything in this field is a reading of it and it has three dimensions rather than two.
Five copulas across, six marginals down, and three rules in depth. The five copulas are the joint laws of the two ranks; the six marginals are what the first variable is relabelled to before the split is taken; the three rules are what a trial balances on.
Of the thirty (copula, marginal) cells, ten are on a margin — the Gaussian column and the two symmetric rows — and twenty are live: a copula that is not the base with a marginal that is not the base. Those twenty are what “eleven cancel and nine compound” counts, and what “four change their answer” counts.
The two margins this essay reads are the Gaussian column and the normal row, and they are what the additivity baseline is built from: a cell’s predicted leak, if the two failures simply added, is its copula’s own leak plus its marginal’s own leak, less the corner where both are at base. Everything downstream is a cell measured against that.
A baseline built from two margins inherits both of their shapes, which is what makes this essay a prerequisite for the other two rather than a preliminary.
The one number that does not move
Across every copula, every marginal and every one of the seven correlations, the interaction a median split leaks is under 10⁻¹⁶.
That is arithmetic rather than a measurement. A centred median split takes the values ±½ and its square is a quarter identically, so the interaction term a rule balancing the two main effects fails to remove is identically zero whatever the joint law is. The field that first established it proves it and this one re-runs it across two hundred and ten combinations because a claim that cannot fail is the cheapest kind to check.
It is also the one thing in this field a practitioner can act on without qualification. Every other reading here is conditional on a correlation; that one is not.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A symmetry that was not enough — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness, symmetry
- Which tail the cut sits in — both name copula, covariate balance, interaction, kendall tau, marginal distribution, median split, quadrature, rank correlation, symmetry, tail dependence
- A copula that halves a marginal — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness, symmetry
- The symmetry the marginals could not show — both name copula, covariate balance, interaction, marginal distribution, parity, quadrature, rank correlation, symmetry, tail dependence
- A split survives what a mean does not — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
- Balancing a skewed covariate — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
Named objects
A flat tag is an object no other essay names yet.
Closed formCopulaCovariate balanceGaussian copulaInteractionKendall tauMarginal distributionMedian splitMonotone transformationParityQuadratureRank correlationSkewnessSymmetryTail dependence