The same table at seven correlations

A margin that turns over

A skewed covariate's leak grows without limit as the dependence strengthens. A copula's own leak does not — it peaks at a rank correlation of 0.6 and falls. The margin of the table turns over before any cell in it does.

Worth reading first: Which series does the moving · Randomisation is not balance.

The field that varied a copula and a marginal together builds a table of thirty cells and reports twenty of them: eleven where two failures of a dependence cancel and nine where they compound. Every one of those numbers is measured at a Spearman rank correlation of 0.40.

Its own last paragraph says why that matters. The leak is known to grow with the correlation as well as with the skew, so whether a cell’s answer is a property of the pair or of the pair at one strength of dependence is a guess.

This field sweeps the correlation from 0.1 to 0.7 in steps of a tenth, recalibrating every copula to each target. Before any cell can be read, the two margins of the table have to be, and one of them does something the sweep was not expecting.

The two margins

A marginal’s own leak is the cell at that marginal under the Gaussian copula, where the dependence is as harmless as this collection’s copulas get. Across the sweep, for a covariate skewed at 0.75:

0.728%, 2.727%, 5.588%, 8.872%, 12.208%, 15.364%, 18.213%.

Rising at every step. The same is true of the three other skewed marginals, and the two symmetric ones leak exactly nothing at every correlation, which is the parity result the earlier fields established.

A copula’s own leak is the cell at that copula with a normal covariate. Across the sweep, for the lower-tail copula:

0.958%, 3.215%, 5.706%, 7.707%, 8.875%, 9.064%, 8.219%.

It rises to a maximum at a Spearman of 0.6 and falls.

Why the copula’s leak has to turn over

The maximum is not an accident of this copula, and the argument for it is short.

The leak is an interaction: the share of a product of two split indicators that a rule balancing the two main effects fails to remove. At a rank correlation of zero the two variables are independent, the product’s expectation factorises, and there is no interaction to leak.

At a rank correlation of one the copula is comonotone — each variable is a deterministic increasing function of the other — so the product of the two split indicators is a function of either one alone, and the main effects span it exactly. There is no interaction to leak there either.

A quantity that is zero at both ends of an interval and positive inside it has a maximum inside it. The only question is where, and the sweep puts it at 0.6 for the copulas that leak at all.

What the rule being balanced is

A word about what “leaks” means, because everything above is a share of one specific thing.

A trial balances a covariate and an outcome-relevant function of it. The functions on this table are splits: an indicator for the covariate being above a threshold, and the same for a second variable. Balancing their two main effects is easy; what a rule cannot balance without being told about it is their product, and the interaction the product carries is what a rule that balanced only the main effects leaves behind.

The leak is the share of that interaction the main effects fail to remove. It is computed as an integral against the joint law of the two ranks, on a grid, in closed form rather than by simulation — which is what makes a table of two hundred and ten cells affordable and what makes a reading of 10⁻¹⁶ meaningful.

Three rules are on the table, and this field reads the first two. A mean split cuts at the mean and leaks nothing when both the copula and the marginal are symmetric. A median split cuts at the median and leaks nothing ever. A threshold at a fixed value has no guarantee at all. The sweep is reported for the mean split, because it is the rule whose zero is conditional and therefore the rule whose answer can change.

What that does to a reading

The consequence for the table is not subtle, and it is the reason this essay comes first.

Every cell of the table is compared against a baseline built from the two margins: the copula’s own leak plus the marginal’s own leak, less the corner where both are at their base. If one margin rises and the other turns over, then the baseline itself turns over, and a cell’s excess against it is a difference between two curves of different shapes.

So a cell whose excess changes sign across the sweep may be doing so because the cell moved or because the baseline did. Distinguishing those is the field’s third essay, and it can only be done cell by cell.

It also means the correlation of 0.40 is not a neutral place to have measured. It is on the rising part of the copula margin, about two thirds of the way to its peak, and the peak is inside the range a practitioner would work in.

Both margins must turn over, and only one does so in range

The comonotone argument is applied to the copula margin and it applies to the other one just as completely, which changes what the asymmetry between them is.

At a rank correlation of one a Gaussian copula is comonotone too. Then the second variable is a deterministic increasing function of the first, so the second split indicator is the first at a shifted threshold, and their product is whichever of the two indicators has the higher cut — a function already in the span of the main effects. The leak is exactly zero.

So the marginal margin is zero at both ends of the interval as well, and it must have an interior maximum for the same reason the copula margin does. The two margins are not one rising curve and one turning curve. They are two turning curves, and the sweep stops before the second one turns.

The deceleration is already visible. The marginal margin’s increments across the sweep are 2.00, 2.86, 3.28, 3.34, 3.16 and 2.85 — rising to a maximum at a Spearman of about 0.45 and falling after it. The curve’s second derivative has changed sign inside the range measured; only its first derivative has not.

Where it turns is outside the sweep and the shape says it must turn sharply. Extending the increments by the quadratic the last three imply gives about 2.4, 1.9 and 1.2 over the remaining three tenths, so the margin would reach roughly 23.7% at a rank correlation of one rather than the zero the comonotone argument requires. The gentle deceleration in the sweep is nowhere near enough, so the fall has to happen abruptly somewhere above 0.7 — in exactly the range the field declines to enter because the copulas’ calibration strains there.

That is worth stating because it changes how the field’s own framing should be read. The margin does not grow without limit; it grows over the range a trial’s covariates plausibly occupy and collapses somewhere beyond it. And the cells built on it inherit the same shape, so the four that change sides inside the sweep are the leading edge of a turnover the whole table has to undergo rather than a special property of the two most skewed covariates.

Four cells change their answer. The four cells of the twenty whose excess changes sign as the dependence strengthens, over 7 recalibrations. Above the line the two failures compound — the cell leaks more than adding the copula's own leak and the marginal's — and below it they cancel. All four start above and end below, and all four are at the two most skewed covariates: skew 0.90 under heavy-tailed, skew 0.95 under heavy-tailed, skew 0.90 under upper tail, skew 0.95 under upper tail. Whether two failures of a dependence compound or cancel is therefore not a property of the pair. It is a property of the pair at a strength of dependence, and a fifth of the table changes its answer inside the range measured here.
Fig. 1 The four cells of the twenty whose excess changes sign as the dependence strengthens, over seven recalibrations. Above the line the two failures compound and below it they cancel; all four start above and end below, and all four are at the two most skewed covariates.

Which correlations a reader is in

Seven points from 0.1 to 0.7 covers what a covariate pair in a trial plausibly looks like, and it is worth saying why the range stops where it does at each end.

Below 0.1 every leak is a fraction of a per cent and the differences between cells are smaller than the third decimal place. The table is uninformative there because there is almost nothing to fail to balance.

Above 0.7 the calibration starts to strain. The copulas are parameterised differently — a correlation-like parameter for two of them, a shape parameter for the lower-tail one, a scalar for the one with no tail dependence — and pushing all five to a rank correlation of 0.9 puts three of them near the edge of the range where a grid of the resolution used here resolves them. The sweep stops before that rather than reporting numbers whose calibration is worse than their differences.

The correlation the earlier field measured at, 0.40, sits in the middle. That is the right place to have chosen one point, and it is also the place where the copula margin is still rising steeply and two thirds of the way to its peak.

How many cells cancel, by how strong the dependence is. How many of the twenty live cells fall below what adding their two halves gives, at each of 7 rank correlations. It rises from 9 at a Spearman of 0.1 to 13 at 0.7, and the eleven the earlier field reports is the reading at 0.4 — the third value along, and not a stable one. Every cell that moves moves the same way, from compounding to cancelling, and every one is at one of the two most skewed covariates. The rest of the table keeps its answer at every correlation in the sweep.
Fig. 2 How many of the twenty live cells fall below what adding their two halves gives, at each of seven rank correlations. It rises from 9 at a Spearman of 0.1 to 13 at 0.7, and the eleven the earlier field reports is the reading at 0.4 — the third value along, and not a stable one.

The reflection, which is a relabelling

One control has to hold across the whole sweep, and it is worth stating because it is the cleanest thing in the field.

A copula and its reflection — the lower-tail one and the upper-tail one — have the same rank correlation, the same Kendall tau and the same radial gap. With a symmetric marginal the reflection is a relabelling: reversing the sign of both variables maps one to the other and leaves a symmetric distribution where it was.

So their leaks with a normal covariate must be identical, not close. They are: 0.958%, 3.215%, 5.706%, 7.707%, 8.875%, 9.064%, 8.219% for both, agreeing to a part in 10¹².

That is checked at every correlation because it is the one way the sweep could be broken without any number looking odd. A calibration that drifted, a grid built at the wrong resolution, a target missed by a hundredth — any of those would show up first as two columns that ought to be identical and are not.

The calibration, checked at every target

The sweep is over a rank correlation, so each copula’s parameter has to be solved for at each target, and a sweep whose copulas are not actually at those correlations is a sweep over a label.

Each copula’s parameter is found by search, its grid is built at that parameter, and the grid’s own Spearman correlation is then computed from the grid rather than from the parameter. The two routes share no arithmetic. At every one of the seven targets and every one of the five copulas the computed correlation is within 0.002 of the target.

At a Spearman of 0.4 the parameters come out at 0.4158 for the Gaussian, 0.4283 for the heavy-tailed one, 2.6099 for the one with no tail dependence, and 0.7587 for the lower-tail copula and its reflection — five different numbers for one rank correlation, which is what makes the calibration necessary in the first place.

One zero is arithmetic and one is a symmetry. What two balancing rules remove of the interaction they are aimed at, on five joint laws of the ranks matched at a Spearman correlation of 0.4, with a normal covariate throughout. A rule holding a median split of each covariate removes exactly nothing of the product of the splits under every one of them, including the two that are not symmetric under reflection — and the reason is not a symmetry at all: a centred median split takes the values ±½, so its square is a quarter identically, and the interaction is orthogonal to both main effects whatever the joint law is. A rule holding the mean of each removes exactly nothing under the three radially symmetric copulas and 7.71% under the two that are not. Bars at the floor are exact zeros; the axis cannot draw 9e-32.
Fig. 3 The two zeros a copula can have, in the field that separated them.

What the five copulas are

The sweep recalibrates five copulas at every target and it is worth knowing what they are, because three of them contribute nothing to this essay’s margin and everything to the next two.

The Gaussian copula is the base: radially symmetric, no dependence left in the tails, and the corner every cell’s additivity baseline is measured against. The heavy-tailed one is radially symmetric too, with dependence surviving into both tails. The one with no tail dependence is radially symmetric and has the least tail dependence of the three. All three leak exactly nothing with a symmetric covariate, at every correlation, because a radially symmetric copula and a symmetric marginal give a parity zero.

The lower-tail copula and its reflection are the two that are not radially symmetric, and they are the only two that appear in this essay’s copula margin. Their leak with a normal covariate is the same number, and it is the number that turns over.

So the copula margin is really one curve, drawn by two copulas that must agree, and the other three sit on zero. That is a thin margin to hang a turning point on — one shape of dependence, seen twice — and it is the honest limit of this essay. What the argument for the turning point rests on is the comonotone limit, which applies to any copula, and what the measurement rests on is one family.

The economy that makes seven correlations affordable

Thirty cells at seven correlations is two hundred and ten evaluations, and the reason that is a few minutes rather than an afternoon is inherited.

The earlier field’s construction builds each copula’s grid once and evaluates it at every marginal, because the grid is a joint law of the ranks and the marginal is a relabelling applied on top of it. Rebuilding the grid per marginal — which is what calling the two earlier fields’ own tables in a loop would do — costs six times as much for identical numbers.

This field inherits that and adds one loop outside it. Each correlation needs five calibrations and five grids; each grid is then read at six marginals and three rules. So the cost is linear in the number of correlations and the factor-of-six economy is preserved at every one.

That is why the sweep is seven points rather than three, and why the second essay can afford a second sweep four times finer over a narrow range.

Where the copula margin’s maximum actually is

The sweep is on a grid of tenths, so “the maximum is at 0.6” is a statement about seven points.

The three readings around it are 8.875% at 0.5, 9.064% at 0.6 and 8.219% at 0.7. The rise from 0.5 to 0.6 is a fifth of a point and the fall from 0.6 to 0.7 is eight tenths, so the curve is asymmetric about its peak — flat on the way up and steeper on the way down, which is what a quantity forced to zero at 1 and rising smoothly from zero at 0 should look like.

That means the true maximum is somewhere between 0.5 and 0.65 rather than at 0.6 exactly, and the grid cannot say where. It is reported as an interior maximum rather than as a location, because the location is not measured and the interiority is what the argument needs.

The same shape is worth noticing on the other side of the table: the marginal margin is still rising at 0.7 — 15.364% to 18.213% — with no sign of turning. It has no comonotone limit forcing it down, because a marginal’s leak is not an interaction between two variables at all; it is what a mean split of a skewed variable fails to be, and skewness does not go away as a dependence strengthens.

One margin has a reason to turn over and the other does not, and both behave accordingly.

A near-perfect cancellation, at one correlation. A covariate skewed at 0.75 under a lower-tail copula, at each of 7 rank correlations. The earlier field measures this cell at a Spearman of 0.4 and reads 0.002% where adding the two halves' own leaks gives 16.579% — a cancellation so near exact that it is that field's headline. Across the sweep the same cell reads 0.0084%, 0.0182%, 0.0097%, 0.0015%, 0.0864%, 0.4693%, 1.5327%. Its smallest value is at 0.4, in the interior, and by 0.7 it is 1007.40 times larger. The near-zero is where two curves cross, and they cross beside the one correlation that was measured.
Fig. 4 One cell across the sweep: a covariate skewed at 0.75 under a lower-tail copula. It reads 0.0084%, 0.0182%, 0.0097%, 0.0015%, 0.0864%, 0.4693%, 1.5327% — smallest in the interior at 0.4, and a thousandfold larger by 0.7. A near-zero is where two curves cross.

What a table of thirty cells is for

It is worth restating the object once, because everything in this field is a reading of it and it has three dimensions rather than two.

Five copulas across, six marginals down, and three rules in depth. The five copulas are the joint laws of the two ranks; the six marginals are what the first variable is relabelled to before the split is taken; the three rules are what a trial balances on.

Of the thirty (copula, marginal) cells, ten are on a margin — the Gaussian column and the two symmetric rows — and twenty are live: a copula that is not the base with a marginal that is not the base. Those twenty are what “eleven cancel and nine compound” counts, and what “four change their answer” counts.

The two margins this essay reads are the Gaussian column and the normal row, and they are what the additivity baseline is built from: a cell’s predicted leak, if the two failures simply added, is its copula’s own leak plus its marginal’s own leak, less the corner where both are at base. Everything downstream is a cell measured against that.

A baseline built from two margins inherits both of their shapes, which is what makes this essay a prerequisite for the other two rather than a preliminary.

The whole table at a correlation of 0.4. Every cell of the copula-by-marginal table at a Spearman correlation of 0.4, measured against what adding the copula's own leak and the marginal's own leak gives. A positive bar is two failures compounding and a negative one is two cancelling; 9 compound and 11 cancel. The largest compounding cell is exponential under upper tail, at 40.288% against 30.489% added; the largest cancellation is skew 0.95 under lower tail, at 6.162% against 33.405%. Read the same table at another correlation and four of these bars are on the other side of the line.
Fig. 5 The whole table at a Spearman of 0.4, each cell against what adding the copula’s own leak and the marginal’s gives. Nine compound and eleven cancel; the largest compounding cell is exponential under an upper-tail copula at 40.288% against 30.489% added, and the largest cancellation is skew 0.95 under a lower-tail copula at 6.162% against 33.405%.

The one number that does not move

Across every copula, every marginal and every one of the seven correlations, the interaction a median split leaks is under 10⁻¹⁶.

That is arithmetic rather than a measurement. A centred median split takes the values ±½ and its square is a quarter identically, so the interaction term a rule balancing the two main effects fails to remove is identically zero whatever the joint law is. The field that first established it proves it and this one re-runs it across two hundred and ten combinations because a claim that cannot fail is the cheapest kind to check.

It is also the one thing in this field a practitioner can act on without qualification. Every other reading here is conditional on a correlation; that one is not.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A symmetry that was not enough — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness, symmetry
  • Which tail the cut sits in — both name copula, covariate balance, interaction, kendall tau, marginal distribution, median split, quadrature, rank correlation, symmetry, tail dependence
  • A copula that halves a marginal — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness, symmetry
  • The symmetry the marginals could not show — both name copula, covariate balance, interaction, marginal distribution, parity, quadrature, rank correlation, symmetry, tail dependence
  • A split survives what a mean does not — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
  • Balancing a skewed covariate — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness

Named objects

A flat tag is an object no other essay names yet.

Closed formCopulaCovariate balanceGaussian copulaInteractionKendall tauMarginal distributionMedian splitMonotone transformationParityQuadratureRank correlationSkewnessSymmetryTail dependence