The same table at seven correlations

An answer that changes

Eleven of twenty cells cancel and nine compound, at one rank correlation. Sweep the correlation and four of the twenty change sides — all four from compounding to cancelling, all four at the most skewed covariates.

Worth reading first: Which series does the moving · Randomisation is not balance.

The field this one re-reads reports a count: of its twenty live cells, eleven fall below what adding the two halves’ own leaks gives and nine rise above. Eleven cancel and nine compound.

Read at seven rank correlations from 0.1 to 0.7, the count of cells that cancel is

9, 9, 9, 11, 11, 12, 13.

Four cells change sides, and the eleven is the reading at 0.4 — the fourth value along, and not a stable one.

What varying the correlation is

One clarification, because “sweeping the correlation” could mean two different operations and only one of them is a sweep.

Every copula here has its own parameter, and those parameters are not comparable: a Gaussian copula’s correlation parameter and a lower-tail copula’s shape parameter are different objects on different scales. Moving all five by the same amount would move them to five different dependences.

So what is swept is the rank correlation, and each copula’s parameter is solved for at each target. The field that established this discipline is explicit about why: five copulas matched on Spearman correlation are five different dependences with one thing held, and any difference between them is then a difference in shape rather than in strength.

At a Spearman of 0.4 the five parameters come out at 0.4158, 0.4283, 2.6099 and 0.7587 twice — and every one is checked, after the grid is built, by computing the grid’s own Spearman correlation and requiring it to be within 0.002 of the target. That check runs at all seven correlations.

The four

cell signs across the sweep
a heavy-tailed copula, covariate skewed at 0.90 +++++--
a heavy-tailed copula, covariate skewed at 0.95 +++----
an upper-tail copula, covariate skewed at 0.90 ++++++-
an upper-tail copula, covariate skewed at 0.95 +++----

Every one goes from compounding to cancelling, and none goes the other way. Every one is at one of the two most skewed covariates in the table.

Four cells change their answer. The four cells of the twenty whose excess changes sign as the dependence strengthens, over 7 recalibrations. Above the line the two failures compound — the cell leaks more than adding the copula's own leak and the marginal's — and below it they cancel. All four start above and end below, and all four are at the two most skewed covariates: skew 0.90 under heavy-tailed, skew 0.95 under heavy-tailed, skew 0.90 under upper tail, skew 0.95 under upper tail. Whether two failures of a dependence compound or cancel is therefore not a property of the pair. It is a property of the pair at a strength of dependence, and a fifth of the table changes its answer inside the range measured here.
Fig. 1 The four cells whose excess changes sign, across the sweep. Above the line the two failures compound; below it they cancel.

What the excesses do

The sign is a summary; the paths are the finding.

The heavy-tailed copula with a covariate skewed at 0.95 runs +9.325%, +5.717%, +1.723%, −1.197%, −2.800%, −3.316%, −3.050% — crossing between 0.3 and 0.4 and then flattening. The upper-tail copula with the same covariate runs +7.957%, +7.712%, +3.979%, −0.141%, −3.475%, −5.545%, −6.174%, crossing at almost exactly 0.4.

The two at a skew of 0.90 cross later: +6.974% down to −0.562% for the heavy-tailed copula, crossing between 0.5 and 0.6, and +6.660% up to +10.866% and back down to −0.615% for the upper-tail one, crossing between 0.6 and 0.7.

So two of the four cross near where the earlier field measured and two cross above it. The correlation of 0.4 is not a special point for these cells; it is a point that happens to have two of them on each side.

Where the crossings are, and where they are not

The four cells cross at four different places, and the spread is worth reading because it says the crossings are not one event seen four times.

The two at a skew of 0.95 cross between 0.3 and 0.4 — the heavy-tailed copula’s excess goes +1.723% to −1.197%, and the upper-tail copula’s +3.979% to −0.141%. The two at a skew of 0.90 cross later: between 0.5 and 0.6 for the heavy-tailed copula, +0.616% to −0.240%, and between 0.6 and 0.7 for the upper-tail one, +1.085% to −0.615%.

So the more skewed the covariate, the earlier its cells cross. That is the mechanism read directly: a more skewed covariate has a faster-growing marginal margin, the baseline overtakes the cell sooner, and the excess crosses at a weaker dependence.

It also means that at any correlation in the sweep at least two of the four are on each side, which is why the count of cancelling cells moves by one or two rather than by four. The four cells do not flip together; they flip in order of skew.

The four crossings, located

The sweep brackets each crossing between two grid points and interpolating puts numbers on them.

The heavy-tailed copula at a skew of 0.95 goes from +1.723% to −1.197% between Spearman 0.3 and 0.4, crossing at 0.359. The upper-tail copula at the same skew goes from +3.979% to −0.141%, crossing at 0.397. At a skew of 0.90 the heavy-tailed cell crosses at 0.572 and the upper-tail one at 0.664.

Read as a two-by-two, the skew moves the crossing by about 0.22 and the copula moves it by 0.04 to 0.09. So the covariate’s shape is three to five times more decisive than the dependence’s shape in deciding where a cell changes sides — which is the opposite of the emphasis a table indexed by copula invites.

That also explains why the count moves by one or two rather than by four. The crossings are spread across a range of 0.3 in Spearman, so no two grid points contain more than two of them, and the four cells flip in order of skew rather than together.

Sixteen cells never move

The instability is confined and it is worth saying how confined.

Of twenty live cells, sixteen keep their sign across the whole sweep — a sevenfold range of dependence — and the four that move are the two most skewed covariates crossed with two of the five copulas. They are a two-by-two corner of the table.

So the earlier field’s eleven-against-nine is 80% a property of the table and 20% a property of the correlation it was read at, and the conditional fifth is identifiable in advance: it is where a covariate’s marginal leak is growing fastest and the copula’s cell is nearest its own peak.

The growth figures say the same thing in one line. Across the sweep the most skewed covariate’s baseline grows by a factor of 9.7 and the heavy-tailed cell built on it by 2.5. A cell growing four times more slowly than the thing it is measured against is a cell that will cross, and the only question is where — which the crossings above answer, and which the skew decides.

Why they all move the same way

Four cells moving in one direction is a mechanism rather than four accidents, and the mechanism is the margin that turns over.

A cell’s baseline is its copula’s own leak plus its marginal’s own leak. The marginal margin rises without limit — 3.727% to 36.056% for the most skewed covariate. The copula margin peaks at a Spearman of 0.6 and falls.

But the two copulas whose cells move are radially symmetric and asymmetric respectively, and neither has the lower-tail copula’s shape. The heavy-tailed copula leaks exactly nothing with a symmetric covariate at every correlation, so its cells’ baselines are the marginal’s leak alone; the upper-tail copula’s is the lower-tail one’s, which turns over.

What both have in common is that their cells grow more slowly than their baselines at the top of the sweep. The heavy-tailed copula with the most skewed covariate runs from 13.052% to 33.006% across the sweep — a factor of 2.5 — while its baseline runs from 3.727% to 36.056%, a factor of 9.7. The baseline overtakes the cell, and the excess crosses.

A cell that saturates against a baseline that does not is a cell that will change sides, and the two most skewed covariates are the ones whose baselines grow fastest.

Why the cells saturate is the other half, and it is the same comonotone argument the copula margin obeys. A cell is a leak of an interaction, and at a rank correlation of one there is no interaction left: the two variables are a deterministic function of each other and the main effects span the product exactly. So every cell in the table is heading back towards zero eventually, and the ones that turn first are the ones already nearest their peak.

What “compound” and “cancel” mean here

The two words carry the field, so it is worth being exact about the baseline they are measured against.

A cell’s predicted leak, if the two failures simply added, is its copula’s own leak with a symmetric covariate, plus its marginal’s own leak under the Gaussian copula, less the corner where both are at base — which for a mean split is zero. The excess is the measured cell less that prediction.

A positive excess means the pair leaks more than adding gives: two failures compounding. A negative excess means it leaks less: two failures cancelling.

Neither word is a statement about how much a cell leaks. A cell can compound at 13.052% against a baseline of 3.727% and be less dangerous than one that cancels at 25.996% against 33.405%, and both of those are on this table at a Spearman of 0.4. The sign is about whether the two failures reinforce; the level is about whether a trial has a problem.

That distinction matters more than usual here because the four cells that change sign do not change level in any interesting way. Every one of them leaks more at 0.7 than at 0.1, by factors of 2.5 to 3.7. What changed is the thing they are being compared against.

Which cells do not move

Sixteen of the twenty keep their answer at every correlation, and the pattern is clean enough to state.

The copula with no tail dependence cancels at all five of its marginals, at every correlation. Its cells run well below their baselines throughout: a covariate skewed at 0.90 leaks 12.118% at 0.4 against a baseline of 21.5%.

The lower-tail copula cancels at all four of its skewed marginals and compounds at the symmetric heavy-tailed one, at every correlation. The symmetric heavy-tailed covariate is the case the parity fields built to show that normality was not what the guarantees needed, and it is the one marginal on which this copula’s cells go the other way.

The upper-tail copula compounds at three of its five, including the exponential one, at every correlation.

And the heavy-tailed copula compounds at three of its five throughout.

So the twenty cells are twelve that never move, four that never move and read zero throughout because of parity, and four that cross. The pattern of who moves is not random and it is not everywhere.

How many cells cancel, by how strong the dependence is. How many of the twenty live cells fall below what adding their two halves gives, at each of 7 rank correlations. It rises from 9 at a Spearman of 0.1 to 13 at 0.7, and the eleven the earlier field reports is the reading at 0.4 — the third value along, and not a stable one. Every cell that moves moves the same way, from compounding to cancelling, and every one is at one of the two most skewed covariates. The rest of the table keeps its answer at every correlation in the sweep.
Fig. 2 The count of cells that cancel, at each correlation, with the point the earlier field measured at marked.
The whole table at a correlation of 0.4Every cell of the copula-by-marginal table at a Spearman correlation of 0.4, measured against what adding the copula's own leak and the marginal's own leak gives. A positive bar is two failures compounding and a negative one is two cancelling; 9 compound and 11 cancel. The largest compounding cell is exponential under upper tail, at 40.288% against 30.489% added; the largest cancellation is skew 0.95 under lower tail, at 6.162% against 33.405%. Read the same table at another correlation and four of these bars are on the other side of the line.the two failures addheavy, symmetric under heavy-tailed0.000%skew 0.75 under heavy-tailed1.732%skew 0.90 under heavy-tailed2.101%skew 0.95 under heavy-tailed-1.197%exponential under heavy-tailed3.625%heavy, symmetric under no tail dependence-0.000%skew 0.75 under no tail dependence-4.276%skew 0.90 under no tail dependence-9.422%skew 0.95 under no tail dependence-10.383%exponential under no tail dependence-9.498%heavy, symmetric under lower tail6.522%skew 0.75 under lower tail-16.577%skew 0.90 under lower tail-25.816%skew 0.95 under lower tail-27.243%exponential under lower tail-25.946%heavy, symmetric under upper tail6.522%skew 0.75 under upper tail9.417%skew 0.90 under upper tail6.966%skew 0.95 under upper tail-0.141%exponential under upper tail9.799%20 cells at a Spearman of 0.4positive is compounding
Fig. 3 And the whole table at one correlation. The slider changes which correlation, and four of the bars change side.

Why the count is a poor summary

“Eleven cancel and nine compound” is the earlier field’s headline for the table, and this sweep is a reason to prefer almost any other summary.

It throws away the size. A cell whose excess is −0.141% and one whose excess is −6.174% both count as one cancel, and the first is within a rounding error of additive while the second is a substantial departure.

It is unstable exactly where the cells are near zero. The four cells that move are the four whose excess passes through zero inside the range, so the count changes by one every time one of them crosses. A count of a sign is at its least informative near the sign change and its most informative far from it, which is the opposite of what a summary should be.

And it hides that the sign is not the interesting axis. Nineteen of the twenty cells leak more at a stronger correlation. That is the fact a trial should act on, and it is invisible in a count of signs because it is true on both sides of the line.

The reading this field would substitute is the paths: each cell’s leak and each cell’s baseline, across the range, with the crossings marked. That is twenty curves rather than two numbers, which is why the earlier field reported two numbers, and it is why the figures here draw the four that move rather than all twenty.

The two symmetric rows, which are the control

Four of the thirty cells never move because they cannot, and they are what says the sweep is measuring the joint law rather than the machinery.

The normal covariate and the heavy-tailed symmetric one are both symmetric, and three of the five copulas are radially symmetric. A radially symmetric copula with a symmetric marginal leaks exactly nothing under a mean split — that is the parity result, and it is an argument rather than a measurement.

Across the sweep those cells read 0.0000% at every one of the seven correlations, to the resolution the quadrature reports. A calibration that had drifted, a grid built at the wrong resolution, or a marginal relabelled with the wrong transform would show up there first, as a zero that stopped being one.

The two asymmetric copulas with a symmetric covariate are not zero — that is the copula margin, which is what the first essay of this field reads — and the two symmetric copulas with a skewed covariate are not zero either, which is the marginal margin. The zeros are exactly where parity says they should be and nowhere else, at every correlation.

The extremes, which also move

The two cells the earlier field singles out are worth following, because one of them is not the same cell at every correlation.

At a Spearman of 0.4 the largest cancellation is the lower-tail copula with a covariate skewed at 0.95 — 6.162% against a baseline of 33.405% — and the largest compounding is the upper-tail copula with an exponential covariate, 40.288% against 30.489%.

At 0.1 the largest cancellation is the same cell, at 0.376% against 4.686%, and the largest compounding is the heavy-tailed copula with a covariate skewed at 0.95, 13.052% against 3.727%.

At 0.7 the largest cancellation is the lower-tail copula with an exponential covariate — 17.575% against 48.656% — and the largest compounding is the lower-tail copula with the symmetric heavy-tailed covariate, 14.478% against 8.219%.

So neither extreme is held by one cell across the sweep. Both change hands, and the cell holding the largest compounding at 0.7 is one of the four that read a parity zero on its own.

What the sweep does not vary

Three things are held at every one of the seven correlations, and each of them could be a further dial.

The marginals. Six of them: normal, a heavy-tailed symmetric one, three from a skew family at 0.75, 0.90 and 0.95, and an exponential. The skewnesses are fixed as the correlation moves, so the sweep separates the two dials the earlier fields varied together — and it means nothing here says what happens when a skew is pushed further at a fixed correlation.

The rules. Three of them, and the field reads the mean split. The threshold at a fixed value has no guarantee at all and its numbers move for reasons that have nothing to do with the sweep; the median split reads zero everywhere and is checked to.

And the grid. Every integral is computed by quadrature at a fixed resolution, so a cell reading 10⁻¹⁶ is a statement about that resolution as much as about the integral. It is the resolution the earlier fields use and it resolves a median split’s zero to sixteen figures, which is the check that says the resolution is adequate for the smallest quantity anybody reads.

What is not held is the copulas’ parameters, and that is the point: five copulas at seven rank correlations is thirty-five calibrations, each solved by search and each checked against the grid it produced.

What a reader should take

The practical form is short and it is not the earlier field’s.

Cancellation is not a property of a pair. It is a property of a pair at a strength of dependence, and a fifth of this table changes its answer inside a range anybody would work in — with where exactly a cell crosses a question of its own.

Whether it cancels or compounds is worth less than how much it leaks. The cell that changes sides from +9.325% to −3.050% is a cell whose leak grows from 13.052% to 33.006% over the same sweep. A practitioner reading “these two failures cancel” would be reading a statement about a baseline they do not care about, while the quantity they do care about tripled.

And the one thing that does not move is the median split. Under every copula, every marginal and every correlation in the sweep, its interaction leak is under 10⁻¹⁶ — because a centred median split takes the values ±½ and its square is a quarter identically. That is arithmetic and no joint law can touch it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A symmetry that was not enough — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness, symmetry
  • The symmetry the marginals could not show — both name copula, covariate balance, interaction, marginal distribution, parity, quadrature, rank correlation, symmetry, tail dependence, worst case
  • A copula that halves a marginal — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness, symmetry
  • A split survives what a mean does not — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
  • A zero that rests on a symmetry — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
  • Balancing a skewed covariate — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness

Named objects

A flat tag is an object no other essay names yet.

Closed formCopulaCovariate balanceGaussian copulaInteractionMarginal distributionMedian splitMonotone transformationParityQuadratureRank correlationSkewnessSymmetryTail dependenceWorst case