Both halves of the dependence at once

Two failures that cancel

A mildly skewed covariate under a lower-tail copula leaks 0.002% of an interaction where each failure alone leaks eight and seven per cent. Turn the copula over and the same pair compounds.

Worth reading first: Which series does the moving · Balancing what is known in advance.

Two fields have now taken apart the same exact zero, and each varied one half of the dependence with the other held at its base.

The parity field holds the joint law of the ranks Gaussian and varies the covariate’s marginal: a mean split’s interaction zero is gone at a skewness of one, and the separating case is a marginal that is heavy-tailed and symmetric, where every zero holds exactly. The copula field holds the covariate normal and varies the copula: the mean’s zero needs the copula to be symmetric under reflection, and is 7.707% under a Clayton and under its reflection.

Both close on the same admission. What a skewed covariate under an asymmetric copula costs is not measured, and the natural guess — that the two leaks compound — is a guess.

It is a wrong guess, and it is wrong in both directions.

The table

Every number below is a share of an interaction that a balancing rule fails to remove — a squared multiple correlation, computed exactly on a grid of the copula rather than estimated from draws, so the zeros are zeros to twenty decimal places rather than small numbers.

Five copulas across, six marginals down, matched at a Spearman rank correlation of 0.40, the leak being the share of a mean split’s interaction the balancing rule fails to remove.

What a mean split leaves, with both halves varyingThe share of a mean split's interaction that survives the rule balancing it, at every copula and every marginal, matched at a Spearman correlation of 0.40. The three radially symmetric copulas leave exactly nothing with a symmetric covariate and rise steeply with the skew. The two asymmetric ones start at 7.707% and go opposite ways: the lower-tail copula falls to 0.002% at a skewness of 0.95 — the two failures cancel almost exactly, and a guarantee both fields report as broken is restored — while the upper-tail one climbs to 40.288%. And the heavy-tailed symmetric covariate, which leaks exactly nothing on its own, doubles what the asymmetric copulas leak: 14.229% against 7.707%.00.2000.400the covariate's marginal, by how skewed it isshare of the interaction leftnormalheavy tailsskew 0.95skew 2.26skew 4.75exponentialGaussiant(4)FrankClaytonturned overmatched at Spearman 0.40five copulas, six marginals
Fig. 1 The share of a mean split’s interaction that survives, at every copula and every marginal. The slider changes which rule the trial balances.

The three radially symmetric copulas leave exactly nothing with a symmetric covariate — that is the guarantee — and rise steeply with the skew: 21.539% under a Gaussian copula at a skewness of 2.26.

The two asymmetric ones start at 7.707% with a normal covariate, which is the copula field’s number, and then go opposite ways.

The lower-tail copula falls. At a skewness of 0.95 it leaks 0.002% — where adding the two failures would give 16.579%, and where each failure alone gives 8.872% and 7.707%. The guarantee both fields report as broken is, at that combination, restored.

The upper-tail copula climbs to 40.288% with an exponential covariate, against 30.489% for the two added.

What the matching holds fixed

Five copulas are not comparable until they are matched on something, and every number above is at a Spearman rank correlation of exactly 0.4000. That leaves each family’s own parameter free — 0.4158 for the Gaussian, 0.4283 for the t on four degrees of freedom, 2.6099 for the Frank, 0.7587 for the Clayton — and it does not match them on Kendall’s tau, which comes out at 0.2730, 0.2817, 0.2722 and 0.2749.

The field that measured how much that choice moves things reports it as under a point on every leak, so the readings here are not artefacts of which rank statistic was matched on. What the matching does buy is the one comparison the whole field rests on: a copula and its reflection have the same parameter, the same Spearman, the same tau and the same radial gap, so nothing separates them but direction.

Against additivity

Twenty cells have both halves of the dependence failing, and every one of them has an additive prediction: what the copula leaks with a symmetric covariate, plus what the covariate leaks under a Gaussian copula.

The two leaks do not add, in both directions. Each cell of the table where both halves of the dependence fail, against what the two failures would give if their leaks added — the copula's leak with a symmetric covariate, plus the covariate's leak under a Gaussian copula. The diagonal is additivity. 11 of the 20 cells fall below it and 9 rise above, so the guess that the two compound is not merely imprecise, it has the wrong sign on more than half the table. The extreme is strongly skewed under a Clayton copula, which leaks 6.162% where adding the two would give 33.405% — a shortfall of 27.24 percentage points. The largest excess is an exponential covariate under the same copula turned over, at 40.288% against 30.489%.
Fig. 2 Each cell against what adding the two failures would give. The diagonal is additivity; below it is cancelling.

Eleven of the twenty fall below the diagonal and nine rise above it. The extreme in each direction is on the same copula and its reflection: a strongly skewed covariate under the lower-tail copula leaks 6.162% against 33.405% predicted, a shortfall of 27.243 percentage points; an exponential covariate under the upper-tail one leaks 40.288% against 30.489%, an excess of 9.799.

So the guess is not merely imprecise. It has the wrong sign on more than half the table, and the errors are larger than most of the quantities being predicted.

What is not in the table

Two things are held fixed across all thirty cells and both would move the numbers.

The rank correlation. Everything is at 0.40. The parity field measures the leak growing with the correlation as well as with the skew, so a table at 0.6 would be a different table with, presumably, the same shape. Nothing here checks that.

And the interaction being balanced. All three rules balance a product of two covariates’ functions, and the outcome shapes the worst case runs over are the six this collection has used since the dictionary field. A different family of interactions would have a different worst case, and the ordering among the cells is not guaranteed to survive it.

Why an asymmetry has a direction

The mechanism is short and it is the whole of the finding.

Both halves of the dependence are asymmetries. A right-skewed marginal stretches the upper tail of the covariate and compresses the lower one. A lower-tail copula concentrates the dependence between the two covariates in the lower tail — where the marginal is compressed — and an upper-tail copula concentrates it where the marginal is stretched.

A mean split’s interaction zero fails when the two centred indicators are not orthogonal to their own product, and what breaks that orthogonality is a mismatch between where the mass is and where the dependence is. Two asymmetries pointing opposite ways put the mass and the dependence back together; two pointing the same way pull them further apart.

The same copula, turned over. A Clayton copula and its reflection, at the same Spearman correlation of 0.40 and the same Kendall tau of 0.275, against the covariate's marginal. With a symmetric covariate the two are the same number to nine decimals — 7.707% apiece — because the leak then depends on how much asymmetry the copula has and not on which way it points. Skew the covariate and they come apart: at a skewness of 2.26 the lower-tail copula leaves 3.431% and the upper-tail one 36.213%, a factor of 10.6. Both halves of the dependence are asymmetries and an asymmetry has a direction; a lower-tail copula concentrates the dependence where a right-skewed marginal is compressed and the two distortions partly undo each other, and an upper-tail one concentrates it where the marginal is stretched.
Fig. 3 A Clayton copula and its reflection against the covariate’s marginal — same rank correlation, same Kendall tau, same amount of asymmetry, opposite direction.

The comparison that isolates it is the copula against its own reflection. The two have identical Spearman correlations of 0.4000, identical Kendall taus of 0.2749, and identical radial gaps of 0.1359 — every rank-based summary of them is the same number, and one is the other turned over. With a symmetric covariate they leak the same 7.707%, because the leak then depends on how much asymmetry there is and not on which way it points. With a covariate skewed at 2.26 they leak 3.431% and 36.213%, a factor of 10.6.

Nothing about the copula’s own statistics says which of those a trial is in. The Spearman correlation is the same, the tau is the same, the amount of asymmetry is the same. What differs is the direction, and the direction only matters once the marginal has one too.

The cell that reads zero

One number deserves more than a line, because it is the only place in this collection where two broken guarantees produce an intact one.

A covariate skewed at 0.95 under a lower-tail copula leaks 0.002% of a mean split’s interaction. The quadrature’s own noise on this grid is around 10⁻¹⁸ for the quantities that are exactly zero, so 0.002% is not zero — it is two parts in a hundred thousand, four orders of magnitude above the noise floor and four orders below either failure alone.

It is a cancellation rather than a restoration, and the difference matters. The zero at the corner of the table is a theorem: a symmetric marginal under a radially symmetric copula gives an odd function whose interaction is orthogonal to both main effects, exactly, for any correlation. The near-zero here is two large quantities happening to be nearly equal at one skewness, and moving the skewness to 2.26 takes it to 3.431% and to 4.75 takes it to 6.162%. It is a crossing, and a crossing is not a guarantee.

Which is worth saying because the practical reading of a cancellation is the opposite of the reading of a theorem. A trial that lands near this cell is not protected; it is at a point where its exposure is small and its sensitivity to the skewness is at its largest, because the leak is passing through zero on its way from one sign of contribution to the other.

The zero set is a rectangle, and that is the theorem made visible

Before any cell is read for size, the table’s pattern of exact zeros is worth counting, because its shape is what two independent necessary conditions look like.

Three of the five copulas are radially symmetric and two of the six marginals are symmetric. The exact zeros are precisely the six cells where both hold — a 3 × 2 block — and every one of the other twenty-four is non-zero.

A rectangle is what a conjunction of two conditions produces, and it is checkable at a glance. Had the zeros formed any other shape — a diagonal, a scatter, a row and a column — the two conditions would not be independent and the parity argument would be wrong about one of them.

So the first thing the joint table establishes is not a size at all. It is that the two guarantees the two earlier fields found separately are exactly the two factors of one condition, with no third requirement hiding in the interaction and no cell where the conjunction fails to deliver.

Three kinds of cell, and additivity fails in all of them

The remaining twenty-four split by which asymmetry is present, and the split is worth having before the sizes are read.

  • Both present — an asymmetric copula with a skewed marginal: 8 cells. These are the ones where cancelling or compounding is even conceivable, and they are where the field’s headline lives.
  • The marginal only — a symmetric copula with a skewed marginal: 12 cells.
  • The copula only — an asymmetric copula with a symmetric marginal: 4 cells.

The natural expectation is that additivity is trivially exact in the last sixteen, since one of the two components is exactly zero and there is nothing for it to interact with.

It is not. The heavy-tailed symmetric covariate leaks 0.000% on its own and the Clayton copula leaks 7.707% on its own, and together they leak 14.229%the row with no skew in it at all, which is nearly double what additivity requires.

So the failure of additivity is not a property of two defects meeting. It happens where one of the two defects is exactly absent, which means the interaction is not a product of two asymmetries. It is a genuine three-way quantity involving the rule, the copula and the marginal, and no accounting that adds two margins can reach it.

Reading the columns instead of the rows

The table has a second reading and it is the one that reaches the most trials.

Read down a column and it says what a marginal costs under a given copula. Read across a row and it says what a copula costs at a given marginal. The second reading is the surprising one for the asymmetric copulas and the first is the surprising one for everything else: the three copulas that break nothing on their own put a factor of two between the same marginal’s leaks. A covariate skewed at 2.26 leaks 12.118% under a Frank copula and 23.640% under a t on four degrees of freedom, at the same rank correlation, with both copulas radially symmetric and both leaking exactly nothing with a symmetric covariate.

That is not an interaction between two failures — neither copula has failed anything — and it is the third essay of this field. It applies to every trial rather than to the ones with an asymmetric dependence, and it says the parity field’s number was as much a fact about the Gaussian copula as about the skew.

What this does to the two fields it joins

Neither field’s numbers change; both fields’ conclusions narrow.

The parity field’s reading was that a skewed covariate costs about a fifth of the interaction. It is 21.539% under a Gaussian copula and 12.118% under a Frank, at the same rank correlation — so the size is not a property of the skewness, it is a property of the skewness and the copula together, and the copula was held fixed at the one case where it contributes nothing.

The copula field’s reading was that an asymmetric copula costs 7.707%. With a mildly skewed covariate the same copula costs 0.002% and its reflection costs 25.996%. Neither is near 7.707, and 7.707 was measured at the one marginal where the marginal contributes nothing.

Both numbers are corners of a table rather than sizes of an effect. That is not a criticism of either measurement — a field that varies one thing has to hold the other somewhere — and it is exactly why the pair had to be run.

The cost of running both at once, and why it was worth it

There is a reason neither field ran this table and it is not oversight: a copula’s grid is expensive, and evaluating six marginals against five copulas by calling either field’s own table in a loop rebuilds each grid six times.

It does not have to. A copula is a joint law of the ranks and a marginal is a relabelling applied on top of it, so the grid is a function of the copula alone and every marginal is an evaluation against it. Calibrating and building each copula’s grid once and looping the marginals inside costs a sixth of the obvious implementation, which is what makes thirty cells affordable where five or six were.

That is a small piece of arithmetic and it is the whole reason the finding exists. Two fields each varied one thing because varying both looked like thirty times the work, and it is five times the work. The structure that makes it cheap — one half of the dependence being a transformation of the other’s output — is the same structure that makes the two halves separable in the first place, which is what both fields relied on to vary one of them.

The row that has no skew in it at all

One row of the table is not about skewness and is the sharpest thing in it.

The heavy-tailed symmetric covariate has a skewness of 0.000 and leaks 0.000% under every radially symmetric copula — it is the case the parity field identifies as the separating one, where every zero holds exactly and the conclusion drawn is that it is not normality the guarantees needed.

Under the two asymmetric copulas it leaks 14.229%, against 7.707% for a normal covariate under the same copula. A covariate that costs exactly nothing on its own doubles what the copula costs.

Nothing in the parity argument sees that coming, because parity is a statement about the marginal being odd and this one is. The next essay of this field is what that does to the conclusion, and it is the case where the two halves interact without either of them being skewed.

Where the two fields would have found it

It is worth asking whether either field could have seen this without the other, because the answer is a small lesson about how a two-way effect hides.

The parity field varies six marginals and reports the leak growing with the skewness. Every one of its readings is at a Gaussian copula, so its column is one of the five here — and nothing in a single column can say that the column depends on which copula it was taken at. A monotone rise in the leak with the skew looks like a complete answer.

The copula field varies five copulas and reports two of them breaking the zero. Every one of its readings is at a normal covariate, so its row is one of the six here — and nothing in a single row can say that the two asymmetric copulas differ from each other, because with a symmetric covariate they do not.

Each field’s variation was invisible to the other’s, and each produced a clean monotone story. That is the ordinary shape of a two-way effect measured one way at a time, and the only thing that finds it is running both.

What a trial can do with it

Less than a reader would like, and the honest version is worth stating.

The direction of a copula’s asymmetry is not something a trial observes. It is a property of the joint law of two covariates, and a trial has one sample of it; estimating which tail the dependence is concentrated in needs more data than estimating a rank correlation does, and a rank correlation is already what the matching here is done on.

The direction of a marginal’s skew is observable — it is a sample skewness, and it is the one quantity in this whole line of fields a practitioner can read off their own covariate.

So the usable form is a warning rather than a correction. A trial with a skewed covariate cannot be told what its exposure is from the skewness alone, because the same skewness costs 0.002% or 25.996% depending on a property of the dependence nobody measures. And the direction it errs in cannot be signed either: the leak is sometimes smaller than the marginal alone would suggest and sometimes twice it.

What a trial is exposed to, over the shapes it does not know. The worst case of the rule every trial runs — a mean and a median split of each covariate — over six shapes the outcome's interaction might take, at three marginals and five copulas. Drawn as what the rule removes of its worst shape, so a full bar is a guarantee and a short one is exposure. With a symmetric covariate under a radially symmetric copula it removes all of it: 100.00% of the worst shape, which is the exact zero the earlier fields report. With both halves failing it removes 95.48% — so 4.52% of the worst interaction survives a rule that was proved to remove all of it. The exposure is not monotone in either half: the lower-tail copula with a skewed covariate is nearer the guarantee than a Gaussian one is.
Fig. 4 What the rule every trial runs is exposed to, over six shapes the outcome might take, at every cell of the table.

The worst case over the outcome shapes a trial does not know is the number to act on, and it runs from 0.0000% left with both halves symmetric to 4.5205% with both failing in the same direction — and back down to 0.0002% where they fail in opposite directions.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A zero that rests on a symmetry — both name covariate adjustment, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, numerical methods, parity, skewness
  • A split survives what a mean does not — both name covariate adjustment, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
  • The cut that is not a quantile — both name continuous covariate, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
  • A zero that is arithmetic — both name covariate adjustment, covariate balance, interaction, marginal distribution, median split, numerical methods, parity, symmetry
  • A cut is not a polynomial, and it does not have to be — both name closed form, continuous covariate, covariate balance, experimental design, interaction
  • A dictionary that is neither — both name closed form, covariate adjustment, covariate balance, experimental design, interaction

Named objects

A flat tag is an object no other essay names yet.

Closed formContinuous covariateCovariate adjustmentCovariate balanceExperimental designGaussian copulaInteractionMarginal distributionMedian splitMonotone transformationNumerical methodsParitySkewnessSymmetryTreatment effect