Both halves of the dependence at once

A copula that halves a marginal

Three copulas break nothing on their own and put a factor of two between the same skewed covariate's leaks — 12.118% under a Frank against 23.640% under a t, at the same rank correlation.

Worth reading first: Which series does the moving · Balancing what is known in advance.

The two essays before this are about copulas that break a guarantee. This one is about three that do not, and it reaches more trials than either.

The parity field reports that a skewed covariate costs a mean split’s interaction zero, and prices it: at a skewness of 2.26 the rule leaves 21.539% of the interaction. Every one of its readings is at a Gaussian copula, because holding the joint law of the ranks fixed is what isolates the marginal.

Change the copula to another one that is also radially symmetric, also leaks exactly nothing with a symmetric covariate, and is also matched at a Spearman rank correlation of 0.40, and the same covariate leaves 12.118% or 23.640%.

Three copulas that break nothing

The three are not chosen to make a point. They are the three radially symmetric members of the family the copula field built to separate symmetric dependence from asymmetric, and the fourth and fifth members — the asymmetric pair — are what the first and second essays of this field are about.

Three copulas that break nothing, and a factor of two between them. The three radially symmetric copulas, at a matched Spearman correlation of 0.40, against the covariate's marginal. All three leave exactly nothing with a symmetric covariate — that is the guarantee, and it holds to twenty decimal places. What they do to a skewed covariate is not the same at all: at a skewness of 2.26 a Frank copula leaves 12.118% where a Gaussian leaves 21.539% and a t on four degrees of freedom leaves 23.640%. A factor of 2.0 between two copulas that are both symmetric, both matched on rank correlation, and both harmless on their own. So the copula matters to the marginal's leak without breaking any symmetry of its own, which is a milder version of the same finding and applies to every trial rather than to the asymmetric ones.
Fig. 1 The three radially symmetric copulas against the covariate’s marginal. All three leave exactly nothing with a symmetric covariate.

The Gaussian, the t on four degrees of freedom and the Frank copula are all symmetric under reflection: their radial gap is 0.0000, which is the statistic the copula field defines to separate them from the asymmetric pair. All three leave exactly nothing with a normal covariate and exactly nothing with the heavy-tailed symmetric one. Every guarantee the parity argument states holds under all three, to twenty decimal places.

What they do to a skewed covariate is not the same at all.

At a skewness of 0.95: 8.872%, 10.603%, 4.595%. At 2.26: 21.539%, 23.640%, 12.118%. At 4.75: 25.697%, 24.500%, 15.314%. With an exponential covariate: 22.782%, 26.407%, 13.284%.

The Frank copula halves it. At a skewness of 2.26 the ratio between the t and the Frank is 1.95, on two copulas that are both symmetric, both harmless alone, and both matched on the same rank statistic.

The comparison, stated so it can be dismissed

Three copulas differing by a factor of two is only interesting if the three are genuinely alternatives rather than a family with one member stretched. Four things are held equal across them.

The rank correlation, at exactly 0.4000 by calibration — each family’s own parameter is solved for rather than set, giving 0.4158, 0.4283 and 2.6099.

The radial symmetry, at a gap of 0.0000 for all three, so none of them has the asymmetry the other two essays of this field are about.

The marginal, which is the same relabelling applied on top of each grid — the copula is a joint law of the ranks and the marginal is a function of them, so the same six marginals are evaluated against all five copulas with nothing else changed.

And the target, which is the same product of two centred indicators throughout.

What is left varying is the shape of the dependence between the ranks, which is what a copula is. So the factor of two is a fact about that shape and about nothing else.

The one number that is the same

Across the three, one quantity does not move at all, and it is the check that the comparison is measuring what it claims.

With the heavy-tailed symmetric covariate, all three leak 0.000% — exactly, to twenty decimal places, the same as with a normal covariate. So the factor of two is a factor of two in what the skew costs, not a factor of two in some general looseness of the Frank copula. Where the parity argument applies, all three copulas honour it identically; where it does not, they differ by two.

That is the cleanest form of the finding: the copula multiplies the marginal’s failure and contributes none of its own.

What a factor of two does to the attribution

The number being halved is worth restating as the claim it was originally made as, because the halving changes what kind of claim it is.

21.539% was reported as what a skewed covariate costs. It was measured with the joint law of the ranks held Gaussian — deliberately, since holding the copula fixed is what isolates the marginal — and so it is a reading at one corner of a two-dimensional table.

A copula that halves it does not contradict the measurement. It shows that the measurement was of a pair: this marginal under that copula. The marginal’s cost is not a property the covariate carries around with it, and a trial that knows its covariate’s skewness knows one of the two things the number depends on.

Which is the awkward half, because the two are not equally observable. A skewness can be estimated from the covariate alone, on as many rows as there are. A copula is a statement about a joint law, and nothing in a marginal reports it.

What is different about them

The three differ in where the dependence sits, and the difference is legible even though none of them is asymmetric.

A Gaussian copula spreads the dependence evenly across the joint range. A t copula on four degrees of freedom concentrates it in both tails — that is what its tail dependence is, and it is symmetric because both tails get it equally. A Frank copula does the opposite: it concentrates the dependence in the middle and has no tail dependence at all.

A skewed covariate stretches one tail and compresses the other. The interaction leak measures how badly the mean split’s product lines up with the two main effects, and that misalignment is largest where the covariate’s mass is most distorted — in the tails. So a copula that puts its dependence in the tails is amplifying the marginal’s distortion, and one that puts it in the middle is putting it where the distortion is smallest.

That is a mechanism rather than a theorem, and it makes a prediction the table can check: the t should be highest, the Gaussian in between, the Frank lowest, at every skewed marginal.

Half of it holds and half does not. The Frank copula is lowest at all four — 4.595, 12.118, 15.314 and 13.284 against every other reading — which is the half the mechanism is really about, since it is the only one of the three with no dependence left in the tails a skewed marginal stretches. Which of the other two is higher is not settled: the t is above the Gaussian at three marginals and below it at the most skewed, where it reads 24.500 against 25.697.

So the reading is tail dependence costs a skewed covariate more, established against the copula that has none, and not more tail dependence costs more, which would need the t and the Gaussian to keep their order and they do not.

Where the middle copula puts its dependence

It is worth being concrete about the Frank copula, because it is the one that halves the leak and it is the least familiar of the three.

A Frank copula has no tail dependence in either tail — conditional on one covariate being extreme, the other is asymptotically independent of it — and it makes up its rank correlation entirely in the middle of the joint range. A Gaussian copula also has no tail dependence, but it approaches independence in the tails slowly, so there is still substantial dependence out where a skewed marginal has stretched its scale. A t copula on four degrees of freedom has genuine tail dependence in both tails.

So the three are ordered by how much of their dependence survives into the region a skewed marginal distorts most, and their leaks come out in that order at three of the four: 4.595, 8.872, 10.603 at a skewness of 0.95, and 12.118, 21.539, 23.640 at 2.26. At the most skewed the Gaussian and the t change places — 15.314, 25.697, 24.500 — and the Frank stays lowest.

The part of the mechanism that survives is therefore the part about having no tail dependence at all, which is the Frank’s alone. Ordering the other two by how much they have is a finer distinction than four points can support.

That is the mechanism and it is a mechanism about the interaction of two features rather than about either. A marginal’s skew decides where the distortion is; the copula decides how much dependence is there to be distorted.

Which makes the parity field’s number a corner

Nothing in the parity field is wrong and its headline is smaller than it reads.

A mean split’s interaction zero is gone at a skewness of one, and the rule leaves about a fifth of the interaction. The fifth is a fifth at a Gaussian copula. It is an eighth at a Frank and a quarter at a t, and those are the two nearest neighbours in a family the field never varied.

So the size of the leak is a property of the pair and the parity field’s figure is one cell of a table, in exactly the way the copula field’s 7.707% is one cell. The difference is that this one applies without either half of the dependence being asymmetric — it is a fact about every trial with a skewed covariate rather than about the ones with an unusual dependence.

The prediction, and where it nearly fails

The mass-in-the-tails argument makes one more prediction and it is worth checking rather than leaving as a story: the gaps between the three should widen as the skew grows.

The gaps do widen and then stop. From a skewness of 0.95 to 2.26 the Frank-to-t gap goes from 6.008 points to 11.523 — it roughly doubles, as predicted. From 2.26 to 4.75 it goes to 9.186, which is smaller, and that is where the t and the Gaussian cross.

The reason is that the most skewed covariate in the family is skewed enough that its mean split is badly unbalanced: the mean of a strongly right-skewed variable sits well above its median, so the indicator is a rare event and the interaction it forms has less to leak. The leak is not monotone in the skewness for any copula — the Gaussian’s own column runs 8.872, 21.539, 25.697 and then 22.782 for the exponential, whose skewness of 2.00 sits between the last two.

So one rung of the ordering is a finding and the rest is not. That is the right level of confidence for a mechanism supported by four points, and it is why the essay’s claim is about the copula with no tail dependence rather than about a ranking of three.

And it is worse than it looks, because nothing identifies the copula

The practical difficulty is that the three copulas are hard to tell apart from data and easy to tell apart in their consequences.

They are matched here on Spearman’s rank correlation, which is what a trial estimates. Their Kendall taus are 0.2730, 0.2817 and 0.2722 — within a hundredth of each other, on quantities a trial would need hundreds of pairs to separate. Their radial gaps are all zero. Every rank statistic anybody computes routinely says they are the same dependence.

Their leaks differ by a factor of two.

A trial cannot distinguish the copulas it is exposed to and the exposure differs by a factor of two between them. That is a sharper version of the difficulty than the asymmetric case, where at least the radial gap is estimable and would flag the problem.

What a mean split leaves, with both halves varyingThe share of a mean split's interaction that survives the rule balancing it, at every copula and every marginal, matched at a Spearman correlation of 0.40. The three radially symmetric copulas leave exactly nothing with a symmetric covariate and rise steeply with the skew. The two asymmetric ones start at 7.707% and go opposite ways: the lower-tail copula falls to 0.002% at a skewness of 0.95 — the two failures cancel almost exactly, and a guarantee both fields report as broken is restored — while the upper-tail one climbs to 40.288%. And the heavy-tailed symmetric covariate, which leaks exactly nothing on its own, doubles what the asymmetric copulas leak: 14.229% against 7.707%.00.2000.400the covariate's marginal, by how skewed it isshare of the interaction leftnormalheavy tailsskew 0.95skew 2.26skew 4.75exponentialGaussiant(4)FrankClaytonturned overmatched at Spearman 0.40five copulas, six marginals
Fig. 2 The whole table, where the three symmetric copulas are the three curves that start at zero. The slider changes which rule the trial balances.

What this says about matching on a rank correlation

Every comparison in this line of fields is made at a matched rank correlation, and that convention now needs a sentence, because this essay is the first place it does real work rather than merely making a table legible.

Matching on Spearman’s rho is the right convention: it is the quantity a trial can estimate, it is invariant to the marginals, and it is what a practitioner means by how dependent the covariates are. The whole point of using it is that two copulas at the same rho are two ways of having the same amount of dependence.

This essay is the measurement that says the amount of dependence is not what the exposure is a function of. Three copulas at the same rho, the same tau to within a hundredth, and the same zero radial gap, produce leaks differing by a factor of two. So how dependent is not enough, and how asymmetric — the radial gap, which the copula field introduced for the other two essays here — is not enough either, because it is zero for all three.

What would be enough is the shape of the dependence across the joint range, which has no one-number summary and is what a copula is. The field that priced the matching shows the choice between two rank statistics moving the leaks by under a point; the choice between two copulas at the same value of either moves them by twelve.

What the worst case does

The leak on one target is not what a trial is exposed to, because a trial does not know the shape of its outcome’s interaction. The quantity to act on is the worst case over the shapes it might take, and it is worth reading here because the ordering does not survive intact.

What a trial is exposed to, over the shapes it does not know. The worst case of the rule every trial runs — a mean and a median split of each covariate — over six shapes the outcome's interaction might take, at three marginals and five copulas. Drawn as what the rule removes of its worst shape, so a full bar is a guarantee and a short one is exposure. With a symmetric covariate under a radially symmetric copula it removes all of it: 100.00% of the worst shape, which is the exact zero the earlier fields report. With both halves failing it removes 95.48% — so 4.52% of the worst interaction survives a rule that was proved to remove all of it. The exposure is not monotone in either half: the lower-tail copula with a skewed covariate is nearer the guarantee than a Gaussian one is.
Fig. 3 What the rule every trial runs removes of its worst shape, at three marginals and five copulas. A full bar is a guarantee.

With a symmetric covariate all three symmetric copulas remove 100.0000% of the worst shape — the guarantee. With a covariate skewed at 2.26 they leave 1.4177%, 0.8561% and 0.8171% for the Gaussian, the t and the Frank.

That ordering is not the leak’s. The t, which leaks most on the mean split’s own interaction, is second best on the worst case, and the Gaussian is worst. A rule’s exposure over a family of shapes is not the same comparison as its leak on one of them, because the worst shape differs by copula and the maximum is taken after the leak rather than before.

So the factor of two is real on the quantity the parity field reports and is a third of that on the quantity a trial acts on: 1.4177% against 0.8171%, a ratio of 1.74. Both are worth having and only the second is a recommendation.

Two readings a reader should not take

That the Frank copula is safer. It leaks least on the mean split at every skewed marginal, and it is not safer in any sense a trial can use: a trial does not choose its copula, and the three are indistinguishable by the statistics anybody computes. What the comparison establishes is a range of exposure, not a preferred member of it.

And that this is a small effect because the numbers are small. Twelve per cent against twenty-four per cent of an interaction is a doubling of a quantity that is itself the whole subject of three fields. The interaction a balancing rule fails to remove is what a covariate-adaptive trial is protecting against; twice as much of it surviving is exactly the kind of difference the rules were designed for, and the reason it reads as small is that both numbers are shares rather than effects.

What a trial should do instead

The recommendation that comes out of this is not a correction to apply and is a choice of rule.

The three copulas differ by a factor of two on the mean split’s leak. On the median split they differ by nothing at all, because that zero is arithmetic and holds at every combination. On a threshold at a value they differ, but the whole quantity is smaller and it never had a guarantee to lose.

So a trial that cannot identify its copula and cannot rule out a skewed covariate has one rule whose exposure it knows: the one whose guarantee does not depend on either half of the dependence. The argument for balancing on a median split rather than a mean has never been that it removes more — the two remove the same thing when everything is symmetric — and it is that its guarantee is the only one in the collection that survives being wrong about the joint law.

That is the same conclusion the parity field reaches, arrived at from the half of the dependence it could not see, which is worth something on its own: two independent variations, one recommendation.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A margin that turns over — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness, symmetry
  • An answer that changes — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness, symmetry
  • The zero that was a crossing — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness, symmetry
  • A split survives what a mean does not — both name covariate adjustment, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness
  • A zero that rests on a symmetry — both name covariate adjustment, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness
  • The cut that is not a quantile — both name continuous covariate, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness

Named objects

A flat tag is an object no other essay names yet.

Closed formContinuous covariateCorrelationCovariate adjustmentCovariate balanceExperimental designGaussian copulaIdentificationInteractionMarginal distributionMedian splitMonotone transformationSkewnessSymmetryTail probability