The other half of the dependence

Which tail the cut sits in

The same copula and its reflection have the same rank correlation, the same Kendall tau and the same marginals. A balancing rule holding a threshold at a dose leaves 5.33% under one and 33.36% under the other.

Worth reading first: Balancing what is known in advance · A design is a number.

A protocol can specify two rules that read identically on the page. Split each covariate at its median, or split it at a value on its own scale — a dose, a temperature, a clinical threshold.

The parity field shows they are different rules: the first has an exact interaction zero under every marginal and the second never had one, because a threshold anywhere but the median is neither odd nor even on the latent scale. What is left for the second is a size, and this essay is about what decides it.

Two copulas, six times the leak

Five copulas at a Spearman rank correlation of 0.4, with a standard normal covariate throughout, and a threshold at 1 on the covariate’s own scale.

The rule removes 16.42% of the interaction between the two thresholds under a Gaussian copula, 22.20% under a t on four degrees of freedom, 13.20% under a Frank, 5.33% under a Clayton whose dependence sits in the lower tail, and 33.36% under the same Clayton turned over.

Which tail the threshold is in. What a rule balancing a threshold at 1 on each covariate's own scale removes of the interaction between the two thresholds, on five copulas matched at a Spearman rank correlation of 0.4 with a normal covariate throughout. This rule never had a zero to lose — the earlier field establishes that under every marginal — so what is left is a size, and the size depends on where the dependence lives. A Clayton copula, whose density piles up in the lower tail, leaves 5.33%; the same copula turned over, so that it piles up in the upper tail where the threshold is, leaves 33.36%. Same rank correlation, same Kendall tau, same marginal, same threshold: 6.26 times the leak, decided by which end of the distribution the dependence and the cut are both in.
Fig. 1 A threshold at a value across the five copulas, with where each one’s dependence lives.

The last two are the same copula. Same parameter, same rank correlation, same Kendall tau to four decimal places, same marginals, same threshold. A factor of 6.3 between them, decided by which end of the distribution the dependence and the cut are both in.

The mean's zero is the copula's symmetry. Five copulas, each at a Spearman rank correlation of 0.4, with a normal covariate throughout — so nothing here is about the marginal, which is the whole of the earlier field. Horizontally: how far the copula's density is from its own reflection through the centre of the unit square, measured rather than read off the family's name. Vertically: what a rule balancing the mean of each covariate removes of their product. The three copulas at zero on the horizontal axis remove exactly nothing, to thirty decimal places. The two that are not symmetric remove 7.71%. A guarantee that held for six marginals turns out to have needed something the marginals could not have told anybody about.
Fig. 2 The property that separates the five copulas, measured rather than read off a family name.

Why the cut’s position is the whole of it

The threshold sits at 1 on a standard normal scale, which is the 84th percentile: it is a cut in the upper tail.

A copula with upper-tail dependence makes units that are high on one covariate much more likely to be high on the other, so the two threshold indicators move together far more than their rank correlation suggests — and a balancing rule holding both of them removes a great deal of their product.

A copula with lower-tail dependence does the opposite. The two indicators are nearly independent up where the cut is, because that copula’s dependence is concentrated somewhere else entirely, and the rule removes little.

Matched on what. Five copulas cannot be compared until they are matched on something, and there are two conventional somethings. Matched at a Spearman rank correlation of 0.4 the threshold rule's leaks are 16.4%, 22.2%, 13.2%, 5.3%, 33.4%; matched instead at the Kendall tau the Gaussian copula has there — 0.273 — they are 16.4%, 21.4%, 13.3%, 5.3%, 33.1%. The largest move is 0.81%, on the t copula, whose two rank statistics disagree most because its dependence is concentrated in the tails where neither statistic looks. Neither matching is more correct. What is not allowed is a table with the matching left out.
Fig. 3 The same five copulas under two conventions for matching them, which is the other way this table could have been wrong.

That is the practical shape of the finding. A rank correlation does not say where the dependence lives, and a threshold rule is entirely a statement about one place. Two studies reporting the same correlation between two covariates and the same threshold in the protocol can have leaks six times apart, and nothing either of them reports would distinguish them.

Matched on what

Five copulas are not comparable until they are matched on something, and there are two conventional somethings. Which one is used moves the table.

Matched at a Spearman rank correlation of 0.4, the threshold leaks are 16.42%, 22.20%, 13.20%, 5.33% and 33.36%. Matched instead at the Kendall tau the Gaussian copula has there — 0.2730 — they are 16.43%, 21.38%, 13.27%, 5.26% and 33.09%.

The largest move is 0.81 points, on the t copula. That is not an accident either: the t copula’s two rank statistics disagree more than any other’s here, because its dependence is concentrated in both tails and neither Spearman nor Kendall looks at tails.

One zero is arithmetic and one is a symmetry. What two balancing rules remove of the interaction they are aimed at, on five joint laws of the ranks matched at a Spearman correlation of 0.4, with a normal covariate throughout. A rule holding a median split of each covariate removes exactly nothing of the product of the splits under every one of them, including the two that are not symmetric under reflection — and the reason is not a symmetry at all: a centred median split takes the values ±½, so its square is a quarter identically, and the interaction is orthogonal to both main effects whatever the joint law is. A rule holding the mean of each removes exactly nothing under the three radially symmetric copulas and 7.71% under the two that are not. Bars at the floor are exact zeros; the axis cannot draw 9e-32.
Fig. 4 The exact zeros, which do not move under either matching — the check that says they are not artefacts of a scale.

Neither matching is more correct. What is not allowed is a table with the matching left out, which is the same discipline the block-window field applies to a comparison at a shared length: a number is a number at a setting, and the setting belongs in the sentence.

The exact zeros are the check on that. They are exactly zero under both matchings, which is what says they are properties of the geometry rather than of the scale it was compared on.

The whole table at a cut in the other tail, for nothing

The threshold sits at the 84th percentile, and the natural next question — what a cut at the 16th percentile would give — needs no second computation.

With a symmetric marginal, reflecting a copula is negating both covariates. The indicator 1{X>1}1\{X > 1\} becomes 1{X<1}1\{X < -1\}, so the leak of a cut at +1 under a copula equals the leak of a cut at −1 under its reflection. The three radially symmetric copulas are their own reflections, so their entries do not move at all.

So the table at a lower-tail cut is this table with the two Clayton entries exchanged:

cut at +1 cut at −1
Gaussian 16.42% 16.42%
t on four 22.20% 22.20%
Frank 13.20% 13.20%
Clayton, lower tail 5.33% 33.36%
Clayton, upper tail 33.36% 5.33%

Which states the finding in its strongest form. Nothing about the copula decides the leak; what decides it is whether the cut and the dependence are in the same tail. The same protocol, run on the same population, with the threshold moved from the 84th percentile to the 16th, moves the leak by a factor of six in whichever direction the dependence is not.

What the two matchings move, all five of them

The t copula’s eight tenths of a point is the largest absolute move between the two matchings, and the other four are worth writing down because one of them is a check rather than a result.

Spearman-matched Kendall-matched move
Gaussian 16.42% 16.43% +0.01
t on four 22.20% 21.38% −0.82
Frank 13.20% 13.27% +0.07
Clayton, lower tail 5.33% 5.26% −0.07
Clayton, upper tail 33.36% 33.09% −0.27

The Gaussian’s hundredth of a point is not a small effect; it is a convergence check. The Kendall target was defined as the Kendall tau the Gaussian copula has at a Spearman of 0.4, so the Gaussian is the fixed point of the reparameterisation and must return the identical number. That it returns 16.43 against 16.42 says the recalibration search is finding its parameter to about four figures, which is what licenses reading the other four rows’ moves as real.

In relative terms the t moves 3.7% of its own value and no other copula moves more than 1.3%. So the sensitivity to the matching convention is not spread across the table; it is one copula’s, and it is the one whose dependence lives where neither rank statistic looks.

How much of the spread is the family and how much is the tail

Two ranges are available in the same five numbers and they are a factor of three apart.

Across the three radially symmetric copulas — Gaussian, t, Frank, which is the choice a modeller usually thinks they are making — the leak runs from 13.20% to 22.20%, a spread of nine points.

Across all five, it runs from 5.33% to 33.36%, a spread of twenty-eight.

Nineteen of those twenty-eight points are one copula against its own mirror image, at identical rank correlation and identical Kendall tau. Choosing the family is worth nine points; choosing which way up the asymmetry sits is worth three times that, and it is the choice no summary statistic in common use records.

The worst case, which is what a trial acts on

An outcome model is not known before a trial, so the reading that matters is what each dictionary removes of the shape it removes least of, over the six shapes an outcome might have.

Under the three radially symmetric copulas, a rule holding a mean and a median split of each covariate has a worst case of exactly 0.00%. Under a Clayton and under its reflection it is 1.74%.

A rule holding a mean and the square of each is better everywhere and moves the other way: 5.34% under the Gaussian, 3.11% under the t, 3.26% under Frank and 4.41% under the two Clayton variants.

A worst case that is better for being asymmetric. The least each dictionary removes over six outcome shapes, which is what a trial can act on because an outcome model is not known before the trial. Under the three radially symmetric copulas the rule holding a mean and a median split of each covariate — the two things every trial balances — has a worst case of exactly 0.00%: a guarantee of no protection at all, put there by parity. Under a Clayton copula it is 1.74%. The asymmetry that broke the zero also removed the shape the zero was protecting, so the worst case improves by losing its exactness — which is the same trade the earlier field found in the marginals and is worth seeing twice, because it is the one direction nobody expects a broken guarantee to move in.
Fig. 5 The least each dictionary removes over six outcome shapes, which is the reading a trial can act on.

Two things in that. The rule every trial actually runs — a mean and a split of each covariate — has a worst case that is exactly a guarantee of nothing under a symmetric copula, and the asymmetry that broke the mean’s zero also removed the shape the zero was protecting, so the rule is better off under a Clayton.

And holding one more function is worth several times either effect. That is the dictionary field’s conclusion reached from outside the Gaussian world it was derived in, which is the strongest form a recommendation in this area gets.

Where a threshold is specified, and why

It is worth saying why a protocol would ever specify a cut at a value rather than at a quantile, since the median split has the better guarantee.

A clinical threshold usually means something. A dose above which a treatment is contraindicated, a temperature above which a process changes, a laboratory value that defines a condition — those are cuts at values because the value is what matters, and moving the cut to whatever the median of this particular sample happens to be would be balancing on a different quantity in every trial.

A median split, by contrast, is a cut at a rank. It has the property that it is the same cut whatever the marginal is — which is why its zero survives every transformation — and it has the disadvantage that it is defined by the sample rather than by the subject.

So the choice is between a rule with a guarantee and a rule with a meaning, and this field’s contribution is to price the second rather than to argue against it. 5.33% to 33.36%, across five dependences at one rank correlation, with the position of the threshold relative to the dependence deciding where in that range a particular trial lands.

The six best bases of 2 functions, and what each protectsEvery cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.linearquadraticcubiccut at 1cut at 2median cut1{x>0} + 1{x>2}0.700.270.360.271.001.000.268x + 1{x>2}1.000.300.230.451.000.660.226x³ + 1{x>1}0.440.221.001.000.320.290.2191{x>1} + 1{x>2}0.460.360.221.001.000.190.189x³ + 1{x>2}0.160.331.000.151.000.220.1541{x>2} + 1{x>−1}0.540.520.200.151.000.200.151worstthe rule readsclosed-form projections, no simulationordered by the worst cell, which is the guarantee
Fig. 6 What a balancing rule is holding, geometrically, which is what the two kinds of cut are two choices of.

The instrument, and where it was wrong first

Everything here is a weighted sum over a grid of each copula’s own density, and the grid is on the latent normal scale rather than on the unit square.

That is not a presentation choice. A uniform grid in ranks has its outermost node at Φ1(1/2N)\Phi^{-1}(1/2N) — 3.3 at 240 points — and the functions integrated here include squares, so the truncated tails carry the answer. The rank grid reports 62.55% where the two constructions this collection already had compute 64.00%, and nothing about 62.55% looks wrong: right order, stable under refinement, and it would have passed a comparison to two significant figures.

On the latent scale the same computation gives 63.99996% and 22.50% against the earlier 64.00% and 22.49%. Three independent routes — a Hermite series, a nested quadrature over the bivariate normal, and a grid over a copula density — agreeing to four figures is what makes every exact zero in this field a statement rather than an artefact.

Two covariates make the dictionary an outer product. Four functions of each covariate, and everything a balancing rule may be handed. The margins are the 8 main effects and the block between them is the 16 interactions, which are 66.7% of the dictionary. Every inner product in it is closed form — ⟨f₁g₁, f₂g₂⟩ = ⟨f₁,f₂⟩⟨g₁,g₂⟩ when the covariates are independent — so nothing about the geometry gets harder. What gets harder is the counting: choosing k of 24 is C(24, k), which is 10,626 at four and 735,471 at eight.
Fig. 7 The construction the grid has to reproduce, computed by a Hermite series.

The rank statistics, checked

The matching rests on two quantities computed from the same grid, and both have closed forms for a Gaussian copula that the grid knows nothing about.

Kendall’s tau is (2/π)arcsinρ(2/\pi)\arcsin\rho and Spearman’s rank correlation is (6/π)arcsin(ρ/2)(6/\pi)\arcsin(\rho/2). At ρ=0.416\rho = 0.416 the grid gives 0.27310 and 0.40017 against 0.27314 and 0.40017.

That check caught a real error. Summing whole cells to get the copula at a midpoint — the obvious reading, and the first one written — biases tau upward by 0.03 at these correlations, which is a tenth of the quantity and looks entirely plausible: it made the Gaussian copula report 0.3014 where the closed form says 0.2731. Reading the mass below a midpoint as every cell strictly below it plus half the row and column and a quarter of its own cell fixes it.

A calibration computed wrongly is five families matched on nothing, so the check is load-bearing rather than decorative.

What a practitioner should take

Three sentences, and the second is the one this field exists for.

A threshold at a value has no exact protection under any dependence, which the parity field establishes across marginals and this one confirms across copulas: the removed share runs from 13.20% to 33.36% over five joint laws with identical rank correlations.

Where the threshold sits relative to the dependence decides the size, by a factor of six, and nothing in a correlation or a marginal reveals it. A protocol that specifies a clinical cut-off in the upper tail should ask whether the covariates are more strongly related up there, and the honest position is that most analyses have no way to answer.

And a median split has none of these problems, because its zero is arithmetic and holds under every copula there is. If the choice between a median split and a clinical threshold is open, the first is the one with a guarantee.

The two zeros beside the sizes

Putting the whole field’s numbers in one place makes the division visible.

The median split’s interaction zero is 102010^{-20} or smaller under all five copulas, because a centred median split’s square is a quarter identically. Nothing can break it.

The mean’s interaction zero is 102910^{-29} or smaller under the three radially symmetric copulas and 7.707% under the two that are not, and it needs a symmetric marginal as well. Two conditions, either of which real data can break.

The threshold rule’s share is between 13.20% and 33.36% under all five and has never been zero anywhere. It is a size rather than a guarantee, and this essay is about which of the five it lands nearest.

One zero is arithmetic and one is a symmetry. What two balancing rules remove of the interaction they are aimed at, on five joint laws of the ranks matched at a Spearman correlation of 0.6, with a normal covariate throughout. A rule holding a median split of each covariate removes exactly nothing of the product of the splits under every one of them, including the two that are not symmetric under reflection — and the reason is not a symmetry at all: a centred median split takes the values ±½, so its square is a quarter identically, and the interaction is orthogonal to both main effects whatever the joint law is. A rule holding the mean of each removes exactly nothing under the three radially symmetric copulas and 9.06% under the two that are not. Bars at the floor are exact zeros; the axis cannot draw 1e-31.
Fig. 8 The same two rules at a higher rank correlation, where the leak is larger and the zeros are still exactly zero.

Three rules, three kinds of statement — an identity, a conditional guarantee, and a number — all of which read as removes the interaction in a protocol.

What the field adds up to

Three essays, and the shortest statement of each.

One zero is arithmetic. A centred median split squares to a quarter identically, so the interaction between two median splits is orthogonal to both main effects for any joint law whatsoever — no symmetry, no marginal, no correlation, nothing to check.

One zero is two symmetries. A mean’s interaction zero needs the covariate to be symmetric, which the parity field measures, and the copula to be symmetric under reflection, which nothing had measured. It is 7.707% under a Clayton with a normal covariate throughout.

And one rule never had a zero and has a size that depends on where its cut is. From 13.20% to 33.36% across five dependences at one rank correlation, with a factor of 6.3 between a copula and its own reflection.

Behind all three is the same observation about how a dependence was being varied. Six marginals are six monotone relabellings of one copula, so a field that varies marginals holds the copula fixed by construction — and a conclusion phrased as what the zeros need means what the zeros need, given this copula. The repair is to run the comparison along the other axis, which is what this field is.

One more thing is worth saying about the six-fold gap, because it is the kind of number that invites over-reading. It is a ratio between two shares of an interaction, and the interaction is one of six shapes an outcome might have. A trial whose outcome does not depend on the product of two thresholded covariates is unaffected by every number in this essay, and a trial that does not know what its outcome depends on should read the worst-case table instead — where the same five copulas differ by two points rather than by a factor of six.

One caveat belongs with the six-fold figure. It is a ratio between two shares of one interaction, and an outcome that does not depend on the product of two thresholded covariates is untouched by any of it.

What is claimed here, and what is not

This essay takes what decides how much a threshold-balancing rule removes. The claims are that at a Spearman rank correlation of 0.4 with a normal covariate and a threshold at 1, the rule removes 16.42%, 22.20%, 13.20%, 5.33% and 33.36% of the interaction under a Gaussian, t, Frank, Clayton and reflected Clayton copula; that the last two are the same copula reflected, with identical rank correlation, Kendall tau and marginals, so the factor of 6.3 between them is decided entirely by which tail the dependence and the cut share; that matching on Kendall’s tau instead moves those figures by up to 0.81 points, on the t copula, and moves the exact zeros not at all; that the worst case of a rule holding a mean and a median split is exactly 0.00% under the three symmetric copulas and 1.74% under the two that are not; and that the grid reproduces this collection’s two earlier constructions at 64.00% and 22.50%.

What stays out, and is named as a decision: a sweep over the threshold’s position. Everything here is at a cut of 1 on a standard normal scale, which is the 84th percentile and is where the parity field puts it. Sweeping the cut from the median outwards would show the leak growing from the exact zero, and it would need a second axis in every figure. The one number that stands in for it is the pair of Claytons, which is a sweep over the dependence’s position at a fixed cut and answers the same question from the other side.

Also out: a copula fitted to real covariates. Whether real dependence is closer to a Gaussian, a t or a Clayton is an empirical question with an empirical answer, and this collection has no data. The five families here are chosen to span the possibilities rather than to represent any of them.

The boundary against the parity field’s version is that it shows a threshold has no zero under any marginal and this one shows what its size depends on.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A margin that turns over — both name copula, covariate balance, interaction, kendall tau, marginal distribution, median split, quadrature, rank correlation, symmetry, tail dependence
  • The other dial — both name copula, covariate balance, interaction, kendall tau, marginal distribution, median split, quadrature, rank correlation, symmetry, tail dependence
  • The zero that was a crossing — both name copula, covariate balance, interaction, marginal distribution, median split, quadrature, rank correlation, symmetry, tail dependence
  • The zero that survives both — both name covariate adjustment, covariate balance, interaction, marginal distribution, median split, numerical methods, symmetry, threshold
  • A split survives what a mean does not — both name covariate adjustment, covariate balance, interaction, marginal distribution, median split, threshold
  • A symmetry that was not enough — both name covariate adjustment, covariate balance, interaction, marginal distribution, median split, symmetry

Named objects

A flat tag is an object no other essay names yet.

CopulaCovariate adjustmentCovariate balanceInteractionKendall tauMarginal distributionMedian splitNumerical methodsQuadratureRank correlationSymmetryTail dependenceThresholdWorst case