The other dial
Worth reading first: Which series does the moving · Randomisation is not balance.
The essay that swept this table across seven rank correlations ends by listing what it held fixed while it did so, and the first item on the list is the covariates. Six marginals, three of them from a skew family; the skewnesses do not move as the correlation does. That is what makes the sweep a sweep of one dial rather than of both at once, and it is also, in that essay’s own words, why nothing there says what happens when a skew is pushed further at a fixed correlation.
This essay turns that dial instead.
The marginal axis was already a ladder
The three skewed covariates the table is read down are Tukey’s transform at , and , applied to a standard normal on the latent scale. So the marginal axis of the table is a ladder in one parameter already. It has three rungs, and nothing has ever read it as one.
Extended to nine rungs at even steps of , and starting at — where the transform is the identity and the covariate is the symmetric normal every parity zero holds at — the ladder runs
The dial is monotone in the quantity anybody would actually measure on their own covariate, and it is strongly convex in it: the last step adds four units of skewness where the second adds half a unit. Whatever is interesting will be at the top.
Three of those rungs are the table’s own. That is what makes the ladder checkable rather than merely plausible: at , and every copula’s leak computed along this ladder must equal the leak the table already reports for the corresponding marginal, and it does — to within , at every one of the five copulas and on the median-split row as well. The two are separate assemblies of the same least-squares projection over the same built grid, and a ladder that agreed to three digits rather than to fifteen would be a ladder measuring something adjacent.
Every row turns over
The first thing the ladder says is not about cancellation at all. It is that the quantity a trial is actually exposed to — how much of the interaction a rule holding two means fails to remove — is not monotone in how skewed the covariate is.
Under the lower-tail copula the leak runs 7.7074% with a symmetric covariate, down through 0.0015% at , up to 6.2007% at and back to 5.5868% at the top of the ladder. Under the upper-tail copula it runs 7.7074%, 25.9963%, 36.2130% at , and down to 23.9686%. Under the heavy-tailed copula it peaks at 25.5802% and falls to 18.0662%.
This is worth stating on its own, because it contradicts the shape of the reading the earlier fields leave. Along the correlation, a marginal’s own leak rises at every step — 0.728%, 2.727%, 5.588%, 8.872%, 12.208%, 15.364%, 18.213% for the mildly skewed covariate under a Gaussian copula — and it is the copula’s own leak that turns over. Along the skew the roles are exactly reversed, and they have to be: a copula’s own leak is measured with a symmetric covariate, which is the rung, so it cannot move along this axis at all. It is the marginal’s own leak that has the interior maximum here, at 25.6973% at , falling to 21.7027% by the end.
A covariate can be too skewed to leak. Past a skewness of about five, pushing the asymmetry further makes a mean split take more of the interaction rather than less. Which margin turns over and which rises is therefore not a fact about either quantity — it is a fact about which dial is being moved, and both fields had only ever moved one.
Why a covariate can be too skewed to leak
The turn is worth an account, because it is the one movement in this field that has no counterpart in the other dial and so has no earlier explanation to borrow.
What a rule holding the two means removes of the interaction is the share of that lies in the span of and . At the transform is odd, the interaction between two odd functions is even, and an odd function is orthogonal to an even one: the share is exactly nought. Skewing the covariate breaks that orthogonality, and at first it breaks it faster the harder it is pushed — which is the rising half of every row, and is what the essays that balanced a skewed covariate already report.
The falling half is a different effect and it is about where the covariate’s own variance lives. At the transform’s standard deviation is 3.07 against a normal’s 1.00, and almost all of that spread comes from the top of the latent scale: the covariate is very nearly its own largest draw. In that régime and are close to the same direction, and so are and the plane , spans — the product is dominated by whichever factor is extreme, and that factor is in the span. A rule holding the two means therefore takes more of the product back, and the share it fails to remove falls.
The two effects run against each other and the maximum is where they balance. On this ladder that is at a skewness of about five, and the argument gives no reason for it to be at five rather than at three or at nine — which is why it is measured rather than derived, and why the rung it falls on is the reading rather than the explanation.
Two of the four cross
The question the sweep was run for is whether a cell’s answer — compounding or cancelling, measured against what adding the two halves’ own leaks gives — moves when the covariate is the dial rather than the dependence.
It does. Of the four copulas that have an answer at all, two change the sign of their excess along the ladder. The heavy-tailed copula runs +0.5810%, +1.7319%, +2.4240%, +2.1010%, +0.7560%, −1.1975%, −2.9627%, −3.6365%; the upper-tail copula runs +6.4892%, +9.4174%, +9.2004%, +6.9661%, +3.6184%, −0.1410%, −3.4741%, −5.4415%. Both cross between and . The other two never cross: the copula with no tail dependence and the lower-tail one cancel at every rung of the ladder.
Two things about that are not what the sweep was run to find.
They are the same two copulas. The four cells that change sides along the correlation are the heavy-tailed and upper-tail copulas at the two most skewed covariates — and those are the two copulas here. The copula with no tail dependence cancels everywhere in both dials, and the lower-tail copula cancels at every skewed covariate in both.
And they cross the same way. Every cell that crosses along the correlation goes from compounding to cancelling as the dependence strengthens. Every copula that crosses along the ladder goes from compounding to cancelling as the covariate is skewed further. Not one crossing in either dial goes the other way.
So the two dials are not two mechanisms. The excess has a zero contour in the plane of the two, and strengthening the dependence and skewing the covariate are two ways of walking across the same contour in the same direction.
And the count that cancel moves with either
The reading a practitioner carries away from the earlier table is a count: eleven of twenty cells cancel and nine compound. That count rises from nine to thirteen across the correlation sweep, which is what that essay is about.
Along the skew ladder at a fixed correlation of 0.4 it rises too — 2 of 4 at every rung up to a skewness of 3.26, and 4 of 4 from a skewness of 4.75 upward. The count that cancel is a reading at a point of both dials, not at a point of one.
That is the practical form of the whole field now, and it is short. A statement of the shape these two failures of a dependence cancel has three arguments in it and is usually quoted with one: the copula, the covariate, and the strength of the dependence. Fixing any two and moving the third can turn the answer over. The earlier field held the covariate and moved the dependence; this one holds the dependence and moves the covariate; the count moves either way.
The near-zero cell is a minimum in both coordinates
The sharpest thing the second dial says about the first is about one number.
The cell that carries the earlier fields’ headline is the mildly skewed covariate under a lower-tail copula. At a Spearman correlation of 0.4 it leaks 0.0015% where adding the two halves’ own leaks gives 16.5789%, and the essay that swept it in the correlation established that the small number is not a cancellation but a crossing: the cell’s leak is smallest at an interior point of the correlation sweep, at 0.4, and two orders of magnitude larger by 0.7.
Along the skew ladder at that same correlation, the same cell reads 7.7074%, 1.7320%, 0.0015%, 1.2378%, 3.4313%, 5.2410%, 6.1619%, 6.2007%, 5.5868%. Its smallest value is at — an interior rung, and exactly the covariate the headline is quoted at.
So the point is a minimum of the leak in both coordinates, and the two thousandths of a per cent is what a two-dimensional minimum looks like when it is read as a property of a pair. Move either dial in either direction and the number is larger — by a factor of eleven in one step down the ladder and a factor of eight hundred by the end of it.
That is a stronger statement than either sweep could make alone, and it is the reason a second dial was worth turning even after the first one had already qualified the number. One sweep says a small cell is fragile in one direction. Two say it is a point.
Where the two sweeps meet, cell by cell
The crossings are a statement about two curves. A finer version of the same question is available at the twelve cells both sweeps actually reach — the three skewed covariates the earlier table names, under each of the four copulas that have an excess at all.
Each of those cells has a local slope in each dial: the excess one rung up the ladder less the excess one rung down, at a fixed correlation, and the excess at a Spearman of 0.5 less the excess at 0.3, at a fixed covariate. The signs are compared rather than the sizes, because a hundredth of rank correlation and a hundredth of the skew dial are not the same step and no common unit exists.
Nine of the twelve carry the same sign in both. The three that do not are the mildly skewed covariate under the heavy-tailed copula, the mildly skewed one under the upper-tail copula, and the most skewed one under the lower-tail copula — and all three are the same way round, with the skew raising the excess where the correlation lowers it. Two of the three sit at the bottom rung of the ladder and one at the top of it, which is where a slope taken from neighbours carries the most of whatever the sweep’s ends are doing.
So the agreement is not universal, and it is not a coincidence either. The reading the measurement supports is that the two dials move the non-additivity the same way at three cells in four, and that the exceptions do not point in random directions.
What does not move, in either dial
One quantity has now been checked at thirty combinations at one correlation, at two hundred and ten across seven, and at thirty-six along this ladder, and it has never once been anything but zero.
A median split leaks nothing. Under every copula, every marginal, every rank correlation and every rung of the skew dial, the interaction a rule holding a median split of each covariate fails to remove is under — which is the quadrature’s own noise rather than a small number. It is arithmetic: a centred median split takes the values ±½ and its square is a quarter identically, whatever the covariate’s scale is and whatever the joint law of the ranks is.
That is why it is checked again at every new setting rather than inherited. A claim that cannot fail is the cheapest kind to verify and the most expensive kind to have wrong, and the one thing that would break it — a marginal transform applied on the wrong scale, so that the split stopped being a split at the median — is exactly the change this ladder makes to the machinery.
The two symmetric rows are the other control and they behave the same way. A radially symmetric copula with a symmetric covariate leaks exactly nothing under a mean split, and the rung of this ladder is that covariate. Every one of those cells reads zero to machine precision, at every copula the parity argument covers, which is what says the ladder starts where the earlier fields’ zeros are and departs from them for the reason the transform gives rather than for a reason the arithmetic gave.
What a reader should take from two dials rather than one
Three sentences, and none of them is the one the sweep was expected to produce.
The non-additivity of two failures is a surface, not a table of answers. It has at least two arguments — the strength of the dependence and the shape of the covariate — and it changes sign along both. A cell quoted with one of them fixed and unstated is a cell whose answer cannot be reproduced.
But the two arguments are not independent stories. The same copulas cross in both dials, they cross the same way in both, and three of the four shared cells move the same way locally. Whoever reads the correlation sweep and generalises stronger dependence makes these asymmetries cancel into more asymmetry makes them cancel has, in this table, made a generalisation that happens to hold.
And the leak itself is not monotone in either dial. That matters more to a practitioner than the sign of an excess, because the excess is measured against a baseline nobody cares about while the leak is the quantity a trial is exposed to. A covariate skewed at 4.75 leaks more than one skewed at 11.16, under every copula on the table, at this correlation. There is no direction in which “worse covariate” reliably means “worse leak”.
What this opens and does not measure
Three things, each named where it could have been run.
The surface has not been swept, only its two axes. Everything above is a cross through a two-dimensional object: one sweep at a fixed skew, one at a fixed correlation, and twelve cells where they cross. The zero contour of the excess is a curve in that plane and nothing here has traced it. Doing so is a grid rather than a line — each point a fresh calibration and a fresh grid — and the crossings would then be a shape rather than a pair of sign changes.
The other four marginals are off the ladder. The heavy-tailed symmetric covariate and the exponential one are in the table and are not members of the Tukey family, so the ladder has nothing to say about them. The exponential one in particular has a skewness of 2.00, which falls between two rungs, and it behaves unlike its neighbours in the earlier table — whether it sits on the ladder’s own curve or off it is a question the ladder was built to be able to answer and does not.
And the threshold rule has been left where it was. This field reads the mean split, because the median split’s zero is arithmetic and the threshold at a value on the covariate’s own scale never had a guarantee at all. But a cut at a fixed value moves the other way with skew — a value on a skewed scale is an extreme quantile, and an unbalanced cut has less interaction to leak — so that rule has a skew dependence of its own and a ladder of its own to be read along. It is the same nine calls per copula and nobody has made them.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A symmetry that was not enough — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness, symmetry
- Which tail the cut sits in — both name copula, covariate balance, interaction, kendall tau, marginal distribution, median split, quadrature, rank correlation, symmetry, tail dependence
- The symmetry the marginals could not show — both name copula, covariate balance, interaction, marginal distribution, parity, quadrature, rank correlation, symmetry, tail dependence
- A split survives what a mean does not — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
- The zero that survives both — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, skewness, symmetry
- A cut is not a polynomial, and it does not have to be — both name closed form, covariate balance, interaction
Named objects
A flat tag is an object no other essay names yet.
Closed formCopulaCovariate balanceGaussian copulaInteractionKendall tauMarginal distributionMedian splitMonotone transformationParityQuadratureRank correlationSkewnessSymmetryTail dependence