The zero that was a crossing
Worth reading first: Which series does the moving · Randomisation is not balance.
The field that varied a copula and a marginal together has a headline, and it is a good one: a covariate skewed at 0.75 under a lower-tail copula leaks 0.002% of the interaction a mean split cannot balance, where the copula alone leaks 7.707%, the marginal alone leaks 8.872%, and adding the two gives 16.6%.
Two failures of a dependence, each substantial on its own, annihilating each other.
It is measured at one rank correlation. Across seven, the same cell reads
0.0084%, 0.0182%, 0.0097%, 0.0015%, 0.0864%, 0.4693% and 1.5327%.
Not small — smallest
The distinction is the whole essay, so it is worth stating before anything else.
That the cell is small at a Spearman of 0.4 is the earlier field’s measurement, and it is correct.
That it is smallest at 0.4, on a grid that reaches either side, is a different claim. It says the cell has an interior minimum, and a quantity with an interior minimum is not small because two things cancel in general; it is small because two curves happen to cross, and where they cross is a fact about the sweep rather than about the pair.
The minimum on the coarse grid is at 0.4, exactly where the number was read. By 0.7 the cell reads 1.5327%, which is a thousand times the reading at 0.4.
Where the crossing actually is
A minimum found on a grid of tenths could sit anywhere inside a tenth, so the same cell is read again on a grid four times finer between 0.30 and 0.50.
0.00969%, 0.00320%, 0.00001%, 0.00152%, 0.00641%, 0.03128%, 0.08643% at 0.30, 0.34, 0.38, 0.40, 0.42, 0.46 and 0.50.
The minimum is at 0.38, where the leak is 0.00001% — one part in ten million, which on a quantity computed by quadrature is a zero rather than a small number.
So the cell crosses zero at a Spearman of about 0.38, and the earlier field measured it at 0.40. Two hundredths of a rank correlation away.
What a crossing is made of
The cell is a measured leak and its baseline is the sum of two margins, and both move with the correlation. The crossing is where the first passes the second.
The previous essay establishes the shapes. The marginal margin rises without turning — 0.728% to 18.213% across the sweep for this covariate. The copula margin rises to a peak at 0.6 and falls. The cell itself rises throughout, from 0.0084% to 1.5327%, but far more slowly than either margin.
So the excess — the cell less the baseline — starts positive, falls through zero somewhere near 0.38, and becomes increasingly negative. The excess crosses at 0.38; the cell happens to have its minimum there too, on this grid, because the cell is so much smaller than either margin that its own shape is dominated by where the two failures cancel.
That is why the cell’s reading of 0.002% at 0.4 is such a striking number and such a fragile one. It is a difference of two quantities each about eight per cent, and a difference of that kind can be made arbitrarily small by choosing where to stand.
What the finer grid costs, and why it is worth it
Seven more recalibrations and seven more grids, which is one more sweep on top of the seven the field already runs.
That is not free — each point solves for a copula parameter by search and then builds a grid at that parameter — and it is worth being clear about what the second sweep buys, because the first one already established that the minimum is interior.
It locates the crossing. A minimum reported at 0.4 on a grid of tenths is consistent with the true minimum being anywhere between 0.3 and 0.5, and the difference matters: a crossing at 0.40 would mean the earlier field stood exactly on it, and a crossing at 0.30 would mean the cell’s small reading at 0.40 was already well off the crossing and the smallness needs another explanation.
It is at 0.38. The earlier field stood two hundredths away, close enough that its reading is dominated by the crossing and far enough that the crossing is not at the round number.
And it shows the crossing is a crossing. On the coarse grid the smallest reading is 0.0015%, which is small but is not obviously a zero. On the fine grid it is 0.00001%, a hundred times smaller, with readings on both sides an order of magnitude above it. That is what a function passing through zero looks like and it is not what a function with a shallow minimum at a small positive value looks like.
Whether it is exact
A cancellation that passes through exactly zero at some correlation is a stronger statement than one that gets small, and this one does pass through zero.
The reading at 0.38 is 0.00001% — a hundredth of a thousandth of a per cent. Everything here is computed by quadrature on a grid rather than by simulation, so there is no sampling noise to hide behind: the number is what the integral gives at that grid resolution, and it is a hundred times smaller than the neighbouring readings.
That the crossing exists at all is not surprising. A continuous function that is positive at 0.30 and negative at 0.42 crosses zero in between, and the excess is continuous in the copula’s parameter. What is worth noticing is that the crossing is near enough to a round number that somebody standing on the round number would report it as a cancellation, and that is not something the function had to do.
The zero is tangential, not a sign change
The fine grid says more than where the crossing is. It says what shape the cell has around it, and the shape is the reason the reading at 0.40 is as small as it is.
Take the distance from 0.38 and read the leak against it. Above the crossing: 0.00152% at two hundredths, 0.00641% at four, 0.03128% at eight, 0.08643% at twelve. Doubling the distance from two hundredths to four multiplies the leak by 4.2, and from four to eight by 4.9. Below the crossing, 0.00320% at four hundredths and 0.00969% at eight, a factor of 3.0.
Those are the factors of a quadratic, not of a line — a leak that fell through zero linearly would double when the distance doubled. So the cell touches zero at 0.38 rather than passing through it, and that is what a share computed as a squared quantity does where the thing being squared crosses zero linearly. The excess changes sign; the leak, being a magnitude, cannot, and bounces.
The square roots make it plainest. √0.00152 = 0.039 at two hundredths and √0.00641 = 0.080 at four, so the root is very nearly proportional to the distance with a constant of about 2.0 per unit of rank correlation above the crossing, and about 1.3 below it. The asymmetry is why the coarse grid’s reading at 0.30 is larger than a symmetric bounce would predict.
How precisely the correlation would have to be known
The quadratic makes the fragility arithmetic rather than an impression.
Taking the leak as roughly 4.0 (ρ − 0.38)² per cent on the steeper side, holding it under 0.01% — the order the earlier field’s headline is at — needs the rank correlation within 0.05 of the crossing. Holding it under a tenth of a per cent needs 0.16.
A Spearman correlation estimated from n rows has a standard error of about (1 − ρ²)/√n, which at ρ = 0.38 is 0.86/√n. Getting that under 0.05 takes about 290 rows, and under 0.016 — which is the two hundredths separating the crossing from where it was read — takes about 2,900.
So the two windows sit on opposite sides of the sample sizes this collection’s trials are built at. A few hundred rows is enough to place a covariate pair inside the window where the cancellation is worth a hundredth of a per cent, and nowhere near enough to place it on the crossing itself. That is a more useful statement than “nobody knows their correlation that well”: the cancellation is usable as a region and not as a point, and the region is about a tenth of a rank correlation wide.
What a practitioner would have concluded
The reading being corrected is not a subtle misinterpretation; it is the natural one, and it is worth spelling out what somebody would have done with it.
A trial balancing on a skewed covariate, worried about an asymmetric dependence, reads a table saying that these two particular failures very nearly annihilate. The conclusion available is: this combination is safe. Two things that each break a guarantee, together, do not.
That conclusion is wrong at every correlation except one. At a Spearman of 0.2 the same combination leaks 0.0182%, at 0.6 it leaks 0.4693%, at 0.7 it leaks 1.5327%. The last is not catastrophic — the baseline there is 48.7%, so the cancellation is still nearly complete in relative terms — but it is a hundred times the number the table reported and it is growing steeply.
A practitioner does not know their rank correlation to two hundredths. It is estimated from the same sample that everything else is estimated from, with a standard error that on a few hundred rows is several hundredths. So the question of whether a given sample sits at the crossing is one nobody can answer, and a cancellation whose location is that precise is not a property anybody can rely on.
The safe reading of the cell is the one the sweep gives: the two failures cancel substantially across the whole range, by an amount that varies by two orders of magnitude, and they cancel exactly at one point nobody can locate.
The two other near-misses
The same cell’s neighbours are worth reading, because they say how narrow the crossing is.
At the same correlation of 0.4, the same copula with a covariate skewed at 0.90 leaks 3.4313% and one skewed at 0.95 leaks 6.1619%. Neither is near zero; both are cancelling substantially against baselines of 30.4% and 33.4%, but neither annihilates.
And the same covariate under the copula’s reflection leaks 25.9963% at the same correlation, against a baseline of about 16.6%. Same rank correlation, same Kendall tau, same radial gap; the direction of the asymmetry reversed; and a factor of seventeen thousand between the two cells.
The reflection’s cell is worth one more line, because it is the sharpest contrast the table offers. Its path across the sweep is 4.5302%, 12.9501%, 20.4479%, 25.9963%, 29.7060%, 31.8266%, 32.5540% — rising steadily and never near zero, while its mirror image passes through zero at 0.38. Two copulas with identical rank statistics, one cell nearly annihilating and the other compounding, at every correlation measured.
One cell of thirty passes through zero, and it passes through it two hundredths from where the table was read. That is not evidence of anything having gone wrong — the number reported is the number the cell has at that correlation — and it is the reason the cell cannot be read as a statement about the pair.
What the cell actually is
It is worth writing the cell out once, because “a covariate skewed at 0.75 under a lower-tail copula” is three choices and each of them is doing something.
The rule is a mean split of each of two variables: an indicator for being above the mean. A trial balancing the two main effects leaves the product’s interaction behind, and the leak is the share of that interaction the main effects fail to remove.
The copula is the lower-tail one, which is radially asymmetric: the two variables cling together in the lower tail and separate in the upper. Its own leak with a symmetric covariate is 7.707% at a Spearman of 0.4, and an earlier field established that a radially symmetric copula would leak nothing there.
The marginal is a covariate skewed at 0.75, which breaks the other symmetry the zero rests on. Its own leak under a Gaussian copula is 8.872%.
So the cell is a mean split of a skewed variable against a mean split of a variable joined to it by an asymmetric dependence, and the two asymmetries are of opposite sign in the sense that matters. That opposition is what produces the cancellation, and turning the copula over — to its reflection — makes the two asymmetries agree and produces 25.9963% instead of 6.1619% for the more skewed covariate.
The cancellation is real and its size is a coordinate.
What survives
The earlier field’s argument does not rest on this cell, and it is worth saying what does survive.
That the leaks do not add survives entirely, and it survives at every correlation in the sweep: eleven cells cancel and nine compound at 0.4, nine and eleven at 0.1, thirteen and seven at 0.7. Additivity is refuted everywhere.
That the direction of a copula’s asymmetry matters once the marginal has one survives, and it is the cleanest result in the earlier field: the same pair under a copula and its reflection reads 6.1619% and 33.2638% at 0.4, and the two copulas are matched on every rank statistic this collection measures.
That a median split’s zero is arithmetic survives, at every one of two hundred and ten combinations, under 10⁻¹⁶. That is the one reading in this whole line of fields that a change of correlation cannot touch, because a centred median split takes the values ±½ and its square is a quarter identically — no joint law enters the argument at all.
What does not survive is reading the 0.002% as how much two particular failures cancel. It is how much they cancel two hundredths of a rank correlation from the point at which they cancel exactly, and a reader who moved the correlation by a tenth in either direction would find 0.0097% or 0.0864%.
The same question asked of the other cells
If one cell of twenty has an interior minimum, the obvious next question is how many others do, and the answer is none.
Every other live cell’s smallest reading on the sweep is at 0.1, the weakest correlation measured, and every one of them rises monotonically from there. The heavy-tailed copula with a covariate skewed at 0.90 runs 9.3308% to 34.7984%; the upper-tail copula with an exponential covariate runs 9.3476% to 50.0057%; the one with no tail dependence and a covariate skewed at 0.95 runs 1.3353% to 33.9684%.
Two cells are flat at exactly zero throughout — the two radially symmetric copulas with the symmetric heavy-tailed covariate — which is the parity result rather than a crossing.
So the near-zero cell is the only one in the table whose leak is not monotone in the correlation, and it is the only one whose reported value is a difference small enough for the difference’s shape to dominate its own. Those are the same fact: a cell whose leak is a per cent or more is a cell whose own growth swamps whatever cancellation is happening, and a cell at two thousandths of a per cent is nothing but the cancellation.
The fragility is a consequence of the smallness, not an additional defect, and it applies to any cell anybody reports as a near-perfect anything.
Why a single point was the right thing to do
It would be easy to read this as a criticism of measuring at one correlation, and it is not.
A table of thirty cells at seven correlations is two hundred and ten evaluations, and the earlier field’s own economy — building each copula’s grid once and reading it at every marginal — is what makes even one such table affordable. Sweeping the correlation as well is the obvious next thing and it is exactly what a deferral is for.
What the earlier field did do is name it. Its closing paragraph says every cell is at a Spearman of 0.40, says the leak is known to grow with the correlation as well as with the skew, and says whether the cancellation survives a stronger dependence is a guess. The number was reported with its conditioning attached, and this field is what happens when somebody reads the conditioning.
The general form is worth carrying. A quantity that is a difference of two larger quantities is near zero on a surface rather than at a point, and a measurement at one point on that surface says nothing about whether the point is special. The way to find out is not to measure more precisely; it is to move.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A symmetry that was not enough — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness, symmetry
- A copula that halves a marginal — both name closed form, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness, symmetry
- The symmetry the marginals could not show — both name copula, covariate balance, interaction, marginal distribution, parity, quadrature, rank correlation, symmetry, tail dependence
- Which tail the cut sits in — both name copula, covariate balance, interaction, marginal distribution, median split, quadrature, rank correlation, symmetry, tail dependence
- A split survives what a mean does not — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
- A zero that rests on a symmetry — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
Named objects
A flat tag is an object no other essay names yet.
Closed formCopulaCovariate balanceGaussian copulaInteractionMarginal distributionMedian splitMonotone transformationParityQuadratureRank correlationSelection effectSkewnessSymmetryTail dependence