Balancing a skewed covariate
Worth reading first: Balancing what is known in advance · A design is a number.
The interaction zeros need symmetry and one of them survives anything. The remaining question is the one a trial designer actually asks: given a covariate that is not symmetric, which functions should the rule hold?
That is a question about the whole table rather than about the zeros, and the table is computable.
The table on a normal covariate
Rows are what the rule holds and columns are what the outcome depends on. Each entry is the share of that shape’s variance a rule holding those functions removes, computed exactly.
The means remove everything of a linear shape, 63.66% of a median-split shape, 25.00% of one covariate times the other’s square, and nothing of the square, the product, or the product of the splits.
The median splits remove 65.65% of a linear shape, everything of a split shape, 4.10% of the mixed term, and nothing of the same three.
Both together remove everything of the first two, 38.26% of the mixed term, and still nothing of the same three — which is the parity result: adding a second odd function to a first leaves the worst case where it was.
The squares are the row that breaks the pattern. They remove everything of the square, 64.00% of the product, and 6.84% of the product of the splits — so one even function moves the worst case off zero.
One entry deserves a note because it is the only place in the table where two apparently unrelated constructions meet. The means remove 63.66% of a median-split shape, and the median splits remove 65.65% of a linear one — and 63.66% is 2/π to four figures, which is the squared correlation between a standard normal and its own sign. That is the number the cut-point field derives as the whole content of a median split’s relationship to its covariate, and it appears here as a table entry with nothing else in it.
The 65.65% is the same quantity plus a small contribution from the other covariate’s split, which is reachable at a correlation of a half. The two entries are not symmetric because the dictionaries are not: a rule holding both covariates’ splits has two functions pointed at one linear shape, and a rule holding both covariates’ means has two pointed at one split shape, and the geometry of those two situations differs by how much of the second covariate the first can see.
And on a skewed one
Three things happen at once.
The zeros fill in. The means now remove 52.01% of the square and 26.30% of the product; the whole right-hand block of exact zeros becomes a block of small-to-moderate numbers.
The split’s power falls. The median splits removed 65.65% of a linear shape on the normal covariate and 48.93% here, because a skewed covariate’s linear part lives in its long tail and a split at the median cannot see where in the tail a unit is.
One zero stays. The median splits still remove exactly nothing of the product of the median splits, and so does every other dictionary that is entirely odd on the latent scale.
63.66% is 2/π
One entry of the table has a closed form, and having it makes the row above it legible.
The share of a median-split shape that a rule holding the mean removes is the squared correlation between a standard normal and its own sign. That correlation is , so the share is
63.66%, exactly, and it does not depend on the correlation between the covariates or on anything else in the design. It is a fact about a normal variable and its own sign.
Which makes the neighbouring number readable. The reverse entry — what a rule holding the median splits removes of a linear shape — is 65.65%, two points higher rather than equal. On independent covariates the two would both be 2/π, by the same argument run backwards. The extra two points are the second covariate: at a correlation of a half, the split of explains a little of that the split of does not.
So the asymmetry between the two off-diagonal entries is the correlation, and its size is two points at ρ = 0.5. Neither dictionary is intrinsically better at the other’s shape; one of them is being helped by a covariate it was not aimed at.
The two odd dictionaries are super-additive
The mixed term is the one column where the four rows genuinely differ, and the numbers there do something a reader might not expect.
The means remove 25.00% of it. The median splits remove 4.10%. Both together remove 38.26% — which is 9.16 points more than the sum of the two taken alone.
That is legitimate and it has a name. A squared multiple correlation on a union of predictors is at least the largest of the parts, and it can exceed their sum whenever the predictors are correlated, because one of them can absorb variation that was obscuring the other’s contribution. On independent predictors the shares would add exactly.
So on this one shape the mean and the median split are not two partial answers being combined; they are two functions each of which makes the other more useful. And it is worth noticing where this does not happen: on the three shapes with an exact zero, adding the second odd function moves nothing at all, because zero plus zero is zero however the predictors are correlated.
The worst case, which is what a trial can act on
An outcome model is not known before a trial, so what a dictionary is worth is what it removes of the shape it removes least of.
On the two symmetric marginals every rule made of odd functions has a worst case of exactly zero. On the four asymmetric ones the rule holding a mean and a median split of each covariate has worst cases of 0.74%, 1.83%, 2.24% and 2.49%, and the rule holding the means alone 0.24%, 0.80%, 1.32% and 1.07%.
Two things are true of those numbers and the second is more important than the first.
They are small. Two and a half per cent of a variance is not a meaningful amount of protection or a meaningful amount of exposure, and a trial designer told their blind spot has moved from nothing to two and a half per cent should not change anything they do.
They are not statable in advance. The exact zero is a sentence: a rule made of odd functions is worth nothing against an interaction between two of them. What replaces it is a calculation that needs the covariate’s marginal and the correlation between the covariates, and a protocol that wants to say what its balancing rule protects against now has to compute rather than cite.
Reading the table as a trade
The two rows a trial chooses between are the means and the median splits, and putting their whole rows side by side says what the choice is.
On the normal covariate the means read 100.00, 0.00, 63.66, 0.00, 0.00, 25.00 across the six shapes and the splits read 65.65, 0.00, 100.00, 0.00, 0.00, 4.10. Each is perfect on the shape it is, good on the other’s shape, and worth nothing on the four the parity argument covers. They are near-mirror images and the choice is a statement about which of a linear and a split outcome is likelier.
On the strongly skewed covariate the means read 100.00, 44.27, 32.85, 29.73, 1.32, 6.33 and the splits read 33.68, 1.84, 100.00, 0.77, 0.00, 0.07. The means have gained across the board and lost their zeros; the splits have lost a great deal of power on everything except their own shape and kept their zero.
One rule trades guarantees for power and the other keeps neither more nor less than it had. Which is the trade the whole field is about, seen in one pair of rows.
Which rule to hold
The table answers the practical question and the answer is uncomfortable.
Adding an even function is the only thing that moves the worst case off zero on a symmetric covariate, and it is what the square does: from 0.00% to 6.84%. That is a tenfold improvement over the odd rules on the binding shape, and it costs power elsewhere — the squares row removes 63.66% of a split shape where the splits row removes 100%.
On a skewed covariate the comparison inverts partly. The squares’ worst case falls to 4.60%, 1.70%, 1.32% and 2.21%, because a skewed covariate’s square is partly odd and stops being a clean complement to the odd functions; meanwhile the odd rules’ worst cases rise. At a skewness of 4.75 the four rows read 1.32%, 0.00%, 2.24% and 1.32%, which is a table with no clear winner in it.
So the advice a normal covariate supports — hold an even function as well — is worth less as the covariate skews, and is worth most exactly where it is least needed. That is not a trap anybody laid; it is what happens when a recommendation derived under a symmetry is applied without it.
What the median-split row keeps
One row of the table is unusual enough to deserve its own reading.
The median-split dictionary has a worst case of exactly zero on all six marginals, attained on the product of the two median splits. It is the only row in the table whose worst case is a constant.
It is also the row with the weakest protection everywhere else: on the strongly skewed covariate it removes 33.68% of a linear shape, 1.84% of a square and 0.07% of the mixed term. So it is a rule whose guarantees are exact and whose guarantees are mostly that it protects against nothing.
That is not a recommendation against it. A rule whose worst case is exactly zero and stated is a rule a protocol can describe, and a trial that also adjusts for the covariate in the analysis is not relying on the balancing rule for its efficiency in the first place. What a balancing rule is for is the shapes the analysis will not model, and the median split’s invariance says precisely which those are.
What the field established, in four numbers
Four essays of geometry come down to four figures and it is worth collecting them.
6.5 × 10⁻³³ and 7.9 × 10⁻³³ — the interaction the means remove on the normal covariate and on a heavy-tailed symmetric one. Two zeros, one of them on a marginal that is not normal, which is what says the guarantee needed symmetry rather than normality.
1.08 × 10⁻³⁰, six times — the interaction the median splits remove, identical under every marginal, because the calculation contains no marginal at all.
22.49% — what a threshold at a value removes on a normal covariate, which is the largest of its six figures and is the number that says the exact zero was about the median rather than about splitting.
2.49% — the largest worst case the rule a trial actually runs has, across the four asymmetric marginals. Which is small, and which is the number that has to be computed rather than cited.
The first three are exact and the fourth is a measurement. That mix is what a field looks like when a symmetry argument is taken apart: the parts that were really about the symmetry go on being exact under whatever preserves it, and the part that was about the law becomes arithmetic.
The correlation, which is the other dial
Everything above is at a correlation of a half between the two covariates, and the leak grows with it: on the mildly skewed marginal the means’ leak runs 2.51%, 8.35%, 11.63%, 14.81% and 20.39% across correlations of 0.2 to 0.8.
So the honest statement of the field’s result is two-dimensional. A rule’s worst case is a function of the covariate’s marginal and of the correlation between the covariates, and it is zero when either the marginal is symmetric or the correlation is zero. Both of those are conditions a trial can check and neither is a condition a trial usually has.
What it would take to do better
The obvious next move is to choose the dictionary for the marginal, and it is worth saying why the field stops short of it.
The machinery exists. The maximin basis picks the k functions whose worst case over a stated list of shapes is largest, by enumeration over subsets, and running it at each marginal is a few seconds of work. What comes out is a recommendation of the form for a covariate of this skewness, hold these functions.
What makes that recommendation weaker than it looks is the list of shapes. The worst case is a worst case over a list somebody wrote down, and the list here is six shapes chosen because they are the ones this collection’s fields have argued about. On a normal covariate the list’s arbitrariness is partly hidden by the parity structure — several entries are exactly zero for a reason that has nothing to do with the list — and off the normal nothing is exactly zero and the answer is a function of the list alone.
So a per-marginal maximin recommendation would be six answers to a question whose premise is a judgement, and presenting six numbers where the input is one opinion is the shape of an over-claim. The field reports the table instead and leaves the choosing to whoever can defend a list.
What is claimed here, and what is not
This essay takes what each balancing dictionary is worth on a covariate that is not symmetric. The claims are that on a normal covariate the rule holding a mean and a median split of each covariate removes everything of a linear shape, everything of a split shape, 38.26% of a mixed term and exactly nothing of the square or of either product; that on a skewed covariate the exact zeros fill in — 52.01% of the square and 26.30% of the product at a skewness of 2.26 — while the median splits’ power against a linear shape falls from 65.65% to 48.93%; that the worst case of the rule a trial actually runs goes from exactly zero to between 0.74% and 2.49% across the four asymmetric marginals; that the median-split dictionary’s worst case is exactly zero on all six; and that the advice to hold an even function is worth 6.84% on a normal covariate and between 1.32% and 4.60% on the skewed ones.
What stays out, and is named as a decision: an optimal dictionary for a given marginal. The maximin machinery for choosing a basis by its worst case exists in this collection and is run at one marginal. Running it per marginal would produce a recommendation of the form for a covariate of skewness two, hold these three functions, and the reason it is not done is that the recommendation would be conditional on the list of shapes worth protecting, which is a judgement rather than a measurement — and a judgement that is already the weakest link at one marginal should not be multiplied by six.
Also out: what any of it does to a trial of a hundred units. These are exact geometric shares, and turning a share into an expected imbalance needs the design. The comparison that does it is run at one marginal, and the translation is not marginal-free.
The boundary against the dictionary field is that it computes this table at one marginal and calls the zeros exact, which they are. This one asks what happens to the table when the condition behind the word exact is removed, and the answer is that everything moves a little and one entry does not move at all.
The checks, and the refusals that make them mean something
Three claims are gated. Every entry of the table is required to lie between zero and one, because a removed share is a share of a variance and an entry outside that range would mean the Gram solve had failed silently — which is exactly the failure the dictionary field found by extending its own machinery, where a share of 151% is what caught a wrong denominator. The worst case of every odd dictionary is required to be exactly zero on every symmetric marginal, and positive on every asymmetric one. And the quadrature is required to agree with four hundred thousand draws on five cases spanning the marginals and the target shapes.
The refusal is the field’s own summary. The exact worst case, quoted for a covariate that is not symmetric, is refused — with the normal covariate’s exactly-zero figure printed beside the skewed ones’ 0.74%, 1.83%, 2.24% and 2.49%, and with the reason: the guarantee did not become a slightly different guarantee, it became a number that depends on a marginal nobody stated.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A margin that turns over — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
- An answer that changes — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
- A zero that is arithmetic — both name basis functions, covariate balance, interaction, marginal distribution, median split, parity, projection
- A zero that was an assumption — both name basis functions, covariate balance, interaction, maximin design, projection, rerandomisation
- The symmetry the marginals could not show — both name basis functions, covariate balance, interaction, marginal distribution, parity
- The zero that survives a cut — both name basis functions, covariate balance, interaction, projection, variance explained
Named objects
A flat tag is an object no other essay names yet.
Basis functionsCovariate balanceGaussian copulaInteractionMarginal distributionMaximin designMedian splitMonotone transformationParityProjectionRerandomisationSkewnessStudy designVariance explained