A guarantee that needed a symmetry

The cut that is not a quantile

A protocol that says split the covariate at a threshold and one that says split it at the median read the same and are different rules. One has an exact guarantee under every marginal and the other has none under any.

Worth reading first: Balancing what is known in advance · A design is a number.

A median split keeps its exact zero under every monotone transformation because it is a statement about the sign of the latent normal, and a monotone map does not move signs.

A protocol does not usually say median split. It says split at 65, or above and below 200 milligrams, or stratify by whether the patient is over sixty-five — a threshold at a value on the covariate’s own scale, chosen because it means something in the subject.

Those two sentences describe different rules with different guarantees, and the difference is not small.

Where a value lands

Where the two kinds of cut sit. Six covariates, each a monotone transformation of the same latent normal. The vertical line at zero is where every median split sits, on every one of them, because a monotone map preserves order: the median of the covariate is the image of the median of the latent normal. The marks on the curves are where a threshold at 1 on the covariate's scale falls — 1.000, 0.881, 0.875, 0.783, 0.713, 0.337 — and none of them is at zero. That is the whole of the difference. A function of the sign of the latent normal is odd, and a rule made of odd functions removes exactly nothing of an interaction between two of them; a threshold anywhere else is neither odd nor even and removes something.
Fig. 1 Six covariates as transformations of one latent normal. The vertical line is every median split; the marks are a threshold at one on the covariate’s own scale.

A threshold at the value v is a threshold on the latent normal at T⁻¹(v), and that point is not zero unless v happens to be the covariate’s median.

At a threshold of one on the covariate’s own scale, the six marginals put the cut at 1.000, 0.881, 0.875, 0.783, 0.713 and 0.337 on the latent scale. The more skewed the covariate, the more of its mass sits below any fixed value, so the same threshold slides towards the median as the skewness rises.

Every one of those is away from zero, so the indicator is neither odd nor even. It is 59.4% odd on the normal covariate and 79.1% odd on the exponential — mixed in every case, which is what the parity argument says cannot be protected.

What it removes

A cut at a quantile, and a cut at a value. Two rules that read identically in a protocol. One splits each covariate at its median; the other splits it at 1 on the covariate's own scale — a dose, a temperature, a clinical threshold. At a correlation of 0.5 the first removes exactly nothing of the interaction between its own two splits, under every marginal here, because a median split is a function of the sign of the latent normal whatever the marginal is. The second removes what the bars show, and it does so on a normal covariate too: the threshold sits at 1.000 on the latent scale rather than at zero, so it is 59.4% odd and 40.6% even. The exact zero was never about the cut; it was about the cut being at the median.
Fig. 2 What a threshold at one removes of the interaction between two such thresholds, on all six marginals.

A rule balancing a threshold at one on each covariate, against an outcome depending on the product of those thresholds, removes 22.49% on the normal covariate, 20.08% on the heavy-tailed symmetric one, 19.92%, 17.69%, 15.80% on the three skewed ones and 4.88% on the exponential.

A rule balancing the median split of each, against the product of the median splits, removes 1.1 × 10⁻³⁰ on all six.

The exact zero is a fact about the cut being at the median, not about it being a cut. That is a sentence about a protocol rather than about a distribution, and it holds on the normal covariate too — where the threshold rule removes the most of the six.

What the odd share is measuring

The parity numbers deserve reading against the ones the mean’s essay reports, because they measure the same thing about a different function and move the opposite way.

The covariate itself is 100% odd on a normal marginal and 72.2% odd on a strongly skewed one: skewing the covariate is what breaks its parity.

A threshold at one is 59.4% odd on the normal marginal and 79.1% odd on the exponential: skewing the covariate is what restores the threshold’s parity, because it moves the fixed value towards the median.

Two functions of the same covariate, one losing its parity as the marginal skews and the other gaining it, because parity here is a fact about where a function sits relative to the latent origin rather than about the covariate’s shape. Nothing about a marginal is good or bad for the guarantees; what matters is what a particular dictionary entry becomes on the latent scale.

That is why the criterion at the end of this essay is stated per entry rather than per covariate, and it is why a rule holding a mean and a threshold on a skewed covariate can have both of its entries partly odd and partly even in different amounts.

How much of a covariate is odd. Each covariate is a monotone transformation of a standard normal, and what decides every guarantee in this field is how much of that transformation is odd — how much of it changes sign when the latent normal does. A normal covariate is entirely odd, and so is a heavy-tailed symmetric one, which is the case that separates symmetry from normality. A mildly skewed covariate is 95.7% odd, and a strongly skewed one 72.2%. The even remainder is what a balancing rule can suddenly remove of an interaction it used to be able to remove exactly none of.
Fig. 3 How much of each covariate is odd — how much of its monotone transformation of the latent normal changes sign when the normal does. A normal covariate is entirely odd and so is a heavy-tailed symmetric one; a mildly skewed covariate is 95.7% odd and a strongly skewed one 72.2%.

The odd share has a closed form, and it is one line

The two odd shares quoted are not measurements that happen to come out where they do. They are a function of where the cut sits and nothing else.

Write the cut on the latent scale as c, and let Φ(c)\Phi(c) be the share of mass below it. Centre the indicator and split it into even and odd parts. The odd part is [1{z>c}1{z<c}]/2\left[\mathbf{1}\{z>c\} - \mathbf{1}\{z<-c\}\right]/2, whose variance is p/2p/2 with p=1Φ(c)p = 1-\Phi(c), and the whole centred indicator has variance p(1p)p(1-p). So

odd share  =  p/2p(1p)  =  12Φ(c).\text{odd share} \;=\; \frac{p/2}{p(1-p)} \;=\; \frac{1}{2\Phi(c)} .

At c = 1.000 that is 1/(2×0.8413)1/(2 \times 0.8413), which is 0.594, and at c = 0.337 it is 1/(2×0.6319)1/(2 \times 0.6319), which is 0.791. The two numbers the section above reports, from one expression.

Applied to all six latent cuts — 1.000, 0.881, 0.875, 0.783, 0.713 and 0.337 — the odd shares are 0.594, 0.617, 0.618, 0.638, 0.656 and 0.791.

Which fixes both ends of the scale

The expression is worth having because it says what the two extremes are, and neither is obvious.

At the median, c = 0 and Φ(0)=12\Phi(0) = \tfrac12, so the odd share is exactly 1. A median split is purely odd — which is the earlier field’s whole result, recovered here as the value of a formula at a point rather than as a separate argument.

Far out in the tail, Φ(c)1\Phi(c) \to 1 and the odd share tends to 1/2. It never goes below a half.

So the parity content of a threshold is bounded between a half and one, and a fixed value on the covariate’s own scale can never be worse than half odd however extreme it is. That is a more forgiving picture than “neither odd nor even” suggests — the rule is always at least half protected by parity, and what it loses is the other half.

It also explains the ordering the marginals come in. A more skewed covariate puts more of its mass below a fixed value, which pushes Φ(c)\Phi(c) down, which pushes the odd share up. The exponential’s cut at 0.337 on the latent scale is the closest to a median of the six, so its threshold is the most nearly odd of the six — 79.1% against the normal’s 59.4% — and it is the covariate the guarantee was least expected to survive on.

Why the ordering runs backwards

The numbers fall as the skewness rises, which is the opposite of everything else in this field, and the reason is the sliding above.

At a threshold of one on a normal covariate the cut is a full standard deviation from the median, so the indicator is far from odd and the rule removes a great deal of its interaction. At a threshold of one on an exponential covariate the cut is at 0.337 on the latent scale — the exponential’s median is log 2, or 0.693, so a threshold at one is much nearer the middle of the distribution — and the indicator is nearly odd, so the rule removes little.

The quantity that matters is where the cut sits as a quantile, and a fixed value is a different quantile under every marginal. So the ordering in the table is an ordering of how close each marginal puts the value to its own median, and it says nothing about skewness as such.

That is worth being careful about, because a reader who has read the zero that rests on a symmetry will expect skew to make things worse and here it makes them better, for a reason that has nothing to do with the mechanism there.

The threshold that is already in the collection

This is not a new object in the collection, and it is worth saying where it already appears, because the earlier appearance is what makes the geometry available at all.

The cut-point field closes the geometry of a threshold at a correlation exactly, by conditioning on the second variable: every mixed inner product is ρ^j times a one-variable answer and the only genuinely two-dimensional object left is an orthant probability. It reports the leak for a cut at one and a correlation of a half as 22.49%, and that is the number this essay reproduces on the normal covariate by a completely different route.

Two routes to a number are two routes only if they can disagree about something, and a Hermite series and a nested Gauss–Legendre rule share no arithmetic at all. The agreement is what licenses the other five figures in the table, which have no series route available to them.

The earlier field also states the parity conclusion, and states it about the median: the interaction guarantee a correlation destroys for powers survives it exactly for median splits, because a two-valued function squares to a constant. That derivation is right and it is about the median, and the sentence travels as though it were about splits. This essay is the correction of the sentence rather than of the derivation.

The same rule, said two ways

There is a version of this that is entirely a fact about protocols and not about statistics.

Two trials balance the same covariate. One writes stratify by the median of the covariate in this trial. The other writes stratify by a value of 65, having looked at a previous cohort in which 65 was the median. If the new cohort’s median is also 65, the two rules are the same rule and both have the exact zero. If it is 60, the second rule cuts at a quantile above the median, the indicator picks up an even part, and the exact zero is gone.

The guarantee depends on a fact about the sample that the protocol does not state. Nothing about the sentence stratify at 65 reveals whether it will be a median split, and the person writing it usually intends one.

The measurement above prices how much that costs at a cut a full standard deviation out: 22.49%, on a normal covariate, at a correlation of a half. A cut nearer the median costs less, continuously, down to zero exactly at the median.

What happens as the cut moves

The measurements above are at one threshold, and the shape of the answer as the threshold moves is worth stating even though it is not drawn.

At the median the removed share is exactly zero. Move the cut away from the median in either direction and the indicator picks up an even part, which grows; the removed share grows with it, continuously, from zero. Far out in a tail the indicator becomes almost a constant — it is one for a handful of units and zero for the rest — and the removed share does not keep growing indefinitely, because a nearly constant function has almost no variance for anything to be a share of.

So the curve rises from zero at the median and turns over somewhere in the tail, and 22.49% at one standard deviation on a normal covariate is a point on its way up.

The practical reading is that the guarantee degrades from the moment the cut leaves the median, without a neighbourhood in which it approximately holds. A trial cutting at the 55th percentile rather than the 50th has a small even part and a small leak; a trial cutting at the 84th has a large one. There is no cut-off below which the protocol is safe.

What a trial should do instead

Two options and neither is free.

Cut at the sample median. The zero comes back exactly. The cut then means nothing in the subject — the threshold is wherever the units happened to land — and a trial reporting high versus low has defined high by its own sample, which is a different claim and a weaker one.

Cut at the value that means something, and hold something else as well. A rule holding both the threshold and a mean of the covariate has a wider span and removes more of everything, and its worst case is decided by whatever is left. That is the general dictionary question and it is the field’s last essay.

There is no third option in which a meaningful threshold keeps the exact zero, because the exact zero is exactly the statement that the threshold is at the median.

A guarantee that stops being a number. The worst case of each dictionary over six outcome shapes, at a correlation of 0.5, against the skewness of the covariate. Under a symmetric marginal every rule made of odd functions has a worst case of exactly zero, and the rule holding a mean and a median split of each covariate — the two things every trial balances — is one of them. Under skew that zero becomes 0.74%, 1.83%, 2.24%, 2.49%: small numbers, each of which depends on a marginal nobody stated. The guarantee has not improved by becoming positive. It has stopped being a guarantee, because it can no longer be written down without the covariate's distribution in it.
Fig. 4 The worst case of each dictionary over six outcome shapes at a correlation of 0.5, against the skewness of the covariate. Under a symmetric marginal every rule made of odd functions has a worst case of exactly zero; under skew that zero becomes 0.74%, 1.83%, 2.24%, 2.49% — small numbers, each depending on a marginal nobody stated.

What the numbers are and are not

The 22.49% is a share of a variance removed by a projection, computed exactly. It is not a share of an imbalance in any particular trial, and the two are related but not equal — the geometry-against-trials comparison is what connects them, and it needs a design, a sample size and an outcome model.

What the geometric number does say is which shapes a rule is worth nothing against, and worth nothing is the property that transfers to a trial without further assumptions. A rule that removes exactly none of a shape’s variance leaves that shape’s imbalance exactly as a coin would, whatever the sample size. A rule that removes 22% of it leaves 78%, on average, in a sense that needs the design to make precise.

So the exact zeros carry further than the non-zero numbers, which is another reason the invariance of the median split’s zero is the finding worth writing down.

The general form

The three rules this field measures are three points on one statement, and the statement is worth having in place of the three.

A dictionary entry participates in an exact interaction zero exactly when it is odd on the latent scale. A median split always is. A mean is when the marginal is symmetric. A threshold at a value never is, on any marginal, unless the value happens to be the median. A quantile split other than the median never is. A square never is. A cube of a symmetric covariate is.

That single criterion replaces every case in the field, and it has one useful property: it can be checked. The odd share of any candidate dictionary entry is a one-dimensional quadrature — decompose the centred function into its odd and even parts under z → −z and report the ratio of squared norms — and it costs nothing to run on a real covariate’s fitted marginal before a trial starts.

A guarantee that can be checked is a different object from one that is asserted, and it is the form this one takes off the normal. That is not as good as the exact zero and it is much better than an assumption, and it is the practical output of the whole field.

What is claimed here, and what is not

This essay takes the difference between a cut at a quantile and a cut at a value. The claims are that a threshold at a fixed value on a covariate’s scale sits at 1.000, 0.881, 0.875, 0.783, 0.713 and 0.337 on the latent scale across the six marginals; that the resulting indicator is 59.4% to 79.1% odd — mixed in every case, so unprotected under every marginal including the normal; that a rule balancing such thresholds removes between 4.88% and 22.49% of the interaction between them, with the largest figure on the normal covariate; that the ordering runs backwards to the rest of the field because a fixed value is a different quantile under every marginal; and that a rule balancing median splits removes 1.1 × 10⁻³⁰ under all six.

What stays out, and is named as a decision: a sweep over the threshold’s position. The essay measures one value on six marginals, which is enough to establish that the position is the quantity and the marginal is not. A curve of the removed share against the cut’s quantile, which is the object a practitioner would actually consult, is one call away and is not drawn — because it would be a curve about the normal covariate with five other curves nearly on top of it, which is a picture that says what this essay says in words.

Also out: the multi-level case. A covariate stratified into three or four bands is a set of indicators rather than one, and the parity of a set is not the parity of its members. The field that gives a continuous covariate levels prices what the banding costs and does not ask this question of it.

The boundary against the essay that established the invariance is that it shows a median split keeps its zero and this one shows what happens one step away from the median. The two together are the statement that the guarantee is a knife edge in the cut’s position and a flat plane in the marginal.

The checks, and the refusals that make them mean something

Two claims are gated. A split at the median is required to be exactly odd on the latent scale under every marginal, and a threshold at a value required to be mixed under every marginal — including the normal, which is the half that makes this a statement about the cut rather than about the covariate. And the rule balancing thresholds is required to remove something of its own interaction under every marginal, because a zero anywhere would mean the parity measurement and the geometry disagreed.

The refusal is the protocol sentence itself. A threshold at a value read as though it were a median split is refused, with the median split’s exact zero and the threshold’s 22.49% printed side by side on the same normal covariate.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A symmetry that was not enough — both name continuous covariate, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
  • The zero that survives both — both name continuous covariate, covariate balance, gaussian copula, interaction, marginal distribution, median split, orthogonality, skewness, threshold
  • Two failures that cancel — both name continuous covariate, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
  • A copula that halves a marginal — both name continuous covariate, covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, skewness
  • A margin that turns over — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness
  • An answer that changes — both name covariate balance, gaussian copula, interaction, marginal distribution, median split, monotone transformation, parity, skewness

Named objects

A flat tag is an object no other essay names yet.

Continuous covariateCovariate balanceDiscretenessGaussian copulaInteractionMarginal distributionMedian splitMonotone transformationOrthogonalityParitySkewnessStratificationStudy designThreshold