Covariate balance — where it appears
Named by 65 essays across 20 fields — each of them below, with the objects they name alongside it.
A basis is a subspace
A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.
A cut is not a polynomial, and it does not have to be
A threshold's expansion never terminates, which is why a balancing dictionary's geometry was closed for powers and taken to draws for cut points. Conditioning on the second variable closes it for both.
A dictionary that is a product
Two covariates make what a balancing rule may read an outer product — eight main effects and sixteen interactions — and every inner product in it is still closed form. What a rule holding all eight main effects removes of a pure interaction is not small. It is zero.
A dictionary that is neither
A rule handed two median splits removes none of their interaction; a rule handed two covariates removes none of their product. Those were two results with two explanations, and they are one result with one — and finding it corrected the number underneath both.
A margin that turns over
A skewed covariate's leak grows without limit as the dependence strengthens. A copula's own leak does not — it peaks at a rank correlation of 0.6 and falls. The margin of the table turns over before any cell in it does.
A proposal that moves more than two units
The walk's autocorrelation is a fact about its step size and not about its acceptance rate. Exchanging three units from each arm mixes nearly twice as fast as exchanging one, and is refused a third more often.
A test rather than a survey
A thin admissible set falls into an arrangement and its mirror image, and the walk that samples it is uniform on half the reference distribution for ever. That was found by enumerating fourteen units, and enumeration stops at twenty-four.
A zero that is arithmetic
A median split's exact zero was explained by a symmetry of the latent normal. It holds under a Clayton copula, which has no such symmetry, because a centred median split squares to a quarter identically.
A zero that rests on a symmetry
A balancing rule removes exactly none of an interaction between two odd functions, at every correlation. The argument needs the joint sign flip to preserve the law, and no real covariate is symmetric about anything.
Balanced on the wrong function
A rule that reads a covariate's numbers halves the variance of the treatment estimate, if the covariate enters the outcome as a straight line. If it enters as a threshold the rule is worth a fifth of that, and if it enters as a curve every rule here is worse than a coin.
Balancing what is known in advance
Four allocation rules, three definitions of balance, and no rule that holds more than one of them. Minimisation keeps the worst factor margin near three patients whether the trial has forty or six hundred and forty — and lets the imbalance in the cross-classified cells climb to 86% of a coin's, because the cells are not what it is watching.
The fourth moment that was missing
Mehler's formula makes the main effects exact at any correlation and stops there, because the interactions need an expectation of four Hermite functions rather than two. A linearisation turns the four into two, and the whole geometry becomes closed again.
The part the rule already took
A diagnostic that reports on what a balancing rule was not handed is run through a column that is 92% inside the span the rule balanced — because orthogonality in the population is not orthogonality on fourteen units.
The statistic the p-value is about
The test for whether a balanced-assignment walk reaches its whole set is run on a covariate function chosen before the trial. Run on the difference in arm means it is the same test, and it is about the number the trial publishes.
Three arms and three scores
Minimisation balances a trial by keeping the arms' counts even inside every prognostic factor. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms.
Two failures that cancel
A mildly skewed covariate under a lower-tail copula leaks 0.002% of an interaction where each failure alone leaks eight and seven per cent. Turn the copula over and the same pair compounds.
What the rule blocks
A balancing rule breaks the admissible set into pieces by refusing exchanges. Which exchanges it refuses is computable from the design and the tolerance alone, before any assignment exists — and it makes a probe.
A score that balances
Weighting each unit by one over its own assignment probability drives the standardised difference between the arms from 0.8310 to 2.8×10⁻¹⁷ — exactly, not nearly. A score fitted without the second covariate leaves that covariate at 0.7057, further apart than doing nothing at all.
The reversal a coin cannot prevent
Randomisation removes Simpson's reversal in expectation, which is not the same as removing it. A correctly randomised trial of eighty units, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials — and stratifying the randomisation takes that to zero at every size.
A model and a count
The share of a unit's exchanges a tolerance box refuses can be modelled from the design or counted over the admissible set. They order the units the same way at a correlation of 0.81 and disagree about the level by 0.027.
A probe chosen from the design
The design's own leverage aligns with the separating direction four times better than a random direction in the same subspace. The concentrated direction the argument invites is worse than random.
A probe nobody chose
On a set that is definitively in two pieces, seven of twenty-four outcomes report nothing at all. Every covariate probe reports it. What separates them is not accuracy — it is that one of them can be chosen and the other is what happened.
A split survives what a mean does not
The two things every trial balances come apart on a skewed covariate. A median split is a function of the sign of the latent normal whatever the marginal is; a mean is not, and its exact zero is gone at a skewness of one.
A symmetry that was not enough
A heavy-tailed symmetric covariate has a skewness of zero and leaks exactly nothing under three copulas. Under the two asymmetric ones it doubles the leak, from 7.707% to 14.229%.
A threshold in the tail
How much of a threshold's imbalance a balanced covariate removes is a correlation, and the correlation is a closed form. At the median it is exactly 2/π — the same 2/π a median split throws away — and two standard deviations out it is an eighth.
A zero that was an assumption
A rule handed every main effect of both covariates removes exactly none of a pure interaction. That is true at machine precision, it is a fact about independence, and it dies as the square of the correlation.
Balancing towards unequal targets
A three-arm trial allocating two to one to one is the ordinary case, and a balancing rule built from raw counts does not know it. It balances the arms towards equality inside every factor level, delivers a third to each arm, and reports that it minimised imbalance.
Randomisation is not balance
A third of all ways to split sixteen units leave the two halves more than half a standard deviation apart on a covariate. What randomisation delivers is not balance but a known reference distribution — and it makes a test exact with no assumption about the data's shape at all.
Stationary is not convergent
A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.
The arcsine that closes it, and the error that was overstated
Two median splits of a correlated pair agree with probability ½ + arcsin(ρ)/π, exactly. And the truncation the field was avoiding falls geometrically in the correlation, not algebraically in the order.
The rule that can be guessed
A balancing rule improves as it becomes more deterministic, and a deterministic rule can be worked out in advance from information the person enrolling the patient already has. At full determinism 87.6% of assignments are guessable, and an investigator who acts on the guess produces a treatment effect of three quarters of a standard deviation where the truth is zero.
The statistic that changes sign
A test for an unreachable half needs a quantity that tells one half from the other. Every symmetric reading of a mirror pair is identical, and a magnitude is the natural thing to reach for.
The symmetry the marginals could not show
A mean's interaction zero needs the covariate to be symmetric and the copula to be symmetric under reflection. Six marginals could only ever test one of those, and the other is broken by the commonest kind of dependence there is.
The zero that was a crossing
A cell that leaks 0.002% where adding its two halves gives 16.6% is a field's headline. On a finer grid it passes through zero at a rank correlation of 0.38 — two hundredths from where it was measured.
What the extra function buys
A rule balancing the mean of each covariate has a worst case of exactly zero. Adding the median split — the other thing every trial balances — leaves it at exactly zero, and one square moves it.
Where the enumeration stops
A maximin over an eight-function dictionary is a walk over seventy subsets. Over twenty-four it is 735,471 at eight functions, and the exchange algorithm that replaces the walk scores 421. What licenses the second curve is four sizes where both exist and agree, which is a weaker warrant than it looks.
Which shapes are worth protecting
Choosing a basis by its worst case is a finite problem with an exact answer. The answer has no tie in it, which a maximin optimum is supposed to have — and the tie comes back, along with twice the guarantee, when the basis is drawn rather than chosen.
A copula that halves a marginal
Three copulas break nothing on their own and put a factor of two between the same skewed covariate's leaks — 12.118% under a Frank against 23.640% under a t, at the same rank correlation.
A count that has to be estimated
At sixteen units the admissible assignments can be counted by walking all 12,870 of them. At four hundred there are about 2^393.70, and the share admitted is 0.31885 against a closed form of 0.31818 that has no trial size in it at all. The exhaustion a small trial runs into is a fact about small trials.
A defect that is about size
The admitted share of a rerandomisation barely moves with the number of units. The number of admissible neighbours grows like the square of it, and that is what decides whether the walk can go everywhere.
A quantity that loses to a heuristic
Leverage is a heuristic about which units a balancing rule has most to say about. The constraint's active set is the thing the rule actually does. As a probe, the heuristic wins by 4.4 paired standard errors.
An answer that changes
Eleven of twenty cells cancel and nine compound, at one rank correlation. Sweep the correlation and four of the twenty change sides — all four from compounding to cancelling, all four at the most skewed covariates.
Half a reference distribution
A walk that reaches half its admissible set reports the two-sided p-value exactly right, to the last digit, for ever. A one-sided one it puts on the wrong side of five per cent about once in thirty.
The cut that is not a quantile
A protocol that says split the covariate at a threshold and one that says split it at the median read the same and are different rules. One has an exact guarantee under every marginal and the other has none under any.
The set a dictionary leaves
A rule constrained on six functions at a loose tolerance leaves a set as thin as one constrained on three at a tight one. Both sampling methods cross over at the same thinness, and the tolerance where that happens moves by a factor of three.
The zero that survives a cut
A rule holding both main effects removes half of a pure interaction between correlated powers and exactly none between correlated median splits. The guarantee that a correlation destroyed was never about interactions.
Three functions of one number
A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.
Walking the admissible set
A rerandomisation test hunts for admissible assignments and throws away the rest. A walk visits them instead — and it is exactly uniform only because it stands still when a proposal fails, which is the step that looks like waste.
What a chosen probe finds
On a chain of eight hundred draws the probe the earlier fields use misses 44% of the sets that are split. Its own residual off the rule's span misses 12%, for one least-squares fit.
Where the gain is, and where the decision is
A bigger proposal is worth a factor of six at a loose tolerance and nothing at a tight one. The tolerances where it helps are the ones where a hunt costs two evaluations a draw, and the crossing barely moves.
Where the guarantee is exactly zero
An experimenter who declines to name the shapes, and asks instead to be protected against anything in a class, is asking for a number that is not small but zero. Bounding the class is unavoidable, and the two ways of doing it choose different bases.
Which tail the cut sits in
The same copula and its reflection have the same rank correlation, the same Kendall tau and the same marginals. A balancing rule holding a threshold at a dose leaves 5.33% under one and 33.36% under the other.
Balancing a skewed covariate
The worst case of the rule every trial runs goes from exactly zero to somewhere between a quarter of a per cent and two and a half. Which is small, and is a number that cannot be stated without the covariate's distribution in it.
Before the trial and after
The same diagnostic run at two moments answers two different questions. Before, a positive verdict changes the design. After, it changes which number gets reported — and only for the numbers the defect can reach.
Counting it exactly does not help
If a modelled active set lost because the model was crude, the exact one would win. It is computed at a cost no trial can pay, and it is worse — so the approximation was never what was costing the probe.
The analysis and the shape
An unadjusted analysis after a rule that read the covariate is too cautious — by a third against a linear outcome, by nothing at all against a quadratic. And an adjustment for the wrong function recovers almost none of the precision the right one would.
The diagnostic at two hundred
Pointed at a trial size no enumeration reaches, the test gives three answers rather than one — and past a certain thinness it stops agreeing with itself, which is the honest reading and the one nothing could give before.
The walk that cannot cross
A thin enough admissible set is not one set. It splits into an assignment and its mirror image, no sequence of admissible single swaps joins them, and the walk that samples it is uniform on half the reference distribution for ever.
The zero that survives both
A median split's interaction leak is under 10⁻¹⁶ at all thirty combinations of copula and marginal. It is the only guarantee in the collection that neither half of the dependence can touch.
What a reference distribution costs to sample
A randomisation test on a trial too large to enumerate has to sample its reference distribution, at 1/p attempts per draw and a p-value resolved to 1/(B + 1). Six constraints cost 9,878 attempts per thousand draws, and a thousand draws resolve p to 9.99·10⁻⁴ and not one digit finer.
When the constraints run out
Every function added to a basis is a constraint the assignment has to satisfy with the same units. At sixteen units and a stated tolerance the admissible assignments run 3,874, then 1,006, then 314, then none — and the count is exact, because the assignment space is finite.
The other dial
The table is swept along the strength of the dependence and never along the shape of the covariate. Swept along the shape at a fixed correlation, the same two copulas cross, the same way — and the near-zero cell turns out to be a minimum in both directions at once.
A set of pairs, not a vector
The active set is a graph on the units, and every probe built from it so far has been its degree. Read as a graph it recovers 0.1326 of the alignment the summary lost — and draws level with leverage rather than passing it.
A weight fitted to balance
Weights fitted so that each arm's weighted covariate means equal the sample's leave a difference of 1.4×10⁻¹⁴ between the arms and give the estimate a third of the variance of weights fitted by likelihood — 0.011883, within a relative 5.8% of the bound no estimator can beat. In the world where the assignment carries a square nobody named, the same exact balance leaves the square further apart than no weighting at all, and where the outcome carries it too the estimate is wrong by 0.6973 with an interval that covers 1.5%.
The moments a balance is told
Weights fitted to balance the covariates' means were wrong by 0.6973 in the world where both the assignment and the outcome carry a square. Told the squares and the product as well, the same construction is off by −0.0045 there and its interval covers 91.0%. The failure moves up a moment rather than away: with a cube in both, the second-moment balance is off by 0.3099 and leaves the cube twice as far apart as no weighting. And where overlap is thin, 37.0% of samples have no such weights at all.
Named alongside it
The objects these essays reach for when they reach for this one.
RerandomisationInteractionRandomisation testClosed formExperimental designImbalanceOrthogonalityProjectionBasis functionsReference distributionAssignment mechanismMarkov chain Monte Carlo