Balancing what has no levels

Balancing more than one number

The criterion generalises to several covariates without a word changing, which makes the question what it is worth rather than whether it can be done. Each one added takes a share of the assignment's freedom, and the imbalance left in every one of them rises.

Worth reading first: The variance removed before the data · Balancing what is known in advance.

A trial records more than one number about each unit. Age, a baseline score, a duration, a laboratory value: four or eight covariates is ordinary, and a balancing rule that handles one of them is not obviously a rule at all.

The criterion generalises without a word changing, which is the useful property of having taken it from optimal design rather than inventing it. The covariate imbalance Sₐₓ becomes a vector, the correction Sₐₓ′Sₓₓ⁻¹Sₐₓ becomes a quadratic form, and the rule still assigns each arrival to whichever arm makes the information about the treatment effect larger. Nothing has to be decided, no weights have to be chosen between covariates, and the arithmetic is the same size as the number of covariates squared.

So the question is not whether it can be done. It is what it is worth.

What each covariate is left with

What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 12.7% of a coin's imbalance in it; balancing eight leaves 23.2% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.
Fig. 1 Two hundred units, three hundred trials per point, independent standard normal covariates. The rule is asked for one, two, four and eight at once, and the imbalance left in each of them is plotted as a share of complete randomisation’s.

Balancing one covariate leaves 11.9% of a coin’s imbalance in it. Balancing two leaves 15.7% in each; four leaves 17.7%; eight leaves 23.1%.

The rule degrades rather than failing. At eight covariates it is still holding every one of them to under a quarter of what a coin leaves, simultaneously rather than in turn, and the cost of the eighth covariate is far smaller than the cost of the second.

Why the cost is sub-linear

The shape of that curve is the interesting part, and it follows from what the rule has to spend.

An assignment of n units is a sequence of n binary choices, and the rule uses them to push a vector of p imbalances towards zero. Each arrival offers one bit — this arm or that — and the imbalance it can correct is whichever component its own covariate values happen to load on. With one covariate, every arrival is useful for the one thing that needs correcting; with eight, an arrival that would help the third component may hurt the fifth, and the rule takes whichever assignment is better by the criterion.

The criterion is what makes that trade non-arbitrary. It weighs the components by how much each contributes to the variance of the treatment effect, which for independent covariates is equally — so the quadratic form is the sum of squared imbalances and the rule is minimising the length of the imbalance vector rather than any one of its components.

Minimising a length of p components with the same number of choices leaves each component larger by about √p, which is what 11.9%, 15.7%, 17.7% and 23.1% approximately are: the ratios against the first are 1.3, 1.5 and 1.9 where √2, √4 and √8 are 1.4, 2.0 and 2.8. Slower than √p, because the choices are being used better than at random, and the same shape.

What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 400 with 200 trials per point, a rule balancing one covariate leaves 9.3% of a coin's imbalance in it; balancing eight leaves 17.7% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.
Fig. 2 The same sweep on four hundred units. Every point falls, the shape does not change, and the rate argument from the previous essay says the whole curve keeps falling with n while a category rule’s would not.

The measurement, and the two things it holds fixed

Two choices in that sweep are worth stating, because both could have produced the shape by themselves and neither did.

The number of units is fixed at two hundred while the number of covariates grows, so what is being measured is the cost of asking more of the same assignment rather than the cost of a larger problem. A rule with eight covariates and sixteen hundred units would hold each of them far better than a rule with one covariate and two hundred.

The covariates are independent, which is the expensive case and is argued for below. Correlated covariates are cheaper because they are fewer things in disguise, and reporting the cost on correlated covariates would understate it.

The reference in every case is complete randomisation on the same covariates, measured on the same draws, which is what makes the ratios comparable across the four points: a coin’s imbalance in each covariate is 2/√n whatever the others are doing, so the denominator is a constant and the numerator is the only thing moving.

Against the alternative, which does not generalise

Stratification is the standard way of balancing several things at once, and it generalises in a way that stops working almost immediately.

Two covariates cut into three categories each is nine strata. Four covariates is eighty-one. Eight is six and a half thousand, at which point a trial of two hundred units has more strata than units and “balanced within stratum” means every stratum contains one unit and the assignment is a coin.

That is the standard argument for minimisation over stratification and it is correct. What this field adds is that the same argument applies to minimisation itself once the covariates are continuous, because minimisation balances margins — one covariate at a time — and a margin is what a category has. The rule here does not have strata, does not have margins, and does not have a count that grows with p at all: it has a p × p matrix.

Four rules, and the floor two of them cannot pass. The standard deviation of the covariate imbalance under each rule at n = 200, over 500 trials, as a share of a coin's — which is exactly 2/√n = 0.1414 and is the one number here that needs no simulation. Blocking inside 2 categories and minimising on the same 2 categories are the same rule to within their noise, 61.7% and 61.2%, and the marked line is why: a 2-category split can see 63.7% of the covariate's variance, so a rule that balanced its categories perfectly would still leave 60.3% of a coin's imbalance. Reading the number instead leaves 13.0%, and the worst imbalance it produced in 500 trials was 0.094 standard deviations against the coin's 0.439.
Fig. 3 The one-covariate comparison, for the scale it sets. Everything in this essay is what happens to the bottom row when the covariate becomes several.

What correlated covariates change

Independence is the hard case, and it is worth saying why, because the instinct runs the other way.

Two covariates correlated at 0.9 are nearly one covariate. Balancing them is nearly one constraint rather than two, the imbalance vector lies nearly along a line, and the rule spends its choices on one direction instead of two. The measurements above use independent covariates because that is where p covariates are genuinely p things and the cost is largest.

The criterion handles correlation without being told: Sₓₓ⁻¹ in the quadratic form is exactly the correction for it, so two nearly-collinear covariates contribute nearly one term and the rule is not double-counting. That is the Mahalanobis distance rather than the Euclidean one, and it arrives from the information matrix rather than from a decision to use it.

Two rates, not a factor. The standard deviation of the covariate imbalance under three rules, at five trial sizes, 260 trials each, on log axes. The upper line is a coin: its slope is -0.489, against a closed form of exactly −½. The middle line is minimisation on a median split; its slope is -0.519 — the same rate — because inside a category the assignment is still a coin, and what it buys is the constant, 0.654 of a coin's at n = 200. The lower line is the rule that reads x and maximises the information about the treatment effect: slope -0.987, nearly twice as steep. Its advantage is therefore not a number that can be quoted — it is 0.258 of a coin's at n = 50 and 0.065 at n = 800, and it keeps going.
Fig. 4 The rate from the previous essay, which is what the sub-linear cost is a cost against. Eight covariates at 23.1% of a coin’s is still a rate of n⁻¹, and a categorising rule at 62% of a coin’s is a rate of 1/√n.

What happens at the strata’s own game

The comparison worth making is not against a stratified design that has enough units — it is against one that does not, because that is the design a trial with several covariates actually has.

Two hundred units and four covariates cut into three categories each gives eighty-one strata and about two and a half units per stratum. Every stratum with one unit in it is assigned by a coin; every stratum with two is a coin followed by its opposite. The design’s balancing is doing almost nothing, and the marginal balance it delivers is the balance a coin delivers, which is what the covariate field’s own measurements of stratification at fine grids show.

The rule here has no such threshold. Its cost per covariate is the √p sharing above, its state is a four-by-four matrix rather than eighty-one counters, and at two hundred units it is holding four covariates to 17.7% of a coin’s each. The stratified design with the same information available to it is at a coin’s, because it has cut its two hundred units into eighty-one pieces before doing anything with them.

That is not an argument that stratification is a bad idea. It is an argument that stratification has a capacity — a number of strata beyond which it stops being a design at all — and that a rule built on a criterion rather than on cells does not.

What the extra covariates cost the analysis

A design that balances eight covariates is followed by an analysis that has to know about eight covariates, and the two costs are different.

The adjusted analysis spends a degree of freedom per covariate, which at sixty units and eight covariates is a real fraction of the residual degrees of freedom and at four hundred is nothing. That cost is unavoidable and is the price of the repair rather than of the design.

The rerandomisation test spends nothing. Its reference distribution is the rule re-run, whatever the rule reads, so a rule balancing eight covariates and a rule balancing one produce reference distributions of the same size and the test costs the same. That is the strongest practical argument for it in this field: the analysis that needs to be told only the rule is the analysis whose cost does not grow with what the rule reads.

Three analyses of the same trials, none of them wrong about the data. 320 trials at n = 60 with no treatment effect at all, so every rejection counted is a false one, and a covariate that drives the outcome with coefficient 1. The unadjusted comparison is at 5.94% after a coin — its level — and at 0.00% after the rule that reads the covariate: the design removed the imbalance and the analysis is still pricing it. Adjusting for the covariate gives 4.06%, and the rule's own reference distribution — hold the outcomes, re-run the rule 199 times, count — gives 3.13% against the 4.5% that 199 draws can deliver. The last of the three has to be told the assignment rule and nothing else, which is the one thing the experimenter certainly knows.
Fig. 5 The three analyses at one covariate, from the previous essay. Two of them get more expensive as covariates are added and one does not.

What eight balanced covariates are worth to the trial

The imbalance ratios are the design’s own scorecard and not the trial’s. Converting them takes one more fact: how much of the outcome’s variance each covariate carries.

If eight covariates jointly explain half the outcome’s variance, the part of the estimate’s error that comes from imbalance is halved before the rule does anything, and holding each covariate to 23% of a coin’s imbalance removes about 95% of what is left of that part. If they explain a tenth, the same design work is worth a tenth as much.

That is why the field’s advice is about which covariates rather than how many. The rule’s cost of carrying an extra covariate is small and falls with the list’s length; the benefit of carrying one is proportional to how much of the outcome it explains, and is zero for a covariate that explains nothing. A list chosen by importance and not truncated by cost is the arrangement both facts point at.

What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 7.8% of a coin's imbalance in it; balancing eight leaves 14.3% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.
Fig. 6 The same sweep with the rule made deterministic. Every point falls and the shape is identical, which says the sub-linear cost is a property of how the choices are shared between covariates rather than of how firmly each choice is taken.

And the thing eight covariates still do not balance

Every covariate here is balanced on its mean, because the criterion contains each of them linearly, because the analysis model does.

What none of the rules balance. Each rule's imbalance in two quantities, both as a share of a coin's: the covariate's mean, which is what the rules are about, and the covariate's spread, which is not. Reading the number takes the mean imbalance to 13.0% of a coin's and leaves the spread at 102.4% — that is, exactly where it found it. Nothing here is a failure of the rules: a rule balances what it is given, and every one of these was given the mean. It is the covadapt finding about cells under balanced margins, on a covariate that has no cells, and it matters for the same reason: an analysis that uses the covariate any way other than linearly is using a quantity the design said nothing about.
Fig. 7 Each rule’s imbalance in the covariate’s spread beside its imbalance in the mean. Adding covariates to the rule adds means to the list of things it holds; it adds nothing at all to this column.

The imbalance in the second moment is what a coin leaves it — 0.969 of a coin’s under the rule that reads the number — and balancing eight covariates instead of one does not change that for any of the eight. A trial with eight balanced covariates has eight balanced means and sixteen unconstrained second moments, and any analysis that uses a square, a threshold or an interaction is using one of them.

This is the multi-arm field’s cells under balanced margins at its most general: a rule holds the functionals it was given, exactly, and holds nothing else, and adding functionals to the list is cheap while guessing which one the analysis will want is the actual difficulty.

The cost is a cube root, not a square root

“Slower than √p” is the right description and the four numbers say how much slower. Against the one-covariate figure of 11.9%, the ratios at two, four and eight covariates are 1.32, 1.49 and 1.94, and the exponent they imply is log(1.94)/log(8) = 0.32.

That is a cube root rather than a square root, and the difference is not cosmetic. At eight covariates √p would predict 33.7% and the counted figure is 23.1%; the choices are being shared between components about a third better than a random allocation of them would manage, because the criterion picks which component each arrival helps rather than letting the arrival’s own covariate values decide.

Extrapolating that exponent — and it is an extrapolation, over a range twice as wide as the one measured — puts twenty covariates at about 31% of a coin’s each and fifty at 42%. So the practical answer to how many covariates the rule can carry is that there is no number: at fifty it is still holding every one of fifty to under half of what a coin leaves, simultaneously, on two hundred units.

The two capacities, and the factor of fifty between them

Stratification’s capacity can be written down. With c categories a covariate and p covariates there are c^p strata, and a design needs at least a couple of units in each before blocking does anything — so p ≤ log(n/2)/log©. At two hundred units and three categories that is 4.2: four covariates is the most a stratified design at this size can carry, and the fifth takes it below one unit a stratum.

Its capacity grows logarithmically. Adding one covariate needs three times the trial; doubling from four covariates to eight needs eighty-one times.

The continuous rule’s capacity grows very differently. Its imbalance goes as p^0.32 and falls as n^(−0.54) relative to a coin, so holding each covariate to a fixed share needs n ∝ p^0.59: doubling the covariate list costs about fifty per cent more units.

Set the two against each other at the same doubling, from four covariates to eight. Stratification needs a factor of 81; the rule needs a factor of 1.5. That is the whole argument of this essay in one ratio, and the ratio is fifty-four.

It also says which of the two constraints an experimenter is actually up against. A stratified design is limited by its own arithmetic long before it is limited by anything about the trial — four covariates at two hundred units — while the continuous rule is limited by nothing in the range anybody records. The question how many covariates can be balanced has a hard answer for one construction and no answer at all for the other, and the two get discussed as though they were the same question with different constants.

What to balance on

The practical question the field ends at is not how to balance but what to balance on, and the measurements answer part of it.

Adding a covariate to the rule is cheap: the eighth costs the first seven about five points of their own imbalance between them, and costs nothing in machinery. So the instinct to be economical — balance on the two or three most important covariates — is buying very little.

Adding one is only worth anything if it drives the outcome. A covariate that is unrelated to y contributes nothing to the variance of τ̂ however imbalanced it is, so balancing on it is spending choices to no purpose — and worse, spending choices that would otherwise have gone to a covariate that matters.

So the list should be short for a reason and long by default. Include what plausibly drives the outcome, exclude what does not, and stop worrying about the length: the arithmetic of eight covariates is the same as the arithmetic of one and the cost per covariate falls as the list grows.

What balancing several numbers at once costs each of themThe criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 100 with 200 trials per point, a rule balancing one covariate leaves 18.3% of a coin's imbalance in it; balancing eight leaves 36.5% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.00.2000.4001248how many covariates the rule is asked to balanceimbalance left in each of them, as a share of a coin's18.3%20.7%25.9%36.5%a coin would be at 1.0, off the top of this frame200 trials per point at n = 1001: 0.183, 2: 0.207, 4: 0.259, 8: 0.365
Fig. 8 Drag the size of the trial. The curve keeps its shape and slides down, so the answer to “how many covariates can this rule hold” is “more, the larger the trial”, which is not how a stratified design behaves.

Where the field ends

Four essays and one sentence: a balancing rule acts on what it is given, and giving it a category is giving it less than the number.

The first essay computes how much less, in closed form, before any data exists. The second replaces the category with the number and finds that the improvement is a rate rather than a constant. The third finds that the improvement is worth nothing — worse than nothing — to an analysis that does not know it happened. And this one finds that the rule generalises to as many numbers as anybody records, cheaply, and still holds only the functionals it was handed.

What the field inherits from the two before it is the whole apparatus: the rules, the predictability argument, the rerandomisation test, the finding that an unadjusted analysis after a balancing rule is conservative. What it adds is what happens when the thing being balanced has no levels, which is the case in every trial that records a measurement rather than a classification — which is most of them.

What is claimed here, and what is not

This essay claims the cost of asking one rule for several covariates at once: that the criterion generalises unchanged, that the imbalance left in each covariate grows sub-linearly in the number of them, and that the correction for correlation between covariates is already in the criterion rather than an addition to it.

What stays out and is named as a decision: covariates of different kinds in one rule — a continuous one and a factor — where the criterion accepts an indicator column and the measurement has not been made here; weighting covariates by importance, which the criterion does automatically through the outcome model and which a practitioner might want to override; and the case where p is large relative to n, where Sₓₓ becomes ill-conditioned and the rule needs a regularised version of the correction it is currently computing exactly.

The checks, and what they are checked against

Two claims are gated in this field’s library. The imbalance left in each covariate is required to rise with every covariate added, monotonically, and to stay under half a coin’s at eight — the pair, because the first alone would be consistent with the rule collapsing and the second alone with it being free. And the running criterion is required to agree with the matrix definition at every arrival, with one covariate and with three, because the multi-covariate path is a different body of code from the one-covariate path and a rule implemented twice is two rules until they are compared.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlockingContinuous covariateCovariate-adaptive randomisationCovariate imbalanceDₐ-optimalityDegrees of freedomExperimental designInformation matrixMahalanobis distanceMonte CarloOptimal designRandomisationStratification