Balanced on the wrong function
Worth reading first: The variance removed before the data · Balancing what is known in advance.
A balancing rule is an optimiser and it has an objective. The rule that reads a continuous covariate maximises S_aa − S_ax′S_xx⁻¹S_ax, which is one over the variance of the treatment estimate in the model
and that is what makes it the right rule: it is minimising exactly the quantity the experiment exists to make small. Minimisation on a median split and blocking inside one are cruder versions of the same idea, and the whole of the previous field is a comparison between them.
Every one of those rules is optimising for a model in which the covariate enters linearly. That is not an assumption about the analysis that could be relaxed afterwards. It is what the criterion is: change the model and S_ax′S_xx⁻¹S_ax stops being the right correction, and the rule goes on maximising it anyway.
The four rules, in one paragraph each
The comparison below uses the same four rules the previous field measures, and they are worth restating because what each of them reads is the whole of what is at stake.
A coin assigns the arms at random with their sizes forced equal. It reads nothing, it balances everything in expectation and nothing in any particular trial, and it is the reference every other row is a fraction of.
Blocks in a median split cut the covariate at its median and randomise in permuted blocks inside each half, so the two halves are balanced by construction and everything within a half is a coin.
Minimisation on that split reads the same two categories and sends each arrival to whichever arm is currently short in its own category, most of the time.
The rule that reads the number does not categorise at all: it computes what each assignment would do to the information about the treatment effect and takes the better one, most of the time. It is the only one of the four whose objective is the variance of the estimate rather than a distance somebody invented, and it is the best of the four in every measurement the previous field makes.
Three shapes, standardised so the comparison is about the shape
Give the outcome a covariate that matters exactly as much in three different ways: the covariate itself, an indicator for exceeding 1, and a quadratic. Each is standardised to unit variance, so an outcome built from any of them has the same amount explained and the same residual noise. The only difference is the shape.
Two of the three numbers in that caption decide everything below. The correlation between x and 1{x > 1} is 0.6623 — a closed form, and the subject of the next essay. The correlation between x and its own standardised square is zero, by the symmetry of the normal distribution, so a rule that balances the mean of a covariate has no purchase whatever on an outcome that depends on the covariate’s spread.
The table
Four rules, three shapes, and the number in each cell is the variance of the unadjusted treatment estimate as a fraction of the variance a coin gives.
| rule | linear in x | a threshold at 1 | quadratic in x |
|---|---|---|---|
| a coin | 1.000 | 1.000 | 1.000 |
| blocks in a median split | 0.685 | 0.896 | 1.008 |
| minimisation on that split | 0.701 | 0.919 | 0.989 |
| the rule that reads the number | 0.476 | 0.755 | 1.010 |
Read down the first column and the previous field’s conclusion is confirmed exactly: the rule that reads the number more than halves the variance, the two rules that read a cut manage about a third, and the ordering between them is what the earlier measurements said it was.
Read across the rows and none of it survives. Against a threshold, the rule that reads the number removes 24.5% of the variance rather than 52.4%. Against a quadratic it removes nothing at all — it is at 1.010, marginally worse than a coin, and it is the worst of the four rules in that column.
Why a rule can be worse than a coin
The last cell is the one worth explaining, because “the rule did nothing” and “the rule did harm” are different claims and only the second is surprising.
A balancing rule constrains the assignment. It sends each arrival to whichever arm makes its own objective larger, so the set of assignments it can produce is not the set a coin can produce, and within that set the arrival’s covariate value and its arm are related. When the outcome depends on x, that relation is exactly what removes the imbalance. When the outcome depends on x², the relation is irrelevant to the imbalance that matters and the constraint is still there.
The mechanism can be measured directly rather than argued. The imbalance each rule leaves in the function the outcome actually uses, again as a fraction of a coin’s:
| rule | linear in x | a threshold at 1 | quadratic in x |
|---|---|---|---|
| blocks in a median split | 0.366 | 0.847 | 1.019 |
| minimisation on that split | 0.403 | 0.848 | 0.975 |
| the rule that reads the number | 0.016 | 0.614 | 1.088 |
The rule that reads the number removes 98.4% of the imbalance in the covariate it reads and leaves 8.8% more of the imbalance in a quadratic than a coin does. Sorting arrivals by the sign of their contribution to a running sum does nothing to balance the count of large deviations in either direction, and the dependence it induces between arm and covariate makes the squared imbalance slightly worse rather than slightly better.
A second reading of the same row makes the size of it concrete. The rule that reads the number removes 98.4% of the imbalance in x and 38.6% of the imbalance in an indicator at 1 — so of the two things a trial might care about, the imbalance left in the one it reports is nearly forty times smaller than the imbalance left in the one it may actually depend on.
Eight per cent is not a disaster and it is not the point. The point is the sign: a rule that is optimal for one shape is not merely suboptimal for another, it is on the wrong side of doing nothing.
Where the threshold column’s number comes from
The middle column is not an arbitrary intermediate and it is worth deriving, because the derivation is what makes the rest of the field predictable rather than merely measured.
A rule that balances x perfectly leaves the part of the outcome’s covariate that is not linear in x. For an indicator at 1 the correlation with x is 0.6623, so the share of the imbalance that survives is 1 − 0.6623² = 0.5614 — and the counted g-imbalance ratio for the rule that reads the number is 0.614, which is that number plus what the rule fails to remove because it balances a sample rather than a population.
The variance of the treatment estimate is not the imbalance, but it moves with it: the estimate’s variance is the noise plus the part of the covariate’s contribution that the imbalance carries into it. So the 0.755 in the table is the 0.614 diluted by the residual noise, and the whole column can be predicted from one correlation and one signal-to-noise ratio.
That is the field’s arithmetic in one paragraph, and the reason to have it is that it says what happens at settings nobody measured: a threshold further out is less correlated with x, so more of its imbalance survives, so the rule is worth less. There is nothing left to discover about the shape of that dependence, only its size — which is what the next essay computes exactly.
The ranking has no fixed order
The consequence for practice is not that one rule is better than another. It is that the question “which of these rules is the right one?” has a different answer in each column, and nothing in a trial says which column it is in.
Against a linear covariate the rule that reads the number beats blocking by 21 points of variance. Against a threshold it beats it by 14. Against a quadratic it is 2 points behind minimisation and 3 points behind a coin. An experimenter choosing a rule is choosing against a distribution of shapes they cannot observe, which is a decision problem with a name — and it is the one the design field has just spent two essays on, arriving in a completely different subject with the same structure.
The maximin reading of the table above is that the rule which reads the number has a worst case of 1.010 across the three shapes and blocking has 1.008. On that criterion the two are indistinguishable and both are a coin. That is a poor answer, and it can be improved by changing what the rule is allowed to read rather than which rule is used.
What the design’s own literature says about this
It is worth being fair about what is and is not new here, because the linearity assumption is not a secret.
Every treatment of covariate-adaptive allocation states the model it optimises for. What is not usually stated is the size of the exposure: that the rule’s advantage falls to nothing at a threshold two standard deviations out, that it reverses sign against a symmetric non-linearity, and that the imbalance measure everybody reports — the one on the covariate itself — is at 0.016 in a trial where the estimate’s variance is barely improved.
Nor is the second-order consequence usually drawn. If the outcome model is non-linear in a covariate and nobody knows it, then the trial is reported as well-balanced, because the balance table shows the covariate; the estimate is more variable than the balance table implies; and the standard error is conservative or not depending on the same unknown shape. Three separate things go wrong together and all three are invisible from inside the trial.
One constant maps the imbalances to the variances
The two tables are described as related — the estimate’s variance is the noise plus what the imbalance carries into it — and the relation is one number, which can be extracted from the tables themselves and then checked against the cells it was not extracted from.
If a share s of a coin’s estimate variance is the imbalance’s doing, a rule leaving a fraction r of that imbalance gives a variance ratio of 1 − s(1 − r). Fitting s to the two split rules’ linear column gives s = 0.50, and that single value then reproduces:
- blocks against a quadratic, r = 1.019, predicted 1.010 against a counted 1.008;
- minimisation against a quadratic, r = 0.975, predicted 0.988 against 0.989;
- minimisation against a threshold, r = 0.848, predicted 0.924 against 0.919.
Four of the six cells belonging to the two category rules land within half a point, on a constant fitted from two of them. So the middle table is not a second measurement: it is the first table divided by a number, and the number is how much of the estimate’s variance the covariate was responsible for.
That is worth having because the imbalance table is the observable one. An experimenter can compute the imbalance in any candidate function of the covariate without an outcome, and if they can also guess what share of the outcome’s variance that function carries, the variance column follows with no further simulation. The unobservable table is the observable one and one guess.
The exception is the rule whose objective is the variance
Three cells do not fit, and they are the three belonging to the rule that reads the number. Its predicted variance ratios under the same constant are 0.508, 0.807 and 1.044 against counted values of 0.476, 0.755 and 1.010 — better than the map says by 3.2, 5.2 and 3.4 points, in every column and always in the same direction.
The direction is the explanation. The two category rules minimise a distance somebody invented, so the imbalance they leave is a complete description of what they did; the rule that reads the number minimises S_aa − S_ax′S_xx⁻¹S_ax, which is the estimate’s variance, so it picks up whatever the imbalance summary does not contain. Chiefly that is the arm sizes and the covariate’s own spread within each arm, both of which enter the variance and neither of which is a mean imbalance.
The quadratic column is where that matters most and where it is easiest to misread. The rule leaves 8.8% more imbalance in a quadratic than a coin does, which under the map would be a 4.4% worse estimate; it is 1.0% worse. So the honest statement about the last cell is narrower than “the rule is worse than a coin”: its objective is still doing something useful even when the imbalance it induces is actively unhelpful, and what it is doing recovers three quarters of the harm.
That does not rescue the column — a rule that halves the variance in one shape and is a coin in another is still a rule chosen against an unknown — and it does sharpen what the failure is. The rule is not broken by a quadratic outcome. It is idle, and the small negative in the imbalance table overstates the cost.
What an experimenter can actually see
Three of the quantities above are observable inside a trial and three are not, and the split is unkind.
Observable: the imbalance in the covariate, which every trial reports; the imbalance in any function of the covariate somebody thinks to compute; and the standard error of the treatment estimate, which the analysis produces.
Not observable: the variance of the treatment estimate across the trials that might have been run, which is what the whole table is about; the shape the covariate enters the outcome by, which is a fact about the world; and therefore which column of the table the trial is in.
The one bridge between the two lists is that the second column of imbalances — the g-ratios — can be computed for any candidate shape without knowing which is right. An experimenter who suspects a threshold at 1 can compute the imbalance in that indicator and see 0.614 rather than 0.016, and knows immediately that the balance table is flattering. That costs nothing and is not done, because the balance table is a report on the covariate rather than on the outcome model.
One more asymmetry is worth naming. Of the three shapes, only one of them — the linear one — is the shape somebody would have to be unlucky to be wrong about. A threshold is what a clinical cutoff, an eligibility rule or a legal limit produces, and a quadratic is what any covariate with an optimum in the middle of its range produces. Neither is exotic, and neither leaves a trace in the data the design is allowed to look at.
What is claimed here, and what is not
This essay takes the shape a covariate enters an outcome by, and the claim is the table: four rules, three shapes, the variance of the estimate relative to a coin’s, and the ordering reversing between columns.
What stays out and is named as a decision: outcomes that depend on several covariates in different shapes, which is the same argument with more bookkeeping; a treatment effect that varies with the covariate, which is measured to inflate every rule’s variance by the same amount because it enters through the sample mean of the covariate and no assignment rule controls that; binary and survival outcomes, where the criterion is not a least-squares variance at all; and any claim about which shape is common in practice, which is not a statistical question and is not answered here.
The boundary against the previous field is the criterion. What each rule is, why the one that reads the number is the subset criterion applied one unit at a time, and what each of them does to the covariate it reads are established there and used here unchanged. What changes is the outcome, and the outcome was never part of any of those rules.
The checks, and the refusals that make them mean something
Two claims are gated in this field’s library. The rule that reads the number is required to be the best of the four against a linear covariate and worth nothing at all against a quadratic — both in one check, because either one alone is half the finding. And the threshold is required to sit between them, which is the correlation doing the work and is what the next essay measures exactly.
The refusal beside them is a rule reported on the covariate it reads. Handed the imbalance in x, the rule looks extraordinary: it removes 99.4% of a coin’s, at a determinism of one. Handed the variance of the treatment estimate, with an outcome that depends on a threshold at 1.5, the same rule removes 22.9%. The check throws when the two disagree by that much, because the first number is a property of the criterion the rule maximises and says nothing about the experiment — and it is the number every balance table in the subject reports.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Which shapes are worth protecting — both name allocation rule, covariate balance, design criterion, efficiency, model misspecification, randomisation, threshold, variance reduction
- A proposal that moves more than two units — both name covariate balance, imbalance, minimisation, randomisation
- Stationary is not convergent — both name covariate balance, imbalance, randomisation, variance reduction
- The arm whose variance is its answer — both name allocation rule, efficiency, treatment effect, variance reduction
- Two contrasts, one split — both name allocation rule, efficiency, treatment effect, variance reduction
- Where the gain is, and where the decision is — both name covariate balance, imbalance, minimisation, randomisation
Named objects
A flat tag is an object no other essay names yet.
Allocation ruleCovariate balanceDesign criterionEfficiencyImbalanceLinearityMinimisationModel misspecificationNuisance parameterRandomisationStratificationThresholdTreatment effectVariance reduction