The shape the covariate enters by

Balanced on the wrong function

A rule that reads a covariate's numbers halves the variance of the treatment estimate, if the covariate enters the outcome as a straight line. If it enters as a threshold the rule is worth a fifth of that, and if it enters as a curve every rule here is worse than a coin.

Worth reading first: The variance removed before the data · Balancing what is known in advance.

A balancing rule is an optimiser and it has an objective. The rule that reads a continuous covariate maximises S_aa − S_ax′S_xx⁻¹S_ax, which is one over the variance of the treatment estimate in the model

y=α+τa+βx+ε,y = \alpha + \tau\,a + \beta x + \varepsilon,

and that is what makes it the right rule: it is minimising exactly the quantity the experiment exists to make small. Minimisation on a median split and blocking inside one are cruder versions of the same idea, and the whole of the previous field is a comparison between them.

Every one of those rules is optimising for a model in which the covariate enters linearly. That is not an assumption about the analysis that could be relaxed afterwards. It is what the criterion is: change the model and S_ax′S_xx⁻¹S_ax stops being the right correction, and the rule goes on maximising it anyway.

The four rules, in one paragraph each

The comparison below uses the same four rules the previous field measures, and they are worth restating because what each of them reads is the whole of what is at stake.

A coin assigns the arms at random with their sizes forced equal. It reads nothing, it balances everything in expectation and nothing in any particular trial, and it is the reference every other row is a fraction of.

Blocks in a median split cut the covariate at its median and randomise in permuted blocks inside each half, so the two halves are balanced by construction and everything within a half is a coin.

Minimisation on that split reads the same two categories and sends each arrival to whichever arm is currently short in its own category, most of the time.

The rule that reads the number does not categorise at all: it computes what each assignment would do to the information about the treatment effect and takes the better one, most of the time. It is the only one of the four whose objective is the variance of the estimate rather than a distance somebody invented, and it is the best of the four in every measurement the previous field makes.

Three shapes, standardised so the comparison is about the shape

Give the outcome a covariate that matters exactly as much in three different ways: the covariate itself, an indicator for exceeding 1, and a quadratic. Each is standardised to unit variance, so an outcome built from any of them has the same amount explained and the same residual noise. The only difference is the shape.

Three ways one covariate can matter. Each curve is a standardised function of the same covariate: the covariate itself, an indicator for exceeding 1, and a quadratic. All three are scaled to unit variance, so an outcome built from any of them is explained to the same degree — what differs is only the shape. A rule that balances the covariate's mean is balancing the first curve exactly and the others only through their correlation with it: 0.6623 for the threshold and 0.000 for the quadratic, which is zero by symmetry.
Fig. 1 Three standardised functions of one covariate. A rule that balances the covariate’s mean is balancing the first exactly, the second through a correlation of 0.66, and the third through a correlation of zero.

Two of the three numbers in that caption decide everything below. The correlation between x and 1{x > 1} is 0.6623 — a closed form, and the subject of the next essay. The correlation between x and its own standardised square is zero, by the symmetry of the normal distribution, so a rule that balances the mean of a covariate has no purchase whatever on an outcome that depends on the covariate’s spread.

The table

Four rules, three shapes, and the number in each cell is the variance of the unadjusted treatment estimate as a fraction of the variance a coin gives.

rule linear in x a threshold at 1 quadratic in x
a coin 1.000 1.000 1.000
blocks in a median split 0.685 0.896 1.008
minimisation on that split 0.701 0.919 0.989
the rule that reads the number 0.476 0.755 1.010
Four rules, three shapes, and no ordering that survives. The variance of the unadjusted treatment estimate under each rule, as a fraction of the variance a coin gives, over 450 trials of 200 units each. Against a covariate that enters linearly the rule that reads the number nearly halves it. Against a threshold at 1 it removes about a fifth. Against a quadratic every rule here is at or worse than a coin — they are all optimising a criterion that is one over the variance of an estimate in a model this outcome does not obey, and a constraint that helps nothing still costs something. Nothing in a trial says which column it is in.
Fig. 2 The same table drawn. The best rule in the first column is the worst in the last, and no experiment knows which column it is in.

Read down the first column and the previous field’s conclusion is confirmed exactly: the rule that reads the number more than halves the variance, the two rules that read a cut manage about a third, and the ordering between them is what the earlier measurements said it was.

Read across the rows and none of it survives. Against a threshold, the rule that reads the number removes 24.5% of the variance rather than 52.4%. Against a quadratic it removes nothing at all — it is at 1.010, marginally worse than a coin, and it is the worst of the four rules in that column.

Why a rule can be worse than a coin

The last cell is the one worth explaining, because “the rule did nothing” and “the rule did harm” are different claims and only the second is surprising.

A balancing rule constrains the assignment. It sends each arrival to whichever arm makes its own objective larger, so the set of assignments it can produce is not the set a coin can produce, and within that set the arrival’s covariate value and its arm are related. When the outcome depends on x, that relation is exactly what removes the imbalance. When the outcome depends on x², the relation is irrelevant to the imbalance that matters and the constraint is still there.

The mechanism can be measured directly rather than argued. The imbalance each rule leaves in the function the outcome actually uses, again as a fraction of a coin’s:

rule linear in x a threshold at 1 quadratic in x
blocks in a median split 0.366 0.847 1.019
minimisation on that split 0.403 0.848 0.975
the rule that reads the number 0.016 0.614 1.088

The rule that reads the number removes 98.4% of the imbalance in the covariate it reads and leaves 8.8% more of the imbalance in a quadratic than a coin does. Sorting arrivals by the sign of their contribution to a running sum does nothing to balance the count of large deviations in either direction, and the dependence it induces between arm and covariate makes the squared imbalance slightly worse rather than slightly better.

A second reading of the same row makes the size of it concrete. The rule that reads the number removes 98.4% of the imbalance in x and 38.6% of the imbalance in an indicator at 1 — so of the two things a trial might care about, the imbalance left in the one it reports is nearly forty times smaller than the imbalance left in the one it may actually depend on.

Eight per cent is not a disaster and it is not the point. The point is the sign: a rule that is optimal for one shape is not merely suboptimal for another, it is on the wrong side of doing nothing.

Four rules, and the floor two of them cannot pass. The standard deviation of the covariate imbalance under each rule at n = 200, over 500 trials, as a share of a coin's — which is exactly 2/√n = 0.1414 and is the one number here that needs no simulation. Blocking inside 2 categories and minimising on the same 2 categories are the same rule to within their noise, 61.7% and 61.2%, and the marked line is why: a 2-category split can see 63.7% of the covariate's variance, so a rule that balanced its categories perfectly would still leave 60.3% of a coin's imbalance. Reading the number instead leaves 13.0%, and the worst imbalance it produced in 500 trials was 0.094 standard deviations against the coin's 0.439.
Fig. 3 The previous field’s own picture of the same rules, measured on the covariate they read. Every rule looks like what it was designed to be, because it is being measured by its own objective.
What each rule leaves behind, at 120 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.
Fig. 4 The rules measured on the factor levels they were designed for, in the field that built them. Every one of them is being scored by its own objective there, which is what this essay is about.

Where the threshold column’s number comes from

The middle column is not an arbitrary intermediate and it is worth deriving, because the derivation is what makes the rest of the field predictable rather than merely measured.

A rule that balances x perfectly leaves the part of the outcome’s covariate that is not linear in x. For an indicator at 1 the correlation with x is 0.6623, so the share of the imbalance that survives is 1 − 0.6623² = 0.5614 — and the counted g-imbalance ratio for the rule that reads the number is 0.614, which is that number plus what the rule fails to remove because it balances a sample rather than a population.

The variance of the treatment estimate is not the imbalance, but it moves with it: the estimate’s variance is the noise plus the part of the covariate’s contribution that the imbalance carries into it. So the 0.755 in the table is the 0.614 diluted by the residual noise, and the whole column can be predicted from one correlation and one signal-to-noise ratio.

That is the field’s arithmetic in one paragraph, and the reason to have it is that it says what happens at settings nobody measured: a threshold further out is less correlated with x, so more of its imbalance survives, so the rule is worth less. There is nothing left to discover about the shape of that dependence, only its size — which is what the next essay computes exactly.

What is left after the covariate's mean is balanced perfectly. The share of a coin's imbalance that survives in each function of the covariate, once the covariate's own mean has been balanced exactly. These are closed forms: 1 − ρ², with ρ the correlation between the covariate and the function. For a median split ρ² is exactly 2/π = 0.6366, the same constant that describes how much of a normal covariate a median split can see — the two statements are one integral read in opposite directions. For a threshold at 2, balancing the mean removes 13.1% and leaves 86.9%. For a quadratic it removes nothing, because the two are uncorrelated by symmetry.
Fig. 5 The closed forms behind the middle column: what a perfectly balanced mean removes from each function of the covariate, before any trial is run.

The ranking has no fixed order

The consequence for practice is not that one rule is better than another. It is that the question “which of these rules is the right one?” has a different answer in each column, and nothing in a trial says which column it is in.

Against a linear covariate the rule that reads the number beats blocking by 21 points of variance. Against a threshold it beats it by 14. Against a quadratic it is 2 points behind minimisation and 3 points behind a coin. An experimenter choosing a rule is choosing against a distribution of shapes they cannot observe, which is a decision problem with a name — and it is the one the design field has just spent two essays on, arriving in a completely different subject with the same structure.

The maximin reading of the table above is that the rule which reads the number has a worst case of 1.010 across the three shapes and blocking has 1.008. On that criterion the two are indistinguishable and both are a coin. That is a poor answer, and it can be improved by changing what the rule is allowed to read rather than which rule is used.

Four rules, three shapes, and no ordering that survivesThe variance of the unadjusted treatment estimate under each rule, as a fraction of the variance a coin gives, over 450 trials of 200 units each. Against a covariate that enters linearly the rule that reads the number nearly halves it. Against a threshold at 1.5 it removes about a fifth. Against a quadratic every rule here is at or worse than a coin — they are all optimising a criterion that is one over the variance of an estimate in a model this outcome does not obey, and a constraint that helps nothing still costs something. Nothing in a trial says which column it is in.linear in xa threshold at x = 1quadratic in xa coin1.0001.0001.000blocks in a median split0.6400.9861.305minimisation on that split0.7130.9931.094the rule that reads the number0.4500.8461.121the rule450 trials of 200 units, variance relative to a cointhe best rule in one column is the worst in another
Fig. 6 The same table with the threshold further into the tail. Drag it: the middle column moves towards the right-hand one, because the further out the threshold sits the less correlated it is with the covariate anybody balanced.
What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514).
Fig. 7 The same measurement with a fourth thing varied: what the rule is allowed to read. The top row is this essay’s table and the bottom row is the repair, which the third essay of this field is about.

What the design’s own literature says about this

It is worth being fair about what is and is not new here, because the linearity assumption is not a secret.

Every treatment of covariate-adaptive allocation states the model it optimises for. What is not usually stated is the size of the exposure: that the rule’s advantage falls to nothing at a threshold two standard deviations out, that it reverses sign against a symmetric non-linearity, and that the imbalance measure everybody reports — the one on the covariate itself — is at 0.016 in a trial where the estimate’s variance is barely improved.

Nor is the second-order consequence usually drawn. If the outcome model is non-linear in a covariate and nobody knows it, then the trial is reported as well-balanced, because the balance table shows the covariate; the estimate is more variable than the balance table implies; and the standard error is conservative or not depending on the same unknown shape. Three separate things go wrong together and all three are invisible from inside the trial.

Three ways one covariate can matter. Each curve is a standardised function of the same covariate: the covariate itself, an indicator for exceeding 2, and a quadratic. All three are scaled to unit variance, so an outcome built from any of them is explained to the same degree — what differs is only the shape. A rule that balances the covariate's mean is balancing the first curve exactly and the others only through their correlation with it: 0.3621 for the threshold and 0.000 for the quadratic, which is zero by symmetry.
Fig. 8 A threshold at two standard deviations, where 2.3% of units are above the line. It is a perfectly ordinary way for a covariate to matter — a clinical cutoff, a legal limit, an eligibility rule — and its correlation with the covariate is 0.36.

One constant maps the imbalances to the variances

The two tables are described as related — the estimate’s variance is the noise plus what the imbalance carries into it — and the relation is one number, which can be extracted from the tables themselves and then checked against the cells it was not extracted from.

If a share s of a coin’s estimate variance is the imbalance’s doing, a rule leaving a fraction r of that imbalance gives a variance ratio of 1 − s(1 − r). Fitting s to the two split rules’ linear column gives s = 0.50, and that single value then reproduces:

  • blocks against a quadratic, r = 1.019, predicted 1.010 against a counted 1.008;
  • minimisation against a quadratic, r = 0.975, predicted 0.988 against 0.989;
  • minimisation against a threshold, r = 0.848, predicted 0.924 against 0.919.

Four of the six cells belonging to the two category rules land within half a point, on a constant fitted from two of them. So the middle table is not a second measurement: it is the first table divided by a number, and the number is how much of the estimate’s variance the covariate was responsible for.

That is worth having because the imbalance table is the observable one. An experimenter can compute the imbalance in any candidate function of the covariate without an outcome, and if they can also guess what share of the outcome’s variance that function carries, the variance column follows with no further simulation. The unobservable table is the observable one and one guess.

The exception is the rule whose objective is the variance

Three cells do not fit, and they are the three belonging to the rule that reads the number. Its predicted variance ratios under the same constant are 0.508, 0.807 and 1.044 against counted values of 0.476, 0.755 and 1.010 — better than the map says by 3.2, 5.2 and 3.4 points, in every column and always in the same direction.

The direction is the explanation. The two category rules minimise a distance somebody invented, so the imbalance they leave is a complete description of what they did; the rule that reads the number minimises S_aa − S_ax′S_xx⁻¹S_ax, which is the estimate’s variance, so it picks up whatever the imbalance summary does not contain. Chiefly that is the arm sizes and the covariate’s own spread within each arm, both of which enter the variance and neither of which is a mean imbalance.

The quadratic column is where that matters most and where it is easiest to misread. The rule leaves 8.8% more imbalance in a quadratic than a coin does, which under the map would be a 4.4% worse estimate; it is 1.0% worse. So the honest statement about the last cell is narrower than “the rule is worse than a coin”: its objective is still doing something useful even when the imbalance it induces is actively unhelpful, and what it is doing recovers three quarters of the harm.

That does not rescue the column — a rule that halves the variance in one shape and is a coin in another is still a rule chosen against an unknown — and it does sharpen what the failure is. The rule is not broken by a quadratic outcome. It is idle, and the small negative in the imbalance table overstates the cost.

What an experimenter can actually see

Three of the quantities above are observable inside a trial and three are not, and the split is unkind.

Observable: the imbalance in the covariate, which every trial reports; the imbalance in any function of the covariate somebody thinks to compute; and the standard error of the treatment estimate, which the analysis produces.

Not observable: the variance of the treatment estimate across the trials that might have been run, which is what the whole table is about; the shape the covariate enters the outcome by, which is a fact about the world; and therefore which column of the table the trial is in.

The one bridge between the two lists is that the second column of imbalances — the g-ratios — can be computed for any candidate shape without knowing which is right. An experimenter who suspects a threshold at 1 can compute the imbalance in that indicator and see 0.614 rather than 0.016, and knows immediately that the balance table is flattering. That costs nothing and is not done, because the balance table is a report on the covariate rather than on the outcome model.

One more asymmetry is worth naming. Of the three shapes, only one of them — the linear one — is the shape somebody would have to be unlucky to be wrong about. A threshold is what a clinical cutoff, an eligibility rule or a legal limit produces, and a quadratic is what any covariate with an optimum in the middle of its range produces. Neither is exotic, and neither leaves a trace in the data the design is allowed to look at.

What is claimed here, and what is not

This essay takes the shape a covariate enters an outcome by, and the claim is the table: four rules, three shapes, the variance of the estimate relative to a coin’s, and the ordering reversing between columns.

What stays out and is named as a decision: outcomes that depend on several covariates in different shapes, which is the same argument with more bookkeeping; a treatment effect that varies with the covariate, which is measured to inflate every rule’s variance by the same amount because it enters through the sample mean of the covariate and no assignment rule controls that; binary and survival outcomes, where the criterion is not a least-squares variance at all; and any claim about which shape is common in practice, which is not a statistical question and is not answered here.

The boundary against the previous field is the criterion. What each rule is, why the one that reads the number is the subset criterion applied one unit at a time, and what each of them does to the covariate it reads are established there and used here unchanged. What changes is the outcome, and the outcome was never part of any of those rules.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. The rule that reads the number is required to be the best of the four against a linear covariate and worth nothing at all against a quadratic — both in one check, because either one alone is half the finding. And the threshold is required to sit between them, which is the correlation doing the work and is what the next essay measures exactly.

The refusal beside them is a rule reported on the covariate it reads. Handed the imbalance in x, the rule looks extraordinary: it removes 99.4% of a coin’s, at a determinism of one. Handed the variance of the treatment estimate, with an outcome that depends on a threshold at 1.5, the same rule removes 22.9%. The check throws when the two disagree by that much, because the first number is a property of the criterion the rule maximises and says nothing about the experiment — and it is the number every balance table in the subject reports.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Allocation ruleCovariate balanceDesign criterionEfficiencyImbalanceLinearityMinimisationModel misspecificationNuisance parameterRandomisationStratificationThresholdTreatment effectVariance reduction