The rule that reads the number
Worth reading first: Balancing what is known in advance · A design is a number.
A rule that balances categories is subject to the floor its categories impose. The alternative is to stop categorising, and the question that immediately arises is what such a rule should minimise.
The obvious answers are all inventions. Minimise the difference in covariate means; minimise a t-statistic; minimise a Mahalanobis distance between the arms. Each is a plausible measure of imbalance, none of them is the quantity the trial exists to make small, and choosing between them by argument is exactly the kind of decision this site tries not to make.
There is a non-invented answer available, and it comes from the field two doors down.
What the rule should minimise
The analysis at the end of the trial fits y = α + τ·arm + βx and reports τ̂. Its variance is
with a the ±1 arm indicator: the spread of the assignment, less the part of it the covariate explains. So there is nothing to invent. The rule assigns each arrival to whichever arm makes that denominator larger, and it is minimising the variance of the estimate rather than a proxy for it.
That is Dₐ-optimality — the design field’s criterion for a subset of the parameters, which the design fields build for a non-linear model — evaluated one unit at a time instead of once before the experiment. It handles arm-size balance and covariate balance in one number rather than trading them off by hand: an arm that is short of units raises Sₐₐ, and an arm that is short of high-x units raises the subtracted term, and the criterion weighs them against each other in the only currency that matters.
Deterministic assignment would make the rule guessable, so the arrival goes to the preferred arm with probability 0.8 and to the other with probability 0.2 — the same construction, and the same reason, as the rule that can be guessed.
What it leaves
At two hundred units it leaves 12.9% of a coin’s covariate imbalance, against 62.2% for minimisation on a median split and a floor of 60.3% for any rule that reads only that split. Its worst case in eight hundred trials was 0.094 standard deviations, against the coin’s 0.439.
Those are large numbers and they are the wrong way to state the result.
The finding is a rate
Run the same comparison at four sizes and the ratios do not settle.
The slopes are −0.478 for complete randomisation, −0.488 for minimisation on a median split, and −1.015 for the rule that reads the number.
The first is the closed form: 2/√n has slope exactly −½, and measuring −0.478 on five hundred trials is what agreement looks like. The second is the same rate, because inside a category the assignment is still a coin — the rule fixes the between-category part and the within-category part is left to average out at the usual √n speed.
The third is a different rate. A rule that reads the number can, in principle, place every arrival to cancel the imbalance the previous arrivals left, so the imbalance is a bounded quantity being divided by n rather than a random walk being divided by n.
That is why the ratios keep moving: 25.4% at fifty units, 18.5% at a hundred, 12.8% at two hundred, 8.7% at four hundred. There is no number to quote for how much better the rule is, because the answer depends on how large the trial is, and the gap grows with it.
Why the rate is different, in one paragraph
The mechanism is worth having explicitly, because “it uses more information” does not explain a change of exponent.
Under complete randomisation the imbalance after n units is a sum of n independent contributions divided by n — a random walk over the sample size — so it goes as √n/n = 1/√n. Balancing categories subtracts the between-category part of each contribution and leaves the rest, which is still a random walk over the same n, so the exponent survives and the constant falls.
A rule that reads the number is doing something structurally different: it observes the accumulated imbalance before each assignment and acts against it. The imbalance is then not a random walk but a process with a restoring force, and a process with a restoring force does not grow with n at all. It stays bounded, and the standardised imbalance — the accumulated total divided by n — therefore falls as 1/n.
The restoring force is imperfect, because the rule follows its preference only four times in five and because a single arrival cannot always cancel what is there. That is why the measured exponent is −1.015 rather than exactly −1, and why the deterministic version has the same exponent with a smaller constant: how firmly the rule acts changes the strength of the restoring force and not whether there is one.
What the exponent is worth
Two quantities scaling differently is the kind of statement that is easy to write and easy to overstate, so it is worth converting.
At fifty units the rule that reads the number is worth about 15 times the sample size as far as covariate imbalance is concerned — the ratio of variances, 1/0.254². At four hundred units it is worth 130 times. Over a range of trial sizes that a single research group might run, the value of the same rule changes by an order of magnitude.
The comparison against the categorising rule is the one that matters in practice, and it moves the same way: 2.4 times the units at fifty, and 50 times at four hundred.
What it costs
Three costs, and the first two are smaller than expected.
Predictability. A deterministic version of this rule is guessable in the sense the covariate field measures: an investigator who knows the rule and the covariates can often predict the next assignment. At p = 0.8 the rule is a coin one time in five, which is the same protection minimisation uses and buys the same thing.
Arm sizes. The rule is off by 0.94 units on average, against 1.11 for minimisation and 0 for a forced-equal split. It is not trying to equalise the arms; it is trying to maximise a criterion that arm-size imbalance happens to reduce, and it gets there without being told.
Complexity. The criterion is a ratio of running sums and costs the same per arrival as a table of counts. The version used here maintains six running quantities and is checked against the matrix definition at every arrival of a real experiment — they agree to fourteen decimals, which is the site’s standing requirement whenever one quantity has two implementations.
The probability is the only dial
The rule takes one input beyond the covariate: how often it follows its own preference. That single number decides everything about how it behaves, and the two ends of it are two different objects.
At p = 1 the rule is deterministic. It leaves the least imbalance, it produces an assignment that is a function of the covariates alone, and it is guessable: anybody who knows the rule and has seen the previous arrivals can predict the next assignment exactly. The covariate field measures what that gives away and it is not small.
At p = ½ the rule is a coin and reads nothing.
In between, the rule is a biased coin whose bias is computed from the criterion. At p = 0.8 — the value used throughout this field, and the value the covariate field uses for its own rules — one arrival in five is assigned against the rule’s preference, which is enough to make prediction unreliable and cheap in imbalance: the exponent does not move at all and the constant moves by less than the difference between two sample sizes.
That is the whole tuning surface. There is no bandwidth, no threshold, no cut and no schedule — the criterion is fixed by the analysis model and the only choice is how firmly to act on it, which is a choice about predictability rather than about balance.
What it does not do
A rule balances what it is given, and this one is given the mean of x.
The imbalance in the covariate’s spread is what complete randomisation leaves it under every rule here, including the one that reads the number: the ratios are 1.000, 0.999, 0.956 and 0.969. A rule that has removed six sevenths of the imbalance in the mean has removed nothing of the imbalance in the second moment.
This is not a defect and it is not a surprise once stated: the criterion contains x linearly because the model contains x linearly, and a design optimised for a model is optimised for that model. An analysis that uses x any other way — a quadratic term, a threshold effect, a subgroup at the extremes — is using a quantity the design said nothing about.
It is the multi-arm field’s finding about cells under balanced margins in a setting with no cells, and the general form is worth stating once: a balancing rule is a promise about a stated functional of the covariates and about nothing else. Whichever functional the analysis turns out to use is the one that had to be declared.
The exponent, read forwards
The four ratios — 25.4%, 18.5%, 12.8% and 8.7% at fifty, a hundred, two hundred and four hundred units — are the two exponents’ difference in the only form a designer can use, and they confirm it. Over the eightfold range they fall by a factor of 2.92, which is an exponent of log(2.92)/log(8) = 0.515 against the fitted 1.015 − 0.478 = 0.537. Four per cent apart, from a ratio of two quantities neither of which was fitted to the other.
Run it forwards and it says when the two kinds of rule stop being comparable. At two hundred units the continuous rule leaves 12.8% of a coin’s imbalance against the categorising rule’s 62.2%, a ratio of 0.21. Extrapolating at n^(−0.537), that ratio reaches a tenth by about eight hundred units and a hundredth by twenty thousand.
So there is a size at which the comparison stops being about degree. Below a few hundred units the two rules are the same kind of object with different constants; above about a thousand they are not, because one of them is still dividing a random walk by n and the other has stopped having a random walk to divide. A rule that is twice as good on a trial of fifty is fifty times as good on a trial of four hundred, and no protocol distinguishes them by name.
The balance table gets worse as the trial gets bigger
The spread column is read above as a limitation of the criterion, and combining it with the exponent makes it a limitation that grows.
Every rule leaves the covariate’s second moment where a coin leaves it — the ratios are 1.000, 0.999, 0.956 and 0.969 — so the spread imbalance falls at the coin’s rate, 1/√n, for the continuous rule as much as for anything else. Its mean imbalance falls at 1/n. The two therefore diverge: at two hundred units the mean is balanced 7.6 times better than the spread relative to a coin, at four hundred 11.1 times, and at twenty thousand it would be about a hundred.
The better the rule, the more misleading its own balance table. A trial reporting that the arms are matched on age to three decimal places is reporting the one functional the rule was optimising, and the larger the trial the more spectacular that number looks and the less it says about anything else. A reader who takes a very small covariate imbalance as evidence that the arms are comparable is reading a quantity that was driven to zero on purpose beside quantities that were not touched.
The cheap repair follows from the same arithmetic and needs no outcome: report the imbalance in a second functional beside the first. On this rule at two hundred units those are 0.129 and 0.969 of a coin’s, and the gap between them is the whole content of the previous section stated in two numbers a trial already has.
The arm sizes are worth one line in the same spirit. The rule is off by 0.94 units against minimisation’s 1.11, so it balances the arms better than the rule that is trying to — because arm imbalance enters the variance it minimises and does not enter a count of categories at all.
What it is worth to the trial, rather than to the covariate
Everything above is measured on the covariate, which is where a design’s work is visible and not where its value is.
What the imbalance costs is decided by how strongly the covariate drives the outcome. If it drives it not at all, a perfectly balanced trial and a coin-tossed one estimate the same effect with the same precision, and every rule here has spent its complexity on nothing. If it drives the outcome entirely, the imbalance is the error, and removing it removes the error.
The honest summary is that a balancing rule converts a known source of variation into a removed one, and that the analysis can do the same thing afterwards by including the covariate in the model. Which of the two is worth more, and whether they overlap, is the question the next essay measures — and the answer is not the reassuring one, because an analysis that does not know what the design did prices an imbalance the design has already removed.
That is the reason this field has four essays rather than two. A rule that leaves 12.9% of a coin’s imbalance is a good rule by the only measure applied so far, and what happens to a trial afterwards depends on something the rule never sees.
The same picture at every cut
One consequence of the previous essay is worth restating here from the other side. The rule that reads the number does not care how the covariate would have been categorised, so its imbalance is the same whatever cut a protocol proposes — 0.0772 of a coin’s at two categories, at three, at four and at eight, in a sweep where the categorising rule’s own number moves from 0.61 to 0.26.
That is the practical form of “the floor is a fact about categories”. Cutting more finely is a real improvement to a rule that reads categories and is not an approach to the rule that does not; the two are not points on one scale, and the finest cut anybody uses in practice still leaves the category rule at three times the continuous one’s imbalance.
Where the rule comes from
The construction here is not new and its provenance is worth naming, because the interesting part is which field it comes from.
Optimal design is a subject about choosing settings before an experiment, and its criteria are functionals of an information matrix. A covariate-adaptive assignment rule is normally treated as a subject about balance, and its criteria are measures of imbalance. What connects them is that the information matrix for the treatment effect is a function of the assignment, so a sequential assignment rule is a sequential design and the criterion is already available.
Once that connection is made, the choice between the invented measures of imbalance stops being a matter of preference. Minimising is minimising the wrong thing when the arms are unequal in size; minimising a t-statistic is closer and still not the variance; maximising Sₐₐ − Sₐₓ²/Sₓₓ is minimising the variance itself, and there is nothing left to argue about.
What is claimed here, and what is not
This essay claims a covariate-adaptive rule built on the variance of the treatment effect: the criterion, the imbalance it leaves, and the exponent that separates it from every rule that balances categories.
What stays out and is named as a decision: the analysis after this rule, which is the next essay and is where the interesting cost is; more than two arms, where the criterion generalises and the arm-size term stops being a single number; and covariates observed with error or arriving late, where a rule that reads the number is reading a number that is not the one the analysis will use.
The checks, and what they are checked against
Three claims are gated in this field’s library. The coin’s imbalance is required to match 2/√n at three sizes. The three slopes are required to come out at −½, −½ and steeper than −0.85, which is the finding stated as the only thing it can be stated as. And the running criterion is required to agree with the matrix definition at every arrival of a real experiment, with one covariate and with three, because a rule implemented twice is two rules until they are compared.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Guessing one arm in three — both name covariate-adaptive randomisation, experimental design, minimisation, monte carlo, randomisation
- A proposal that moves more than two units — both name experimental design, minimisation, monte carlo, randomisation
- Balancing towards unequal targets — both name covariate-adaptive randomisation, experimental design, minimisation, monte carlo
- Where the gain is, and where the decision is — both name experimental design, minimisation, monte carlo, randomisation
- A cut is not a polynomial, and it does not have to be — both name continuous covariate, experimental design, randomisation
- Augmenting a design that has already run — both name experimental design, information matrix, optimal design
Named objects
A flat tag is an object no other essay names yet.
Biased coinContinuous covariateConvergence rateCovariate-adaptive randomisationCovariate imbalanceDₐ-optimalityExperimental designInformation matrixMinimisationMonte CarloOptimal designRandomisationStandardised difference