Balancing what has no levels

The rule that reads the number

Stop categorising and let the rule read the covariate itself. What it should minimise is not an invented distance but the variance of the effect being estimated — and what comes back is not a better constant but a different rate.

Worth reading first: Balancing what is known in advance · A design is a number.

A rule that balances categories is subject to the floor its categories impose. The alternative is to stop categorising, and the question that immediately arises is what such a rule should minimise.

The obvious answers are all inventions. Minimise the difference in covariate means; minimise a t-statistic; minimise a Mahalanobis distance between the arms. Each is a plausible measure of imbalance, none of them is the quantity the trial exists to make small, and choosing between them by argument is exactly the kind of decision this site tries not to make.

There is a non-invented answer available, and it comes from the field two doors down.

What the rule should minimise

The analysis at the end of the trial fits y = α + τ·arm + βx and reports τ̂. Its variance is

σ2SaaSax2/Sxx\frac{\sigma^2}{S_{aa} - S_{ax}^2/S_{xx}}

with a the ±1 arm indicator: the spread of the assignment, less the part of it the covariate explains. So there is nothing to invent. The rule assigns each arrival to whichever arm makes that denominator larger, and it is minimising the variance of the estimate rather than a proxy for it.

That is Dₐ-optimality — the design field’s criterion for a subset of the parameters, which the design fields build for a non-linear model — evaluated one unit at a time instead of once before the experiment. It handles arm-size balance and covariate balance in one number rather than trading them off by hand: an arm that is short of units raises Sₐₐ, and an arm that is short of high-x units raises the subtracted term, and the criterion weighs them against each other in the only currency that matters.

Deterministic assignment would make the rule guessable, so the arrival goes to the preferred arm with probability 0.8 and to the other with probability 0.2 — the same construction, and the same reason, as the rule that can be guessed.

What it leaves

Four rules, and the floor two of them cannot pass. The standard deviation of the covariate imbalance under each rule at n = 200, over 500 trials, as a share of a coin's — which is exactly 2/√n = 0.1414 and is the one number here that needs no simulation. Blocking inside 2 categories and minimising on the same 2 categories are the same rule to within their noise, 61.7% and 61.2%, and the marked line is why: a 2-category split can see 63.7% of the covariate's variance, so a rule that balanced its categories perfectly would still leave 60.3% of a coin's imbalance. Reading the number instead leaves 13.0%, and the worst imbalance it produced in 500 trials was 0.094 standard deviations against the coin's 0.439.
Fig. 1 Four rules at two hundred units, eight hundred trials, as a share of complete randomisation’s exact 2/√n. The bottom row is the rule that reads the number.

At two hundred units it leaves 12.9% of a coin’s covariate imbalance, against 62.2% for minimisation on a median split and a floor of 60.3% for any rule that reads only that split. Its worst case in eight hundred trials was 0.094 standard deviations, against the coin’s 0.439.

Those are large numbers and they are the wrong way to state the result.

The finding is a rate

Run the same comparison at four sizes and the ratios do not settle.

Two rates, not a factor. The standard deviation of the covariate imbalance under three rules, at five trial sizes, 260 trials each, on log axes. The upper line is a coin: its slope is -0.489, against a closed form of exactly −½. The middle line is minimisation on a median split; its slope is -0.519 — the same rate — because inside a category the assignment is still a coin, and what it buys is the constant, 0.654 of a coin's at n = 200. The lower line is the rule that reads x and maximises the information about the treatment effect: slope -0.987, nearly twice as steep. Its advantage is therefore not a number that can be quoted — it is 0.258 of a coin's at n = 50 and 0.065 at n = 800, and it keeps going.
Fig. 2 The standard deviation of the covariate imbalance against the size of the trial, on log axes, five hundred trials at each of five sizes. Two of the three lines are parallel and one is not.

The slopes are −0.478 for complete randomisation, −0.488 for minimisation on a median split, and −1.015 for the rule that reads the number.

The first is the closed form: 2/√n has slope exactly −½, and measuring −0.478 on five hundred trials is what agreement looks like. The second is the same rate, because inside a category the assignment is still a coin — the rule fixes the between-category part and the within-category part is left to average out at the usual √n speed.

The third is a different rate. A rule that reads the number can, in principle, place every arrival to cancel the imbalance the previous arrivals left, so the imbalance is a bounded quantity being divided by n rather than a random walk being divided by n.

That is why the ratios keep moving: 25.4% at fifty units, 18.5% at a hundred, 12.8% at two hundred, 8.7% at four hundred. There is no number to quote for how much better the rule is, because the answer depends on how large the trial is, and the gap grows with it.

Two rates, not a factor. The standard deviation of the covariate imbalance under three rules, at five trial sizes, 260 trials each, on log axes. The upper line is a coin: its slope is -0.489, against a closed form of exactly −½. The middle line is minimisation on a median split; its slope is -0.516 — the same rate — because inside a category the assignment is still a coin, and what it buys is the constant, 0.541 of a coin's at n = 200. The lower line is the rule that reads x and maximises the information about the treatment effect: slope -1.025, nearly twice as steep. Its advantage is therefore not a number that can be quoted — it is 0.157 of a coin's at n = 50 and 0.037 at n = 800, and it keeps going.
Fig. 3 The same three slopes with the rule made deterministic. The rate is unchanged and the constant improves, which is what says the exponent belongs to what the rule reads rather than to how firmly it acts.

Why the rate is different, in one paragraph

The mechanism is worth having explicitly, because “it uses more information” does not explain a change of exponent.

Under complete randomisation the imbalance after n units is a sum of n independent contributions divided by n — a random walk over the sample size — so it goes as √n/n = 1/√n. Balancing categories subtracts the between-category part of each contribution and leaves the rest, which is still a random walk over the same n, so the exponent survives and the constant falls.

A rule that reads the number is doing something structurally different: it observes the accumulated imbalance before each assignment and acts against it. The imbalance is then not a random walk but a process with a restoring force, and a process with a restoring force does not grow with n at all. It stays bounded, and the standardised imbalance — the accumulated total divided by n — therefore falls as 1/n.

The restoring force is imperfect, because the rule follows its preference only four times in five and because a single arrival cannot always cancel what is there. That is why the measured exponent is −1.015 rather than exactly −1, and why the deterministic version has the same exponent with a smaller constant: how firmly the rule acts changes the strength of the restoring force and not whether there is one.

What the exponent is worth

Two quantities scaling differently is the kind of statement that is easy to write and easy to overstate, so it is worth converting.

At fifty units the rule that reads the number is worth about 15 times the sample size as far as covariate imbalance is concerned — the ratio of variances, 1/0.254². At four hundred units it is worth 130 times. Over a range of trial sizes that a single research group might run, the value of the same rule changes by an order of magnitude.

The comparison against the categorising rule is the one that matters in practice, and it moves the same way: 2.4 times the units at fifty, and 50 times at four hundred.

The floor, as a function of how finely the covariate is cut. The curve is a closed form with no simulation in it: k equal-probability categories of a normal covariate see k Σ (φ(zᵢ₋₁) − φ(zᵢ))² of its variance — 63.7% at 2, 79.3% at 3, 86.1% at 4, 89.7% at 5, 94.5% at 8, 96.8% at 12 — so a rule that balanced them perfectly would leave the square root of what is left, 0.603, 0.455, 0.373, 0.321, 0.234, 0.180 of a coin's. The marks are minimisation actually run at each of those cuts, 400 trials each, and they sit on the curve. The lower marks are the rule that reads the number, which is below the floor everywhere because the floor is a fact about categories rather than about balancing. Two categories is the common choice and the worst one: it throws away 36.3% of the covariate before the rule has done anything.
Fig. 4 Why the category rules cannot follow: their floor is a ratio, so it holds at every size, and it is a ratio of two quantities that both fall at the same rate.

What it costs

Three costs, and the first two are smaller than expected.

Predictability. A deterministic version of this rule is guessable in the sense the covariate field measures: an investigator who knows the rule and the covariates can often predict the next assignment. At p = 0.8 the rule is a coin one time in five, which is the same protection minimisation uses and buys the same thing.

Arm sizes. The rule is off by 0.94 units on average, against 1.11 for minimisation and 0 for a forced-equal split. It is not trying to equalise the arms; it is trying to maximise a criterion that arm-size imbalance happens to reduce, and it gets there without being told.

Complexity. The criterion is a ratio of running sums and costs the same per arrival as a table of counts. The version used here maintains six running quantities and is checked against the matrix definition at every arrival of a real experiment — they agree to fourteen decimals, which is the site’s standing requirement whenever one quantity has two implementations.

Four rules, and the floor two of them cannot pass. The standard deviation of the covariate imbalance under each rule at n = 400, over 500 trials, as a share of a coin's — which is exactly 2/√n = 0.1000 and is the one number here that needs no simulation. Blocking inside 2 categories and minimising on the same 2 categories are the same rule to within their noise, 60.8% and 60.2%, and the marked line is why: a 2-category split can see 63.7% of the covariate's variance, so a rule that balanced its categories perfectly would still leave 60.3% of a coin's imbalance. Reading the number instead leaves 9.1%, and the worst imbalance it produced in 500 trials was 0.042 standard deviations against the coin's 0.339.
Fig. 5 Four hundred units. The two category-balancing rules are exactly where they were as a share of the coin, and the rule that reads the number has moved down again.

The probability is the only dial

The rule takes one input beyond the covariate: how often it follows its own preference. That single number decides everything about how it behaves, and the two ends of it are two different objects.

At p = 1 the rule is deterministic. It leaves the least imbalance, it produces an assignment that is a function of the covariates alone, and it is guessable: anybody who knows the rule and has seen the previous arrivals can predict the next assignment exactly. The covariate field measures what that gives away and it is not small.

At p = ½ the rule is a coin and reads nothing.

In between, the rule is a biased coin whose bias is computed from the criterion. At p = 0.8 — the value used throughout this field, and the value the covariate field uses for its own rules — one arrival in five is assigned against the rule’s preference, which is enough to make prediction unreliable and cheap in imbalance: the exponent does not move at all and the constant moves by less than the difference between two sample sizes.

Four rules, and the floor two of them cannot pass. The standard deviation of the covariate imbalance under each rule at n = 200, over 500 trials, as a share of a coin's — which is exactly 2/√n = 0.1414 and is the one number here that needs no simulation. Blocking inside 2 categories and minimising on the same 2 categories are the same rule to within their noise, 61.7% and 61.6%, and the marked line is why: a 2-category split can see 63.7% of the covariate's variance, so a rule that balanced its categories perfectly would still leave 60.3% of a coin's imbalance. Reading the number instead leaves 7.8%, and the worst imbalance it produced in 500 trials was 0.027 standard deviations against the coin's 0.439.
Fig. 6 The four rules with the sequential ones made deterministic. Both improve, neither changes its character, and the protection given up is the same protection in both cases.

That is the whole tuning surface. There is no bandwidth, no threshold, no cut and no schedule — the criterion is fixed by the analysis model and the only choice is how firmly to act on it, which is a choice about predictability rather than about balance.

What it does not do

A rule balances what it is given, and this one is given the mean of x.

What none of the rules balance. Each rule's imbalance in two quantities, both as a share of a coin's: the covariate's mean, which is what the rules are about, and the covariate's spread, which is not. Reading the number takes the mean imbalance to 13.0% of a coin's and leaves the spread at 102.4% — that is, exactly where it found it. Nothing here is a failure of the rules: a rule balances what it is given, and every one of these was given the mean. It is the covadapt finding about cells under balanced margins, on a covariate that has no cells, and it matters for the same reason: an analysis that uses the covariate any way other than linearly is using a quantity the design said nothing about.
Fig. 7 Each rule’s imbalance in the covariate’s spread — measured as the arms’ difference in mean squared value — beside its imbalance in the mean. The rules differ enormously in one column and not at all in the other.

The imbalance in the covariate’s spread is what complete randomisation leaves it under every rule here, including the one that reads the number: the ratios are 1.000, 0.999, 0.956 and 0.969. A rule that has removed six sevenths of the imbalance in the mean has removed nothing of the imbalance in the second moment.

This is not a defect and it is not a surprise once stated: the criterion contains x linearly because the model contains x linearly, and a design optimised for a model is optimised for that model. An analysis that uses x any other way — a quadratic term, a threshold effect, a subgroup at the extremes — is using a quantity the design said nothing about.

It is the multi-arm field’s finding about cells under balanced margins in a setting with no cells, and the general form is worth stating once: a balancing rule is a promise about a stated functional of the covariates and about nothing else. Whichever functional the analysis turns out to use is the one that had to be declared.

The exponent, read forwards

The four ratios — 25.4%, 18.5%, 12.8% and 8.7% at fifty, a hundred, two hundred and four hundred units — are the two exponents’ difference in the only form a designer can use, and they confirm it. Over the eightfold range they fall by a factor of 2.92, which is an exponent of log(2.92)/log(8) = 0.515 against the fitted 1.015 − 0.478 = 0.537. Four per cent apart, from a ratio of two quantities neither of which was fitted to the other.

Run it forwards and it says when the two kinds of rule stop being comparable. At two hundred units the continuous rule leaves 12.8% of a coin’s imbalance against the categorising rule’s 62.2%, a ratio of 0.21. Extrapolating at n^(−0.537), that ratio reaches a tenth by about eight hundred units and a hundredth by twenty thousand.

So there is a size at which the comparison stops being about degree. Below a few hundred units the two rules are the same kind of object with different constants; above about a thousand they are not, because one of them is still dividing a random walk by n and the other has stopped having a random walk to divide. A rule that is twice as good on a trial of fifty is fifty times as good on a trial of four hundred, and no protocol distinguishes them by name.

The balance table gets worse as the trial gets bigger

The spread column is read above as a limitation of the criterion, and combining it with the exponent makes it a limitation that grows.

Every rule leaves the covariate’s second moment where a coin leaves it — the ratios are 1.000, 0.999, 0.956 and 0.969 — so the spread imbalance falls at the coin’s rate, 1/√n, for the continuous rule as much as for anything else. Its mean imbalance falls at 1/n. The two therefore diverge: at two hundred units the mean is balanced 7.6 times better than the spread relative to a coin, at four hundred 11.1 times, and at twenty thousand it would be about a hundred.

The better the rule, the more misleading its own balance table. A trial reporting that the arms are matched on age to three decimal places is reporting the one functional the rule was optimising, and the larger the trial the more spectacular that number looks and the less it says about anything else. A reader who takes a very small covariate imbalance as evidence that the arms are comparable is reading a quantity that was driven to zero on purpose beside quantities that were not touched.

The cheap repair follows from the same arithmetic and needs no outcome: report the imbalance in a second functional beside the first. On this rule at two hundred units those are 0.129 and 0.969 of a coin’s, and the gap between them is the whole content of the previous section stated in two numbers a trial already has.

The arm sizes are worth one line in the same spirit. The rule is off by 0.94 units against minimisation’s 1.11, so it balances the arms better than the rule that is trying to — because arm imbalance enters the variance it minimises and does not enter a count of categories at all.

What it is worth to the trial, rather than to the covariate

Everything above is measured on the covariate, which is where a design’s work is visible and not where its value is.

What the imbalance costs is decided by how strongly the covariate drives the outcome. If it drives it not at all, a perfectly balanced trial and a coin-tossed one estimate the same effect with the same precision, and every rule here has spent its complexity on nothing. If it drives the outcome entirely, the imbalance is the error, and removing it removes the error.

The honest summary is that a balancing rule converts a known source of variation into a removed one, and that the analysis can do the same thing afterwards by including the covariate in the model. Which of the two is worth more, and whether they overlap, is the question the next essay measures — and the answer is not the reassuring one, because an analysis that does not know what the design did prices an imbalance the design has already removed.

That is the reason this field has four essays rather than two. A rule that leaves 12.9% of a coin’s imbalance is a good rule by the only measure applied so far, and what happens to a trial afterwards depends on something the rule never sees.

The same picture at every cut

One consequence of the previous essay is worth restating here from the other side. The rule that reads the number does not care how the covariate would have been categorised, so its imbalance is the same whatever cut a protocol proposes — 0.0772 of a coin’s at two categories, at three, at four and at eight, in a sweep where the categorising rule’s own number moves from 0.61 to 0.26.

That is the practical form of “the floor is a fact about categories”. Cutting more finely is a real improvement to a rule that reads categories and is not an approach to the rule that does not; the two are not points on one scale, and the finest cut anybody uses in practice still leaves the category rule at three times the continuous one’s imbalance.

Where the rule comes from

The construction here is not new and its provenance is worth naming, because the interesting part is which field it comes from.

Optimal design is a subject about choosing settings before an experiment, and its criteria are functionals of an information matrix. A covariate-adaptive assignment rule is normally treated as a subject about balance, and its criteria are measures of imbalance. What connects them is that the information matrix for the treatment effect is a function of the assignment, so a sequential assignment rule is a sequential design and the criterion is already available.

Once that connection is made, the choice between the invented measures of imbalance stops being a matter of preference. Minimising xˉAxˉB|\bar{x}_A - \bar{x}_B| is minimising the wrong thing when the arms are unequal in size; minimising a t-statistic is closer and still not the variance; maximising Sₐₐ − Sₐₓ²/Sₓₓ is minimising the variance itself, and there is nothing left to argue about.

What each rule leaves behind, at 120 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.
Fig. 8 The covariate field’s rules, on covariates that have levels. Each holds one margin flat and lets the others grow like a coin’s — which is the same statement as this essay’s last section, on quantities where “flat” and “like a coin’s” are counts rather than moments.
What balancing several numbers at once costs each of themThe criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 12.7% of a coin's imbalance in it; balancing eight leaves 23.2% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.00.1000.2000.3001248how many covariates the rule is asked to balanceimbalance left in each of them, as a share of a coin's12.7%15.2%18.4%23.2%a coin would be at 1.0, off the top of this frame200 trials per point at n = 2001: 0.127, 2: 0.152, 4: 0.184, 8: 0.232
Fig. 9 Drag the number of units. What the rule can hold when it is asked to balance several covariates at once is the last essay of this field, and the criterion generalises to it without a word changing.

What is claimed here, and what is not

This essay claims a covariate-adaptive rule built on the variance of the treatment effect: the criterion, the imbalance it leaves, and the exponent that separates it from every rule that balances categories.

What stays out and is named as a decision: the analysis after this rule, which is the next essay and is where the interesting cost is; more than two arms, where the criterion generalises and the arm-size term stops being a single number; and covariates observed with error or arriving late, where a rule that reads the number is reading a number that is not the one the analysis will use.

The checks, and what they are checked against

Three claims are gated in this field’s library. The coin’s imbalance is required to match 2/√n at three sizes. The three slopes are required to come out at −½, −½ and steeper than −0.85, which is the finding stated as the only thing it can be stated as. And the running criterion is required to agree with the matrix definition at every arrival of a real experiment, with one covariate and with three, because a rule implemented twice is two rules until they are compared.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Biased coinContinuous covariateConvergence rateCovariate-adaptive randomisationCovariate imbalanceDₐ-optimalityExperimental designInformation matrixMinimisationMonte CarloOptimal designRandomisationStandardised difference