Balancing on what was recorded first

Balancing what is known in advance

Four allocation rules, three definitions of balance, and no rule that holds more than one of them. Minimisation keeps the worst factor margin near three patients whether the trial has forty or six hundred and forty — and lets the imbalance in the cross-classified cells climb to 86% of a coin's, because the cells are not what it is watching.

Worth reading first: Randomisation is not balance · The variance removed before the data.

The design field established, and measured exactly, that randomisation does not buy balance. The standardised imbalance on any covariate is 2/√n over every possible assignment, whatever the covariate is, and that result holds no matter how carefully anybody shuffles — it is a fact about sampling rather than about diligence. What randomisation buys is a reference distribution, which is a different and better thing, and the field left the imbalance itself unrepaired.

There is a whole family of rules that repairs it, and none of it had been measured on this site. They read the baseline covariates — sex, age band, centre, whatever was recorded before anything was given to anybody — and use them to decide where the next arrival goes.

That is a different thing from the rules the adaptive field measured, which read outcomes and broke the error rate before any complication was added. Covariates are fixed, observed, and have nothing to do with the response. Almost every conclusion the adaptive field reached comes out the other way here, and this essay is the first of four saying how.

What each rule leaves behind, at 120 patientsFour allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.the arm totalsa coin at every arrival9.30permuted blocks of four0.00blocks inside every cell3.55minimisation1.38the worst factor levela coin at every arrival11.50permuted blocks of four8.66blocks inside every cell4.57minimisation3.41the worst of 24 cellsa coin at every arrival5.27permuted blocks of four5.14blocks inside every cell1.91minimisation4.38500 cohorts of 120, three factors, 24 cellsno rule holds all three columns
Fig. 1 Four rules over the same cohorts and the same seeds, each scored on three imbalances. No rule holds all three columns, and which one each holds is not something its name says. Drag the trial size and watch which bars move and which do not.

Four rules

Complete randomisation is a coin at every arrival. It balances nothing on purpose, and the purpose is real: every allocation is equally likely, which is what makes the reference distribution simple.

Permuted blocks fill blocks of four with two of each arm in random order. This holds the number of patients in each arm to within one at every point in the trial, which matters if the trial might stop early or if the recruitment might drift.

Stratified blocks run a separate block sequence inside each cell of the cross-classification — each combination of sex, age band and centre gets its own. With three factors at two, three and four levels that is 24 cells.

Minimisation does the arithmetic instead. For each arrival, compute what the total imbalance across all the factor margins would be if this patient went to A, and what it would be if they went to B, and send them to whichever is better — with probability p, keeping a coin’s worth of randomness in reserve.

The important word in the last one is margins. Minimisation balances the nine factor levels one at a time, not the 24 cells, which is exactly what lets it work at a sample size where the cells are nearly empty. It is also, as will turn out, the whole of what it cannot promise.

Three definitions of balance, and no rule that holds all three

The three quantities are the arm totals, the worst of the nine factor levels, and the worst of the 24 cells. Every rule is run over the same cohorts.

At 120 patients:

  • Permuted blocks hold the totals at 0.00 — exactly level, always — and leave the worst margin at 8.66 against a coin’s 11.50 and the worst cell at 5.14 against a coin’s 5.27.
  • Stratified blocks hold the worst cell at 1.91 and leave the totals at 3.55.
  • Minimisation holds the worst margin at 3.41 and the totals at 1.38, and leaves the worst cell at 4.38.

Each rule is a coin, or nearly one, on something. That is not a criticism of any of them; it is what happens when a rule watches one quantity and lets a random walk take care of the rest.

What each rule leaves behind, at 40 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 85% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.
Fig. 2 The same three columns at forty patients. Stratification’s totals are worse here in relative terms — 24 part-filled blocks do not have to end level, and at forty patients most of them hold one or two people.

Six numbers against twenty-three

The word margins is doing the work in the paragraph above, and the arithmetic behind it is worth writing down, because it is the whole reason two rules with the same inputs behave differently.

Balancing the levels of a factor with L of them is L − 1 independent constraints once the total is fixed. Three factors at two, three and four levels give 1 + 2 + 3 = 6.

Balancing the cells of the cross-classification is 24 − 1 = 23.

On a hundred and twenty patients that is twenty patients per constraint against a little over five. And the second number is not a comfortable one: a cell holding five patients, filled by blocks of four, is one complete block and one patient left over, so the block structure that guarantees balance inside a cell has nothing to guarantee it with.

That is the whole of why stratification runs out. It is not that twenty-four is a large number in itself; it is that twenty-four constraints on a hundred and twenty units leaves each constraint about five units to be satisfied with, and a block of four cannot be half-filled without leaving an imbalance behind.

What the incomplete blocks add up to

The consequence is quantifiable to an order of magnitude, and the order is the interesting part.

Each of the twenty-four cells ends its recruitment part-way through a block, and a part-filled block of four carries an arm-size imbalance whose variance is of order one. Twenty-four of them, independent across cells, give a total whose standard deviation is about five patients.

Set that against the two rules on either side. Permuted blocks over the whole trial hold the arm-size difference to at most one, at every point, by construction. A coin leaves it with a standard deviation of √120 ≈ eleven.

So stratifying on twenty-four cells recovers roughly half of the arm-size imbalance that blocking removes — the imbalances inside the cells do not cancel, they accumulate, and there are twenty-four of them. The rule that was chosen to improve balance on the covariates has given back a substantial part of the balance on the simplest quantity there is.

Minimisation’s six margins do not have that problem for the same reason they do not have the cells’ sparsity: with twenty patients per constraint and a score that reads every margin at every arrival, there is no block boundary to be caught behind. What it gives up instead is any promise about the cells at all, which is the next section.

What stops growing

The comparison that turns this from a table into a result is what happens as the trial grows.

A coin’s imbalance grows like √n and never stops: the worst margin runs 6.32 at forty patients and 26.28 at six hundred and forty. That is the random walk doing what random walks do, and no amount of patience fixes it.

A rule that repairs an imbalance as it appears leaves a residue that depends on how many things it is tracking, not on how many patients arrive. So it goes flat. Minimisation’s worst margin is 3.00 at forty patients and 3.23 at six hundred and forty. Stratification’s worst cell is 1.86 at one end of that range and 1.90 at the other.

What stops growing, and what does not. The worst imbalance on any of the nine factor levels, averaged over cohorts, against the number of patients. A coin's grows like √n forever: 6.32 at 40 patients and 26.28 at 640. A rule that repairs an imbalance as it appears leaves a residue that depends on how many things it is tracking rather than on how many patients arrive, so it goes flat — minimisation runs 3.00 to 3.23 across a sixteenfold range. Blocking the arm totals is the odd one: it improves the margins by a constant factor and leaves the growth untouched, because the only thing it is watching is the total.
Fig. 3 The worst factor margin against the number of patients, doubling five times. Two lines climb and two are flat, and the flat ones are flat because they are corrective rather than random. The same axes on the cross-classified cells are two figures below.

This is worth stating as a general shape, because it is what the whole family is for. A random rule’s imbalance is a sum of n independent decisions and grows like its square root. A corrective rule’s imbalance is bounded by how far it can fall behind before it starts pushing back, which is a function of the number of margins and the coin it keeps in reserve. The difference between the two is not that one is better, it is that one has a ceiling.

And the ceiling is what makes the rules worth having at large n rather than small. At forty patients minimisation halves the worst margin; at six hundred and forty it divides it by eight.

The blind spot, and it is minimisation’s

The growth picture has a second axis and it is where this essay’s finding is.

Minimisation balances margins. It does not balance cells, and it is not merely worse at it — its cell imbalance grows like a coin’s. At forty patients it leaves 2.61 against a coin’s 3.12; at six hundred and forty it leaves 9.98 against a coin’s 11.61. Eighty-six per cent of the way back to no rule at all, on a quantity most people would assume a “balancing” rule was balancing.

On the cells, minimisation is most of the way back to a coin. The worst imbalance in any cell of the cross-classification, averaged over cohorts, against the number of patients. A coin's grows like √n forever: 3.12 at 40 patients and 11.61 at 640. A rule that repairs an imbalance as it appears leaves a residue that depends on how many things it is tracking rather than on how many patients arrive, so it goes flat — blocks inside every cell runs 1.86 to 1.90 across a sixteenfold range. Minimisation is not that rule here: it balances the nine margins and the 24 cells are not among them, so its cell imbalance reaches 9.98 against a coin's 11.61. A rule chosen for balance can promise nothing about an interaction.
Fig. 4 The same picture on the cross-classified cells. Stratification is the flat line here and minimisation is climbing with the coin, which is the reverse of the previous figure. Each rule is a coin on whatever it is not watching.

The consequence is specific and worth stating in the terms an experimenter would use. Minimisation guarantees that the two arms have similar numbers of women, similar numbers of over-sixties and similar numbers from each centre. It guarantees nothing about whether the two arms have similar numbers of over-sixty women at centre three. So a trial balanced by minimisation and analysed for a main effect is on firm ground, and the same trial analysed for an interaction — does the treatment work differently in older women — is on ground no better than a coin’s.

That is not a bug. Minimisation was defined on margins, its arithmetic is about margins, and it delivers margins. What is missing is anywhere in its description that says so.

What each rule leaves behind, at 80 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.
Fig. 5 Eighty patients. Stratification has begun to hold its cells and its totals are still adrift; the two columns move at different rates because they are limited by different things.

Why twenty-four cells is too many

Stratification’s weakness has a name in the literature — over-stratification — and the arithmetic behind it is worth setting out, because it is the reason minimisation exists at all.

A stratified design runs a block sequence inside each cell. A block of four ends level; a part-filled block does not. So the imbalance a stratified design leaves is, roughly, one part-filled block per cell, and the number of cells is fixed by the factors rather than by the trial. With 24 cells the design carries up to 24 unfinished blocks whatever the trial size — which is why stratification’s totals sit at 3.55 at 120 patients and are still there at six hundred and forty.

The cells also fill unevenly. The three factors here have unequal level probabilities on purpose, because real prognostic factors are not uniform, and the rarest combination — a patient in the smallest age band at the smallest centre — turns up in about four per cent of arrivals. At forty patients that cell holds one or two people and its block never closes.

So the condition under which stratification works is not “enough patients”; it is enough patients per cell, which is n divided by a number that grows multiplicatively with every factor added. Two more factors at three levels each takes 24 cells to 216, and a trial would need thousands of patients for that to behave. Minimisation’s whole design decision is to give up the cells in exchange for scaling with the number of levels — nine here — rather than with their product.

What each rule leaves behind, at 640 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 82% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.
Fig. 6 Six hundred and forty patients, where stratification finally has enough people per cell for its totals to look reasonable in relative terms — and where the coin’s margins have grown to twenty-six. The two rules that repair are unchanged from the forty-patient picture, which is the ceiling again.
What each rule leaves behind, at 320 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 80% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.
Fig. 7 Three hundred and twenty patients, between the two sizes above. Stratification’s cells hold; the coin’s three columns have all grown; and minimisation’s cell column has grown with them.

The coin kept in reserve

Minimisation’s description above contained a parameter that has done no work yet: the arrival goes to the preferred arm with probability p, and to the other arm otherwise. Every number in this essay uses p = 0.8.

Setting p = 1 makes the rule deterministic and improves the balance considerably: the worst factor margin at 120 patients falls from 3.43 to 1.75. It also makes the next assignment a function of information the person enrolling the patient already has, which is the subject of the next essay — and at p = 0.5 the rule is a coin, at a worst margin of 11.42, so the whole of what minimisation buys sits between those two settings of one number.

That is the trade this whole family lives on, and it is worth naming here because the balance numbers above are meaningless without it. A rule can be made to balance arbitrarily well by making it arbitrarily predictable, and at the limit the perfectly balanced design is the one where anybody can name every assignment in advance. Balance and unpredictability are the two things an allocation rule has to buy, and they are bought with the same currency.

Every way of splitting 16 units into two halves. All 12,870 assignments, enumerated. The spread of the standardised imbalance is exactly 2/√n = 0.500, whatever the covariate's own distribution, and 33.3% of assignments differ by more than 0.5 standard deviations. Randomisation does not deliver balance; it delivers a known distribution of imbalance.
Fig. 8 The design field’s picture of what a coin leaves on a single covariate — the imbalance across every possible split of a set of units, with the standardised spread that comes out at exactly 2/√n. Every rule in this essay is an attempt to move off that distribution, and every one of them pays for the move somewhere else.

The one number that is about the outcome

All three imbalances are counts of patients, and a count of patients is not what anybody actually cares about. What matters is whether the arms differ on prognosis — the part of the outcome the covariates predict.

That is measurable directly: give each patient a prognostic score built from their factor levels, and report the standardised difference between the arms’ mean scores. At 120 patients a coin leaves 0.149 of a standard deviation, permuted blocks 0.148, stratification 0.065 and minimisation 0.056.

Two things follow. Minimisation is doing what it was built to do on the quantity that matters, better than any other rule here. And the numbers are all small — a tenth of a standard deviation is not what wrecks a trial — which is the honest context for everything above. The design field’s 2/√n gives 0.183 at 120 patients and the coin’s measured 0.149 is the same statement with a real covariate in place of a worst case.

So the case for these rules is not that a coin produces disasters. It is that the rules cost nothing, remove most of a real effect, and — as the next three essays measure — change what the analysis has to do in ways that are not optional.

There is a version of that argument that this site has already refuted and it is worth keeping separate. Removing prognostic variation is also what blocking does, and blocking’s benefit is a variance reduction with an exact size: the block variance comes out of the residual, and the gain is σ²/(σ² + β²), measured at 0.212 against a predicted 0.200. That is not what these rules buy. A balancing rule does not remove the covariate’s variance from anything — the covariate is still in the outcomes, and if the analysis ignores it the residual still contains it. What the rule removes is the correlation between the covariate and the assignment, which is a different quantity with different consequences, and the third essay in this field is about what those consequences do to a test.

The distinction matters because the two are routinely described in the same words. “Balancing on a covariate improves precision” is true of blocking, where the covariate enters the analysis, and false of a balancing rule whose covariate the analysis then ignores — where, as it turns out, the effect on precision runs the other way.

The allocations this trial could have made, and the ones it could not. One 120-patient trial allocated by minimisation at p = 1, re-randomised 399 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the rule again over the same patients in the same order with the same recorded factors, so what is drawn is the set of experiments that could have happened. The bars are that set; the outline is what shuffling the labels gives, which is the reference distribution of a coin and is what every off-the-shelf permutation routine assumes. The coin's is wider — its 5% point is 1.95 against the rule's 1.09 — because a coin's allocations are less balanced and a less balanced allocation gives a larger statistic. Reading this trial against it makes the test conservative rather than anti-conservative, which is the opposite error from the outcome-adaptive case and for the same structural reason.
Fig. 9 What a rule that balances does to the set of experiments that could have happened. The bars are the allocations minimisation could have produced from one cohort; the outline is a coin’s. The distance between them is what the last two essays in this field are about.

What is being claimed here, and what is not

This field takes allocation rules that read baseline covariates: what each balances, what each gives away in predictability, what the analysis afterwards has to know, and the exact test the design supplies. That was named as not claimed when the adaptive field was written, and it is claimed now.

What stays out: response-adaptive rules, which are the adaptive field’s and read outcomes rather than covariates; continuous covariates, where minimisation needs a discretisation and the choice of discretisation is a design decision this field does not measure; unequal allocation ratios; and covariate-adaptive rules for more than two arms, where the margin score generalises in more than one way and the ways disagree.

The boundary against the design field is the one worth stating. That field measures what randomisation is — an exact 2/√n imbalance and an exact reference distribution — and this one measures what happens when a rule is added on top of it. Nothing here contradicts the 2/√n; every rule below simply stops being the thing that result is about.

The checks

Two claims are gated in this field’s library.

Every rule is a coin on what it does not balance, asserted as four movements at once: minimisation under half a coin’s worst margin at sixty patients, minimisation’s worst margin no larger at six hundred and forty than at sixty while the coin’s has trebled, minimisation’s worst cell above 60% of a coin’s at the large size, and stratification’s worst cell flat across the whole range. Any one of those alone would read as a complaint about a rule; the four together are the shape.

There is one more thing worth saying about how those numbers are produced, because it is what makes them comparable at all. Every rule is run over the same cohort — the same patients, in the same arrival order, with the same factor levels — and only the assignment differs. A rule scored on its own cohort would be a statement about that cohort, and with unequal level probabilities and 24 cells the cohorts vary enough for that to matter at forty patients. It is the same discipline the whole site applies to comparisons of tests and intervals, moved to a place where the thing held fixed is a list of people rather than a list of numbers.

And blocking the totals is not balancing a covariate, which is the refusal. What the covariates themselves supply as a reference distribution is the other half of this field. Permuted blocks improve the margins — they must, since fixing the totals removes one source of margin imbalance — and the improvement is a constant factor with the √n growth untouched, at 78% of a coin’s worst margin at sixty patients and 74% at six hundred and forty. The check requires that ratio to be stable and above 0.6, and throws explicitly if blocking ever beats minimisation on the margins, because that would mean the margins are not being measured.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

ConfoundingCovariate-adaptive randomisationCovariate balanceInteractionMinimisationPermuted blocksPredictabilityRandomisationRandomised block designStratified randomisationStudy design