Balancing on what was recorded first

The rule that can be guessed

A balancing rule improves as it becomes more deterministic, and a deterministic rule can be worked out in advance from information the person enrolling the patient already has. At full determinism 87.6% of assignments are guessable, and an investigator who acts on the guess produces a treatment effect of three quarters of a standard deviation where the truth is zero.

Worth reading first: Randomisation is not balance.

Every rule in the previous essay improves as it becomes more deterministic. Minimisation with the coin removed balances better than minimisation with it kept; a stratified block sequence balances better than a coin. The improvement is real and measured.

It is also, exactly, a description of how much of the allocation is a function of things that are already known. And the person who knows those things is standing in front of the patient deciding whether to enrol them.

What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 120 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 49.7% at p = 0.5, which is a coin and cannot be beaten, and 87.6% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.5 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.75 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).
Fig. 1 Two curves against the rule’s own randomisation probability. The upper one is the share of assignments an investigator can name in advance; the lower is what that is worth to somebody willing to act on it. They move together because they are the same quantity twice.

What can be guessed

The guesser knows the rule, the factors of everyone already enrolled, and every assignment so far. That is not an unreasonable adversary; it is the ordinary situation in an open-label trial, where the allocation is not concealed after the fact and the investigator has the enrolment log.

For minimisation the guesser computes the same margin score the rule computes and names the arm it prefers. Correct with probability p by construction — except at a tie, where the rule falls back on a coin and cannot be beaten.

Counted over cohorts of 120: 49.7% at p = 0.5, 66.4% at p = 0.7, 73.8% at p = 0.8, and 87.6% at p = 1.

The last number is the informative one. A rule with no random element at all is guessable 87.6% of the time rather than 100%, and the twelve-point shortfall is entirely ties — arrivals where both arms give the same margin score and the rule flips a coin. Ties are common because the score is a sum of small integers, so a deterministic minimisation retains a real amount of unpredictability for free.

Permuted blocks are the surprise. The guesser tracks the current block and names whichever arm it is short of; at the end of a block of four the answer is forced. Counted: 71.0%.

That is more predictable than minimisation at p = 0.7. The rule most trials actually use, chosen for reasons that have nothing to do with covariates, gives away more than the rule that gets criticised for being predictable.

The upper curve is the probability itself

The share an investigator can name in advance has a closed form, and it is the simplest one available.

On a two-arm trial, a rule that follows its own preference with probability p sends the patient to the other arm with probability 1p1 - p. So an investigator who knows the score and names the preferred arm is right with probability exactly p.

The upper curve is the line p. At p = 1 the guess is certain; at p = 0.5 the rule is a coin and the guess is worth nothing over chance; and everywhere in between the guessable share and the randomisation probability are the same number under two names.

Which is why the reserved coin cannot be tuned away. Whatever p buys in balance it hands over in predictability, unit for unit, and the only thing that changes along the dial is which of the two the trial would rather have.

What a guess is worth

Predictability is only a problem if somebody can act on it, and the mechanism by which they can is worth setting out precisely, because it does not involve breaking any rule.

The investigator predicts the arm they favour — perhaps they believe the treatment works, perhaps they want it to. When the prediction is their favoured arm, they enrol a patient with a slightly better prognosis: a marginally healthier candidate from the waiting list, an earlier stage, a case they judge more likely to do well. When the prediction is the other arm, they wait.

Nothing about the randomisation has been violated. The rule was followed exactly, the assignment was whatever the rule said, and no envelope was opened early. The bias entered through who was enrolled, which is the one thing an allocation rule has no control over at all.

Model it as δ: a selected patient is δ of a standard deviation better than average, and a non-selected one δ worse. Then the estimated treatment effect has expectation

2δ(2g − 1)

with g the guess rate. The factor of two is the two arms moving in opposite directions, and it is the part most easily left out — leaving it out halves the answer.

At δ = ½ and 120 patients, counted: 0.0442 of a standard deviation at p = 0.5, 0.3229 at p = 0.7, 0.4694 at p = 0.8, and 0.7474 at p = 1. The closed form gives 0.3275, 0.4754 and 0.7527 at the same three settings — two routes, one counting correct guesses and one counting what the estimate did, agreeing to within half a per cent.

A coin gives 0.0084, which is zero within its own standard error.

What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 120 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 49.7% at p = 0.5, which is a coin and cannot be beaten, and 87.6% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 1 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 1.50 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).
Fig. 2 The same picture with a selected patient a full standard deviation better than average. The guess curve is unchanged — it is a property of the rule — and the bias curve doubles, because the bias is linear in what selection is worth.

Three quarters of a standard deviation is not a subtle effect. It is larger than most treatment effects anybody is looking for, and it has been produced under a null with a rule that was followed correctly throughout.

Where the closed form stops

At p = 0.5 minimisation is a coin, the mechanism supplies nothing, and the closed form gives zero. The count gives 0.0442, which is small and is not noise.

The residue is an artefact of the estimator rather than of the rule, and it is worth naming rather than absorbing. The effect is a difference of two means whose group sizes are themselves driven by the assignments the preference was tracking, so the numerator and denominator of each arm’s mean are correlated, and a ratio of correlated quantities is not unbiased at finite n. It is about 0.07δ at 120 patients and it shrinks with the trial.

It is drawn rather than asserted on, and the field’s check is applied only where p ≥ 0.7 — with the reason stated, because a check quietly loosened to accommodate a residue nobody understood would be worse than no check.

What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 400 cohorts of 240 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 50.0% at p = 0.5, which is a coin and cannot be beaten, and 88.0% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.5 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.77 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).
Fig. 3 The same two curves on a trial twice the size. Neither moves: the guess rate is a property of the rule rather than of the trial, and the bias it buys is a property of the guess rate.

Where the other rules sit

Minimisation gets the attention and the other three are worth measuring, because the ranking is not the one the discussion implies.

  • A coin: guessable 49.7%, bias 0.0084. Nothing available and nothing extracted.
  • Stratified blocks: guessable 67.3%, bias 0.3439.
  • Permuted blocks: guessable 71.0%, bias 0.4119.
  • Minimisation at p = 0.8: guessable 73.8%, bias 0.4694.

The three that are not a coin are within a few points of each other. Permuted blocks — the default in a large share of trials, chosen for reasons of administration rather than of covariates — are more predictable than a minimisation that keeps a coin one time in five, and deliver 88% of its bias.

The reason is the same in both cases and it is worth stating plainly, because it means this is not a criticism of minimisation specifically. Any rule that corrects an imbalance is predictable to the extent that it corrects. A block that is short of one arm will supply that arm; a margin that is behind will be topped up. Correcting and being guessable are the same property described from opposite sides, and a rule that could correct without being predictable would have to be reading something the guesser cannot see.

Stratified blocks sitting slightly below plain blocks is the one detail that needs explaining, and it is a size effect: with 24 cells at 120 patients most cells hold two to six people, most blocks are part-filled at any moment, and a part-filled block early in its sequence is close to a coin. The guesser does better against a rule whose blocks actually complete.

The allocations this trial could have made, and the ones it could not. One 120-patient trial allocated by minimisation at p = 0.85, re-randomised 399 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the rule again over the same patients in the same order with the same recorded factors, so what is drawn is the set of experiments that could have happened. The bars are that set; the outline is what shuffling the labels gives, which is the reference distribution of a coin and is what every off-the-shelf permutation routine assumes. The coin's is wider — its 5% point is 1.95 against the rule's 1.15 — because a coin's allocations are less balanced and a less balanced allocation gives a larger statistic. Reading this trial against it makes the test conservative rather than anti-conservative, which is the opposite error from the outcome-adaptive case and for the same structural reason.
Fig. 4 The set of allocations a rule at p = 0.85 could have produced. It is wider than a deterministic rule’s and narrower than a coin’s, and its width is the same quantity the guess rate measures from the other side: an allocation nobody can predict is one the rule could have made many ways.

Blinding is the answer and it is not always available

The repair is obvious and it is not statistical: conceal the allocation. If the investigator does not know what has already been assigned, the guess has nothing to work with, and the whole mechanism collapses regardless of how deterministic the rule is.

That is why this is not usually a scandal. A double-blind trial with central allocation gives the enrolling clinician nothing, and minimisation at p = 1 is then perfectly safe.

It is worth noticing that concealment does not make the rule less predictable. The guess rate is a property of the rule and the covariates, and it is 87.6% whether or not anybody is in a position to compute it. What concealment removes is the input: the guesser needs the history of assignments, and without it the margin score cannot be evaluated at all. So the two defences are of different kinds — keeping a coin in the rule reduces what can be known, and concealment reduces who can know it — and they compose. A trial with both is protected twice, and the numbers above say how much each half is worth on its own.

There is also an asymmetry in when each defence has to be decided. The coin in the rule is a design decision, fixed in the protocol before anybody arrives, and cannot be revisited. Concealment is an operational arrangement that can be tightened during a trial. So a trial that is unsure about its own blinding should keep more coin than one that is confident, and the cost of doing so is the balance column of the previous essay — which, at p = 0.8, is a few tenths of a patient on the worst margin.

The cases where it is not available are the ones worth listing, because they are common: surgery against medicine, a device against a drug, anything where the two arms are visibly different; trials where the site does its own allocation; and trials where the outcome is assessed by the same person who enrolled. In those the coin kept in reserve is the only defence, and how much of it to keep is a design decision with numbers attached.

What a rule gives away by being predictableMinimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 120 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 49.7% at p = 0.5, which is a coin and cannot be beaten, and 87.6% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.25 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.37 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).00.2500.5000.75010.5000.6000.7000.8000.9001probability the rule follows its own preferenceshare guessed, and bias in standard deviationshalf, which is what a coin gives awayupper: share of assignments guessed · lower: bias extracted · dashed: 2δ(2g − 1)500 cohorts of 120 at each of 6 settings, δ = 0.2587.6% guessable buys 0.37 of bias
Fig. 5 At a quarter of a standard deviation — a selection effect small enough that nobody would notice themselves exercising it. The bias at p = 1 is still a third of a standard deviation, which is a treatment effect worth publishing. Drag δ to see the whole family.

What it costs to keep the coin

The trade is between the two essays, and it is a real trade rather than a free lunch.

Dropping p from 1 to 0.7 takes the guess rate from 87.6% to 66.4% and the bias from 0.7474 to 0.3229 — well over half of it gone. It also takes the worst factor margin from 1.75 patients to 5.00, which is nearly three times the imbalance. Neither half of that is negligible.

What makes it a genuine choice rather than an optimisation is that the two curves have different shapes, and the difference is the opposite way round from what the word “trade-off” suggests.

The guess rate is close to linear in p — 49.7%, 66.4%, 80.8%, 87.6% at p = 0.5, 0.7, 0.9 and 1, which is between 6.9 and 8.3 points per tenth all the way along. The worst margin is not: 11.42, 5.00, 2.36 and 1.75 at the same four settings, which is 3.21 patients per tenth at the coin end and 0.61 at the deterministic end, a factor of five apart.

So the balance is bought cheaply at the start and expensively at the finish, while the predictability is given away at a flat rate throughout. The last tenth of determinism costs 6.9 points of guessability and buys 0.61 of a patient; the first costs 8.3 points and buys 3.21. There is no knee in the sense of a free lunch, and there is a clear direction: the part of minimisation worth having is the part near a coin, and the part that makes it notorious is the part that buys least.

Where to sit is still a judgement about which failure matters in a particular trial rather than a number the arithmetic supplies. But the arithmetic does say that a trial worried about predictability should give up the top of the range, where it is paying the most for the least.

What each rule leaves behind, at 120 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.
Fig. 6 The balance side of the same trade, from the previous essay. Minimisation’s advantage over a coin on the margins is a factor of three, and almost all of it survives keeping a coin in reserve one time in five.

The same shape, twice, in two other fields

This is the third time this site has measured a defect that enters through a decision nobody thought of as part of the analysis, and putting the three side by side is the clearest way to see what they have in common.

The garden of forking paths is an analyst choosing between analyses after seeing the data: twenty available analyses turn a 5% procedure into a 57% one, and every individual analysis is correct. The winner’s curse is an estimate reported from the arm chosen for being ahead: inflated by a factor of 2.07 at the sample size measured, and every individual estimate is unbiased for the arm it came from. And this is an investigator choosing who to enrol after predicting where they will go.

All three are selections made on information correlated with the outcome, by somebody following every stated rule, and in none of the three is there a step anybody would identify as wrong. What differs is only where the selection sits: after the data in the first, after the interim in the second, and before the patient exists in the third — which is the earliest any of them has been found and the hardest to see afterwards, because the selected-against patients are not in the dataset at all.

The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.
Fig. 7 The forking-paths measurement, for the comparison. There the choice is between analyses of one dataset; here it is between candidates for one slot. The arithmetic that turns a choice into a bias does not care which.

That last point is what makes this one undetectable by any analysis of the trial’s own data. A forking-paths problem leaves the other nineteen analyses computable; a winner’s-curse problem leaves the discarded arms’ data on file. A selection at enrolment leaves nothing — the patients who were not enrolled are not recorded, the arms look comparable on every measured covariate because the rule balanced them, and the bias sits in a prognostic quantity nobody wrote down.

The other thing predictability does

There is a second consequence that is not about bias and is easy to miss.

A rule that can be guessed can also be checked. If the assignment sequence is reproducible from the covariates and the rule, then a reader given the covariates and the rule can reconstruct it, which is the difference between a trial whose allocation can be audited and one that has to be trusted.

That cuts both ways and this field takes the second half seriously: the reconstructibility that makes a rule auditable is the same reconstructibility that makes it guessable, and it is also, as the last essay in this field shows, exactly what makes an exact test available. The rule being a known function of known quantities is the problem in this essay and the solution in that one, which is the sharpest example this site has of a property whose sign depends entirely on who is holding it.

What is being claimed here, and what is not

This essay claims the predictability of four allocation rules and the bias an optimal guesser can extract, with the bias’s closed form checked against the count.

What stays out: partial concealment, where the investigator knows some of the history and not all of it; guessers who are wrong about the rule, which reduces the guess rate by an amount that depends on how wrong; and any claim about how often this actually happens in practice, which is a question about people rather than about arithmetic and has a literature this site does not enter. The measurement here is what is available to a guesser, not what guessers do.

One further boundary, because it is the obvious next question and the answer is not in this field. The bias measured here is in the estimate. What it does to the test — whether a trial biased this way rejects a true null more often, and by how much — depends on what the analysis does with the covariates, and that is the subject of the next essay. The two interact: a selection effect that inflates the estimate and an analysis that is conservative for unrelated reasons can produce a test whose size looks reassuring while both of its components are wrong. Measuring either alone would be misleading, which is why they are measured in that order.

What stops growing, and what does not. The worst imbalance on any of the nine factor levels, averaged over cohorts, against the number of patients. A coin's grows like √n forever: 6.32 at 40 patients and 26.28 at 640. A rule that repairs an imbalance as it appears leaves a residue that depends on how many things it is tracking rather than on how many patients arrive, so it goes flat — minimisation runs 3.00 to 3.23 across a sixteenfold range. Blocking the arm totals is the odd one: it improves the margins by a constant factor and leaves the growth untouched, because the only thing it is watching is the total.
Fig. 8 The balance side once more, drawn here as the thing being paid for. Every point on the flat lines was bought with predictability, and the price is on the first figure of this essay at the matching value of p.
20 randomised trials, 45% against 25%. Each line is one trial allocating patients one at a time by a fair coin. The average final share on the better arm is 50.9%, with a standard deviation of 2.9 points across these 20 trials, and 7 of them finished with most of the units on the worse arm. The rule does not deliver a fixed advantage: it delivers one that depends on how the first few patients came out.
Fig. 9 The adaptive field’s picture of twenty allocations under a fair coin, for the contrast this essay turns on: nothing about a coin can be worked out in advance, which is the one thing it is unbeatable at and the whole of what these rules are trading away.

The checks

Two claims are gated in this field’s library.

A deterministic rule is guessable and a coin is not, with minimisation at p = 1 required above 85% and a coin required within two points of half and with a bias inside four standard errors of zero. The coin half is the control: if the guessing machinery produced a bias against a fair coin, every other number in this essay would be an artefact of the simulation rather than a property of a rule.

And the bias is 2δ(2g − 1), checked against the count at p = 1 within the larger of two hundredths and four standard errors, with the softer rule required to deliver materially less bias than the deterministic one. The closed form is the second route: one side counts correct guesses and the other counts what the estimate did, and neither can confirm itself.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlindingConfoundingCovariate-adaptive randomisationCovariate balanceMinimisationPermuted blocksPredictabilityRandomisationSelection biasStudy design