A covariate with no levels
Worth reading first: Balancing what is known in advance · The variance removed before the data.
Every balancing rule in the two fields before this one reads a level. Minimisation counts how many units of each sex, each centre and each disease stage have gone to each arm and sends the arrival wherever the counts are more uneven; stratification blocks within the same cells; the whole apparatus is arithmetic on counts.
Age in years, blood pressure, a baseline score, a tumour diameter and a household income have no levels. What happens to them is universal and almost never recorded: somebody cuts them. Under sixty-five and over. Low, medium and high. Quartiles of the baseline score. That choice is made before the trial starts, is not usually reported, and has a cost that can be computed before a single unit arrives.
What a cut can see
Take a covariate with a standard normal distribution and split it at its median. A rule that balances the two halves is balancing the counts in each half, so the only part of the covariate it can act on is the part that differs between them: the two half-normal means, at ±√(2/π) = ±0.7979.
Everything else is variation inside a category, and no rule that reads only the category can see it. The variance decomposes exactly:
between the two halves: 2/π = 0.6366. Inside them: 1 − 2/π = 0.3634.
So a rule that balanced the two categories perfectly would still leave a covariate imbalance whose standard deviation is √(1 − 2/π) = 0.6028 of what complete randomisation leaves — before it has done anything wrong, and whatever it does next.
That is a floor rather than a performance. It does not depend on the rule, on how deterministically it assigns, on the sample size or on anything else the experimenter controls except the cut itself.
The general cut
The same arithmetic runs for any number of equal-probability categories, and it stays a closed form. The mean of the ith category is k(φ(zᵢ₋₁) − φ(zᵢ)) — because ∫xφ(x)dx = −φ(x), so a category’s mean is a difference of densities — and the between-category variance is the weighted sum of their squares:
k Σ (φ(zᵢ₋₁) − φ(zᵢ))²
which comes to 0.6366 at two categories, 0.7932 at three, 0.8606 at four, 0.8970 at five, 0.9450 at eight and 0.9677 at twelve.
The floors those imply — the square roots of what is left — are 0.6028, 0.4547, 0.3734, 0.3210, 0.2344 and 0.1796. Cutting more finely is worth a great deal at first and then very little: going from two categories to four halves the floor, and going from four to twelve halves it again.
The measurement sits on the closed form
None of that is a statement about any particular rule, so it has to be checked against one. Run four rules on two hundred units and count.
Complete randomisation measures 0.1378 against the exact 2/√200 = 0.1414, which is the anchor the whole picture is quoted against and the one number in it that is a theorem: over a random split of n units into two halves the difference in covariate means has standard deviation 2S/√n whatever the covariate values are, so its standardised version is 2/√n exactly.
Blocking inside two categories leaves 61.5% of that. Minimisation on the same two categories leaves 62.2%. The floor is 60.3%.
Two rules that look quite different — one blocks in advance, one assigns sequentially by a score — land within a point of each other and within two points of a number computed from a density function. They are not two rules; they are one cut, wearing two rules.
At four categories the measured values are 39.9% and the floor is 37.3%; at eight, 26.5% against 23.4%. The rules track the floor at every cut, a couple of points above it, and the gap is the residual randomness a rule with a probability of 0.8 keeps deliberately.
Why the two rules land in the same place
That blocking and minimisation agree to within a point is not a coincidence and is worth taking apart, because the two are usually presented as alternatives with different properties.
Blocking within a category is a pre-randomisation: the units in a category are shuffled and assigned in pairs, so the category’s two counts differ by at most one at the end. Minimisation on the same category is a sequential rule: each arrival goes to whichever arm has fewer of its category so far, with probability 0.8. Their mechanics are different, their predictability is different — one is a permutation nobody can guess and the other is a rule with a known bias — and their arm-size behaviour is different.
What is the same is the information each of them reads. Both see a unit’s category and nothing else, both can drive the category counts to near-equality, and neither can do anything about where inside a category a unit sits. Once the counts are equal, the residual imbalance is the average of the within-category deviations of whichever units happened to land in each arm, and that is a coin flip for both rules.
So the agreement is the floor asserting itself. Two rules that read the same thing and both do it nearly perfectly must land at the same place, and the place is decided by the cut. Anywhere the two differ — predictability, arm sizes, what happens with several factors at once — is a difference in how they get to the floor rather than in where it is.
The choice nobody declares
The uncomfortable part is not that a cut costs something. It is that the cut is a design decision of the same size as the choice of rule, and it is made and reported completely differently.
A trial protocol says “minimisation on centre, sex and age group”. The rule is named, its probability is often given, and the whole apparatus of balancing what is known in advance applies to it. The words “age group” carry a cut that decides more than the choice between minimisation and stratification does — the difference between two categories and four is 22 points of residual imbalance, and the difference between the two rules at a fixed cut is under one point.
And the cut is usually chosen for reasons that have nothing to do with balance: a clinical threshold, a convention, the way the variable was collected. Those are legitimate reasons for a reporting category and they are not reasons for a balancing category, and the two get the same number because nobody separates them.
Where the cut is made, and where it is not
One more property of the closed form is worth having, because it decides how much of this survives contact with a real covariate.
The 2/π is a fact about a normal covariate cut at its median. Neither half of that is essential to the argument and both change the number. Cut a normal covariate at ±1 rather than at 0 and the three categories that result are not equally likely; the between-category share is still a sum of squared density differences and is still available in closed form, and it is smaller than the equal-probability cut into three would give. Cut a skewed covariate — a duration, a count, a concentration — and the categories at the long end are wide and heterogeneous, so the within-category variance is concentrated where the cut can least afford it.
What does not change is the structure: a rule that reads a category is acting on the between-category variance and nothing else, so the residual is always the within-category part, always computable in advance from the covariate’s distribution and the cut, and always a floor rather than a performance. The number 0.6028 is an example. The sentence around it is the finding.
The rule that does not have a floor
One row of those figures has been ignored so far, and it is the field’s second essay. A rule that reads the covariate’s actual value rather than its category leaves 12.9% of a coin’s imbalance at two hundred units — a fifth of what the best category-balancing rule manages, and below the floor at every cut, because the floor is a fact about categories and not about balancing.
That the alternative exists is what makes the cut a cost rather than a constraint. If categorising were forced by the arithmetic, 2/π would be a fact of life; it is forced only by a rule that was written for factors and then applied to something that is not one.
What this does to the sample size
The floor is stated as a ratio of standard deviations, which is the right unit for comparing rules and the wrong one for deciding anything. Converted, it is a statement about how many units a balancing rule is worth.
The variance of the covariate imbalance is what a design can reduce, and reducing a variance by a factor is worth the same as multiplying the sample size by that factor for the part of the analysis that depends on it. A rule at the two-category floor leaves 0.6028² = 36.3% of the coin’s variance, so as far as covariate imbalance goes it is worth about 2.8 times the units. At four categories, 0.3734² = 13.9%, worth 7.2 times. And the rule that reads the number leaves 0.1286² = 1.65% at two hundred units, worth sixty times.
Those multipliers are not the trial’s effective sample size — most of the outcome’s variance has nothing to do with the covariate, and what the imbalance costs the analysis depends on how strongly the covariate drives the outcome. They are the honest way to read the ratios: the difference between a median split and reading the number is not 62% against 13%, it is a factor of three in units against a factor of sixty, and stating it in standard deviations makes it look smaller than it is.
What the rules do balance, and what the count hides
Two other numbers in that first figure are worth reading, because they say what each rule is actually optimising.
The arm sizes: complete randomisation forces them exactly equal, blocking within categories is off by 0.52 units on average, minimisation by 1.11 and the continuous rule by 0.94. A sequential rule cannot force equal arms without giving up its balancing, and each of these three has made a slightly different trade without being asked.
The worst case: over eight hundred trials the largest standardised imbalance complete randomisation produced was 0.439 standard deviations; minimisation’s was 0.288 and the continuous rule’s was 0.094. The worst case is what a single trial actually risks, and it is a more useful number than the standard deviation for a decision that will be taken once.
How fine a cut would have to be
“Cutting more finely is worth a great deal at first and then very little” is a description of the floors — 0.6028, 0.4548, 0.3734, 0.3209, 0.2345, 0.1797 — and differencing them turns it into a rate.
Each doubling of the number of categories multiplies the floor by about 0.62: two to four takes 0.6028 to 0.3734, four to eight takes 0.3734 to 0.2345, and three to twelve — two doublings — takes 0.4548 to 0.1797, which is 0.62 twice over. So the floor falls as roughly k^(−0.68) rather than the 1/k that evenly spaced cuts of a bounded covariate would give.
The shortfall from 1/k is the tails. Equal-probability categories are narrow in the middle of a normal and very wide at the ends, so the outermost bins keep most of their internal variance however many bins there are, and cutting finely spends its effort where the variance already is not.
Read forwards, that rate says what a cut would have to be to make the floor negligible. Getting to 0.1 — a tenth of a coin’s imbalance, which is roughly what the rule that reads the number achieves at two hundred units — takes about twenty-eight categories. Twenty-eight categories of a continuous covariate is not a categorisation; it is the number with rounding, arrived at by a route that spent six essays’ worth of machinery avoiding it.
That is the strongest form of the argument this field is making. The cut is not a coarse approximation that finer cutting repairs; it is a construction whose cost declines so slowly that removing it entirely is easier than reducing it. Two categories give up 40% of the covariate, four give up 37% of what two kept, and no number of categories anybody would write into a protocol gets within sight of reading the number.
The rule’s own coin becomes the binding term
The measured rules sit slightly above their floors at every cut — 62.2% against 60.3% at two categories, 39.9% against 37.3% at four, 26.5% against 23.4% at eight — and the gaps are the deliberate randomness a rule with probability 0.8 keeps.
Those gaps are 1.9, 2.6 and 3.1 points, so they grow as the floor falls; and as a share of the floor they grow much faster, from 3.2% to 7.0% to 13.2%, roughly doubling with each finer cut. The residual randomness is nearly a constant of the rule while the thing it is being added to shrinks.
So finer cutting has diminishing returns twice over, and the second one is not in any of the closed forms above. The floor falls at k^(−0.68), and the excess over the floor does not fall at all — which means that past about eight categories a trial is paying for finer categories and receiving an increasing fraction of the coin the rule was told to keep. At the twenty-eight categories the previous section prices, the rule’s own randomness would be the dominant term and the cut would have stopped mattering.
The lever that exists for that is the rule’s probability rather than the cut, and it is the one the rule a guesser can work out prices from the other side: raising p towards 1 closes the gap and buys predictability with it. The two costs trade against each other, and only one of them is in the protocol.
One number, before the trial starts
The practical form of all of this is a single calculation that takes a minute and is almost never done.
Before a trial begins, the covariate’s distribution is usually known well enough — from a registry, a pilot, or the last trial in the same population. The proposed cut is in the protocol. Those two things give the between-category share, and its complement’s square root is the best any category-balancing rule can do, whatever else the protocol says.
If that number is 0.60, the protocol is proposing to remove 40% of the imbalance in a covariate it has named as important enough to balance on. That is worth knowing while the cut is still a draft, because the alternatives are all cheap: cut more finely, or stop cutting.
What is claimed here, and what is not
This essay claims the cost of categorising a continuous covariate before balancing it: the between-category share in closed form, the floor it imposes on any rule that reads only the category, and the measurement showing that two standard rules sit on that floor at every cut.
What stays out and is named as a decision: unequal-probability cuts, where the arithmetic is the same and the numbers are different, and where a clinically-motivated threshold usually sits; covariates that are not normal, where the between-category share has to be computed for the actual distribution and 2/π is replaced by something else; and the question of whether the analysis should use the category or the number, which is a modelling question and is taken up later in this field. A covariate with a second number beside it is a separate question, and a rule reading more than two arms is another.
The checks, and the refusal that makes them mean something
Three claims are gated in this field’s library. The coin’s own imbalance is required to match 2/√n at three sizes, which anchors everything else. The between-category share at two categories is required to equal 2/π to twelve decimals. And the two category-balancing rules are required to sit at or above the floor their categories impose at four different cuts, while the rule that reads the number is required to be below it at all four — because the floor is a statement about what a category can see, and a rule that does not read categories is not subject to it.
The refusal is the reporting that follows: a rule that balances the counts either side of the median meets its own description exactly, with its category margins even to within one unit, and leaves 62% of a coin’s imbalance in the number the analysis will use. A design described as having balanced the covariate, on that evidence, is refused.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Guessing one arm in three — both name covariate-adaptive randomisation, experimental design, minimisation, monte carlo, randomisation
- A count that has to be estimated — both name blocking, closed form, monte carlo, randomisation
- A proposal that moves more than two units — both name experimental design, minimisation, monte carlo, randomisation
- Balancing towards unequal targets — both name covariate-adaptive randomisation, experimental design, minimisation, monte carlo
- The arcsine that closes it, and the error that was overstated — both name closed form, continuous covariate, experimental design, monte carlo
- The zero that survives a cut — both name closed form, continuous covariate, experimental design, randomisation
Named objects
A flat tag is an object no other essay names yet.
BlockingCategorisationClosed formContinuous covariateCovariate-adaptive randomisationCovariate imbalanceExperimental designMinimisationMonte CarloRandomisationStandardised differenceStratificationVariance decomposition