Balancing what has no levels

A covariate with no levels

Every balancing rule on this site reads a level. Age and blood pressure have none, so somebody cuts them into categories — and a median split can see exactly 2/π of a normal covariate, whatever the rule does with the halves.

Worth reading first: Balancing what is known in advance · The variance removed before the data.

Every balancing rule in the two fields before this one reads a level. Minimisation counts how many units of each sex, each centre and each disease stage have gone to each arm and sends the arrival wherever the counts are more uneven; stratification blocks within the same cells; the whole apparatus is arithmetic on counts.

Age in years, blood pressure, a baseline score, a tumour diameter and a household income have no levels. What happens to them is universal and almost never recorded: somebody cuts them. Under sixty-five and over. Low, medium and high. Quartiles of the baseline score. That choice is made before the trial starts, is not usually reported, and has a cost that can be computed before a single unit arrives.

What a cut can see

Take a covariate with a standard normal distribution and split it at its median. A rule that balances the two halves is balancing the counts in each half, so the only part of the covariate it can act on is the part that differs between them: the two half-normal means, at ±√(2/π) = ±0.7979.

Everything else is variation inside a category, and no rule that reads only the category can see it. The variance decomposes exactly:

between the two halves: 2/π = 0.6366. Inside them: 1 − 2/π = 0.3634.

What a median split can see. A standard normal covariate with its median marked, and the mean of each category as a vertical rule: -0.7979, 0.7979. A rule that balances the categories is balancing those numbers and nothing else, so the part of the covariate it can act on is the variance between them — 0.6366 of the total, which at two categories is exactly 2/π because the two half-normal means are ±√(2/π). The rest, 0.3634, is variation inside the categories that the rule cannot see and does not touch: the assignment within a category is still a coin. Everything the next figure measures is a consequence of this one, and it is available before any unit has arrived.
Fig. 1 A standard normal covariate cut at its median, with the mean of each half marked. A rule that balances the halves is balancing those two numbers, and the spread of everything around them is what it cannot reach.

So a rule that balanced the two categories perfectly would still leave a covariate imbalance whose standard deviation is √(1 − 2/π) = 0.6028 of what complete randomisation leaves — before it has done anything wrong, and whatever it does next.

That is a floor rather than a performance. It does not depend on the rule, on how deterministically it assigns, on the sample size or on anything else the experimenter controls except the cut itself.

The general cut

The same arithmetic runs for any number of equal-probability categories, and it stays a closed form. The mean of the ith category is k(φ(zᵢ₋₁) − φ(zᵢ)) — because ∫xφ(x)dx = −φ(x), so a category’s mean is a difference of densities — and the between-category variance is the weighted sum of their squares:

k Σ (φ(zᵢ₋₁) − φ(zᵢ))²

which comes to 0.6366 at two categories, 0.7932 at three, 0.8606 at four, 0.8970 at five, 0.9450 at eight and 0.9677 at twelve.

What a cut into 4 can see. A standard normal covariate with its 4 equal-probability categories marked, and the mean of each category as a vertical rule: -1.2711, -0.3247, 0.3247, 1.2711. A rule that balances the categories is balancing those numbers and nothing else, so the part of the covariate it can act on is the variance between them — 0.8606 of the total, which at two categories is exactly 2/π because the two half-normal means are ±√(2/π). The rest, 0.1394, is variation inside the categories that the rule cannot see and does not touch: the assignment within a category is still a coin. Everything the next figure measures is a consequence of this one, and it is available before any unit has arrived.
Fig. 2 The same covariate cut into quartiles. Four category means rather than two, further apart at the ends and closer in the middle, and 86.1% of the variance now visible to a rule that balances them.

The floors those imply — the square roots of what is left — are 0.6028, 0.4547, 0.3734, 0.3210, 0.2344 and 0.1796. Cutting more finely is worth a great deal at first and then very little: going from two categories to four halves the floor, and going from four to twelve halves it again.

The floor, as a function of how finely the covariate is cut. The curve is a closed form with no simulation in it: k equal-probability categories of a normal covariate see k Σ (φ(zᵢ₋₁) − φ(zᵢ))² of its variance — 63.7% at 2, 79.3% at 3, 86.1% at 4, 89.7% at 5, 94.5% at 8, 96.8% at 12 — so a rule that balanced them perfectly would leave the square root of what is left, 0.603, 0.455, 0.373, 0.321, 0.234, 0.180 of a coin's. The marks are minimisation actually run at each of those cuts, 400 trials each, and they sit on the curve. The lower marks are the rule that reads the number, which is below the floor everywhere because the floor is a fact about categories rather than about balancing. Two categories is the common choice and the worst one: it throws away 36.3% of the covariate before the rule has done anything.
Fig. 3 The floor as a function of how finely the covariate is cut, with minimisation actually run at each of those cuts. The curve is a closed form and the marks on it are a simulation.

The measurement sits on the closed form

None of that is a statement about any particular rule, so it has to be checked against one. Run four rules on two hundred units and count.

Four rules, and the floor two of them cannot pass. The standard deviation of the covariate imbalance under each rule at n = 200, over 500 trials, as a share of a coin's — which is exactly 2/√n = 0.1414 and is the one number here that needs no simulation. Blocking inside 2 categories and minimising on the same 2 categories are the same rule to within their noise, 61.7% and 61.2%, and the marked line is why: a 2-category split can see 63.7% of the covariate's variance, so a rule that balanced its categories perfectly would still leave 60.3% of a coin's imbalance. Reading the number instead leaves 13.0%, and the worst imbalance it produced in 500 trials was 0.094 standard deviations against the coin's 0.439.
Fig. 4 Four rules at two hundred units, eight hundred trials, each shown as the standard deviation of the covariate imbalance it leaves — as a share of complete randomisation’s, which is exactly 2/√n and needs no simulation.

Complete randomisation measures 0.1378 against the exact 2/√200 = 0.1414, which is the anchor the whole picture is quoted against and the one number in it that is a theorem: over a random split of n units into two halves the difference in covariate means has standard deviation 2S/√n whatever the covariate values are, so its standardised version is 2/√n exactly.

Blocking inside two categories leaves 61.5% of that. Minimisation on the same two categories leaves 62.2%. The floor is 60.3%.

Two rules that look quite different — one blocks in advance, one assigns sequentially by a score — land within a point of each other and within two points of a number computed from a density function. They are not two rules; they are one cut, wearing two rules.

Four rules, and the floor two of them cannot pass. The standard deviation of the covariate imbalance under each rule at n = 200, over 500 trials, as a share of a coin's — which is exactly 2/√n = 0.1414 and is the one number here that needs no simulation. Blocking inside 4 categories and minimising on the same 4 categories are the same rule to within their noise, 37.6% and 41.0%, and the marked line is why: a 4-category split can see 86.1% of the covariate's variance, so a rule that balanced its categories perfectly would still leave 37.3% of a coin's imbalance. Reading the number instead leaves 12.3%, and the worst imbalance it produced in 500 trials was 0.062 standard deviations against the coin's 0.439.
Fig. 5 The same four rules with the covariate cut into quartiles instead. Both category-balancing rules drop to about 40% of a coin’s, which is the new floor, and the rule that reads the number does not move — because the cut was never anything to it.

At four categories the measured values are 39.9% and the floor is 37.3%; at eight, 26.5% against 23.4%. The rules track the floor at every cut, a couple of points above it, and the gap is the residual randomness a rule with a probability of 0.8 keeps deliberately.

Why the two rules land in the same place

That blocking and minimisation agree to within a point is not a coincidence and is worth taking apart, because the two are usually presented as alternatives with different properties.

Blocking within a category is a pre-randomisation: the units in a category are shuffled and assigned in pairs, so the category’s two counts differ by at most one at the end. Minimisation on the same category is a sequential rule: each arrival goes to whichever arm has fewer of its category so far, with probability 0.8. Their mechanics are different, their predictability is different — one is a permutation nobody can guess and the other is a rule with a known bias — and their arm-size behaviour is different.

What is the same is the information each of them reads. Both see a unit’s category and nothing else, both can drive the category counts to near-equality, and neither can do anything about where inside a category a unit sits. Once the counts are equal, the residual imbalance is the average of the within-category deviations of whichever units happened to land in each arm, and that is a coin flip for both rules.

So the agreement is the floor asserting itself. Two rules that read the same thing and both do it nearly perfectly must land at the same place, and the place is decided by the cut. Anywhere the two differ — predictability, arm sizes, what happens with several factors at once — is a difference in how they get to the floor rather than in where it is.

The choice nobody declares

The uncomfortable part is not that a cut costs something. It is that the cut is a design decision of the same size as the choice of rule, and it is made and reported completely differently.

A trial protocol says “minimisation on centre, sex and age group”. The rule is named, its probability is often given, and the whole apparatus of balancing what is known in advance applies to it. The words “age group” carry a cut that decides more than the choice between minimisation and stratification does — the difference between two categories and four is 22 points of residual imbalance, and the difference between the two rules at a fixed cut is under one point.

What a cut into 3 can seeA standard normal covariate with its 3 equal-probability categories marked, and the mean of each category as a vertical rule: -1.0908, -0.0000, 1.0908. A rule that balances the categories is balancing those numbers and nothing else, so the part of the covariate it can act on is the variance between them — 0.7932 of the total, which at two categories is exactly 2/π because the two half-normal means are ±√(2/π). The rest, 0.2068, is variation inside the categories that the rule cannot see and does not touch: the assignment within a category is still a coin. Everything the next figure measures is a consequence of this one, and it is available before any unit has arrived.00.2000.400-202the covariate, in standard deviationsdensitymean 1.091between the categories: 0.7932 of the variance · inside them: 0.2068closed form: category means k(φ(zᵢ₋₁) − φ(zᵢ))3 cuts see 79.32%, leaving 45.5%
Fig. 6 Drag the number of categories. The category means move apart and multiply, the share of the covariate they can see climbs, and none of this is a property of any rule.

And the cut is usually chosen for reasons that have nothing to do with balance: a clinical threshold, a convention, the way the variable was collected. Those are legitimate reasons for a reporting category and they are not reasons for a balancing category, and the two get the same number because nobody separates them.

Where the cut is made, and where it is not

One more property of the closed form is worth having, because it decides how much of this survives contact with a real covariate.

The 2/π is a fact about a normal covariate cut at its median. Neither half of that is essential to the argument and both change the number. Cut a normal covariate at ±1 rather than at 0 and the three categories that result are not equally likely; the between-category share is still a sum of squared density differences and is still available in closed form, and it is smaller than the equal-probability cut into three would give. Cut a skewed covariate — a duration, a count, a concentration — and the categories at the long end are wide and heterogeneous, so the within-category variance is concentrated where the cut can least afford it.

What does not change is the structure: a rule that reads a category is acting on the between-category variance and nothing else, so the residual is always the within-category part, always computable in advance from the covariate’s distribution and the cut, and always a floor rather than a performance. The number 0.6028 is an example. The sentence around it is the finding.

The rule that does not have a floor

One row of those figures has been ignored so far, and it is the field’s second essay. A rule that reads the covariate’s actual value rather than its category leaves 12.9% of a coin’s imbalance at two hundred units — a fifth of what the best category-balancing rule manages, and below the floor at every cut, because the floor is a fact about categories and not about balancing.

Two rates, not a factor. The standard deviation of the covariate imbalance under three rules, at five trial sizes, 260 trials each, on log axes. The upper line is a coin: its slope is -0.489, against a closed form of exactly −½. The middle line is minimisation on a median split; its slope is -0.519 — the same rate — because inside a category the assignment is still a coin, and what it buys is the constant, 0.654 of a coin's at n = 200. The lower line is the rule that reads x and maximises the information about the treatment effect: slope -0.987, nearly twice as steep. Its advantage is therefore not a number that can be quoted — it is 0.258 of a coin's at n = 50 and 0.065 at n = 800, and it keeps going.
Fig. 7 The three rules against the size of the trial, on log axes. Two of them have one slope and one has another, which is the next essay and is a different kind of statement from a ratio.

That the alternative exists is what makes the cut a cost rather than a constraint. If categorising were forced by the arithmetic, 2/π would be a fact of life; it is forced only by a rule that was written for factors and then applied to something that is not one.

What this does to the sample size

The floor is stated as a ratio of standard deviations, which is the right unit for comparing rules and the wrong one for deciding anything. Converted, it is a statement about how many units a balancing rule is worth.

The variance of the covariate imbalance is what a design can reduce, and reducing a variance by a factor is worth the same as multiplying the sample size by that factor for the part of the analysis that depends on it. A rule at the two-category floor leaves 0.6028² = 36.3% of the coin’s variance, so as far as covariate imbalance goes it is worth about 2.8 times the units. At four categories, 0.3734² = 13.9%, worth 7.2 times. And the rule that reads the number leaves 0.1286² = 1.65% at two hundred units, worth sixty times.

Those multipliers are not the trial’s effective sample size — most of the outcome’s variance has nothing to do with the covariate, and what the imbalance costs the analysis depends on how strongly the covariate drives the outcome. They are the honest way to read the ratios: the difference between a median split and reading the number is not 62% against 13%, it is a factor of three in units against a factor of sixty, and stating it in standard deviations makes it look smaller than it is.

What the rules do balance, and what the count hides

Two other numbers in that first figure are worth reading, because they say what each rule is actually optimising.

The arm sizes: complete randomisation forces them exactly equal, blocking within categories is off by 0.52 units on average, minimisation by 1.11 and the continuous rule by 0.94. A sequential rule cannot force equal arms without giving up its balancing, and each of these three has made a slightly different trade without being asked.

The worst case: over eight hundred trials the largest standardised imbalance complete randomisation produced was 0.439 standard deviations; minimisation’s was 0.288 and the continuous rule’s was 0.094. The worst case is what a single trial actually risks, and it is a more useful number than the standard deviation for a decision that will be taken once.

Four rules, and the floor two of them cannot pass. The standard deviation of the covariate imbalance under each rule at n = 50, over 500 trials, as a share of a coin's — which is exactly 2/√n = 0.2828 and is the one number here that needs no simulation. Blocking inside 2 categories and minimising on the same 2 categories are the same rule to within their noise, 57.3% and 62.6%, and the marked line is why: a 2-category split can see 63.7% of the covariate's variance, so a rule that balanced its categories perfectly would still leave 60.3% of a coin's imbalance. Reading the number instead leaves 25.1%, and the worst imbalance it produced in 500 trials was 0.291 standard deviations against the coin's 0.851.
Fig. 8 Fifty units rather than two hundred. Every rule’s imbalance is larger, the floor is the same fraction of a larger number, and the category rules are relatively no better than they were.

How fine a cut would have to be

“Cutting more finely is worth a great deal at first and then very little” is a description of the floors — 0.6028, 0.4548, 0.3734, 0.3209, 0.2345, 0.1797 — and differencing them turns it into a rate.

Each doubling of the number of categories multiplies the floor by about 0.62: two to four takes 0.6028 to 0.3734, four to eight takes 0.3734 to 0.2345, and three to twelve — two doublings — takes 0.4548 to 0.1797, which is 0.62 twice over. So the floor falls as roughly k^(−0.68) rather than the 1/k that evenly spaced cuts of a bounded covariate would give.

The shortfall from 1/k is the tails. Equal-probability categories are narrow in the middle of a normal and very wide at the ends, so the outermost bins keep most of their internal variance however many bins there are, and cutting finely spends its effort where the variance already is not.

Read forwards, that rate says what a cut would have to be to make the floor negligible. Getting to 0.1 — a tenth of a coin’s imbalance, which is roughly what the rule that reads the number achieves at two hundred units — takes about twenty-eight categories. Twenty-eight categories of a continuous covariate is not a categorisation; it is the number with rounding, arrived at by a route that spent six essays’ worth of machinery avoiding it.

That is the strongest form of the argument this field is making. The cut is not a coarse approximation that finer cutting repairs; it is a construction whose cost declines so slowly that removing it entirely is easier than reducing it. Two categories give up 40% of the covariate, four give up 37% of what two kept, and no number of categories anybody would write into a protocol gets within sight of reading the number.

The rule’s own coin becomes the binding term

The measured rules sit slightly above their floors at every cut — 62.2% against 60.3% at two categories, 39.9% against 37.3% at four, 26.5% against 23.4% at eight — and the gaps are the deliberate randomness a rule with probability 0.8 keeps.

Those gaps are 1.9, 2.6 and 3.1 points, so they grow as the floor falls; and as a share of the floor they grow much faster, from 3.2% to 7.0% to 13.2%, roughly doubling with each finer cut. The residual randomness is nearly a constant of the rule while the thing it is being added to shrinks.

So finer cutting has diminishing returns twice over, and the second one is not in any of the closed forms above. The floor falls at k^(−0.68), and the excess over the floor does not fall at all — which means that past about eight categories a trial is paying for finer categories and receiving an increasing fraction of the coin the rule was told to keep. At the twenty-eight categories the previous section prices, the rule’s own randomness would be the dominant term and the cut would have stopped mattering.

The lever that exists for that is the rule’s probability rather than the cut, and it is the one the rule a guesser can work out prices from the other side: raising p towards 1 closes the gap and buys predictability with it. The two costs trade against each other, and only one of them is in the protocol.

One number, before the trial starts

The practical form of all of this is a single calculation that takes a minute and is almost never done.

Before a trial begins, the covariate’s distribution is usually known well enough — from a registry, a pilot, or the last trial in the same population. The proposed cut is in the protocol. Those two things give the between-category share, and its complement’s square root is the best any category-balancing rule can do, whatever else the protocol says.

If that number is 0.60, the protocol is proposing to remove 40% of the imbalance in a covariate it has named as important enough to balance on. That is worth knowing while the cut is still a draft, because the alternatives are all cheap: cut more finely, or stop cutting.

What is claimed here, and what is not

This essay claims the cost of categorising a continuous covariate before balancing it: the between-category share in closed form, the floor it imposes on any rule that reads only the category, and the measurement showing that two standard rules sit on that floor at every cut.

What stays out and is named as a decision: unequal-probability cuts, where the arithmetic is the same and the numbers are different, and where a clinically-motivated threshold usually sits; covariates that are not normal, where the between-category share has to be computed for the actual distribution and 2/π is replaced by something else; and the question of whether the analysis should use the category or the number, which is a modelling question and is taken up later in this field. A covariate with a second number beside it is a separate question, and a rule reading more than two arms is another.

The checks, and the refusal that makes them mean something

Three claims are gated in this field’s library. The coin’s own imbalance is required to match 2/√n at three sizes, which anchors everything else. The between-category share at two categories is required to equal 2/π to twelve decimals. And the two category-balancing rules are required to sit at or above the floor their categories impose at four different cuts, while the rule that reads the number is required to be below it at all four — because the floor is a statement about what a category can see, and a rule that does not read categories is not subject to it.

The refusal is the reporting that follows: a rule that balances the counts either side of the median meets its own description exactly, with its category margins even to within one unit, and leaves 62% of a coin’s imbalance in the number the analysis will use. A design described as having balanced the covariate, on that evidence, is refused.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlockingCategorisationClosed formContinuous covariateCovariate-adaptive randomisationCovariate imbalanceExperimental designMinimisationMonte CarloRandomisationStandardised differenceStratificationVariance decomposition