Splitting the units

Not half and half

The same units, the same measurements, the same analysis — and a different variance, decided before anything is measured. When the two arms have different spreads the best split is σ₁ : σ₂, equal allocation costs 2(σ₁²+σ₂²)/(σ₁+σ₂)², and at three to one that is a quarter of the experiment.

Worth reading first: What a p-value does not say · The variance removed before the data.

Every other improvement in experimental design is bought with something. Blocking needs a nuisance factor to block on. A factorial arrangement needs the factors to be settable together. More power needs more units, and the exchange rate is unforgiving.

The split between the arms is free. The same units, the same measurements, the same analysis, the same budget — and a different variance, fixed before anything is measured. It is the cheapest decision in this field and it is the one most often made by reflex.

Every split of 100 units, σ = 1 against 3Each point is one integer split, with its variance computed exactly rather than simulated. The minimum is at 25:75, which is the ratio of the spreads 25:75, and equal allocation costs 25% more variance — the same as throwing away 20 of the 100 units. The shaded band is every split within 5% of the best, and it runs from 17% to 35%: sharp to state, flat to sit on.00.1000.2000.3000.40000.2000.4000.6000.8001share of the 100 units given to the arm with spread 1variance of the estimated difference25:75equal allocation, 25% worsewithin 5% of the best97 integer splits, enumeratedoptimum at σ₁ : σ₂ = 1 : 3
Fig. 1 Every integer split of a hundred units between two arms whose spreads are 1 and 3, with the variance of the estimated difference computed exactly rather than simulated.

Where the rule comes from

The estimate is a difference of two means, and its variance is

σ₁²/n₁ + σ₂²/n₂

with n₁ + n₂ fixed. Minimise it and the answer is

n₁ : n₂ = σ₁ : σ₂.

Sample the noisier arm more. It is one line of calculus and it is not what anybody does: equal allocation is the default in every field that runs comparative experiments, and the reason is not that the rule is unknown but that the quantity it depends on is.

The variance at the optimum is (σ₁ + σ₂)²/N — the sum of the spreads, squared, over the total. That is the number worth carrying, because it makes the comparison with equal allocation immediate: equal allocation gives 2(σ₁² + σ₂²)/N, so the ratio is

2(σ₁² + σ₂²) / (σ₁ + σ₂)²

which depends on nothing but the ratio of the spreads. Not on N, not on the effect, not on the noise level in absolute terms.

Enumerated, because allocations are integers

The calculus answer is n₁ = N·σ₁/(σ₁ + σ₂), which is almost never a whole number, and what a designer allocates is a whole number. So this site computes the whole curve.

At a hundred units with spreads of 1 and 3 the best integer split is 25:75, its variance is 0.1600, and equal allocation gives 0.200025% more. Twenty-five per cent more variance is the same precision as 20 fewer units: a hundred units split evenly is worth eighty split properly, and the twenty were spent on the arm that did not need them.

The whole table, at a hundred units:

σ₂/σ₁ best split what equal allocation costs units effectively wasted
1 50:50 0% 0
1.5 40:60 4% 4
2 33:67 11% 10
3 25:75 25% 20
5 17:83 44% 31
10 9:91 67% 40
What splitting the units evenly costs. The line is 2(σ₁² + σ₂²)/(σ₁ + σ₂)², which is what equal allocation costs relative to the σ₁ : σ₂ split, and the points are the same quantity read off an enumeration of every integer split of 200 units. At a ratio of 2 it is 11%, at 3 it is 25%, and at 10 it is 67%. Below about 1.5 the rule is not worth the trouble of applying, which matters because that is where an estimate of the ratio usually lands.
Fig. 2 The closed form drawn as a line with the enumerated integer optima as points on it. Two routes to the same number, and they agree to machine precision at every ratio.

And a third route, which is the one that matters

Two of those routes are algebra. A variance formula that is minimised correctly is still a claim about an experiment nobody ran, so the third route runs it: two hundred units split fifty against a hundred and fifty, twenty thousand times, with the spread of the estimated difference recorded.

Measured 0.08109. Predicted 0.08000. The agreement is within the simulation’s own error, and it is the check that would catch the failure the algebra cannot — a variance formula that is right about the wrong estimator.

The rule is sharp and the loss is flat

Here is the part that decides whether any of this is worth doing, and it is the part the calculus hides.

The variance is minimised at 25:75, and the loss from missing that optimum is second order: the curve is flat at its bottom, like every well-behaved objective at its minimum. Every split from 17% to 35% on the first arm is within five per cent of the best available variance.

So the rule is sharp to state and forgiving to follow. A designer who knows the spreads differ by about a factor of three and allocates a third against two thirds — an easier number to work with than 25:75 — gives up almost nothing.

That flatness is doing two jobs at once and they point in opposite directions. It is why the rule is practical: a rough estimate of the ratio is enough. It is also why the rule is so often skipped: if being roughly right is nearly as good as being exactly right, being entirely wrong — 50:50 — sounds like it should be nearly as good too, and at a ratio of three it costs a quarter of the experiment.

Every split of 100 units, σ = 1 against 1.5. Each point is one integer split, with its variance computed exactly rather than simulated. The minimum is at 40:60, which is the ratio of the spreads 40:60, and equal allocation costs 4% more variance — the same as throwing away 4 of the 100 units. The shaded band is every split within 5% of the best, and it runs from 30% to 51%: sharp to state, flat to sit on.
Fig. 3 The same enumeration at a spread ratio of 1.5, where the band within 5% of the best runs from 30% to 51% and includes equal allocation. Here the rule genuinely does not matter.
Every split of 100 units, σ = 1 against 8. Each point is one integer split, with its variance computed exactly rather than simulated. The minimum is at 11:89, which is the ratio of the spreads 11:89, and equal allocation costs 60% more variance — the same as throwing away 38 of the 100 units. The shaded band is every split within 5% of the best, and it runs from 6% to 20%: sharp to state, flat to sit on.
Fig. 4 And at a ratio of eight, where equal allocation sits far outside the band and costs more than half the experiment.

The split decides the test’s size, and not in the direction anyone expects

Everything so far has been about the variance of the estimate. The calibration of the test run on it is a separate question, and it turns out to be a much larger one.

The two-sample t test taught first pools the two arms’ variances into one estimate of σ, weighting by degrees of freedom, and then applies it to a standard error weighted by 1/n. When the spreads are equal those two weightings agree. When they are not, they disagree, and which way they disagree is decided by the allocation.

At σ = 1 against 3 in a hundred units, forty thousand experiments with no difference between the arms at all, tested at a nominal 5%:

split share of true nulls rejected
25:75 — the variance-optimal split 0.35%
50:50 5.44%
90:10 — the reverse 37.1%

The last row is not a typo. A design that puts ninety of its hundred units on the quieter arm rejects a true null more than a third of the time, at a nominal 5%, using the standard test with nothing done wrong in the analysis.

The split decides the test's size, σ = 1 against 3. 6,000 experiments at each split with no difference between the arms at all, analysed two ways. The pooled t test rejects 0.00% of true nulls at 10:90 and 37.4% at 90:10, because it builds one estimate of σ from both arms and weights it by degrees of freedom while the standard error weights by 1/n. Welch's test, which keeps the two variances apart, holds 5.17% across the whole range. The variance-optimal split is marked, and it is in the conservative half.
Fig. 5 The same axis as the variance curve: the share of the units on the quieter arm. The pooled test’s size runs from under one per cent to over a third across it, while Welch’s stays flat.

Two inequalities, and neither alone does it

The obvious reading — unequal variances break the t test — is not what the measurement says. With equal spreads, the same machinery at 10:90, 50:50 and 90:10 gives 4.93%, 5.16% and 5.05%: exact at every split. Imbalance alone costs nothing.

And equal allocation with spreads of 1 and 3 gives 5.44%, which is off but not alarming. Unequal variance alone costs little.

It is the interaction that does the damage, and the direction is the part worth carrying. When the larger arm is the noisier one, the pooled estimate of σ is dominated by the noisy arm while the standard error is dominated by the small one — so the test overstates its own uncertainty and becomes conservative. When the larger arm is the quieter one, the same two weightings pull the other way and the test understates its uncertainty by a factor.

So the allocation rule points in the safe direction. The design that minimises the variance also happens to make the standard test conservative rather than liberal, which is a piece of luck rather than a piece of theory — and a piece of luck that costs power, since 0.35% at the nominal 5% is a test that is throwing away most of its rejection region.

The split decides the test's size, σ = 1 against 1. 6,000 experiments at each split with no difference between the arms at all, analysed two ways. The pooled t test rejects 4.50% of true nulls at 10:90 and 5.8% at 90:10, because it builds one estimate of σ from both arms and weights it by degrees of freedom while the standard error weights by 1/n. Welch's test, which keeps the two variances apart, holds 5.11% across the whole range. The variance-optimal split is marked, and it is in the conservative half.
Fig. 6 The same measurement with equal spreads, where both tests are exact at every split. The whole effect above is the interaction of unequal variances with unequal allocation, and neither one alone produces it.

Which is why an unequal design needs Welch

The repair is not to abandon the allocation. It is to run the analysis the design implies.

Welch’s test estimates each arm’s variance separately, forms the standard error as √(s₁²/n₁ + s₂²/n₂), and adjusts the degrees of freedom to match. On the same three designs it gives 4.98%, 5.27% and 5.10% — the level it claims, at every split, with the spreads three to one.

That is a general lesson this field keeps producing in different forms: a design decision and an analysis decision are one decision, and separating them is where the failure lives. An experimenter who allocates unequally and analyses as though they had not is not making a small mistake; at 90:10 they are running a test whose size is seven times its label. And an experimenter who allocates optimally and analyses with a pooled t has done nothing wrong and has quietly given away most of the power the allocation bought.

The site has met the shape before, in randomisation: what a design buys is a reference distribution, and using a different reference distribution than the design produced is the whole of the error.

Coverage of four nominal 95% intervals, n = 20. Computed exactly by summing over all 21 possible counts, not simulated. The Wald interval drops to 18.2% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.
Fig. 7 Why the spreads differ so often without anyone arranging it: a proportion’s variance is a function of its mean, so two arms with different rates have different spreads before anything unusual has happened.

The penalty is bounded, and by a familiar factor

The cost column climbs steeply — 4%, 11%, 25%, 44%, 67% — and it does not climb without limit. As the spread ratio grows, 2(1 + r²)/(1 + r)² tends to 2, so equal allocation never costs more than a doubling of the variance and never wastes more than half the experiment.

At a ratio of ten it is already at 67% of a possible 100%, and the last third takes another factor of ten in the spreads: at a hundred to one the cost is 96% and the waste 49%. So the table’s steepness is misleading about where it is going. The interesting range is entirely below a ratio of about five, and past that the penalty is asymptoting rather than accelerating.

Two is a number this site’s allocation arithmetic reaches from a second direction. The penalty for splitting units evenly across a shared control and k treatment arms is 2(k + 1)/(1 + √k)² − 1, and it tends to the same limit for the same reason: an equal split is the optimum of the wrong problem, and the wrong problem’s optimum is never worse than twice the right one’s when the quantity being minimised is a sum of reciprocals. Whatever an equal allocation is being asked to do, it is at worst a doubling — which is a reassuring ceiling and is exactly why the rule is so easy to skip.

How wide the forgiving band is

The band within 5% of the best variance is quoted at two ratios and it is the root of a quadratic, so it can be stated for any ratio without further enumeration. Setting (1/p + r²/(1 − p))/(1 + r)² = 1.05 and solving gives 16.8p² − 8.8p + 1 = 0 at r = 3, whose roots are 0.1667 and 0.3571, and 6.5625p² − 5.3125p + 1 = 0 at r = 1.5, whose roots are 0.2976 and 0.5115 — the 17-to-35 and 30-to-51 the enumeration reports.

Written as a multiple of the optimum, those bands are 0.67 to 1.43 times p* at a ratio of three, and 0.74 to 1.28 at a ratio of one and a half. The tolerance widens as the spreads diverge: a designer facing a threefold difference may be out by a third in either direction, and one facing a modest difference may be out by only a quarter.

That is the opposite of what caution suggests, and it has a practical form. The setting where the rule matters most is also the setting where it is easiest to follow — a rough guess at a large ratio is enough, and the arithmetic tolerates it. The setting where the rule is fiddly to get right is the one where getting it wrong costs 4%.

The same quadratic says where equal allocation falls out of the band. Putting p = ½ into it gives 0.95r² − 2.1r + 0.95 = 0, whose root is r = 1.58. So equal allocation is within 5% of optimal for every spread ratio up to about 1.6, and outside it thereafter — which is the one number a designer needs to decide whether the whole question arises.

Where unequal spreads come from

The rule is worth applying only where the spreads genuinely differ, so it is worth being concrete about when they do. Four cases, and none of them is unusual.

A treatment that changes the variance as well as the mean. A drug that helps some patients a great deal and others not at all has a larger spread in the treated arm than the control by construction. So does a teaching method whose effect depends on preparation, or a process change that works only on some machines.

A control arm that is a mixture. “Usual care” is not one thing; it is whatever each site does, which is several things, and its spread is the spread of a mixture. The treated arm is a single protocol and is usually tighter.

Counts and proportions, where the variance is a function of the mean and the arms therefore differ in variance whenever the effect is real at all. A proportion near 0.5 has three times the variance of one near 0.1 — a fact this site’s interval essays live on — so a comparison of 10% against 40% has unequal spreads before anyone has done anything unusual.

A comparison against a fixed standard. Where one arm is a reference material, a calibrated instrument or a simulation, its spread can be an order of magnitude below the other’s, and the rule then says something an experimenter would resist: run the reference a handful of times and spend everything else on the thing being measured.

And a control arm that already exists. Registry data, historical controls, a standing cohort: the control observations may be cheaper and noisier and there may be a great many of them, which is the budget version of the same question.

What it costs in power, and what it buys back

An experimenter reading the previous section might reasonably conclude that unequal allocation is more trouble than it is worth. The variance table and the size table point in different directions, and the second is much more dramatic than the first.

They are the same decision seen twice, and the arithmetic reconciles them.

Equal allocation at a spread ratio of three has 25% more variance and a test whose size is 5.4% rather than 5%. The optimal allocation has the smaller variance and, if analysed with a pooled t, a size of 0.35% — which is not a safety margin, it is a test that has thrown away seven eighths of its rejection region and with it most of the power the allocation just bought.

Analysed with Welch, the optimal allocation has the smaller variance and the level it claims. So the ordering is unambiguous once both decisions are made together: optimal allocation with Welch beats equal allocation with anything, and optimal allocation with a pooled t is the worst of the three.

That is worth stating as a rule because the failure mode it names is a silent one. A conservative test does not announce itself. It produces fewer findings, which looks like an honest experiment producing an honest null, and there is no diagnostic anywhere in the output that says the design and the analysis disagreed with each other.

The split decides the test's size, σ = 1 against 2. 6,000 experiments at each split with no difference between the arms at all, analysed two ways. The pooled t test rejects 0.17% of true nulls at 10:90 and 25.8% at 90:10, because it builds one estimate of σ from both arms and weights it by degrees of freedom while the standard error weights by 1/n. Welch's test, which keeps the two variances apart, holds 5.15% across the whole range. The variance-optimal split is marked, and it is in the conservative half.
Fig. 8 At a spread ratio of two, where everything is smaller and the shape is the same. The pooled test is already three times its nominal size at the reverse split, and already conservative at the optimum.

What the rule does not say

Three misreadings, one of which is subtle enough to have its own essay.

It is not about which arm matters more. The allocation is decided by the spreads and by nothing else. An arm that is scientifically the point of the experiment and an arm that is a formality get allocated on the same basis, because what is being estimated is a difference and the difference is as uncertain as its noisier half.

It does not change the estimator. The difference of means is still the difference of means, still unbiased, still analysed the same way. The design is doing all the work, which is what makes it free — and also what makes it invisible in the analysis, since nothing in the output of an unequal design says that a decision was made.

And it depends on quantities nobody has. σ₁ and σ₂ are what the experiment is being run to estimate, or near enough. Substituting estimates of them and proceeding as though the rule were the rule is the site’s standing failure in a new place, and it has an essay of its own: at a spread ratio of 1.5, a pilot-driven allocation is worse than the equal split it replaced.

The refusal

The check behind this essay asserts the optimum at four spread ratios, the closed form for the variance at each, and the simulation against the algebra. None of those is the interesting half.

The interesting half is that with equal spreads the rule must say half. A rule that recommended an unequal split when there is nothing to be unequal about would be a rule fitting noise, and the failure would be invisible — an unequal design looks exactly like a deliberate one.

So the enumeration is run at σ₁ = σ₂ and required to land on 50:50 exactly, and the penalty for equal allocation is required to be exactly 1 — not approximately, since (σ₁ + σ₂)² = 2(σ₁² + σ₂²) is an identity when the two are equal, and a claim that is exact and gated loosely is a claim nobody will notice breaking. It is checked to 10⁻¹².

That refusal is what makes the 25% at a ratio of three mean something. A measurement of how much a rule gains is worth having only from machinery that has been shown to report no gain where there is none.

The size table has a refusal of the same shape, and it is the one that decides how the whole section above is read. If the pooled test’s 37.1% at 90:10 were an artefact of imbalance rather than of the spreads, the essay’s argument would be wrong in a way that no amount of re-running would reveal — the number would still be 37.1%. So the same machinery is given two arms with equal spreads at the same three splits and required to report 5% at each of them, which it does: 4.93%, 5.16%, 5.05%.

One measurement is a finding. The same measurement with its own control beside it is an explanation, and the control costs one extra call.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AllocationAllocation ratioExperimental designNeyman allocationSample sizeStandard deviationStatistical powerVariance reduction