Allocating on a guess
Worth reading first: Not half and half.
The three rules in this field are all functions of quantities nobody has. The optimal split needs σ₁ and σ₂. The budget rule needs those and the marginal costs. The square-root rule needs a spread for every arm. In each case the quantity is what the experiment is being run to learn something about, and the usual answer is to run a small pilot first, estimate it, and proceed.
This site has met that pattern before, in a field that looks nothing like this one. The plug-in interval estimates a population spread from eight groups and then uses it as though it were known; counted over sixteen thousand intervals, it covers 78.8% against a claimed 95%. The failure is not that the estimate is bad. It is that a rule which is correct given a quantity is a different rule when it is fed an estimate of that quantity, and the difference is measurable.
Here it is measurable in the currency the design is about: variance.
What is being compared
Three designs, all spending a hundred units.
Equal allocation: fifty and fifty. It needs nothing, estimates nothing, and its variance has a closed form — 2σ₁²/N + 2σ₂²/N — so it is computed exactly rather than simulated, which matters because every number below is a ratio to it.
The oracle: the σ₁ : σ₂ split, computed from the true spreads by a designer who somehow knows them. Nobody can run this. It is the ceiling.
The two-stage design: m units per arm first, an estimate of each spread from them, and the remaining 100 − 2m split by the estimated ratio. The pilot’s units count against it, because they were spent.
Where the rule costs rather than pays
Start with the case that has nothing in it. Both arms have the same spread, so the oracle is equal allocation and the ceiling is exactly 1.000: there is no gain available.
| pilot per arm | variance ÷ equal allocation |
|---|---|
| 3 | 1.204 |
| 4 | 1.114 |
| 8 | 1.050 |
| 25 | 1.018 |
| 35 | 0.999 |
The rule splits unequally anyway, because the ratio of two sample standard deviations on a handful of observations is rarely 1, and every departure from equal is pure loss. A four-unit pilot makes the experiment 11% worse than doing nothing at all.
That is the shape of the thing. A rule that cannot help can still hurt, and it hurts most when the information it is fed is worst.
The excess falls like the degrees of freedom
The three readings in the table are enough to identify what the excess is a function of, and it is not the pilot’s size directly.
The excess over equal allocation is 0.204, 0.114 and 0.050 at pilots of three, four and eight. Multiplying each by — the degrees of freedom in one arm’s spread estimate — gives 0.408, 0.342 and 0.350.
The last two agree to within two per cent, so over the range where the estimate is not degenerate
At a pilot of four that predicts 1.117 against 1.114 measured, and at eight it predicts 1.050 against 1.050. At three it predicts 1.175 against 1.204 — seventeen per cent of the excess unaccounted for, which is what a variance estimate on two degrees of freedom does: its distribution has a tail heavy enough that the allocation it produces is occasionally extreme, and an extreme allocation costs more than a symmetric argument allows for.
So the smallest pilot is not merely on the curve; it is off it, in the direction that makes it worse.
What the excess costs, in units
A variance ratio converts into the currency the whole design is denominated in, and the conversion makes the size of the failure legible.
A design whose variance is 1.204 times equal allocation’s is the design a hundred units would have given at 83 units. So a pilot of three per arm on a problem with nothing in it has thrown away 17 units of a hundred. At four per arm it is 10, and at eight per arm 5.
The pilots themselves cost six, eight and sixteen units — and they are not lost, since those units are still measured and still enter the estimate.
So the smallest pilot destroys nearly three times its own size, and the largest destroys less than a third of its own size. The loss is not the units spent on the pilot; it is the damage a noisy ratio does to the allocation of the eighty or ninety units that come after, and that damage is largest exactly when the pilot is smallest.
Which reverses the intuition a two-stage design is usually chosen on. The instinct is to keep the pilot small so that most of the budget is allocated well. The arithmetic says a small pilot allocates the rest of the budget badly, and that the units saved are worth a fraction of what the bad allocation costs.
And the crossing point is further out than expected
The interesting question is not whether the rule can lose but where it starts winning, and the answer is uncomfortable.
At a spread ratio of 1.5, where the oracle would gain 3.8%, the two-stage design is at 1.150 with a three-unit pilot, 1.074 at four, 1.003 at eight, and 0.980 at thirty-five. Spend a third of the experiment establishing the ratio and it recovers half of a gain that was never large.
At a ratio of 2 — the oracle gains 10% — the design first beats equal allocation at a four-unit pilot and reaches 0.925 at twelve, capturing three quarters of what was available.
At a ratio of 3 the oracle gains 20%, the best pilot is eight units, and the design reaches 0.824 against a ceiling of 0.800: 88% of the available gain.
So the practical rule is not “estimate the spreads and allocate”. It is: allocate unequally when the arms are expected to differ by a factor of two or more, and split evenly otherwise. Between 1 and 1.5 the rule is a liability whatever the pilot; past 2 it works and the pilot can be small.
Why there is a best pilot size
The curve at a ratio of three turns: 0.917 at three units per arm, 0.824 at eight, 0.832 at twelve, 0.906 at thirty-five. Both ends are bad and the reasons are different.
Small pilots estimate badly. The ratio comes from two sample standard deviations on m − 1 degrees of freedom each, and at m = 3 that is a wild quantity.
Large pilots have nothing left to allocate, and the mechanism is worth stating exactly because it is structural rather than statistical. The pilot’s own units are split equally — they have to be, since there is no estimate yet when they are spent. So the design’s overall split is m/N equal plus the rest at the estimated ratio, and as m grows the overall split is dragged back towards 50:50 whatever the estimate says.
Measured: at a spread ratio of 3, where the best overall share for the first arm is 25%, the two-stage design’s average share is 29.5% with a pilot of eight and 33.2% with a pilot of sixteen. The larger pilot knows the ratio better — the spread of the chosen share falls from 6.4 percentage points to 3.4 — and it aims at a worse target.
How noisy the decision is, from first principles
The spread of the chosen share is the quantity everything above turns on, and it can be predicted rather than only measured — which makes it a two-routes claim of the kind this site prefers.
A sample standard deviation from m observations has a relative standard error of about 1/√(2(m−1)). At m = 8 that is 0.267 — each spread estimate is out by a quarter, typically. The ratio of two of them compounds: √2 × 0.267 = 0.378, so a ratio whose true value is ⅓ arrives with a standard deviation of about 0.126.
The allocation share is r/(1 + r), whose derivative at r = ⅓ is 0.5625, so the chosen share should have a standard deviation of about 7.1 percentage points.
Measured across six thousand experiments: 6.4. The prediction is a first-order expansion of a nonlinear function of two chi variables, so agreeing to within a point is as close as it should get, and it settles where the noise comes from — not from the design, not from the simulation, but from the fourteen degrees of freedom the pilot has.
That formula is also the honest answer to “how big should the pilot be to decide this properly”. To know the ratio to within 10%, 1/√(2(m−1)) × √2 must be about 0.1, which needs m ≈ 100 per arm. On a hundred-unit experiment there is no such pilot: the design cannot afford to know its own allocation parameter, at any split of its budget, which is why the curve turns where it does.
The bias nobody looks for
There is a second, smaller effect in those figures and it is worth naming because it survives any pilot size.
The rule allocates by s₁/(s₁ + s₂), which is a nonlinear function of two random variables, and a nonlinear function of an unbiased estimate is not an unbiased estimate of the function. The chosen share comes out above the target on average — 29.5% against 25.0% at a pilot of eight — and it would do so even if the pilot’s own units were not dragging the average towards a half.
The effect is small here, and it is the same mechanism that makes an estimated population spread of zero so consequential in a different field: what a plug-in rule does with an estimate is decided by the whole distribution of that estimate, not by its centre.
The consolation is the one this field keeps returning to. The loss surface is flat near its optimum, so a share that is 4.5 points off target costs a fraction of a per cent, and the whole of the bias is worth less than the noise. It is worth knowing about because it is the reason a two-stage design cannot be tuned to reach the oracle: the target it aims at is not where the oracle is.
What a pilot is actually for
None of this says pilots are a bad idea. It says that allocation is a poor thing to spend one on, and the distinction matters because pilots are usually run for other reasons entirely — whether the protocol is followable, whether recruitment works, whether the measurement instrument does what it should, whether anything has been forgotten.
Those purposes are qualitative and a handful of units answers them. The statistical uses are the ones with an arithmetic behind them, and they differ sharply in how well they are served:
Estimating σ for a sample-size calculation is the standard use and it is on firmer ground than allocation, because n depends on σ² through a formula whose input is one spread rather than a ratio of two. It has its own failure, which is what an interim look at σ does to the error rate, and it is a smaller one.
Estimating the effect size is the use with no defence at all. An effect estimated from a pilot is the winner’s curse waiting to happen if the pilot’s result decides whether the trial proceeds, and it is far too noisy to power anything even when it does not.
And estimating the allocation ratio is this essay: a rule that needs a hundred units per arm to be worth applying, on experiments that have a hundred units in total.
The ordering is worth remembering, because a single pilot is often asked to do all three and the three have quite different prospects.
The same shape, four fields apart
It is worth setting the three plug-in failures on this site beside each other, because they are the same failure and they arrive in fields that share no machinery.
The population spread. Partial pooling needs τ, estimates it from the group means, and then treats it as known: the interval covers 78.8% against a claimed 95%, and on 32.6% of eight-group datasets the estimate is exactly zero, which is an instruction to pool completely.
The allocation ratio. This essay. The rule needs σ₁/σ₂, estimates it from a pilot, and treats it as known: at equal spreads the result is 11% worse than doing nothing.
And the sample size. An interim look that re-estimates σ and recomputes n from it — the mildest of the three, because a variance enters that formula directly rather than as a ratio, and the one with a closed form for exactly what it costs.
Three fields, three quantities, one structure: a formula derived under the assumption that a nuisance parameter is known, applied with an estimate substituted, and evaluated as though the substitution were free. What varies between them is only how much the substitution costs, and that is a measurement rather than a matter of opinion in every case.
The repair also has the same shape in all three, and only one of the three fields on this site has taken it: integrate over the uncertainty rather than substituting a point. For the allocation problem that would mean averaging the design’s variance over the posterior for the ratio and choosing the split that minimises the average — which is a well-posed calculation, is not what anybody does, and is not attempted here.
What to do with all this
Four things, in order of how much they are worth.
Decide unequal allocation from what is known in advance, not from a pilot. Prior data, a published standard deviation, the structure of the outcome — a proportion near 0.1 against one near 0.5 has a spread ratio of about 1.5 by arithmetic, before any data. That kind of knowledge costs no units and does not have the noise this essay is about.
If the expected ratio is under about 1.5, split evenly. The measurement is unambiguous: the rule cannot recover the small gain available and reliably costs more than it.
If a pilot is being run anyway, use it, and keep it small. Eight per arm out of a hundred captured 88% of what was available at a ratio of three; twenty-five per arm captured less.
Report the split and the reason. An unequal design is invisible in the analysis — the estimator is the same, the standard error is the same arithmetic — so a reader cannot tell a deliberate allocation from an accident of recruitment, and the two imply different things about what the analysis should be. The pooled t test’s size depends on which one happened.
And do not spend units to learn the ratio. That is the finding a designer would least expect and it falls straight out of the curve: the pilot is not an investment that pays back, it is a cost with a diminishing benefit and a structural penalty attached, and the best of these designs is the one that gets the pilot out of the way.
Where the check refuses
Two refusals, and the first one had to be written after the measurement contradicted the essay.
The claim this machinery was built to make was that small pilots are worse than equal allocation. At a spread ratio of three that is false — a four-unit pilot already reaches 0.870 — and the check said so on its first run. What is true is the narrower and sharper statement the essay now makes: the rule costs where there is nothing to gain, at every pilot size, and the crossing point in the spread ratio rather than in the pilot size is the number worth quoting.
The second refusal is the oracle. If knowing the true spreads did not beat equal allocation on this machinery, everything above would be measuring the simulation rather than the pilot. The check requires the oracle to win by more than the pilot ever recovers, and requires the two-stage design never to reach it — because a two-stage design that matched the oracle would mean the estimate was costing nothing, which is the thing this whole essay is about not being true.
There is a third check that exists only because a figure caught something the prose had not. Comparing a simulated two-stage variance against a simulated equal-allocation variance put Monte Carlo noise in both halves of every ratio, and at equal spreads — where the honest answer is exactly 1 — the sweep printed 0.988 at one pilot size and 1.028 at another with nothing but the seed between them. Equal allocation has a closed form; there was never a reason to simulate it. The simulated value is kept and required to agree with the closed form, which is this site’s rule rather than a precaution, and the ratios in this essay now have one source of noise instead of two.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Balancing towards unequal targets — both name allocation, allocation ratio, experimental design
- Dropping the losers — both name experimental design, sample size, two-stage design
- How many subjects — both name allocation, sample size, standard deviation
- Stopping when it is precise enough — both name experimental design, sample size, two-stage design
- The design that stops guessing — both name experimental design, pilot study, two-stage design
- The variance removed before the data — both name allocation, experimental design, variance reduction
Named objects
A flat tag is an object no other essay names yet.
AllocationAllocation ratioExperimental designPilot studyPlug in estimateSample sizeStandard deviationTwo-stage designVariance reduction