The cost of a unit
Worth reading first: Not half and half.
The allocation rule in the previous essay minimises the variance subject to a fixed number of units. Almost no experiment is constrained that way. What runs out is money, or time, or a machine, and the arms rarely consume it at the same rate: a laboratory assay against a field measurement, a specialist procedure against usual care, a simulation against a physical trial.
Change the constraint and the rule changes with it, in a way that can reverse the answer.
The rule, and where the square root comes from
Minimise σ₁²/n₁ + σ₂²/n₂ subject to c₁n₁ + c₂n₂ = C, and the answer is
n₁ : n₂ = (σ₁/√c₁) : (σ₂/√c₂).
The spread enters linearly and the cost enters as a square root, and the asymmetry is the whole content of the rule. Doubling an arm’s noise doubles the units it deserves. Doubling its price cuts them by only √2 — because the variance of a mean falls as 1/n while the money falls as n, so an expensive arm is still worth buying, just less of it.
That single square root is the difference between a rule anyone would guess and a rule worth deriving. The guess — spend proportionally less on the expensive arm, in proportion to its price — over-corrects by a factor of √c and gives away real precision.
Where the two rules disagree
Take the case in the figure: spreads of 1 and 3, costs of 1 and 20. The unit rule says buy in the ratio 1:3 — three times as many of the noisy arm. The budget rule says (1/1) : (3/√20) = 1.49 : 1 — half as many of the noisy arm as of the cheap one.
The two rules point in opposite directions, and both are correct about their own constraint. Which one applies is a fact about the experiment rather than about the statistics.
Enumerated over every affordable pair of integers at a budget of four thousand:
| rule followed | units bought | variance of the estimate |
|---|---|---|
| the budget rule | 280 and 186 | 0.05196 |
| equal units | 190 and 190 | 0.05263 |
| the unit rule, made affordable | 66 and 197 | 0.06084 |
Following the σ-ratio rule when money is the constraint costs 17% more variance than the σ/√c rule, for exactly the same spend. Equal allocation, which both rules say is wrong, happens to sit between them and costs 1.3% — a coincidence of these numbers rather than a general fact, and a good example of why the arithmetic is worth doing rather than reasoning about.
The direction is not always the one the example suggests
The case above has the noisy arm also being the expensive one, which is what makes the reversal dramatic. Turn it around — the cheap arm is the noisy one — and the two rules agree in direction and differ in size.
With spreads of 1 and 3 and costs of 5 and 1, the budget rule buys 342 and 2,290: enormously more of the noisy arm, because it is noisy and cheap and both terms push the same way. The unit rule’s affordable version, 500 and 1,500, costs 17% more variance.
So the general statement is not “spend less on the expensive arm”. It is that each arm’s claim on the budget is σᵢ/√cᵢ, and that quantity has to be computed rather than reasoned about, because two terms that can point in opposite directions do not have an intuition attached.
What the penalty is in the currency that binds
A variance penalty is an abstraction and a budget is not, so it is worth converting.
The variance of the estimate falls as 1/C — buy twice as much of everything and the variance halves, which is true here because both arms scale together and nothing else in the design changes. So a 17% variance penalty is exactly a 17% larger budget. Reaching the budget rule’s variance while following the unit rule takes 4,684 rather than 4,000: six hundred and eighty-four spent on nothing.
That conversion is the argument for doing the arithmetic at all. Nobody is moved by the difference between 0.05196 and 0.06084. A line in a proposal saying that the last 684 of the 4,684 buys no precision at all is a different kind of statement, and it is the same statement.
It also gives the honest test of whether the rule is worth applying to a particular experiment. Work out σ₁/√c₁ and σ₂/√c₂, take their ratio, and read the penalty for equal allocation off the curve from the previous essay — the same curve applies, because the algebra is identical with σᵢ/√cᵢ in place of σᵢ. If the ratio is inside about 1.5, the whole question is worth a few per cent and can be ignored. Outside 3, it is worth a quarter of the budget.
The whole table is two closed forms
The three enumerated variances are worth deriving rather than only reporting, because the two expressions behind them are the previous field’s with one substitution and they say something the table does not.
Minimising σ₁²/n₁ + σ₂²/n₂ subject to c₁n₁ + c₂n₂ = C gives an optimal variance of
(σ₁√c₁ + σ₂√c₂)² / C
which at the figure’s numbers is (1 + 3√20)²/4000 = 14.4164²/4000 = 0.051957, against the enumerated 0.05196. And equal units spends C/(c₁ + c₂) on each arm, for a variance of (σ₁² + σ₂²)(c₁ + c₂)/C = 10 × 21/4000 = 0.0525, against the enumerated 0.05263.
So the budget problem is the unit problem with σᵢ replaced by σᵢ√cᵢ throughout — the optimal variance is the squared sum of those, over the budget, exactly as it was the squared sum of the spreads over the total. Setting c₁ = c₂ collapses one expression into the other, which is the check the essay’s third figure makes numerically.
The penalty has no ceiling here
That substitution changes one thing the previous field established, and it changes it completely.
Under a unit constraint, equal allocation costs 2(σ₁² + σ₂²)/(σ₁ + σ₂)² − 1, which never exceeds a factor of two however far apart the spreads are. Under a budget constraint the same quantity is
(σ₁² + σ₂²)(c₁ + c₂) / (σ₁√c₁ + σ₂√c₂)² − 1
and its limit as one arm dominates is (c₁ + c₂)/c₁ or (c₁ + c₂)/c₂, whichever arm dominates.
At the figure’s costs of 1 and 20 that upper limit is 21. So equal units under a budget can cost twenty times the optimal variance, where equal units under a unit constraint can never cost more than twice it. The reassuring bound the previous field ends on does not survive the change of constraint, and it fails in the direction that matters: the wider the price gap, the more a convention that ignores prices can cost.
The 1.3% measured in the table is therefore a coincidence of these particular spreads rather than a reassurance. At σ = (1, 3) and costs of 1 and 20 the two quantities σ√c are 1 and 13.4, a ratio of 13.4 — far outside the factor of 1.5 the previous field’s rule of thumb calls safe — and equal units happens to land near the optimum only because it is wrong about the spreads and the prices in opposite directions. Change either and the cancellation goes.
Integers, and where the algebra stops describing the design
The rule is continuous and an allocation is not, and at these prices the gap between them is not negligible.
At a budget of one thousand with the same costs, the enumerated best pair is 80 and 46, a ratio of 1.74 against the algebra’s 1.49. At twenty thousand it is 1.48. The rule is right and the design is coarse: one unit of the expensive arm costs twenty of the cheap one, so the affordable pairs form a lattice with wide spacing in one direction, and whatever budget is left over after buying the last expensive unit has nowhere to go but the cheap arm.
That is worth knowing rather than smoothing over, because it is the regime real experiments live in. A trial that can afford eleven sites and has money left over does not buy 11.4 sites. It buys eleven and spends the remainder on more patients at the ones it has, which is exactly the lattice above.
The check requires the convergence rather than the value: the integer optimum has to sit further from the algebraic ratio at a small budget than at a large one, and it does. A tolerance that passed at both budgets would be a tolerance hiding a rule that was simply wrong.
What a cost actually is
The arithmetic is easy and the inputs are not, so it is worth being explicit about what belongs in cᵢ.
Marginal cost, not average cost. The quantity that matters is what one more unit in that arm costs. Setting up an assay, training a site, building a rig: those are paid once and do not enter the allocation at all, even when they dominate the budget. A rule computed from average costs will systematically under-buy whichever arm has the larger fixed component.
Everything that runs out, in one currency. If the binding constraint is technician-hours rather than money, cᵢ is hours. If two things bind at once, the problem is not the one solved here — it is a linear program with two constraints, and the σ/√c rule is the special case where one of them is slack.
Cost per unit is not always constant. Sites, batches and machines come in blocks: the twentieth patient at a site is cheap and the first patient at a new site is not. Where that is the structure, cᵢ in this arithmetic is a smoothed version of something with steps in it, and the smoothing is another reason the integer lattice above is the honest picture rather than a caveat on it.
And cost is a design variable too. An arm that is expensive because of how it is measured can sometimes be made cheaper, and cutting c₂ by a factor of four buys the same precision as cutting σ₂ by a factor of two — which follows directly from the square root, and is the one useful thing that asymmetry tells an experimenter who is not allocating anything.
The analysis has to match this design too
The previous essay ended with a warning that applies here with more force, because a budgeted design is unequal by a larger factor than a variance-optimal one.
At spreads of 1 and 3 and costs of 1 and 20, the design buys 280 units of the quiet arm and 186 of the noisy one. That is the reverse of the variance-optimal direction — more of the quieter arm — which is precisely the direction in which the pooled-variance t test becomes liberal. On a hundred units at 90:10 that test rejected a true null 37.1% of the time; this design is less extreme and pointing the same way.
So a budgeted allocation makes Welch’s test not a refinement but a requirement. The pooled test’s assumption is that the two arms have the same spread, and the whole reason this design is unequal is that they do not.
There is a general shape here that is worth stating once. A design decision made for one reason — money — has changed which analysis is valid, through a route that has nothing to do with money. Nothing in the budget arithmetic mentions the test, and nothing in the test’s assumptions mentions the budget. The two meet only in the finished experiment, which is where this kind of failure is always found and never looked for.
What this does to power calculations
Sample-size calculations are written in units, which makes them the wrong shape for a budgeted experiment and quietly builds the equal-allocation assumption into everything downstream.
The standard formula gives n per arm for a stated power at a stated effect. Under a budget the question is not “how many per arm” but “what is the best precision C buys” — and the answer is a variance, from which power follows through the non-central t, rather than a number of units.
The two calculations differ in their answer to a question the unit version cannot even ask: is this experiment worth running at all? A budget that cannot buy a variance small enough for the effect of interest is a budget that should not be spent, and that comparison is only visible when the constraint is the money.
The same rule, in the field that invented it
This rule is older than response surfaces and did not come from experiments at all. It comes from survey sampling, where a population is divided into strata and the question is how many people to interview in each — and where both terms are naturally unequal: a stratum with high variance deserves more of the sample, and a stratum that is expensive to reach deserves less, in proportion to the square root of what it costs to reach.
Two things carry across and one does not, and the one that does not is worth flagging.
What carries: the arithmetic is identical. Minimising a sum of σᵢ²/nᵢ subject to a linear constraint gives nᵢ ∝ σᵢ/√cᵢ whether the index runs over arms of a trial or strata of a population, and the enumeration in the figures above is indifferent to which is meant.
What carries: the flatness. A stratified sample allocated roughly right loses a few per cent, which is why proportional allocation — sample each stratum in proportion to its size, ignoring both σ and c — survives as a default in survey practice despite being optimal only when every stratum has the same variance.
What does not carry: the estimand. A survey is estimating a total or a mean over the whole population, so each stratum’s contribution is weighted by its size, and the rule picks up that weight. An experiment is estimating a difference, and the two arms enter symmetrically. Applying a survey sampling formula to a trial, or the reverse, gives an answer that differs by exactly those weights and looks perfectly reasonable.
That is worth a sentence of caution because the two literatures use the same words. “Optimal allocation”, “Neyman allocation” and the σ/√c rule mean the same arithmetic and are applied to different estimands, and a formula lifted across the boundary is the kind of mistake this site has recorded before: every quantity correct, and the wrong one being used.
The rule that keeps arriving
Three of the four essays in this field end at the same place, and it is worth naming here because this is the one where it is least expected.
Every allocation rule so far is a function of quantities that are not known when the allocation is decided. σ₁ and σ₂ are the point of the experiment. And now c₁ and c₂ join them — marginal costs are better known than variances, usually, but “better known” is not “known”, and a cost estimate that is out by a factor of two moves the optimal ratio by √2.
The saving grace is the flatness the previous essay measured. The variance curve is quadratic at its minimum, so a ratio wrong by 40% costs a few per cent, and an experimenter who knows the costs to within a factor of two is doing well enough. That is the practical reason this rule survives contact with real budgets, and it is the same second-order argument that makes steepest ascent work with a direction estimated from four runs.
What does not survive is the plug-in habit itself, applied without noticing. The last essay in this field is what happens when the estimate is bad enough that the rule built on it is worse than no rule at all — and where the crossing point is turns out to be much further from equal spreads than anyone would guess.
There is one difference worth holding on to between the two inputs, though, and it is in the designer’s favour. A variance has to be estimated from data of the kind the experiment will collect, which is why it is not known until the experiment is over. A cost is usually known — it is on an invoice — and where it is not, it is knowable by asking rather than by sampling. So of the two quantities this rule needs, the one this essay added is the reliable one, and an experiment that knows its costs exactly and its variances roughly is in a better position than the previous essay’s arithmetic suggests.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The arm whose variance is its answer — both name allocation ratio, experimental design, neyman allocation, sample size, variance reduction
- Two contrasts, one split — both name allocation ratio, experimental design, neyman allocation, variance reduction
- Stationary is not convergent — both name discreteness, experimental design, variance reduction
- The same draws for both methods — both name discreteness, sample size, variance reduction
- The variance removed before the data — both name allocation, experimental design, variance reduction
- A coverage table with its own error — both name discreteness, sample size
Named objects
A flat tag is an object no other essay names yet.
AllocationAllocation ratioCostDiscretenessExperimental designNeyman allocationSample sizeVariance reduction