The arm whose variance is its answer
Worth reading first: Not half and half · Balanced on the wrong function.
The essay that opened this question established that the variance-minimising split of a fixed number of units is , and the third established the difficulty: every rule in this field is a function of quantities the experiment is being run to find out, and fed a pilot’s estimate of them the rule can make the experiment worse than not bothering.
With a continuous outcome the spreads are separate parameters and a pilot can estimate them badly or well. With a binary outcome they are not separate at all. A proportion of p has a variance of p(1 − p) and nothing else, so the quantity the rule needs is the quantity the trial exists to estimate — exactly, with no slack between them.
That sounds like the problem in its purest form and it is the case where the problem nearly vanishes.
A rule that needs the answer, and that it costs under five per cent to ignore.
That comparison is the one the rest of the field has been building towards without being able to make it. The first three essays in this field treated the allocation rule as a thing to get right, and priced getting it wrong; this is the first case where the rule can be priced against the alternative of not using it at all, because the value of perfect information is bounded and computable.
Why the rule barely moves
The reason is the shape of one function.
is as flat as a function on the unit interval gets while still going to zero at both ends. It is 0.500 at p = 0.5, 0.458 at 0.3, 0.300 at 0.1 — a proportion five times smaller with a weight only 40% smaller. Two arms whose proportions both lie between 0.067 and 0.933 differ in weight by at most a factor of two, so the rule they drive asks for a split no further from even than two to one.
So the rule’s range is what makes it safe. A rule whose answer is always between about 40% and 52% of the units cannot be got very wrong by guessing, because the guess and the truth are both in a narrow band.
That is a different reason from the one the third essay in this field found for continuous outcomes. There the plug-in rule was dangerous because the spreads could differ by any factor at all and a pilot could be wrong by a large multiple. Here the spreads cannot differ by more than a bounded factor unless one proportion is extreme, and the danger is confined to the ends.
The same dependence, in a form that cannot be escaped
It is worth being precise about how tight the circularity is here, because “the rule needs the answer” is a phrase that covers a range of situations and this is the extreme of it.
For a continuous outcome the mean and the variance are separate parameters. A trial estimating a difference in means needs and to allocate, and those are nuisance parameters — knowable in principle from a pilot, from previous studies, from a measurement instrument’s specification, without knowing the means at all. The rule needs something, and the something is not the answer.
For a binary outcome there is one parameter per arm and it does both jobs. Knowing enough to allocate well is knowing the answer, and no pilot, instrument specification or previous study can supply the variance without supplying the effect. The circularity is complete rather than awkward.
That is exactly why the flatness matters. If the weight function had any appreciable curvature, this would be the one design problem in the field with no honest solution — and it is instead the one with the smallest stake.
What the cost curve says exactly
The cost of an even split is a pure number — it is a ratio of variances, so the trial size cancels — and it has features worth reading.
Two zeros. It is exactly zero at = 0.3 and at = 0.7, given = 0.3. The second is the interesting one: p(1 − p) is symmetric about a half, so an arm at 0.7 has exactly the same weight as an arm at 0.3 and the optimal split is exactly even between two arms whose proportions are mirror images. A trial comparing 30% against 70% — a large, obvious effect — gains nothing whatever from unequal allocation.
A shallow interior. At = 0.5 against 0.3 the cost is 0.19%, at 0.2 it is 0.46% and at 0.1 it is 4.36%. Under one per cent across most of the range.
The shallowness has the same cause as the flatness of the weight, one level up. A variance as a function of the split is smooth and has a minimum, so near the minimum it is quadratic: moving the split by a small amount from the optimum costs the square of that amount. With the optimum never further than about 12 points from even inside the ordinary range, the square of a 12-point displacement is a small number, and that is the whole of why the interior of the curve is shallow. It is the same reason a design criterion is flat near its optimum, applied to a design whose optimum is known to be nearby.
And the cost is at the ends. At = 0.06 it is 10.07% and it keeps climbing. That region is where one arm’s event is rare, and it is also where the trial has other problems: an arm with a 6% event rate on a few hundred units is estimating a proportion from a handful of events, where the interval taught first covers 87.6% of the time and the allocation is not the binding constraint.
What the four per cent is worth in units
A variance ratio is hard to price, so it is worth converting once. A 4.36% reduction in variance is a 2.16% reduction in standard error, and since the standard error falls like , the same gain is available by adding 4.36% to the trial. On four hundred units that is seventeen more.
So the optimal split, at the worst point in the ordinary range, is worth seventeen units out of four hundred. Everywhere else it is worth fewer: at = 0.5 against 0.3 it is worth 0.19%, which is less than one unit.
That is the comparison the decision should be made on, because units are the currency an allocation decision is denominated in, and it makes the answer obvious in a way a percentage of variance does not. A design meeting that would not spend an hour arguing about seventeen units should not spend one arguing about the split.
The rule that needs a pilot, and no longer does
The practical consequence changes the shape of the advice this field has been giving.
For a continuous outcome the third essay’s finding stands: allocate by a pilot’s estimate of the spreads only when the arms are expected to differ by about a factor of two or more, because below that the estimation error costs more than the optimisation gains.
For a binary outcome the calculation is different, because the prize is smaller. The most an optimal split can win over an even one, anywhere between a tenth and nine tenths, is 4.36% of variance — about 2.2% of standard error, which is about what adding 4% to the trial would buy. Set against that, an estimated split carries its own error and the risk of allocating the wrong way round if the pilot’s ordering is wrong.
There is a second consideration that points the same way and is not about variance. An even split is the split at which the test that follows is least sensitive to the two arms having different spreads — the corner where the pooled two-sample statistic loses its level is entered by making the split uneven, and a binary outcome with different proportions has different spreads by construction. So the even split that costs at most 4.36% of variance also avoids a separate problem that an unequal one walks into.
So for a binary outcome the honest advice is to split evenly and stop thinking about it, unless one arm’s proportion is expected below about a tenth — where the gain rises and the pilot’s estimate of a rare rate is at its least reliable, which is the unhelpful combination.
Where the extreme end actually bites
The cost at = 0.06 is 10.07% and it goes on rising, so the “split evenly” advice needs the region where it fails to be described rather than waved at.
Two things happen together as a proportion falls. The optimum moves away from even — at 0.06 the rule asks for 34.1% of the units in the rare-event arm, since that arm’s outcomes vary less — and the cost of ignoring it grows. So the advice reverses at the end of the range, and the reversal is gradual rather than sudden: 0.46% at a fifth, 4.36% at a tenth, 10.07% at about a sixteenth.
What pushes against taking the optimum there is that a pilot estimate of a rare rate is itself poor. A pilot of a hundred units at a true rate of 0.06 sees six events on average and estimates the rate with a relative standard error of about 40%, which propagates into the weight as about 20% — and the third essay’s arithmetic about plug-in rules applies with full force. The region where the optimisation is worth something is the region where the input to it is worst measured.
A trial in that position has a better lever than allocation, and it is the one a fixed-width design uses: decide what precision the rare arm needs and size it for that, rather than splitting a fixed total in a ratio.
What a mirror pair says about effect size
The zero at = 0.7 against = 0.3 deserves more than a line, because it is the case a reader is most likely to have in mind and it goes the opposite way to the intuition.
A trial comparing 30% against 70% has a large effect — a risk difference of 0.4, an odds ratio of 5.4 — and the intuition is that a large effect means the two arms are very different and therefore want very different numbers of units. They do not: p(1 − p) is symmetric about a half, so 0.3 and 0.7 have exactly the same variance, and the optimal split is exactly even.
Meanwhile a trial comparing 0.3 against 0.35 — almost no effect — also wants a nearly even split, because the two weights are nearly identical. The split is not a function of the effect at all. It is a function of how far each arm’s proportion is from a half, in either direction, and two arms on opposite sides of a half can be far apart in effect and identical in weight.
That is why the cost curve has two zeros rather than one, and it is the cleanest available statement of what a Neyman rule is responding to: not the difference between the arms, but their individual spreads.
The variance the rule minimises, checked
An allocation rule is derived from a variance function, and that function is a delta-method approximation to the variance an estimator actually has. A rule minimising the wrong function is minimising nothing.
The two routes agree, which is what licenses reading a cost of 0.19% off a formula rather than off a simulation. The agreement is not perfect and does not need to be: what the rule uses is the shape of the variance in the split, and a multiplicative constant common to every split does not move the minimum.
Where this leaves the rule
Three essays in this field have now said the same thing about the same rule from three directions, and collecting them gives a decision procedure rather than a caution.
The rule itself is and is worth a quarter of the experiment when the spreads differ three to one. Fed a guess, it is worth using only when the spreads are expected to differ by about a factor of two or more. And here, where the outcome is binary, the spreads cannot differ by more than a bounded factor unless a proportion is extreme — so the condition under which the rule is worth using is a condition a binary outcome mostly cannot satisfy.
The three together say that the rule’s value is governed by one quantity, the ratio of the spreads, and that a binary outcome bounds it. That is not a special case of the rule. It is a statement about which experiments the rule was ever going to help.
What is claimed here and what is not
The contrast is a risk difference. Everything above minimises , which is the variance of . A trial reporting a ratio rather than a difference minimises something else, and the something else is not a small perturbation of this — which is the subject of the next essay.
The proportions are the true ones. The cost curve compares an even split against the split that is optimal knowing the proportions, so it is the value of perfect information, and an estimated split is worth strictly less. That makes 4.36% an upper bound on what any procedure could gain, which is the form the comparison needs to be in.
The weight’s flatness is a claim about a bounded interval. is flat in the sense that matters here — it varies by at most a factor of two over 87% of the unit interval — and it is not flat near the ends, where it falls to zero. Both halves of that are used above: the first is why the rule is safe in the ordinary range and the second is why it is not at the extremes.
The trial size does not enter. The cost is a ratio of variances, both proportional to 1/N, so N cancels exactly. What does not cancel is that a small trial at an extreme proportion may have an arm with no events at all, where the variance formula describes nothing — and the sweep is at 400 units, where that is rare at a tenth and not at a fiftieth.
The counted check is at one pair of proportions. The simulation is at = 0.6 and = 0.1 across five splits, which is enough to say the variance function describes the estimator and is not a sweep over the whole range. The closed form is exact arithmetic on the delta-method variance; what the simulation checks is that the delta-method variance is the right function to be doing exact arithmetic on.
And “under 4.36%” is a statement about a range that was chosen. From a tenth to nine tenths is a convention, not a theorem, and the cost rises without bound as a proportion goes to zero. The range is stated because a bound with no range attached to it would be false.
Still open: which contrast the split is for
The whole of this essay minimises the variance of a difference of proportions, because that is what a two-arm trial with a binary outcome is usually described as estimating. A great many such trials report a ratio instead — a relative risk, or an odds ratio — and some report two of them in the same table.
Those are different quantities with different variances, and a split that minimises one need not minimise another. There is a specific reason to expect the disagreement to be large rather than technical: the weight a difference allocates in proportion to is , which is largest at a half, and the weight an odds ratio allocates in proportion to is , which is smallest there. Two rules whose weights are reciprocals do not agree about which arm should be larger. What a trial reporting more than one of them can do about it is the next thing this field has to settle.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A threshold in the tail — both name allocation rule, closed form, efficiency, treatment effect, variance reduction
- The check worth more than the check — both name binomial proportion, closed form, efficiency, monte carlo, variance reduction
- The same draws for both methods — both name binomial proportion, closed form, monte carlo, sample size, variance reduction
- A basis is a subspace — both name allocation rule, monte carlo, treatment effect, variance reduction
- A coverage table with its own error — both name binomial proportion, closed form, monte carlo, sample size
- Balancing towards unequal targets — both name allocation ratio, experimental design, monte carlo, neyman allocation
Named objects
A flat tag is an object no other essay names yet.
Allocation ratioAllocation ruleBinomial proportionClosed formDelta methodEfficiencyExperimental designMonte CarloNeyman allocationPilot studyPlug in estimateSample sizeTreatment effectVariance reduction