Two contrasts, one split
Worth reading first: Not half and half · Balanced on the wrong function.
The essay before this one minimised the variance of a risk difference and found the rule barely worth using. It took for granted that the trial is estimating a risk difference.
A great many two-arm trials with a binary outcome report a ratio instead — a relative risk, or an odds ratio — and a good many report more than one in the same table. Those are different quantities with different variances, and a split that minimises one has no reason to minimise another.
Three rules, three answers, one dataset. And the first and the third do not merely disagree — they point in opposite directions, exactly.
That exactness is worth pausing on, because it rules out the reading a designer would prefer. Two rules that disagree by an amount could be split between; two rules that are reflections of each other in the half line have the property that any split favouring one arm is unfavourable to the other by a matching amount, and there is no arrangement that is nearly right for both except the one exactly in the middle.
Which contrast a trial is actually for
Before the arithmetic, the question the arithmetic makes unavoidable: what is the trial estimating?
The three contrasts are not three summaries of one number. They answer different questions and a treatment can be impressive on one and unremarkable on another. A treatment taking a rate from 1% to 2% has a risk difference of one point — negligible for a population — and a risk ratio of two, which reads as a doubling. A treatment taking 40% to 60% has a risk difference of twenty points and a risk ratio of 1.5. Which of those is the larger effect depends entirely on which contrast is asked for, and the two orderings are opposite.
So “the contrast” is a decision about what the trial means, not a presentational choice made at write-up. What follows is that the allocation cannot be decided without it, and that a protocol which fixes the sample size and the split without fixing the contrast has left a design decision to be made after the data arrive.
Three variances, three weights
Each rule allocates units in proportion to the square root of a weight, and the weight is whatever the arm contributes to the variance of the contrast. The delta method gives all three:
- risk difference — , so ;
- log risk ratio — , so ;
- log odds ratio — , so .
The first and the third are reciprocals. An arm with a larger p(1 − p) gets more units if the trial reports a difference and fewer if it reports an odds ratio, and the two shares add to one exactly — checked to twelve decimal places at every proportion in the sweep.
The reason is the shape of the two estimands. A difference is measured on the probability scale, where an arm near a half is the noisy one and deserves the units. An odds ratio is measured on the log-odds scale, where an arm near a half is the stable one: the transformation’s derivative is 1/(pq), which is smallest at a half, so an arm near a half contributes least to the variance and deserves fewer units. The same arm is the difficult one on one scale and the easy one on the other.
The risk ratio sits between them in form and not in position: its weight q/p is monotone rather than symmetric, so it always wants more units in the arm with the smaller proportion, however far from a half the other arm is.
The weight functions, side by side
The three weights are worth looking at directly, because their shapes explain everything the optimal-share curves do.
The difference’s weight is bounded above by 0.5 and goes to zero at both ends. Its reciprocal is bounded below by 2 and goes to infinity at both ends, so the odds ratio’s rule is the aggressive one: an arm at a twentieth gets 2.2 times the weight of an arm at a half, where under the difference’s rule it would get 0.44 times.
That asymmetry is why the two rules’ penalties for using the wrong one are not equal. The difference’s rule asks for splits in a narrow band, so using it when an odds ratio was wanted is a modest error; the odds ratio’s rule asks for splits that can be extreme, so using it when a difference was wanted can be a large one. The cross-costs below show the asymmetry as numbers.
What the disagreement costs
The three optima are far apart and the space between them is not free.
At = 0.6 and = 0.1 the three rules ask for 62.0%, 21.4% and 38.0% of the units in the first arm. Reading a contrast off a split chosen for another costs:
- allocating for the risk ratio, reporting the difference: +98.1% of variance, which is nearly doubling it;
- allocating for the difference, reporting the risk ratio: +70.1%;
- allocating for the difference, reporting the odds ratio: +24.5%;
- allocating for the odds ratio, reporting the risk ratio: +11.7%.
Doubling a variance is halving the effective trial size. A trial that allocated for one contrast and is read for another has thrown away half its units, and nothing in the analysis says so, because each analysis is correct on its own terms.
The asymmetry of the penalties is worth reading off that list. The two largest costs both involve the risk ratio — allocating for it and reporting the difference, or the reverse — and that is because its weight q/p is the only one of the three that is not symmetric about a half. It wants the low-proportion arm to be large without limit as that proportion falls, so its optimum runs away from the other two rather than sitting between them.
Between the difference and the odds ratio alone the disagreement is more contained: +24.5% either way, because the two optima are reflections and each is exactly as far from the other’s as the other is from its. That symmetry is a consequence of the reciprocal relation and it is the one piece of good news in the table.
The compromise, and the one that is not the obvious one
A trial that will report more than one contrast has to choose a split that is optimal for none. The sensible criterion is the worst case across the contrasts it will report.
The best compromise is the odds ratio’s own split, at 38.0% of the units in the first arm, with a worst case of +24.5%. That it wins is not an accident: it sits between the other two optima by construction, because it is the reciprocal of one and the risk ratio’s is on the same side as the other.
An even split is second, at +32.7% — eight points behind the best and far ahead of the two obvious choices. A trial that splits evenly has not chosen well, and it has not chosen badly either, which is more than can be said for a trial that allocates for the contrast it leads with.
And the worst compromise is the difference’s split, at +70.1%, which is the split most trials would reach for, because the risk difference is what a protocol usually states as the primary outcome.
Why an even split does as well as it does
That the even split places second is worth an explanation, because it was chosen by nothing and beats two rules each chosen by an optimisation.
A minimax criterion rewards being close to every optimum rather than on any of them, and the even split is the point that no contrast’s rule ever strays very far from: the difference’s rule stays within about twelve points of it, the odds ratio’s within twelve on the other side, and only the risk ratio’s runs away. So the even split is nearly the centre of the set of optima, and the centre of a set is what a minimax over that set rewards.
The two rules that lose do so by being extremal. The difference’s optimum at 62.0% is the furthest right of the three and the risk ratio’s at 21.4% the furthest left, so each is maximally far from the other, and their worst cases are that distance squared through the variance function.
This is the same reason a design built for the worst case rarely looks like a design built for any particular truth: the criterion is about the whole set of possibilities and the answer sits in the middle of it, which is a place no single possibility would have chosen.
What a defensible design looks like
State the split’s justification in the protocol. A split of 62 : 38 with no reason attached is indistinguishable from a split chosen for recruitment convenience, and the reader cannot reconstruct which contrast it was optimal for. One sentence — “allocated to minimise the variance of the risk difference at an assumed 0.6 and 0.1” — makes the design auditable and makes the secondary contrasts’ inflated variances predictable rather than surprising.
Name the contrast before choosing the split. The three rules disagree by forty points of allocation at these proportions, so the choice of split is not separable from the choice of estimand — and the estimand is a decision about what the trial means rather than about how it is analysed. This is the same discipline a two-arm rule that may not pool needs: the analysis that will be run is part of the design.
Fix the contrast in the protocol, with the split beside it. The two decisions are one decision and separating them in a document is what lets the second be made by default. That is the same discipline a stopping rule needs stated before the data: a choice made in advance and written down is a different object from the same choice made later.
Keep the split inside the band the rules agree on. All three optima lie between 21.4% and 62.0% here, and anything outside that interval is worse than every rule for every contrast — a split of 80 : 20, chosen for recruitment or for ethics, is dominated. The interval is computable from assumed proportions before the trial, which is the same discipline an admissible set imposes: the first question is which designs are not ruled out, and only then which is best.
Where more than one contrast will be reported, allocate for the odds ratio or split evenly. Both have worst cases under a third; the two obvious alternatives have worst cases near or above two-thirds. Between them, the even split has the advantage of needing no estimate of the proportions, which is what the previous essay was about.
Report the split’s cost for each contrast in the table. The numbers are closed-form and need only the assumed proportions, so a trial can state that its secondary risk ratio carries 70% more variance than an optimally allocated one would. That is a second number beside the interval, and it is the kind a width reported beside a coverage supplies elsewhere here.
And do not read a secondary contrast as though it were free. A trial powered for its primary difference and reporting a risk ratio alongside is reporting the second with an effective sample size substantially smaller than its first. That is not a reason to leave it out; it is a reason for its interval to be read as wider than the primary one, which it will be, and for the reader to know why.
The previous essay’s conclusion, revisited
The essay before this one concluded that for a binary outcome a trial should split evenly and stop thinking about it, because the most an optimal split can win is 4.36% of variance.
That conclusion was about a risk difference and it does not survive intact. Across the three contrasts the stakes are an order of magnitude larger: +98.1% rather than +4.36%, because the disagreement between two contrasts’ optima is far wider than the disagreement between one contrast’s optimum and an even split.
What survives, and survives more strongly, is the recommendation. An even split costs at most 4.36% if the trial reports a difference and at most 32.7% across all three, and both of those are better than what a trial gets by allocating confidently for the wrong contrast. The previous essay said an even split was cheap; this one says it is also safe, and the two arguments are independent.
The one thing that changes is the reason. Splitting evenly is not a way of avoiding a decision about the proportions — it is a way of avoiding a decision about the estimand, which is the more consequential of the two and the one a protocol is less likely to have made.
The estimand decides more than the split
One consequence runs past allocation and is worth stating, because it is the reason the choice of contrast is not a presentational one.
The estimand decides the analysis. A risk difference is estimated by a difference of proportions and an odds ratio by a ratio of odds, and the two have different sampling distributions, different small-sample behaviour and different intervals. An odds ratio’s log is much closer to normal than a risk difference is at small counts, which is one of the reasons it is used; a risk difference is bounded and interpretable on the scale a decision is made on, which is one of the reasons it is preferred.
The estimand decides the sample size. Power for a difference and power for a ratio are computed from the two variances above, and a trial sized for one is not sized for the other. At = 0.6 and = 0.1 the two differ by whatever the ratio of their variances is at the chosen split — the same factors the cross-cost table reports.
And the estimand decides what “no effect” means. A risk difference of zero and a risk ratio of one are the same hypothesis, so the null is common to all three; everything else about the three analyses is not. This is the same distinction an interval built for a different question turns on: the construction is shared and the meaning is not.
What is claimed here and what is not
One pair of proportions for the cross-costs. The +98.1% and the rest are at = 0.6 and = 0.1, which is a large effect with one arm far from a half — the case where the three rules are furthest apart. At proportions closer together the disagreement shrinks and at it vanishes, since every rule then says half. The reciprocal relation between the difference and the odds ratio holds everywhere and the size of the penalty does not.
The variances are delta-method variances. All three are first-order approximations to the variance of a transformed estimator, and the previous essay checked that the difference’s one describes its estimator. The two log contrasts are checked the same way and to the same tolerance, and at small counts in either arm all three are approximations that a trial would be unwise to trust.
The compromise is a minimax over three contrasts that were chosen. A trial reporting only a difference and an odds ratio has a different best compromise — the two optima are reflections, so their minimax is exactly an even split, which is a neat special case and not the general one. The ranking above is for the three contrasts together.
The worst case is over contrasts, not over proportions. A trial does not know its proportions in advance, so a fuller treatment would take a worst case over both — over which contrast is reported and over what the proportions turn out to be. That is a harder problem with a different answer, and the even split’s case would strengthen in it, since the even split is the only candidate here that does not need the proportions to be computed at all.
And the reciprocal relation is exact and is about these two contrasts. and give shares that sum to one because each share is a/(a + b) with the roles of a and b exchanged. It is an identity, checked to , not a numerical coincidence, and it does not extend to any other pair of contrasts in this essay.
Still open: what a trial should allocate for when the analysis is adaptive
Everything here assumes the contrast is fixed before the trial runs. A design that decides which contrast to report after seeing the data has no fixed target to allocate for, and the split it chose is optimal for whichever contrast it turns out to select — which is a selection the allocation cannot anticipate.
That is a shape priced elsewhere and not here: a rule chosen before the data, evaluated against a target chosen after it, is a charge the analysis has not paid. Whether the worst-case split above is also the right split under a selection rule, how much the selection costs on top of the compromise, and whether a design that pre-registers a pair of contrasts does better than one that pre-registers one, are questions with a clear shape and no measurement behind them here.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A threshold in the tail — both name allocation rule, closed form, efficiency, treatment effect, variance reduction
- Three functions of one number — both name allocation rule, efficiency, maximin design, treatment effect, variance reduction
- The check worth more than the check — both name binomial proportion, closed form, efficiency, variance reduction
- The cost of a unit — both name allocation ratio, experimental design, neyman allocation, variance reduction
- Which shapes are worth protecting — both name allocation rule, efficiency, maximin design, variance reduction
- A basis is a subspace — both name allocation rule, treatment effect, variance reduction
Named objects
A flat tag is an object no other essay names yet.
Allocation ratioAllocation ruleBinomial proportionClosed formDelta methodEfficiencyEstimandExperimental designMaximin designNeyman allocationOddsTreatment effectVariance reductionWorst case