A contrast chosen after the data
Worth reading first: Not half and half.
Two contrasts, one split found that a two-arm trial with a binary outcome cannot be split optimally for more than one way of reporting its result. With proportions of 0.6 and 0.1, a risk difference wants 62.0% of the units in the first arm, a log risk ratio 21.4% and a log odds ratio 38.0%; the difference’s rule and the odds ratio’s are exact reflections; and a trial that will report all three does best at the odds ratio’s own split, with its worst contrast 24.5% above that contrast’s best variance.
Every contrast in that essay was chosen before the trial ran. The essay ended on the design that chooses afterwards: a trial that reports whichever contrast the data favour has no fixed target to allocate for, and the selection is a choice the allocation cannot anticipate. Whether the compromise split is still right under selection, how much the selection costs on top of it, and whether pre-registering a pair of contrasts does better than pre-registering one, were left unmeasured. They turn out to have answers that depend almost entirely on the split.
Three statistics that ask one question
The trial here is four hundred units with a binary outcome, the first arm’s proportion 0.18 and the second’s 0.1 — a difference a trial of that size has about a two-in-three chance of detecting, which is where a choice of contrast can matter. Each contrast is estimated from the counts with the usual half added to every cell and divided by its delta-method standard error at the estimated proportions:
and the log risk ratio’s likewise. The trial reports the one with the largest .
All three test the same null hypothesis — the two proportions are equal — and every one of them is zero exactly when the others are. They differ only in the scale on which the difference is measured and in how each scale’s variance is estimated from the data. So the selection is not three independent chances at a false positive; it is one chance seen through three slightly different windows.
What the selection costs when nothing is there
At an equal split the chosen contrast rejects 5.1%, 4.8% and 5.2% of the time with both arms at 0.05, 0.1 and 0.3 — exactly the rate of the risk difference read alone, to the trial. Moving towards a 30–70 split, the rate rises: at 30% in the first arm, 6.5%, 6.0% and 5.8%; at 70%, 6.3%, 5.8% and 5.7%. The rarest outcome inflates most, because there the three variance estimates differ most from one another.
The inflation is small because the windows are close. Three tests of one hypothesis on one data set are correlated far more strongly than the correlation between three different hypotheses would be, and the 5% point of their maximum — counted under the null at the same split — is between 1.90 and 1.96 at an equal split and between 2.00 and 2.07 at the outer splits, against Bonferroni’s 2.39 for three independent tests. A family of three statistics that ask one question is worth little more than one.
At an equal split there is nothing to choose
The panels show why the equal split is special. With two hundred units in each arm the difference’s is at least the odds ratio’s in all four hundred trials drawn, and in every one of twenty thousand counted at each of three outcome rates, the difference’s is the largest of the three. The points lie on one side of the diagonal without exception. With 30% in the first arm the cloud has moved across it, and the difference is the larger in seven of four hundred trials.
That the difference wins whenever the arms are equal is counted here, not proved. It is consistent with how the three standard errors are built: with equal arms, the difference’s averages the two arms’ variances while the odds ratio’s averages their reciprocals, and the odds ratio’s numerator is the difference divided by a variance from between the two proportions. With unequal arms each standard error weights the arms differently, and which statistic comes out largest then depends on which arm’s variance the weighting favours — which is what the right-hand panel shows happening.
So the choice the trial makes is set by its split. With proportions of 0.18 and 0.1, the difference is chosen in 100.0% of trials at an equal split and in 98.3% to 99.0% of trials with more units in the first arm; the log risk ratio in 98.1% with 30% in the first arm and 56.4% at 40%; the odds ratio in 56.7% at 45% and in 42.4% at 40%. A trial reporting its most significant contrast has, in effect, pre-registered one of three contrasts by choosing its split, and the choice is close to deterministic except near 40% to 45%, where the risk ratio and the odds ratio trade places.
Power, and the pair
The risk difference alone has 63.8% power at an equal split, 53.7% at 30% in the first arm and 58.3% at 70%. The log risk ratio peaks at 63.4% with 40% in the first arm, the odds ratio likewise at 63.4%. The most significant contrast, read at its own 5% point so that its false-positive rate is honest, has 64.4% at an equal split — the best of every reading at every split — and between 55.5% and 64.0% elsewhere.
The pre-registered pair is the expensive option. Testing the difference and the odds ratio each at 2.24, which holds the family at 5% whatever the two statistics’ correlation, has 52.6% power at an equal split and between 47.5% and 53.3% across the range: eight to twelve points below the chosen contrast at its own critical value, and below either contrast alone. Bonferroni charges the pair as if it were two independent hypotheses. It is one hypothesis seen twice, and the correlation that makes the selection nearly free makes the Bonferroni charge nearly all waste.
What naming the right contrast would have bought
Selection’s cost has two parts, and the split decides which is present. One is the false-positive inflation at 1.96, removed by reading the maximum at its own 5% point. The other is the power a trial gives up by not knowing in advance which contrast its split favours: at its own critical value the chosen contrast must clear a slightly higher bar than a single pre-registered contrast, and it clears it with whichever statistic happened to be largest rather than the one that would have been largest on average.
At an equal split neither part exists. The chosen contrast is the difference on every trial, and the 5% point of the maximum at the pooled proportion is 1.94 rather than 1.96 — slightly below the textbook value, because the difference’s own null distribution at these sizes is a little narrower than a normal — so its power, 64.4%, is a little above the difference’s 63.8% read at 1.96. At 30% in the first arm both parts are present. The risk ratio alone, the contrast that split favours, has 61.1%; the chosen contrast at its own critical value has 58.0%. Choosing after the data costs three points against having named the right contrast, and a trial that splits unequally on purpose — for cost, for ethics, for the reasons a split is not half and half — should name the contrast its split serves.
A rarer outcome
The arithmetic sharpens as the outcome becomes rare, because the three variance estimates diverge more when the counts are small. With proportions of 0.1 and 0.05 — a trial of four hundred with about thirty events in all — the chosen contrast at 1.96 rejects 6.3% of the time with both arms at 0.075 and 30% of the units in the first arm, 4.9% at an equal split and 5.9% at 70%.
Power moves with the split more sharply too. At an equal split the difference alone has 47.1%, the chosen contrast at its own critical value 48.4%, and the pre-registered pair 35.7% — almost thirteen points less. At 30% in the first arm the risk ratio and odds ratio alone have 46.3% each, the difference 36.5%, and the chosen contrast 42.6%; at 70%, the difference has 43.6% and the ratios 27.2% and 29.9%. A rare outcome punishes a split that does not suit the contrast, and punishes the Bonferroni pair more than either.
Allocating on a guess found that a split set from a guessed proportion inherits the guess’s error. Here the guess that matters is not the proportion but the contrast, and an equal split is the one allocation that needs no guess about either: it is near-optimal for every contrast’s variance at these proportions and it makes the choice among them irrelevant.
The compromise split, under selection
The earlier essay’s compromise for these proportions — the split that keeps every contrast’s variance closest to its own best — is 48.9% in the first arm, nearly an equal split, because the three contrasts’ optimal splits for proportions of 0.18 and 0.1 are close together: 56.2% for the difference, 41.6% for the risk ratio, 43.8% for the odds ratio. Under selection, the best split is also at or near half, where the chosen contrast’s power peaks at 64.4%.
So the compromise and the selection agree, and for a reason that is not a coincidence. A split far from half makes the variance of one contrast much worse than its best, which is what a design built for the worst case avoids by sitting in the middle of the optima; it also makes the three statistics disagree with one another, which is what makes selection cost anything. Near an equal split both problems vanish together. With proportions of 0.6 and 0.1, where the earlier essay found the optima spread from 21.4% to 62.0%, no split avoids both; what selection costs at that compromise of 38.0% has not been counted here.
Trials whose verdict depends on the scale
A different way to see how much the choice matters is to count the trials in which it changes the verdict: some contrast significant at 5% and another not. With both arms at 0.1, that happens in 0.5% of trials at an equal split and 2.8% at either 30–70 split. With proportions of 0.18 and 0.1 it happens in 1.5% of trials at an equal split, 7.3% with 30% in the first arm and 10.8% with 70%.
Those are the trials in which a reader of the published contrast and a reader of a different one would draw different conclusions from the same data, and the reader of the published one has no way to know. At an equal split there are almost none of them: one trial in seventy under the alternative, one in two hundred under the null. At an unequal split one trial in nine can be reported as a finding or as nothing depending on the scale the authors preferred. That is the real cost of an unequal split when the contrast is not named, and it is larger than the false-positive inflation suggests, because most of the discordant trials are trials where something is genuinely there.
It is also a cost no correction removes. Reading the maximum at its own critical value repairs the rate of false findings; it does not make the three scales agree in the trials where they disagree, and a trial that reports the difference’s non-significance when the risk ratio was significant has made a choice about what to say as surely as one that reported the risk ratio. The cost of a unit priced an unequal split in units; here the same split carries a price in how often the trial’s conclusion depends on a sentence in its analysis plan.
What this says about the choice of contrast
The interval after the choice found that choosing a model from the data costs a forecast interval more coverage than estimating the model does. A choice of contrast is a choice of model too — three link functions for one binary outcome — and here it costs almost nothing, because the three models are not competing descriptions of the data; they are three scales for one comparison, and any of them is zero exactly when the others are.
That stops being true the moment the choice is between hypotheses rather than scales. A trial that reports the most significant of three outcomes, or of three subgroups, is choosing among questions, and what the correction corrects prices that at very nearly three chances. The same selection rule — report the most significant — is nearly free or close to a threefold charge depending on whether its candidates share a null hypothesis.
What a protocol can take from it
Split equally if the contrast may be chosen later. At an equal split the most significant contrast is the risk difference on every trial, so reporting “the most significant” is reporting the difference and costs nothing. The earlier essay found an equal split a respectable compromise for variances; here it is also the split that makes selection harmless.
At an unequal split, name the contrast or use the maximum’s own critical value. The inflation is a point or a point and a half, and the honest critical value for the maximum of three — about 2.00 to 2.07 at a 30–70 split — removes it at a cost of a few points of power.
Do not pre-register a pair at Bonferroni’s threshold. It charges a correlated pair as two hypotheses and pays eight to twelve points of power for protection against an inflation that was a point and a half. A pair that is to be reported should be tested at the joint critical value of its maximum.
Allocate for the contrast whose scale the decision uses. The arm whose variance is its answer found the difference’s allocation rule costing little to ignore across most proportions; the selection result adds that the difference is also the scale a balanced trial ends up reporting whenever it is free to choose.
Counted, and what was not
Every rate is over twenty thousand simulated trials of four hundred units at each split and set of proportions; the maximum’s 5% point at each split is counted over twenty thousand trials with both arms at the pooled proportion; the four hundred trials drawn in the scatter are seeded. The proportions are 0.18 and 0.1 for power, and 0.05, 0.1 and 0.3 in both arms for the false-positive rates; the rarer outcome uses 0.1 and 0.05, with 0.075 in both arms for its null. The discordant trials are counted on the same draws as the power and the false-positive rates, so every percentage in a row comes from one set of trials and the differences between readings are differences between methods.
Not claimed: anything about contrasts estimated by likelihood rather than with half a unit added to each cell, or with variances computed under the null rather than at the estimated proportions — a score test of each contrast pools the arms and the three tests then coincide more closely still. Nor anything about covariate-adjusted contrasts, where the odds ratio is non-collapsible and the choice between scales stops being a choice of window on one comparison.
Still open: an adaptive split that follows the chosen contrast
The question the earlier essay posed in its title was about adaptive analysis. A trial that splits half and half for its first hundred units, looks at which contrast is leading, and allocates the rest for that contrast’s optimum, makes the split a function of the data and the contrast a function of the split — two selections chained.
Whether that design gains power over a fixed equal split, how much it inflates the false-positive rate when the arms do not differ, and whether the chosen contrast’s own critical value still restores the level when the split itself was chosen by the data, are measurable on the same trials and have not been measured here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- One control, many arms — both name allocation, bonferroni, multiple comparisons, statistical power
- A coverage table with its own error — both name bonferroni, multiple comparisons, statistical power
- A slow return across many pairs — both name bonferroni, multiple comparisons, statistical power
- An order that spends the error rate — both name bonferroni, multiple comparisons, statistical power
- One characteristic that really matters — both name bonferroni, multiple comparisons, statistical power
- Sixteen subgroups and one effect — both name bonferroni, multiple comparisons, statistical power
Named objects
A flat tag is an object no other essay names yet.
AllocationBonferroniMultiple comparisonsOdds ratioRelative riskRisk differenceSpecification searchStatistical power