Dropping the losers
Worth reading first: When the looking happens · What the correction corrects.
Eight candidate treatments and one control. Running all nine groups to the end is the honest design and it is enormously wasteful: most of those arms are not going to work, and the units spent on them after that is apparent are units spent learning nothing.
So run every arm for a while, drop all but the best, and put the remaining budget into that one and the control. It is the right instinct, it saves a quarter of the experiment, and the test at the end is a test of a hypothesis chosen by looking at the data.
What the naive analysis does
The design: eight arms and a control at sixty each, the arm with the largest observed difference carried forward, thirty more on it and on the control, and a final comparison using all the data on the two surviving groups.
That final comparison looks like a two-sample test. It has two groups, ninety observations each, an estimate and a standard error. Compare it with 1.96 and, under a global null where nothing is real, it fires 10.3% of the time.
The reason is visible in the histogram: the distribution is shifted right. The arm being tested is the one that came out ahead of seven others at the interim, so its first-stage contribution is a maximum rather than a draw — and the final statistic carries that maximum forward inside it, diluted by the second stage but not removed.
This is the sequential field’s principle in a new place. A p-value is defined relative to a sampling plan, and the plan here includes “and then the hypothesis was chosen”. The set of outcomes that could have produced this statistic is not the set 1.96 was computed for.
The value that does hold the rate
There is no table for the right value, and the reason is worth being precise about: the null distribution depends on k, on n₁ and on n₂ together, because the selection is made on a fraction of the data that then reappears inside the final estimate. Change how much of the trial happens before the interim and the whole distribution moves.
So it is solved for, the way this site solves the Pocock boundary: draw one set of null paths against a fixed seed, bisect on the boundary over them, and cache the answer. For this design it is 2.313, against 1.96, and the counted rate at that value is 5.01% on fresh seeds.
Two pins make that believable rather than merely computed. With k = 1 there is no selection and the solver must return the ordinary value: it gives 1.983, and the naive rate at k = 1 is 5.3% — the design is essentially unadapted, as it should be. With eight arms it gives 2.313. A solver that returned something above 1.96 when nothing had been selected would be manufacturing a correction out of its own machinery, and every number in this essay would be that artefact.
The correction depends on when, not just how many
Here is the part that makes this different from an ordinary multiplicity correction.
Bonferroni’s value depends on how many comparisons there were. Dunnett’s, from the allocation field, depends on that and on their correlation. This one depends on how much of the trial the selection saw, and the dependence is not small:
| interim at | share of the trial | naive rate | critical value |
|---|---|---|---|
| 30 of 90 | a third | 5.9% | 2.031 |
| 60 of 90 | two thirds | 7.0% | 2.113 |
| 75 of 90 | five sixths | 7.3% | 2.128 |
All three rows are four arms. What changes is when the choice was made, and selecting on a third of the data and then diluting it with two thirds collected afterwards produces a much smaller inflation than selecting on five sixths.
That is the same arithmetic as the estimate’s bias and it is worth carrying as a design principle: an early interim is a cheap interim. It costs less error rate to correct for, because the selection is made on less information and the stage that follows it is large enough to swamp what the selection did.
Two arms, where the rate looks fine and both tails are wrong
The clearest thing this design taught was found by placing the figure above at two arms, which was meant to be an unremarkable intermediate case.
The two-sided rate comes out at 4.79% — below its nominal 5% — on a design whose statistic has a mean of +0.33 under the null. It is not calibrated; it is two errors cancelling.
| tail | nominal | counted, two arms |
|---|---|---|
| above +1.96 | 2.5% | 4.11% |
| below −1.96 | 2.5% | 0.69% |
| either | 5.0% | 4.79% |
Selection moves mass into the upper tail and out of the lower one. At two arms the second effect is slightly the larger, the two nearly cancel, and a check that looked only at the two-sided rate would report the design as safe while its upper tail was 64% too heavy and its lower tail a quarter of what it should be.
And these designs are tested one-sided in practice — the question is whether the surviving arm beats the control, not whether it differs from it. The one-sided rate at two arms is 8.34%, against a nominal 5%. At eight arms it is 17.79% where the two-sided rate is 10.32%.
So the honest headline for this design is the one-sided number, and the two-sided one is included here because of what it demonstrates: a rate that matches its claim is not evidence of calibration when it is the sum of two errors. This site has made that argument before about procedures controlling different quantities at the same nominal level; this is the same failure inside a single number.
What it saves
The other side of the ledger, and the reason anybody runs this design.
Eight arms and a control at sixty each, then thirty more on two groups, is 600 units. Running all nine groups to ninety is 810. The selection saves 25.9% of the experiment, and it saves it on exactly the arms that were not going to be the answer.
Both curves rise with k, which is the whole design problem in one picture. More arms means more saved — there is more to drop — and a higher boundary, because the winner of a larger tournament is expected to be further ahead. A design that counts the first and forgets the second has kept the saving and spent the error rate.
Why the saving is not the whole of the efficiency
The 25.9% is units, and units are not the only thing a two-stage design saves.
Time. Dropping seven arms at the interim frees the sites, the staff and the supply chain that were running them, and a trial that finishes sooner is worth more than the same trial finishing later. None of that is in the arithmetic here and all of it is in the decision.
And the arms that were never worth running. The design’s real economy is that it converts a decision made badly in advance — which of eight candidates deserves a full trial — into a decision made with data. An all-arms design commits to eight full-sized comparisons before anything is known; this one commits to eight small ones and then to a single large one.
Against that, one cost that is not units either: the seven dropped arms are dropped on very little evidence. Sixty observations each is enough to rank them and not nearly enough to establish that any of them is inferior. An arm dropped at the interim has not been shown not to work — it has been shown to be behind, on a comparison that is itself noisy — and reporting it as a negative result would be a much stronger claim than the design supports.
That is worth stating because it is how these trials are read afterwards. The surviving arm gets a p-value and an estimate. The seven others get a sentence, and the sentence is usually stronger than the sixty observations behind it.
Where the units should go
Given a total, the design has two knobs: how many arms, and where to put the interim. The measurements above constrain both.
More arms is better than it looks. The saving grows and the boundary grows, but the boundary grows slowly — from 2.031 at four arms to 2.313 at eight — while the saving grows nearly linearly. Screening eight candidates costs about a quarter of a standard error more evidence than screening four, which is a low price for four more chances.
And the interim should be early, for two reasons that come from different measurements. It costs less correction, from the table above, and it saves more units — dropping seven arms after thirty units each saves more than dropping them after seventy-five. Both point the same way, and the constraint on how early is the one this essay cannot measure: the interim has to be late enough that the selection picks the right arm, and how late that is depends on how different the arms actually are.
That last point is the honest limit of the arithmetic here. Everything measured in this essay is under a global null, where every arm is identical and no selection rule can pick correctly — which is the right setting for an error rate and the wrong setting for a design decision. A design that selects at thirty units and gets the wrong arm has saved 40% of the experiment and lost the thing it was for.
What the one-sided rates look like across k
Since the one-sided rate is the honest one, here it is beside the two-sided rate that this essay’s figures draw:
| arms | two-sided | one-sided |
|---|---|---|
| 1 | 5.3% | 5.1% |
| 2 | 4.8% | 8.3% |
| 3 | 6.0% | 10.5% |
| 4 | 7.0% | 12.3% |
| 8 | 10.3% | 17.8% |
The one-sided column is monotone and the two-sided one is not, which is the cancellation above showing up as a shape rather than a number. And the size of the one-sided inflation is worth sitting with: at four arms, a design that reports a one-sided p below 0.05 is running at 12.3%, which is not a correction anyone would describe as a technicality.
The correction for it is the same solved value, computed against the one-sided boundary instead. Nothing about the method changes; what changes is which quantity the boundary is solved to hold, and that is a decision the design has to state rather than inherit.
The two knobs, priced against each other
The essay ends on two design decisions — how many arms, and when to look — and the numbers already in it put the two on one scale.
Doubling the arms from four to eight at a fixed interim moves the boundary from 2.113 to 2.313: a cost of 0.200. Moving the interim from a third of the trial to five sixths at four arms moves it from 2.031 to 2.128: a cost of 0.097. An extra doubling of the candidate list costs about twice as much boundary as delaying the interim across the whole range this design admits.
That is the ratio worth carrying, because the two decisions are usually taken by different people for different reasons. The number of arms is decided by how many candidates exist; the interim’s timing is decided by logistics. The arithmetic says the first is the expensive knob and the second is not — so a design under pressure to move its interim later for operational reasons is giving up about half a doubling of its own screening capacity, which is a defensible trade and is not a free one.
The early interim is not a preference, it is a dominance
Working the unit counts out at four arms makes the timing decision sharper than “both point the same way”.
An all-arms design is five groups at ninety, which is 450 units. The two-stage design with the interim two thirds of the way through is 5 × 60 + 2 × 30 = 360, a saving of 20%, at a boundary of 2.113. With the interim at a third it is 5 × 30 + 2 × 60 = 270, a saving of 40%, at a boundary of 2.031. With it at five sixths it is 5 × 75 + 2 × 15 = 405, a saving of 10%, at a boundary of 2.128.
So across the three timings measured, the earliest interim saves four times as many units as the latest and requires the lowest boundary of the three. There is no axis on which the late interim wins: it is strictly dominated, in the same sense the derived weighting is dominated elsewhere on this site, and a design that chose it has traded nothing for nothing.
That is worth stating because the intuition runs the other way. Waiting longer before dropping arms feels like the cautious choice — more evidence behind the selection, less risk of discarding the right candidate — and on the error rate it is the reverse: a selection made on more of the trial is a larger fraction of the final statistic, so it needs a larger correction, and it leaves less of the budget to concentrate on the survivor.
The caution the intuition is reaching for is real and it is the one quantity none of this measures. An early interim picks the wrong arm more often, and every number in this essay is counted under a global null where there is no right arm to pick. The error rate and the units both say look early; the only thing that says look late is the one thing the global null cannot see, and a design has to supply that from what it believes about how far apart the candidates are.
What this design is and is not
Two boundaries are worth drawing, because the phrase “adaptive design” covers several unrelated things and this essay is about one of them.
It is not a stopping rule. Nothing here stops early; every trial runs to the same total. What is adapted is which arms the second stage spends on, and the correction is for the selection rather than for repeated testing. A design that also stopped early would need both, and the two do not simply add.
And it is not the multiplicity correction the arms would have needed anyway. Running all nine groups to the end and comparing each with the control would need Dunnett’s value — 2.652 at eight arms. This design needs 2.313, which is less, because the selection means only one comparison is made at the end and the second stage’s data was collected after the choice. The two-stage design is cheaper in units and cheaper at the boundary than the design it replaces.
That comparison is the strongest thing that can be said for it, and it is the one usually left out of the argument. The debate about these designs is normally framed as efficiency against rigour. Measured, this one is more efficient and holds a tighter boundary than the all-arms design — provided the boundary is computed for the design that was actually run, which is the entire condition.
The refusal, and the one that had to be added
The check requires the naive rate to exceed its claim and the solved value to hold it, on fresh seeds rather than the ones the value was solved on. That much is the ordinary shape.
The refusal is the k = 1 pin above: with a single arm the correction must vanish, and it does — 1.983 against 1.96, which is the solver’s own Monte Carlo error on twenty thousand paths.
The assertion that had to be added was about the shape of the dependence. The first version of this essay’s claim was “selecting the best of k arms inflates the error rate”, full stop, with a number attached from one design. Placed at a different interim time the same design gives 5.9% rather than 10.3%, and a reader given only the first number would take away a fact about arms when the fact is about arms and timing together. The check now requires the late-interim design to inflate more than the early one, so the dependence is asserted rather than incidentally true — and the standard error used for that comparison is computed at the rates observed rather than at the conservative √(0.25/n), because at rates near 6% the wide version is a band three times too big to see a real difference through.
The same problem produced the essay’s other correction, and it is the more instructive one. The figure that draws the null distribution originally asserted its own counted rate — that it exceed 5% by four standard errors — which is true of the design in the hero and false at an early interim, where 5.9% and eight thousand trials cannot be told apart from 5%. It now asserts the solved critical value instead, which is deterministic, cached, and computed on twenty thousand paths. Asserting the thing the machinery computes precisely rather than the thing it estimates noisily is the general form of that repair, and it is available more often than it is taken.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A boundary for giving up — both name error rate, interim analysis, sample size
- A simulation that stops when it looks settled — both name error rate, interim analysis, selection bias
- Allocating on a guess — both name experimental design, sample size, two-stage design
- An order that spends the error rate — both name error rate, familywise error rate, multiple comparisons
- Eight forecasters and one benchmark — both name error rate, familywise error rate, multiple comparisons
- False discoveries that arrive together — both name error rate, familywise error rate, multiple comparisons
Named objects
A flat tag is an object no other essay names yet.
Adaptive designCritical valueError rateExperimental designFamilywise error rateInterim analysisMultiple comparisonsSample sizeSelection biasTwo-stage design