Designs that change while they run

Dropping the losers

Carrying the best of eight arms forward and testing it at 1.96 rejects a true null 10.3% of the time — the hypothesis was chosen by looking at the data, so the statistic is a maximum wearing a single comparison's clothes. The value that holds the rate is 2.313, and it has to be solved for.

Worth reading first: When the looking happens · What the correction corrects.

Eight candidate treatments and one control. Running all nine groups to the end is the honest design and it is enormously wasteful: most of those arms are not going to work, and the units spent on them after that is apparent are units spent learning nothing.

So run every arm for a while, drop all but the best, and put the remaining budget into that one and the control. It is the right instinct, it saves a quarter of the experiment, and the test at the end is a test of a hypothesis chosen by looking at the data.

The best of 8 arms, tested as though it were the only one8,000 trials with no effect in any arm. Stage one runs 8 arms at 60 each, the best is carried forward, and stage two adds 30 more to it and to the control. The histogram is where the final statistic lands and the curve is the standard normal it is being read against — shifted right, because the arm was chosen for being ahead. 10.5% of these trials clear 1.96 against a claimed 5%, and the value that actually holds the rate for this design is 2.313.0200400600-4-2024the final statistic, read as a ztrials out of 8,0001.96, from the table2.31, solved for this design8,000 trials, 8 arms at 60, then 30 more10.5% beyond 1.96
Fig. 1 Twenty thousand trials with no effect in any arm. The histogram is where the final statistic lands and the curve is the standard normal it is being read against.

What the naive analysis does

The design: eight arms and a control at sixty each, the arm with the largest observed difference carried forward, thirty more on it and on the control, and a final comparison using all the data on the two surviving groups.

That final comparison looks like a two-sample test. It has two groups, ninety observations each, an estimate and a standard error. Compare it with 1.96 and, under a global null where nothing is real, it fires 10.3% of the time.

The reason is visible in the histogram: the distribution is shifted right. The arm being tested is the one that came out ahead of seven others at the interim, so its first-stage contribution is a maximum rather than a draw — and the final statistic carries that maximum forward inside it, diluted by the second stage but not removed.

This is the sequential field’s principle in a new place. A p-value is defined relative to a sampling plan, and the plan here includes “and then the hypothesis was chosen”. The set of outcomes that could have produced this statistic is not the set 1.96 was computed for.

The value that does hold the rate

There is no table for the right value, and the reason is worth being precise about: the null distribution depends on k, on n₁ and on n₂ together, because the selection is made on a fraction of the data that then reappears inside the final estimate. Change how much of the trial happens before the interim and the whole distribution moves.

So it is solved for, the way this site solves the Pocock boundary: draw one set of null paths against a fixed seed, bisect on the boundary over them, and cache the answer. For this design it is 2.313, against 1.96, and the counted rate at that value is 5.01% on fresh seeds.

Two pins make that believable rather than merely computed. With k = 1 there is no selection and the solver must return the ordinary value: it gives 1.983, and the naive rate at k = 1 is 5.3% — the design is essentially unadapted, as it should be. With eight arms it gives 2.313. A solver that returned something above 1.96 when nothing had been selected would be manufacturing a correction out of its own machinery, and every number in this essay would be that artefact.

The best of 1 arms, tested as though it were the only one. 8,000 trials with no effect in any arm. Stage one runs 1 arms at 60 each, the best is carried forward, and stage two adds 30 more to it and to the control. The histogram is where the final statistic lands and the curve is the standard normal it is being read against — shifted right, because the arm was chosen for being ahead. 5.4% of these trials clear 1.96 against a claimed 5%, and the value that actually holds the rate for this design is 1.983.
Fig. 2 The same two-stage design with one arm, where there is nothing to select. The histogram sits on the standard normal and the solved value is 1.983 — the refusal that makes the 2.313 mean something.

The correction depends on when, not just how many

Here is the part that makes this different from an ordinary multiplicity correction.

Bonferroni’s value depends on how many comparisons there were. Dunnett’s, from the allocation field, depends on that and on their correlation. This one depends on how much of the trial the selection saw, and the dependence is not small:

interim at share of the trial naive rate critical value
30 of 90 a third 5.9% 2.031
60 of 90 two thirds 7.0% 2.113
75 of 90 five sixths 7.3% 2.128

All three rows are four arms. What changes is when the choice was made, and selecting on a third of the data and then diluting it with two thirds collected afterwards produces a much smaller inflation than selecting on five sixths.

That is the same arithmetic as the estimate’s bias and it is worth carrying as a design principle: an early interim is a cheap interim. It costs less error rate to correct for, because the selection is made on less information and the stage that follows it is large enough to swamp what the selection did.

Two arms, where the rate looks fine and both tails are wrong

The clearest thing this design taught was found by placing the figure above at two arms, which was meant to be an unremarkable intermediate case.

The two-sided rate comes out at 4.79%below its nominal 5% — on a design whose statistic has a mean of +0.33 under the null. It is not calibrated; it is two errors cancelling.

tail nominal counted, two arms
above +1.96 2.5% 4.11%
below −1.96 2.5% 0.69%
either 5.0% 4.79%

Selection moves mass into the upper tail and out of the lower one. At two arms the second effect is slightly the larger, the two nearly cancel, and a check that looked only at the two-sided rate would report the design as safe while its upper tail was 64% too heavy and its lower tail a quarter of what it should be.

And these designs are tested one-sided in practice — the question is whether the surviving arm beats the control, not whether it differs from it. The one-sided rate at two arms is 8.34%, against a nominal 5%. At eight arms it is 17.79% where the two-sided rate is 10.32%.

So the honest headline for this design is the one-sided number, and the two-sided one is included here because of what it demonstrates: a rate that matches its claim is not evidence of calibration when it is the sum of two errors. This site has made that argument before about procedures controlling different quantities at the same nominal level; this is the same failure inside a single number.

The best of 2 arms, tested as though it were the only one. 8,000 trials with no effect in any arm. Stage one runs 2 arms at 60 each, the best is carried forward, and stage two adds 30 more to it and to the control. The histogram is where the final statistic lands and the curve is the standard normal it is being read against — shifted right, because the arm was chosen for being ahead. 4.7% of these trials clear 1.96 against a claimed 5%, and the value that actually holds the rate for this design is 1.947.
Fig. 3 Two arms, where the shift is visible and the two-sided rate is 4.79%. The histogram’s right tail is heavy, its left tail is thin, and the two nearly cancel in any statistic that adds them.

What it saves

The other side of the ledger, and the reason anybody runs this design.

Eight arms and a control at sixty each, then thirty more on two groups, is 600 units. Running all nine groups to ninety is 810. The selection saves 25.9% of the experiment, and it saves it on exactly the arms that were not going to be the answer.

What dropping arms saves, and what it costs at the boundary. Two quantities against the number of arms. The units saved against running every arm to the end rise from 11% at 2 arms to 28% at 12, because the arms that are dropped are not paid for in stage two. The critical value rises too, from 1.95 to 2.42, because the more arms were available the further ahead the winner is expected to be. A design that counts the first and ignores the second has kept the saving and spent the error rate.
Fig. 4 Both sides of the trade against the number of arms: the units saved by dropping, and the critical value the survivor then has to clear.

Both curves rise with k, which is the whole design problem in one picture. More arms means more saved — there is more to drop — and a higher boundary, because the winner of a larger tournament is expected to be further ahead. A design that counts the first and forgets the second has kept the saving and spent the error rate.

Why the saving is not the whole of the efficiency

The 25.9% is units, and units are not the only thing a two-stage design saves.

Time. Dropping seven arms at the interim frees the sites, the staff and the supply chain that were running them, and a trial that finishes sooner is worth more than the same trial finishing later. None of that is in the arithmetic here and all of it is in the decision.

And the arms that were never worth running. The design’s real economy is that it converts a decision made badly in advance — which of eight candidates deserves a full trial — into a decision made with data. An all-arms design commits to eight full-sized comparisons before anything is known; this one commits to eight small ones and then to a single large one.

Against that, one cost that is not units either: the seven dropped arms are dropped on very little evidence. Sixty observations each is enough to rank them and not nearly enough to establish that any of them is inferior. An arm dropped at the interim has not been shown not to work — it has been shown to be behind, on a comparison that is itself noisy — and reporting it as a negative result would be a much stronger claim than the design supports.

That is worth stating because it is how these trials are read afterwards. The surviving arm gets a p-value and an estimate. The seven others get a sentence, and the sentence is usually stronger than the sixty observations behind it.

Where the units should go

Given a total, the design has two knobs: how many arms, and where to put the interim. The measurements above constrain both.

More arms is better than it looks. The saving grows and the boundary grows, but the boundary grows slowly — from 2.031 at four arms to 2.313 at eight — while the saving grows nearly linearly. Screening eight candidates costs about a quarter of a standard error more evidence than screening four, which is a low price for four more chances.

And the interim should be early, for two reasons that come from different measurements. It costs less correction, from the table above, and it saves more units — dropping seven arms after thirty units each saves more than dropping them after seventy-five. Both point the same way, and the constraint on how early is the one this essay cannot measure: the interim has to be late enough that the selection picks the right arm, and how late that is depends on how different the arms actually are.

That last point is the honest limit of the arithmetic here. Everything measured in this essay is under a global null, where every arm is identical and no selection rule can pick correctly — which is the right setting for an error rate and the wrong setting for a design decision. A design that selects at thirty units and gets the wrong arm has saved 40% of the experiment and lost the thing it was for.

The best of 4 arms, tested as though it were the only one. 8,000 trials with no effect in any arm. Stage one runs 4 arms at 30 each, the best is carried forward, and stage two adds 60 more to it and to the control. The histogram is where the final statistic lands and the curve is the standard normal it is being read against — shifted right, because the arm was chosen for being ahead. 5.9% of these trials clear 1.96 against a claimed 5%, and the value that actually holds the rate for this design is 2.031.
Fig. 5 Four arms selected on a third of the trial. The shift is much smaller — the naive rate is 5.9% — and the value that holds the rate is 2.031.

What the one-sided rates look like across k

Since the one-sided rate is the honest one, here it is beside the two-sided rate that this essay’s figures draw:

arms two-sided one-sided
1 5.3% 5.1%
2 4.8% 8.3%
3 6.0% 10.5%
4 7.0% 12.3%
8 10.3% 17.8%

The one-sided column is monotone and the two-sided one is not, which is the cancellation above showing up as a shape rather than a number. And the size of the one-sided inflation is worth sitting with: at four arms, a design that reports a one-sided p below 0.05 is running at 12.3%, which is not a correction anyone would describe as a technicality.

The correction for it is the same solved value, computed against the one-sided boundary instead. Nothing about the method changes; what changes is which quantity the boundary is solved to hold, and that is a decision the design has to state rather than inherit.

Four boundaries for 5 looks, all spending 5% in total. test at 0.05 every look: 1.96, 1.96, 1.96, 1.96, 1.96. Pocock — a constant, higher boundary: 2.41, 2.41, 2.41, 2.41, 2.41. O'Brien–Fleming — strict early, nearly nominal at the end: 4.55, 3.22, 2.63, 2.27, 2.03. Bonferroni across looks: 2.58, 2.58, 2.58, 2.58, 2.58. Every one except the first spends the same total error rate; they differ in when they spend it.
Fig. 6 The sequential field’s boundaries, solved for rather than quoted. The critical value in this essay is solved the same way and for the same reason: the design decides the null distribution, and no table is indexed by the design.

The two knobs, priced against each other

The essay ends on two design decisions — how many arms, and when to look — and the numbers already in it put the two on one scale.

Doubling the arms from four to eight at a fixed interim moves the boundary from 2.113 to 2.313: a cost of 0.200. Moving the interim from a third of the trial to five sixths at four arms moves it from 2.031 to 2.128: a cost of 0.097. An extra doubling of the candidate list costs about twice as much boundary as delaying the interim across the whole range this design admits.

That is the ratio worth carrying, because the two decisions are usually taken by different people for different reasons. The number of arms is decided by how many candidates exist; the interim’s timing is decided by logistics. The arithmetic says the first is the expensive knob and the second is not — so a design under pressure to move its interim later for operational reasons is giving up about half a doubling of its own screening capacity, which is a defensible trade and is not a free one.

The early interim is not a preference, it is a dominance

Working the unit counts out at four arms makes the timing decision sharper than “both point the same way”.

An all-arms design is five groups at ninety, which is 450 units. The two-stage design with the interim two thirds of the way through is 5 × 60 + 2 × 30 = 360, a saving of 20%, at a boundary of 2.113. With the interim at a third it is 5 × 30 + 2 × 60 = 270, a saving of 40%, at a boundary of 2.031. With it at five sixths it is 5 × 75 + 2 × 15 = 405, a saving of 10%, at a boundary of 2.128.

So across the three timings measured, the earliest interim saves four times as many units as the latest and requires the lowest boundary of the three. There is no axis on which the late interim wins: it is strictly dominated, in the same sense the derived weighting is dominated elsewhere on this site, and a design that chose it has traded nothing for nothing.

That is worth stating because the intuition runs the other way. Waiting longer before dropping arms feels like the cautious choice — more evidence behind the selection, less risk of discarding the right candidate — and on the error rate it is the reverse: a selection made on more of the trial is a larger fraction of the final statistic, so it needs a larger correction, and it leaves less of the budget to concentrate on the survivor.

The caution the intuition is reaching for is real and it is the one quantity none of this measures. An early interim picks the wrong arm more often, and every number in this essay is counted under a global null where there is no right arm to pick. The error rate and the units both say look early; the only thing that says look late is the one thing the global null cannot see, and a design has to supply that from what it believes about how far apart the candidates are.

What this design is and is not

Two boundaries are worth drawing, because the phrase “adaptive design” covers several unrelated things and this essay is about one of them.

It is not a stopping rule. Nothing here stops early; every trial runs to the same total. What is adapted is which arms the second stage spends on, and the correction is for the selection rather than for repeated testing. A design that also stopped early would need both, and the two do not simply add.

And it is not the multiplicity correction the arms would have needed anyway. Running all nine groups to the end and comparing each with the control would need Dunnett’s value — 2.652 at eight arms. This design needs 2.313, which is less, because the selection means only one comparison is made at the end and the second stage’s data was collected after the choice. The two-stage design is cheaper in units and cheaper at the boundary than the design it replaces.

That comparison is the strongest thing that can be said for it, and it is the one usually left out of the argument. The debate about these designs is normally framed as efficiency against rigour. Measured, this one is more efficient and holds a tighter boundary than the all-arms design — provided the boundary is computed for the design that was actually run, which is the entire condition.

What dropping arms saves, and what it costs at the boundary. Two quantities against the number of arms. The units saved against running every arm to the end rise from 22% at 2 arms to 56% at 12, because the arms that are dropped are not paid for in stage two. The critical value rises too, from 1.95 to 2.22, because the more arms were available the further ahead the winner is expected to be. A design that counts the first and ignores the second has kept the saving and spent the error rate.
Fig. 7 The same trade with the interim at a third of the trial rather than two thirds. Everything saved is larger and every boundary is lower, which is the case for an early interim in one figure.
The best of 12 arms, tested as though it were the only one. 8,000 trials with no effect in any arm. Stage one runs 12 arms at 60 each, the best is carried forward, and stage two adds 30 more to it and to the control. The histogram is where the final statistic lands and the curve is the standard normal it is being read against — shifted right, because the arm was chosen for being ahead. 12.7% of these trials clear 1.96 against a claimed 5%, and the value that actually holds the rate for this design is 2.415.
Fig. 8 Twelve arms, where the shift is largest and the boundary that holds the rate is furthest from the table. The design saves most here and needs the most correction, which are the same fact.

The refusal, and the one that had to be added

The check requires the naive rate to exceed its claim and the solved value to hold it, on fresh seeds rather than the ones the value was solved on. That much is the ordinary shape.

The refusal is the k = 1 pin above: with a single arm the correction must vanish, and it does — 1.983 against 1.96, which is the solver’s own Monte Carlo error on twenty thousand paths.

The assertion that had to be added was about the shape of the dependence. The first version of this essay’s claim was “selecting the best of k arms inflates the error rate”, full stop, with a number attached from one design. Placed at a different interim time the same design gives 5.9% rather than 10.3%, and a reader given only the first number would take away a fact about arms when the fact is about arms and timing together. The check now requires the late-interim design to inflate more than the early one, so the dependence is asserted rather than incidentally true — and the standard error used for that comparison is computed at the rates observed rather than at the conservative √(0.25/n), because at rates near 6% the wide version is a band three times too big to see a real difference through.

The same problem produced the essay’s other correction, and it is the more instructive one. The figure that draws the null distribution originally asserted its own counted rate — that it exceed 5% by four standard errors — which is true of the design in the hero and false at an early interim, where 5.9% and eight thousand trials cannot be told apart from 5%. It now asserts the solved critical value instead, which is deterministic, cached, and computed on twenty thousand paths. Asserting the thing the machinery computes precisely rather than the thing it estimates noisily is the general form of that repair, and it is available more often than it is taken.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive designCritical valueError rateExperimental designFamilywise error rateInterim analysisMultiple comparisonsSample sizeSelection biasTwo-stage design