Designs that change while they run

A second stage sized by the first

A drop-the-losers trial can size its second stage by what the first stage says about the arms' shape — more patients when the chosen arm's lead looks like luck, fewer when it looks real. The rule does what it was built to do: at sixteen arms with a genuine winner it spends 51.0 patients a group when the interim chose the wrong arm and 24.2 when it chose the right one. And it buys nothing. Against a fixed second stage of the same average size, every estimate of the chosen arm errs by the same to within a few thousandths; the published combination's bias rises from 0.014 to 0.022, and the second stage alone, still unbiased, errs by 0.319 against 0.281.

Worth reading first: Five times in six.

A prior that lets one arm stand out corrected the winner’s curse in a drop-the-losers trial with a mixture prior — the eight arms a bell with probability four fifths, one arm genuinely different with probability one fifth — and found that the prior’s posterior weight on “the chosen arm stands out” is the trial’s own verdict on whether its winner’s lead was real. It ended on a design that would act on that verdict: size the second stage by it, with more patients when the first stage cannot tell what it has found and fewer when it can, so the patients go where the curse is worst.

It asked whether that would keep the second stage unbiased, since the stage’s size would depend on the first stage’s data, and whether the saving at sixteen arms would be worth the complexity at eight. The rule turns out to be well aimed and to buy nothing, and why it buys nothing is a fact about what a second-stage patient is worth that holds for any rule of this kind.

The error of the chosen arm's estimate against a fixed second stage's size, and a second stage sized by the interim, 8 arms, one ahead by half a standard deviationFixed second stages of 10, 25, 40, 55, 70 patients: both stages combined err by 0.172, 0.155, 0.143, 0.133, 0.126; the mixture-shrunk estimate by 0.208, 0.183, 0.164, 0.151, 0.140. The rule that sizes the second stage by the interim spends 41.4 on average and errs by 0.142 and 0.159; its second stage alone errs by 0.225.0.1000.2000.3001025405570second-stage patients a grouproot-mean-square error of the chosen arm's estimatethe sized rule's average, 41.4both stages, by informationmixture-shrunk first stagesecond stage alone10,000 trials a point; 8 arms of 60lines: fixed sizes; points: the sized rule
Fig. 1 The root-mean-square error of three estimates of the chosen arm’s effect against the size of a fixed second stage, for a trial of eight arms with one ahead by half a standard deviation: both stages combined as published, the mixture-shrunk first stage combined with the second, and the second stage alone. The points are the rule that sizes the second stage at the interim, drawn at its own average size. The slider sets the arms’ true effects.

The trial and the rule

The trial is the one these essays have used throughout: eight arms of sixty patients and a shared control; at the interim the arm with the largest mean is chosen and carried into a second stage, and its effect is estimated. Here the second stage’s size is set at the interim, by the mixture’s posterior weight ww that the chosen arm stands out:

n2=round⁡(10+60(1−w)2),n_2 = \operatorname{round}\big(10 + 60(1 - w)^2\big),

seventy patients a group when the first stage has no evidence that the winner is different from the rest, ten when it is sure. The weight depends only on first-stage data — the chosen arm’s mean, the eight arms’ mean and their spread — so it is known at the interim, and the rule can be written into a protocol.

Four estimates of the chosen arm are compared. The published estimate combines both stages by the information each actually carried, so a larger second stage gets more weight. The prespecified estimate combines them with the weights a planned second stage of thirty would have had, whatever was run. The mixture estimate shrinks the first stage under the mixture prior before combining. And the second stage alone uses only the patients recruited after the choice. Each is computed on the same simulated trials, ten thousand a setting, with the second stage drawn at its largest size and truncated, so that a fixed rule and a sized one see the same patients in the same order.

Where the rule spends

How many second-stage patients the sized rule spends, by the number of arms, the world, and whether the interim chose the real winner. Average second stage per group: 8 arms, all alike, 60.2; 8 arms, winner chosen, 41.1; 8 arms, another arm chosen, 58.0; 16 arms, all alike, 55.1; 16 arms, winner chosen, 24.2; 16 arms, another arm chosen, 51.0. With one arm ahead by half a standard deviation, the interim chose it in 98.2% of trials at eight arms and 96.5% at sixteen.
Fig. 2 The average second stage the rule runs, in patients a group, at eight and sixteen arms: when all arms are alike, and when one is ahead by half a standard deviation, split by whether the interim chose that arm or another.

The rule behaves as designed. When all eight arms are alike the weight averages 0.086 and the second stage averages 60.2 patients a group; with one arm ahead by half a standard deviation the weight averages 0.284 and the stage 41.4. Within that second world it spends 41.1 patients when the interim chose the real winner and 58.0 when it chose another arm, which is exactly the targeting the design was meant to achieve: a chosen arm whose lead was luck gets more patients to expose it.

At sixteen arms the targeting is sharper, because sixteen arms describe their bell well enough for an outlier to be seen, as the earlier essay found — and as the fewest groups that can borrow found for groups estimating their own spread, eight is near the smallest number that says anything at all. With all sixteen alike the stage averages 55.1; with a real winner, 25.1, and it splits into 24.2 when the winner was chosen and 51.0 when it was not. Whatever else is true of the rule, it knows roughly which trials need the patients.

What the patients buy

The hero figure’s lines are the error of each estimate at a fixed second stage, from ten to seventy patients a group, and the point on each line is the sized rule at its own average size. At eight arms with a real winner the published estimate errs by 0.1422 under the sized rule and 0.1420 with a fixed stage of 41, the sized rule’s average. With all eight alike, 0.1537 against 0.1513 at a fixed 60. At sixteen arms, 0.1576 against 0.1568 with a winner and 0.1754 against 0.1703 without. In every case the sized rule’s point sits on or a little above the fixed line: it spends its patients differently from a fixed rule and ends where a fixed rule of the same average spend ends.

The mixture estimate shows the one small gain the rule produces. With a real winner it errs by 0.1591 against 0.1631 at eight arms and 0.1768 against 0.1842 at sixteen, because the trials the rule gives extra patients are the ones where the interim chose a loser, and those are the trials in which the mixture estimate is furthest from the truth. When all arms are alike it gives the gain back — 0.1213 against 0.1204, 0.1257 against 0.1235 — because then there is no winner to have missed and the extra patients are spread where they help least.

The reason the rule cannot do better is visible once the trials are split by whether the interim chose the real winner.

The published estimate's error within trials that chose the real winner and trials that did not, under a fixed second stage and the sized rule. One arm ahead by half a standard deviation. At 8 arms the interim chose it in 98.2% of trials; there the published estimate errs by 0.1397 with a fixed stage of 41 and 0.1406 sized, and in the rest by 0.2370 and 0.2117. At 16 arms the interim chose it in 96.5% of trials; there the published estimate errs by 0.1513 with a fixed stage of 25 and 0.1544 sized, and in the rest by 0.2675 and 0.2298.
Fig. 3 The published estimate’s root-mean-square error within the trials whose interim chose the arm that was genuinely ahead and within the trials that chose another, at eight and sixteen arms with one arm ahead by half a standard deviation, under a fixed second stage of the sized rule’s average size and under the sized rule itself. The shares in brackets are how often each kind of trial happened.

The rule is aimed at the trials whose interim chose a loser, and it helps them: at eight arms their published estimate errs by 0.2117 instead of 0.2370, and at sixteen by 0.2298 instead of 0.2675. But those trials are rare. An arm ahead by half a standard deviation leads the others by nearly four standard errors of an arm mean, and the interim chooses it in 98.2% of trials at eight arms and 96.5% at sixteen. The rule’s rescue touches the other 1.8% and 3.5%.

The patients it spends on them come from the trials that chose well, and the rule takes them exactly when those trials’ leads were largest — and a large lead is partly luck even when the arm is real. In the trials that chose the winner the published estimate errs by 0.1406 under the sized rule against 0.1397 fixed at eight arms, and 0.1544 against 0.1513 at sixteen: the undiluted luck of the right choices is the bias the next section measures. A stage whose size varies also costs precision on its own, since an average of the inverse sizes is larger than the inverse of the average size.

So a large gain on a few trials is paid for with a small loss on nearly all of them, and at eight and at sixteen arms, with a winner and without, the net is within a few thousandths of zero. The expected error is set by the average spend. The winner’s curse in a trial that has a real winner is mostly not a wrong choice at all; it is a right choice made with a lead that luck inflated, and the rule, by shortening the second stage on exactly those trials, feeds it.

The bias the sizing adds

The lead asked whether a second stage sized by the first would stay unbiased. It depends on which estimate is meant, and the answer is different for each.

What sizing the second stage by the interim does to each estimate's bias, 16 arms, one ahead by half a standard deviation. Bias of the chosen arm's estimate: fixed second stage, by information +0.014 (error 0.157); sized, combined by information +0.022 (error 0.158); sized, combined with planned weights +0.012 (error 0.165); sized, second stage alone +0.000 (error 0.319). The fixed stage has 25 patients a group, the sized rule's average.
Fig. 4 The bias of four estimates of the chosen arm at sixteen arms with one ahead by half a standard deviation: a fixed second stage of the sized rule’s average size, combined by information; and the sized rule’s second stage combined by information, combined with the weights planned for a fixed stage, and used alone.

The second stage alone stays exactly unbiased. Its size depends on the first stage, but its patients do not, and the mean of nn independent patients is unbiased for any nn however it was chosen. Measured, its bias is +0.000 at sixteen arms, and within a few thousandths in every setting. What it loses is precision: a stage whose size varies from ten to seventy has more error than one fixed at the same average, because an average of the inverse sizes is larger than the inverse of the average size. At sixteen arms with a winner, 0.3192 against 0.2814.

The published combination gains a bias. Its weight on the first stage is n1/(n1+n2)n_1/(n_1 + n_2), and n2n_2 is now small precisely when the first stage produced a large lead. A large lead is partly a large first-stage error, so the combination leans on the first stage most when the first stage is most flattering. At sixteen arms with a real winner the published estimate overstates by +0.022 under the sized rule against +0.014 with a fixed stage of the same average; at eight arms, +0.0098 against +0.0053. The overstatement is small beside the winner’s curse the design exists to correct, and it is new: the fixed design did not have it.

Weights fixed in advance remove it and cost precision. Combining the stages with the weights a planned second stage of thirty would have had breaks the link between the weight and the first stage’s luck, and the bias at sixteen arms returns to +0.012. But a second stage of seventy is then given the weight of thirty, and the estimate’s error rises to 0.1654 against the published combination’s 0.1576. This is the estimation version of the trade choosing n after looking described for testing: a rule that changes the sample size at an interim keeps its error rate only by committing in advance to how the stages will be weighed, and the commitment wastes whatever the extra patients carried.

Why sizing after the choice cannot touch the choice

There is a simpler way to see that no rule of this shape can do much, and it is the part of the winner’s curse the second stage never reaches. The curse is a property of the choice, and the choice is made at the interim from first-stage data alone. With one arm ahead by half a standard deviation among eight, the interim chooses that arm in a fixed share of trials, and nothing done afterwards changes that share. The second stage can only measure the arm that was chosen more precisely; it cannot make a wrong choice right.

So the only thing the sizing rule can move is how precisely the chosen arm is measured, and the estimate after the choice already established what precision can do for an estimate whose first stage is contaminated by selection: the second stage alone is unbiased and noisy, the combination is precise and biased, and more second-stage patients slide the combination towards the second stage alone. That slide happens with any extra patients, sized or fixed. A design that wants to reduce the curse rather than dilute it has to spend patients before the choice — more patients in each arm’s first stage, which sharpens the choice itself — and dropping the losers and the essays after it measured what that buys.

Eight arms or sixteen

The lead expected the rule to be worth its complexity at sixteen arms and not at eight, on the grounds that sixteen arms tell their shape and eight do not. The first half of that is right — the rule’s targeting is much sharper at sixteen, 24.2 against 51.0 patients by whether the choice was right, against 41.1 and 58.0 at eight — and the second half follows, but not for the expected reason. The rule is not worth its complexity at either size, because sharper targeting of the second stage buys nothing when a second-stage patient is worth the same to every trial. At sixteen arms it is slightly worse than a fixed stage, since its sizes vary more and the variability costs precision.

What sixteen arms do change is where the patients go on average, and that is a genuine operational difference if not a statistical one. A program running many sixteen-arm screens with a mixture of real winners and none would, under the rule, spend 55 patients a group on screens with nothing to find and 25 on screens with a winner — an allocation some sponsors would prefer for reasons of their own, such as getting a real winner into a confirmatory trial sooner. The measurements here say only that the estimate of each screen’s winner is no better for it.

What the measurements support

A second stage sized by the interim’s verdict on the arms spends its patients where the winner’s curse is worst, and buys no accuracy that a fixed stage of the same average size does not. At eight and at sixteen arms, in a world with a winner and one without, both combined estimates err within about seven thousandths of the fixed-stage value, in either direction.

The second stage alone stays unbiased under any sizing rule that depends only on the first stage, and loses precision to the variability of its size — 0.3192 against 0.2814 at sixteen arms.

Combining the stages by the information collected adds a bias under sizing, +0.022 against +0.014 at sixteen arms with a winner, because the second stage is short exactly when the first stage flattered; combining them with planned weights removes the bias and costs 0.008 of error.

The mixture-shrunk estimate is the only one the rule helps — the shrinkage that pulls the chosen arm towards the others is furthest wrong on the wrong choices the rule rescues —, by a few thousandths when a winner is real, and it gives the gain back when none is.

The interim’s verdict is worth reporting, not acting on. The rule sizes its second stage on effect information — how far the chosen arm stands out — and that is what lets it introduce a bias at all. Choosing n after looking and blinded, and still exact found the one adaptation that is safe to be the one that reads only a nuisance quantity, a spread with no effect in it; a stand-out weight is the opposite, a direct reading of the effect’s shape. Published beside the chosen arm’s estimate, the same weight tells a reader how much of the winner’s curse to expect — low weight, a lead that may be luck — at no cost to any estimate. Written into the protocol as a sizing rule, it spends patients on the rare trials that chose wrongly and shortens the common ones that chose rightly but luckily, and the two cancel.

Each comparison is between ten thousand trials under the sized rule and the same ten thousand trials with a fixed second stage of the rule’s average size, rounded, so the two differ only in how the second stage was sized. The standard error of a root-mean-square error from ten thousand trials is about a hundredth of its value, so differences of a few thousandths are at the edge of what the comparison can resolve, and the conclusion rests on their being small in every setting, not on any one of them. Not measured: rules that size by something other than the mixture weight, such as the gap between the chosen arm and the runner-up; a rule that also spends patients before the choice; and arms whose effects differ by more than half a standard deviation, where the choice is nearly always right and the question hardly arises.

Still open: patients spent before the choice

Every rule here spends its adaptivity after the choice, where it cannot touch the choice. The design question the result points at is the other one: with a fixed total budget, how many patients should go into the first stage, sharpening which arm is chosen, and how many into the second, sharpening the estimate of the arm that was chosen. More first-stage patients reduce the curse at its source, by making a lucky choice less likely, and cost second-stage precision; fewer do the reverse.

At eight arms of sixty and a second stage of thirty, the trial spends 540 patients before the choice and 60 after. Whether moving some of the first stage’s patients to the second, or the reverse, reduces the chosen arm’s error, and whether the answer changes with how many arms have a real effect, is a measurement of the same kind as these, and it would say whether the familiar two-stage shape of a selection trial puts its patients in the right place for the estimate it publishes.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive designEmpirical BayesMean squared errorSample sizeSelection biasShrinkageTwo-stage designThe winner's curse