Designs that change while they run

Patients spent before the choice

A trial that picks the best of eight arms and then confirms it splits its patients between choosing and estimating, and the familiar shape — sixty an arm before the choice, thirty after — spends ninety per cent of a six-hundred-patient budget on the choice. Measured against the effect a reader is told about, the best arm's, that split is nearly right when one arm is modestly ahead and twice as wrong as it needs to be when the arms are alike. A split of thirty-two an arm before the choice is never more than 21.3% worse than the best split for whichever world holds; sixty an arm can be 98.8% worse. The two targets an estimate could serve — the chosen arm's effect and the best arm's — pull the split in opposite directions, and only one of them is what gets published.

Worth reading first: The winner's curse · What a p-value does not say.

A second stage sized by the first let a drop-the-losers trial set its second stage’s size from what the first stage said about the arms, and found that the patients it added bought nothing a fixed second stage of the same average size did not. The reason was where they were spent: after the choice, where they can sharpen the estimate of the arm that was chosen but cannot change which arm that is. Every adaptive rule in that series spent its adaptivity in the same place.

The essay ended on the decision that comes before any of them. With a fixed total, how many patients should go into the first stage, sharpening which arm is chosen, and how many into the second, sharpening the estimate of the arm that was chosen? At eight arms of sixty and a second stage of thirty, a trial spends 540 patients before the choice and 60 after. That shape is familiar enough to look like a recommendation. Whether it is one depends on what the trial is for, and on a fact about the arms nobody knows in advance.

How far the published estimate of the chosen arm sits from the best arm's effect, by how six hundred patients are split between choosing and estimatingEight arms and a control, six hundred patients: a stated number an arm in the first stage and the rest split equally between the chosen arm and the control in the second. Root mean squared error of the published estimate about the best arm's true effect, over 5,000 trials a split. All eight alike: best at 16 an arm (0.092), 0.183 at the familiar sixty; One modestly ahead: best at 56 an arm (0.151), 0.153 at the familiar sixty; One clearly ahead: best at 48 an arm (0.146), 0.154 at the familiar sixty; Two clearly ahead: best at 40 an arm (0.128), 0.151 at the familiar sixty; Spread evenly: best at 40 an arm (0.131), 0.146 at the familiar sixty. The split that guards every world puts 32 an arm before the choice and costs at most 21.3% over each world's best.0.00.10.20.31624324048566064patients an arm in the first stage (the rest of six hundred go to the second)error about the best arm's true effectsixty and thirtyminimaxall eight alikeone modestly aheadone clearly aheadtwo clearly aheadspread evenly5,000 trials a split and worldsix hundred patients however they are split
Fig. 1 The published estimate’s error about the best arm’s true effect, against how many patients an arm the first stage takes out of a fixed six hundred, in five configurations of the eight arms’ effects. Dashed lines mark the familiar sixty an arm and the split that guards every configuration. The slider switches the estimate.

Six hundred patients, two uses

The trial is the one the earlier essays ran: eight arms and a control, each first-stage group of n1n_1 patients, the arm with the best first-stage mean chosen, and a second stage of n2n_2 patients on that arm and n2n_2 more on the control. The total is held at six hundred, 9n1+2n2=6009n_1 + 2n_2 = 600, so every patient moved out of the first stage buys four and a half for the second: sixteen an arm before the choice leaves 228 a group after it, sixty leaves thirty, sixty-four leaves twelve. Outcomes are normal with unit variance, and every split is run five thousand times in each world.

The worlds are the five the series has used. All eight arms alike. One arm modestly ahead, by 0.25. One arm clearly ahead, by 0.5. Two arms clearly ahead, by 0.4. And the eight spread evenly from zero to 0.5. They stand for the things a selection trial might be facing, and the trial does not know which.

Three estimates of the chosen arm are scored, all from the earlier essays. The published estimate pools the chosen arm’s two stages by their sizes and subtracts the pooled control, which is what a report prints. The mixture-shrunk estimate shrinks the first-stage mean towards the other arms under a prior that lets one arm stand out, then pools. The second stage alone discards the first stage and is unbiased whatever happened at the choice.

The target the published number is read against

An estimate of the chosen arm can be wrong in two ways a reader would care about. It can misstate the chosen arm’s own effect, which is the winner’s curse every essay in the series measured. And it can misstate the effect a reader takes it to be about — the best treatment’s — because the chosen arm is not always the best. A report that says “the best of eight arms improved the outcome by 0.4” is read as a statement about the best arm, and its error against that is the sum of the two: the estimate’s error about the chosen arm, and the effect the choice gave up.

The hero figure scores the published estimate against the best arm. With the arms alike, every arm is the best, the choice cannot be wrong, and the error is purely the estimate’s: 0.092 with sixteen an arm before the choice, rising steadily to 0.183 at the familiar sixty. Moving patients before the choice buys nothing when there is nothing to choose, and every patient moved there is taken from the estimate.

With one arm clearly ahead the curve has a minimum. At sixteen an arm the error about the best arm is 0.269, because the choice is often wrong; at forty-eight it is 0.146, its lowest; at sixty, 0.154. With one modestly ahead the best split is fifty-six an arm and sixty is within 1.5% of it. With two clearly ahead, or the eight spread evenly, the best is forty an arm, and sixty costs 17.4% and 10.8% more.

What the first stage buys

The minimum exists because the first stage buys something the second cannot: the right arm.

How often the first stage chooses the best of eight arms, by how many patients it takes from a fixed six hundred. One modestly ahead: 39%, 45%, 52%, 58%, 62%, 67%, 68%, 70%; One clearly ahead: 71%, 82%, 89%, 94%, 96%, 98%, 98%, 98%; Two clearly ahead: 81%, 88%, 93%, 96%, 97%, 99%, 99%, 99%; Spread evenly: 36%, 42%, 44%, 49%, 51%, 54%, 53%, 55% — at 16, 24, 32, 40, 48, 56, 60, 64 patients an arm. The effect given up against the best arm, on average: one modestly ahead, 0.153 at sixteen an arm and 0.081 at sixty; one clearly ahead, 0.144 at sixteen an arm and 0.009 at sixty; two clearly ahead, 0.078 at sixteen an arm and 0.006 at sixty; spread evenly, 0.108 at sixteen an arm and 0.055 at sixty.
Fig. 2 How often the first stage chooses the best of the eight arms, against how many patients an arm it takes from the six hundred, in the four configurations that have a best arm. 5,000 trials a split.

With one arm clearly ahead, sixteen patients an arm choose it 71% of the time, thirty-two choose it 89%, forty-eight 96% and sixty 98%. With two clearly ahead, 81% at sixteen and 99% at sixty. With one modestly ahead the choice never becomes reliable — 39% at sixteen an arm and 68% at sixty — because an effect of 0.25 is small against eight chances for luck. The arithmetic is short. With sixteen patients an arm, each arm’s first-stage mean has a standard error of 0.25 — the size of the modest arm’s whole advantage — and seven null arms each get a draw at overtaking it. At sixty an arm the standard error is 0.129, half the advantage, and seven draws still overtake it about one time in three. An advantage has to be several standard errors wide before eight-way selection stops being a lottery, and a quarter of a standard deviation is not, at any first stage this budget can afford. And with the arms spread evenly the choice of the very best is a coin toss at any size, 36% to 55%, though the arm chosen is usually one of the good ones.

What the choice gives up is the other side of the same numbers. With one arm clearly ahead, the chosen arm’s effect falls short of the best by 0.144 on average at sixteen an arm and by 0.009 at sixty. That is most of the error at the small first stages: the estimate itself is precise there, but it is precise about the wrong arm a quarter of the time. The returns flatten quickly. Past forty-eight an arm the clear winner is found 96% of the time or more, and the patients still being spent on the choice would do more good estimating it.

Two targets, one budget

The two halves of the error do not just add; they pull the split in opposite directions.

Two errors of one estimate, by how the budget is split, one clearly ahead. The published estimate's root mean squared error about the chosen arm's own effect: 0.091, 0.097, 0.103, 0.113, 0.125, 0.139, 0.152, 0.165; about the best arm's effect: 0.269, 0.214, 0.177, 0.154, 0.146, 0.147, 0.154, 0.164 — at 16, 24, 32, 40, 48, 56, 60, 64 patients an arm before the choice. The first falls the more patients go after the choice; the second is smallest at 48 an arm.
Fig. 3 With one arm clearly ahead, the published estimate’s error about the chosen arm’s own effect and about the best arm’s effect, against how many patients an arm the first stage takes.

Against the chosen arm’s own effect — the quantity the earlier essays scored — the published estimate’s error only rises as patients move before the choice: 0.091 at sixteen an arm, 0.113 at forty, 0.152 at sixty and 0.165 at sixty-four, with one arm clearly ahead. On that target the best split puts as few patients as possible before the choice, in every world, because the estimate’s precision comes from the second stage and the winner’s curse comes from the first. Against the best arm’s effect the same estimate is best at forty-eight an arm. The same opposition holds in every world with a best arm: with two clearly ahead, the error about the chosen arm is 0.091 at sixteen an arm and 0.152 at sixty, while the error about the best arm is smallest at forty.

So a trial optimised to estimate whatever arm it chose and one optimised to say how good the best treatment is are different trials. The first spends almost nothing on the choice; the second spends enough to make the choice reliable and no more. The familiar split serves neither exactly. It is more choice than the first target wants, and in most worlds more choice than the second needs.

Where the curse comes from in each split

The published estimate’s error about the chosen arm has a bias in it, and the bias is the winner’s curse: the chosen arm’s first-stage mean was chosen because it was high, and pooling it with the second stage carries some of that luck into the published number. How much depends directly on the split, because the first stage’s share of the pooled estimate is its share of the chosen arm’s patients.

With the arms alike, the published estimate overstates the chosen arm’s effect by 0.023 at sixteen an arm before the choice, 0.043 at thirty-two and 0.126 at sixty. At sixty the first stage contributes two thirds of the chosen arm’s patients, so two thirds of the first-stage luck survives into the number; at sixteen it contributes one fifteenth. With two arms clearly ahead the bias is 0.015, 0.022 and 0.055 at the same splits, and with the eight spread evenly 0.019, 0.031 and 0.079. Only with one arm clearly ahead is the bias small at every split — 0.012, 0.011 and 0.007 — because the arm chosen is almost always the one whose lead is real, and a real lead carries little luck.

The second stage alone carries none, at any split: its bias is zero to within its counting error in every world. That is the property the earlier essays built it for, and the split does not touch it. What the split decides is how many patients the unbiased estimate gets, and at sixty an arm it gets thirty.

So the familiar split is the one that makes the curse largest, in every world but the one where a single arm is clearly better. It is not a coincidence that the curse was discovered in trials shaped that way. A design that spends most of its budget choosing, and then publishes a pooled estimate, has built the curse into its own arithmetic: the estimate after the choice measured how much a reader should discount it, and the measurement here says how much of the discount a design could have avoided by spending its patients the other way.

A split chosen without knowing the world

Each world’s best split is known only to someone who knows the world. A trial has to fix its split first, and the honest comparison is the worst case: for each split, the largest excess of its error over the best split’s error, across the five worlds.

For the published estimate that worst case is smallest at thirty-two an arm before the choice — 288 patients choosing, 312 estimating. There it costs 17.5% more than the best split with the arms alike, 13.6% with one modestly ahead, 21.3% with one clearly ahead, 3.3% with two clearly ahead and 5.5% with the arms spread evenly. The familiar sixty an arm costs 98.8% more with the arms alike — twice the error — and between 1.5% and 17.4% in the other four worlds. Its worst case is almost five times the minimax split’s.

How far the mixture-shrunk estimate of the chosen arm sits from the best arm's effect, by how six hundred patients are split between choosing and estimating. Eight arms and a control, six hundred patients: a stated number an arm in the first stage and the rest split equally between the chosen arm and the control in the second. Root mean squared error of the mixture-shrunk estimate about the best arm's true effect, over 5,000 trials a split. All eight alike: best at 16 an arm (0.089), 0.136 at the familiar sixty; One modestly ahead: best at 56 an arm (0.178), 0.183 at the familiar sixty; One clearly ahead: best at 48 an arm (0.161), 0.182 at the familiar sixty; Two clearly ahead: best at 40 an arm (0.131), 0.150 at the familiar sixty; Spread evenly: best at 40 an arm (0.143), 0.153 at the familiar sixty. The split that guards every world puts 32 an arm before the choice and costs at most 16.1% over each world's best.
Fig. 4 The mixture-shrunk estimate’s error about the best arm’s true effect against the first stage’s size, in the same five configurations.

The other two estimates move the numbers and not the conclusion. The mixture-shrunk estimate, which already discounts a lead that looks like luck, is least hurt by a large first stage — sixty an arm costs it at most 51.8% — and its minimax split is again thirty-two an arm, at a worst case of 16.1%. The second stage alone, which ignores the first stage entirely, is the most hurt: at sixty an arm it has only thirty patients a group to estimate from, and it errs 178.4% more than its best split with the arms alike. Its minimax split is thirty-two an arm too, at 19.5%.

Why the familiar shape is the familiar shape

The sixty-and-thirty shape comes from a different question. Selection trials were designed to pick a winner reliably, and on that question alone a large first stage is right: the pick rate rises with every patient spent on the choice, and the second stage is there to confirm rather than to estimate. The winner’s curse and the estimate after the choice are about what happens when such a trial’s number is then published as an effect size, and the measurement here says the design was not built for that number.

That is not a criticism of the shape for the job it was designed for. With one arm clearly ahead it picks the winner 98% of the time and its published estimate errs by 0.154 about the best arm, within 5.1% of the best any split achieves. The cost lands in the worlds where the choice matters less than the shape assumes — arms that are alike, or nearly alike — and those are exactly the worlds where a published effect size is most likely to mislead, because the curse is largest when there is little to find.

What the five worlds stand for

The five configurations are not a sample of anything, and the minimax split is only as good as the list. They were chosen in a prior that lets one arm stand out to span the situations a selection trial is designed for: nothing to find, something small to find, one thing clearly to find, two things to find, and a gradient. A trial whose arms are all variations on one mechanism is probably nearer the alike world; one comparing genuinely different treatments, nearer the clear-winner worlds.

That matters because the worlds disagree about the best split by a factor of three — sixteen an arm against forty-eight — and the minimax split is a compromise between the alike world and the clear-winner one, which are the two whose best splits are furthest apart. A trial confident that one arm is clearly better can spend more before the choice; one that expects the arms to be similar should spend less. What no design should do is spend sixty an arm before the choice on the assumption that the choice is the hard part and then publish the pooled estimate as an effect size.

What a design can take from this

Decide which number the trial is for. If the report will print an effect size for the chosen arm, the error about the best arm is the target, and a split near a third of the budget before the choice guards every world measured to within about a fifth. If the trial only has to pick, the familiar shape picks well.

Spend before the choice only what makes the choice reliable. With a clear winner, forty-eight an arm finds it 96% of the time; every patient past that point does more good estimating it. With no winner, every patient spent choosing is wasted.

Prefer an estimate that discounts the first stage when the first stage is large. The mixture-shrunk estimate loses least to a large first stage, as shrinking the arm that was chosen would suggest, and its minimax worst case is the smallest of the three.

Every error is a root mean squared error over five thousand trials for each split and world, with the same seeds across estimates so the comparison between estimates is on identical trials. Every split is checked to spend exactly six hundred patients, the second stage alone to stay unbiased at every split, and the published estimate’s error with the arms alike to rise with every patient moved before the choice. The familiar split is refused as right for every world: with the arms alike its published estimate errs by twice the best split’s.

Still open: a first stage that stops early

Every split here is fixed before the trial. A first stage can instead be sequential: look after a few patients an arm, and stop choosing when one arm’s lead is clear — moving the remaining patients into the second stage only when the choice has been made reliably. That is a rule that spends patients before the choice in proportion to how hard the choice is, which is exactly what the five worlds disagree about.

Whether such a rule beats the fixed minimax split’s worst case of 21.3%, and what it does to the published estimate’s bias — a choice stopped early on a clear lead is a choice made on a lead that may be partly luck, the curse at its sharpest — is computable on the same trials, and has not been computed.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive designEmpirical BayesMean squared errorSample sizeSelection biasShrinkageTwo-stage designThe winner's curse