Designs that change while they run

Shrinking the arm that was chosen

A trial that carries the best of eight arms forward reports it as better than it is — by +0.123 when the eight arms are alike — and its second stage alone is unbiased but the least accurate of the estimates. Pulling the chosen arm's first stage back towards the eight arms' mean before combining removes most of the bias when the arms are alike (+0.020) and gives the smallest error of any estimate, 0.133 against the published 0.182. When one arm really is ahead by half a standard deviation, the same shrinkage pulls the true winner back towards the losers and understates it by 0.103, with more error than the published estimate. The weight is estimated from eight numbers, and it is right when the selection was noise and wrong when the selection was right.

Worth reading first: Five times in six.

The estimate after the choice measured what a drop-the-losers trial reports about the arm it carries forward. Eight arms run a first stage of sixty patients each; the one that looks best continues with thirty more; the trial publishes the chosen arm’s effect using all ninety of its patients. With every arm truly alike, the first stage reports the chosen arm as better than it is, because it was chosen for being ahead; the second stage, collected after the choice, is unbiased; and the published combination inherits two thirds of the first stage’s bias. The unbiased estimate is built from a third of the data and is the least accurate of the three.

The essay named three ways to correct the published estimate and attempted none. The one it called most promising was to shrink: pull the chosen arm’s first-stage result back towards the other arms, by a weight estimated from how spread out the eight arms are, since the pull is in exactly the direction the selection pushed. It also named the difficulty — the weight is estimated from eight numbers — and that difficulty turns out to be the whole story.

How far three estimates of the chosen arm sit from its true effect, with all eight alikeEight arms of sixty, the best carried forward with thirty more. Average error: both stages, as published +0.123; stage two alone −0.001; stage one shrunk, then both +0.020. The arms' spread estimates to zero on 57.1% of trials.both stages, as published+0.123stage two alone−0.001stage one shrunk, then both+0.020← understatedoverstated →20,000 trials, eight arms of 60 then 30 morebias of the chosen arm's estimate
Fig. 1 How far three estimates of the chosen arm’s effect sit from its truth on average, when all eight arms are alike: the published combination of both stages, the second stage alone, and the first stage shrunk towards the eight arms’ mean before combining. The slider changes the arms’ true effects.

The shrunk estimate

Each arm’s first-stage mean is a noisy measurement of its true mean, with a standard error set by its sixty patients. If the eight true means are themselves spread around a common value with some spread τ, the best estimate of any one of them is its own mean pulled towards the common value by the weight se2/(se2+τ2)\mathrm{se}^2/(\mathrm{se}^2 + \tau^2) — the weight eight groups, one population derived for groups borrowing from each other. The spread is not known, so it is estimated from the eight first-stage means themselves: their observed variance less the part the sampling noise accounts for, floored at zero.

The chosen arm’s first-stage mean is shrunk this way, then combined with its thirty second-stage patients by their share of the information, and the control mean is subtracted. The arms’ means are shrunk rather than their differences from control, because the eight differences share one control group and are not eight independent draws from anything.

The logic of the correction is appealing. The selection picked the arm with the largest first-stage mean, and the largest of eight noisy means is too large for two reasons: its true value may be large, and its noise is probably positive. Shrinkage removes the second part in proportion to how much of the arms’ spread is noise, and a selection made on noise is corrected in full.

When the arms are alike, shrinkage wins

With all eight arms truly alike, the published estimate overstates the chosen arm by +0.123 standard deviations on average; the second stage alone is unbiased, −0.001; the shrunk estimate is off by +0.020. The arms’ spread is estimated as exactly zero on 57.1% of trials, and in those the chosen arm’s first stage is pulled all the way to the eight arms’ mean, which is exactly right when there is nothing to choose between them.

In root-mean-square error the shrunk estimate is the best of the three, 0.133 against the published 0.182 and the second stage’s 0.261. It is both less biased than the published estimate and less noisy than the second stage alone, because it keeps all ninety of the chosen arm’s patients and uses the other seven arms’ sixty each to judge how much of the chosen arm’s lead to believe.

The same holds when the arms genuinely differ but gradually — true effects spread evenly from nothing to half a standard deviation. The published estimate overstates by +0.075; the shrunk estimate is off by −0.003, with an error of 0.145 against 0.160. In a trial of eight arms whose differences are real but modest, the correction does what the winner’s curse asks of any estimate reported after a selection: it takes back the part of the winner’s margin the selection manufactured.

When one arm is really ahead, shrinkage loses

The trial was run to find the best arm, and the configuration it hopes for is the one where one arm is clearly better than the others — the result that would justify having run eight arms rather than one, and the one a successful programme will be built on. Set one arm’s true effect half a standard deviation above the rest.

How far three estimates of the chosen arm sit from its true effect, with one arm ahead by 0.5. Eight arms of sixty, the best carried forward with thirty more. Average error: both stages, as published +0.006; stage two alone −0.001; stage one shrunk, then both −0.103. The arms' spread estimates to zero on 2.6% of trials.
Fig. 2 The same three estimates when one arm is truly ahead by half a standard deviation and the other seven are alike. The published estimate is nearly unbiased; the shrunk one understates the winner.

The first stage now picks the right arm 98.2% of the time, and because the right arm is far enough ahead that it would have won without luck, the selection adds little: the published estimate is off by only +0.006. The shrunk estimate is off by −0.103. It has pulled the true winner a fifth of the way back towards seven arms it is genuinely better than, and in root-mean-square error it is the worst of the two estimates that use all the data: 0.189 against the published 0.151.

The reason is in the weight. With one arm far from seven identical ones, the eight means’ spread is real but lopsided, and the moment estimate of τ treats it as the spread of a bell. The weight it produces — about 0.42 on the mean of the eight, on average — is the right shrinkage for a typical arm in a population with that spread and the wrong one for the arm that is the population’s outlier. A group from the population’s own tail found the same thing for partial pooling generally: the group whose true value is far from the centre is estimated worse by pooling than by its own mean. A trial’s chosen arm is, whenever the trial succeeds, exactly that group.

The weight is estimated from eight numbers

How strongly the chosen arm is pulled back towards the others, when the arms are alike and when one is ahead. With the arms alike the weight on the arms' mean averages 0.879 and is exactly one — the spread estimated as zero — on 56.6% of trials. With one arm ahead by 0.5 it averages 0.416.
Fig. 3 The weight the shrinkage puts on the eight arms’ mean, as a distribution over trials, when the arms are alike and when one arm is ahead by half a standard deviation.

With the arms alike, the weight on the arms’ mean averages 0.879 and is exactly one on 56.6% of trials, because the spread of eight noisy means with nothing behind them usually falls below what sampling noise alone would produce, and the moment estimate floors at zero. That is the spread that estimates to zero, and here it is the correction working: with nothing to choose between, full pooling is right.

With one arm ahead, the weight averages 0.416, and it is almost never zero or one. The eight means are spread by a real difference, the estimate of τ says so, and the chosen arm is pulled back by a weight calibrated to the spread rather than to the chosen arm’s position in it. Nothing in eight numbers can tell a population of eight moderately spread arms from seven alike arms and one outlier; the fewest groups that can borrow found how little eight groups can say about their own spread, and saying which shape the spread has is harder still.

No estimate wins everywhere

Across five configurations of the eight arms, the three estimates trade places.

The root-mean-square error of each estimate of the chosen arm, in five configurations of the eight arms. all eight alike: published 0.182, stage two 0.261, shrunk 0.133; spread evenly from 0 to 0.5: published 0.160, stage two 0.261, shrunk 0.145; one arm ahead by 0.25: published 0.165, stage two 0.261, shrunk 0.157; two arms ahead by 0.4: published 0.152, stage two 0.261, shrunk 0.152; one arm ahead by 0.5: published 0.151, stage two 0.261, shrunk 0.189.
Fig. 4 The root-mean-square error of the three estimates of the chosen arm’s effect in five configurations: all alike, spread evenly, one arm ahead by a quarter, two arms ahead by 0.4, and one arm ahead by a half.

The second stage alone has the same error, 0.261, in every configuration: it is unbiased and uses thirty patients, and nothing about the other arms affects it. The shrunk estimate is best when the arms are alike (0.133) or spread evenly (0.145), and slightly better when one arm is ahead by a quarter (0.157 against 0.165). With two arms ahead by 0.4 the shrunk and published estimates tie at 0.152. With one arm ahead by a half, the published estimate wins, 0.151 against 0.189.

The biases behind those errors follow the same pattern. With one arm ahead by a quarter, the published estimate overstates the chosen arm by +0.072 and the shrunk one understates it by 0.038; with two arms ahead by 0.4, +0.052 and −0.035. In the middle configurations each estimate is wrong by a few hundredths in its own direction, and neither is clearly better; the extremes are where the choice matters, and they are opposite extremes.

So the choice of estimate depends on the configuration the trial is in, and the configuration is what the trial is trying to find out. A design that expects its arms to be similar, and runs eight to find a modest improvement among them, is served well by shrinkage. A design that expects one arm to stand out, and runs the others as the field it will be judged against, is served better by the published estimate — whose bias, in that case, is small precisely because the selection had a real difference to find.

Why the corrections point in opposite directions

Three estimates of the same effect, and the honest one is the worst. 8,000 trials in which every arm has an effect of exactly 0. The arm carried forward is reported by its first stage at +0.187 above the truth — it was chosen for being ahead — and by its second stage at -0.0032, which is unbiased by construction because the selection could not see it. The combination that gets published is at 0.123, exactly the share of the first stage's bias the first stage contributes. Root mean squared error: 0.240, 0.261, 0.182 — the unbiased estimate is the least accurate of the three.
Fig. 5 The three estimates the essay on the choice compared, eight arms alike: the first stage that chose the arm, the second stage collected after it, and the published combination. The shrunk estimate is a fourth, built from the first stage’s distance to the other arms.

The published estimate’s bias and the shrunk estimate’s bias have a simple relationship. The published estimate is biased upwards by the part of the winner’s lead that was luck. The shrunk estimate is biased downwards by the part of the winner’s lead that was real but was treated as luck. When the lead is all luck, the first bias is large and the second is zero; when the lead is all real, the first is zero and the second is large. Every configuration sits somewhere between, and the two biases trade against each other along the way.

It would be convenient if the size of the correction revealed which world a trial is in, and it does not. The shrinkage moves the chosen arm’s estimate down by 0.103 on average when the eight arms are alike and by 0.108 when one arm is truly ahead by half a standard deviation — the same distance, with a spread of about 0.02 to 0.03 across trials in both. The correction is a nearly fixed pull, set by how far the chosen arm’s first stage sits from the other seven, and the chosen arm sits about as far from them whether it got there by luck or by being better. In one world that pull removes a bias; in the other it creates one of the same size.

So the gap between the shrunk and published estimates is not evidence about the selection. What would be evidence is the second stage, which the selection could not see: when the chosen arm’s thirty later patients reproduce its first-stage lead, the lead was real, and when they fall back, it was luck. The shrinkage uses the second stage only as extra data, weighted by its share of the information, and so it cannot let the second stage decide how much to believe the first.

What the second stage can and cannot settle

The second stage is thirty patients against thirty controls, and its estimate of the chosen arm’s effect has a standard error of 2/30=0.258\sqrt{2/30} = 0.258 standard deviations. That is enough to lean one way and not enough to decide. When the arms are alike, the chosen arm’s second stage exceeds a quarter of a standard deviation 16.6% of the time; when the chosen arm is truly ahead by half a standard deviation, 83.4% of the time. A trial whose second stage comes in above a quarter is five times as likely to be in the second world as in the first, which is real evidence and falls well short of certainty.

That is also why the design is built the way it is. A second stage large enough to settle the question would have been large enough to run as a trial of its own, and the attraction of carrying the best of eight forward is that the second stage need only confirm, not discover. The price is that it confirms weakly, and every estimate that uses the first stage has to decide how much of the first stage’s lead to believe with the second stage able to contribute only a little evidence about it.

The configuration a design expects

The configurations above are not equally likely in practice, and which one a trial should be designed for is a judgement about its field. A dose-finding or formulation screen, choosing among variants of one intervention that are expected to differ modestly, is closest to the evenly spread case, where shrinkage is the better estimate. A screen of genuinely different candidates, run in the hope that one of them works and the rest do not, is closest to the one-clear case, where the published estimate is nearly unbiased and shrinkage harms the one result that matters.

A design can state which it expects before the trial runs, and the choice of estimate can follow from the statement, the way a protocol states the analysis before the data. What it should not do is choose the estimate after seeing which arm won and by how much, because that choice would be one more selection on the same data.

What a drop-the-losers trial should report

The second-stage estimate, always. It is the only one of the three that is unbiased whatever the arms’ configuration, and its interval is honest. It is wide — dropping the losers already found the design’s power concentrated in the selection — and its width is the true price of having chosen the arm from the same trial.

The published and shrunk estimates beside it, and the stage-two estimate as the arbiter. The shrunk estimate moves the chosen arm down by about a tenth whatever the truth; whether that tenth was bias or signal is what the second stage, compared with the first, can speak to and the shrinkage cannot.

The estimated spread of the arms, including when it is zero. A zero spread means the first stage could not distinguish the arms from each other at all, which is a finding about the selection that the published estimate hides entirely.

No single corrected number presented as the effect. An interval for the winner showed that a correction that conditions on the selection can be exact for its coverage; a shrinkage correction is a different bargain — lower error on average over configurations of similar arms, higher error for the configuration a successful trial is in — and it should be reported as that bargain, not as the answer.

What is measured here and what is not

With eight alike arms of sixty and thirty more for the chosen one, the published estimate overstates it by 0.123, the shrunk estimate by 0.020, and the shrunk estimate has the least error, 0.133 against 0.182.

With one arm ahead by half a standard deviation, the published estimate overstates it by 0.006 and the shrunk estimate understates it by 0.103, with error 0.189 against 0.151.

Every number is counted over twenty thousand simulated trials a configuration, with outcomes normal of known unit variance and the arms’ spread estimated by the moment equation floored at zero.

Not measured: shrinkage towards the control rather than towards the arms’ mean, which assumes the arms are no better than nothing and would over-correct further; fully Bayesian shrinkage with a prior on the spread, which would soften the floor at zero; and designs whose number of arms or stage sizes differ, where the balance between the published estimate’s bias and the shrunk estimate’s changes.

Still open: a shrinkage that knows one arm might stand out

The correction failed because it assumed the eight arms were a bell and the successful trial’s arms are a bell with an outlier. A prior that allows for the second shape — a mixture of a common distribution and a small chance that one arm is genuinely different, or a heavy-tailed distribution of arm effects — would shrink the chosen arm less when its lead is too large for the bell, in the way a heavy-tailed prior let go of a history that disagreed.

Whether such a prior keeps most of the shrinkage’s advantage when the arms are alike while losing little when one stands out, how its behaviour depends on the number of arms, and whether eight arms carry enough information to tell which shape they have, are the measurements that would make the correction safe to recommend. None of them has been made here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive designEmpirical BayesInterim analysisMean squared errorPartial poolingSelection biasShrinkageThe winner's curse