Shrinking the arm that was chosen
Worth reading first: Five times in six.
The estimate after the choice measured what a drop-the-losers trial reports about the arm it carries forward. Eight arms run a first stage of sixty patients each; the one that looks best continues with thirty more; the trial publishes the chosen arm’s effect using all ninety of its patients. With every arm truly alike, the first stage reports the chosen arm as better than it is, because it was chosen for being ahead; the second stage, collected after the choice, is unbiased; and the published combination inherits two thirds of the first stage’s bias. The unbiased estimate is built from a third of the data and is the least accurate of the three.
The essay named three ways to correct the published estimate and attempted none. The one it called most promising was to shrink: pull the chosen arm’s first-stage result back towards the other arms, by a weight estimated from how spread out the eight arms are, since the pull is in exactly the direction the selection pushed. It also named the difficulty — the weight is estimated from eight numbers — and that difficulty turns out to be the whole story.
The shrunk estimate
Each arm’s first-stage mean is a noisy measurement of its true mean, with a standard error set by its sixty patients. If the eight true means are themselves spread around a common value with some spread τ, the best estimate of any one of them is its own mean pulled towards the common value by the weight — the weight eight groups, one population derived for groups borrowing from each other. The spread is not known, so it is estimated from the eight first-stage means themselves: their observed variance less the part the sampling noise accounts for, floored at zero.
The chosen arm’s first-stage mean is shrunk this way, then combined with its thirty second-stage patients by their share of the information, and the control mean is subtracted. The arms’ means are shrunk rather than their differences from control, because the eight differences share one control group and are not eight independent draws from anything.
The logic of the correction is appealing. The selection picked the arm with the largest first-stage mean, and the largest of eight noisy means is too large for two reasons: its true value may be large, and its noise is probably positive. Shrinkage removes the second part in proportion to how much of the arms’ spread is noise, and a selection made on noise is corrected in full.
When the arms are alike, shrinkage wins
With all eight arms truly alike, the published estimate overstates the chosen arm by +0.123 standard deviations on average; the second stage alone is unbiased, −0.001; the shrunk estimate is off by +0.020. The arms’ spread is estimated as exactly zero on 57.1% of trials, and in those the chosen arm’s first stage is pulled all the way to the eight arms’ mean, which is exactly right when there is nothing to choose between them.
In root-mean-square error the shrunk estimate is the best of the three, 0.133 against the published 0.182 and the second stage’s 0.261. It is both less biased than the published estimate and less noisy than the second stage alone, because it keeps all ninety of the chosen arm’s patients and uses the other seven arms’ sixty each to judge how much of the chosen arm’s lead to believe.
The same holds when the arms genuinely differ but gradually — true effects spread evenly from nothing to half a standard deviation. The published estimate overstates by +0.075; the shrunk estimate is off by −0.003, with an error of 0.145 against 0.160. In a trial of eight arms whose differences are real but modest, the correction does what the winner’s curse asks of any estimate reported after a selection: it takes back the part of the winner’s margin the selection manufactured.
When one arm is really ahead, shrinkage loses
The trial was run to find the best arm, and the configuration it hopes for is the one where one arm is clearly better than the others — the result that would justify having run eight arms rather than one, and the one a successful programme will be built on. Set one arm’s true effect half a standard deviation above the rest.
The first stage now picks the right arm 98.2% of the time, and because the right arm is far enough ahead that it would have won without luck, the selection adds little: the published estimate is off by only +0.006. The shrunk estimate is off by −0.103. It has pulled the true winner a fifth of the way back towards seven arms it is genuinely better than, and in root-mean-square error it is the worst of the two estimates that use all the data: 0.189 against the published 0.151.
The reason is in the weight. With one arm far from seven identical ones, the eight means’ spread is real but lopsided, and the moment estimate of τ treats it as the spread of a bell. The weight it produces — about 0.42 on the mean of the eight, on average — is the right shrinkage for a typical arm in a population with that spread and the wrong one for the arm that is the population’s outlier. A group from the population’s own tail found the same thing for partial pooling generally: the group whose true value is far from the centre is estimated worse by pooling than by its own mean. A trial’s chosen arm is, whenever the trial succeeds, exactly that group.
The weight is estimated from eight numbers
With the arms alike, the weight on the arms’ mean averages 0.879 and is exactly one on 56.6% of trials, because the spread of eight noisy means with nothing behind them usually falls below what sampling noise alone would produce, and the moment estimate floors at zero. That is the spread that estimates to zero, and here it is the correction working: with nothing to choose between, full pooling is right.
With one arm ahead, the weight averages 0.416, and it is almost never zero or one. The eight means are spread by a real difference, the estimate of τ says so, and the chosen arm is pulled back by a weight calibrated to the spread rather than to the chosen arm’s position in it. Nothing in eight numbers can tell a population of eight moderately spread arms from seven alike arms and one outlier; the fewest groups that can borrow found how little eight groups can say about their own spread, and saying which shape the spread has is harder still.
No estimate wins everywhere
Across five configurations of the eight arms, the three estimates trade places.
The second stage alone has the same error, 0.261, in every configuration: it is unbiased and uses thirty patients, and nothing about the other arms affects it. The shrunk estimate is best when the arms are alike (0.133) or spread evenly (0.145), and slightly better when one arm is ahead by a quarter (0.157 against 0.165). With two arms ahead by 0.4 the shrunk and published estimates tie at 0.152. With one arm ahead by a half, the published estimate wins, 0.151 against 0.189.
The biases behind those errors follow the same pattern. With one arm ahead by a quarter, the published estimate overstates the chosen arm by +0.072 and the shrunk one understates it by 0.038; with two arms ahead by 0.4, +0.052 and −0.035. In the middle configurations each estimate is wrong by a few hundredths in its own direction, and neither is clearly better; the extremes are where the choice matters, and they are opposite extremes.
So the choice of estimate depends on the configuration the trial is in, and the configuration is what the trial is trying to find out. A design that expects its arms to be similar, and runs eight to find a modest improvement among them, is served well by shrinkage. A design that expects one arm to stand out, and runs the others as the field it will be judged against, is served better by the published estimate — whose bias, in that case, is small precisely because the selection had a real difference to find.
Why the corrections point in opposite directions
The published estimate’s bias and the shrunk estimate’s bias have a simple relationship. The published estimate is biased upwards by the part of the winner’s lead that was luck. The shrunk estimate is biased downwards by the part of the winner’s lead that was real but was treated as luck. When the lead is all luck, the first bias is large and the second is zero; when the lead is all real, the first is zero and the second is large. Every configuration sits somewhere between, and the two biases trade against each other along the way.
It would be convenient if the size of the correction revealed which world a trial is in, and it does not. The shrinkage moves the chosen arm’s estimate down by 0.103 on average when the eight arms are alike and by 0.108 when one arm is truly ahead by half a standard deviation — the same distance, with a spread of about 0.02 to 0.03 across trials in both. The correction is a nearly fixed pull, set by how far the chosen arm’s first stage sits from the other seven, and the chosen arm sits about as far from them whether it got there by luck or by being better. In one world that pull removes a bias; in the other it creates one of the same size.
So the gap between the shrunk and published estimates is not evidence about the selection. What would be evidence is the second stage, which the selection could not see: when the chosen arm’s thirty later patients reproduce its first-stage lead, the lead was real, and when they fall back, it was luck. The shrinkage uses the second stage only as extra data, weighted by its share of the information, and so it cannot let the second stage decide how much to believe the first.
What the second stage can and cannot settle
The second stage is thirty patients against thirty controls, and its estimate of the chosen arm’s effect has a standard error of standard deviations. That is enough to lean one way and not enough to decide. When the arms are alike, the chosen arm’s second stage exceeds a quarter of a standard deviation 16.6% of the time; when the chosen arm is truly ahead by half a standard deviation, 83.4% of the time. A trial whose second stage comes in above a quarter is five times as likely to be in the second world as in the first, which is real evidence and falls well short of certainty.
That is also why the design is built the way it is. A second stage large enough to settle the question would have been large enough to run as a trial of its own, and the attraction of carrying the best of eight forward is that the second stage need only confirm, not discover. The price is that it confirms weakly, and every estimate that uses the first stage has to decide how much of the first stage’s lead to believe with the second stage able to contribute only a little evidence about it.
The configuration a design expects
The configurations above are not equally likely in practice, and which one a trial should be designed for is a judgement about its field. A dose-finding or formulation screen, choosing among variants of one intervention that are expected to differ modestly, is closest to the evenly spread case, where shrinkage is the better estimate. A screen of genuinely different candidates, run in the hope that one of them works and the rest do not, is closest to the one-clear case, where the published estimate is nearly unbiased and shrinkage harms the one result that matters.
A design can state which it expects before the trial runs, and the choice of estimate can follow from the statement, the way a protocol states the analysis before the data. What it should not do is choose the estimate after seeing which arm won and by how much, because that choice would be one more selection on the same data.
What a drop-the-losers trial should report
The second-stage estimate, always. It is the only one of the three that is unbiased whatever the arms’ configuration, and its interval is honest. It is wide — dropping the losers already found the design’s power concentrated in the selection — and its width is the true price of having chosen the arm from the same trial.
The published and shrunk estimates beside it, and the stage-two estimate as the arbiter. The shrunk estimate moves the chosen arm down by about a tenth whatever the truth; whether that tenth was bias or signal is what the second stage, compared with the first, can speak to and the shrinkage cannot.
The estimated spread of the arms, including when it is zero. A zero spread means the first stage could not distinguish the arms from each other at all, which is a finding about the selection that the published estimate hides entirely.
No single corrected number presented as the effect. An interval for the winner showed that a correction that conditions on the selection can be exact for its coverage; a shrinkage correction is a different bargain — lower error on average over configurations of similar arms, higher error for the configuration a successful trial is in — and it should be reported as that bargain, not as the answer.
What is measured here and what is not
With eight alike arms of sixty and thirty more for the chosen one, the published estimate overstates it by 0.123, the shrunk estimate by 0.020, and the shrunk estimate has the least error, 0.133 against 0.182.
With one arm ahead by half a standard deviation, the published estimate overstates it by 0.006 and the shrunk estimate understates it by 0.103, with error 0.189 against 0.151.
Every number is counted over twenty thousand simulated trials a configuration, with outcomes normal of known unit variance and the arms’ spread estimated by the moment equation floored at zero.
Not measured: shrinkage towards the control rather than towards the arms’ mean, which assumes the arms are no better than nothing and would over-correct further; fully Bayesian shrinkage with a prior on the spread, which would soften the floor at zero; and designs whose number of arms or stage sizes differ, where the balance between the published estimate’s bias and the shrunk estimate’s changes.
Still open: a shrinkage that knows one arm might stand out
The correction failed because it assumed the eight arms were a bell and the successful trial’s arms are a bell with an outlier. A prior that allows for the second shape — a mixture of a common distribution and a small chance that one arm is genuinely different, or a heavy-tailed distribution of arm effects — would shrink the chosen arm less when its lead is too large for the bell, in the way a heavy-tailed prior let go of a history that disagreed.
Whether such a prior keeps most of the shrinkage’s advantage when the arms are alike while losing little when one stands out, how its behaviour depends on the number of arms, and whether eight arms carry enough information to tell which shape they have, are the measurements that would make the correction safe to recommend. None of them has been made here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Estimates that are too alike — both name empirical bayes, mean squared error, partial pooling, shrinkage
- Where the borrowing goes — both name empirical bayes, mean squared error, partial pooling, shrinkage
- One population, or two — both name empirical bayes, partial pooling, shrinkage
- The correction that makes the estimate worse — both name mean squared error, selection bias, the winner's curse
- The effect a stopped trial reports — both name interim analysis, selection bias, the winner's curse
- The small groups one centre protects — both name empirical bayes, partial pooling, shrinkage
Named objects
A flat tag is an object no other essay names yet.
Adaptive designEmpirical BayesInterim analysisMean squared errorPartial poolingSelection biasShrinkageThe winner's curse