A prior that lets one arm stand out
Worth reading first: Five times in six.
Shrinking the arm that was chosen corrected the winner’s curse in a drop-the-losers trial by pulling the chosen arm’s first stage back towards the mean of all eight arms, with a weight estimated from their spread. When the eight arms were alike it removed most of the published estimate’s bias and had the least error of any estimate. When one arm was genuinely ahead by half a standard deviation it pulled that arm back a fifth of the way towards seven arms it was better than, and understated it by 0.103.
The failure had a diagnosis: the correction assumed the eight arms were a bell, and the successful trial’s arms are a bell with an outlier. The essay ended on the prior that would allow for the second shape — heavy tails, or a mixture with a small chance that one arm is genuinely different — and asked three things of it: whether it keeps the shrinkage’s advantage when the arms are alike, what it recovers when one stands out, and whether eight arms carry enough information to tell which shape they have.
Two priors that allow an outlier
The trial is the one the earlier essays used: eight arms of sixty patients and a shared control, the best arm carried into a second stage with thirty more, and its effect estimated from both stages. Everything below is computed on the same draws as the normal shrinkage, so its numbers are the earlier essay’s exactly.
The heavy-tailed prior replaces the normal distribution of arm effects with a t distribution on three degrees of freedom, centred on the eight arms’ mean and scaled to have the same variance as the moment estimate of their spread. Near the centre it shrinks almost as the normal prior does; far out, its tail makes a large lead a plausible arm rather than an implausible piece of luck, and the posterior mean follows the data. That is the behaviour a lead that a heavy tail keeps measured for selected units in general: under a t on four degrees of freedom, the top one per cent kept 78.4% of their lead where normal scores kept 60%.
The mixture prior is the normal one with probability four fifths and, with probability one fifth, a component whose spread is wider by a full standard deviation — an arm that is genuinely different from the rest. It is the construction a history that agrees with itself used to give a borrowed control a way out: the data decide how much weight the vague component gets, and a chosen arm far from the others shifts weight onto it and is shrunk less.
The mixture’s verdict is a single ratio. With the eight arms’ mean, their estimated spread and the chosen arm’s first-stage mean, the weight on the vague component is
the prior odds times how much better a wide normal predicts the observed lead than the bell does. The chosen arm’s estimate is then the two components’ posterior means averaged with weights and : nearly unshrunk under the first, shrunk exactly as before under the second. Every number the mixture produces below is that average, and its behaviour is entirely a matter of how large gets.
What each keeps when the arms are alike
A correction for a real winner is worth nothing if it gives back the bias it was built to remove, so the first test is the configuration the normal shrinkage handles well.
With all eight arms alike, the published estimate overstates the chosen arm by +0.123. The normal shrinkage is off by +0.020; the t prior by +0.016; the mixture by +0.030. In root-mean-square error the three shrunk estimates are 0.133, 0.131 and 0.135, against the published 0.182. Allowing an outlier costs almost nothing here, because when the arms are alike the chosen arm’s lead is small enough that neither prior’s allowance comes into play: the mixture’s average weight on “stands out” is 0.086, and it never exceeds one half.
The middle configurations tell the same story with smaller margins. With the arms spread evenly from nothing to half a standard deviation the three shrunk estimates all err by 0.145 or 0.146; with one arm ahead by a quarter the mixture is the best of all four estimates at 0.154; with two arms ahead by 0.4, again the mixture, at 0.150. The t prior is never the best and, in the configurations where more than one arm is ahead, slightly worse than the normal: two arms ahead by 0.4 are understated by 0.048 against 0.035.
What each recovers when one arm is ahead
With one arm ahead by half a standard deviation — the configuration the trial was run in the hope of finding — the normal shrinkage understates the winner by 0.103. The t prior understates it by 0.101: it has recovered nothing. The mixture understates it by 0.073, recovering about three tenths of the loss. In root-mean-square error the three are 0.189, 0.192 and 0.176, and all three are worse than the published estimate’s 0.151, whose bias in this world is only +0.006 because the selection had a real difference to find.
The t prior fails for a reason that is plain once the numbers are in standard errors. The winner’s true lead over the eight arms’ mean is 0.44 — half a standard deviation, less its own eighth share of the mean — and a first-stage arm mean has a standard error of 0.129, so the lead is about three and a half standard errors. The moment estimate of the spread puts the t prior’s scale near a tenth, and at four scales out a t on three degrees of freedom is still shrinking hard: its tail only lets go of a value much further out than the likelihood’s own width. Heavy tails protect an outlier that is far beyond the noise, and a trial designed to find a half-standard-deviation difference with sixty patients an arm produces one that is not.
How far ahead an arm must be
The question can be turned round: how large must the winner’s lead be before each prior lets it keep it?
At a lead of a quarter of a standard deviation all three shrunk estimates understate by 0.023 to 0.045, and the published one overstates by 0.071; the lead is too small for the selection to be reliable, and pulling back is right. At a half, the normal and t priors are at their worst. At three quarters they part: normal −0.090, t −0.072, mixture −0.057. At a full standard deviation, −0.073, −0.050 and −0.043. At one and a half, −0.052, −0.032 and −0.031.
None of the shrunk estimates reaches the published estimate’s accuracy in this range, and none stops understating. Both robust priors behave as they should — the further out the winner, the less they pull it back — but they begin to behave that way at leads of three quarters of a standard deviation and more, which is six standard errors of an arm mean. A trial that could see a lead that large did not need eight arms to find it.
Whether eight arms know their shape
The mixture’s weight on its vague component is the prior’s own verdict on whether the chosen arm stands out. If the eight arms carried enough information to tell a bell from a bell with an outlier, the weight would be low in one world and high in the other.
It is higher when an arm really stands out — an average of 0.285 against 0.086 — but it almost never says so. Over twenty thousand trials with a genuine half-standard-deviation winner, the weight exceeds one half in 0.4%. The two distributions overlap on most of their range, and a trial reading its own posterior would conclude “probably a bell” nearly every time, in either world.
A larger lead is visible. With the winner a full standard deviation ahead the weight exceeds one half in 47.1% of trials, and at one and a half standard deviations in 94.2%. But those are the leads at which the question hardly matters, since the selection is reliable and the published estimate nearly unbiased. At the lead the design was sized for, eight arms cannot tell which shape they have — which is the same boundary the fewest groups that can borrow found for eight groups’ ability to say anything about their spread, one level further out.
Sixteen arms are a different problem
The number of arms changes the picture in both directions at once.
The normal shrinkage’s loss on a real winner grows with the number of arms: −0.055 at four, −0.104 at eight, −0.131 at twelve, −0.156 at sixteen. More alike arms make the bell look tighter, so the one real outlier is pulled harder towards it. The mixture’s loss does not grow: −0.047 at four, −0.075 at eight, −0.069 at twelve, −0.063 at sixteen. At sixteen arms it recovers three fifths of what the normal prior takes, because fifteen alike arms describe their bell well enough for the sixteenth to be seen as outside it: the weight on “stands out” exceeds one half in 59.9% of trials with a real winner, and in 1.2% with all sixteen alike.
The price rises too. When all sixteen are alike the mixture overstates the chosen arm by +0.039 against the normal’s +0.019, because the largest of sixteen noisy means is far enough out to draw some weight onto the vague component by luck. At four arms there is almost nothing to choose between the two priors, since four arms say little about any shape.
So the mixture is a correction for a trial with many arms. With eight it is a modest improvement that still loses to the published estimate when the trial succeeds; with sixteen it is most of the way to the estimate a trial would want, in both worlds.
The slope every prior is guessing
Every correction here is a way of estimating one number. For an arm mean observed with standard error , the posterior mean under any prior is the observation plus times the slope of the log of the observations’ marginal density at that point — Tweedie’s formula. A normal prior makes that slope a straight line through the arms’ mean, so the correction grows with the lead; a t prior makes it bend back towards zero far out; a mixture bends it back at the distance where the vague component starts to explain the data better than the bell.
So the three priors are three guesses at the shape of a density from eight points, and the chosen arm sits exactly where eight points say least: in the tail, beyond all but one of them. The slope of a density nobody can see estimated that slope directly from a study’s readings, and needed about a thousand of them before its estimate beat the linear rule, losing to it at 250. Eight arms are two orders of magnitude short of that, which is why each prior’s answer is mostly the prior’s shape rather than the data’s — and why sixteen arms, still far short, already change what the mixture can do.
The same reading explains the winner’s curse in one line. The winner’s curse and dropping the losers measured how far a selected estimate overstates; Tweedie’s formula says the overstatement is times the density’s slope at the winner, and a trial’s design fixes while nature fixes the slope. A trial with more patients an arm shrinks the correction quadratically, whatever the prior; a trial with more arms learns the slope. The first is the reliable remedy, and the second is the one the mixture needs.
Why the published estimate keeps winning
In every configuration with a real winner the published estimate has the least error, and it is worth saying why that is not a failure of the corrections. The published estimate’s bias is the part of the winner’s lead that was luck. When one arm is genuinely ahead by half a standard deviation, that part is small — the arm was going to win anyway — and the estimate’s error is mostly its own sampling noise, the same 0.15 it would have without any selection. Every shrinkage adds a pull to that noise, and a pull that is not needed is pure error.
The corrections win only where the selection manufactured the lead, and the trouble is that a trial does not know which case it is in. A group from the population’s own tail found that partial pooling does worse than a group’s own mean for every group more than 1.73 population widths from the centre; a trial’s winner is, when the trial succeeds, the population’s tail by construction. A prior that allows an outlier moves that boundary outwards without removing it.
What a trial reporting its chosen arm can do
Report the published estimate beside a shrunk one and say what each assumes. The published estimate is nearly unbiased if the winner was real and biased by up to +0.123 if it was not; the shrunk estimates are the reverse. The pair brackets the answer in both worlds.
Prefer the mixture to the normal prior when shrinking, and prefer neither to heavy tails at this scale. The mixture costs 0.010 of bias when the arms are alike and recovers three tenths of the loss when one is ahead; a t prior with the arms’ own spread recovers nothing at the leads a trial of this size produces.
With many arms, shrink with a mixture. At sixteen arms it recovers three fifths of the loss and can usually say which world the trial is in.
And let the second stage speak. The estimate after the choice and its successor found the second stage alone unbiased in every configuration, with an error of 0.261 from thirty patients. It is the one piece of evidence the selection did not touch, and a larger second stage buys an estimate that needs no prior at all.
What was counted
Counted over twenty thousand trials of eight arms, on the same draws as the normal shrinkage: the average errors +0.123, +0.020, +0.016 and +0.030 of the published, normal, t and mixture estimates with the arms alike, and +0.006, −0.103, −0.101 and −0.073 with one arm ahead by half a standard deviation; the mixture’s weight on “stands out” exceeding one half in 0.4% of trials with a real winner.
Counted over four thousand trials a point: the leads at which the priors let go, and the error of the normal and mixture shrinkage at four to sixteen arms, −0.156 and −0.063 at sixteen with a real winner.
Not claimed: that a fifth and a full standard deviation are the right settings for the mixture. They are the settings used across the borrowing essays, and a larger vague weight would recover more of the winner at more cost when the arms are alike; the trade has the same shape. Not claimed either that the t prior fails in general — a t prior whose scale was set from outside the trial, or with fewer degrees of freedom, would let go sooner; one scaled by the arms’ own spread cannot, because that spread is dominated by the arms that are alike.
Still open: a second stage sized by the first
Every estimate here treats the second stage’s thirty patients as fixed. The mixture’s weight on “stands out” is available at the interim, before the second stage is run, and it is exactly the quantity that says how much the published estimate can be trusted: low weight means the winner’s lead may be luck and more second-stage data would be worth having; high weight means the lead is probably real.
A design that sizes the second stage by that weight — more patients when the first stage cannot tell its shape, fewer when it can — would spend its patients where the winner’s curse is worst. Whether it would keep the second stage unbiased, since its size would now depend on the first stage’s data, and whether the saving at sixteen arms is worth the complexity at eight, are measurements of the same kind as these and have not been made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Estimates that are too alike — both name empirical bayes, mean squared error, partial pooling, shrinkage
- Where the borrowing goes — both name empirical bayes, mean squared error, partial pooling, shrinkage
- One population, or two — both name empirical bayes, partial pooling, shrinkage
- The correction that makes the estimate worse — both name mean squared error, selection bias, the winner's curse
- The small groups one centre protects — both name empirical bayes, partial pooling, shrinkage
- What the plug-in forgets — both name empirical bayes, partial pooling, shrinkage
Named objects
A flat tag is an object no other essay names yet.
Adaptive designEmpirical BayesMean squared errorPartial poolingPrior sensitivitySelection biasShrinkageThe winner's curse