An interval for the winner
Worth reading first: Twenty analyses of nothing.
The correction that makes the estimate worse found the uncomfortable trade at the end of every corrected exploratory analysis. Correcting for twenty analyses repairs the p-value by raising the threshold, and a result that clears a higher threshold is more selected, so the estimate reported beside the honest p-value is more inflated than before: 1.69 times the truth at two standard errors, against 1.35 uncorrected. The repairs it listed were partial, and the question it ended on was what an honest report of the winner would print.
One answer exists in the literature on selective inference and is simple to state. Build the interval from the distribution the winner actually has — the distribution of a statistic given that it won — rather than from the distribution it would have had if nobody had looked. This essay builds that interval for the family measured throughout, prices it, and finds that it does exactly what it promises and that the promise is far smaller than the one the ordinary interval makes.
The distribution of a winner
Twenty analyses of one dataset are correlated, at 0.6 here, and one of them holds an effect of two standard errors. The analyst reports the largest in absolute value if it clears the exact family-wise threshold, 2.846, which holds the chance of any false report at 5%. That happens in 22.34% of datasets, and in 87.24% of those the winner is the analysis that holds the effect.
The ordinary interval, the winner’s statistic plus or minus 1.96, is built for a statistic that was going to be reported whatever value it took. The winner was not. It was reported because it was large, so its distribution is its ordinary normal distribution with everything below the threshold cut away — a truncated normal — and an interval built for the uncut distribution sits too high.
The truncation has an exact form, and a simple one. Write each of the other statistics as its part that moves with the winner and its part that does not, ; those parts are independent of the winner. Given them, the event “this statistic won and cleared the threshold” is a single inequality, , where is the threshold or the point at which some other statistic would have overtaken the winner, whichever is higher. Given the rest of the data, the winner is restricted to , and — the true effect in the winning analysis — is the only unknown in it. The other nineteen effects drop out.
Inverting that truncated normal in gives an interval whose coverage, conditional on the analysis having won, is 95% by construction. At a correlation of 0.6 the threshold is the binding part of in all but 3.42% of winners; the rival statistics rarely come close enough to the winner to matter.
Coverage, and what it costs
The two intervals, counted over twenty thousand datasets at each effect size:
| effect in one analysis | winners reported | ordinary interval covers | conditional interval covers | conditional interval, times as wide |
|---|---|---|---|---|
| none | 4.95% | 0.00% | 95.25% | 4.59 |
| 1 | 7.08% | 9.10% | 94.50% | 3.74 |
| 2 | 22.34% | 75.96% | 94.90% | 2.48 |
| 3 | 58.33% | 91.93% | 94.68% | 1.80 |
| 4 | 89.35% | 95.14% | 95.09% | 1.29 |
The ordinary interval never covers the winner’s mean when there is no effect anywhere. Every reported winner cleared 2.846, so every ordinary interval lies entirely above 0.886, and the true mean is zero. That is the family-wise correction working as designed — it holds the chance of reporting anything to 5% — and it does nothing for the interval reported when something is reported. At two standard errors the ordinary interval covers three times in four; only when the effect is so large that winning required no luck does it recover its 95%.
The conditional interval covers within a point of 95% at every effect, and it covers equally well whether the winner was the right analysis or one of the nineteen nulls: at two standard errors 94.95% of the time for the right analysis and 94.56% for a wrong one, where the ordinary interval covers 87.07% and never. It pays in width, a median of 2.48 times the ordinary interval at two standard errors and 4.59 times when nothing is there. The width is what the selection cost; the ordinary interval was simply not charging for it.
The null picture is the cleanest statement of the difference between the two intervals. Every one of the thirty grey bars misses, and misses in the same direction, because every one of them belongs to a statistic that was reported for being large. The heavy bars cross zero in every row and, in more than a quarter of datasets, run off the left of the frame: a winner with no effect behind it usually cleared the threshold by very little, and a narrow win is precisely the case the conditional interval refuses to bound. A reader shown only the grey bars would see thirty findings. A reader shown the heavy bars would see thirty analyses that were largest by chance, most of them saying so.
What the interval declines to say
The coverage is restored and the price is not only width. It is the finding.
At an effect of two standard errors, every reported winner has passed a family-wise test at 5%, which a paper would print as a corrected significant result. The conditional interval excludes zero for 14.10% of those winners. For the other 85.90% it includes zero — and for 14.32% it has no lower limit at all. At three standard errors, 33.04% exclude zero; at four, a majority, 61.76%. With no effect anywhere, 2.63% exclude zero, which is the one-sided 2.5% an honest interval should leave, plus noise.
That looks like a contradiction and is not one. The family-wise test answers “is anything in these twenty non-zero?” and it answers yes, correctly, at a controlled rate. The conditional interval answers “what is the effect in the analysis that happened to come out largest, given that it came out largest?” — and being the largest of twenty correlated statistics is something a null analysis does often enough that clearing the threshold by a little is weak evidence about this analysis. The first question is the one the correction was designed for. The second is the one the reported estimate claims to answer.
The consequence for reading a corrected exploratory result is sharp. “Significant after correction for twenty analyses” licenses the statement that something in the family is non-zero. It does not, most of the time, license a statement that the reported analysis is, and an interval that respected the selection would say so in print.
The margin decides everything
The conditional interval depends on the winner’s statistic and on , and since is almost always the threshold, it depends on one number: how far the winner cleared it.
A winner that cleared by a hair carries almost no information about its mean. The truncated normal just above its cut-off looks nearly the same whatever is below it — a very negative and a zero both make “just above the threshold” the most likely place to land, given that the statistic landed above it — so the data cannot rule out a mean far below zero. The lower limit does not exist until the winner clears the threshold by 0.10, and the interval does not exclude zero until the margin is 1.03, a winner of about 3.88.
Take a winner at 3.4, which a reader would call convincingly significant: 0.55 above the family-wise threshold and far above 1.96. The ordinary interval is 1.44 to 5.36. The conditional interval is −3.38 to 5.24, with a median estimate of 2.48, and the conditional test of a zero effect in that analysis has a p-value of 0.152. The upper limit has barely moved — the selection says little about how large an effect could be — and the lower limit has moved almost five units. The shape is the one the winner’s curse predicts, drawn as an interval: selection pushes estimates up, so an honest interval stretches down, and it stretches furthest exactly where the estimate was most likely to have been pushed.
The asymmetry also tells a reader where the ordinary interval’s error lives. At the top it is nearly right; the whole error is that its bottom sits two or three units too high, which is exactly the region a claim of a positive effect is read from.
The estimate the interval comes with
The conditional interval has a natural point estimate, its median: the at which the observed winner sits at the middle of its truncated distribution. It is median-unbiased conditional on selection — it lies above the truth in 49.62% of reported winners at two standard errors, which is a half up to the counting — and that repairs exactly what the correction made worse. The median of those estimates is 1.796 against an average true mean of 1.745 among the winners.
It is not mean-unbiased and should not be averaged. Near the threshold the median estimate falls without limit, since a winner on the cut-off is equally consistent with any mean far enough below it, and a few of those drag any average far below zero. That is a fact about what the data say rather than a defect of the estimator: a winner that barely won says that its own analysis might hold no effect at all, or less than none, and a single number cannot carry that without being either misleading or unstable. The median chooses unstable.
For a reader the practical form is simpler than the theory. The reported effect of a winner should be shrunk towards the threshold by more the closer it is to it, and a winner within a tenth of a standard error of the threshold should be reported with no point estimate at all — the interval is the finding, and it is unbounded below.
What an exploratory report should print
The question left open was what an honest exploratory paper looks like, and the conditional interval answers most of it. The selection is visible in the interval rather than corrected for and then forgotten: the width is the price of having looked twenty times, and a reader sees it on the page next to the estimate.
A report of the winner, then, carries four numbers where it now carries two. The family-wise p-value, which is still correct for the question it answers. The ordinary estimate and interval, for comparison and because readers will compute them anyway. The conditional interval and median. And the margin by which the winner cleared the threshold, which is what makes the other three readable, since the conditional interval is a function of it.
Two features of the construction make this cheaper than it sounds. It needs no knowledge of the other nineteen effects, because they drop out given the parts of the other statistics that do not move with the winner; and it needs only the one statistic and the threshold whenever, as here in 96.58% of winners, the threshold rather than a close rival sets . Where a rival was close, the rival’s value enters, and the interval is correspondingly wider — which is right, since a close second means the winner was nearly somebody else.
It does not solve the problem naming analyses in advance solves, and it is worth being precise about the difference. A prespecified analysis needs no conditioning, because it was going to be reported whatever it showed; its ordinary interval is correct and its estimate is not selected. The conditional interval is the honest report of an analysis that was not prespecified. Its width, set against the ordinary interval a prespecified analysis would have earned, is a direct measure of what the lack of a plan cost — at two standard errors, a factor of 2.48 in interval width, and at no effect a factor of 4.59.
What the next study should be sized from
A reported winner is rarely the end of its story. It becomes the effect a confirmatory study is sized to detect, and the inflation measured in the essay on the correction goes straight into that sizing: an effect reported at 1.69 times its true size produces a replication with barely a third of the sample it needs, which is why the p-value a replication gets is so often disappointing.
The conditional median is the better number to size from, for a reason the table of coverage makes clear: it is centred on the truth given selection, so a study sized from it is sized for an effect that is as likely to be too small as too large. The conditional interval is better still, because a planner who sizes for its lower half — say, its lower quartile — is buying protection against exactly the case the ordinary estimate hides. When that lower half includes zero, as it does for most winners at two standard errors, the honest conclusion is that the exploratory study has not established an effect worth sizing for, and that the next study’s first job is to find out whether there is one. That is a more modest plan than the ordinary estimate would suggest and a much more likely one to succeed.
Where the correction’s two failures meet
The earlier essays in this line of argument found the family-wise correction in two minds. Twenty analyses of nothing needed it to stop a search from reporting noise; how many analyses there really were priced it exactly; and then the correction was found to make the reported effect worse while making the p-value honest.
The conditional interval shows the two failures were one. Both come from reporting the winner as though it had not been chosen. The correction repairs half of that by moving the threshold, which fixes the error rate of the decision to report; the conditional interval repairs the other half by moving the interval, which fixes the coverage of what is reported. Neither alone is a complete account of an exploratory search, and together they are consistent — the interval excludes zero for about 2.5% of winners when nothing is there, the one-sided rate a 95% interval should have, and the test’s 5% is the two-sided rate of the same kind of mistake.
What they cannot do together is manufacture evidence the search did not collect. A winner that cleared a family-wise threshold by a tenth of a standard error has been shown, by a correct test, to come from a family in which something is probably going on, and the correct interval for its own effect runs from nowhere to a little above where it landed. That is not a weakness of the method. It is an accurate description of what twenty looks at one dataset produce.
What conditioning on the win buys, and the selections it cannot write down
Conditional on selection, the interval inverted from the truncated normal covers the winner’s mean within a point of 95% at every effect from none to four standard errors, for the right analysis and for a wrong one alike; the ordinary interval covers 0.00% at no effect and 75.96% at two standard errors.
At two standard errors the conditional interval is a median 2.48 times as wide as the ordinary one, excludes zero for 14.10% of reported winners and has no lower limit for 14.32%.
Its lower limit exists only for winners that clear the threshold by more than 0.10, and excludes zero only beyond a margin of 1.03. A winner at 3.4 has a conditional interval of −3.38 to 5.24.
Its median estimate is median-unbiased given selection — above the truth in 49.62% of winners — and unbounded below near the threshold.
The coverage figures are counts over twenty thousand seeded families at each effect, drawn from the same one-factor model of twenty analyses correlated at 0.6 used throughout; the threshold is exact, from the one-dimensional integral; and each interval is an exact inversion of the truncated normal, computed in logarithms of the tail so that limits far below zero are exact rather than underflowed.
Not claimed: that the conditional interval is the only honest report — hybrid intervals that condition less strongly and give finite limits in exchange for a little coverage exist, and are not measured here. Not claimed either that the construction extends unchanged to selections an analyst cannot write down. It needs the selection event as an inequality, which a “largest of twenty” rule provides and a garden of forking paths, whose paths were never enumerated, does not.
Still open: an interval for a selection nobody wrote down
The construction here conditions on a rule — report the largest if it clears 2.846 — and the rule has to be known exactly for the truncation point to be computed. A real exploratory analysis rarely has one. The analyst looked at several outcomes, tried two exclusion rules, dropped an outlier after seeing it, and reported what looked cleanest; the selection happened, and it was not a maximum over a list.
Conditional inference needs the selection as an event in the data’s space. What is not known is how much of the benefit survives when the event is approximated — conditioning on “the largest of the analyses the analyst admits to” when the true selection was over more, or different ones. If the admitted family is a subset of the real one, the truncation point is too low and the interval too narrow, by an amount set by how much of the real selection went unreported. That is the counterfactual family problem again, asked of an interval rather than a threshold, and it is the one form of it where a sensitivity analysis — how wrong would the admitted family have to be for this interval to lose its lower limit — could be computed from the report alone.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A ratio whose interval has to be the whole line — both name confidence interval, coverage, the winner's curse
- Five times in six — both name confidence interval, coverage, the winner's curse
- Intervals for the findings — both name confidence interval, coverage, the winner's curse
- A block size that changes — both name confidence interval, coverage
- A coverage table with its own error — both name confidence interval, coverage
- A flat point with more than one direction — both name confidence interval, coverage
Named objects
A flat tag is an object no other essay names yet.
Confidence intervalCoverageFamilywise error rateThe garden of forking pathsMedian-unbiased estimateSelective inferenceTruncated normalThe winner's curse