Intervals for the findings
Worth reading first: What the correction corrects.
The procedures in this series decide which of twenty hypotheses to report. A report does not stop at the decision. Each finding goes out with an estimate and an interval — an effect of 2.9 standard errors, with a 95% interval from 0.9 to 4.9 — and the interval is the part a reader carries on into a review, a sample-size calculation or a decision about what to test next.
The ordinary interval is computed as though the finding had not been selected. It was. The effect reported by a study that reached significance is inflated because the selection kept the lucky draws, and an interval centred on a lucky draw misses for the same reason. Benjamini and Yekutieli named the quantity that counts it: the false coverage-statement rate, the expected share of the reported intervals that do not contain their effects. This essay counts it for the ordinary intervals and for the intervals built to hold it at 5%.
Thirty families, eighty-five findings
Thirty families of twenty tests, ten of them real effects of two standard errors, each run through Benjamini–Hochberg at 5%, make 85 findings between them. Drawn with its ordinary interval, its estimate at the centre and its true effect as a tick, each is a claim of the form a results table makes.
12 of the 85 ordinary intervals miss. Every one of them misses on the far side — the estimate further from zero than the effect it estimates — and none on the near side. 6 of the wider adjusted intervals miss. Six of eighty-five is 7.1%, above the 5% the adjusted intervals promise, and that is thirty families rather than a failure: the promise is about the average over families, and counted over twenty thousand of them at the same effect size the adjusted intervals miss 3.74% of the time. A figure of thirty families shows the shape of the misses, not their rate.
The side is the argument. An interval computed from an estimate nobody selected misses high and low equally often, by the construction that makes it a 95% interval. An interval computed from a selected estimate misses away from zero, because the finding was selected for being far from zero, and the draws that put it there are the draws that overshoot. Counted over twenty thousand families at the same effect size, the ordinary intervals miss away from zero in every case.
A false discovery makes the point in its sharpest form. A true null that Benjamini–Hochberg reports has a p-value at or below 5%, so its statistic is at least 1.96 standard errors from zero, and its ordinary interval — the statistic plus and minus 1.96 — cannot contain zero. The interval reported around a false discovery never contains the truth. Not usually; never. Every false finding on a list brings with it an interval that excludes its true effect, and the false discovery rate of a list is therefore a floor under the share of its intervals that miss.
How many reported intervals miss
With the ten real effects fixed at one standard error, 20.55% of the ordinary intervals reported with the findings miss. At two standard errors it is 11.59%; at three, 5.84%; at four, 5.07%. With the real effects spread about those sizes, it is 13.74%, 9.51%, 6.37% and 6.05%.
The shape is the winner’s curse read through intervals. At large effects nearly every real effect is found, so the findings are barely selected and their intervals barely conditioned; the share that miss approaches the 5% an unselected interval would. At small effects few are found, the ones found are the ones that drew far above their effect, and a fifth of the intervals attached to them are wrong — all wrong in the same direction, towards a larger effect.
The adjusted intervals, built for this rate, miss 3.42%, 3.74%, 3.78% and 3.75% of the time with the effects fixed, and 4.06%, 3.89%, 3.80% and 4.01% with them spread. The guarantee is kept at every size, with a margin that is itself a statement about how conservative the construction is.
Spreading the real effects lowers the share that miss at small sizes and raises it at large ones: at one standard error from 20.55% to 13.74%, because some of the real effects are then large enough to be found without much luck, and at four from 5.07% to 6.05%, because some are then small enough to be found only with it.
The finding reported first
A list of findings is read from the top, and a press release is written from its first line. So the interval that matters most is the one around the most prominent finding — the statistic furthest from zero among those selected.
That interval is the most selected of all, and it covers accordingly. With real effects of one standard error, the ordinary interval around the top finding contains its effect 2.38% of the time. At one and a half standard errors, 53.87%; at two, 72.36%; at three, 77.38%; at four, 77.43%. At one standard error Benjamini–Hochberg makes only 0.31 findings a family, each one far out, and many of them are false discoveries whose intervals, as above, cannot cover at all.
Even at four standard errors, where nearly every real effect is found and selection has almost stopped mattering for the list, the top finding covers only 77.43%. It is the largest of about ten estimates of effects of the same size, and the largest of ten normal draws lies 1.54 standard errors above their common mean on average — so its interval, centred on that draw, is centred on an overshoot built into the act of ranking.
The adjusted intervals do better and do not reach 95% for the top finding either: 83.43% at one standard error, 91.92% at one and a half, 93.48% at two, 90.29% at three and 88.24% at four. They hold the average over the list, which is what they promise, and the most prominent entry on a list is not an average entry. Their coverage of it falls as the effects grow because more findings make each adjusted interval narrower while the top finding remains the maximum.
How far the midpoint overshoots
Every miss above is a midpoint too far from zero, and the distance can be counted directly: the mean, over the reported findings, of how far each estimate lies beyond its effect in the direction of its sign. With real effects of one standard error it is 2.380 standard errors — the findings sit, on average, more than twice their effect’s size further out than the effect. At two standard errors the overshoot is 1.276; at three, 0.493; at four, 0.159.
That is the winner’s curse measured on a list rather than a literature, and the numbers say the same thing in the unit a reader uses. An interval of plus and minus 1.96 around an estimate that has overshot by 1.276 on average reaches back to the effect only when the overshoot happened to be small, and the far-side misses are the cases where it was not. The adjusted interval does not move the midpoint. It widens around the same overshot estimate until the overshoot is covered on the average list — which is why it holds the rate and why it is not a better estimate of any single effect.
Even at four standard errors the overshoot does not vanish. Nearly every real effect is found there, but the few that fall short of the threshold are the ones whose estimates came out low, and leaving them out raises the average of those that remain. Selection with a pass rate near one still selects.
What the wider interval costs
The false-coverage-rate interval replaces 1.96 by the point that leaves R times 5% divided by forty in each tail, R being the number of findings among twenty. With one finding that is the 99.875% point of the normal, and the interval is 6.047 standard errors wide against the ordinary 3.92. With two findings it is 5.614; with five, 4.995; with ten, 4.483; with all twenty a finding, 3.920 — the ordinary interval exactly, since nothing was selected out.
Reported at the effect sizes counted, the mean width runs from 5.778 at one standard error to 4.490 at four. It is a Bonferroni correction applied to the intervals and scaled by the share of tests that became findings, and its direction is the right one: fewer findings means harder selection, and harder selection needs a wider interval to say something true about what was selected.
The cost is real and it is paid in a currency results tables rarely show. The price of control counted what a correction costs in findings; this one costs nothing in findings — the list is Benjamini–Hochberg’s list either way — and costs precision in every interval on it. A list of ten findings at two standard errors, reported with adjusted intervals, reports each effect with an interval about half as wide again as the ordinary one.
There is a way to read the width that needs no table. With R findings among m tests at 5%, the adjusted interval is the Bonferroni interval for a family of m divided by R tests. One finding among twenty gets the interval Bonferroni would give twenty tests, a half-width of 3.023; ten findings get the interval for two tests, 2.241; twenty findings get the interval for one test, 1.960. A reader handed a list of three findings from twenty tests can therefore rebuild the adjusted interval by hand — the Bonferroni interval for six and two thirds tests, a half-width of 2.674 — and see at a glance how much of the reported precision the selection has already used.
Stricter selection, and a longer list
Benjamini–Hochberg is the lenient selector. With the same families — ten real effects of two standard errors among twenty — Holm reports 1.61 findings a family and Bonferroni 1.55, against Benjamini–Hochberg’s 2.77, and the ordinary intervals attached to their findings miss 14.13% and 14.38% of the time, against 11.59%. Their findings overshoot by 1.555 and 1.573 standard errors on average rather than 1.276. The adjusted intervals keep the rate for all three, at 3.74%, 2.34% and 2.30%.
So the correction that protects the decision makes the reported numbers worse. The price of control observed that a corrected analysis reports only the larger effects; here that observation has a size. A familywise procedure’s findings have cleared a higher bar, which means they needed luckier draws to clear it, and their intervals are centred further from the truth for exactly the reason the findings are more trustworthy as findings.
A longer list works the same way through the top finding. Among a hundred tests with fifty real effects of two standard errors, Benjamini–Hochberg reports 11.96 findings a family and 14.22% of their ordinary intervals miss; the interval around the top finding covers 27.08% of the time. The top finding is the largest of fifty estimates, and the chance that fifty estimates all stay within 1.96 of their effects is 0.975 to the fiftieth power, 28.2%. The same arithmetic with ten estimates gives 77.6%, beside the 77.43% counted for the top finding at four standard errors among twenty. With ten real effects among a hundred, the ordinary intervals miss 19.56% of the time, the findings overshoot by 1.904, and the adjusted interval is 6.494 standard errors wide.
Why reporting every interval is not enough
The position that answers corrections for multiple testing — report every test with its ordinary interval and let the reader judge — has an answer for intervals too, and it is the same answer the price of control gave for p-values. Within one paper, with all twenty intervals printed, a careful reader can see that the largest estimate is the one most likely to overshoot. Between papers, the findings travel without their neighbours: a review extracts the significant effects, a sample-size calculation uses the top one, and the ordinary interval that went with each finding goes with it, carrying its 95% label and its 72% coverage.
Pooling does not rescue the extracted findings either. A review that averages the reported effects of selected findings averages their overshoots, and at real effects of two standard errors those average 1.276 standard errors each, every one of them in the direction of the finding’s sign. Errors that share a sign do not cancel when they are averaged; they add. The pooled estimate settles, as studies accumulate, on the effect plus the overshoot, and its interval narrows around that sum rather than around the effect. The winner’s curse found the same thing about a literature assembled from significant studies, and a review assembled from the significant entries of many lists is that literature built one list at a time.
The adjusted interval is the version that survives extraction. It says, attached to each finding, how much the selection that produced the finding has to be allowed for, and it says it in the only place a travelling number keeps: its own width.
What a list of findings with intervals should carry
Intervals adjusted for the selection, or a plain statement that they are not. An ordinary interval beside a selected finding is a 95% interval in name. At two standard errors, one in nine of them misses, and every miss is an overstatement.
The number of findings the adjustment was computed from. The adjusted interval’s width depends on how many findings there were and how many tests they were selected from, and neither is recoverable from the interval alone.
The rule that selected them. The same effects reported from a Holm selection miss more often than from a Benjamini–Hochberg selection, and a reader cannot correct for a selection whose strictness is not stated.
A half-width a reader can check. The adjusted interval for R findings among m tests is the Bonferroni interval for m/R tests. Stating R and m beside a list lets anyone recompute it, and a list reported without them cannot be read either way.
Particular caution about the first finding. The most prominent finding’s interval covers least under either construction, and it is the one most likely to be quoted. Where a single effect is going to be carried forward, a fresh estimate from data that played no part in the selection is the only interval that makes no allowance at all.
What is proved here and what is counted
Proved, and quoted rather than derived. For Benjamini–Hochberg’s selection on independent tests, the adjusted intervals keep the false coverage-statement rate at or below 5% (Benjamini and Yekutieli, 2005). That a false discovery’s ordinary interval excludes its true effect of zero follows from the selection rule directly: a reported null has a statistic at least 1.96 from zero.
Counted. Every miss rate and coverage above, over twenty thousand families at each effect size, and the thirty families drawn in the first figure. The largest of ten normal draws averaging 1.54 standard deviations above their mean is an integral, computed rather than counted.
Particular to these families. Twenty or a hundred independent tests, ten or fifty real effects in standard-error units, selection by Benjamini–Hochberg, Holm or Bonferroni at 5%. Correlated tests would, by the counts for correlated families, make the misses lumpier from one list to the next, and whether they move the average share was not counted.
Still open: an interval for each finding that knows it was selected
The adjusted interval holds an average over a list. It does not promise 95% for any finding in particular, and the counts say the most prominent finding gets well under 95%. There is a construction that does promise it: an interval computed from the distribution of the estimate given that it was selected, which for a single finding just past the threshold is wide in one direction without limit and for a finding far past the threshold is close to ordinary. What such an interval covers for each finding, and how wide it becomes near the threshold where most small-effect findings sit, is the question these counts leave — and it is the one a reader holding a single finding from a long list is actually asking.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name closed form, confidence interval, coverage, monte carlo, multiple comparisons
- A flat point with more than one direction — both name closed form, confidence interval, coverage, monte carlo
- A schedule that reads the mean — both name confidence interval, coverage, monte carlo, selection effect
- A simulation that stops when it looks settled — both name closed form, confidence interval, coverage, monte carlo
- A standard error that knows about the instruments — both name closed form, coverage, interval width, monte carlo
- An interval that carries its scale — both name closed form, confidence interval, coverage, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Benjamini–HochbergClosed formConfidence intervalCoverageEffect sizeFalse discovery rateInterval widthMonte CarloMultiple comparisonsSelection effectThe winner's curse