An interval that knows its finding was found
Worth reading first: What the correction corrects.
Intervals for the findings took twenty tests, ten of them real, ran Benjamini–Hochberg over them, and gave each finding the interval it usually goes out with. With real effects of two standard errors, 11.59% of those ordinary intervals missed their effect, every miss on the far side, and the interval around the most prominent finding — the first line of the press release — covered 72.36% of the time. Intervals widened for the number of findings held the share that miss under 5%, which is what they promise: an average over the list.
The essay ended on the construction that promises something about each finding rather than about the list. An estimate that was selected for clearing a threshold is no longer normally distributed around its effect; it is a normal given that it cleared the threshold, and an interval can be read off that conditional distribution exactly as the ordinary one is read off the unconditional. What it covers, finding by finding, and what it costs to cover it, are the subject here.
The threshold a finding had to clear
Benjamini–Hochberg looks like a procedure about a whole list, and it is. But for any one test, with the other nineteen statistics held as they came out, the procedure’s verdict on that test is a plain threshold: it is rejected if and only if its p-value is at or below some that the other p-values determine — the largest level at which enough of the others sit below it. Converted to the statistic, test is a finding exactly when , with fixed by the other tests.
That makes the selected estimate’s distribution simple. The test’s statistic is with its effect in standard-error units; given that it was selected, it is that normal truncated to . For a candidate effect the conditional distribution function is
and it decreases as grows. The 95% interval is the set of with , read off by bisection — the same inversion that gives the ordinary interval when .
What the interval looks like
The shape is the whole argument. Far past its threshold a finding’s interval is the ordinary one: at an estimate 3.5 standard errors past a threshold of 2.8 it runs from 4.31 to 8.26, against the ordinary 4.34 to 8.26. The truncation is so far below that it removes almost nothing, and the estimate’s distribution given selection is very nearly its distribution.
Just past the threshold the interval is unrecognisable. At an estimate of 2.81 — a hundredth past the line — it runs from −0.58 to 0.95. It is 1.53 wide, narrower than the ordinary 3.92, and it lies wholly below its own estimate, around zero. The earlier essay expected an interval for a finding just past the line to be wide in one direction without limit; it is the opposite, narrow and in a place the estimate is not.
That is not a numerical accident. A statistic that lands a hundredth past a threshold carries almost no information about its size beyond the fact that it cleared — and among effects that clear 2.8 only by a hair, small effects are far more common than large ones, since a large effect would usually have cleared it by more. All that is left to read is which side of zero the estimate fell, and how likely an effect of each size is to put it barely past the line. Effects near zero do that as often as anything, and the interval sits there.
Between the two, at an estimate one standard error past the threshold, the interval runs from 0.18 to 5.73: wider than the ordinary interval, shifted towards zero, and only just excluding it. The interval is asymmetric everywhere near the line, because the truncation acts on one side.
Each finding, covered
The same thirty families the earlier essay drew make eighty-five findings. Twelve of their ordinary intervals miss, every one of them on the far side. Five of the selective intervals miss. Sixty-three of the eighty-five selective intervals contain zero.
Over twenty thousand families the counts settle. With real effects of two standard errors, the selective interval covers its finding’s effect 94.8% of the time; with effects of one standard error, 95.8%; with three, 95.0%. Those are per-finding rates — the chance that a finding, once found, has an interval containing its effect — and they sit on 95% whatever the effects are, because the construction conditions on exactly the event that made the finding a finding. The ordinary interval’s per-finding coverage at two standard errors is 87.0%; the false-coverage-rate interval keeps the list’s average miss rate at 3.7%.
Coverage by finding also includes the false discoveries. Among findings whose effect is truly zero — 3.9% of findings at two standard errors — the selective interval covers zero 94.8% of the time. An ordinary interval around a false discovery cannot contain zero at all: the estimate cleared a threshold above 1.96, and the interval is centred on it with half-width 1.96.
The promise is an average over where the finding landed
The 95% is not uniform across findings, and it is not meant to be. Half of all findings land within half a standard error of their threshold, and among them the selective interval covers 95.2%, the ordinary one 94.0%. Those that land between half a standard error and one past it are covered 99.8% of the time. Between one and one and a half, 99.1%. Between one and a half and two, 83.5% — and the 1.7% of findings that land two to three standard errors past their threshold are covered 0.4% of the time.
Those last findings are the overshoots. With every real effect two standard errors and thresholds near 2.7, an estimate two to three standard errors beyond its threshold sits more than two and a half standard errors from its effect, and no interval that is close to ordinary that far out can reach back. The selective interval spends its 5% of misses there, on the rare findings that overshot most, and over-covers the common ones nearer the line. The guarantee is the average over where a finding lands, given that it was selected, which is the guarantee a reader holding one finding without knowing whether it overshot can use.
The ordinary interval’s pattern is the reverse of what the earlier essay’s averages suggested. It is close to 95% on the findings just past the line — 94.0% — and fails on the ones further out, 78.3% between one and one and a half standard errors past and 5.1% between one and a half and two. The findings a reader is most impressed by are the ones whose ordinary intervals are least to be trusted.
The most prominent finding
The most prominent finding is the largest statistic among those selected, and the earlier essay found its ordinary interval covering 2.38% of the time with real effects of one standard error, 72.36% at two and 77.38% at three. The selective interval around the same finding covers 95.3%, 91.3% and 82.7%.
That is much better and not 95%, and the reason is the same one that made the ordinary interval fail: the most prominent finding has been selected twice. Once by Benjamini–Hochberg, which the selective interval accounts for, and once by being the largest of the findings, which it does not. At effects of one standard error there is usually only one finding a family, so the second selection is empty and the interval covers 95.3%; at three standard errors there are nearly eight, and being the largest of eight is a selection of its own. An interval for “the top finding” would have to condition on being top as well — a truncation that depends on every other selected statistic — and the construction here does not attempt it.
What it costs
The selective interval is wider than the ordinary one. Its median width is 5.15 standard errors with real effects of one, 5.06 at two and 4.90 at three, against the ordinary 3.92. Near the threshold, where half the findings are, it is 4.79; far past it, 4.21. The width is not spread evenly: it comes almost entirely on the side towards zero, where the truncation has made the estimate uninformative.
And it contains zero often. With real effects of one standard error, 90.3% of selective intervals reach past zero; at two, 76.2%; at three, 49.8%. These are findings — Benjamini–Hochberg rejected each of their null hypotheses at a false discovery rate of 5% — and most of them, read with an interval that accounts for how they were found, cannot exclude an effect of nothing.
That is not a contradiction, and it is the most useful thing the interval says. The false discovery rate is a statement about the list: at most a twentieth of these rejections are expected to be false. The selective interval is a statement about each entry: given that this one was selected, its effect is somewhere in this range. A list can be mostly right while each entry is individually uncertain, and with effects of two standard errors that is the situation. The procedure found ten real effects in twenty tests, and it cannot say, entry by entry, which of the nearly three findings a family made is the large one.
The same move made for a single winner
This construction has been made once before in this collection, for a different selection. An interval for the winner took the largest of twenty correlated analyses, reported because it cleared a family-wise threshold, and built its interval from the distribution of that largest statistic given that it won. The ordinary interval covered 75.96% at an effect of two standard errors; the conditioned one 94.90%, at 2.48 times the width, and it excluded zero for only 14.10% of the winners.
The two results agree in their shape and differ in their numbers for a clear reason. The winner was selected by a single, severe event — being the largest of twenty and clearing a family-wise line — and conditioning on it costs a great deal of width and almost all of the winners’ claims to be non-zero. A Benjamini–Hochberg finding is selected by a milder event, clearing a threshold the false discovery rate sets lower, and conditioning on it costs a third more width and leaves a quarter of findings at two standard errors excluding zero. How much an interval must give back is set by how hard the finding had to work to be found.
It is also the same accounting as the interval after the choice, which found a forecast interval losing four and a half points of coverage to a model chosen from the same forty observations. Wherever a choice is made with the data and then reported without it, an interval that ignores the choice is centred on the choice’s favourite outcome; an interval that conditions on it is centred on what the choice leaves possible.
Three intervals, three promises
The ordinary interval promises 95% for a statistic nobody selected, and is used for statistics that were. The false-coverage-rate interval promises that the share of a list’s intervals that miss averages at most 5% — a promise about the list, which it keeps, at 3.7% here, while the most prominent finding’s interval covers far less. The selective interval promises 95% for each finding given that it was found, and keeps that too, at 94.8%, while the most prominent finding — selected again — is covered 91.3%.
They are three different questions, and the choice between them is a choice about what the reader will do. A reader who will act on the whole list, following up every finding, is served by the false-coverage-rate interval. A reader who will act on one finding — the one in their own field, the one that will be replicated — is served by the selective interval. The winner’s curse is the case that neither fully repairs: the reader who acts on the most prominent finding has selected again, and every interval here under-covers what that reader picked.
What the correction corrects and two different promises set out the same division for tests: family-wise control for the reader who cannot tolerate one false finding, false discovery control for the reader who reads the list. Intervals inherit the division, with a third column for the reader who reads one line.
What a list of findings should report
Report each finding’s selective interval beside its estimate. It needs nothing the analysis does not already have — the finding’s statistic and the threshold the procedure applied to it, which the other p-values fix — and it is computed by a one-dimensional bisection. It covers each finding at the stated rate whatever the effects are.
Say how far past its threshold each finding landed. A finding that cleared its line by a tenth of a standard error has an interval that sits below its own estimate and usually includes zero; one that cleared it by two has an interval close to the ordinary one. The margin is the most informative single number a list can add to a finding.
Do not treat a zero-containing selective interval as a failed finding. The finding was made at a false discovery rate, and that rate still holds for the list. What the interval adds is that this entry’s effect may be small, and that its estimate is likely to shrink on replication, by about the amount how far the midpoint overshoots measured.
Condition again for the top finding, or say that it was not done. The most prominent finding is selected twice, and an interval that accounts for the first selection covers it 82.7% to 95.3% of the time depending on how many findings there are.
What is exact and what is counted
Exact: given the other statistics, Benjamini–Hochberg rejects a test exactly when its statistic clears a threshold the others determine; the selected statistic is then a normal truncated at that threshold, and the selective interval covers each finding with probability 95% conditional on its selection, at any configuration of effects.
Counted, over twenty thousand families of twenty tests at each effect size: per-finding coverage of 95.8%, 94.8% and 95.0%; the most prominent finding’s coverage of 95.3%, 91.3% and 82.7% against the ordinary interval’s 2.4%, 72.4% and 77.4%; the median widths and the shares containing zero; and the coverage by margin past the threshold.
Not claimed: anything about correlated tests. The threshold is a fixed function of the others only because the tests are independent here; with correlated statistics the selection event for one test depends on its own statistic through the others, and the truncation is no longer a single cut.
Still open: an interval for the finding at the top of the list
The most prominent finding was covered 82.7% of the time with effects of three standard errors, because being the largest of eight findings is a second selection the interval did not account for. Conditioning on that second event as well — the statistic being both past its own threshold and larger than every other selected statistic — is a truncation whose boundary depends on the others’ values, and its interval can be computed by the same inversion with a more complicated cut.
Whether that interval covers the top finding at 95% at every effect size, how much wider it is than the once-conditioned interval, and whether it too reaches past zero for most top findings at moderate effects, are measurable on the same families and have not been measured here. It is the interval a press release would need, since a press release reports the largest finding and not a finding chosen at random from the list, and the coverage that matters to its reader is the coverage of that one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Right for the wrong reason — both name conditional inference, confidence interval, coverage, interval width, selection effect
- Testing for a flat point first — both name confidence interval, coverage, interval width, selection effect
- The cleanest of the significant — both name conditional inference, confidence interval, coverage, multiple comparisons
- A block at every starting row — both name confidence interval, coverage, interval width
- A count that bets against its interval — both name conditional inference, confidence interval, coverage
- A coverage table with its own error — both name confidence interval, coverage, multiple comparisons
Named objects
A flat tag is an object no other essay names yet.
Benjamini–HochbergConditional inferenceConfidence intervalCoverageFalse discovery rateInterval widthMultiple comparisonsSelection effectThe winner's curse