An interval for the analyses admitted to
Worth reading first: Twenty analyses of nothing.
An interval for the winner built the interval a reported winner deserves. Twenty correlated analyses, the largest reported because it cleared the family-wise threshold of 2.846, and an interval inverted from the winner’s distribution given that it won — a normal truncated at a point set by the threshold and by the other nineteen statistics. That interval covered 94.90% of the time where the ordinary one covered 75.96%, and it paid for its honesty in width.
It needed the selection written down exactly. The essay ended on the case every exploratory analysis is really in: the analyst looked at more than the report admits, and the interval is conditioned on the family the analyst is willing to name. If the admitted family is a subset of the real one, it said, the truncation point is too low and the interval too narrow, by an amount set by how much of the real selection went unreported — and that is the one form of the counterfactual-family problem where a sensitivity analysis could be computed from the report alone.
Both halves turn out to be true, and the first is much smaller than expected. The second is the one worth having.
The selection as written and as it happened
The family is the one the earlier essays used: twenty analyses of one dataset, correlated at 0.6, one of them holding an effect of two standard errors. The analyst looks at all twenty and takes the largest in absolute value. The report admits to of them — the winner and others — and applies the family-wise threshold for : 1.960 if it admits to one, 2.199 for two, 2.480 for five, 2.671 for ten, 2.846 for all twenty. If the winner clears that threshold it is reported, with an interval conditioned on the admitted family.
The conditioning follows the earlier construction. Each other statistic is split into a part that moves with the winner and a part that does not, and the winner’s truncation point is the larger of the threshold and the largest of and over the others. Conditioned on the admitted family, those others are the admitted; conditioned on what happened, they are all nineteen. The true truncation point is therefore never below the admitted one, and the admitted interval is always the more generous to the winner.
Coverage barely moves
With one analysis admitted and an effect present, the interval conditioned on the admitted family covers the winner’s true mean 95.3% of the time. Conditioned on all twenty under the same reporting rule it covers 95.6%. With two admitted, 94.9% against 95.1%; with five, 94.9% against 95.1%; with ten, 95.1% against 95.2%. The ordinary interval around the same winners covers 67.7% at one admitted and 76.1% at twenty.
Independent analyses tell the same story. With one admitted, 94.4% against 94.9% for the interval that knows all twenty, and the ordinary interval 47.2%. With no effect anywhere, the admitted interval covers 94.1% at one admitted analysis and correlation 0.6, against 95.3% for the full conditioning.
So the predicted failure is real and small. Conditioning on nineteen fewer analyses than were run costs the interval about a point of coverage at most, at every admitted size and both correlations measured. The selective interval, which looked as though it needed the whole selection written down, needs mainly one number: the threshold the winner was actually reported against.
Where the truncation matters
The reason is in how the interval depends on its truncation point.
A winner observed at four, conditioned on a truncation at 1.96, has a lower limit of 1.654. Raise the truncation point, and the lower limit barely moves until the truncation approaches the observation: it reaches zero only when the truncation point is 3.019, and it falls to minus infinity as the two meet. A winner at five keeps a lower limit near three until the truncation passes about four. The truncation point matters when it is close to the winner and hardly at all when it is not.
And the unadmitted analyses rarely put it close. Their contribution to the truncation is , a rival’s excess over what the winner’s own value predicts for it, and a rival big enough to push the truncation near the winner is a rival nearly as large as the winner, after the part of it the correlation explains is taken away. Near-ties of that kind are uncommon, and when they happen the rival is often one of the admitted analyses anyway. Averaged over reported winners, hiding nineteen analyses of twenty moves the truncation point by 0.061 with correlated analyses and 0.170 with independent ones; hiding fifteen of twenty, by 0.012 and 0.019. The threshold, not the rivals, is doing the truncating.
What hiding does instead
If the hidden analyses do not spoil the interval, the natural conclusion is that an unreported family is harmless to selective inference. It is not, and the harm is in the step before the interval.
With no effect anywhere, a family of twenty correlated analyses produces a reported winner in 4.8% of datasets when all twenty are admitted — the family-wise rate, as designed. Admitting to ten, 7.4%; to five, 11.9%; to two, 22.9%; to one, 35.5%. The lower threshold of a smaller admitted family lets through the winner of a larger real one, and the report exists far more often than its stated error rate allows.
The interval then does its job on whatever is reported. Among the reported null winners, the admitted interval excludes zero 4.2% of the time at one admitted analysis against 2.8% at twenty — a little more often, because its truncation is a little low — and multiplied through, a family with nothing in it produces a reported winner whose interval excludes zero in 1.48% of datasets at one admitted analysis and 0.13% at twenty. The interval’s own promise, 95% coverage given that something was reported, is kept either way. The report’s promise, that a finding this strong arises one time in twenty from nothing, is broken by a factor of seven before the interval is drawn, and the claim that survives the interval by a factor of eleven.
So an honest interval around a dishonest threshold is honest about the wrong thing. The selective interval corrects for the winner’s having won; it cannot correct for there having been a contest the report does not mention, because the contest decides whether there is a winner to put an interval on.
Independent analyses hide more
The damage from an unreported family grows as the analyses become less alike, because a family of independent analyses offers more separate chances to clear a low threshold. With twenty independent analyses and no effect anywhere, a winner is reported in 63.8% of datasets against a threshold for one admitted analysis, and in 5.1% against the threshold for all twenty. The reported winner’s admitted interval excludes zero 3.9% of the time, against 2.5% for the interval that knows the whole family, and the rate at which a null family yields a reported winner whose interval excludes zero is 2.50% against 0.11% — a factor of about twenty-three, where the correlated family’s was eleven.
With an effect present the independent family shows the interval’s one real weakness. The admitted interval still covers 94.4%, but it excludes zero for 15.8% of reported winners where the interval conditioned on all twenty excludes zero for 11.3%. The truncation point moves by 0.170 on average when nineteen independent analyses are hidden, nearly three times the correlated family’s shift, and a lower truncation point buys a higher lower limit. The coverage stays honest because the extra claims are mostly about real effects; what grows is the share of reports that claim more than the full selection would have allowed.
The correction that makes the estimate worse found that a stricter threshold makes a reported estimate more inflated, not less, because the winner has been selected harder; the admitted-family threshold runs the other way, selecting less hard than the real contest did, and the selective interval, which corrects for whichever selection it is told about, corrects for too little of it.
What the reported p-value says
The same arithmetic applies to the p-value that usually stands beside the interval. A winner reported against the threshold for one analysis is quoted with that analysis’s own p-value, and under the null that p-value is supposed to be uniform: below 0.05 one time in twenty. A p-value that is not flat is not a p-value made the requirement exact. For the reported winner of an unreported family of twenty, the “p below 0.05” event happens in 35.5% of correlated null families and 63.8% of independent ones. The p-value is computed correctly for the analysis it names and is not a p-value for the selection that produced it, which is the whole of what the winner’s curse and what the correction corrects measured, arriving here as a rate of reports.
The selective interval is better placed than the p-value because it conditions on the report having happened. It cannot tell a reader how many reports like it were never going to be made, and that count — the denominator of the family — is exactly the quantity that went unreported.
A sensitivity analysis from the report
The earlier essay’s lead was that the counterfactual family could be priced from the report alone, and the measurement above says what the price should be stated in. Not in coverage, which barely moves, but in the one quantity the interval does depend on when the truncation is near the winner: whether the lower limit stays above zero.
A reader has the winner’s z and the threshold it was reported against. The interval’s lower limit reaches zero at a truncation point that depends on those two alone — 3.019 for a winner at four reported against 1.96 — and the truncation that unreported null analyses would impose has a distribution that depends on and the correlation alone. Setting one against the other gives the number of hidden analyses the finding can survive. A winner at 3.25 loses its lower limit’s sign at the median once seven analyses went unreported; at 3.5, ten; at four, 25; at 4.5, 60; at five, 171. A winner at three, reported against 1.96 with nothing else admitted, does not exclude zero even with no hidden analyses at all: its lower limit is −0.801.
That is a statement a report can print about itself and a reader can check. “This interval excludes zero unless more than twenty-five analyses of this kind were run and not reported” is a claim with content, and it is exactly the claim how many analyses there really were found an analyst is least able to make from memory — twenty analyses of one dataset were worth 11.37 independent ones at a correlation of 0.6, and the analyst who ran them is usually counting a smaller number.
Why the threshold is the thing to write down
The practical reading of all this is narrower and more useful than “write down every analysis”. A selective interval conditioned on the threshold the winner was actually reported against, and on whatever rivals are admitted, is nearly exact whatever else was run. What must be written down, before the data are seen, is the threshold — which is to say the size of the family the report will be judged as — because the threshold is what decides how often there is a report at all.
That is what naming a handful in advance priced from the other side, and what the price of control put an exchange rate on: a preregistration names analyses and earns the threshold for , and an analysis that ranges over more than it named has spent a threshold it did not pay for. The measurement here shows where the bill arrives. It does not arrive in the interval, which is robust to the unreported rivals; it arrives in the rate of reports, which is not.
What a report of a selected finding can state
The threshold, and the family size it was computed for. Everything else about the selection can be approximate; this cannot. A winner reported against 1.96 from a family of twenty is reported seven times as often from nothing as one reported against 2.846.
An interval conditioned on the admitted family. It keeps its coverage within a point of 95% whatever went unreported, at the correlations measured.
The correlation the analyses share, or a range for it. It enters twice: in the threshold a family of a given size earns, and in the sensitivity count below. Correlated analyses of one dataset — the same outcome under different covariate sets, the same comparison with different exclusions — are the usual case, and at a correlation of 0.6 a hidden family of twenty costs the report a factor of seven in its false-report rate; independent analyses, such as different outcomes, cost a factor of twelve and a half at the same threshold. A report that says which kind of family it came from lets a reader choose the right factor.
The number of hidden analyses the finding survives. It needs only the winner’s z, the threshold and an assumed correlation, and it turns the question “how many analyses were really run?” into a number a reader can hold the report to.
Counted here
Counted over twelve thousand families of twenty at each setting: at a correlation of 0.6 and an effect of two standard errors, the admitted-family interval covers 95.3%, 94.9%, 94.9% and 95.1% with one, two, five and ten analyses admitted, against 95.6%, 95.1%, 95.1% and 95.2% conditioned on all twenty; with no effect, a winner is reported in 35.5% of families at one admitted analysis and 4.8% at twenty, and one whose interval excludes zero in 1.48% and 0.13%.
Exact, given the construction: the truncated-normal lower limit’s dependence on the truncation point, and the zero points — 3.019 for a winner at four and 4.234 for a winner at five, reported against 1.96.
Not claimed: that real exploratory analyses are maxima over lists. An analyst who drops an outlier after seeing it, or switches outcome because the first looked flat, has made a selection that is not a single inequality on the reported statistic, and the robustness found here — which rests on the truncation point being set mostly by the threshold — has not been shown for those. The sensitivity count also assumes the unreported analyses were null and correlated like the reported ones; unreported analyses with effects of their own would be larger rivals. And the admitted rivals here were simply the first ones listed; an analyst who admits to the weakest rivals and hides the strongest moves the truncation point as far as it can move, which the average shift reported above does not describe.
Still open: a selection on something other than the statistic
Every selection here chose the largest statistic. A common selection chooses on something else: the analysis whose residuals passed a normality check, the model whose diagnostics looked cleanest, the exclusion rule that left the tidiest data — a choice made on a second statistic that is correlated with the one reported, not on the reported one itself.
The conditional interval then needs the selection event in terms of both statistics, and the robustness found here need not carry over: a diagnostic correlated with the reported statistic can truncate it in ways no threshold on the statistic describes. How far an interval conditioned only on the reporting threshold misses when the real selection ran through a correlated diagnostic, and whether the report can state enough about the diagnostic for a reader to redo the conditioning, have not been measured here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A ratio whose interval has to be the whole line — both name confidence interval, coverage, the winner's curse
- Five times in six — both name confidence interval, coverage, the winner's curse
- Intervals for the findings — both name confidence interval, coverage, the winner's curse
- A block size that changes — both name confidence interval, coverage
- A coverage table with its own error — both name confidence interval, coverage
- A flat point with more than one direction — both name confidence interval, coverage
Named objects
A flat tag is an object no other essay names yet.
Confidence intervalCoverageFamilywise error rateThe garden of forking pathsSelective inferenceSensitivity analysisTruncated normalThe winner's curse