The analyses that were available and not run

A winner chosen by its diagnostic

An analyst runs twenty correlated analyses and reports the one whose diagnostics look cleanest, if it clears 1.96 — and a diagnostic is rarely independent of the statistic it checks. With no effect anywhere, a diagnostic correlated 0.9 with its statistic lets a winner through in 13.07% of families against 2.52% for an unrelated one. The interval conditioned on clearing 1.96 still covers 93.27%, but its misses have moved: it excludes zero in 4.69% of reported winners, where a 95% interval should do so in 2.5%. Conditioning on the diagnostic as well repairs it, and so does something a reader can do alone — weight the reported statistic by the chance an analysis like it wins the diagnostic contest, from three numbers a report can state.

Worth reading first: Twenty analyses of nothing.

An interval for the analyses admitted to conditioned a reported winner’s interval on the family the analyst admitted to rather than the one actually run, and found the coverage barely moved: the truncation that selection imposes on a winner is set by the reporting threshold far more than by its unseen rivals. Every selection there chose the largest statistic. The essay ended on the selection most exploratory analyses actually make, which is a choice on something else: the specification whose residuals passed a check, the model whose diagnostics looked cleanest, the exclusion rule that left the tidiest data.

Such a diagnostic is rarely independent of the statistic it sits beside. A specification that fits the data better usually has smaller residuals, and smaller residuals make a larger test statistic; a choice of outlier rule that tidies the residuals also moves the estimate. So the selection runs through a second statistic correlated with the reported one, and the question was whether the threshold’s dominance survives it. It does, as far as coverage goes. What it does not survive is the side the misses fall on, which is the side a reader acts on.

How often a diagnostically chosen winner's interval excludes zero, against the diagnostic's correlation with the statistic, no effect anywhereTwenty analyses correlated 0.6; the one with the best diagnostic is reported if it clears 1.96. A winner is reported in 2.52%, 4.16%, 6.81%, 13.07% of families at γ = 0, 0.3, 0.6, 0.9. Its interval excludes zero: conditioned on clearing 1.96, 1.69%, 2.28%, 3.30%, 4.69%; conditioned on the diagnostic as well, 1.69%, 1.86%, 1.83%, 2.31%; rebuilt from the report's three numbers, 1.69%, 1.80%, 2.13%, 2.52%.0%1%2%3%4%5%6%00.3000.6000.900the diagnostic's correlation with its analysis's statistic, γshare of reported winners whose interval excludes zeroconditioned on clearing 1.96conditioned on the diagnostic toorebuilt from K, ρ and γ40,000 families of twenty a pointthe target line is 2.5%, one side of 95%
Fig. 1 How often a reported winner’s interval excludes zero, against the correlation γ between each analysis’s diagnostic and its statistic, when twenty analyses correlated at 0.6 are run, the one with the best diagnostic is reported if it clears 1.96, and nothing has any effect. Three intervals: conditioned on clearing 1.96 alone, conditioned on the diagnostic as well, and rebuilt by a reader from three reported numbers. The dashed line is 2.5%. The slider puts an effect of two standard errors into one analysis.

The selection, and why it is still a truncation

The family is the one these essays have used throughout: twenty analyses of one dataset, their statistics correlated at 0.6. Each analysis now also produces a diagnostic score, higher meaning cleaner,

di=γzi+1−γ2 ei,d_i = \gamma z_i + \sqrt{1 - \gamma^2}\, e_i,

correlated γ\gamma with its own statistic and carrying independent noise eie_i of its own. The analyst picks the analysis with the highest diagnostic, and reports its statistic if it clears 1.96, with the interval an analyst who knows about selective inference would give: the normal truncated at 1.96, inverted in its mean.

The event “analysis jj has the best diagnostic” can be written down exactly. Hold fixed everything independent of zjz_j — each other statistic’s part wi=zi−ρzjw_i = z_i - \rho z_j that does not move with it, and every noise term eie_i — and for γ>0\gamma > 0 the event is a set of lower bounds on zjz_j:

zj  ≥  γ wi−1−γ2 (ej−ei)γ(1−ρ)for every other analysis i.z_j \;\ge\; \frac{\gamma\, w_i - \sqrt{1-\gamma^2}\,(e_j - e_i)}{\gamma(1 - \rho)} \qquad \text{for every other analysis } i.

Given those, zjz_j is a normal truncated at the larger of 1.96 and the largest of the bounds, whatever the other analyses’ means are. That is the exact conditional interval, the same construction as an interval for the winner with the diagnostic in the place the rivals’ statistics occupied. It needs every other analysis’s statistic and diagnostic, so it is the interval a reader of the report can never compute; it is the yardstick.

Why a diagnostic tracks its statistic

The correlation γ\gamma is the whole of the problem, so it is worth saying where it comes from, because it is almost never zero and the analyst choosing on diagnostics rarely thinks about it.

A specification that fits the data better leaves smaller residuals, and a regression’s test statistic is its estimate divided by a standard error built from those residuals. A choice among models by residual variance, by an information criterion, or by a goodness-of-fit statistic is therefore partly a choice of the model whose standard error is smallest, and so, at a fixed estimate, whose statistic is largest. A choice by a normality check on the residuals is subtler. The quantile plot of residuals that look too normal is dominated by the largest residuals, and a specification or an exclusion rule that removes a large residual both passes the check and changes the estimate — in whichever direction that point was pulling it. Across specifications, the tidiest-looking residuals and the largest statistic tend to come from the same handful of choices.

An outlier rule is the plainest case. Excluding the points that most disagree with the fitted line makes the residuals cleaner by construction and moves the estimate towards whatever the remaining points say, and a robust loss and a far x found how much a single point at the edge of a design can move a fitted slope. The analyst who chooses the exclusion rule that “cleans the data best” is choosing, with a correlation that depends on the data, the rule under which the estimate moved most in the direction of the points kept. None of this requires looking at the p-value, which is why it survives a preregistration that fixes the hypothesis but leaves the specification to be chosen by diagnostics.

How often a winner is reported

The first effect of choosing by a correlated diagnostic has nothing to do with intervals. With no effect in any analysis and a diagnostic unrelated to the statistic, the chosen analysis is effectively picked at random and clears 1.96 in 2.52% of families — one side of a 5% test. As the correlation rises the diagnostic increasingly picks analyses whose statistics happen to be large: a winner is reported in 4.16% of families at γ=0.3\gamma = 0.3, 6.81% at 0.6 and 13.07% at 0.9.

That is the garden of forking paths with the forks hidden inside a criterion that sounds unrelated to significance. An analyst who never looked at a p-value while choosing, and chose only for clean residuals, reports a significant result five times as often as the one-analysis rate when the diagnostic tracks the statistic closely. How many analyses there really were priced this for a choice made on the statistic; here the same inflation arrives through a choice made on something else, and it scales with how much the something else shares with the statistic.

Coverage holds and the misses move

Which side a diagnostically chosen winner's interval misses on, with no effect anywhere, against the diagnostic's correlation with the statistic. The threshold-only interval misses below in 1.69%, 2.28%, 3.30%, 4.69% of reported winners at γ = 0, 0.3, 0.6, 0.9, and above in 3.28%, 3.06%, 2.61%, 2.05%. The rebuilt interval misses below in 1.69%, 1.80%, 2.13%, 2.52% and above in 3.28%, 2.88%, 2.42%, 1.91%.
Fig. 2 With no effect anywhere, the share of reported winners whose true mean lies below the interval and above it, for the interval conditioned on clearing 1.96 alone and for the one a reader rebuilds from the report, against the diagnostic’s correlation with the statistic. A 95% interval misses 2.5% of the time on each side.

The interval conditioned only on clearing 1.96 keeps most of its coverage, as the earlier essay’s admitted-family interval did: 94.65% at γ=0.3\gamma = 0.3, 94.09% at 0.6 and 93.27% at 0.9. On coverage alone it would pass any check a reader could run.

Its misses are no longer balanced. At γ=0.9\gamma = 0.9 the true mean lies below the interval for 4.69% of reported winners and above it for 2.05%. Below the interval, with no effect anywhere, means the interval’s lower limit is above zero: the report says the effect is positive and it is not. A 95% interval should say that 2.5% of the time; this one says it nearly twice as often, while its total coverage reads 93.27% and looks nearly right. The exact interval that conditions on the diagnostic splits its misses 2.31% and 2.70%.

The reason is the direction of the truncation the diagnostic adds. A diagnostic that tracks the statistic favours the analyses whose statistics are large, so the chosen statistic’s distribution is pushed upwards beyond what the threshold alone accounts for. An interval that knows only about the threshold reads the pushed-up statistic as evidence of a larger mean, and its lower limit sits too high. The upper misses fall at the same time, because the same over-reading lifts the upper limit too. Coverage — the sum — barely moves; the split shifts from the upper side to the lower, and the lower is the side that turns into a claim.

A truncation that is usually nothing

How far a diagnostic correlated 0.9 with the statistic moves a reported winner's truncation point, with no effect anywhere. Of 5,229 reported winners, 76.0% have no truncation beyond 1.96. The rest are moved by up to 2.30; 6.4% of all winners by more than half a standard error and 1.6% by more than one. The mean move is 0.090.
Fig. 3 How far the diagnostic moves a reported winner’s truncation point above 1.96, with no effect anywhere and a diagnostic correlated 0.9 with its statistic: the share of reported winners it does not move at all, and the distribution of the move for the rest.

The truncation the diagnostic adds is small on average, 0.090 standard errors at γ=0.9\gamma = 0.9, and its median is exactly zero: for 76.0% of the reported winners the bounds it imposes all lie below 1.96, and the exact interval is the threshold-only interval. For 6.4% of them it moves the truncation point by more than half a standard error, and for 1.6% by more than one. The damage comes from the families where the chosen analysis beat its rivals on the diagnostic narrowly and the bounds bite, sometimes by a large amount. That is why a sensitivity analysis built on the typical truncation — the median — reproduces the threshold-only interval and repairs nothing, and why the earlier essay’s finding that the threshold does the truncating can be true on average and still leave a one-sided failure in a minority of reports that is large enough to double the rate of false claims.

More specifications, more of it

Twenty is a round number, and the effect grows with the number of specifications the diagnostic is asked to choose among. At a correlation of 0.9, five specifications with nothing in them produce a reported winner in 7.12% of families, twenty in 13.07% and fifty in 17.85%; the threshold-only interval’s lower-side misses run 3.97%, 4.69% and 5.56%. At 0.6 the same three families give 4.75%, 6.81% and 8.38% reported, and lower-side misses of 3.68%, 3.30% and 4.21%.

The mechanism is the familiar one of a maximum over more candidates, with a twist. A choice on the statistic itself grows its inflation with the number of candidates without limit, which is what twenty analyses of nothing counted; a choice on a diagnostic grows it more slowly, because each additional specification adds a candidate whose diagnostic is only partly its statistic, and at fifty specifications nearly seven reported winners in ten still carry no truncation beyond 1.96 (69.8% at a correlation of 0.9). But it grows, and a data set that admits fifty reasonable specifications — a handful of outlier rules, several covariate sets, two or three transformations of the outcome — is not unusual. An analyst choosing among fifty on residual diagnostics, in a correlated family with nothing to find, reports a significant winner in nearly one family in five, with an interval that claims a positive effect wrongly more than twice as often as it should.

What a reader can do with three numbers

The exact interval needs the unreported analyses. A reader has the reported statistic, and a report could state three more numbers without revealing anything about the others’ values: how many analyses were run, how correlated their statistics are, and how correlated the diagnostic used to choose among them is with the statistic. Those three are enough to compute the chance π(z)\pi(z) that an analysis whose statistic is zz also has the best diagnostic when the other analyses have no effect.

The chance an analysis with statistic z also has the best of twenty diagnostics, for three correlations between diagnostic and statistic. Twenty analyses correlated 0.6, the others null. At z = 2 the chance is 0.074 for γ = 0.3, 0.111 for γ = 0.6, 0.194 for γ = 0.9; at z = 4, 0.114, 0.242, 0.576. With no correlation it would be one in twenty at every z.
Fig. 4 The chance that an analysis with statistic z also has the best of twenty diagnostics, with the other nineteen null and correlated at 0.6, for three correlations between diagnostic and statistic. Without any correlation it would be one in twenty at every z.

With that chance in hand, the reported statistic’s density given the selection is φ(z−μ) π(z)\varphi(z - \mu)\,\pi(z) above 1.96, with the unreported analyses integrated out rather than held at the values they actually took. Inverting its upper tail in μ\mu gives an interval conditioned on the selection a reader can reconstruct. It covers 95.32%, 95.45% and 95.56% at γ=0.3\gamma = 0.3, 0.6 and 0.9, and at 0.9 it misses 2.52% below and 1.91% above: the false positive rate is back at its nominal 2.5% without the unreported analyses ever being seen.

The two conditional intervals answer slightly different questions, and the difference is why the reader’s version is possible at all. The exact interval conditions on everything the selection depended on, including the unreported analyses’ actual values; it is valid given those values, which nobody outside the analysis has. The reader’s interval conditions on the selection having happened and on the reported statistic, and averages over the values the unreported analyses might have taken. Both are valid statements about the reported mean given what each conditions on, and both have the nominal coverage in repeated use; the second is wider, because it cannot use the information in the unseen analyses, but it needs nothing the report does not contain. It is the same move naming a handful in advance made from the other direction — there the analyst declared which analyses count before looking, here the report declares how the one chosen was chosen after the fact — and in both cases what makes the inference honest is a small number of stated facts about the selection, not the data that were hidden.

It is not free. Built on the assumption that the other analyses have no effect, it is conservative when one of them does. With an effect of two standard errors in one analysis and γ=0.9\gamma = 0.9, a winner is reported in 48.00% of families and is the right analysis 90.3% of the time; the threshold-only interval excludes zero for 21.10% of reported winners, the exact interval for 18.60%, and the reader’s reconstruction for 15.10%. Some of the threshold-only interval’s apparent power was the lower-side failure, and the reconstruction gives up a little real power on top of removing it.

What a report of a chosen specification owes

The criterion used to choose among specifications, and its correlation with the reported statistic. A choice on diagnostics is a choice on the statistic to the extent the two are correlated. At a correlation of 0.9 it multiplies the chance of a significant winner fivefold, from 2.52% to 13.07%, in a family with nothing to find.

The number of specifications and their correlation, which the earlier essays asked for already. With the three numbers a reader can rebuild an interval whose lower limit crosses zero at the nominal rate, and without them a reader cannot tell a threshold-only interval that is right from one that claims a positive effect twice as often as it should.

Or a rule for the choice written before the data. A diagnostic criterion fixed in a protocol — use the specification whose residuals pass the check, and if several do, the first on a list written in advance — is still a selection correlated with the statistic, but it is a selection a reader can simulate exactly, which is the step from the three-number reconstruction to the exact interval. What cannot be repaired from the report is a choice made by looking at diagnostics without a rule: then neither the number of specifications considered nor the criterion’s correlation with the statistic is known, even to the analyst, and the interval that conditions only on clearing 1.96 is the only one left.

And no reliance on coverage as the check. The threshold-only interval covers 93.27% at the worst correlation measured and would pass any coverage check; the failure is entirely in the split between the sides. What a p-value does not say made the same point about a single number standing in for a distribution. A coverage figure is one number standing in for two, and here the two moved in opposite directions.

The rates are over 40,000 families of twenty with no effect and 10,000 with one; with no effect the reported winners number 1,007 at γ=0\gamma = 0 and 5,229 at 0.9, so a miss rate near 2.5% carries a standard error of about half a point at the smallest correlation and a fifth of a point at the largest. At γ=0\gamma = 0 the threshold-only interval is the exact one, and its misses of 1.69% below and 3.28% above are within that error of 2.5% each; a direct check of the truncated-normal inversion on five thousand draws gives 2.32% and 2.50%. The reader’s selection chance is estimated from 4,000 conditional draws of the other nineteen analyses at each of 81 values of zz. Not measured: a diagnostic correlated with the other analyses’ statistics as well as its own; a choice that combines the diagnostic with significance, such as the cleanest among the specifications that clear 1.96; and negative correlations, where a diagnostic that dislikes large statistics lets a winner through in under 1.4% of null families and the question mostly disappears.

Still open: the cleanest of the significant

The selection measured here chooses on the diagnostic first and applies the threshold after. Many analysts do it the other way round: they look at the specifications that reached significance and report the cleanest of those. That selection is a union over which analyses cleared the threshold, not a single set of linear bounds, and its truncation is the maximum over a random subset of the rivals, which the polyhedral conditioning used here handles only piece by piece.

Whether the order matters — whether “significant, then clean” inflates the false-claim rate more than “clean, then significant” at the same correlation — and whether the three-number reconstruction still repairs it, are measurable on this family by enumerating which rivals cleared the threshold. The practical stakes are that the second order is the one analysts describe as prudent, since it never reports an unclean specification, and it may be the one that most often reports a significant specification because it was significant.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Confidence intervalCoverageThe garden of forking pathsModel diagnosticsSelective inferenceSensitivity analysisSpecification searchTruncated normal