The analyses that were available and not run

The cleanest of the significant

An analyst with twenty correlated analyses can choose in two orders: take the one whose diagnostics look cleanest and report it if it clears 1.96, or take the analyses that cleared 1.96 and report the cleanest of them. The second is the one described as prudent, because it never reports an analysis whose diagnostics look poor — and with no effect anywhere it reports a significant analysis in about 18% of families whatever the diagnostic is, against 2.55% to 13.15% for the first. The diagnostic decides which significant analysis is shown and never whether one is. Each report's own interval can still be made exact, by conditioning on a selection set that is a union of pieces rather than a half-line; the family's claim count cannot.

Worth reading first: Twenty analyses of nothing · What the correction corrects.

A winner chosen by its diagnostic measured an analyst who runs twenty correlated analyses of one dataset, picks the one whose diagnostics look cleanest, and reports it if its statistic clears 1.96. A diagnostic is rarely independent of the statistic it checks, and with a correlation of 0.9 between them that procedure let a winner through in 13.07% of families with no effect anywhere, against 2.52% for an unrelated diagnostic.

It ended on the other order, the one analysts usually describe when asked. Look first at which specifications reached significance, then report the cleanest of those. It sounds stricter: it never reports a specification whose residuals are a mess. It is a different selection, though, and its selection event is a union over which rivals cleared the threshold rather than a single bound. Whether the order matters, in which direction, and whether the interval conditioned on the selection can still be built, are all measurable on the same families.

How often two orders of selection report a winner from twenty analyses, no effect anywhereTwenty analyses correlated 0.6, each with a diagnostic correlated γ with its own statistic, over 20,000 families a point. Clean then significant — the best diagnostic, reported if it clears 1.96 — reports a winner in 2.55%, 4.25%, 6.91%, 13.15% of families at γ = 0, 0.3, 0.6, 0.9. Significant then clean — the best diagnostic among those that cleared 1.96 — reports one in 18.04%, 18.04%, 18.04%, 18.05%: in every family where any analysis cleared 1.96, whatever γ is.0%5%10%15%20%00.3000.6000.900the diagnostic's correlation with its analysis's statistic, γshare of families reporting a winnerclean, then significant: a winner reportedsignificant, then clean: a winner reported20,000 families of twenty a pointevery analysis null
Fig. 1 How often each order of selection reports a winner from twenty analyses correlated 0.6, against the correlation γ between each analysis’s diagnostic and its statistic, with no effect anywhere — 20,000 families a point. The slider puts an effect of two into one analysis.

Two orders, one family

The family is the earlier essay’s. Twenty statistics z1,…,z20z_1, \dots, z_{20} are standard normal under the null, correlated 0.6 with each other, the way twenty analyses of one dataset that share most of their data are. Each has a diagnostic di=γzi+1−γ2 eid_i = \gamma z_i + \sqrt{1-\gamma^2}\,e_i, correlated γ\gamma with its own statistic and with nothing else: a residual check that looks better when the fit happens to be stronger, a balance table that looks better when the subgroup is larger.

Clean, then significant takes the analysis with the largest diagnostic and reports it if its statistic exceeds 1.96. Significant, then clean takes the analyses whose statistics exceed 1.96 and reports the one among them with the largest diagnostic. Both orders are applied to the same families, so a difference between them is the order and nothing else.

The hero figure shows the first consequence, and it is the decisive one. With no effect anywhere, clean-then-significant reports a winner in 2.55% of families when the diagnostic is unrelated to the statistic, 4.25% at a correlation of 0.3, 6.91% at 0.6 and 13.15% at 0.9 — the earlier essay’s rising curve. Significant-then-clean reports a winner in 18.04% to 18.05% of families across those correlations, and the figure is flat because it has to be: the order reports a winner exactly when any of the twenty analyses cleared 1.96, and whether any did has nothing to do with the diagnostics.

That is the whole of the difference in one sentence. In the prudent-sounding order the diagnostic decides which significant analysis is shown. It never decides whether one is. The family’s false-claim rate is the rate at which twenty correlated null analyses produce at least one significant result, which is what twenty analyses of nothing measured, and the diagnostic, however good, cannot lower it.

Why the stricter-sounding order is the looser one

The intuition that the second order is stricter comes from looking at a single report. Under significant-then-clean, every reported analysis has both a significant statistic and a diagnostic that was the best among its significant rivals; under clean-then-significant, the reported analysis was the cleanest overall and also happened to be significant. Report by report, the second sounds more demanding.

The first order, though, makes significance a test the cleanest analysis has to pass. With an unrelated diagnostic it is one pre-chosen analysis tested at 5% in one tail — 2.55%, the nominal one-sided rate — because the diagnostic picks an analysis without looking at its statistic. As the diagnostic’s correlation with the statistic rises, it starts to pick analyses whose statistics are large, and the rate rises with it; at 0.9 the diagnostic is nearly the statistic, and the first order is nearly reporting the largest of twenty. The second order is reporting the largest from the start, because “any of them is significant” is the event “the largest is significant”, and then using the diagnostic to choose among the ones that made it.

So the two orders converge as the diagnostic becomes the statistic — at γ = 0.9 the first reports in 13.15% of families and the second in 18.05% — and they are furthest apart when the diagnostic is honest. At an unrelated diagnostic, the second order reports a significant analysis seven times as often as the first.

The set a reported statistic lies in

A report from either order can still carry an honest interval: one conditioned on the selection that produced it, so that its 95% is a statement about the analyses that would have been reported, not about all analyses. For the first order the earlier essay built it: holding everything independent of the winner’s statistic fixed, the selection is a single lower bound on that statistic, and the winner is a normal truncated there.

The values the reported statistic could have taken and still been reported, under each order of selection, one family. One family of twenty analyses correlated 0.6, diagnostics correlated 0.6 with their statistics. Under significant then clean the reported analysis's statistic is 2.31, and given everything else about the family it would have been the one reported for any value in [1.96, 2.49] ∪ [3.65, ∞): above the gap another analysis becomes significant and has the cleaner diagnostic. Its exact interval, conditioned on that set, is [−7.83, 4.16]; conditioned only on clearing 1.96, [−8.30, 4.01]. Under clean then significant the analysis with the best diagnostic has a statistic of 1.85 and nothing is reported.
Fig. 2 One family, with the diagnostic correlated 0.6 with each statistic: the values the reported analysis’s statistic could have taken, everything else about the family held fixed, and still been the one reported under significant-then-clean, with the observed value marked. The clean-then-significant order reports nothing in this family.

For the second order the selection set has a different shape. Hold fixed each rival’s statistic net of the winner, wi=zi−0.6 zjw_i = z_i - 0.6\,z_j, and every diagnostic noise eie_i. Rival ii is significant once the winner’s statistic passes (1.96−wi)/0.6(1.96 - w_i)/0.6, because the rivals move with the winner, and it beats the winner on the diagnostic below a second point, because a larger winner has a better diagnostic. Between those two points the rival is both significant and cleaner, and the winner would not have been the one reported. So the set the winner’s statistic lies in is [1.96,∞)[1.96, \infty) with up to nineteen intervals cut out of it: a union of pieces rather than a half-line.

The figure shows one family where that matters. The reported statistic is 2.31; it would have been reported for any value between 1.96 and 2.49, not between 2.49 and 3.65 — where another analysis becomes significant and is cleaner — and again from 3.65 up, where the winner is clean enough to beat it. In that same family the first order reports nothing at all, because the cleanest analysis overall has a statistic of 1.85.

The cut-out pieces are not rare. At a diagnostic correlation of 0.3 the set has more than one piece in 71.44% of reports, at 0.6 in 51.72% and at 0.9 in 19.81%. With an unrelated diagnostic the set is a single piece again, [1.96, u], because a rival with the better diagnostic beats the winner at every value once it is significant, so its cut runs to infinity and simply caps the set.

What each interval claims

Given a set of pieces, the winner’s statistic is a normal truncated to that set, and the interval for its mean is found as the earlier ones were, by inverting the truncated distribution’s tail probability — now summed over the pieces. It is the same polyhedral argument an interval for the analyses admitted to used for a threshold that admits several analyses at once, with the one change that the constraints now come in pairs, and a pair can remove a piece from the middle of the line rather than raise its floor.

How often a winner reported as the cleanest of the significant has an interval excluding zero, with no effect anywhere. Significant then clean, twenty analyses correlated 0.6, 20,000 families a point. The reported winner's interval excludes zero — wrongly, since every analysis is null: conditioned only on clearing 1.96, in 0.64%, 1.05%, 1.69%, 3.35% of reports at γ = 0, 0.3, 0.6, 0.9; conditioned on the whole selection set, 2.05%, 2.30%, 2.13%, 2.19%; rebuilt from the number of analyses, their correlation and the diagnostic's correlation, 2.19%, 2.36%, 2.52%, 2.63%. A 95% interval should exclude it 2.5% of the time.
Fig. 3 Under significant-then-clean with no effect anywhere, the share of reported winners whose interval excludes zero: conditioned only on clearing 1.96, conditioned on the whole selection set, and rebuilt from the number of analyses, their correlation and the diagnostic’s correlation. A 95% interval should exclude it 2.5% of the time.

Conditioned on the whole selection set, the interval covers 95.20%, 95.01%, 95.32% and 95.07% of reported winners at the four correlations, and excludes zero in 2.05%, 2.30%, 2.13% and 2.19% of them — exact, to within the counting error of about three and a half thousand reports a point. The union is harder to write down than a single bound and no harder to invert.

The interval a report would actually carry, conditioned only on clearing 1.96, is the surprise. It excludes zero in only 0.64% of reports at an unrelated diagnostic, 1.05% at 0.3, 1.69% at 0.6 and 3.35% at 0.9 — below the nominal 2.5% at three of the four. Under the second order the reported statistic is not usually the largest of the twenty; it is whichever significant analysis had the cleanest diagnostic, which is often one that only just cleared the threshold, and an interval truncated at 1.96 for a statistic just above it reaches well below zero. Report by report, the second order’s intervals look modest. The family has still reported a significant analysis in about 18% of cases where there was nothing to find.

The reader’s reconstruction carries over. Built from the three numbers a report can state — twenty analyses, correlated 0.6, a diagnostic correlated γ — as the chance that an analysis with a given statistic is the one reported, with the unseen analyses integrated out, it excludes zero in 2.19%, 2.36%, 2.52% and 2.63% of reports. It covers more than 95% at the three weaker correlations, 97.81%, 97.64% and 97.48%, because integrating the rivals out averages over sets that are sometimes much wider than the one that occurred, and 94.71% at 0.9.

None of this is about the reported estimate, which is too large under either order for the familiar reason: it was reported because it was large, and the correction that makes the estimate worse found that shrinking it towards zero by a rule fitted to the selection can make it worse on average when the effect is real. The intervals here are the honest part of the report. They answer what the reported analysis’s effect could be, given that it was reported; they do not answer how often a report like it is a report of nothing, and when every null is true the second is the question that matters.

When one analysis is real

The reporting rate under the null is a cost. Under an effect it is a benefit, and the slider in the hero figure shows what the second order buys with it.

How often two orders of selection report a winner from twenty analyses, one with an effect of two. Twenty analyses correlated 0.6, each with a diagnostic correlated γ with its own statistic, over 20,000 families a point. Clean then significant — the best diagnostic, reported if it clears 1.96 — reports a winner in 5.13%, 11.63%, 25.42%, 47.88% of families at γ = 0, 0.3, 0.6, 0.9. Significant then clean — the best diagnostic among those that cleared 1.96 — reports one in 52.93%, 52.91%, 52.92%, 52.92%: in every family where any analysis cleared 1.96, whatever γ is. The analysis with the real effect is the one reported in 53.6%, 68.4%, 81.2%, 90.1% of clean-first reports and 77.4%, 79.4%, 82.7%, 89.1% of significant-first reports.
Fig. 4 How often each order reports a winner, and how often the winner reported is the one analysis with a real effect of two, against the diagnostic’s correlation with its statistic. 20,000 families a point.

Put an effect of two into one of the twenty analyses. The second order reports a winner in 52.9% of families at every correlation, because at least one analysis — usually the real one — clears 1.96 that often, and the winner it reports is the real one in 77.39%, 79.36%, 82.70% and 89.13% of its reports. So it reports the real effect in 41.0% of families with an unrelated diagnostic and 47.2% with a diagnostic correlated 0.9. The first order reports a winner in 5.13% of families at an unrelated diagnostic and 47.88% at 0.9, the real analysis in 53.56% and 90.11% of those reports — the real effect in 2.7% of families and 43.1%.

Two details of those numbers are worth separating. With an unrelated diagnostic the first order reports the real analysis in only 53.56% of its reports, because a diagnostic that knows nothing picks the real analysis one time in twenty and the rest of its reports come from null analyses that cleared 1.96 by chance; the real effect is found almost only when the noise happens to point the diagnostic at it. The second order’s reports are the real analysis 77.39% of the time even then, because the real analysis is usually among the significant ones and often the only one. A diagnostic that cannot tell real from null is worse than useless as a filter in front of a test, and harmless as a tie-break after one.

So the trade is plain. With a weak diagnostic, the second order finds a real effect fifteen times as often as the first, at the price of reporting a significant analysis seven times as often when there is none. With a strong diagnostic the two orders converge on both counts, because a strong diagnostic is nearly the statistic and both orders are then nearly reporting the largest. Neither order is the prudent one. The first is a test of one analysis chosen by its diagnostic; the second is a search over all of them, and its error rate is the search’s.

What a diagnostic is for

The comparison says something about diagnostics in general, which is easy to lose in the arithmetic. A diagnostic is information about whether an analysis deserves to be believed — whether its model fits, whether its assumptions hold — and that information can enter a decision in two places. Before the test, as a filter, it decides which analysis is tested, and every analysis it rules out is one fewer chance of a false claim; that is what naming a handful in advance does with a document instead of a diagnostic. After the test, as a tie-break, it decides which of the analyses that already passed is shown, and the number of chances has already been spent.

The first order is the filter, and its weakness is that a diagnostic correlated with the statistic is a filter that leaks: at 0.9 it passes the analyses with large statistics, and the filter becomes a search. The second order is the tie-break, and it has no such weakness to have, because it never used the diagnostic to limit anything. What the correction corrects is the count of chances, and only the first order changes that count.

That is why the second order is the one analysts describe as careful. It visibly uses the diagnostic on every report, and the report never contains an analysis with poor diagnostics. Neither fact has anything to do with how often a report is made, which is the number a reader’s 5% is about.

What a report owes, by order

Say which order was used. The two produce the same kind of sentence — “the specification with the best diagnostics was significant” — and their false-claim rates differ by a factor of seven at an honest diagnostic. A reader cannot recover the order from the report.

Under significant-then-clean, correct for the search, not for the diagnostic. The family’s error rate is the rate at which any of the analyses is significant, so the correction it needs is the one an interval for the winner built for the largest of twenty, whatever the diagnostic was. For twenty analyses correlated 0.6 that threshold is 2.85, and with it in place of 1.96 the second order reports a significant analysis in 2.53% of null families, the one-sided share of the familywise 5%. A diagnostic used only to choose among the significant is a tie-break, and costs and buys nothing at the family level.

For the report’s own interval, condition on the whole selection. The union-shaped set gives an interval that covers 95% given the report, at every correlation measured; the threshold-only interval is conservative at weak diagnostics and liberal at 0.9; and the three-number reconstruction is conservative where it errs. None of the three repairs the 18% — an interval honest about the winner given that one was reported says nothing about how often one is reported.

Every rate is counted over 20,000 families of twenty analyses for each setting, with both orders applied to the same families. The selection set is computed exactly for each report by cutting each rival’s interval out of [1.96,∞)[1.96, \infty), and is checked to contain the observed statistic in every report. The truncated-normal inversion over a union of pieces is checked against the single-piece inversion and against counted coverage for a two-piece set. The second order’s flat reporting rate is checked to equal the chance that any analysis clears 1.96, exactly, at every correlation. The claim that significant-then-clean controls the family’s false claims through its diagnostic is refused: its reporting rate does not move with the diagnostic at all.

Still open: a diagnostic that vetoes

Both orders here use the diagnostic to rank. A third common practice uses it to veto: run the planned analysis, and if its diagnostic fails a fixed check, switch to a stated alternative — a transformation when the residuals are skewed, a robust estimator when there are outliers, a different model when a test of fit rejects. The veto is one decision rather than a search, and it is the version most analysis plans would call pre-specified.

Its selection set is two pieces as well — the planned analysis’s statistic given that its diagnostic passed, the alternative’s given that it failed — and when the diagnostic is correlated with the statistic, a planned analysis that “passed” is a planned analysis whose statistic was nudged up. Whether a veto with a stated alternative holds its 5% when the diagnostic is correlated 0.6 with the statistic, or whether the planned-and-passed analyses are the ones that carry an excess, is computable exactly on these families, and has not been computed.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Conditional inferenceConfidence intervalCoverageForking pathsMultiple comparisonsP hackingSelective inferenceSpecification search