Tests, and the second number

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 58% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

Generate data with no effect in it whatsoever. Analyse it twenty ways — twenty outcomes, or twenty subgroups, or twenty ways of excluding an observation — each analysis correct, each p-value honest. Report the one that worked.

Something reaches p < 0.05 58% of the time.

The false-positive rate against the number of analyses, on pure noiseThe data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 57% of the time.00.2000.4000.6000.8005101520analyses available to the researcherchance of finding something significantthe nominal 5%if the analyses were independentcorrelated, as real ones areno effect present anywhereevery individual analysis is correct
Fig. 1 The chance of finding something significant against the number of analyses available, on data containing no effect. The dashed line is what independent analyses would give.

Nobody cheated

This is worth insisting on, because the phenomenon is usually discussed under headings that imply misconduct.

Every analysis in the simulation is correctly specified and correctly computed. Every p-value is exactly right for the analysis that produced it — uniform under the null, as it must be. No data was excluded, no test was repeated until it worked, and nothing was hidden.

The only thing that happened is that several analyses were available, and the one that reached significance is the one that got written up. That is not fraud; it is the default behaviour of anyone exploring a dataset, and it inflates the false-positive rate by an order of magnitude.

20,000 p-values from a true null, n = 12Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0165 (p = 0.13). That flatness is the check that catches an error a single rejection rate would miss.05001e+30.2000.4000.6000.800p-valuecountflatKS distance 0.0165a p-value that is not uniform is not a p-value
Fig. 2 A single analysis of the same null data. Flat, as it must be — the inflation comes entirely from taking the minimum of several.

The arithmetic

If the analyses were independent, the chance of at least one false positive with k of them is 1 − 0.95^k. At k = 20 that is 64%.

The measured rate is 58% — a little lower, because the analyses here are correlated, as real ones are. Twenty outcome variables measured on the same subjects share the subjects’ noise, so their p-values move together and the effective number of independent chances is smaller than twenty.

That correlation is why the naive Bonferroni correction is conservative: it assumes independence, so it over-corrects when the analyses are related. The measurement above puts a number on the gap.

The effect behind a p-value of 0.04, at three sample sizesA p of 0.04 at ten observations needs an effect of 0.76 standard deviations; at two thousand it needs 0.046. The p-value alone does not say which of these happened, which is why it should never be reported alone.ten observations0.758 sdp = 0.04a hundred0.208 sdp = 0.04two thousand0.046 sdp = 0.04the effect that gives the same p-valueall three reach p = 0.04the number does not say how big the effect is
Fig. 3 What a surviving p-value does and does not say about the size of the effect behind it.

Why it is invisible

The reason this is so much harder to deal with than an ordinary multiple-comparisons problem: the number of analyses is not recorded anywhere.

A paper reporting one analysis looks identical whether it was pre-specified as the only analysis, or chosen from five, or chosen from fifty. The reader cannot tell, the reviewer cannot tell, and often the author genuinely cannot reconstruct it — the decisions were made one at a time, each for a defensible reason, and none of them felt like a choice at the time.

Gelman and Loken called this the garden of forking paths, and the important part of their framing is that it applies even when only one analysis was actually run. If the analysis would have been different had the data been different, then the space of analyses is larger than one, and the p-value from the single path taken does not have its stated meaning.

What a one-sample t test at n = 20 can detectAt an effect of 0.5 standard deviations the test finds it 56% of the time. Below that, a non-significant result is the expected outcome of a real effect — which is why "no significant difference" is not evidence of no difference.00.2500.5000.750100.5001true effect, in standard deviationsprobability of rejecting80%, the usual target56% at d = 0.54,000 simulated studies per pointthe second number a p-value needs
Fig. 4 Power, which decides how much of a real effect the surviving analyses would have found anyway.

What actually helps

Pre-registration. Stating the analysis before seeing the data collapses the garden to a single path. It is the only measure that addresses the mechanism directly rather than patching the consequences.

Reporting every analysis performed, so the reader can apply their own discount. Cheap, and it turns an invisible problem into a visible one.

Multiplicity corrections, which help when the number of comparisons is known and finite — a genome scan, a panel of endpoints. They do nothing for the forking-paths case, because the count they need is exactly the thing that is unknown.

Split-sample or hold-out analysis. Explore freely on one half, confirm on the other. The confirmatory half has one pre-specified analysis by construction.

12,000 studies of a real effect of 0.3, n = 16Power is 20%. The studies that reached significance report a mean effect of 0.613 — 2.04 times the truth. Every one of them is honest; the selection did the inflating.02505007501e+3-0.50000.5001effect the study reportsstudiesthe truth, 0.3what gets published, 0.61shaded: reached p below 0.05power 20%, inflation 2.04×
Fig. 5 The second mechanism, operating on which studies survive rather than which analysis is reported.

The relation to the other inflation

This mechanism and the winner’s curse are different and they compound.

The winner’s curse operates on which studies get published: significance filters upward, so surviving effects are inflated. The forking paths operate on which analysis within a study gets reported: the minimum of several p-values is not a p-value.

A literature subject to both — underpowered studies, flexible analysis, publication conditional on significance — produces effects that are inflated twice over and a false-positive rate far above the nominal one. That is a sufficient explanation for a replication crisis without anyone behaving badly, which is both the reassuring and the alarming reading of it.

Coverage of four nominal 95% intervals, n = 30Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.0.4000.6000.80010.2000.4000.6000.800the true proportionactual coverage of a nominal 95% intervalnominal 95%Wald — the textbook oneWilsonAgresti–CoullClopper–Pearsonsummed over all 31 outcomesno interval is 95% until it is counted
Fig. 6 A measurement of a stated procedure across its range, which is the form every claim here takes.

What the figure is not

It is not a claim about any particular field or paper. It is a simulation of a mechanism, with the parameters stated, showing what that mechanism produces on data guaranteed to contain nothing.

Whether a given literature is affected is an empirical question about that literature. The value of the simulation is that it establishes the size of the effect the mechanism can produce — 58% against a nominal 5% — so that the question can be asked with a number attached.

Where the choices actually come from

The phrase “twenty analyses” sounds like deliberate fishing. In practice the multiplicity is assembled from decisions that each felt forced at the time.

Which outcome. Most studies measure several things; one is designated primary, often after the data is seen.

Which subgroup. The effect is checked in men and women, in older and younger, in the compliant and the whole sample. Each split doubles the space.

Which exclusions. An implausible value, a participant who misunderstood, a site with a protocol deviation. Every defensible exclusion rule is another path.

Which transformation. Raw, log, ranked, winsorised. Each is a defensible choice for skewed data and each produces a different p-value.

Which covariates. Adjusted for age, or age and sex, or the full set. All standard.

Five decisions with two or three options each is between thirty and two hundred and forty paths, and no individual decision looks like a choice to the person making it.

Why “only one analysis was run” is not a defence

The subtlest part of Gelman and Loken’s argument, and the part most often missed.

The inflation does not require that multiple analyses were run. It requires that the analysis would have been different if the data had been different — that the choices were data-dependent, even implicitly.

If a researcher would have reported the subgroup effect had the main effect been null, then the space contains both paths whether or not both were computed. The p-value from the path taken is calibrated for a procedure that was never followed.

That is why pre-registration addresses this and honesty does not. Nobody is lying; the number simply does not mean what it says when the procedure that produced it was contingent on the data.

The corrections, and what they can and cannot do

Bonferroni divides the threshold by the number of comparisons. It controls the chance of any false positive, it is conservative when the tests are correlated — as the measurement here shows they usually are — and it requires knowing the number.

Holm is uniformly better than Bonferroni at no cost and should replace it by default.

Benjamini–Hochberg controls the false discovery rate — the expected proportion of rejections that are false — rather than the chance of any false positive. That is the right target for screening thousands of hypotheses, where some false positives are acceptable if the proportion is bounded.

All three need the count. None of them addresses the forking-paths case, where the count is unknown and unknowable, and presenting them as a solution to it is the commonest misapplication.

What the simulation does and does not establish

Being precise about the claim, since the number is large enough to be quoted carelessly.

It establishes that a specific mechanism, with stated parameters, produces a 58% false-positive rate on data containing nothing. The mechanism is: twenty correlated outcomes, one dataset, report the smallest p-value.

It does not establish that any real literature has that rate. Real research has fewer paths in some places and more in others, and the correlation structure differs.

What the simulation provides is a scale. Anyone claiming the effect is small has to explain why their situation differs from this one, and anyone claiming it explains everything has to justify a paths count. That is the value of putting a number on a mechanism: it converts an argument about whether something matters into an argument about parameters.

What p-curve does with this

One constructive use of the mechanism, since the essay is otherwise diagnostic.

Under the null, p-values are uniform. Under a real effect they pile up near zero, and the steeper the pile the larger the effect.

So the distribution of significant p-values across a literature is informative. A set of studies of a real effect produces many p-values well below 0.05 and few just under it. A set produced by selection or flexible analysis produces the opposite — a lump immediately below the threshold, because that is where the selection stops.

That gives a diagnostic applicable to a published literature without access to the underlying data, and it is one of the few tools that can distinguish “these findings are real but small” from “these findings are artefacts of the filter”.

It has limits — it needs enough studies, and heterogeneous effects blur it — and it is a genuine use of the mechanism this essay describes rather than a complaint about it.

The multiverse alternative

If the number of paths cannot be counted, one response is to walk all of them and report the distribution.

A multiverse analysis specifies every defensible choice at every decision point, runs all the resulting analyses, and reports the whole set of results rather than one. If the conclusion holds across the multiverse it is robust to the choices; if it appears in a small corner of it, that is visible.

The cost is that it produces a distribution rather than a number, which resists summarising and is harder to publish. The benefit is that it converts the invisible multiplicity this essay is about into something a reader can see.

It also gives a direct estimate of the quantity the simulation here assumes: how many paths there actually were, in a real analysis, rather than a stipulated twenty.

What honest exploration looks like

The essay risks implying that exploratory analysis is illegitimate, and it is not — most discoveries begin there.

The distinction that matters is between generating a hypothesis and testing it. Exploration generates; it is supposed to look at everything, and demanding pre-registration of it would prevent the thing it is for.

The failure is reporting an exploratory result with the statistical furniture of a confirmatory one. A p-value attached to a hypothesis that was generated by the same data is not a p-value in the sense its threshold assumes.

So the practice is not “do not explore”. It is: explore freely, label it as exploration, and confirm on data that had no part in generating the hypothesis. Split-sample analysis does that within one dataset; a replication does it across two.

That is a workable discipline rather than a counsel of perfection, and it costs the researcher a stated caveat rather than a method.