Twenty analyses of nothing
Generate data with no effect in it whatsoever. Analyse it twenty ways — twenty outcomes, or twenty subgroups, or twenty ways of excluding an observation — each analysis correct, each p-value honest. Report the one that worked.
Something reaches p < 0.05 58% of the time.
Nobody cheated
This is worth insisting on, because the phenomenon is usually discussed under headings that imply misconduct.
Every analysis in the simulation is correctly specified and correctly computed. Every p-value is exactly right for the analysis that produced it — uniform under the null, as it must be. No data was excluded, no test was repeated until it worked, and nothing was hidden.
The only thing that happened is that several analyses were available, and the one that reached significance is the one that got written up. That is not fraud; it is the default behaviour of anyone exploring a dataset, and it inflates the false-positive rate by an order of magnitude.
The arithmetic
If the analyses were independent, the chance of at least one false positive with k of them is 1 − 0.95^k. At k = 20 that is 64%.
The measured rate is 58% — a little lower, because the analyses here are correlated, as real ones are. Twenty outcome variables measured on the same subjects share the subjects’ noise, so their p-values move together and the effective number of independent chances is smaller than twenty.
That correlation is why the naive Bonferroni correction is conservative: it assumes independence, so it over-corrects when the analyses are related. The measurement above puts a number on the gap.
Why it is invisible
The reason this is so much harder to deal with than an ordinary multiple-comparisons problem: the number of analyses is not recorded anywhere.
A paper reporting one analysis looks identical whether it was pre-specified as the only analysis, or chosen from five, or chosen from fifty. The reader cannot tell, the reviewer cannot tell, and often the author genuinely cannot reconstruct it — the decisions were made one at a time, each for a defensible reason, and none of them felt like a choice at the time.
Gelman and Loken called this the garden of forking paths, and the important part of their framing is that it applies even when only one analysis was actually run. If the analysis would have been different had the data been different, then the space of analyses is larger than one, and the p-value from the single path taken does not have its stated meaning.
What actually helps
Pre-registration. Stating the analysis before seeing the data collapses the garden to a single path. It is the only measure that addresses the mechanism directly rather than patching the consequences.
Reporting every analysis performed, so the reader can apply their own discount. Cheap, and it turns an invisible problem into a visible one.
Multiplicity corrections, which help when the number of comparisons is known and finite — a genome scan, a panel of endpoints. They do nothing for the forking-paths case, because the count they need is exactly the thing that is unknown.
Split-sample or hold-out analysis. Explore freely on one half, confirm on the other. The confirmatory half has one pre-specified analysis by construction.
The relation to the other inflation
This mechanism and the winner’s curse are different and they compound.
The winner’s curse operates on which studies get published: significance filters upward, so surviving effects are inflated. The forking paths operate on which analysis within a study gets reported: the minimum of several p-values is not a p-value.
A literature subject to both — underpowered studies, flexible analysis, publication conditional on significance — produces effects that are inflated twice over and a false-positive rate far above the nominal one. That is a sufficient explanation for a replication crisis without anyone behaving badly, which is both the reassuring and the alarming reading of it.
What the figure is not
It is not a claim about any particular field or paper. It is a simulation of a mechanism, with the parameters stated, showing what that mechanism produces on data guaranteed to contain nothing.
Whether a given literature is affected is an empirical question about that literature. The value of the simulation is that it establishes the size of the effect the mechanism can produce — 58% against a nominal 5% — so that the question can be asked with a number attached.
Where the choices actually come from
The phrase “twenty analyses” sounds like deliberate fishing. In practice the multiplicity is assembled from decisions that each felt forced at the time.
Which outcome. Most studies measure several things; one is designated primary, often after the data is seen.
Which subgroup. The effect is checked in men and women, in older and younger, in the compliant and the whole sample. Each split doubles the space.
Which exclusions. An implausible value, a participant who misunderstood, a site with a protocol deviation. Every defensible exclusion rule is another path.
Which transformation. Raw, log, ranked, winsorised. Each is a defensible choice for skewed data and each produces a different p-value.
Which covariates. Adjusted for age, or age and sex, or the full set. All standard.
Five decisions with two or three options each is between thirty and two hundred and forty paths, and no individual decision looks like a choice to the person making it.
Why “only one analysis was run” is not a defence
The subtlest part of Gelman and Loken’s argument, and the part most often missed.
The inflation does not require that multiple analyses were run. It requires that the analysis would have been different if the data had been different — that the choices were data-dependent, even implicitly.
If a researcher would have reported the subgroup effect had the main effect been null, then the space contains both paths whether or not both were computed. The p-value from the path taken is calibrated for a procedure that was never followed.
That is why pre-registration addresses this and honesty does not. Nobody is lying; the number simply does not mean what it says when the procedure that produced it was contingent on the data.
The corrections, and what they can and cannot do
Bonferroni divides the threshold by the number of comparisons. It controls the chance of any false positive, it is conservative when the tests are correlated — as the measurement here shows they usually are — and it requires knowing the number.
Holm is uniformly better than Bonferroni at no cost and should replace it by default.
Benjamini–Hochberg controls the false discovery rate — the expected proportion of rejections that are false — rather than the chance of any false positive. That is the right target for screening thousands of hypotheses, where some false positives are acceptable if the proportion is bounded.
All three need the count. None of them addresses the forking-paths case, where the count is unknown and unknowable, and presenting them as a solution to it is the commonest misapplication.
What the simulation does and does not establish
Being precise about the claim, since the number is large enough to be quoted carelessly.
It establishes that a specific mechanism, with stated parameters, produces a 58% false-positive rate on data containing nothing. The mechanism is: twenty correlated outcomes, one dataset, report the smallest p-value.
It does not establish that any real literature has that rate. Real research has fewer paths in some places and more in others, and the correlation structure differs.
What the simulation provides is a scale. Anyone claiming the effect is small has to explain why their situation differs from this one, and anyone claiming it explains everything has to justify a paths count. That is the value of putting a number on a mechanism: it converts an argument about whether something matters into an argument about parameters.
What p-curve does with this
One constructive use of the mechanism, since the essay is otherwise diagnostic.
Under the null, p-values are uniform. Under a real effect they pile up near zero, and the steeper the pile the larger the effect.
So the distribution of significant p-values across a literature is informative. A set of studies of a real effect produces many p-values well below 0.05 and few just under it. A set produced by selection or flexible analysis produces the opposite — a lump immediately below the threshold, because that is where the selection stops.
That gives a diagnostic applicable to a published literature without access to the underlying data, and it is one of the few tools that can distinguish “these findings are real but small” from “these findings are artefacts of the filter”.
It has limits — it needs enough studies, and heterogeneous effects blur it — and it is a genuine use of the mechanism this essay describes rather than a complaint about it.
The multiverse alternative
If the number of paths cannot be counted, one response is to walk all of them and report the distribution.
A multiverse analysis specifies every defensible choice at every decision point, runs all the resulting analyses, and reports the whole set of results rather than one. If the conclusion holds across the multiverse it is robust to the choices; if it appears in a small corner of it, that is visible.
The cost is that it produces a distribution rather than a number, which resists summarising and is harder to publish. The benefit is that it converts the invisible multiplicity this essay is about into something a reader can see.
It also gives a direct estimate of the quantity the simulation here assumes: how many paths there actually were, in a real analysis, rather than a stipulated twenty.
What honest exploration looks like
The essay risks implying that exploratory analysis is illegitimate, and it is not — most discoveries begin there.
The distinction that matters is between generating a hypothesis and testing it. Exploration generates; it is supposed to look at everything, and demanding pre-registration of it would prevent the thing it is for.
The failure is reporting an exploratory result with the statistical furniture of a confirmatory one. A p-value attached to a hypothesis that was generated by the same data is not a p-value in the sense its threshold assumes.
So the practice is not “do not explore”. It is: explore freely, label it as exploration, and confirm on data that had no part in generating the hypothesis. Split-sample analysis does that within one dataset; a replication does it across two.
That is a workable discipline rather than a counsel of perfection, and it costs the researcher a stated caveat rather than a method.