Tests, and the second number

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

Worth reading first: A p-value that is not flat is not a p-value.

Generate data with no effect in it whatsoever. Analyse it twenty ways — twenty outcomes, or twenty subgroups, or twenty ways of excluding an observation — each analysis correct, each p-value honest. Report the one that worked.

Something reaches p < 0.05 57% of the time.

The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.
Fig. 1 The chance of finding something significant against the number of analyses available, on data containing no effect. The dashed line is what independent analyses would give.

Nobody cheated

This is worth insisting on, because the phenomenon is usually discussed under headings that imply misconduct.

Every analysis in the simulation is correctly specified and correctly computed. Every p-value is exactly right for the analysis that produced it — uniform under the null, as it must be. No data was excluded, no test was repeated until it worked, and nothing was hidden.

The only thing that happened is that several analyses were available, and the one that reached significance is the one that got written up. That is not fraud; it is the default behaviour of anyone exploring a dataset, and it inflates the false-positive rate by an order of magnitude.

20,000 p-values from a true null, n = 12Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0090 (p = 0.81). That flatness is the check that catches an error a single rejection rate would miss.05001e+30.2000.4000.6000.800p-valuecountflatKS distance 0.0090a p-value that is not uniform is not a p-value
Fig. 2 A single analysis of the same null data. Flat, as it must be — the inflation comes entirely from taking the minimum of several.

The arithmetic

If the analyses were independent, the chance of at least one false positive with k of them is 10.95k1 - 0.95^k. At k=20k = 20 that is 64%.

The measured rate is 57% — a little lower, because the analyses here are correlated, as real ones are. Twenty outcome variables measured on the same subjects share the subjects’ noise, so their p-values move together and the effective number of independent chances is smaller than twenty.

That correlation is why the naive Bonferroni correction is conservative: it assumes independence, so it over-corrects when the analyses are related. The measurement above puts a number on the gap.

The effect behind a p-value of 0.04, at three sample sizes. A p of 0.04 at ten observations needs an effect of 0.76 standard deviations; at two thousand it needs 0.046. The p-value alone does not say which of these happened, which is why it should never be reported alone.
Fig. 3 What a surviving p-value does and does not say about the size of the effect behind it.

What the dashed line is worth as a comparison

The independent case is the right benchmark and the gap between it and the measured curve is a measurement in its own right.

Twenty independent analyses at five per cent give 10.95201 - 0.95^{20}, which is 64.2%. The measured figure is 57%.

Solving 10.95m=0.571 - 0.95^m = 0.57 gives m=16.5m = 16.5.

So twenty forking paths on one dataset are worth about sixteen and a half independent chances. The overlap between analyses — they share rows, share a response, share most of their design — removes about three and a half of the twenty, and no more.

That is the number worth carrying, because the reflex is to suppose the overlap does most of the work. It does not. Analyses of one dataset are so nearly independent, as far as their p-values are concerned, that treating twenty of them as twenty separate chances overstates the problem by seven points and understates nothing.

Why it is invisible

The reason this is so much harder to deal with than an ordinary multiple-comparisons problem: the number of analyses is not recorded anywhere.

A paper reporting one analysis looks identical whether it was pre-specified as the only analysis, or chosen from five, or chosen from fifty. The reader cannot tell, the reviewer cannot tell, and often the author genuinely cannot reconstruct it — the decisions were made one at a time, each for a defensible reason, and none of them felt like a choice at the time.

The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 50 correlated outcomes available, something reaches p below 0.05 84% of the time.
Fig. 4 The same sweep run out to fifty analyses rather than stopping at twenty. The curve does not turn over; it keeps climbing, and the gap to the independent benchmark stays roughly where it was.

Extending the sweep is what makes the invisibility quantitative rather than rhetorical. Five analyses put the chance of something significant at roughly one in five — bad, but a reader who suspected five and discounted accordingly would not be badly wrong. Fifty puts it high enough that a significant finding carries almost no information at all, and the paper reporting it is indistinguishable, on the page, from the pre-registered one.

The two ends of that curve are the same document. This is the reason the honest responses to the problem are all procedural — pre-registration, a reported analysis count, a held-out half — rather than statistical: there is no correction that can be applied to a number when the quantity it would need is the one the write-up does not contain. A correction needs to know k. The reader’s whole difficulty is that k is exactly what was never written down.

Gelman and Loken called this the garden of forking paths, and the important part of their framing is that it applies even when only one analysis was actually run. If the analysis would have been different had the data been different, then the space of analyses is larger than one, and the p-value from the single path taken does not have its stated meaning.

Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.
Fig. 5 Power, which decides how much of a real effect the surviving analyses would have found anyway.

What actually helps

Pre-registration. Stating the analysis before seeing the data collapses the garden to a single path. It is the only measure that addresses the mechanism directly rather than patching the consequences.

Reporting every analysis performed, so the reader can apply their own discount. Cheap, and it turns an invisible problem into a visible one.

Multiplicity corrections, which help when the number of comparisons is known and finite — a genome scan, a panel of endpoints. They do nothing for the forking-paths case, because the count they need is exactly the thing that is unknown.

Split-sample or hold-out analysis. Explore freely on one half, confirm on the other. The confirmatory half has one pre-specified analysis by construction.

12,000 studies of a real effect of 0.3, n = 16. Power is 21%. The studies that reached significance report a mean effect of 0.621 — 2.07 times the truth. Every one of them is honest; the selection did the inflating.
Fig. 6 The second mechanism, operating on which studies survive rather than which analysis is reported.

The relation to the other inflation

This mechanism and the winner’s curse are different and they compound.

The winner’s curse operates on which studies get published: significance filters upward, so surviving effects are inflated. The forking paths operate on which analysis within a study gets reported: the minimum of several p-values is not a p-value.

A literature subject to both — underpowered studies, flexible analysis, publication conditional on significance — produces effects that are inflated twice over and a false-positive rate far above the nominal one. That is a sufficient explanation for a replication crisis without anyone behaving badly, which is both the reassuring and the alarming reading of it.

What the figure is not

It is not a claim about any particular field or paper. It is a simulation of a mechanism, with the parameters stated, showing what that mechanism produces on data guaranteed to contain nothing.

Whether a given literature is affected is an empirical question about that literature. The value of the simulation is that it establishes the size of the effect the mechanism can produce — 57% against a nominal 5% — so that the question can be asked with a number attached.

Where the choices actually come from

The phrase “twenty analyses” sounds like deliberate fishing. In practice the multiplicity is assembled from decisions that each felt forced at the time.

Which outcome. Most studies measure several things; one is designated primary, often after the data is seen.

Which subgroup. The effect is checked in men and women, in older and younger, in the compliant and the whole sample. Each split doubles the space.

Which exclusions. An implausible value, a participant who misunderstood, a site with a protocol deviation. Every defensible exclusion rule is another path.

Which transformation. Raw, log, ranked, winsorised. Each is a defensible choice for skewed data and each produces a different p-value.

Which covariates. Adjusted for age, or age and sex, or the full set. All standard.

Five decisions with two or three options each is between thirty and two hundred and forty paths, and no individual decision looks like a choice to the person making it.

Why “only one analysis was run” is not a defence

The subtlest part of Gelman and Loken’s argument, and the part most often missed.

The inflation does not require that multiple analyses were run. It requires that the analysis would have been different if the data had been different — that the choices were data-dependent, even implicitly.

If a researcher would have reported the subgroup effect had the main effect been null, then the space contains both paths whether or not both were computed. The p-value from the path taken is calibrated for a procedure that was never followed.

That is why pre-registration addresses this and honesty does not. Nobody is lying; the number simply does not mean what it says when the procedure that produced it was contingent on the data.

The familywise error rate with no correction, α = 0.05. Two routes: the curve is 1 − (1 − α)^m and the points are counted over 6,000 simulated families of true nulls. With twenty tests the chance of at least one false positive is 64.1%.
Fig. 7 The independent version of this page’s count, in closed form and counted, against the number of analyses.

The corrections, and what they can and cannot do

Bonferroni divides the threshold by the number of comparisons. It controls the chance of any false positive, it is conservative when the tests are correlated — as the measurement here shows they usually are — and it requires knowing the number.

Holm is uniformly better than Bonferroni at no cost and should replace it by default.

Benjamini–Hochberg controls the false discovery rate — the expected proportion of rejections that are false — rather than the chance of any false positive. That is the right target for screening thousands of hypotheses, where some false positives are acceptable if the proportion is bounded.

All three need the count. None of them addresses the forking-paths case, where the count is unknown and unknowable, and presenting them as a solution to it is the commonest misapplication.

What the simulation does and does not establish

Being precise about the claim, since the number is large enough to be quoted carelessly.

It establishes that a specific mechanism, with stated parameters, produces a 57% false-positive rate on data containing nothing. The mechanism is: twenty correlated outcomes, one dataset, report the smallest p-value.

It does not establish that any real literature has that rate. Real research has fewer paths in some places and more in others, and the correlation structure differs.

What the simulation provides is a scale. Anyone claiming the effect is small has to explain why their situation differs from this one, and anyone claiming it explains everything has to justify a paths count. That is the value of putting a number on a mechanism: it converts an argument about whether something matters into an argument about parameters.

What p-curve does with this

One constructive use of the mechanism, since the essay is otherwise diagnostic.

Under the null, p-values are uniform. Under a real effect they pile up near zero, and the steeper the pile the larger the effect.

So the distribution of significant p-values across a literature is informative. A set of studies of a real effect produces many p-values well below 0.05 and few just under it. A set produced by selection or flexible analysis produces the opposite — a lump immediately below the threshold, because that is where the selection stops.

That gives a diagnostic applicable to a published literature without access to the underlying data, and it is one of the few tools that can distinguish “these findings are real but small” from “these findings are artefacts of the filter”.

It has limits — it needs enough studies, and heterogeneous effects blur it — and it is a genuine use of the mechanism this essay describes rather than a complaint about it.

The multiverse alternative

If the number of paths cannot be counted, one response is to walk all of them and report the distribution.

A multiverse analysis specifies every defensible choice at every decision point, runs all the resulting analyses, and reports the whole set of results rather than one. If the conclusion holds across the multiverse it is robust to the choices; if it appears in a small corner of it, that is visible.

The cost is that it produces a distribution rather than a number, which resists summarising and is harder to publish. The benefit is that it converts the invisible multiplicity this essay is about into something a reader can see.

It also gives a direct estimate of the quantity the simulation here assumes: how many paths there actually were, in a real analysis, rather than a stipulated twenty.

What honest exploration looks like

The essay risks implying that exploratory analysis is illegitimate, and it is not — most discoveries begin there.

The distinction that matters is between generating a hypothesis and testing it. Exploration generates; it is supposed to look at everything, and demanding pre-registration of it would prevent the thing it is for.

The failure is reporting an exploratory result with the statistical furniture of a confirmatory one. A p-value attached to a hypothesis that was generated by the same data is not a p-value in the sense its threshold assumes.

So the practice is not “do not explore”. It is: explore freely, label it as exploration, and confirm on data that had no part in generating the hypothesis. Split-sample analysis does that within one dataset; a replication does it across two.

That is a workable discipline rather than a counsel of perfection, and it costs the researcher a stated caveat rather than a method.

Why the measured rate is below the independent one

The essay quotes two numbers — 57% measured against 64% if the analyses were independent — and the gap between them is more informative than either alone.

The independent calculation is 1 − 0.95²⁰, which assumes twenty analyses that share nothing. The simulation deliberately does not: the twenty outcome variables are measured on the same subjects, so they share the subjects’ noise, and their p-values move together.

That correlation reduces the false-positive rate, and the direction is worth understanding because intuition often gets it backwards. Correlated tests give fewer independent chances to get lucky. In the limit where all twenty analyses are perfectly correlated, they are one analysis, and the rate falls to 5%. In the limit where they share nothing, the rate is 64%. Real analyses sit between, and this one sits at 57%.

Two consequences follow.

The independent calculation is an upper bound, not an estimate. Anyone computing 10.95k1 - 0.95^k for a real set of analyses is overstating the problem, sometimes considerably. The overstatement is largest when the analyses are most similar — twenty slight variations of one model are nearly one analysis, and the naive bound treats them as twenty.

The bound is still the right thing to quote when the correlation is unknown, because the error is in the safe direction and because the correlation is rarely estimable from what gets reported. A paper that describes one analysis gives a reader no way to know how many were available or how alike they were.

The gap also explains why corrections behave the way they do. Bonferroni assumes the worst case and is therefore conservative exactly in proportion to how correlated the analyses actually are, which is why it feels punishing on a set of near-identical models and reasonable on a set of genuinely distinct ones.

What the number is and is not evidence for

The 57% is a specific measurement of a specific mechanism, and the essay should be careful about how far it travels.

It is evidence that a plausible, honest analysis process with no misconduct in it can produce a false-positive rate more than ten times the nominal one. That is the argument, and it is established by construction: the mechanism is written down, its parameters are stated, and the rate is counted.

It is not an estimate of the false-positive rate in any real literature. Real investigators do not run exactly twenty analyses, their analyses are not correlated exactly the way these are, and their choices are made with a great deal more knowledge of the data than a simulation can represent. The number here is what this mechanism produces, and the mapping to any actual field is an argument nobody on either side can currently make with numbers.

The distinction is worth insisting on because both misreadings are common. Treating the 57% as an estimate of how much published work is wrong overstates what a simulation establishes. Dismissing it because it is only a simulation misses that it establishes the thing it was built to establish — that the inflation requires no dishonesty, only choice.

What makes the argument stick is the second half rather than the first. Anyone who wants to maintain that a flexible analysis process keeps its nominal error rate now has a concrete mechanism to explain away, with its parameters written down and its arithmetic reproducible.

The smallest of twenty is not a p-value

A framing that makes the whole effect obvious once adopted, and which the arithmetic above supports rather than replaces.

A p-value is defined as the probability, under the null, of a result at least as extreme as the one observed — where “the one observed” means the analysis that was specified in advance. The quantity actually reported after twenty analyses is the minimum of twenty p-values, and the minimum of twenty draws from a uniform distribution is not uniform. It has a mean near 1/21 and its whole distribution is crushed toward zero.

So the reported number is a different statistic with a different distribution, and calling it a p-value is a naming error rather than a subtle bias. Its correct null distribution is computable — that is exactly what the simulation counts — and against that distribution a value of 0.03 is entirely unremarkable.

This is why the defence that each individual analysis was correct does not help. Every one of them is correct. The reported quantity is not one of them; it is a function of all twenty, and its distribution was never the one the threshold was chosen for.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BonferroniCorrelationFalse positiveThe garden of forking pathsMultiple comparisonsp-valueResearcher degrees of freedom