Field

Tests, and the second number

A p-value alone cannot be read: the same 0.04 means different things at different sample sizes, and nothing at all without knowing how many analyses were available. Under a real effect it is a draw from a distribution several orders of magnitude wide, a replication of a result at 0.05 succeeds exactly half the time under any model centred on that result, and the standard ways of combining several p-values disagree about which evidence counts. Every figure here carries the number that makes it interpretable.
20,000 p-values from a true null, n = 12. Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0090 (p = 0.81). That flatness is the check that catches an error a single rejection rate would miss.

A p-value that is not flat is not a p-value

Under a true null, p-values are uniform. That is stronger than saying the test rejects 5% of the time, it constrains the whole distribution rather than one point of it, and it catches implementation errors that a rejection rate sails past.

Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.

What a p-value does not say

The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.

12,000 studies of a real effect of 0.3, n = 16. Power is 21%. The studies that reached significance report a mean effect of 0.621 — 2.07 times the truth. Every one of them is honest; the selection did the inflating.

The winner's curse

Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.

The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

The p-value of a study with 80% power, twenty thousand times. Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.5603 standard deviations from the null, a noncentrality of 2.802. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, the median is 0.0051, and 80.0% fall below 0.05, which is what the power means.

The p-value a replication gets

Under a true null a p-value is flat. Under a real effect its distribution is closed form and wide — a study with 80% power returns anything from 4.4×10⁻⁵ to 0.13 in eight runs of ten — and the chance that an exact replication of a p = 0.05 result is significant again is exactly one half, under both of the models people use without naming them.

Three ways to reject with two studies, drawn where the two z statistics live. Two one-sided studies, each summarised by its z statistic. Fisher's combination rejects outside a curve that runs parallel to both axes, so one study past z = 2.378 decides it alone; Stouffer's rejects above the straight line z₁ + z₂ = 2.326; Tippett's rejects when either z passes 1.955. Each region holds exactly 5% of the standard bivariate normal — Fisher's in closed form, e^(−c/2)(1 + c/2) at c = 9.488 — and of 100,000 counted null pairs they catch 4.95%, 4.88% and 5.04%. Two alternatives carry the same Stouffer evidence: one study at 2.326 and the other at nothing, where the powers are 62.7%, 50.0% and 65.4%; and both at 1.163, where they are 47.7%, 50.0% and 38.3%.

Two ways to combine p-values

Fisher's and Stouffer's combinations are both exactly right when every null is true, for the single reason that each p-value is flat. Under a real effect they disagree about which evidence counts: with Stouffer held at 50% power across ten studies, Fisher is the more powerful while the signal sits in six or fewer of them and the less powerful from seven.

How often each combination, and each union of them, rejects ten studies of nothing. Each combination alone rejects exactly 5% of null sets. Counted on 1,000,000 sets of ten null studies: Fisher or Stouffer 6.63%, Fisher or Tippett 8.05%, Stouffer or Tippett 8.96%, any of the three 9.66% — enclosed on a two-dimensional lattice between 9.18% and 10.09% — and all three together 1.05%. The three sizes add to 15%.

The smallest of three combinations

Reporting whichever of Fisher's, Stouffer's and Tippett's combinations is smallest is a test of its own, and on ten studies of nothing it rejects 9.66% of the time — not 5%, and nowhere near the 15% the three sizes add to, because the statistics are correlated at up to 0.903. Read at 2.448% each it is exact, and then it trails the best single combination by at most 7.45 points and leads the worst by at least 10.30.

All essays