a-p-value-and-what-it-hides

20,000 p-values from a true null, n = 12

Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0090 (p = 0.81). That flatness is the check that catches an error a single rejection rate would miss.

Tests, and the second numberslider: sample size, 5 positionswide18 views

What else it draws

The same object, drawn to answer the other questions the essays put to it.

Power is 21%. The studies that reached significance report a mean effect of 0.621 — 2.07 times the truth. Every one of them is honest; the selection did the inflating.

The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.

A p of 0.04 at ten observations needs an effect of 0.76 standard deviations; at two thousand it needs 0.046. The p-value alone does not say which of these happened, which is why it should never be reported alone.

Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.5603 standard deviations from the null, a noncentrality of 2.802. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, the median is 0.0051, and 80.0% fall below 0.05, which is what the power means.

Quantiles of a two-sided z-test's p-value in closed form, at powers from 5% to 99%, on a scale of powers of ten. The pale band is the middle ninety-five per cent, the darker one the middle eighty per cent, and the line the median. At 5% power the null is true and the middle eighty per cent is 0.1 to 0.9. At 80% power it is 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, and at 99% it is 5.01: more power moves a p-value down and widens the range it lands in.

Two models for the effect behind a first two-sided p-value. Taking the estimate as the truth gives a replication significant in the same direction 73.1% of the time after p = 0.01 and 90.8% after 0.001. A flat prior, which carries the estimate's own uncertainty into the prediction, gives 66.8% and 82.7%. At p = 0.05 both are exactly one half. The points are counted from 600,000 simulated pairs of studies with effects spread flat, keeping the pairs whose first p-value fell near 0.05, 0.01 and 0.001, and they agree with the curve within their error bars of two standard errors.

The predictive distribution of a replication's two-sided p-value after a first p of 0.05, under three models for the effect, on a scale of powers of ten. Each bar is the middle eighty per cent, the thin line the middle ninety-five, the tick the median. Taking the estimate as the truth: 0.0012 to 0.48. A flat prior: 1.6×10⁻⁴ to 0.65. A sceptical prior with τ = 2: 0.001 to 0.74, median 0.11. The open circles on the flat row are the counted tenth, fiftieth and ninetieth percentiles from 7,223 simulated pairs whose first p fell near 0.05.

Two one-sided studies, each summarised by its z statistic. Fisher's combination rejects outside a curve that runs parallel to both axes, so one study past z = 2.378 decides it alone; Stouffer's rejects above the straight line z₁ + z₂ = 2.326; Tippett's rejects when either z passes 1.955. Each region holds exactly 5% of the standard bivariate normal — Fisher's in closed form, e^(−c/2)(1 + c/2) at c = 9.488 — and of 100,000 counted null pairs they catch 4.95%, 4.88% and 5.04%. Two alternatives carry the same Stouffer evidence: one study at 2.326 and the other at nothing, where the powers are 62.7%, 50.0% and 65.4%; and both at 1.163, where they are 47.7%, 50.0% and 38.3%.

Each trial draws 10 uniform p-values, as true nulls produce, and combines them by Fisher's method against a chi-square on 20 degrees of freedom and by Stouffer's against a standard normal. Left half of each bin Fisher, right half Stouffer. Both are flat: Kolmogorov–Smirnov distances 0.0045 and 0.0044, rejection rates 4.92% and 4.90%. Each is exact because its inputs are uniform, and for no other reason.

The same total evidence — enough to give Stouffer's combination 50% power at every spread — placed in m of 10 one-sided studies: a shift of 5.201 standard errors in one study, down to 0.520 in each of all 10. Fisher's power, computed on a lattice whose error is enclosed rather than estimated, falls from 96.1% to 45.0%; Tippett's minimum-p test falls from 99.6% to 18.5%. Fisher beats Stouffer while the signal sits in 6 or fewer of the 10 studies and loses from 7 on. Open circles are Fisher's power counted on 20,000 trials a point.

For each number of studies k, the largest number m of them the same total signal can be spread over while Fisher's combination is still more powerful than Stouffer's, with Stouffer held at 50% power. Every point is located by a bracketed power computation that does not straddle the comparison. It is 1 of 2, 6 of 10 and 18 of 40: a falling share of the studies — 75% at four, 60% at ten, 45.0% at 40 — and a count growing faster than the square root of k. Neither a fixed fraction nor √k describes it.

The exact one-sided binomial test on ten trials at a null probability of one half can return only eleven p-values, and alone it rejects 1.07% of the time at a nominal 5%; its mid-p version rejects 5.47%. Combining k of them by Fisher's method, with every multiset of outcomes enumerated, the true size is 1.07%, 1.59%, 1.67%, 1.33%, 1.12%, 0.94%, 0.71%, 0.55% at k = 1, 2, 3, 4, 5, 6, 8, 10 from exact p-values, and 5.47%, 3.36%, 3.76%, 4.26%, 4.04%, 4.01%, 3.98%, 3.89% from mid-p values. At ten studies: 0.55% and 3.89%. Open circles are sizes counted on 40,000 simulated sets, with two standard errors.

Each combination alone rejects exactly 5% of null sets. Counted on 1,000,000 sets of ten null studies: Fisher or Stouffer 6.63%, Fisher or Tippett 8.05%, Stouffer or Tippett 8.96%, any of the three 9.66% — enclosed on a two-dimensional lattice between 9.18% and 10.09% — and all three together 1.05%. The three sizes add to 15%.

The threshold giving a 5% family-wise error rate, read back as a number of independent analyses. At no correlation it is 20.05; at 0.6 it is 11.37; at 0.95 it is 2.58. Bonferroni divides by 20 throughout.

The prespecified analysis detects an effect that is in it 52% of the time at two standard errors and an effect elsewhere 5% of the time. The corrected slate detects it 23% of the time wherever it is. The two are worth the same when the chance of having named the right analysis is 38% — and that figure rises to 91% at four standard errors.

At two standard errors the estimate that clears an uncorrected 5% threshold averages 1.35 times the truth, and the one that clears the family-wise threshold averages 1.69 times it. The correction fixes the error rate by demanding a larger estimate, and a larger estimate is a more selected one.

Each bar is the reported effect as a multiple of the truth; each note is the share of families that clear that threshold at all. The uncorrected threshold lets 68.5% through at 1.35 times the truth; Bonferroni lets 17.0% through at 1.76 times it.

Where it is used

16 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 16 different questions.

All 80 figures

FieldsThreadsSeriesConceptsAll essaysSearch