Tests, and the second number

The p-value a replication gets

Under a true null a p-value is flat. Under a real effect its distribution is closed form and wide — a study with 80% power returns anything from 4.4×10⁻⁵ to 0.13 in eight runs of ten — and the chance that an exact replication of a p = 0.05 result is significant again is exactly one half, under both of the models people use without naming them.

Worth reading first: A p-value that is not flat is not a p-value · What a p-value does not say.

A p-value that is not flat is not a p-value establishes one fact and builds a discipline on it: when the null hypothesis is true, the p-value is uniform on the unit interval. That is a statement about one world. Every study anybody cares about is run because the null might be false, and in that world the p-value has a distribution too — not a flat one, and not a spike at zero either.

For the simplest test in the subject that distribution is closed form, and it answers two questions that are usually argued about rather than computed. The first is how precise a p-value is: if a study were run again exactly as designed, how different would its p-value be? The second is the one a reader of a single result actually has: given that a study reported p = 0.01, what will a replication report?

The first has an answer with no assumptions in it beyond the design. The second has no answer at all until something is said about the effect the first study measured, and there are two ways of saying it that are in constant use and almost never named. They agree at exactly one point, and the point is p = 0.05.

A p-value under a real effect has a closed form

Take a two-sided z-test with a known standard deviation. Its statistic is normal with unit variance and a mean δ, the noncentrality, which is the true effect measured in standard errors of the study. The p-value is p=2Q(z)p = 2Q(|z|), where QQ is the upper normal tail. A p-value falls at or below a level aa exactly when z|z| is at least the two-sided critical value for aa, so

P(pa)=Q(za/2δ)+Q(za/2+δ).P(p \le a) = Q(z_{a/2} - \delta) + Q(z_{a/2} + \delta).

At δ = 0 this is aa for every aa, which is the uniformity of the null case recovered as the special case it is. At any other δ it is the whole distribution of the study’s p-value, and one point of it has a familiar name: P(p0.05)P(p \le 0.05) is the power. The parallel with the null is exact. A rejection rate is one point of the null distribution and uniformity is the rest of it; power is one point of the alternative distribution, and the rest of it is almost never drawn.

The p-value of a study of a true null, twenty thousand timesTwenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.0000 standard deviations from the null, a noncentrality of 0.000. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 0.1 to 0.9, 0.95 orders of magnitude, the median is 0.5, and 5.0% fall below 0.05, which is what the power means.00.2000.4000.600the study's two-sided p-value, spaced by powers of tenshare of the studies10⁻⁷10⁻⁶10⁻⁵10⁻⁴0.0010.010.110.05middle 80%: 0.1 to 0.9bars counted, line closed form5% below 0.05
Fig. 1 The null case on the axis this essay uses. Twenty thousand z-tests on data with no effect in it, their p-values binned by powers of ten. Flat on a linear scale is this rising staircase on a logarithmic one, since each decade holds ten times the width of the one to its left. The slider moves the same study through powers of 20%, 50%, 80% and 95%.

The axis is spaced by powers of ten because that is the only scale on which a p-value’s distribution can be read at all. On a linear axis, every p-value below 0.01 shares one sliver beside zero. Those are the values that decide whether a result is reported as strong, and a linear axis erases the differences between them.

The figure below is the same study run at 80% power, which is the power most study protocols promise. Each of its twenty thousand runs draws 25 observations whose true mean sits 0.5603 standard deviations from the null, a noncentrality of 2.802, and applies the test a reader would apply. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13. That is 3.46 orders of magnitude. The median is 0.0051; one run in forty is below 1.9×10⁻⁶ and one in forty is above 0.40.

The p-value of a study with 80% power, twenty thousand times. Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.5603 standard deviations from the null, a noncentrality of 2.802. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, the median is 0.0051, and 80.0% fall below 0.05, which is what the power means.
Fig. 2 The same study at 80% power, twenty thousand runs binned by powers of ten. The middle eighty per cent of its p-values spans 3.46 orders of magnitude, from 4.4×10⁻⁵ on the left to 0.13 on the right, around a median of 0.0051.

Those are closed-form quantiles, and the bars were counted rather than derived. At each of the five quantiles the counted distribution function has to agree with the closed form within four binomial standard errors. It does at every one, and the worst is 1.54 standard errors. The share of counted p-values below 0.05 is 80.05%, against a power of 80% that was solved for before any data was drawn. Two routes that share nothing but the definition agree, which is the habit every number here is held to.

A study designed to be a coin flip lands its median on the line

At 50% power the noncentrality is 1.960, which is the critical value itself. The statistic is centred on the threshold, so half of all runs fall on each side of it. The median p-value is 0.050 and the middle eighty per cent runs from 0.0012 to 0.48.

The p-value of a study with 50% power, twenty thousand times. Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.3920 standard deviations from the null, a noncentrality of 1.960. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 0.0012 to 0.48, 2.61 orders of magnitude, the median is 0.05, and 50.0% fall below 0.05, which is what the power means.
Fig. 3 The same study at half the effect-to-noise ratio needed for 80% power. The median sits on 0.05 and the middle eighty per cent spans 2.61 orders of magnitude, from 0.0012 on the left to 0.48 on the right. Every one of these studies measured a real effect.

The spread is not a small-sample artefact. Nothing in the closed form depends on the number of observations except through δ. A study of ten thousand at 50% power has exactly this distribution of p-values, because 50% power at ten thousand observations means an effect small enough that the statistic is again centred on 1.96. The count agrees here too: at the five quantiles the worst departure is 1.31 standard errors, and the counted median is 0.049.

Reading the figure as a population of honest studies gives the sentence it exists for. In a literature of experiments that all had the same real effect and the same 50% power, one in ten reports p below 0.0012 and one in ten reports p above 0.48. The first would be written up as compelling evidence and the second as a null result. They differ by nothing except the draw, and what a p-value does not say makes the same point from the other side: one p-value, many possible effects. Here it is one effect and many possible p-values.

More power moves a p-value down and widens the range it lands in

The intuition this breaks is that the imprecision is a symptom of weak studies, and that a well-powered study pins its p-value down. It does the opposite, measured in the unit a p-value is actually read in.

Where a study's p-value lands, against the power of the study. Quantiles of a two-sided z-test's p-value in closed form, at powers from 5% to 99%, on a scale of powers of ten. The pale band is the middle ninety-five per cent, the darker one the middle eighty per cent, and the line the median. At 5% power the null is true and the middle eighty per cent is 0.1 to 0.9. At 80% power it is 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, and at 99% it is 5.01: more power moves a p-value down and widens the range it lands in.
Fig. 4 The quantiles of a z-test’s p-value at every power from 5% to 99%, closed form. The pale band is the middle ninety-five per cent, the darker band the middle eighty, the line the median. At the left edge the null is true and the eighty per cent band is 0.1 to 0.9. Moving right, the whole band falls and grows taller.

The middle eighty per cent spans 0.95 orders of magnitude at 5% power, 1.69 at 20%, 2.61 at 50%, 3.46 at 80%, 4.29 at 95% and 5.01 at 99%. It widens at every step of the sweep.

The mechanism is one derivative. The statistic’s own spread is fixed: whatever the power, zz has standard deviation one. What changes is how many decades of p-value one unit of zz is worth, and that is

dlog10pdz=φ(z)Q(z)ln10,-\frac{d \log_{10} p}{dz} = \frac{\varphi(z)}{Q(z)\,\ln 10},

the Mills ratio’s reciprocal divided by ln 10. It is 1.02 decades per unit at the critical value and 1.84 at z=4z = 4, growing roughly in proportion to zz because the normal tail falls like ez2/2e^{-z^2/2}. A well-powered study sits further out, where the same unit of statistical noise is stretched across more decades. The p-value’s spread in orders of magnitude is the statistic’s fixed spread seen through a lens whose magnification rises with the effect.

So the claim that a study with 95% power returns a p-value known to within an order of magnitude is refused by arithmetic rather than by argument. At 95% power the eighty per cent range runs from 1.0×10⁻⁶ to 0.020.

None of this makes a p-value useless. It makes it a draw from a distribution whose width is worth knowing before reading anything into the difference between 0.003 and 0.0003. Under a fixed real effect at 80% power, both are well inside the middle eighty per cent of the same study.

What the first p-value says about the second

Now the question a reader has. A study reports a two-sided p-value p1p_1, with a z statistic z1z_1 taken as positive. An exact replication — same design, same size, same population — is run. Its statistic is z2N(δ,1)z_2 \sim N(\delta, 1) for the same unknown δ. Call the replication a success if it is significant at 5% in the same direction, z2>1.96z_2 > 1.96. A significant result in the opposite direction is not a replication of anything.

The chance of success depends on δ, and δ is not known. Something has to be put in its place, and the two things that usually are give different answers.

The estimate taken as the truth. Set δ = z1z_1: the effect is whatever the first study measured. Then z2N(z1,1)z_2 \sim N(z_1, 1) and the chance of success is Φ(z11.96)\Phi(z_1 - 1.96). After p1=0.01p_1 = 0.01 that is 73.1%; after 0.005 it is 80.2%; after 0.001 it is 90.8%.

This number has another name. It is observed power: the power the study would have had if the true effect were the one it observed. Observed power is a function of the p-value and of nothing else — at p=0.05p = 0.05 it is 50% plus the 0.004 points a significant result in the wrong direction contributes, at p=0.01p = 0.01 it is 73.1% — so reporting it beside a p-value adds a relabelled copy of the same number. Power is a property of a design and an assumed effect, which is why how many subjects a trial needs is computed before any data exists.

A flat prior on the effect. Treat δ as unknown with every value equally plausible beforehand. The posterior for δ after seeing z1z_1 is then N(z1,1)N(z_1, 1). A replication adds its own unit of noise on top of that uncertainty, so z2N(z1,2)z_2 \sim N(z_1, 2). The chance of success is Φ((z11.96)/2)\Phi\big((z_1 - 1.96)/\sqrt{2}\big): 66.8% after 0.01, 72.5% after 0.005, 82.7% after 0.001.

This second model is a Bayesian statement, whether or not it is presented as one. Its variance of two is one unit for the replication and one unit for not knowing δ, and the second unit exists only because δ was given a distribution. The flat prior is also the prior under which a credible interval and a confidence interval coincide endpoint for endpoint, which is why the model reads as if it assumed nothing. The first model is not a frequentist statement about the replication either. It is a probability computed as if an estimate were a known parameter, and its only merit is that it is easy.

The chance an exact replication reaches significance, given the first p-value. Two models for the effect behind a first two-sided p-value. Taking the estimate as the truth gives a replication significant in the same direction 73.1% of the time after p = 0.01 and 90.8% after 0.001. A flat prior, which carries the estimate's own uncertainty into the prediction, gives 66.8% and 82.7%. At p = 0.05 both are exactly one half. The points are counted from 600,000 simulated pairs of studies with effects spread flat, keeping the pairs whose first p-value fell near 0.05, 0.01 and 0.001, and they agree with the curve within their error bars of two standard errors.
Fig. 5 The chance of a same-direction significant replication against the first study’s p-value, under both models. The dashed curve takes the estimate as the truth and the solid curve carries the estimate’s uncertainty. They meet at 0.05, marked, and separate on either side. The points are counted from simulated pairs of studies, with bars of two standard errors.

One half at the threshold, whatever is assumed about symmetry

At p1=0.05p_1 = 0.05 both models give exactly one half, and computed, both agree with one half to twelve decimal places. The reason takes one sentence. Both predictive distributions for z2z_2 are symmetric and centred on z1z_1, and at p1=0.05p_1 = 0.05 the first statistic is the threshold, so half of each predictive lies on either side of it.

That makes one half a property of any model whose prediction for the replication is centred on the first result, however wide it is. It is the most generous answer available to anybody who does not believe effects are systematically larger than they are measured. The familiar misreading — that p = 0.05 leaves a 95% chance the result is real, so a replication will very probably succeed — is therefore refused by the most optimistic model in the family.

The counted route is a population of study pairs, not a formula. Six hundred thousand effects are drawn uniformly between −7 and 7 standard errors, which stands in for the flat prior since a flat prior cannot be sampled. Each effect is studied twice, and the pairs whose first p-value falls in a narrow window are kept. For p1p_1 between 0.045 and 0.055 that leaves 7,223 pairs, and 50.34% of them replicate, with a standard error of 0.59 points. The comparison is not with the ideal one half but with the same truncated model integrated over the same window by quadrature, which gives 50.05%; the idealised value sits beside it at exactly 50%. The windows near 0.01 and 0.001 give 67.23% on 5,994 pairs against 66.88%, and 82.72% on 4,838 pairs against 82.68%.

The truncation at seven is chosen rather than hidden. At the largest first statistic used, 3.29, the posterior mass it removes is about one part in ten thousand.

The range a replication’s p-value lands in

A success rate is one point of a distribution again, and the whole distribution is as closed as the first. The replication’s two-sided p-value has the quantiles of the folded predictive, and they are wide.

The p-value an exact replication gets, after a first p of 0.05. The predictive distribution of a replication's two-sided p-value after a first p of 0.05, under three models for the effect, on a scale of powers of ten. Each bar is the middle eighty per cent, the thin line the middle ninety-five, the tick the median. Taking the estimate as the truth: 0.0012 to 0.48. A flat prior: 1.6×10⁻⁴ to 0.65. A sceptical prior with τ = 2: 0.001 to 0.74, median 0.11. The open circles on the flat row are the counted tenth, fiftieth and ninetieth percentiles from 7,223 simulated pairs whose first p fell near 0.05.
Fig. 6 Where a replication’s p-value is predicted to land after a first p of 0.05, one row per model. Bars are the middle eighty per cent, thin lines the middle ninety-five, ticks the medians. The open circles on the flat-prior row are percentiles counted from the simulated pairs.

After p1=0.05p_1 = 0.05, taking the estimate as the truth predicts a replication p between 0.0012 and 0.48 eight times in ten. The flat prior predicts between 1.6×10⁻⁴ and 0.65. The median of both is the first p-value itself, to within rounding: whatever the first study reported, the replication is as likely to report something smaller as something larger. The counted pairs put the tenth, fiftieth and ninetieth percentiles at 1.5×10⁻⁴, 0.047 and 0.63. These were counted from a window rather than read at a point, and they sit where the closed form puts them.

The row that is not centred on the first result is the third, and it is the subject of the next section. Under it the median replication p-value after 0.05 is 0.11.

The p-value an exact replication gets, after a first p of 0.001. The predictive distribution of a replication's two-sided p-value after a first p of 0.001, under three models for the effect, on a scale of powers of ten. Each bar is the middle eighty per cent, the thin line the middle ninety-five, the tick the median. Taking the estimate as the truth: 4.8×10⁻⁶ to 0.045. A flat prior: 3.3×10⁻⁷ to 0.14. A sceptical prior with τ = 2: 1.4×10⁻⁵ to 0.35, median 0.0085. The open circles on the flat row are the counted tenth, fiftieth and ninetieth percentiles from 4,838 simulated pairs whose first p fell near 0.001.
Fig. 7 The same three predictions after a first p of 0.001, which most readers would call a strong result. The flat prior still puts a tenth of replications above 0.14.

After p1=0.001p_1 = 0.001 the plug-in range is 4.8×10⁻⁶ to 0.045 and the flat-prior range is 3.3×10⁻⁷ to 0.14. A tenth of replications of a genuinely strong first result are predicted to come back at a p-value above 0.14, by a model that assumes nothing about the effect except that no value of it was preferred beforehand. The counted percentiles are 3.0×10⁻⁷, 0.0010 and 0.14.

A replication that returns p = 0.2 after an original p = 0.001 is not, on its own, evidence that anything went wrong. Under the flat prior it sits a little above the ninetieth percentile of what a real, unchanged effect produces. What it is evidence of depends on how many such failures there are against how many were expected. That is a count over a literature, not a verdict on a pair.

When most effects are small

Both models so far are centred on the first result, and that is the assumption to question. A field does not study effects drawn from a flat distribution. It studies effects mostly near zero with a few large ones, and a result that reaches p = 0.05 is more likely a small effect measured generously than a large effect measured accurately. That is the winner’s curse seen from the side of the next study rather than the reported estimate.

The third model writes it down. The noncentrality is drawn from N(0,τ2)N(0, \tau^2), so the posterior for δ is N(sz1,s)N(s z_1, s) with s=τ2/(1+τ2)s = \tau^2/(1+\tau^2), and the replication’s predictive is N(sz1,1+s)N(s z_1, 1 + s). The factor ss is the shrinkage weight: its complement is the BB that decides how far a group’s estimate is pulled in a hierarchical model, and here it pulls the first study’s statistic towards zero before the replication is predicted.

τ is easier to read as a property of a literature than as a variance. When noncentralities are drawn this way the first statistic is N(0,1+τ2)N(0, 1+\tau^2), so the share of studies reaching significance is 2Q(1.96/1+τ2)2Q\big(1.96/\sqrt{1+\tau^2}\big). That is 16.58% at τ = 1, 38.07% at τ = 2 and 63.45% at τ = 4. A literature with τ = 2 is one in which about two studies in five are significant.

The same chance, when most effects are assumed to be small. The chance of a same-direction significant replication when the effect is drawn from a normal prior centred on zero with standard deviation τ, measured in standard errors of one study. After a first p of 0.05 it is 46.7% at τ = 4, 38.5% at τ = 2 and 21.2% at τ = 1, against one half under the flat prior drawn pale. The points are counted from 600,000 pairs of studies with effects drawn from the τ = 2 prior.
Fig. 8 The replication chance under priors that put most effects near zero, with the flat prior drawn pale for reference. Every sceptical curve sits below it at every first p-value. The points are counted from pairs of studies whose effects were drawn from the τ = 2 prior.

After p1=0.05p_1 = 0.05 the chance of a same-direction significant replication is 46.7% at τ = 4, 38.5% at τ = 2 and 21.2% at τ = 1. After p1=0.001p_1 = 0.001 it is 79.3%, 69.2% and 39.9%. The counted route here needs no truncation, since the prior is proper. Near 0.05, 12,306 pairs replicate 38.10% of the time with a standard error of 0.44, against 38.54% by quadrature over the window. Near 0.01, 7,526 pairs replicate 53.16% against 53.02%. Near 0.001, 4,043 pairs replicate 68.49% against 69.21%, a departure of one standard error.

The consequence for reading replication projects is the uncomfortable one. A literature with τ = 2, in which every effect is real in the sense of being non-zero and none of the analyses is wrong, replicates its p = 0.05 findings about 38.5% of the time. A replication rate near two in five does not by itself separate a literature of honest small effects from a literature with many true nulls. Separating the two needs the distribution of effects, which is the base rate the screening arithmetic runs on, and it cannot be read off the replication rate alone.

Which of these numbers are arithmetic and which are assumptions

Three kinds of claim sit in this essay, and they deserve different amounts of trust.

The spread of a p-value under a stated effect is arithmetic. The distribution function is two normal tails, its quantiles come from bisection on it, and twenty thousand real tests per power agree with it at five quantiles. The growth of the spread with power is a derivative. None of it depends on anything but the test.

The replication chances are arithmetic given a model, and the model is an assumption. The plug-in model assumes the effect is known, which is false. The flat prior assumes no effect size was more plausible than another beforehand, which is generous. The sceptical prior assumes a normal spread of effects centred on zero with a stated width. No study can check that assumption, and the three τ values are shown side by side rather than one being chosen. The counted routes check the arithmetic under each model, never the model. Choosing the prior is a component with a size, and here the size moves the answer at p = 0.05 from one half to one fifth.

One number is independent of the choice among centred models: one half at the threshold. It is exact, and it is enough to refuse the reading of a p-value as a replication probability.

Everything here is about an exact replication of a z-test with a known standard deviation. A real replication estimates its own variance, often uses a different sample size, and is rarely exact in population or procedure. Each of those changes widens the predictive, and none of them is measured here.

Where this goes next

A replication is two p-values about one effect, read one after the other. The natural continuation reads several at once. Given k studies with k p-values, what do they say together? The answer again rests on uniformity, since the standard ways of combining p-values are exact only because each p-value is flat when its null is true. And the combinations disagree, measurably, about which arrangement of evidence counts: one very small p-value among many large ones, or many moderately small ones. Two ways to combine p-values takes that up.

Two narrower questions are left open by this essay and are distinct from it. The first is the replication with a different sample size, where the predictive mean scales by the square root of the size ratio and a replication designed for “90% power at the observed effect” inherits both the plug-in model’s optimism and the first study’s selection. The second is the same analysis for a t-test at small n, where not knowing the variance thickens every tail used above. The p-value’s spread at the sizes where that matters has not been computed here, and the z-test’s figures should not be read as a bound on it. The same caution applies to a sampling plan that looks at the data more than once, where the first p-value is no longer a draw from the distribution this essay draws.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Effect sizeFlat priorMonte CarloObserved powerp-valuePosteriorPredictive distributionPriorQuadratureReplication probabilitySampling distributionShrinkageStatistical powerUniformityThe winner's curse