Tests, and the second number

What a p-value does not say

The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.

A p-value answers one question: if there were no effect, how surprising would this data be? It is a useful question. It is not the question anyone actually has, and the gap between the two is where most misreading of statistics lives.

What a one-sample t test at n = 20 can detectAt an effect of 0.5 standard deviations the test finds it 56% of the time. Below that, a non-significant result is the expected outcome of a real effect — which is why "no significant difference" is not evidence of no difference.00.2500.5000.750100.5001true effect, in standard deviationsprobability of rejecting80%, the usual target56% at d = 0.54,000 simulated studies per pointthe second number a p-value needs
Fig. 1 What a one-sample t test at a given sample size can actually detect. At small effects a real difference is missed most of the time, and the miss is the expected outcome rather than an anomaly.

The first missing number: how big

The same p-value corresponds to wildly different effects depending on the sample size.

To reach p = 0.04 needs an effect of about 0.71 standard deviations at ten observations, and about 0.045 at two thousand. Those are not similar findings. One would be transformative in most fields and the other would be a rounding error, and the p-value is identical.

The effect behind a p-value of 0.04, at three sample sizesA p of 0.04 at ten observations needs an effect of 0.76 standard deviations; at two thousand it needs 0.046. The p-value alone does not say which of these happened, which is why it should never be reported alone.ten observations0.758 sdp = 0.04a hundred0.208 sdp = 0.04two thousand0.046 sdp = 0.04the effect that gives the same p-valueall three reach p = 0.04the number does not say how big the effect is
Fig. 2 The effect behind a fixed p-value at three sample sizes. All three would be reported as significant at the same threshold.

This is why an effect size and an interval should be reported and the p-value should not be reported alone. The interval carries both pieces of information at once: where the effect is, and how precisely it is pinned down.

12,000 studies of a real effect of 0.3, n = 16Power is 20%. The studies that reached significance report a mean effect of 0.613 — 2.04 times the truth. Every one of them is honest; the selection did the inflating.02505007501e+3-0.50000.5001effect the study reportsstudiesthe truth, 0.3what gets published, 0.61shaded: reached p below 0.05power 20%, inflation 2.04×
Fig. 3 What low power does to the studies that survive the significance filter.

The second missing number: how likely it was to be found

Power is the probability of detecting an effect that is really there. It depends on the effect size, the sample size and the threshold, and it is computable in advance.

The consequence people find hardest is the one about negative results. At n = 20, an effect of 0.3 standard deviations is detected about 20% of the time. So if that effect is real and the study is run, the most likely outcome is a non-significant result.

“No significant difference” therefore does not mean “no difference.” In an underpowered study it means almost nothing at all — the study was never capable of distinguishing the two possibilities, and reporting its failure to do so as evidence of absence is an error the design guaranteed.

The measured comparison: the same real effect of 0.3 standard deviations is found 13% of the time at n = 10 and 97% at n = 160. Nothing about the effect changed.

The false-positive rate against the number of analyses, on pure noiseThe data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 57% of the time.00.2000.4000.6000.8005101520analyses available to the researcherchance of finding something significantthe nominal 5%if the analyses were independentcorrelated, as real ones areno effect present anywhereevery individual analysis is correct
Fig. 4 The false-positive rate against the number of analyses available, on data with no effect.

The third missing number: how many chances there were

A p-value assumes one pre-specified analysis. If several were available and the best was reported, the number no longer means what it says — and it does not require any dishonesty for this to happen. Twenty correct analyses of data with no effect in it find something 58% of the time.

This is the hardest of the three to fix, because unlike sample size and effect size it is not visible in the output. A paper reporting one analysis looks the same whether it was the only one considered or the best of thirty.

20,000 p-values from a true null, n = 12Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0165 (p = 0.13). That flatness is the check that catches an error a single rejection rate would miss.05001e+30.2000.4000.6000.800p-valuecountflatKS distance 0.0165a p-value that is not uniform is not a p-value
Fig. 5 What a p-value is when the null is true: uniform, and therefore uninformative on its own.

What to report instead

The recommendation is not “abandon p-values”. It is that a p-value is a fragment of a report rather than a report.

The effect size, in units a reader can judge.

An interval, which is what the data is compatible with, and which should have had its coverage checked for the procedure and sample size in question.

The power the study had for effects worth caring about — computed in advance, which also stops the study being run if it cannot answer the question.

How many analyses were possible, and whether the reported one was pre-specified.

With those four, the p-value adds very little, which is the clearest evidence that it was never sufficient on its own.

Coverage of four nominal 95% intervals, n = 30Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.0.4000.6000.80010.2000.4000.6000.800the true proportionactual coverage of a nominal 95% intervalnominal 95%Wald — the textbook oneWilsonAgresti–CoullClopper–Pearsonsummed over all 31 outcomesno interval is 95% until it is counted
Fig. 6 An interval, which carries both the location and the precision the p-value leaves out.

Why it survives

Worth a paragraph, because the criticism above is a century old and the practice has not changed much.

A p-value is a single number with a conventional threshold, and that combination is administratively irresistible. It converts a judgement into a decision, it is comparable across studies that have nothing else in common, and it lets someone who has not read the work sort the results.

Everything better is worse in those specific ways: an effect size needs domain knowledge to interpret, an interval requires the reader to think about a range, and a power calculation has to be done before the data exists. All of them ask more of the reader, and the p-value’s whole appeal is that it asks nothing.

That is not an argument for keeping it. It is an argument for understanding why the correction is a hundred years old and still needs making.

What the number actually is

Worth stating carefully, because most of the misreadings are misreadings of the definition rather than of the mathematics.

A p-value is P(data at least this extreme | the null is true). Every part of that matters.

It is a probability of the data, not of the hypothesis. It says nothing directly about whether the null is true, and converting between the two requires a prior probability that the p-value does not contain — the same reversal that makes a positive test misleading for a rare condition.

It is conditional on the null being true, so it can never be evidence about how likely that is on its own.

And “at least this extreme” means it includes data that was not observed. That is a real philosophical oddity — the number depends on results that did not happen — and it is why the stopping rule matters: a value that would have been extreme under one sampling plan is not extreme under another.

The three misreadings

Each of these is common enough to appear in published work.

“p = 0.03 means there is a 3% chance the null is true.” This inverts the conditional. The actual posterior probability depends on the prior and on the power, and under reasonable assumptions a p of 0.03 corresponds to a much weaker case against the null than 3% would suggest.

“p = 0.20 means there is no effect.” It means the data is unsurprising under the null. At low power that is the expected outcome of a real effect, and the study has not distinguished the possibilities.

“p = 0.001 means the effect is large.” It means the effect is well-resolved relative to the noise, which at a large sample size can happen for an effect of no practical size at all. This is the confusion the three-sample-size figure above exists to settle.

The threshold, and where it came from

The 0.05 convention has no theoretical basis and its origin is documented: Fisher chose it as a convenient round number, remarked that it was arbitrary, and expected it to be adjusted to context.

The consequences of that arbitrary choice becoming a rule are structural rather than statistical.

It creates a cliff where the underlying quantity is continuous: 0.049 and 0.051 are essentially identical evidence and are treated as opposite results.

It creates an incentive at the cliff edge, which is the mechanism behind the lump of published p-values just below 0.05 that the replication literature documented.

And it converts a judgement into a decision, which is administratively convenient and is the reason the convention survives a century of criticism.

None of that is an argument against having thresholds where a decision genuinely has to be made. It is an argument against a single threshold, chosen for convenience, applied to every field regardless of what is at stake.

What replaces it

The alternatives are less convenient and each addresses a different part of the problem.

Report the interval. It carries the location and the precision together, and it makes the sample size visible in a way a p-value does not.

Report the effect in units someone can judge. Standard deviations are comparable across studies and meaningless to a practitioner; the natural units are the reverse. Both are worth having.

Pre-specify the analysis, which is the only measure that addresses the forking-paths inflation at its cause.

And where a decision must be made, state the threshold in advance along with the power the design provides, so that a non-significant result means something specific rather than nothing at all.

The number that would actually help

If a p-value cannot be read alone, it is fair to ask what single number could be.

The honest answer is none, and the closest candidate is worth knowing: the likelihood ratio between a stated alternative and the null. Unlike a p-value it is symmetric — it can favour the null — and it can be combined with a prior to give the posterior the reader actually wants.

Its cost is that it requires stating the alternative, which is exactly the work a p-value lets people avoid. “Is there an effect” needs no alternative; “how much more likely is an effect of 0.3 than none” needs a number that has to be defended.

That is the real trade behind the whole debate. The p-value’s popularity comes from requiring no commitment about what would count as an effect worth finding, and everything that improves on it requires exactly that commitment.

Which is a reason to make it. A study that cannot say what size of effect would matter has not specified its question, and the statistics were never going to rescue that.

The proposals to change the threshold

Several groups have proposed lowering the conventional threshold from 0.05 to 0.005, and the argument is worth understanding because it is a partial fix with a stated cost.

The case for it: at 0.05, under plausible priors and typical power, a substantial fraction of significant findings are false. Tightening the threshold reduces that fraction sharply, because the evidence required is much stronger.

The case against: it addresses only the false-positive side. A tighter threshold means lower power at any given sample size, so more real effects are missed, and among those that survive the winner’s curse inflation is worse — because the truncation point has moved further out.

So the proposal trades one error rate against the other, and whether it is an improvement depends on the relative cost of the two errors in a given field. It is not a general correction, and it does nothing at all about the multiplicity problem, which operates whatever the threshold is.

The more radical proposal — abandon thresholds and report estimates with intervals — avoids the trade entirely by declining to make a binary decision. That is right when no decision has to be made and unhelpful when one does, which is the honest summary of a long argument.