What a p-value does not say
A p-value answers one question: if there were no effect, how surprising would this data be? It is a useful question. It is not the question anyone actually has, and the gap between the two is where most misreading of statistics lives.
The first missing number: how big
The same p-value corresponds to wildly different effects depending on the sample size.
To reach p = 0.04 needs an effect of about 0.71 standard deviations at ten observations, and about 0.045 at two thousand. Those are not similar findings. One would be transformative in most fields and the other would be a rounding error, and the p-value is identical.
This is why an effect size and an interval should be reported and the p-value should not be reported alone. The interval carries both pieces of information at once: where the effect is, and how precisely it is pinned down.
The second missing number: how likely it was to be found
Power is the probability of detecting an effect that is really there. It depends on the effect size, the sample size and the threshold, and it is computable in advance.
The consequence people find hardest is the one about negative results. At n = 20, an effect of 0.3 standard deviations is detected about 20% of the time. So if that effect is real and the study is run, the most likely outcome is a non-significant result.
“No significant difference” therefore does not mean “no difference.” In an underpowered study it means almost nothing at all — the study was never capable of distinguishing the two possibilities, and reporting its failure to do so as evidence of absence is an error the design guaranteed.
The measured comparison: the same real effect of 0.3 standard deviations is found 13% of the time at n = 10 and 97% at n = 160. Nothing about the effect changed.
The third missing number: how many chances there were
A p-value assumes one pre-specified analysis. If several were available and the best was reported, the number no longer means what it says — and it does not require any dishonesty for this to happen. Twenty correct analyses of data with no effect in it find something 58% of the time.
This is the hardest of the three to fix, because unlike sample size and effect size it is not visible in the output. A paper reporting one analysis looks the same whether it was the only one considered or the best of thirty.
What to report instead
The recommendation is not “abandon p-values”. It is that a p-value is a fragment of a report rather than a report.
The effect size, in units a reader can judge.
An interval, which is what the data is compatible with, and which should have had its coverage checked for the procedure and sample size in question.
The power the study had for effects worth caring about — computed in advance, which also stops the study being run if it cannot answer the question.
How many analyses were possible, and whether the reported one was pre-specified.
With those four, the p-value adds very little, which is the clearest evidence that it was never sufficient on its own.
Why it survives
Worth a paragraph, because the criticism above is a century old and the practice has not changed much.
A p-value is a single number with a conventional threshold, and that combination is administratively irresistible. It converts a judgement into a decision, it is comparable across studies that have nothing else in common, and it lets someone who has not read the work sort the results.
Everything better is worse in those specific ways: an effect size needs domain knowledge to interpret, an interval requires the reader to think about a range, and a power calculation has to be done before the data exists. All of them ask more of the reader, and the p-value’s whole appeal is that it asks nothing.
That is not an argument for keeping it. It is an argument for understanding why the correction is a hundred years old and still needs making.
What the number actually is
Worth stating carefully, because most of the misreadings are misreadings of the definition rather than of the mathematics.
A p-value is P(data at least this extreme | the null is true). Every part of that matters.
It is a probability of the data, not of the hypothesis. It says nothing directly about whether the null is true, and converting between the two requires a prior probability that the p-value does not contain — the same reversal that makes a positive test misleading for a rare condition.
It is conditional on the null being true, so it can never be evidence about how likely that is on its own.
And “at least this extreme” means it includes data that was not observed. That is a real philosophical oddity — the number depends on results that did not happen — and it is why the stopping rule matters: a value that would have been extreme under one sampling plan is not extreme under another.
The three misreadings
Each of these is common enough to appear in published work.
“p = 0.03 means there is a 3% chance the null is true.” This inverts the conditional. The actual posterior probability depends on the prior and on the power, and under reasonable assumptions a p of 0.03 corresponds to a much weaker case against the null than 3% would suggest.
“p = 0.20 means there is no effect.” It means the data is unsurprising under the null. At low power that is the expected outcome of a real effect, and the study has not distinguished the possibilities.
“p = 0.001 means the effect is large.” It means the effect is well-resolved relative to the noise, which at a large sample size can happen for an effect of no practical size at all. This is the confusion the three-sample-size figure above exists to settle.
The threshold, and where it came from
The 0.05 convention has no theoretical basis and its origin is documented: Fisher chose it as a convenient round number, remarked that it was arbitrary, and expected it to be adjusted to context.
The consequences of that arbitrary choice becoming a rule are structural rather than statistical.
It creates a cliff where the underlying quantity is continuous: 0.049 and 0.051 are essentially identical evidence and are treated as opposite results.
It creates an incentive at the cliff edge, which is the mechanism behind the lump of published p-values just below 0.05 that the replication literature documented.
And it converts a judgement into a decision, which is administratively convenient and is the reason the convention survives a century of criticism.
None of that is an argument against having thresholds where a decision genuinely has to be made. It is an argument against a single threshold, chosen for convenience, applied to every field regardless of what is at stake.
What replaces it
The alternatives are less convenient and each addresses a different part of the problem.
Report the interval. It carries the location and the precision together, and it makes the sample size visible in a way a p-value does not.
Report the effect in units someone can judge. Standard deviations are comparable across studies and meaningless to a practitioner; the natural units are the reverse. Both are worth having.
Pre-specify the analysis, which is the only measure that addresses the forking-paths inflation at its cause.
And where a decision must be made, state the threshold in advance along with the power the design provides, so that a non-significant result means something specific rather than nothing at all.
The number that would actually help
If a p-value cannot be read alone, it is fair to ask what single number could be.
The honest answer is none, and the closest candidate is worth knowing: the likelihood ratio between a stated alternative and the null. Unlike a p-value it is symmetric — it can favour the null — and it can be combined with a prior to give the posterior the reader actually wants.
Its cost is that it requires stating the alternative, which is exactly the work a p-value lets people avoid. “Is there an effect” needs no alternative; “how much more likely is an effect of 0.3 than none” needs a number that has to be defended.
That is the real trade behind the whole debate. The p-value’s popularity comes from requiring no commitment about what would count as an effect worth finding, and everything that improves on it requires exactly that commitment.
Which is a reason to make it. A study that cannot say what size of effect would matter has not specified its question, and the statistics were never going to rescue that.
The proposals to change the threshold
Several groups have proposed lowering the conventional threshold from 0.05 to 0.005, and the argument is worth understanding because it is a partial fix with a stated cost.
The case for it: at 0.05, under plausible priors and typical power, a substantial fraction of significant findings are false. Tightening the threshold reduces that fraction sharply, because the evidence required is much stronger.
The case against: it addresses only the false-positive side. A tighter threshold means lower power at any given sample size, so more real effects are missed, and among those that survive the winner’s curse inflation is worse — because the truncation point has moved further out.
So the proposal trades one error rate against the other, and whether it is an improvement depends on the relative cost of the two errors in a given field. It is not a general correction, and it does nothing at all about the multiplicity problem, which operates whatever the threshold is.
The more radical proposal — abandon thresholds and report estimates with intervals — avoids the trade entirely by declining to make a binary decision. That is right when no decision has to be made and unhelpful when one does, which is the honest summary of a long argument.