Tests, and the second number

What a p-value does not say

The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.

Worth reading first: A p-value that is not flat is not a p-value.

A p-value answers one question: if there were no effect, how surprising would this data be? It is a useful question. It is not the question anyone actually has, and the gap between the two is where most misreading of statistics lives.

Power at an effect of 0.5 standard deviationsThe curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm. 00.2500.5000.7501observations per arm (doubling)chance of detecting the effect64 per arm for 80%34% at 20 per arm4166425680%non-central t, δ = 0.5√(n/2)4,000 experiments per dot
Fig. 1 What a one-sample t test at a given sample size can actually detect. At small effects a real difference is missed most of the time, and the miss is the expected outcome rather than an anomaly.

The first missing number: how big

The same p-value corresponds to wildly different effects depending on the sample size.

To reach p = 0.04 needs an effect of about 0.76 standard deviations at ten observations, and about 0.046 at two thousand. Those are not similar findings. One would be transformative in most fields and the other would be a rounding error, and the p-value is identical.

The effect behind a p-value of 0.04, at three sample sizes. A p of 0.04 at ten observations needs an effect of 0.76 standard deviations; at two thousand it needs 0.046. The p-value alone does not say which of these happened, which is why it should never be reported alone.
Fig. 2 The effect behind a fixed p-value at three sample sizes. All three would be reported as significant at the same threshold.

This is why an effect size and an interval should be reported and the p-value should not be reported alone. The interval carries both pieces of information at once: where the effect is, and how precisely it is pinned down.

A reasonable objection at this point is that 0.04 is an arbitrary number, and that the spread might be an artefact of choosing a p-value nobody has any particular attachment to. It is not, and the cheapest way to show it is to redraw the same three bars at the threshold the convention is actually built on.

The effect behind a p-value of 0.05, at three sample sizes. A p of 0.05 at ten observations needs an effect of 0.72 standard deviations; at two thousand it needs 0.044. The p-value alone does not say which of these happened, which is why it should never be reported alone.
Fig. 3 The identical comparison at p = 0.05 rather than 0.04. Ten observations need 0.72 standard deviations, two thousand need 0.044, and the ratio between them is unchanged.

Nothing moved. Each bar shrank by about five per cent — the critical value at 0.05 is a little smaller than at 0.04 — and the ratio between the shortest bar and the longest barely moved, because it depends on the threshold only through the shape of the t distribution at nine degrees of freedom against two thousand. It is a shade over sixteen at either, and it would be a shade over sixteen at 0.001 or at 0.2.

That is the sharper version of the complaint. The problem is not that 0.05 is the wrong line to draw; moving the line changes every effect size by a few per cent and changes nothing about what the p-value fails to say. A reader who wants to know how big the effect is gets no more help from a p of 0.001 than from a p of 0.04, because the quantity they want was never in the number at all.

12,000 studies of a real effect of 0.3, n = 16. Power is 21%. The studies that reached significance report a mean effect of 0.621 — 2.07 times the truth. Every one of them is honest; the selection did the inflating.
Fig. 4 What low power does to the studies that survive the significance filter.

A p-value fixes a product, not either factor

The two effects quoted are not a coincidence of two sample sizes; they are one relation read at two points, and stating the relation is what makes the omission unmistakable.

A p-value is a statement about a statistic, and the statistic is the effect times the square root of the sample size. So fixing p fixes the product

dnconstant,d\sqrt{n} \approx \text{constant},

and says nothing whatever about either factor separately.

Check it on the two numbers. At ten observations, 0.76×10=2.400.76 \times \sqrt{10} = 2.40; at two thousand, 0.046×2000=2.060.046 \times \sqrt{2000} = 2.06. The same product to within seventeen per cent, the residual being the t multiplier at nine degrees of freedom sitting above the normal’s.

So an effect at a fixed p-value falls like 1/n1/\sqrt{n}, and the sixteen-fold difference between the two readings is nothing but a two-hundredfold difference in sample size, square-rooted.

Which is the sharpest way to state what the number is missing. The p-value reports a product of two things a reader needs separately, and the sample size — the one that is always printed somewhere — is exactly what is needed to recover the other.

The second missing number: how likely it was to be found

Power is the probability of detecting an effect that is really there. It depends on the effect size, the sample size and the threshold, and it is computable in advance.

The consequence people find hardest is the one about negative results. At n = 20, an effect of 0.3 standard deviations is detected about 20% of the time. So if that effect is real and the study is run, the most likely outcome is a non-significant result.

“No significant difference” therefore does not mean “no difference.” In an underpowered study it means almost nothing at all — the study was never capable of distinguishing the two possibilities, and reporting its failure to do so as evidence of absence is an error the design guaranteed.

The measured comparison: the same real effect of 0.3 standard deviations is found 14% of the time at n = 10 and 97% at n = 160. Nothing about the effect changed.

The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.
Fig. 5 The false-positive rate against the number of analyses available, on data with no effect.

The third missing number: how many chances there were

A p-value assumes one pre-specified analysis. If several were available and the best was reported, the number no longer means what it says — and it does not require any dishonesty for this to happen. Twenty correct analyses of data with no effect in it find something 57% of the time.

This is the hardest of the three to fix, because unlike sample size and effect size it is not visible in the output. A paper reporting one analysis looks the same whether it was the only one considered or the best of thirty.

20,000 p-values from a true null, n = 12. Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0090 (p = 0.81). That flatness is the check that catches an error a single rejection rate would miss.
Fig. 6 What a p-value is when the null is true: uniform, and therefore uninformative on its own.

What to report instead

The recommendation is not “abandon p-values”. It is that a p-value is a fragment of a report rather than a report.

The effect size, in units a reader can judge.

An interval, which is what the data is compatible with, and which should have had its coverage checked for the procedure and sample size in question.

The power the study had for effects worth caring about — computed in advance, which also stops the study being run if it cannot answer the question.

How many analyses were possible, and whether the reported one was pre-specified.

With those four, the p-value adds very little, which is the clearest evidence that it was never sufficient on its own.

Why it survives

Worth a paragraph, because the criticism above is a century old and the practice has not changed much.

A p-value is a single number with a conventional threshold, and that combination is administratively irresistible. It converts a judgement into a decision, it is comparable across studies that have nothing else in common, and it lets someone who has not read the work sort the results.

Everything better is worse in those specific ways: an effect size needs domain knowledge to interpret, an interval requires the reader to think about a range, and a power calculation has to be done before the data exists. All of them ask more of the reader, and the p-value’s whole appeal is that it asks nothing.

That is not an argument for keeping it. It is an argument for understanding why the correction is a hundred years old and still needs making.

What the number actually is

Worth stating carefully, because most of the misreadings are misreadings of the definition rather than of the mathematics.

A p-value is P(data at least this extreme | the null is true). Every part of that matters.

It is a probability of the data, not of the hypothesis. It says nothing directly about whether the null is true, and converting between the two requires a prior probability that the p-value does not contain — the same reversal that makes a positive test misleading for a rare condition.

It is conditional on the null being true, so it can never be evidence about how likely that is on its own.

And “at least this extreme” means it includes data that was not observed. That is a real philosophical oddity — the number depends on results that did not happen — and it is why the stopping rule matters: a value that would have been extreme under one sampling plan is not extreme under another.

The three misreadings

Each of these is common enough to appear in published work.

“p = 0.03 means there is a 3% chance the null is true.” This inverts the conditional. The actual posterior probability depends on the prior and on the power, and under reasonable assumptions a p of 0.03 corresponds to a much weaker case against the null than 3% would suggest.

“p = 0.20 means there is no effect.” It means the data is unsurprising under the null. At low power that is the expected outcome of a real effect, and the study has not distinguished the possibilities.

“p = 0.001 means the effect is large.” It means the effect is well-resolved relative to the noise, which at a large sample size can happen for an effect of no practical size at all. This is the confusion the three-sample-size figure above exists to settle.

The threshold, and where it came from

The 0.05 convention has no theoretical basis and its origin is documented: Fisher chose it as a convenient round number, remarked that it was arbitrary, and expected it to be adjusted to context.

The consequences of that arbitrary choice becoming a rule are structural rather than statistical.

It creates a cliff where the underlying quantity is continuous: 0.049 and 0.051 are essentially identical evidence and are treated as opposite results.

It creates an incentive at the cliff edge, which is the mechanism behind the lump of published p-values just below 0.05 that the replication literature documented.

And it converts a judgement into a decision, which is administratively convenient and is the reason the convention survives a century of criticism.

None of that is an argument against having thresholds where a decision genuinely has to be made. It is an argument against a single threshold, chosen for convenience, applied to every field regardless of what is at stake.

What replaces it

The alternatives are less convenient and each addresses a different part of the problem.

Report the interval. It carries the location and the precision together, and it makes the sample size visible in a way a p-value does not.

Report the effect in units someone can judge. Standard deviations are comparable across studies and meaningless to a practitioner; the natural units are the reverse. Both are worth having.

Pre-specify the analysis, which is the only measure that addresses the forking-paths inflation at its cause.

And where a decision must be made, state the threshold in advance along with the power the design provides, so that a non-significant result means something specific rather than nothing at all.

The number that would actually help

If a p-value cannot be read alone, it is fair to ask what single number could be.

The honest answer is none, and the closest candidate is worth knowing: the likelihood ratio between a stated alternative and the null. Unlike a p-value it is symmetric — it can favour the null — and it can be combined with a prior to give the posterior the reader actually wants.

Its cost is that it requires stating the alternative, which is exactly the work a p-value lets people avoid. “Is there an effect” needs no alternative; “how much more likely is an effect of 0.3 than none” needs a number that has to be defended.

That is the real trade behind the whole debate. The p-value’s popularity comes from requiring no commitment about what would count as an effect worth finding, and everything that improves on it requires exactly that commitment.

Which is a reason to make it. A study that cannot say what size of effect would matter has not specified its question, and the statistics were never going to rescue that.

The proposals to change the threshold

Several groups have proposed lowering the conventional threshold from 0.05 to 0.005, and the argument is worth understanding because it is a partial fix with a stated cost.

The case for it: at 0.05, under plausible priors and typical power, a substantial fraction of significant findings are false. Tightening the threshold reduces that fraction sharply, because the evidence required is much stronger.

The case against: it addresses only the false-positive side. A tighter threshold means lower power at any given sample size, so more real effects are missed, and among those that survive the winner’s curse inflation is worse — because the truncation point has moved further out.

So the proposal trades one error rate against the other, and whether it is an improvement depends on the relative cost of the two errors in a given field. It is not a general correction, and it does nothing at all about the multiplicity problem, which operates whatever the threshold is.

The more radical proposal — abandon thresholds and report estimates with intervals — avoids the trade entirely by declining to make a binary decision. That is right when no decision has to be made and unhelpful when one does, which is the honest summary of a long argument.

The three numbers, put together on one study

The essay has taken the missing quantities one at a time. Their interaction is the thing that actually bites, and it is best seen on a single worked case.

A study of 16 subjects looking for an effect of 0.3 standard deviations. Its power is 21%, so about one study in five reaches significance. Among the ones that do, the reported effect averages 2.07 times the truth. And if the investigators had twenty outcomes available rather than one, the chance of something reaching significance under a true null would be 57% rather than 5%.

Now read a significant result from that study. The p-value is below 0.05 and every one of those three facts bears on what it means:

  • the effect it reports is, in expectation, twice the size of anything real
  • the study had a four-in-five chance of missing a real effect entirely
  • and if the analysis was one of several available, the threshold it crossed was not the one it appears to have crossed

None of that is visible in the p-value, and none of it is a criticism of the p-value’s arithmetic. Each number is correct. The reported quantity simply does not carry the information required to interpret itself.

The familywise error rate with no correction, α = 0.05. Two routes: the curve is 1 − (1 − α)^m and the points are counted over 6,000 simulated families of true nulls. With twenty tests the chance of at least one false positive is 64.1%.
Fig. 7 How many chances there were, as a curve. The nominal level is a per-test statement and this is what a family of them produces.

What a sufficient report looks like

The constructive version, stated as a short list rather than a complaint. A result is interpretable when four things are present, and three of them are usually absent.

The effect size, with its interval. The magnitude and the uncertainty, in units a reader can reason about. This is the number the study was conducted to estimate and it is the one most often reduced to an asterisk.

The power the design had for an effect worth detecting. Not post-hoc power computed from the observed effect, which is a transformation of the p-value and adds nothing, but the power the design was built with against a pre-stated effect. This is what tells a reader whether a null result means anything and how much to discount a positive one.

How many analyses were available. Not how many were reported — how many were possible. The outcomes measured, the subgroups defined, the models considered. This is the number that converts a nominal threshold into a real one, and it is the one nobody records because recording it requires deciding in advance.

The pre-registration, or its explicit absence. Which of the above were fixed before the data were seen. A study that specified its outcome in advance and one that chose it afterwards can report identical numbers and mean entirely different things, and only the second needs the multiplicity correction.

The first is cheap and increasingly common. The second is cheap and rare. The third and fourth require a decision made before the data exist, which is why they are the ones that changed practice where they were adopted and the ones that meet the most resistance.

The p-value survives all of this, and should. It answers one narrow question correctly — how surprising this data would be if nothing were going on — and the argument here is not that it should be abolished but that a number answering one question should not be asked to stand in for four.

Why the threshold is the part that does the damage

The essay has treated the p-value as a number carrying too little information. There is a sharper version of the complaint, and it is about the threshold rather than the number.

A p-value is a continuous quantity. Nothing in its definition suggests a cut, and 0.05 has no property that 0.04 or 0.06 lacks. The cut is a convention, adopted for a table’s convenience and hardened into a decision rule.

Everything expensive in this essay follows from the cut rather than from the number. Selection on significance is what produces the inflation of published effects — with no threshold there is nothing to select on. Multiplicity matters because the minimum of twenty p-values is compared with a fixed cut; against a continuous reading there is no cliff to fall off. And the dichotomy between “significant” and “not significant” is what allows a study with 21% power to report a result rather than an estimate with an interval around it.

That suggests where the leverage is. Reporting the p-value as a continuous measure of surprise, beside an effect size and its interval, removes most of the pathology described here without changing any arithmetic. What it costs is the ability to say a study “worked”, which is a narrative convenience rather than a statistical one.

The proposals to lower the threshold rather than abandon it are worth reading in that light. Moving the cut to 0.005 reduces the false-positive rate and leaves every mechanism described here intact — selection still happens, it simply happens at a different place, and the inflation of surviving effects gets worse rather than better, because a stricter cut selects more severely.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Effect sizeFalse positivep-valuePriorSample sizeStandard deviationStatistical power