Tests, and the second number

The winner's curse

Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.

Take a real effect. Run twelve thousand honest studies of it. Keep the ones that reached significance, which is what a literature does. The effects those studies report are too large — and at realistic sample sizes they are too large by a factor of two.

12,000 studies of a real effect of 0.3, n = 16Power is 20%. The studies that reached significance report a mean effect of 0.613 — 2.04 times the truth. Every one of them is honest; the selection did the inflating.02505007501e+3-0.50000.5001effect the study reportsstudiesthe truth, 0.3what gets published, 0.61shaded: reached p below 0.05power 20%, inflation 2.04×
Fig. 1 Twelve thousand simulated studies of an effect that really is 0.3. The shaded ones reached p < 0.05. Their average is not 0.3.

The measurement

The true effect is 0.3 standard deviations. At n = 16 the power is around 20%, so about one study in five reaches significance.

Those studies report a mean effect of about 0.62 — roughly 2.05 times the truth.

At n = 250 the power is high, nearly every study reaches significance, and the inflation disappears: the surviving studies report 1.00 times the truth.

The inflation is therefore not a property of the effect or of the field. It is a property of power, and it is severe exactly where power is low.

What a one-sample t test at n = 20 can detectAt an effect of 0.5 standard deviations the test finds it 56% of the time. Below that, a non-significant result is the expected outcome of a real effect — which is why "no significant difference" is not evidence of no difference.00.2500.5000.750100.5001true effect, in standard deviationsprobability of rejecting80%, the usual target56% at d = 0.54,000 simulated studies per pointthe second number a p-value needs
Fig. 2 The power behind the filter. Everything on this page is a function of this curve.

Why selection alone does it

No study here is fraudulent, p-hacked or badly run. Each draws its sample honestly and computes correctly.

The mechanism is that a study reaches significance when its estimate is far enough from zero. When power is low, the only way to clear that bar is to get an unusually large estimate — a sample that happened to overstate the effect. Studies that estimated the effect accurately did not clear the bar and are not in the surviving set.

So the filter is not neutral with respect to the quantity being filtered. It selects on the estimate, and therefore selects upward.

This is the same mechanism as regression to the mean, which is worth noticing: selecting on a noisy measurement and then looking at the underlying quantity gives a biased picture, in a direction determined entirely by which end was selected.

The false-positive rate against the number of analyses, on pure noiseThe data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 57% of the time.00.2000.4000.6000.8005101520analyses available to the researcherchance of finding something significantthe nominal 5%if the analyses were independentcorrelated, as real ones areno effect present anywhereevery individual analysis is correct
Fig. 3 A second inflation mechanism that compounds with this one.

What it predicts about replication

The prediction is specific and it has been tested at scale.

If a literature consists largely of underpowered significant findings, then the published effects overstate the truth. A replication with adequate power should therefore find smaller effects, systematically, and should fail to reach significance often even when the original effect is real.

That is what large replication projects have found: replication effect sizes consistently smaller than originals, and a substantial fraction of replications not reaching significance. Some of that is fraud and some is p-hacking, and this figure shows that none of it needs to be. Honest underpowered research produces exactly this pattern on its own.

Two measurements of the same thing, correlated 0.60Pick the worst 15% on the first measurement and their average rises by 0.55 on the second. Pick the best and theirs falls by 0.59. No treatment was given to anybody.first measurementsecond+0.55-0.59predicted 0.64 from the correlation alonenobody was treated
Fig. 4 The same selection effect on individuals rather than studies.

What follows for reading a result

Three consequences, in order of how often they are ignored.

A significant result from a small study should be discounted, not celebrated. The smaller the study, the more its significant findings are inflated. The intuition that a significant result from a small sample is impressive — because it “overcame” the small n — is exactly wrong: it cleared the bar by being lucky.

A large effect from a small study is the most suspect combination of all. It is the signature the selection produces.

The first published estimate of anything is likely to be the largest. Not because of misconduct, but because it was the one that got through the filter first.

20,000 p-values from a true null, n = 12Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0165 (p = 0.13). That flatness is the check that catches an error a single rejection rate would miss.05001e+30.2000.4000.6000.800p-valuecountflatKS distance 0.0165a p-value that is not uniform is not a p-value
Fig. 5 The unconditioned distribution, before any filter is applied.
The effect behind a p-value of 0.04, at three sample sizesA p of 0.04 at ten observations needs an effect of 0.76 standard deviations; at two thousand it needs 0.046. The p-value alone does not say which of these happened, which is why it should never be reported alone.ten observations0.758 sdp = 0.04a hundred0.208 sdp = 0.04two thousand0.046 sdp = 0.04the effect that gives the same p-valueall three reach p = 0.04the number does not say how big the effect is
Fig. 6 And why the surviving p-value cannot be read without its sample size.

What fixes it

Power the study properly. Everything above is a function of power. At adequate power the inflation vanishes because the filter stops selecting.

Report the estimate and its interval regardless of significance, and publish null results. The inflation comes entirely from conditioning on significance; remove the condition and the bias goes with it.

Pre-register, which removes the additional inflation from having many analyses available — a separate mechanism that stacks on top of this one.

And treat a meta-analysis of underpowered studies with care. Averaging biased estimates does not remove the bias. It produces a very precise estimate of the wrong number, with a narrow interval that excludes the truth — which is a worse failure than any individual study, because it looks so much more authoritative.

The arithmetic of the inflation

The size of the effect is predictable before any simulation, which is worth showing because it makes the phenomenon a calculation rather than a warning.

A study reaches significance when its estimate exceeds roughly 1.96 standard errors. The published estimate is therefore the mean of a normal distribution truncated below at that point, and the mean of a truncated normal has a closed form: it is the untruncated mean plus the standard deviation times the inverse Mills ratio at the truncation point.

When power is high the truncation point sits far below the distribution’s centre, so almost nothing is cut and the mean barely moves. When power is low the truncation point sits above the centre, only the upper tail survives, and the surviving mean is pulled a long way up.

That is why the inflation is a function of power alone and not of the field, the effect, or anything about how the research was conducted. Power is the fraction surviving, and the fraction surviving determines how far up the survivors’ mean sits.

The sign of the estimate

A detail that follows and is worse than the inflation itself.

At very low power, some of the studies reaching significance do so in the wrong direction — their estimate is far enough below zero to be significant, even though the true effect is positive. That rate has a name, the Type S error rate, and at 10% power it is not negligible.

So an underpowered literature does not merely overstate effects. It contains a minority of published, significant, honestly-conducted findings pointing the opposite way to the truth, and nothing distinguishes them from the rest except that they disagree.

That is the strongest argument against reading any individual small significant study, and it is a consequence of the same truncation.

Why meta-analysis does not rescue it

The instinct on hearing all this is to pool: many small studies, averaged, should approximate the truth.

They do not, because the bias is in each study’s inclusion, not in its measurement. Averaging unbiased estimates removes noise. Averaging estimates that were selected for being large removes noise and leaves the selection intact — producing a very precise estimate of the wrong number, with a narrow interval that excludes the truth.

That is worse than any individual study, because the precision makes it look authoritative. A meta-analysis of published underpowered work is a meta-analysis of a filter.

Methods exist to correct for it — funnel plots, trim-and-fill, selection models, p-curve. All of them attempt to model the filter and undo it, all of them require assumptions about the filter’s shape, and none of them recovers what an unfiltered literature would have shown. The reliable fix is upstream: publish the null results, and the filter never operates.

What a reader can do with a single paper

Since most readers meet this literature one paper at a time, the practical version.

Find the sample size and the effect size. Together they give the standard error, and the standard error against the reported effect gives an idea of power. A large effect with a wide interval is the signature of a study that cleared the bar by luck.

Ask whether the effect is plausible. An implausibly large effect from a small study is much more likely to be a lucky draw than a discovery, and the prior does real work here even when nobody writes it down.

Wait for the replication, and expect it to be smaller. Not because the original was dishonest, but because the mechanism above operates on every literature that publishes conditional on significance.

What an honest literature would look like

Worth sketching, because the corrective measures are easier to justify against a picture of the alternative.

If every study were published regardless of outcome, the distribution of reported effects would be centred on the truth. Small studies would scatter widely and large ones narrowly, and a meta-analysis would recover the right answer by weighting them.

The scatter would look alarming. Many studies would report effects near zero or the wrong sign, and readers accustomed to a filtered literature would read that as a field in disarray rather than as a field behaving correctly.

That is worth anticipating: removing the filter makes the literature look worse and be better, and the appearance is a large part of why the filter persists.

The registered report

The structural fix that addresses both this and the forking paths at once.

The design and analysis are reviewed and accepted before the data is collected. Publication is guaranteed on the basis of the question and the method, not the result.

That removes the significance filter entirely — nothing is conditioned on — and it removes the analysis flexibility, because the analysis was specified in advance. Both inflation mechanisms are cut off at their cause rather than corrected for afterwards.

The empirical result, where it has been tried, is what the mechanism predicts: the proportion of published positive findings drops sharply, and the effects that are reported are smaller. That is the correct outcome and it reads as a decline in productivity, which is the political difficulty rather than the statistical one.

Estimating the inflation from a published result

Since most readers meet this after the fact, it is worth knowing that the correction is partially computable from what a paper reports.

Given the reported effect and its standard error, the truncation point is known — it is the significance threshold — so the relationship between the observed estimate and the underlying effect is the truncated-normal relation stated above. Inverting it gives a shrunk estimate.

That is the idea behind several published correction methods, and they share a limitation worth being honest about: the correction depends on the prior distribution of true effects in the field, which nobody knows. Different assumptions give different shrinkage.

What is robust is the direction and the rough magnitude. If the study was underpowered for the effect it reports, the effect is overstated, and the shrinkage is substantial rather than marginal. That is enough to change how a single result is read, which is most of what a reader needs.

The other side: what a large study buys

The essay is about small studies, and the complement is worth stating because it is the practical recommendation.

At n = 250 in the simulation the inflation is 1.00 — no bias at all. Not reduced, gone. Once power is high enough that almost every study of a real effect reaches significance, there is no filter, because nothing is being filtered out.

So adequate power does not merely increase the chance of detecting an effect. It removes the bias in the effects that are detected, removes the wrong-sign errors, and makes the published literature an unbiased sample of the studies conducted.

That is a much stronger argument for powering studies properly than the usual one about avoiding wasted effort, and it is the argument this figure is for: power is not only about finding things, it is about the findings being right.