The winner's curse
Take a real effect. Run twelve thousand honest studies of it. Keep the ones that reached significance, which is what a literature does. The effects those studies report are too large — and at realistic sample sizes they are too large by a factor of two.
The measurement
The true effect is 0.3 standard deviations. At n = 16 the power is around 20%, so about one study in five reaches significance.
Those studies report a mean effect of about 0.62 — roughly 2.05 times the truth.
At n = 250 the power is high, nearly every study reaches significance, and the inflation disappears: the surviving studies report 1.00 times the truth.
The inflation is therefore not a property of the effect or of the field. It is a property of power, and it is severe exactly where power is low.
Why selection alone does it
No study here is fraudulent, p-hacked or badly run. Each draws its sample honestly and computes correctly.
The mechanism is that a study reaches significance when its estimate is far enough from zero. When power is low, the only way to clear that bar is to get an unusually large estimate — a sample that happened to overstate the effect. Studies that estimated the effect accurately did not clear the bar and are not in the surviving set.
So the filter is not neutral with respect to the quantity being filtered. It selects on the estimate, and therefore selects upward.
This is the same mechanism as regression to the mean, which is worth noticing: selecting on a noisy measurement and then looking at the underlying quantity gives a biased picture, in a direction determined entirely by which end was selected.
What it predicts about replication
The prediction is specific and it has been tested at scale.
If a literature consists largely of underpowered significant findings, then the published effects overstate the truth. A replication with adequate power should therefore find smaller effects, systematically, and should fail to reach significance often even when the original effect is real.
That is what large replication projects have found: replication effect sizes consistently smaller than originals, and a substantial fraction of replications not reaching significance. Some of that is fraud and some is p-hacking, and this figure shows that none of it needs to be. Honest underpowered research produces exactly this pattern on its own.
What follows for reading a result
Three consequences, in order of how often they are ignored.
A significant result from a small study should be discounted, not celebrated. The smaller the study, the more its significant findings are inflated. The intuition that a significant result from a small sample is impressive — because it “overcame” the small n — is exactly wrong: it cleared the bar by being lucky.
A large effect from a small study is the most suspect combination of all. It is the signature the selection produces.
The first published estimate of anything is likely to be the largest. Not because of misconduct, but because it was the one that got through the filter first.
What fixes it
Power the study properly. Everything above is a function of power. At adequate power the inflation vanishes because the filter stops selecting.
Report the estimate and its interval regardless of significance, and publish null results. The inflation comes entirely from conditioning on significance; remove the condition and the bias goes with it.
Pre-register, which removes the additional inflation from having many analyses available — a separate mechanism that stacks on top of this one.
And treat a meta-analysis of underpowered studies with care. Averaging biased estimates does not remove the bias. It produces a very precise estimate of the wrong number, with a narrow interval that excludes the truth — which is a worse failure than any individual study, because it looks so much more authoritative.
The arithmetic of the inflation
The size of the effect is predictable before any simulation, which is worth showing because it makes the phenomenon a calculation rather than a warning.
A study reaches significance when its estimate exceeds roughly 1.96 standard errors. The published estimate is therefore the mean of a normal distribution truncated below at that point, and the mean of a truncated normal has a closed form: it is the untruncated mean plus the standard deviation times the inverse Mills ratio at the truncation point.
When power is high the truncation point sits far below the distribution’s centre, so almost nothing is cut and the mean barely moves. When power is low the truncation point sits above the centre, only the upper tail survives, and the surviving mean is pulled a long way up.
That is why the inflation is a function of power alone and not of the field, the effect, or anything about how the research was conducted. Power is the fraction surviving, and the fraction surviving determines how far up the survivors’ mean sits.
The sign of the estimate
A detail that follows and is worse than the inflation itself.
At very low power, some of the studies reaching significance do so in the wrong direction — their estimate is far enough below zero to be significant, even though the true effect is positive. That rate has a name, the Type S error rate, and at 10% power it is not negligible.
So an underpowered literature does not merely overstate effects. It contains a minority of published, significant, honestly-conducted findings pointing the opposite way to the truth, and nothing distinguishes them from the rest except that they disagree.
That is the strongest argument against reading any individual small significant study, and it is a consequence of the same truncation.
Why meta-analysis does not rescue it
The instinct on hearing all this is to pool: many small studies, averaged, should approximate the truth.
They do not, because the bias is in each study’s inclusion, not in its measurement. Averaging unbiased estimates removes noise. Averaging estimates that were selected for being large removes noise and leaves the selection intact — producing a very precise estimate of the wrong number, with a narrow interval that excludes the truth.
That is worse than any individual study, because the precision makes it look authoritative. A meta-analysis of published underpowered work is a meta-analysis of a filter.
Methods exist to correct for it — funnel plots, trim-and-fill, selection models, p-curve. All of them attempt to model the filter and undo it, all of them require assumptions about the filter’s shape, and none of them recovers what an unfiltered literature would have shown. The reliable fix is upstream: publish the null results, and the filter never operates.
What a reader can do with a single paper
Since most readers meet this literature one paper at a time, the practical version.
Find the sample size and the effect size. Together they give the standard error, and the standard error against the reported effect gives an idea of power. A large effect with a wide interval is the signature of a study that cleared the bar by luck.
Ask whether the effect is plausible. An implausibly large effect from a small study is much more likely to be a lucky draw than a discovery, and the prior does real work here even when nobody writes it down.
Wait for the replication, and expect it to be smaller. Not because the original was dishonest, but because the mechanism above operates on every literature that publishes conditional on significance.
What an honest literature would look like
Worth sketching, because the corrective measures are easier to justify against a picture of the alternative.
If every study were published regardless of outcome, the distribution of reported effects would be centred on the truth. Small studies would scatter widely and large ones narrowly, and a meta-analysis would recover the right answer by weighting them.
The scatter would look alarming. Many studies would report effects near zero or the wrong sign, and readers accustomed to a filtered literature would read that as a field in disarray rather than as a field behaving correctly.
That is worth anticipating: removing the filter makes the literature look worse and be better, and the appearance is a large part of why the filter persists.
The registered report
The structural fix that addresses both this and the forking paths at once.
The design and analysis are reviewed and accepted before the data is collected. Publication is guaranteed on the basis of the question and the method, not the result.
That removes the significance filter entirely — nothing is conditioned on — and it removes the analysis flexibility, because the analysis was specified in advance. Both inflation mechanisms are cut off at their cause rather than corrected for afterwards.
The empirical result, where it has been tried, is what the mechanism predicts: the proportion of published positive findings drops sharply, and the effects that are reported are smaller. That is the correct outcome and it reads as a decline in productivity, which is the political difficulty rather than the statistical one.
Estimating the inflation from a published result
Since most readers meet this after the fact, it is worth knowing that the correction is partially computable from what a paper reports.
Given the reported effect and its standard error, the truncation point is known — it is the significance threshold — so the relationship between the observed estimate and the underlying effect is the truncated-normal relation stated above. Inverting it gives a shrunk estimate.
That is the idea behind several published correction methods, and they share a limitation worth being honest about: the correction depends on the prior distribution of true effects in the field, which nobody knows. Different assumptions give different shrinkage.
What is robust is the direction and the rough magnitude. If the study was underpowered for the effect it reports, the effect is overstated, and the shrinkage is substantial rather than marginal. That is enough to change how a single result is read, which is most of what a reader needs.
The other side: what a large study buys
The essay is about small studies, and the complement is worth stating because it is the practical recommendation.
At n = 250 in the simulation the inflation is 1.00 — no bias at all. Not reduced, gone. Once power is high enough that almost every study of a real effect reaches significance, there is no filter, because nothing is being filtered out.
So adequate power does not merely increase the chance of detecting an effect. It removes the bias in the effects that are detected, removes the wrong-sign errors, and makes the published literature an unbiased sample of the studies conducted.
That is a much stronger argument for powering studies properly than the usual one about avoiding wasted effort, and it is the argument this figure is for: power is not only about finding things, it is about the findings being right.