The winner's curse
Worth reading first: What a p-value does not say.
Take a real effect. Run twelve thousand honest studies of it. Keep the ones that reached significance, which is what a literature does. The effects those studies report are too large — and at realistic sample sizes they are too large by a factor of two.
The measurement
The true effect is 0.3 standard deviations. At n = 16 the power is around 20%, so about one study in five reaches significance.
Those studies report a mean effect of about 0.62 — roughly 2.07 times the truth.
At n = 250 the power is high, nearly every study reaches significance, and the inflation disappears: the surviving studies report 1.00 times the truth.
The second picture is the same simulation with one number changed, and it is worth dwelling on how little of the machinery moved. The true effect is 0.3 in both. The threshold is 0.05 in both. Nothing about how a study is run, analysed or reported differs between them. The only difference is that the second set of studies is large enough that a study reaching significance is unremarkable, so surviving the filter says almost nothing about the estimate — and an uninformative filter cannot bias what passes through it.
That is the whole mechanism, visible as a picture rather than as an argument: the inflation is the selectivity of the filter, and a filter that admits everything is not selective. It follows that the inflation is not a property of the effect or of the field. It is a property of power, and it is severe exactly where power is low.
Why selection alone does it
No study here is fraudulent, p-hacked or badly run. Each draws its sample honestly and computes correctly.
The mechanism is that a study reaches significance when its estimate is far enough from zero. When power is low, the only way to clear that bar is to get an unusually large estimate — a sample that happened to overstate the effect. Studies that estimated the effect accurately did not clear the bar and are not in the surviving set.
So the filter is not neutral with respect to the quantity being filtered. It selects on the estimate, and therefore selects upward.
This is the same mechanism as regression to the mean, which is worth noticing: selecting on a noisy measurement and then looking at the underlying quantity gives a biased picture, in a direction determined entirely by which end was selected.
The inflation as a function of power, in closed form
The measurement at 20% power is one point on a curve that can be written down, and having the curve is what turns the finding into something a reader can apply to a paper in front of them.
An estimate that survives a two-sided threshold at , when the true effect is standard errors from zero, has conditional mean
so the inflation factor is .
At the field’s own setting — a true effect of 0.3 with a standard error of 0.25, so — that gives a conditional mean of 0.63 against the 0.62 measured across twelve thousand studies.
Run across the range of power a reader will meet:
- 20% power: inflation 2.07×
- 50%: 1.41×
- 80%: 1.13×
- 90%: 1.06×
- 95%: 1.03×
Which puts a number on the conventional target
The row worth stopping at is the third. A study designed to the conventional 80% power still reports, on average, an effect thirteen per cent larger than the truth — conditional on having been significant, which is the condition under which it gets published.
Thirteen per cent is small beside the doubling at 20% power and it is not nothing. It is comparable to the width of the confidence interval such a study reports, so a reader taking the point estimate at face value is being misled by about half an interval, systematically, in the same direction, in every adequately powered significant study they read.
Getting the inflation under five per cent takes about 92% power, which is a sample size roughly a third larger again than the 80% target.
So there is no power at which the curse switches off, and the received framing — that it is a problem of underpowered research — is a statement about where it is severe rather than about where it exists. At n = 250 the inflation reads 1.00× because the power there is so close to one that the filter admits essentially every study, and a filter that removes nothing cannot select on anything.
What it predicts about replication
The prediction is specific and it has been tested at scale.
If a literature consists largely of underpowered significant findings, then the published effects overstate the truth. A replication with adequate power should therefore find smaller effects, systematically, and should fail to reach significance often even when the original effect is real.
That is what large replication projects have found: replication effect sizes consistently smaller than originals, and a substantial fraction of replications not reaching significance. Some of that is fraud and some is p-hacking, and this figure shows that none of it needs to be. Honest underpowered research produces exactly this pattern on its own.
What follows for reading a result
Three consequences, in order of how often they are ignored.
A significant result from a small study should be discounted, not celebrated. The smaller the study, the more its significant findings are inflated. The intuition that a significant result from a small sample is impressive — because it “overcame” the small n — is exactly wrong: it cleared the bar by being lucky.
A large effect from a small study is the most suspect combination of all. It is the signature the selection produces.
The first published estimate of anything is likely to be the largest. Not because of misconduct, but because it was the one that got through the filter first.
What fixes it
Power the study properly. Everything above is a function of power. At adequate power the inflation vanishes because the filter stops selecting.
Report the estimate and its interval regardless of significance, and publish null results. The inflation comes entirely from conditioning on significance; remove the condition and the bias goes with it.
Pre-register, which removes the additional inflation from having many analyses available — a separate mechanism that stacks on top of this one.
And treat a meta-analysis of underpowered studies with care. Averaging biased estimates does not remove the bias. It produces a very precise estimate of the wrong number, with a narrow interval that excludes the truth — which is a worse failure than any individual study, because it looks so much more authoritative.
The arithmetic of the inflation
The size of the effect is predictable before any simulation, which is worth showing because it makes the phenomenon a calculation rather than a warning.
A study reaches significance when its estimate exceeds roughly 1.96 standard errors. The published estimate is therefore the mean of a normal distribution truncated below at that point, and the mean of a truncated normal has a closed form: it is the untruncated mean plus the standard deviation times the inverse Mills ratio at the truncation point.
When power is high the truncation point sits far below the distribution’s centre, so almost nothing is cut and the mean barely moves. When power is low the truncation point sits above the centre, only the upper tail survives, and the surviving mean is pulled a long way up.
That is why the inflation is a function of power alone and not of the field, the effect, or anything about how the research was conducted. Power is the fraction surviving, and the fraction surviving determines how far up the survivors’ mean sits.
The sign of the estimate
A detail that follows and is worse than the inflation itself.
At very low power, some of the studies reaching significance do so in the wrong direction — their estimate is far enough below zero to be significant, even though the true effect is positive. That rate has a name, the Type S error rate, and at 10% power it is not negligible.
So an underpowered literature does not merely overstate effects. It contains a minority of published, significant, honestly-conducted findings pointing the opposite way to the truth, and nothing distinguishes them from the rest except that they disagree.
That is the strongest argument against reading any individual small significant study, and it is a consequence of the same truncation.
Why meta-analysis does not rescue it
The instinct on hearing all this is to pool: many small studies, averaged, should approximate the truth.
They do not, because the bias is in each study’s inclusion, not in its measurement. Averaging unbiased estimates removes noise. Averaging estimates that were selected for being large removes noise and leaves the selection intact — producing a very precise estimate of the wrong number, with a narrow interval that excludes the truth.
That is worse than any individual study, because the precision makes it look authoritative. A meta-analysis of published underpowered work is a meta-analysis of a filter.
Methods exist to correct for it — funnel plots, trim-and-fill, selection models, p-curve. All of them attempt to model the filter and undo it, all of them require assumptions about the filter’s shape, and none of them recovers what an unfiltered literature would have shown. The reliable fix is upstream: publish the null results, and the filter never operates.
What a reader can do with a single paper
Since most readers meet this literature one paper at a time, the practical version.
Find the sample size and the effect size. Together they give the standard error, and the standard error against the reported effect gives an idea of power. A large effect with a wide interval is the signature of a study that cleared the bar by luck.
Ask whether the effect is plausible. An implausibly large effect from a small study is much more likely to be a lucky draw than a discovery, and the prior does real work here even when nobody writes it down.
Wait for the replication, and expect it to be smaller. Not because the original was dishonest, but because the mechanism above operates on every literature that publishes conditional on significance.
What an honest literature would look like
Worth sketching, because the corrective measures are easier to justify against a picture of the alternative.
If every study were published regardless of outcome, the distribution of reported effects would be centred on the truth. Small studies would scatter widely and large ones narrowly, and a meta-analysis would recover the right answer by weighting them.
The scatter would look alarming. Many studies would report effects near zero or the wrong sign, and readers accustomed to a filtered literature would read that as a field in disarray rather than as a field behaving correctly.
That is worth anticipating: removing the filter makes the literature look worse and be better, and the appearance is a large part of why the filter persists.
The registered report
The structural fix that addresses both this and the forking paths at once.
The design and analysis are reviewed and accepted before the data is collected. Publication is guaranteed on the basis of the question and the method, not the result.
That removes the significance filter entirely — nothing is conditioned on — and it removes the analysis flexibility, because the analysis was specified in advance. Both inflation mechanisms are cut off at their cause rather than corrected for afterwards.
The empirical result, where it has been tried, is what the mechanism predicts: the proportion of published positive findings drops sharply, and the effects that are reported are smaller. That is the correct outcome and it reads as a decline in productivity, which is the political difficulty rather than the statistical one.
Estimating the inflation from a published result
Since most readers meet this after the fact, it is worth knowing that the correction is partially computable from what a paper reports.
Given the reported effect and its standard error, the truncation point is known — it is the significance threshold — so the relationship between the observed estimate and the underlying effect is the truncated-normal relation stated above. Inverting it gives a shrunk estimate.
That is the idea behind several published correction methods, and they share a limitation worth being honest about: the correction depends on the prior distribution of true effects in the field, which nobody knows. Different assumptions give different shrinkage.
What is robust is the direction and the rough magnitude. If the study was underpowered for the effect it reports, the effect is overstated, and the shrinkage is substantial rather than marginal. That is enough to change how a single result is read, which is most of what a reader needs.
The other side: what a large study buys
The essay is about small studies, and the complement is worth stating because it is the practical recommendation.
At n = 250 in the simulation the inflation is 1.00 — no bias at all. Not reduced, gone. Once power is high enough that almost every study of a real effect reaches significance, there is no filter, because nothing is being filtered out.
So adequate power does not merely increase the chance of detecting an effect. It removes the bias in the effects that are detected, removes the wrong-sign errors, and makes the published literature an unbiased sample of the studies conducted.
That is a much stronger argument for powering studies properly than the usual one about avoiding wasted effort, and it is the argument this figure is for: power is not only about finding things, it is about the findings being right.
The inflation as a function of power
The essay quotes one sample size. The relationship across sample sizes is the useful object, because it converts the phenomenon from a warning into a correction factor.
Twelve thousand simulated studies of a real effect of 0.3 standard deviations, at each of six sample sizes, keeping only those that reached significance:
| n | power | inflation | wrong sign |
|---|---|---|---|
| 8 | 12% | 2.57× | 42 of 1,451 |
| 16 | 21% | 2.07× | 10 of 2,499 |
| 30 | 36% | 1.62× | none |
| 60 | 63% | 1.25× | none |
| 120 | 90% | 1.06× | none |
| 250 | 100% | 1.00× | none |
The inflation is a function of power and nothing else. At 12% power the published effect is two and a half times the truth; at 90% power it is within six per cent of it; at full power there is no inflation, because selection on significance is not selecting anything when everything is significant.
That is the whole mechanism in one column. The curse is not a property of small samples as such — it is a property of selecting on a threshold that most studies fail, and sample size enters only by determining what fraction fail.
The sign errors, and why they are the worse failure
The last column is the one that deserves its own attention, and it is usually left out of discussions of this effect.
At n = 8, forty-two of the 1,451 significant studies — about one in thirty-five — report an effect in the wrong direction. Not merely overstated: opposite. The true effect is positive, the study is honest, the analysis is correct, and the published conclusion says the effect is negative and significant.
By n = 30 that column is empty and stays empty. So the sign error, like the inflation, is governed by power, and it disappears well before the inflation does.
The reason it matters more than the inflation is what a reader can do about it. An overstated effect is still an effect, and a reader who discounts published effects from underpowered literatures will be roughly right. A sign error is not correctable by discounting — it points the wrong way, and any amount of shrinking toward zero leaves it pointing the wrong way.
This also sharpens what a replication failure means. A replication that finds a smaller effect is consistent with the original being an inflated estimate of something real. A replication that finds an effect of the opposite sign, in a field where the original studies had power around 10%, is consistent with everything being exactly as the theory here predicts and nothing being wrong anywhere.
The uncomfortable summary is that at low power, a literature of honest, correct, statistically significant studies will contain overstated effects as a matter of routine and reversed effects at a rate of a few per cent — and it will look, from the inside, like a literature of replicated findings.
What this implies for reading a single result
The table gives a reader something concrete to do with a published estimate, provided one number is available.
If the study reports its power — or enough detail to compute it — the inflation factor can be read off directly and the estimate discounted by it. A significant result from a study with 20% power is, on average, twice the truth, and treating it as such is better calibrated than either believing it or dismissing it.
If the study does not report its power, and most do not, the sample size and the effect being looked for are usually enough to reconstruct it approximately. That reconstruction is the single most informative thing a reader can do with a paper in an underpowered field, and it takes about a minute.
The one thing that does not work is reading the confidence interval as though it accounted for this. It does not: the interval is correct conditional on the study having been run, and the selection happened afterwards.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Not half and half — both name sample size, standard deviation, statistical power
- A boundary for giving up — both name sample size, statistical power
- A coverage table with its own error — both name sample size, statistical power
- A tenth as wide, and both of them right — both name sample size, standard deviation
- Choosing n after looking — both name sample size, statistical power
- How slow a return a sample can see — both name statistical power, the winner's curse
Named objects
A flat tag is an object no other essay names yet.
Publication biasReplicationSample sizeStandard deviationStatistical powerThe winner's curse