The correction that makes the estimate worse
Worth reading first: Twenty analyses of nothing · The winner's curse.
A multiple-comparison analysis has two things wrong with it. The p-value is computed as though one analysis had been run, and the effect size is the largest of twenty. The first is the false-positive inflation and the second is the winner’s curse, and they are usually discussed as the same problem with one repair.
They are not, and the repair for the first makes the second worse.
The mechanism is one sentence
A correction works by raising the bar. The estimate that clears a higher bar is larger, and an estimate is selected on the same noise it is measuring, so the estimate that clears a higher bar is further from the truth.
At an effect of two standard errors, in a family of twenty analyses correlated at 0.6:
| threshold | critical value | families that clear it | reported effect, as a multiple of the truth |
|---|---|---|---|
| none | 1.960 | 68.5% | 1.348 |
| family-wise, measured | 2.846 | 22.5% | 1.690 |
| Bonferroni | 3.023 | 17.0% | 1.762 |
Nothing about the estimator changed between the rows. What changed is which families get to report one.
The middle column is the one that makes the trade visible. Correcting takes the share of families reporting anything from 68.5% to 22.5%, which is the power cost, and it takes the reported effect from 1.348 to 1.690 times the truth, which is the estimate cost. Both costs are paid for the same purchase — an honest error rate on the null families, which are not in this table at all, because this table conditions on an effect of two standard errors being present.
That is worth holding onto, because it is why the trade is so easy to miss. The benefit of correcting is measured on families with nothing in them and the costs are measured on families with something in them, so no single simulation shows both, and a paper reporting a corrected analysis of one dataset shows neither. The error rate’s own essay counts the benefit and this page counts the costs, and they are the same procedure seen under two conditions.
The unconditional average of the largest of the twenty is 2.365 — 1.18 times the truth — and that is the inflation the selection to a maximum causes, before any threshold is applied. Every row above is that number conditioned further.
The trade is worse where the research is
The inflation depends on the effect size, and it depends on it steeply.
| effect, in standard errors | uncorrected | family-wise | families clearing the family-wise threshold |
|---|---|---|---|
| 0.5 | 4.861 | 6.366 | 5.3% |
| 1.0 | 2.464 | 3.208 | 7.1% |
| 1.5 | 1.698 | 2.184 | 12.1% |
| 2.0 | 1.348 | 1.690 | 22.5% |
| 2.5 | 1.165 | 1.405 | 38.6% |
| 3.0 | 1.072 | 1.227 | 58.3% |
| 4.0 | 1.010 | 1.054 | 89.3% |
At half a standard error the corrected analysis reports an effect six and a third times the truth, on the 5.3% of occasions it reports anything. At four standard errors it reports it 5% too large, on 89% of occasions. The correction is nearly free at the top of the table and ruinous at the bottom, and the bottom is where underpowered research lives.
So a corrected analysis of an underpowered study is not a conservative analysis. Its error rate is right and its effect size is wrong by a larger factor than the uncorrected version’s would have been, and the effect size is what gets carried into the next study’s power calculation, the meta-analysis and the abstract.
Two selections, and only one of them is the familiar one
The inflation on this page is the product of two conditionings, and separating them is what makes the size of it legible.
Taking a maximum inflates the estimate whether or not a threshold is applied. The largest of twenty statistics, with the effect sitting in one of them, averages 2.365 against a truth of 2 — a factor of 1.18 — and that figure conditions on nothing at all. Every family reports; the reported number is simply the largest one.
Applying a threshold inflates it again, conditional on the report happening. From 1.18 to 1.348 at the uncorrected threshold, and to 1.690 at the family-wise one.
The classic winner’s curse is the second selection alone: honest studies, one analysis each, filtered to the ones that reached significance. A family of analyses adds the first, and the correction strengthens the second. So an exploratory analysis carries a curse with two sources, and only one of them is discussed.
The decomposition also says which repair addresses which. Prespecifying removes the first selection entirely and leaves the second. Publishing regardless of the result removes the second and leaves the first. Neither removes both, and a literature that did both would report estimates inflated by nothing — which is a statement about how much of the problem is structural and how much is practice.
What the estimate is used for next
The reason the inflation matters more than the p-value is that the estimate is the thing carried forward, and the first thing it is carried into is the design of the next study.
A replication powered at 80% for a reported effect needs a sample proportional to the inverse square of that effect. If the reported effect is inflated by a factor f, the replication is designed with 1/f² of the sample it needed, and its actual power against the true effect is what that reduced sample buys:
| inflation of the reported effect | replication’s sample, as a share of what is needed | replication’s actual power |
|---|---|---|
| 1.18 (maximum only) | 72% | 66.1% |
| 1.348 (uncorrected) | 55% | 54.7% |
| 1.690 (family-wise) | 35% | 38.1% |
| 1.762 (Bonferroni) | 32% | 35.6% |
| 6.366 (small effect, corrected) | 2.5% | 7.2% |
A replication designed carefully for 80% power, against an effect reported from a corrected analysis of twenty, has 38% power. It will fail more often than it succeeds, and the failure will be read as evidence against the original finding when it is a consequence of the original finding’s arithmetic.
That is the mechanism by which a correction intended to make a literature more reliable makes its replications less informative, and it is entirely invisible in either study. The first study’s p-value is honest. The second study’s power calculation is arithmetically correct. The input to the second is a number the first was structurally unable to report accurately.
Why the two problems cannot share a repair
The reason is that they are failures of different objects, conditioned on different events.
The error rate is a property of the procedure over all families, including those that report nothing. Raising the threshold reduces the share of null families that report, which is the whole mechanism, and it works.
The estimate is a property of the families that report. It is conditional on the reporting event by construction — an unreported estimate is not in the average — so changing the reporting event changes the conditional distribution. Making the event rarer makes it more extreme.
Those two pull opposite ways and there is no threshold at which both are satisfied. The threshold that makes the estimate unbiased is no threshold at all, and no threshold gives a 5% error rate on a family of twenty.
That is a structural statement rather than a limitation of the corrections available. Any filter applied to honest studies inflates what passes it, and a multiple-comparison correction is a filter, applied on purpose, to a set that was already filtered by taking a maximum.
What actually repairs the estimate
Nothing on the p-value side, and three things elsewhere.
Report the analysis that was prespecified, whether or not it won. Its estimate is unconditional and therefore unbiased. That is a separate virtue of prespecification from the power trade that naming an analysis in advance costs, and it is the larger one: the prespecified estimate is right, and the exploratory one is inflated by a factor the table above puts between 1.05 and 6.4.
Report all twenty estimates. The set of twenty is unconditional; only the maximum is selected. A paper printing all twenty effect sizes has given a reader everything needed to see how extreme the reported one is within its own family, and the shape of the other nineteen says more about whether the winner is real than its p-value does.
Shrink it. An estimate selected as a maximum has a known conditional bias given the threshold it cleared, and it can be corrected — by conditioning on the selection event explicitly, or by borrowing from the rest of the family, which shrinks every estimate towards their common centre and shrinks the largest most. The second is the one that generalises, and it has the property the first lacks: it does not need the threshold to be known.
The three are not equally available and they fail differently. The first needs a prespecification that may not exist. The second needs journal space and a reader willing to use it. The third needs a population of analyses that can reasonably be treated as exchangeable, which twenty outcome measures on the same subjects usually are and twenty exclusion rules are not — the rules are not twenty draws from anything, they are one analysis perturbed, and shrinking towards the mean of twenty perturbations borrows from a population that does not exist. The first two are always correct and rarely done; the third is often the best available and needs an argument each time.
What the trade looks like in total error
Inflation is only half of an estimate’s quality, so it is worth asking whether the correction is better or worse on a measure that counts both halves.
It is worse on both, which is the part that makes this more than a curiosity, and the arithmetic is worth following because the second half is the one that is usually assumed to go the other way. Raising the threshold increases the bias of what survives — the table above — and it does not reduce the variance of what survives, because the surviving set is a tail of the same distribution rather than an average over more of it. A narrower, more extreme tail has a larger squared bias and a comparable spread, so the mean squared error of the reported estimate rises with the threshold.
The comparison that would make the correction look good on this measure is against an analysis that reports its maximum with no threshold at all, and that comparison is not available: an analysis with no threshold reports on every family, including the ones with nothing in them, so its error rate is 64% and its estimate on a null family is pure noise.
There is one measure on which the correction does win, and naming it keeps the page honest. The unconditional mean squared error of the whole procedure — counting families that report nothing as reporting the null — falls with the threshold, because null families stop reporting noise. That is the quantity a decision-theoretic account of testing optimises, and it is not the quantity a reader of a published effect size has in front of them. A reader sees the reports and not the silences, so the conditional measure is the one that describes their experience, and the two point opposite ways.
The honest statement is therefore narrow. Among analyses that control their error rate, a more severe correction produces a more inflated estimate. It says nothing about whether to correct — the answer to that is yes — and everything about what a corrected estimate should be trusted to be.
What a meta-analysis does with it
The second place the estimate is carried is into a pooled analysis, and pooling does not fix a bias that every input shares.
A meta-analysis averages the reported effects, weighted by their precision. If each reported effect is inflated by a factor between 1.3 and 1.8, the pooled estimate is inflated by a factor in that range, and the pooling has narrowed its confidence interval around the wrong value. The interval’s width falls like in the number of studies and the bias does not fall at all, so the more studies are pooled, the more confidently the wrong number is reported.
That is the ordinary criticism of publication bias, arriving here from a source nobody adjusts for. A funnel plot detects asymmetry produced by missing small studies; the inflation on this page is present in every study that appears, small and large, and it leaves no asymmetry to detect — the large studies are inflated less, which is the same direction a funnel plot reads as selective publication, so the two are confounded in exactly the diagnostic meant to separate them.
The consequence for a reader of a meta-analysis is one question. Were the pooled estimates prespecified analyses or selected ones? If the primary outcomes were named in advance, the inflation from the maximum-selection is absent and only the publication filter remains. If the included studies each reported their best of several outcomes, every input carries the factor in the table above and the pooled result carries it too, with a narrower interval to make it look more certain.
Nothing in a standard meta-analysis asks that question, and the information needed to answer it is usually in the included papers.
What is claimed here, and what is not
Two statements, and the second is what keeps the first from being trivial.
The corrected threshold’s surviving estimate is inflated by more than the uncorrected one’s, at every effect size drawn. Stated across the whole sweep rather than at one setting, because a comparison at a single effect is a fact about that effect and this is a claim about the family of them.
The inflation falls as the effect grows. A large effect does not need to be lucky, so the selection has little to select on. That is worth stating because it distinguishes the measurement from one that had confused the conditional and unconditional means — the unconditional inflation also falls with the effect, but it falls to 1.00 from 1.18, and the conditional one falls from 6.37, so the wrong one would look qualitatively similar and be an order of magnitude out.
The reading that does not survive is a family-wise-corrected estimate taken as unbiased. The standard every estimate here is held to is that it should average the thing it estimates, and at two standard errors this one averages 1.690 times it. An estimate that did average the truth would mean the conditioning was not happening — an average over all families rather than over the ones that reported, which is the slip that makes a selection effect vanish from a simulation entirely.
Still open: what an honest exploratory report looks like
The three repairs above are each partial and none is standard practice, so the question of what such a paper should actually print has no settled answer.
Printing all twenty estimates is cheap and rare. Printing the corrected p-value beside a shrunk estimate is coherent and has an awkwardness the others do not: the p-value and the estimate then disagree about how strong the finding is, because they are conditioned differently, and a reader has no convention for holding two numbers that disagree on purpose.
What is missing is a report format in which the selection is visible rather than corrected for. The elements exist — the family, the threshold, the distribution of the other nineteen, the conditional distribution of the maximum given the threshold — and nothing combines them into something a reader can take in. Whether such a format is possible at the length an abstract has, or whether the inflation simply has to be carried as a caveat, is a question about scientific writing that this measurement makes precise and does not answer.
Still open: whether a reader can hold two estimates at once
One candidate is worth naming because it costs nothing and is testable. Report the effect that would have been reported had the threshold been the uncorrected one, beside the corrected p-value. That number is available — it is the same maximum — and the pair carries the information: an honest error rate, and an estimate conditioned on the weakest filter the analysis could have used. It is still inflated, by the factor in the first column of the table rather than the second, and the difference between the two columns is precisely what a reader currently cannot see.
Whether that pairing reads as coherent or as two numbers in an argument with each other is the thing that would have to be tried, and it is the same difficulty a shrunk estimate raises beside its own interval.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An order that spends the error rate — both name bonferroni, familywise error rate, multiple comparisons, statistical power
- One control, many arms — both name bonferroni, familywise error rate, multiple comparisons, statistical power
- The price of control — both name bonferroni, effect size, multiple comparisons, statistical power
- A coverage table with its own error — both name bonferroni, multiple comparisons, statistical power
- Dropping the losers — both name familywise error rate, multiple comparisons, selection bias
- Eight forecasters and one benchmark — both name bonferroni, familywise error rate, multiple comparisons
Named objects
A flat tag is an object no other essay names yet.
BonferroniEffect sizeFamilywise error rateMean squared errorMultiple comparisonsSelection biasStatistical powerThe winner's curse