The analyses that were available and not run

The correction that makes the estimate worse

Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.

Worth reading first: Twenty analyses of nothing · The winner's curse.

A multiple-comparison analysis has two things wrong with it. The p-value is computed as though one analysis had been run, and the effect size is the largest of twenty. The first is the false-positive inflation and the second is the winner’s curse, and they are usually discussed as the same problem with one repair.

They are not, and the repair for the first makes the second worse.

What the family-wise correction does to the effect it lets throughAt two standard errors the estimate that clears an uncorrected 5% threshold averages 1.35 times the truth, and the one that clears the family-wise threshold averages 1.69 times it. The correction fixes the error rate by demanding a larger estimate, and a larger estimate is a more selected one.2461234the effect, in standard errorsreported effect, as a multiple of the truthafter the correctionbefore itthe truth40,000 families at each effectthe p-value is repaired and the estimate is not
Fig. 1 The reported effect as a multiple of the truth, at the uncorrected threshold and at the family-wise one. The correction moves the upper curve, and it moves it upwards.

The mechanism is one sentence

A correction works by raising the bar. The estimate that clears a higher bar is larger, and an estimate is selected on the same noise it is measuring, so the estimate that clears a higher bar is further from the truth.

At an effect of two standard errors, in a family of twenty analyses correlated at 0.6:

threshold critical value families that clear it reported effect, as a multiple of the truth
none 1.960 68.5% 1.348
family-wise, measured 2.846 22.5% 1.690
Bonferroni 3.023 17.0% 1.762

Nothing about the estimator changed between the rows. What changed is which families get to report one.

The middle column is the one that makes the trade visible. Correcting takes the share of families reporting anything from 68.5% to 22.5%, which is the power cost, and it takes the reported effect from 1.348 to 1.690 times the truth, which is the estimate cost. Both costs are paid for the same purchase — an honest error rate on the null families, which are not in this table at all, because this table conditions on an effect of two standard errors being present.

That is worth holding onto, because it is why the trade is so easy to miss. The benefit of correcting is measured on families with nothing in them and the costs are measured on families with something in them, so no single simulation shows both, and a paper reporting a corrected analysis of one dataset shows neither. The error rate’s own essay counts the benefit and this page counts the costs, and they are the same procedure seen under two conditions.

The unconditional average of the largest of the twenty is 2.365 — 1.18 times the truth — and that is the inflation the selection to a maximum causes, before any threshold is applied. Every row above is that number conditioned further.

Three thresholds on one family, at an effect of 2 standard errors. Each bar is the reported effect as a multiple of the truth; each note is the share of families that clear that threshold at all. The uncorrected threshold lets 68.5% through at 1.35 times the truth; Bonferroni lets 17.0% through at 1.76 times it.
Fig. 2 The same three thresholds with the share of families that clear each beside the inflation of what clears it. More correction, fewer reports, and each one further from the truth.

The trade is worse where the research is

The inflation depends on the effect size, and it depends on it steeply.

effect, in standard errors uncorrected family-wise families clearing the family-wise threshold
0.5 4.861 6.366 5.3%
1.0 2.464 3.208 7.1%
1.5 1.698 2.184 12.1%
2.0 1.348 1.690 22.5%
2.5 1.165 1.405 38.6%
3.0 1.072 1.227 58.3%
4.0 1.010 1.054 89.3%

At half a standard error the corrected analysis reports an effect six and a third times the truth, on the 5.3% of occasions it reports anything. At four standard errors it reports it 5% too large, on 89% of occasions. The correction is nearly free at the top of the table and ruinous at the bottom, and the bottom is where underpowered research lives.

So a corrected analysis of an underpowered study is not a conservative analysis. Its error rate is right and its effect size is wrong by a larger factor than the uncorrected version’s would have been, and the effect size is what gets carried into the next study’s power calculation, the meta-analysis and the abstract.

Three thresholds on one family, at an effect of 0.5 standard errors. Each bar is the reported effect as a multiple of the truth; each note is the share of families that clear that threshold at all. The uncorrected threshold lets 36.8% through at 4.86 times the truth; Bonferroni lets 3.3% through at 6.68 times it.
Fig. 3 The bottom row drawn out. At half a standard error the three thresholds report 4.86, 6.37 and 6.68 times the truth, and the most careful of the three is the most wrong.

Two selections, and only one of them is the familiar one

The inflation on this page is the product of two conditionings, and separating them is what makes the size of it legible.

Taking a maximum inflates the estimate whether or not a threshold is applied. The largest of twenty statistics, with the effect sitting in one of them, averages 2.365 against a truth of 2 — a factor of 1.18 — and that figure conditions on nothing at all. Every family reports; the reported number is simply the largest one.

Applying a threshold inflates it again, conditional on the report happening. From 1.18 to 1.348 at the uncorrected threshold, and to 1.690 at the family-wise one.

The classic winner’s curse is the second selection alone: honest studies, one analysis each, filtered to the ones that reached significance. A family of analyses adds the first, and the correction strengthens the second. So an exploratory analysis carries a curse with two sources, and only one of them is discussed.

The decomposition also says which repair addresses which. Prespecifying removes the first selection entirely and leaves the second. Publishing regardless of the result removes the second and leaves the first. Neither removes both, and a literature that did both would report estimates inflated by nothing — which is a statement about how much of the problem is structural and how much is practice.

What the estimate is used for next

The reason the inflation matters more than the p-value is that the estimate is the thing carried forward, and the first thing it is carried into is the design of the next study.

A replication powered at 80% for a reported effect needs a sample proportional to the inverse square of that effect. If the reported effect is inflated by a factor f, the replication is designed with 1/f² of the sample it needed, and its actual power against the true effect is what that reduced sample buys:

inflation of the reported effect replication’s sample, as a share of what is needed replication’s actual power
1.18 (maximum only) 72% 66.1%
1.348 (uncorrected) 55% 54.7%
1.690 (family-wise) 35% 38.1%
1.762 (Bonferroni) 32% 35.6%
6.366 (small effect, corrected) 2.5% 7.2%

A replication designed carefully for 80% power, against an effect reported from a corrected analysis of twenty, has 38% power. It will fail more often than it succeeds, and the failure will be read as evidence against the original finding when it is a consequence of the original finding’s arithmetic.

That is the mechanism by which a correction intended to make a literature more reliable makes its replications less informative, and it is entirely invisible in either study. The first study’s p-value is honest. The second study’s power calculation is arithmetically correct. The input to the second is a number the first was structurally unable to report accurately.

The chance an exact replication reaches significance, given the first p-value. Two models for the effect behind a first two-sided p-value. Taking the estimate as the truth gives a replication significant in the same direction 73.1% of the time after p = 0.01 and 90.8% after 0.001. A flat prior, which carries the estimate's own uncertainty into the prediction, gives 66.8% and 82.7%. At p = 0.05 both are exactly one half. The points are counted from 600,000 simulated pairs of studies with effects spread flat, keeping the pairs whose first p-value fell near 0.05, 0.01 and 0.001, and they agree with the curve within their error bars of two standard errors.
Fig. 4 What a replication of a selected finding is up against before any of this is counted. The inflation measured on this page enters as a further reduction in the size the second study thought it needed.

Why the two problems cannot share a repair

The reason is that they are failures of different objects, conditioned on different events.

The error rate is a property of the procedure over all families, including those that report nothing. Raising the threshold reduces the share of null families that report, which is the whole mechanism, and it works.

The estimate is a property of the families that report. It is conditional on the reporting event by construction — an unreported estimate is not in the average — so changing the reporting event changes the conditional distribution. Making the event rarer makes it more extreme.

Those two pull opposite ways and there is no threshold at which both are satisfied. The threshold that makes the estimate unbiased is no threshold at all, and no threshold gives a 5% error rate on a family of twenty.

That is a structural statement rather than a limitation of the corrections available. Any filter applied to honest studies inflates what passes it, and a multiple-comparison correction is a filter, applied on purpose, to a set that was already filtered by taking a maximum.

What actually repairs the estimate

Nothing on the p-value side, and three things elsewhere.

Report the analysis that was prespecified, whether or not it won. Its estimate is unconditional and therefore unbiased. That is a separate virtue of prespecification from the power trade that naming an analysis in advance costs, and it is the larger one: the prespecified estimate is right, and the exploratory one is inflated by a factor the table above puts between 1.05 and 6.4.

Report all twenty estimates. The set of twenty is unconditional; only the maximum is selected. A paper printing all twenty effect sizes has given a reader everything needed to see how extreme the reported one is within its own family, and the shape of the other nineteen says more about whether the winner is real than its p-value does.

Shrink it. An estimate selected as a maximum has a known conditional bias given the threshold it cleared, and it can be corrected — by conditioning on the selection event explicitly, or by borrowing from the rest of the family, which shrinks every estimate towards their common centre and shrinks the largest most. The second is the one that generalises, and it has the property the first lacks: it does not need the threshold to be known.

The three are not equally available and they fail differently. The first needs a prespecification that may not exist. The second needs journal space and a reader willing to use it. The third needs a population of analyses that can reasonably be treated as exchangeable, which twenty outcome measures on the same subjects usually are and twenty exclusion rules are not — the rules are not twenty draws from anything, they are one analysis perturbed, and shrinking towards the mean of twenty perturbations borrows from a population that does not exist. The first two are always correct and rarely done; the third is often the best available and needs an argument each time.

12,000 studies of a real effect of 0.3, n = 16. Power is 21%. The studies that reached significance report a mean effect of 0.621 — 2.07 times the truth. Every one of them is honest; the selection did the inflating.
Fig. 5 The same selection without any correction in it. Filtering honest studies down to the ones that reached significance inflates what they report, and the correction on this page raises the filter.

What the trade looks like in total error

Inflation is only half of an estimate’s quality, so it is worth asking whether the correction is better or worse on a measure that counts both halves.

It is worse on both, which is the part that makes this more than a curiosity, and the arithmetic is worth following because the second half is the one that is usually assumed to go the other way. Raising the threshold increases the bias of what survives — the table above — and it does not reduce the variance of what survives, because the surviving set is a tail of the same distribution rather than an average over more of it. A narrower, more extreme tail has a larger squared bias and a comparable spread, so the mean squared error of the reported estimate rises with the threshold.

The comparison that would make the correction look good on this measure is against an analysis that reports its maximum with no threshold at all, and that comparison is not available: an analysis with no threshold reports on every family, including the ones with nothing in them, so its error rate is 64% and its estimate on a null family is pure noise.

There is one measure on which the correction does win, and naming it keeps the page honest. The unconditional mean squared error of the whole procedure — counting families that report nothing as reporting the null — falls with the threshold, because null families stop reporting noise. That is the quantity a decision-theoretic account of testing optimises, and it is not the quantity a reader of a published effect size has in front of them. A reader sees the reports and not the silences, so the conditional measure is the one that describes their experience, and the two point opposite ways.

The honest statement is therefore narrow. Among analyses that control their error rate, a more severe correction produces a more inflated estimate. It says nothing about whether to correct — the answer to that is yes — and everything about what a corrected estimate should be trusted to be.

What a meta-analysis does with it

The second place the estimate is carried is into a pooled analysis, and pooling does not fix a bias that every input shares.

A meta-analysis averages the reported effects, weighted by their precision. If each reported effect is inflated by a factor between 1.3 and 1.8, the pooled estimate is inflated by a factor in that range, and the pooling has narrowed its confidence interval around the wrong value. The interval’s width falls like 1/k1/\sqrt{k} in the number of studies and the bias does not fall at all, so the more studies are pooled, the more confidently the wrong number is reported.

That is the ordinary criticism of publication bias, arriving here from a source nobody adjusts for. A funnel plot detects asymmetry produced by missing small studies; the inflation on this page is present in every study that appears, small and large, and it leaves no asymmetry to detect — the large studies are inflated less, which is the same direction a funnel plot reads as selective publication, so the two are confounded in exactly the diagnostic meant to separate them.

The consequence for a reader of a meta-analysis is one question. Were the pooled estimates prespecified analyses or selected ones? If the primary outcomes were named in advance, the inflation from the maximum-selection is absent and only the publication filter remains. If the included studies each reported their best of several outcomes, every input carries the factor in the table above and the pooled result carries it too, with a narrower interval to make it look more certain.

Nothing in a standard meta-analysis asks that question, and the information needed to answer it is usually in the included papers.

What is claimed here, and what is not

Two statements, and the second is what keeps the first from being trivial.

The corrected threshold’s surviving estimate is inflated by more than the uncorrected one’s, at every effect size drawn. Stated across the whole sweep rather than at one setting, because a comparison at a single effect is a fact about that effect and this is a claim about the family of them.

The inflation falls as the effect grows. A large effect does not need to be lucky, so the selection has little to select on. That is worth stating because it distinguishes the measurement from one that had confused the conditional and unconditional means — the unconditional inflation also falls with the effect, but it falls to 1.00 from 1.18, and the conditional one falls from 6.37, so the wrong one would look qualitatively similar and be an order of magnitude out.

The reading that does not survive is a family-wise-corrected estimate taken as unbiased. The standard every estimate here is held to is that it should average the thing it estimates, and at two standard errors this one averages 1.690 times it. An estimate that did average the truth would mean the conditioning was not happening — an average over all families rather than over the ones that reported, which is the slip that makes a selection effect vanish from a simulation entirely.

Still open: what an honest exploratory report looks like

The three repairs above are each partial and none is standard practice, so the question of what such a paper should actually print has no settled answer.

Printing all twenty estimates is cheap and rare. Printing the corrected p-value beside a shrunk estimate is coherent and has an awkwardness the others do not: the p-value and the estimate then disagree about how strong the finding is, because they are conditioned differently, and a reader has no convention for holding two numbers that disagree on purpose.

What is missing is a report format in which the selection is visible rather than corrected for. The elements exist — the family, the threshold, the distribution of the other nineteen, the conditional distribution of the maximum given the threshold — and nothing combines them into something a reader can take in. Whether such a format is possible at the length an abstract has, or whether the inflation simply has to be carried as a caveat, is a question about scientific writing that this measurement makes precise and does not answer.

Still open: whether a reader can hold two estimates at once

One candidate is worth naming because it costs nothing and is testable. Report the effect that would have been reported had the threshold been the uncorrected one, beside the corrected p-value. That number is available — it is the same maximum — and the pair carries the information: an honest error rate, and an estimate conditioned on the weakest filter the analysis could have used. It is still inflated, by the factor in the first column of the table rather than the second, and the difference between the two columns is precisely what a reader currently cannot see.

Whether that pairing reads as coherent or as two numbers in an argument with each other is the thing that would have to be tried, and it is the same difficulty a shrunk estimate raises beside its own interval.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BonferroniEffect sizeFamilywise error rateMean squared errorMultiple comparisonsSelection biasStatistical powerThe winner's curse