Corrections, and what each controls

The price of control

Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.

Worth reading first: What the correction corrects · What a p-value does not say.

Twenty tests, ten real effects of three standard errors each. Uncorrected testing finds 85.1% of the real ones. Holm finds 52.5%. The correction cost thirty-three percentage points of the real effects in the data.

Power to find a real effect of 3 standard errors, 10 of 20 realno correction finds 85.1%, Bonferroni finds 49.1%, Holm finds 52.5%, Benjamini–Hochberg finds 74.9%. The uncorrected procedure finds the most and controls nothing.no correction85.1%familywise 41%Bonferroni49.1%familywise 3%Holm52.5%familywise 4%Benjamini–Hochberg74.9%familywise 20%4,000 families of 20 testspower is what the control is bought with
Fig. 1 What each procedure finds, at the same effect size on the same families. The uncorrected procedure finds the most and guarantees nothing.

The trade, stated once

Every procedure in this field works by raising the threshold, and raising the threshold rejects fewer things. Some of the things not rejected are true nulls, which is the point, and some are real effects, which is the cost.

There is no procedure that removes false positives without removing some true ones, because a procedure cannot tell them apart — it sees p-values, and a true null with a small p-value looks exactly like a real effect with a moderate one.

So the question is never whether to pay but how much, and against what.

The measured exchange rate

At twenty tests with ten real effects of three standard errors:

procedure power familywise false discovery
no correction 85.1% 40.8% 5.3%
Bonferroni 49.1% 2.8% 0.5%
Holm 52.5% 3.6% 0.6%
Benjamini–Hochberg 74.9% 20.0% 2.6%

Read as prices:

Holm buys the familywise guarantee for 33 points of power. From 85.1% down to 52.5%, and the familywise rate falls from 40.8% to 3.6%.

Benjamini–Hochberg buys the false discovery guarantee for 10 points. From 85.1% to 74.9%, and it leaves the familywise rate at 20%.

And Bonferroni charges 36 points for what Holm delivers for 33, which is the same guarantee at a worse price and the reason the recommendation in the first essay was unconditional.

The interesting comparison is the second row against the fourth. Twenty-two points of power separate Holm from BH, and the difference between them is entirely the strength of the promise. Nothing else changed — same data, same tests, same nominal 5%.

Where the three power figures come from

Every power in the table is a normal tail and it is worth checking, because the arithmetic explains the one number that looks like a coincidence.

At an effect of three standard errors, an uncorrected two-sided test at 5% rejects when the statistic exceeds 1.96, so its power is Φ(31.96)=Φ(1.04)\Phi(3 - 1.96) = \Phi(1.04), which is 85.1% — the table’s first row exactly.

Bonferroni compares against 0.05/20 = 0.0025, whose two-sided critical value is 3.0233. So its power is Φ(33.0233)=Φ(0.023)\Phi(3 - 3.0233) = \Phi(-0.023), which is 49.1% — the table’s second row exactly.

Which explains the round-looking figure. Bonferroni’s power is a hair under a half because its critical value, 3.02, is essentially the effect size, 3.00. A test whose threshold sits on the alternative’s mean rejects half the time, and the whole of Bonferroni’s cost at this effect size is that it has moved the threshold onto the effect.

The cost peaks near the effect this field measures at

That framing makes the exchange rate a function rather than a number, and locating its maximum says how representative the table is.

The power cost is Φ(δ1.96)Φ(δ3.0233)\Phi(\delta - 1.96) - \Phi(\delta - 3.0233). Differentiating, it is largest where φ(δ1.96)=φ(δ3.0233)\varphi(\delta-1.96) = \varphi(\delta-3.0233), which is exactly midway between the two critical values:

δ=1.96+3.02332=2.49.\delta = \frac{1.96 + 3.0233}{2} = 2.49 .

There the cost is 40.5 points. At the field’s δ = 3 it is 36.0.

So the measurement sits near the top of the curve, and the curve falls away steeply on both sides. At δ = 3.87 Bonferroni has 80% power against the uncorrected test’s 97% — a cost of 17 points. At δ = 5 it is 97.6% against 99.9%, a cost of 2 points.

The correction is nearly free against large effects and nearly halves the study against effects the size of its own threshold. That is the sentence to carry into a design, and it says the number to compute is not the correction’s severity but where the effect being hunted sits relative to the corrected critical value — at three standard errors, right on it.

Why the cost depends on the effect size

The exchange rate is not a constant, and dragging the effect size makes the dependence visible.

At a large effect, every procedure finds nearly everything, because the real p-values are far below every threshold. The correction costs almost nothing and the debate is empty.

At a small effect, no procedure finds much, because the real p-values are not far below the uncorrected threshold either. The correction costs little in absolute terms because there was little to lose.

The cost is largest in the middle, where real effects produce p-values in the neighbourhood of the uncorrected threshold — which is exactly the region most applied research operates in, and exactly the region where the choice of procedure changes which findings get reported.

That has a consequence for how these comparisons should be read. A demonstration run at a large effect will show the corrections as nearly free, and a demonstration run at a tiny one will show them as irrelevant. Both are misleading about the case that matters, and the site’s gate compares across effect sizes rather than asserting anything at one of them.

Power to find a real effect of 2 standard errors, 10 of 20 real. no correction finds 51.7%, Bonferroni finds 15.3%, Holm finds 15.9%, Benjamini–Hochberg finds 26.8%. The uncorrected procedure finds the most and controls nothing.
Fig. 2 The same procedures at a smaller effect. Everything finds less and the gaps between the procedures change shape.

The response that actually works

Faced with the trade, the usual moves are to argue about which procedure to use or to declare fewer tests. There is a third and it dominates both.

Design the study with enough power that the correction is affordable.

The corrected threshold is known before the data is collected — it is α/m for Bonferroni, and something similar for the others. A power calculation performed at the corrected threshold rather than at 0.05 gives the sample size that makes the whole apparatus work, and that calculation is no harder than the standard one.

A study powered at 80% for a single test at 0.05 has considerably less than 80% power for the same effect after correcting for twenty tests. The measurement above puts it at about 52%, which means such a study will miss half its real effects no matter how carefully the procedure is chosen.

So the argument about procedures is, in most cases, an argument about how to allocate a shortfall that was created at the design stage. Choosing Benjamini–Hochberg over Holm recovers twenty-two points of power; powering the study for its actual analysis plan recovers all of it.

This is the second missing number arriving in a new setting. Power is the quantity that determines whether a study can support its own conclusions, it is computable in advance, and it is the one most often left out.

What the choice looks like when it is made honestly

Three cases, and the reasoning is different in each.

A confirmatory trial with four outcomes. The findings will be quoted individually and one false claim is the failure mode. Holm, familywise at 5%, powered at the corrected threshold. The correction is small at four tests and the design can absorb it.

A screen of ten thousand features. Familywise control is unusable — the per-test threshold would be 5 × 10⁻⁶ and nothing would survive. Benjamini–Hochberg, false discovery rate at 5% or 10%, with the understanding that the output is a candidate list and the second stage is where the real error control happens.

A study with twenty outcomes reported individually. The genuinely hard case. Both promises are arguable, the power cost of the strict one is large, and the honest practice is to declare a primary outcome that carries the familywise guarantee and treat the rest as exploratory with a false discovery rate — reporting both, and labelling which is which.

The third arrangement is the one most protocols should be using and few do, because it requires deciding in advance which finding the study is actually about.

Twenty tests, 10 of them real — what each procedure holds. no correction: familywise 40.8%, false discovery 5.3%, power 85%. Bonferroni: familywise 2.8%, false discovery 0.5%, power 49%. Holm: familywise 3.6%, false discovery 0.6%, power 53%. Benjamini–Hochberg: familywise 20.0%, false discovery 2.6%, power 75%.
Fig. 3 The two rates for every procedure, which is what the choice above is choosing between.

What no procedure can buy

The limit of everything in this field, and it is worth ending on because it bounds how much the machinery is worth.

Every procedure here takes m as an input, and m is the number of tests declared. None of them addresses the tests that were available and not declared, the outcomes measured and not analysed, or the analysis that was chosen after the data was seen.

A study that ran twenty tests, declared twenty, and applied Holm has controlled its familywise rate at 5%. A study that considered two hundred analyses, ran twenty, declared twenty and applied Holm has controlled nothing, and the two are indistinguishable in print.

So the procedures buy exactly one thing: they convert a known multiplicity into a controlled error rate. They do not convert an unknown multiplicity into anything, and the unknown kind is the larger problem by every measurement this site has taken.

That is not an argument against using them. It is an argument for pre-registration as the thing that makes them work — and it explains why the two reforms are usually proposed together and why either alone underdelivers.

Twenty p-values sorted, with the three thresholds, 6 real. Bonferroni is a flat line at α/m = 0.0025. Holm starts there and rises. Benjamini–Hochberg is the steepest line, iα/m. On this family they reject 5, 5 and 6 hypotheses respectively.
Fig. 4 The three thresholds one last time. Each is a line on a plot, and none of them knows how many tests were considered before twenty were written down.
The familywise error rate with no correction, α = 0.05. Two routes: the curve is 1 − (1 − α)^m and the points are counted over 6,000 simulated families of true nulls. With twenty tests the chance of at least one false positive is 64.1%.
Fig. 5 And the quantity all of it is spent on: the chance of at least one false positive, climbing with every test in the family.

Where the power goes

A closer look at what “losing 33 points of power” means in a study, because the aggregate figure hides something.

The power lost is not spread evenly across the real effects. The corrections raise a threshold, so the effects that stop being detected are the ones whose p-values sat just below the old threshold — the smallest real effects in the family.

The large effects survive every correction. The marginal ones do not survive any of them.

So the systematic consequence of correcting is that a study’s reported findings are biased toward larger effects, and the bias is stronger the stricter the procedure. That is the winner’s curse arriving through the correction rather than through low power, and it compounds with it: a study that is underpowered and heavily corrected reports only the effects that were both real and lucky, and overstates them by more than either mechanism alone would.

The measurement in the winner’s curse essay says a study at 21% power inflates its published effects by 2.07×. Correction lowers the effective power further, so a corrected analysis in an underpowered study is operating in the region where the inflation is largest.

Two things follow.

A corrected analysis needs its effect estimates read with more discounting, not less. The rigour of the correction applies to the false positive rate and does nothing about the magnitude of what survives.

And reporting the effects that did not survive matters more than usual, because those are systematically the smaller ones and their absence distorts the picture the study leaves behind.

The comparison nobody runs

A gap worth naming, since this essay has been comparing procedures at a fixed design.

The comparison that would actually inform a study is not “which correction, given this data” but “which combination of sample size and correction, given this budget”. Those are different questions and the second is answerable in advance.

Doubling a sample size raises power substantially and costs money. Switching from Holm to Benjamini–Hochberg raises power by twenty-two points and costs a weaker guarantee. The two are alternatives on the same axis and they are never set against each other, because one is a design decision made months before the other is an analysis decision.

Where the comparison has been made, the answer is usually that the sample size is the better purchase, because it improves the estimates as well as the detection rate — a larger study finds more real effects and reports them less inflated, where a weaker correction finds more and reports them just as inflated.

That is the strongest argument available for treating the multiplicity correction as a design input rather than an analysis choice, and it is the practical conclusion of this field.

What the three essays establish

The field’s findings, each counted rather than argued.

Twenty uncorrected tests of true nulls produce a false positive 64% of the time, matching the closed form at every family size.

Holm dominates Bonferroni — the same guarantee, more power, at every configuration tested — so one of the two most-used procedures has no remaining justification.

Benjamini–Hochberg controls a different quantity, holding the false discovery rate at 2.6% while allowing a familywise rate of 20%, and both behaviours are correct.

The corrections cost between 10 and 36 percentage points of power at a realistic effect size, and the cost is largest in the region where most research operates.

And none of it applies to tests that were not declared, which is the larger source of false positives and the one no procedure addresses.

The through-line is the field’s thread: every procedure holds some quantity at 5%, the quantities differ, and the only way to know what has been promised is to ask which one.

Twenty tests, 15 of them real — what each procedure holds. no correction: familywise 23.0%, false discovery 1.9%, power 85%. Bonferroni: familywise 1.3%, false discovery 0.2%, power 49%. Holm: familywise 2.5%, false discovery 0.3%, power 55%. Benjamini–Hochberg: familywise 15.8%, false discovery 1.3%, power 80%.
Fig. 6 A last configuration: fifteen of twenty effects real, where Benjamini–Hochberg is at its most permissive and its familywise rate at its highest.

Why “just report everything uncorrected” is not the answer

The position deserves a hearing, because it is held by serious people and it is not obviously wrong.

The argument: corrections destroy power, the family is arbitrary, and a reader is better served by seeing all twenty p-values with no adjustment and drawing their own conclusions. Reporting is transparent, nothing is hidden, and the reader has more information than any corrected summary would give them.

Three things are right about it. Reporting all twenty is better than reporting the significant ones. The family boundary is a judgement rather than a fact. And a sophisticated reader with all twenty p-values can apply whatever correction they prefer, which is more than a corrected summary allows.

What is wrong with it is the assumption about what happens next. Findings do not stay attached to the family they came from. A significant result travels into an abstract, a press release, a citation and a meta-analysis, and at every step the other nineteen p-values are left behind. By the time the finding is being used, the context that would let a reader correct it has gone.

So the transparency argument holds within a paper and fails between papers, and the damage happens between papers.

The synthesis both sides can accept: report all the tests, and report the corrected conclusion. The full list serves the sophisticated reader and preserves the information; the corrected conclusion is what travels, and it should be the one that carries a guarantee. That is two extra lines in a results section and it resolves most of the dispute.

The number to compute before designing anything

If this field produces one instruction, it is a calculation rather than a preference.

Before the study runs, take the number of tests intended, work out the corrected threshold, and compute the power at that threshold for the smallest effect worth detecting. That number is the study’s real power, and it is the one that determines whether the analysis plan can support a conclusion.

If it is well below what was intended, the choices are a larger sample, fewer tests, or an explicit decision to run an exploratory study whose findings are candidates rather than claims. All three are legitimate and all three are better than discovering the shortfall at the analysis stage, where the only remaining lever is which correction to apply — and that lever, as the table at the top shows, is worth twenty-two points at most.

The correction is a small decision made late. The power at the corrected threshold is a large decision made early, and it is the one this field is really about.

A last measurement worth carrying

One number from the table deserves to be remembered on its own, because it reframes what a corrected analysis is doing.

At twenty tests with ten real effects, Holm finds 52.5% of them. Not a marginal reduction — barely half. A study that correctly declares its tests, applies the best available familywise procedure, and is powered the way most studies are powered will fail to detect one real effect in every two that are present.

Those missed effects are not errors. Nothing went wrong, no assumption was violated, and the procedure did exactly what it promises. They are the price, and the price is being paid silently in every corrected analysis that was powered for a single test.

Making that number visible is most of what this field is for. A researcher who knows their analysis has 52% power after correction will either fix the design or say plainly that a null result means very little, and both are better than the current default, which is to report the corrected findings as though the ones that did not survive had been shown to be absent. Half of them were simply not looked for hard enough, and the arithmetic saying so was available before the study began.

That last point is the one worth carrying out of the field. Every quantity in this essay — the corrected threshold, the power at it, the number of tests, the effect worth detecting — is available before a single observation is collected. The correction is the only part of it that happens at the end, and it is the part that gets all the attention.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BonferroniEffect sizeFalse discovery rateFalse positiveHolmMultiple comparisonsSample sizeStatistical powerStudy design