Two different promises
Worth reading first: What the correction corrects.
Twenty tests, ten of them with a real effect. Benjamini–Hochberg holds the false discovery rate at 2.6% against its nominal 5%, and its familywise error rate is 20.0%.
Both of those are correct behaviour. It never promised the second number.
The two quantities
They differ in what is in the denominator, and everything follows from that.
The familywise error rate is the probability that the family contains at least one false positive. Its denominator is the family: one number per study, and it does not care how many findings there were.
The false discovery rate is the expected proportion of the rejections that are false. Its denominator is the set of findings: if a study reports a hundred discoveries and five are false, the rate is 5% regardless of how many tests were run.
Consider a study with one significant finding, and it is false. Familywise error: it happened. False discovery rate: 100%. Now a study with a hundred findings, five of them false. Familywise error: it happened, identically. False discovery rate: 5%.
The two situations are treated as equivalent by the first criterion and as forty-fold different by the second, and neither is wrong. They are answers to different questions about what a study is for.
The measurement, on the same families
Four procedures, four thousand families of twenty tests with ten real effects, both rates counted:
| procedure | familywise | false discovery | power |
|---|---|---|---|
| no correction | 40.8% | 5.3% | 85.1% |
| Bonferroni | 2.8% | 0.5% | 49.1% |
| Holm | 3.6% | 0.6% | 52.5% |
| Benjamini–Hochberg | 20.0% | 2.6% | 74.9% |
Read the last row against the second. Benjamini–Hochberg lets the familywise rate reach seven times Bonferroni’s, and buys 26 percentage points of power with it. Both hold what they claim: BH’s false discovery rate is 2.6% and Bonferroni’s familywise rate is 2.8%, and both are under 5%.
The row that should be looked at hardest is the first. Uncorrected testing has a false discovery rate of 5.3% on this family — essentially at nominal. When half the nulls are false, doing nothing at all controls the false discovery rate approximately, because most of the rejections are true ones and they dilute the false ones.
That is not an argument for doing nothing. It is a demonstration that the false discovery rate is a weak promise when many effects are real, and its weakness is exactly where it is most often invoked.
Three of the four rows are predictable in advance
Half the family is real here, and that single fact reproduces most of the table without any simulation.
No correction, familywise. Only the ten true nulls can produce a false positive, so the rate is . Measured 40.8%, within one standard error.
Bonferroni, familywise. The threshold is 0.05/20 = 0.0025 and there are ten nulls, so . Measured 2.8%.
Benjamini–Hochberg, false discovery. BH controls the false discovery rate at , which here is . Measured 2.6%.
Each of those matches, and each says the same thing about where the conservatism comes from. A correction that divides by the number of tests is spending half its budget on the ten tests that have no null to reject. That is not about dependence, or about the tests being poorly chosen; it is arithmetic, and it is why Bonferroni’s familywise rate comes in at half of nominal rather than at nominal.
The fourth row, and what doing nothing already controls
The row that is not predictable is the uncorrected false discovery rate, and it is the one worth dwelling on.
Ten true nulls at 5% produce 0.5 expected false positives; ten real effects at 85.1% power produce 8.5 true ones. The ratio is 5.5%, and the measured figure is 5.3%.
So on this family, doing nothing at all controls the false discovery rate at five per cent. Not approximately — it is the procedure’s actual rate, and it is the same five per cent every correction in the table is trying to reach.
That is the strongest possible statement of the essay’s point. The two promises are not a strict and a lenient version of one thing. On a family that is half real and well powered, one of them is violated forty per cent of the time by the uncorrected procedure and the other is already held.
The two corrections’ exchange rates
Read as trades, the three corrections are buying familywise protection with power at very different prices.
Holm over Bonferroni: +3.4 points of power for +0.8 points of familywise rate — 4.2 points of power per point spent, with the rate still under 5%. That is not a trade so much as a free improvement.
BH over Holm: +22.4 points of power for +16.4 points of familywise rate — 1.4 points per point, a third of Holm’s rate, and it takes the familywise rate to 20%.
So the three procedures are not a sequence of increasingly relaxed versions of one idea. Two of them are on the same curve at very different points, and the third is buying something else entirely at a price that only makes sense if the familywise rate is not what the study is for.
What each is for
The choice is not a matter of taste and it follows from what happens next to the findings.
Familywise control is right when each finding will be acted on individually. A confirmatory trial with a primary outcome, a regulatory submission, a safety signal, any claim that will be quoted on its own. Here one false positive is a failure whatever else the study found, and that is precisely what the familywise rate measures.
False discovery control is right when the findings are a set to be followed up. A screen of ten thousand genes, a scan of every region of an image, a list of candidate features. Nobody will act on any one of them individually; they will all be taken to a second stage. What matters is that the list is not mostly rubbish, and “mostly” is a proportion.
The distinguishing question is therefore: will these findings be used one at a time or as a list? That is usually easy to answer and it is almost never asked, because both procedures are described in the methods section with the same phrase.
Why the familywise rate is unbounded under BH
The 20% deserves an explanation rather than only a measurement, because it is the number that surprises people who use BH routinely.
Benjamini–Hochberg rejects everything up to the largest k with p₍ₖ₎ ≤ kα/m. When many effects are real, the real ones supply small p-values, k gets large, and the threshold kα/m becomes generous — approaching α itself at the top of the list.
So a true null whose p-value happens to be moderately small gets rejected, because it is riding on a threshold that the real effects raised. The more real effects there are, the more permissive the boundary and the more likely at least one false positive.
That is the mechanism, and it means BH’s familywise rate rises with the number of true effects. Measured at two real effects it is well below its value at fifteen, and the site’s gate asserts the direction rather than describing it.
This is worth understanding because it inverts an intuition. A study with many real findings is a study where BH is most permissive, so a rich dataset is where the familywise guarantee is furthest from holding — the opposite of the reassurance a long list of discoveries provides.
The phrase that causes the trouble
“Corrected for multiple comparisons” appears in methods sections and carries no information, because it is consistent with a familywise rate of 3% and with one of 20%.
A reader who sees it and infers that false positives have been controlled to 5% has inferred something that may be off by a factor of four, in a direction determined by which procedure was chosen and how many effects are real — neither of which is usually stated.
The repair is one clause: name the procedure and name the rate it controls. “Holm, controlling the familywise error rate at 5%” or “Benjamini–Hochberg, controlling the false discovery rate at 5%”. Both are shorter than most of the sentences they would replace, and they are the difference between a claim and a gesture.
What BH assumes
One caveat that matters and is frequently omitted.
The Benjamini–Hochberg procedure controls the false discovery rate exactly under independence, and under a condition called positive regression dependence which covers many realistic cases — one-sided tests of correlated normal quantities, for instance.
Under arbitrary dependence it does not, and the fix is a variant that divides the threshold by the harmonic sum of 1 through m, which for twenty tests is a factor of about 3.6. That is a substantial penalty and it is rarely applied, mostly because the dependence structure is rarely examined.
Bonferroni and Holm need no such condition. Their guarantee holds for any dependence whatever, which is the compensation for their conservatism, and it is a real advantage in settings where nobody knows how the tests relate.
So the comparison is not simply “BH gives more power”. It is: BH gives more power, controls a weaker quantity, and needs an assumption the familywise procedures do not.
Choosing, in three questions
The field’s practical output, and it takes less than a minute.
Will each finding be acted on separately, or is this a list for follow-up? Separately means familywise. A list means false discovery.
How many tests, and were they declared in advance? If they were not declared, no procedure here applies honestly, and the problem is the one that has no correction.
Is the dependence structure known? If not, the familywise procedures still hold and BH may not.
Answering those three names the procedure. Not answering them and writing “corrected for multiple comparisons” leaves a reader unable to tell whether the false positive rate being controlled is the one they care about.
The false discovery rate is an expectation
A property that is easy to miss and changes how the guarantee should be read.
The false discovery rate is the expected proportion of false findings, averaged across repetitions of the whole study. It is not a statement about the study in hand.
A particular run can have a much worse proportion. With few rejections the proportion is coarse — one false finding out of three rejections is 33%, and no procedure controlling an average at 5% prevents that outcome occurring. The guarantee is about the long run, and a single study’s list can be considerably worse than the nominal rate without anything having gone wrong.
That is the same distinction this site draws about what a coverage claim attaches to: a property of the procedure, evaluated across the repetitions that did not happen, and not a property of the output in front of anyone.
Two consequences.
Small numbers of rejections make the guarantee coarse. A study with four discoveries controlling the FDR at 5% is not promising that 0.2 of them are false; it is promising an average across studies, and the realised proportion will be 0, 25%, 50% or worse.
And the variability is not reported. There are procedures that control the probability that the false discovery proportion exceeds some bound — a stronger and more useful guarantee — and they are almost never used, largely because the software defaults do not offer them.
Where the two promises coincide
Worth identifying, because it explains why the distinction can be ignored for a long time.
When all the nulls are true, the two rates are the same number. If there are no real effects, every rejection is false, so the proportion of false rejections is either 0 or 1 — and its expectation is exactly the probability of at least one rejection, which is the familywise rate.
The measurement bears this out: on twenty tests with nothing real, the familywise rate and the false discovery rate agree for every procedure, at 5.27% for Bonferroni and Holm and 5.42% for BH.
So under the global null the procedures are indistinguishable, and every difference between them appears only when some effects are real. That is why the distinction is invisible in the textbook demonstration, which is always run on pure noise, and why it becomes important in exactly the applications the procedures were built for.
It also means BH controls the familywise rate in the one case where the two coincide — a property with a name, weak familywise control — and loses it as soon as anything real appears.
What the numbers say about the default choice
A conclusion the table supports and the literature does not always draw.
For a study with a small number of pre-declared tests — a trial with a primary and three secondary outcomes, say — Holm is the right default. The tests are few, so the conservatism is small; each finding will be quoted separately, so familywise control is the relevant promise; and it needs no assumption about dependence.
For a large screen — thousands of tests, results going to a second stage — BH is the right default, and the familywise procedures are not merely conservative but unusable. Holding the chance of any false positive to 5% across ten thousand tests requires a per-test threshold of 5 × 10⁻⁶, and almost nothing survives it.
The awkward middle is a study with twenty or fifty tests whose findings will be reported individually and followed up selectively. That is most of applied research, both promises are arguable, and the honest practice is to report both rates rather than to pick one and describe it vaguely.
Reporting both costs nothing: the same p-values, two procedures, four numbers instead of two.
The general lesson, which is not about multiplicity
This field’s finding generalises past its own subject, and it is worth stating in the general form because the site keeps meeting it.
Procedures promise to control specific quantities, the quantities differ, and the nominal level is the same 5% in every case. A confidence interval’s coverage, a test’s size, a familywise rate, a false discovery rate, the per-look error rate of a stopping rule: all are held at 5%, all are different things, and a method that holds one can violate another by a wide margin while behaving exactly as designed.
The habit that follows is a small one. Whenever a method is described as controlling something at 5%, ask which quantity — and if the answer is not immediately available, the guarantee has not actually been communicated.
That question is the thread running through this field, and answering it is most of what separates a procedure that protects a conclusion from one that decorates it.
One table, four sentences
The measurement in this essay reduces to four statements, and each is counted rather than argued.
Uncorrected testing on twenty tests with ten real effects has a familywise rate of 40.8% and a false discovery rate of 5.3% — so it fails one criterion badly and passes the other, which is why doing nothing can be defended when the criterion is not named.
Bonferroni and Holm hold the familywise rate at 2.8% and 3.6%, and pay for it with power of 49% and 53%.
Benjamini–Hochberg holds the false discovery rate at 2.6%, allows a familywise rate of 20.0%, and achieves power of 75%.
And under the global null all three corrected procedures give the same answer, which is why the difference between them is invisible in every demonstration run on pure noise.
Why this is the field’s central essay
The other two essays here are about machinery: what the procedures do and what they cost. This one is about the question that has to be answered before either matters, and it is the question the standard vocabulary obscures.
A researcher who knows exactly how Bonferroni works, can derive Holm’s step-down argument, and applies Benjamini–Hochberg correctly can still choose the wrong one — because choosing requires deciding what a false positive costs relative to a missed effect, and that is a judgement about the study rather than a fact about the statistics.
The measurements here do not make that judgement. What they do is put a price on it: 26 percentage points of power against a sevenfold increase in the chance of any false claim, at this family size and this effect. Both numbers are counted, both are specific to the configuration, and the trade is now visible instead of hidden inside a phrase.
That is as far as arithmetic can take the question, and it is further than it usually gets taken.
The one thing arithmetic does settle is that the phrase covering both options is not adequate. Two procedures that differ by a factor of seven in the rate a reader most likely cares about should not be described by the same six words, and replacing those words with the name of the procedure and the name of the rate costs nothing at all.
Naming the rate is the whole reform, and it fits in a subordinate clause.
Everything else in the field is a consequence of getting that clause right or leaving it out, and the two outcomes look identical until somebody counts. This essay counted, and the gap it found was a factor of seven in the rate most readers assume is the one being held. A factor of seven is not a nuance, and it is invisible in the sentence that currently describes both procedures.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Estimating how many nulls are true — both name benjamini–hochberg, dependence, error rate, false discovery rate, multiple comparisons
- Intervals for the findings — both name benjamini–hochberg, false discovery rate, multiple comparisons
- The correction that makes the estimate worse — both name bonferroni, familywise error rate, multiple comparisons
- The models that were never in the running — both name error rate, familywise error rate, multiple comparisons
- What naming it in advance costs — both name false positive, familywise error rate, multiple comparisons
- A coverage table with its own error — both name bonferroni, multiple comparisons
Named objects
A flat tag is an object no other essay names yet.
Benjamini–HochbergBonferroniDependenceError rateFalse discovery rateFalse positiveFamilywise error rateHolmMultiple comparisons