How many analyses there really were
Worth reading first: Twenty analyses of nothing · What the correction corrects.
Twenty honest analyses of nothing find something significant 57% of the time, against the 64% that twenty independent analyses would give. The gap is the correlation between them, it was reported there as a comparison, and nothing was done with it.
Something can be done with it. The gap says the correction everybody applies is for the wrong number of analyses.
What the number means
Bonferroni’s rule is a bound: test each of k analyses at α/k and the chance of at least one false positive is at most α. The bound is exact when the analyses are independent and conservative otherwise, and “otherwise” is every real family, because the analyses share a dataset.
The threshold that is exactly right can be measured. Build the null distribution of the largest of the k statistics, take its 95th percentile, and that is the critical value giving a 5% family-wise rate by construction. Read back the per-analysis α that critical value corresponds to, and the ratio αfamily/α* is how many independent analyses the family is worth.
At a correlation of zero the measurement returns 20.05 for twenty analyses, and the critical value is 3.024 against Bonferroni’s 3.023. That agreement is the first thing to check and it is worth stating: under independence the measurement reproduces Bonferroni to three decimal places, so what follows is not a criticism of the bound but a measurement of how much slack it has.
How much slack there is
| correlation | effective analyses | critical value | uncorrected error rate |
|---|---|---|---|
| 0.0 | 20.05 | 3.024 | 64.2% |
| 0.2 | 19.17 | 3.011 | 59.0% |
| 0.4 | 15.81 | 2.951 | 48.3% |
| 0.6 | 11.37 | 2.848 | 35.5% |
| 0.8 | 6.30 | 2.655 | 22.5% |
| 0.95 | 2.58 | 2.338 | 11.5% |
Two things in that table are worth separating, because they are usually conflated.
The right-hand column is the problem and the middle columns are the repair. The uncorrected error rate falls as the analyses become more alike — twenty near-identical analyses are nearly one analysis, so running them costs nearly nothing. That is the familiar observation.
The effective count falls faster than the error rate does. Going from no correlation to 0.6 takes the uncorrected rate from 64.2% to 35.5%, a factor of 1.8, and takes the effective count from 20.05 to 11.37, a factor of 1.8 as well — but the threshold only moves from 3.024 to 2.848, which is 6%. The correction’s severity is a logarithmic function of the count, so halving the effective number of analyses buys back very little threshold.
That is the honest summary of what this measurement is worth, and it is less than it first looks. Against a correlation of 0.6 the corrected per-analysis level rises from 0.0025 to 0.0044 — a factor of 1.8 in the α and a factor of 1.06 in the z. A study that failed Bonferroni at 2.9 standard errors passes the exact correction; a study that failed at 2.5 fails both.
Where the correlation comes from
Twenty analyses correlated at 0.6 is not a hypothetical. It is what a family of analyses looks like when the analyses are variations on one question.
Twenty outcome measures on the same subjects share the subjects. Two scales measuring overlapping constructs on the same people correlate at whatever the constructs correlate at, which is routinely 0.5 to 0.8.
Twenty subgroup analyses overlap in membership. Analysing men, then over-fifties, then over-fifties who are men, produces statistics computed on partly the same observations, and the correlation is roughly the overlap — which is also why a subgroup result is not an independent replication of the main one.
Twenty exclusion rules — drop the outliers, drop the first week, drop the non-compliers — produce twenty datasets differing in a few per cent of their rows. Those statistics correlate at 0.9 and above, and twenty of them are worth about two.
Twenty covariate adjustments of the same regression, with nested covariate sets, are the extreme case: each is a small perturbation of the last.
The forking-paths construction uses twenty outcome variables sharing a common component, which puts it at a correlation near a quarter — the measurement here returns a 56.6% uncorrected rate at 0.25 against its 57%, which is the family reproduced with the correlation made explicit rather than implied.
Measuring it does not require knowing it
The table above sweeps a correlation, which raises the obvious objection: nobody knows the correlation between their twenty analyses.
They do not have to. The measurement never uses the correlation — it uses the null distribution of the largest statistic, and that distribution can be built from the data itself by permutation. Break the association between the exposure and the outcomes by shuffling the exposure column, recompute all twenty analyses on the shuffled data, keep the largest, and repeat. The resulting distribution carries whatever correlation structure the twenty analyses really have — including structure nobody could have written down, since the analyses are correlated through the observations rather than through a model.
That is the max-statistic permutation correction, and it has three properties worth stating together. It is the same device as the shuffled reference distribution a randomisation test uses, pointed at a maximum rather than at a single statistic.
It is exact rather than conservative, at the family-wise level, up to the number of permutations.
It requires no model of the dependence. A formula-based correction needs the correlation matrix of the statistics; the permutation needs only the ability to recompute the analyses.
It costs one recomputation of everything per permutation, which for twenty analyses and a thousand permutations is twenty thousand analyses. That is the whole of its cost, and it is the reason it is not the default: Bonferroni is a division.
What the correction is not
The exactness is about one quantity and it is worth naming which.
The family-wise error rate is the probability of at least one false positive among the twenty. It is the right target when a single false claim would be the damage — one drug approved, one association reported as the finding. It is the wrong target when twenty analyses are a screen, and the damage is proportional to how many of the reported hits are wrong. That is the false discovery rate, and it is a different quantity with different arithmetic.
Which rate a correction controls is the distinction that matters more than the choice of correction within a rate, and nothing on this page changes it: the measurement here makes the family-wise correction exact, and an exact correction to the wrong target is still the wrong target.
The other thing it is not is a repair to the analysis. The corrected p-value is honest about the family; the estimate reported beside it is not, and correcting the first makes the second worse — which is what the correction does to the estimate beside it.
The count is not proportional to the number of analyses either
Sweeping the number of analyses rather than the correlation gives the other half of the picture, and it is the half that decides whether the measurement is worth making on a given family.
At a correlation of 0.6, five analyses are worth 3.83, ten are worth 6.76, twenty are worth 11.37, forty are worth 18.90 and eighty are worth 31.44.
The ratio of effective to actual falls steadily — 77%, 68%, 57%, 47%, 39% — which is the useful shape. A shared component of fixed size explains a growing share of a growing family’s joint behaviour, so the larger the family, the more of it is redundant. The measurement is worth most exactly where Bonferroni is most severe, and that is a convenient rather than an accidental relationship: both are driven by the same quantity.
The practical threshold follows. On a family of five, Bonferroni divides by 5 and the exact correction by 3.83; the difference in the critical value is under 3%, and nobody should spend a thousand permutations on it. On a family of eighty, Bonferroni divides by 80 and the exact correction by 31.44, and the difference in the critical value is 8% — still not large, and now worth having, because eighty analyses of one dataset is a screen and a screen is run at the margin.
What a genomics-scale family does to the same arithmetic
The extreme case is worth following, because it is where the correction stops being a footnote.
A study testing a million positions across a genome has k = 10⁶ and Bonferroni’s threshold is 5 × 10⁻⁸, which is the number that field uses. The positions are correlated in blocks — neighbouring positions are inherited together — so the effective count is the number of independent blocks rather than the number of positions, and it is roughly a million in that setting because the blocks are small relative to the genome.
That is the reason the convention survived: in the one place where the family is large enough for the difference to matter, the correlation happens to be local and the effective count happens to be close to the actual one. It is a fact about that subject rather than about multiplicity, and importing the convention into a setting where twenty outcome measures correlate at 0.6 imports a threshold computed under an assumption that does not hold there.
The general rule is short. The effective count is a property of the family, and it has to be measured on the family. Neither the number of analyses nor a rule of thumb about them is a substitute, and the two families above — a million positions and twenty scales — differ by six orders of magnitude in size and by a factor of two in how much of their number is real.
What is claimed here, and what is not
Two statements bracket the measurement from both sides, and they are arithmetic rather than statistical.
With the analyses independent, the effective count is the count. It returns k to within 8%, which is the strongest available statement that the arithmetic is calibrated: under independence Bonferroni is exactly right, so a measurement disagreeing there would be wrong rather than informative. It is also what catches the commonest possible slip, which is building the null distribution of the largest signed statistic rather than the largest absolute one — that produces a critical value about 0.2 too low and an effective count about a third too small, at every correlation, including zero.
Correlated analyses are worth materially fewer than they number. Stated as at least a factor of two between no correlation and 0.95, which would not hold if the family had lost its shared component. That is a plausible slip, since a family drawn with the correlation applied to the wrong term is still twenty statistics with the right marginal distribution, and only the joint behaviour reveals it.
The reading that does not survive is Bonferroni applied to analyses that are all the same analysis. At a correlation of 0.995 the twenty statistics are one statistic and a family-wise correction has nothing to correct; the effective count is under three, against the twenty Bonferroni divides by. If that came back at twenty, the whole measurement would be returning its own input.
What it costs in power
The point of a less severe correction is power, so it is worth saying how much is recovered rather than leaving it as a direction.
Against an effect of two standard errors sitting in one of twenty analyses correlated at 0.6, the uncorrected 5% threshold detects it 51.5% of the time and the exact family-wise threshold 22.5%. Bonferroni’s threshold is 6% higher than the exact one and detects it a little less often still.
So the exact correction recovers a few percentage points of power out of the twenty-nine that the correction as a whole costs. That is the honest accounting: the expensive part is correcting at all, not correcting by the wrong number. Halving the effective count recovers almost nothing because the relationship between the divisor and the threshold is logarithmic, and the relationship between the threshold and the power is what a normal tail does at 2.8 standard errors, which is steep.
The consequence for practice is the opposite of what a table of effective counts suggests. A study worried about the power cost of multiplicity should reduce the family, not refine the divisor: naming three outcomes instead of twenty moves the threshold from 2.85 to 2.39 standard errors, which is worth several times what measuring the correlation among the twenty buys.
The cost of getting it wrong in the other direction
Everything above says Bonferroni is too severe on a correlated family. The symmetric error is worth the same attention, because it is the one that produces false claims rather than missed ones.
A correction computed for a family of twenty and applied to a family that was actually forty is too lenient by exactly the amount the table quantifies. That happens whenever the analyses reported are a subset of the analyses run — which is the original problem the forking paths are about, and which no correction can repair, because the permutation distribution can only be built over the analyses it is given.
So the ordering of the two errors is worth being explicit about. Over-correcting costs power and is visible: a study reports a threshold and a reader can see how severe it was. Under-correcting costs error control and is invisible: the reported threshold looks identical whether the family it was computed over was the family that was run.
The measurement on this page reduces the first cost and does nothing about the second. That is a smaller contribution than a first reading of the table suggests, and it is the honest one.
Still open: the family nobody can enumerate
The permutation correction needs a list of the analyses. A preregistered study has one. An exploratory one does not, and the analyses that were considered and not run belong in the family by the same logic that puts the ones that were run in it.
That is the garden of forking paths in its strict sense, and it is not a multiplicity problem of the kind anything here addresses: the multiplicity is counterfactual, so no correction can be applied after the fact and no permutation can enumerate it. What can be said is that the effective count for such a family is bounded below by the count for the analyses that were run, since those are a subset.
Whether that bound is useful — whether the analyses an author would have run on different data are much more numerous than the ones they ran, or only slightly — is an empirical question about how analysts behave rather than a question about statistics. It has an answer, it would require watching people analyse data they have not seen before, and nothing in the table above bears on it.
One thing can be said about the shape of that answer from here. The analyses an author would have run on different data are more correlated with the ones they did run than twenty arbitrary analyses would be, because they are all variations on the same idea pursued by the same person. So the counterfactual family is larger in number and smaller in effective count than a naive multiplication suggests, and the two errors point in opposite directions. Whether they cancel, and at what ratio, is the measurement — and until it exists, the honest description of a family-wise correction applied to an exploratory analysis is that it corrects for a family that is known to be incomplete, by an amount that is known to be approximately right for the part that is known.
That is not a reason to skip the correction. It is a reason to report the list of analyses the correction was computed over, which is the one piece of information that makes the number readable and which almost no paper carries — the same omission the winner’s curse depends on and the same one a stopping rule depends on.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An order that spends the error rate — both name bonferroni, false positive, familywise error rate, multiple comparisons
- False discoveries that arrive together — both name correlation, false positive, familywise error rate, multiple comparisons
- One control, many arms — both name bonferroni, correlation, familywise error rate, multiple comparisons
- Eight forecasters and one benchmark — both name bonferroni, familywise error rate, multiple comparisons
- Estimating how many nulls are true — both name correlation, multiple comparisons, p-value
- The price of control — both name bonferroni, false positive, multiple comparisons
Named objects
A flat tag is an object no other essay names yet.
BonferroniCorrelationFalse positiveFamilywise error rateThe garden of forking pathsMultiple comparisonsp-valuePermutation