Multiple comparisons — where it appears
Named by 25 essays across 12 fields — each of them below, with the objects they name alongside it.
Two searches that share nothing
Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.
What the correction corrects
Twenty tests of true nulls produce at least one false positive 64% of the time, and the closed form and the count agree. Bonferroni holds it at 5% and Holm holds it at 5% while finding more. Nobody should still be using Bonferroni.
How many analyses there really were
Bonferroni divides by twenty because twenty analyses were run. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, and the threshold that controls exactly the same error rate is measurable rather than assumed.
The band the eye was standing in for
The confidence band software draws on a quantile plot holds each point at 95%, and a genuinely normal sample of forty has forty chances to leave it — so 45.0% of them do. The band a reader is actually using is that one widened by a factor of 1.502, and nothing draws it.
Eight forecasters and one benchmark
A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.
Two different promises
Bonferroni bounds the chance of any false positive. Benjamini–Hochberg bounds the share of the findings that are false. Both are called correcting for multiple comparisons, and one of them lets the familywise rate reach 20%.
What naming it in advance costs
Preregistration is argued for as free. Against an effect of two standard errors hiding in one of twenty analyses, naming the right one detects it 51.5% of the time and naming the wrong one detects it 4.7% of the time; correcting all twenty detects it 22.5% wherever it is. The two are worth the same when the chance of having named correctly is 38%.
Dropping the losers
Carrying the best of eight arms forward and testing it at 1.96 rejects a true null 10.3% of the time — the hypothesis was chosen by looking at the data, so the statistic is a maximum wearing a single comparison's clothes. The value that holds the rate is 2.313, and it has to be solved for.
One control, many arms
The control appears in every comparison, so it is worth √k treatment arms — and the same sharing makes the k tests correlated at n/(n+n₀), which is the quantity Bonferroni ignores. Both facts come out of one design decision, and it is the size of the control.
The models that were never in the running
A reference distribution for a set has to assume something about every candidate in it. Assuming that all of them are as good as the benchmark is what makes the reality check honest, and it is what sixteen hopeless candidates use to destroy it.
The price of control
Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.
Three quarters of the way to one search
The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.
The interval at the end of the curve
The interval most software prints around a survival curve covers 89.7% at five years, where 3.3 of forty subjects are still being watched and where the curve is actually read. The same variance carried on a log–log scale covers 94.8% there — and the failure was never the width.
A coverage table with its own error
Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.
The correction that makes the estimate worse
Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.
Significant in one, not in the other
Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.
Twenty analyses of nothing
Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.
False discoveries that arrive together
Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.
Estimating how many nulls are true
Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.
The rank is a decision
The sequential procedure's 5% bounds one of its two errors. Over-counting reads between 4.2% and 7.2% at every sample length from fifty observations to three hundred; under-counting reads 69.5% at fifty and 0.0% at three hundred, and nothing in the procedure bounds it.
Two ways to combine p-values
Fisher's and Stouffer's combinations are both exactly right when every null is true, for the single reason that each p-value is flat. Under a real effect they disagree about which evidence counts: with Stouffer held at 50% power across ten studies, Fisher is the more powerful while the signal sits in six or fewer of them and the less powerful from seven.
An order that spends the error rate
Test twenty hypotheses in a declared order, each at the full 5% and each only if every one before it was rejected, and the first is found 85.3% of the time where Holm finds it 52.5%. The tenth is found 20.4% of the time, the product of the powers before it. Move one true null to the head of the list and every real effect behind it is found no more than 4.3% of the time.
Intervals for the findings
Benjamini–Hochberg's findings usually go out each with its ordinary 95% interval. With ten real effects of two standard errors among twenty tests, 11.59% of those intervals miss their effect, every miss on the far side, and the interval around the most prominent finding covers 72.36% of the time — 2.38% when the effects are one standard error. Intervals widened for the number of findings hold the share that miss under 5%.
The smallest of three combinations
Reporting whichever of Fisher's, Stouffer's and Tippett's combinations is smallest is a test of its own, and on ten studies of nothing it rejects 9.66% of the time — not 5%, and nowhere near the 15% the three sizes add to, because the statistics are correlated at up to 0.903. Read at 2.448% each it is exact, and then it trails the best single combination by at most 7.45 points and leads the worst by at least 10.30.
A horizon chosen after looking
A difference in restricted mean survival read at whichever of eleven horizons looks most convincing rejects 11.24% of trials in which the treatment does nothing, against 4.70% at a horizon fixed in advance. The correlation of the differences across horizons is closed, and the Gaussian process it defines prices the choice at a critical value of 2.317 — which brings the counted size back to 4.99% and keeps 96.92% of the power that a horizon nobody could have known to fix would have had.
Named alongside it
The objects these essays reach for when they reach for this one.
Statistical powerMonte CarloFamilywise error rateBonferroniError rateFalse positiveCorrelationp-valueFalse discovery rateHolmBenjamini–HochbergClosed form