Concept

Multiple comparisons — where it appears

Asking more than one question of the same data, so that the chance of some answer looking extreme is larger than the chance for any one of them. How much larger depends on how much the questions overlap, which is a property of a table nobody usually writes down.

Named by 25 essays across 12 fields — each of them below, with the objects they name alongside it.

How much of one search the other has already found. Five pairs of searches on one sample, on a scale whose zero and one are both fixed by construction. Zero is two searches over disjoint sets of independent columns: they remove shares of the residual sum that add, at 0.8 standard errors from exactly additive, and they read 0.004. One is a break search paired with a step column it contains, which reads exactly one on every draw because the step adds nothing at all. Between them: two dictionaries of step columns cut a few rows apart read 0.125, and the pair the earlier field measured — a break and a whitening window, both reading the same residual series — reads 0.762, three quarters of the way to one search containing the other. And below zero, a break paired with a search over independent columns reads -0.306: the joint search finds configurations neither half of it contains, so charging the two separately under-charges.

Two searches that share nothing

Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.

apart · Criterion
The familywise error rate with no correction, α = 0.05. Two routes: the curve is 1 − (1 − α)^m and the points are counted over 6,000 simulated families of true nulls. With twenty tests the chance of at least one false positive is 64.1%.

What the correction corrects

Twenty tests of true nulls produce at least one false positive 64% of the time, and the closed form and the count agree. Bonferroni holds it at 5% and Holm holds it at 5% while finding more. Nobody should still be using Bonferroni.

multiplicity · Multiplicity
What 20 analyses of one dataset are worth. The threshold giving a 5% family-wise error rate, read back as a number of independent analyses. At no correlation it is 20.05; at 0.6 it is 11.37; at 0.95 it is 2.58. Bonferroni divides by 20 throughout.

How many analyses there really were

Bonferroni divides by twenty because twenty analyses were run. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, and the threshold that controls exactly the same error rate is measurable rather than assumed.

paths · Forking
Two 95% bands for a quantile plot of 40 points. The outer band is left by 5% of genuinely normal samples — which is what a reader is using a band for. The inner one holds each point separately at 95%, which is what software draws, and 45.0% of genuinely normal samples step outside it. The outer is the inner widened by a factor of 1.502.

The band the eye was standing in for

The confidence band software draws on a quantile plot holds each point at 95%, and a genuinely normal sample of forty has forty chances to leave it — so 45.0% of them do. The band a reader is actually using is that one widened by a factor of 1.502, and nothing draws it.

lineup · Qq
Eight candidates, one of them exactly as good as the benchmark. The candidate set: moving averages of the last 1, 2, 3, 5, 8, 13, 21 and 34 observations, each drawn as its expected squared error divided by the benchmark's — the mean of all 60. The persistence is not chosen, it is solved for: at φ = 0.4895 the best candidate in the set, the average of 2, has exactly the benchmark's expected squared error, and every other candidate is worse by between 0.5% and 5.1%. So the null that no candidate beats the benchmark is true, with one candidate on its boundary. Everything a set comparison claims about its own error rate has to be measured here, because anywhere further inside the null every procedure flatters itself.

Eight forecasters and one benchmark

A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.

ranking · Multiplicity
Twenty tests, 10 of them real — what each procedure holds. no correction: familywise 40.8%, false discovery 5.3%, power 85%. Bonferroni: familywise 2.8%, false discovery 0.5%, power 49%. Holm: familywise 3.6%, false discovery 0.6%, power 53%. Benjamini–Hochberg: familywise 20.0%, false discovery 2.6%, power 75%.

Two different promises

Bonferroni bounds the chance of any false positive. Benjamini–Hochberg bounds the share of the findings that are false. Both are called correcting for multiple comparisons, and one of them lets the familywise rate reach 20%.

multiplicity · Multiplicity
Naming the analysis in advance, against correcting for all 20 of them. The prespecified analysis detects an effect that is in it 52% of the time at two standard errors and an effect elsewhere 5% of the time. The corrected slate detects it 23% of the time wherever it is. The two are worth the same when the chance of having named the right analysis is 38% — and that figure rises to 91% at four standard errors.

What naming it in advance costs

Preregistration is argued for as free. Against an effect of two standard errors hiding in one of twenty analyses, naming the right one detects it 51.5% of the time and naming the wrong one detects it 4.7% of the time; correcting all twenty detects it 22.5% wherever it is. The two are worth the same when the chance of having named correctly is 38%.

paths · Forking
The best of 8 arms, tested as though it were the only one. 8,000 trials with no effect in any arm. Stage one runs 8 arms at 60 each, the best is carried forward, and stage two adds 30 more to it and to the control. The histogram is where the final statistic lands and the curve is the standard normal it is being read against — shifted right, because the arm was chosen for being ahead. 10.5% of these trials clear 1.96 against a claimed 5%, and the value that actually holds the rate for this design is 2.313.

Dropping the losers

Carrying the best of eight arms forward and testing it at 1.96 rejects a true null 10.3% of the time — the hypothesis was chosen by looking at the data, so the statistic is a maximum wearing a single comparison's clothes. The value that holds the rate is 2.313, and it has to be solved for.

adaptive · Stopping
3 arms against one control, 360 units in all. Every control size, enumerated. The best is 132 on the control and 76 on each arm — a ratio of 1.74, against √3 = 1.73. Splitting the units evenly over all 4 groups costs 7.2%, which is small; what the larger control also does is lower the correlation between the comparisons, from 0.50 to 0.37, and that changes which multiplicity correction is right.

One control, many arms

The control appears in every comparison, so it is worth √k treatment arms — and the same sharing makes the k tests correlated at n/(n+n₀), which is the quantity Bonferroni ignores. Both facts come out of one design decision, and it is the size of the control.

allocation · Multiplicity
Sixteen candidates nobody would have run, and what they cost. At φ = 0.65 one candidate in the original set is genuinely better than the benchmark, and the question is how often each procedure finds it. The added candidates are stale copies of the last value — read two, four, six … steps late — every one of them worse than the benchmark by at least 43%, and not one of them is ever the best candidate in a sample. The reality check goes from 35.5% to 0.0% as they are added, because its reference distribution has to assume every candidate is exactly as good as the benchmark and sixteen such assumptions is a critical value nothing reaches. The recentred version, which drops from the recentring the candidates the data has already ruled out — 15.1 of 24 of them — goes from 25.5% to 24.5%. The top line never moves: reporting the winner's own p-value cannot notice a change to a set it never looks at.

The models that were never in the running

A reference distribution for a set has to assume something about every candidate in it. Assuming that all of them are as good as the benchmark is what makes the reality check honest, and it is what sixteen hopeless candidates use to destroy it.

ranking · Multiplicity
Power to find a real effect of 3 standard errors, 10 of 20 real. no correction finds 85.1%, Bonferroni finds 49.1%, Holm finds 52.5%, Benjamini–Hochberg finds 74.9%. The uncorrected procedure finds the most and controls nothing.

The price of control

Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.

multiplicity · Multiplicity
The ladder is the same ladder under every law. Each pair's overlap under each of the four laws, over 400 draws apiece. The scale is fixed at both ends by construction: two searches over disjoint sets of independent columns read -0.001, 0.013, -0.002, 0.004, and a break search paired with a step column it already contains reads exactly one under every law. Between them the pair that reads one residual series twice runs 0.763, 0.805, 0.577, 0.752 — lowest under long memory, where a whitening has most to do and the break search has least left to find that the whitening has not taken. The rung that moves most is the pair of step dictionaries, from -0.148 under a moving average to 0.419 under a break; and the pair that is negative is negative under all four.

Three quarters of the way to one search

The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.

apart · Criterion
One cohort of 40, and two intervals around the end of its curve. A single simulated study of 40 subjects with exponential survival at rate 0.35, dropout at rate 0.15 and follow-up to 6 — the first seed from 8811 upward whose plain band reaches below −0.05, chosen to show the failure rather than its frequency. The step curve is Kaplan–Meier and the smooth curve the truth. The plain band, the estimate plus and minus 1.96 Greenwood standard errors, first dips below zero at t = 3.78 and reaches −0.052; early on it also rises to 1.023, above one. At t = 5 the estimate is 0.069 with 1 subject still under observation, the plain interval runs from −0.052 to 0.191 and the log-log interval from 0.006 to 0.251, against a truth of 0.174. The log-log band is built on a scale that cannot leave [0, 1], and it bends away from the edge rather than through it.

The interval at the end of the curve

The interval most software prints around a survival curve covers 89.7% at five years, where 3.3 of forty subjects are still being watched and where the curve is actually read. The same variance carried on a log–log scale covers 94.8% there — and the failure was never the width.

survival · Censoring
Twenty cells of an interval that is exactly 95%, 1,000 replications each. The t interval covers exactly 95% in every cell. Estimated at 1,000 replications its cells read 93.9% to 96.5%, and 2 of the twenty are flagged by their own ±1.96 standard errors.

A coverage table with its own error

Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.

method · Seeds
What the family-wise correction does to the effect it lets through. At two standard errors the estimate that clears an uncorrected 5% threshold averages 1.35 times the truth, and the one that clears the family-wise threshold averages 1.69 times it. The correction fixes the error rate by demanding a larger estimate, and a larger estimate is a more selected one.

The correction that makes the estimate worse

Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.

paths · Forking
Two studies of the same effect, z statistics with mean 1.96: where one is significant and the other is not. Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.

Significant in one, not in the other

Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.

alongside · Repetition
The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

testing · Forking
The false discovery rate of twenty correlated tests, against the correlation. BH, every null true: 5.08% at 0, 4.86% at 0.3, 3.70% at 0.6, 2.34% at 0.9. BH, 10 of 20 real: 2.55% at 0, 2.53% at 0.3, 2.26% at 0.6, 1.66% at 0.9. BY, every null true: 1.46% at 0, 1.31% at 0.3, 1.03% at 0.6, 0.69% at 0.9. BY, 10 of 20 real: 0.72% at 0, 0.75% at 0.3, 0.64% at 0.6, 0.50% at 0.9. 20,000 families at each correlation.

False discoveries that arrive together

Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.

multiplicity · Multiplicity
Storey's estimate of the share of true nulls over twenty thousand families, independent and correlated at 0.6. The true share is 0.5. Independent tests: mean 0.610, spread 0.160, below half the truth in 0.92% of families. Correlated at 0.6: mean 0.609, spread 0.240, below half the truth in 9.33%.

Estimating how many nulls are true

Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.

multiplicity · Multiplicity
The bounded error and the unbounded one. How the sequential trace procedure's answer is distributed, against the sample length, for a three-series system with 2 genuine relations. Over-counting — claiming a stationary combination that is a random walk — reads 4.9%, 7.2%, 5.7%, 6.2%, 5.9%, 4.2% across the six lengths, never far from the 5% of a single test. Under-counting reads 69.5%, 40.2%, 14.0%, 0.5%, 0.0%, 0.0%. The procedure is described as a 5% rule and the 5% applies to one of those columns.

The rank is a decision

The sequential procedure's 5% bounds one of its two errors. Over-counting reads between 4.2% and 7.2% at every sample length from fifty observations to three hundred; under-counting reads 69.5% at fifty and 0.0% at three hundred, and nothing in the procedure bounds it.

systems · Rank
Three ways to reject with two studies, drawn where the two z statistics live. Two one-sided studies, each summarised by its z statistic. Fisher's combination rejects outside a curve that runs parallel to both axes, so one study past z = 2.378 decides it alone; Stouffer's rejects above the straight line z₁ + z₂ = 2.326; Tippett's rejects when either z passes 1.955. Each region holds exactly 5% of the standard bivariate normal — Fisher's in closed form, e^(−c/2)(1 + c/2) at c = 9.488 — and of 100,000 counted null pairs they catch 4.95%, 4.88% and 5.04%. Two alternatives carry the same Stouffer evidence: one study at 2.326 and the other at nothing, where the powers are 62.7%, 50.0% and 65.4%; and both at 1.163, where they are 47.7%, 50.0% and 38.3%.

Two ways to combine p-values

Fisher's and Stouffer's combinations are both exactly right when every null is true, for the single reason that each p-value is flat. Under a real effect they disagree about which evidence counts: with Stouffer held at 50% power across ten studies, Fisher is the more powerful while the signal sits in six or fewer of them and the less powerful from seven.

testing · Uniformity
Twenty hypotheses tested in a declared order, the ten real effects listed first. Effects of three standard errors, ten real, familywise 5%. fixed sequence: 85.3% at position 1, 45.0% at 5, 20.4% at 10; overall power 46.10%; fallback: 49.1% at position 1, 56.4% at 5, 57.9% at 10; overall power 55.87%; Holm: 52.5% at position 1, 52.2% at 5, 52.5% at 10; overall power 52.53%.

An order that spends the error rate

Test twenty hypotheses in a declared order, each at the full 5% and each only if every one before it was rejected, and the first is found 85.3% of the time where Holm finds it 52.5%. The tenth is found 20.4% of the time, the product of the powers before it. Move one true null to the head of the list and every real effect behind it is found no more than 4.3% of the time.

multiplicity · Multiplicity
Every finding Benjamini–Hochberg made in thirty families, with its interval, at real effects of 2. 85 findings, sorted by their estimate. 12 of their ordinary 95% intervals miss the true effect, every one of them on the far side; 6 of the wider false-coverage-rate intervals miss.

Intervals for the findings

Benjamini–Hochberg's findings usually go out each with its ordinary 95% interval. With ten real effects of two standard errors among twenty tests, 11.59% of those intervals miss their effect, every miss on the far side, and the interval around the most prominent finding covers 72.36% of the time — 2.38% when the effects are one standard error. Intervals widened for the number of findings hold the share that miss under 5%.

multiplicity · Multiplicity
How often each combination, and each union of them, rejects ten studies of nothing. Each combination alone rejects exactly 5% of null sets. Counted on 1,000,000 sets of ten null studies: Fisher or Stouffer 6.63%, Fisher or Tippett 8.05%, Stouffer or Tippett 8.96%, any of the three 9.66% — enclosed on a two-dimensional lattice between 9.18% and 10.09% — and all three together 1.05%. The three sizes add to 15%.

The smallest of three combinations

Reporting whichever of Fisher's, Stouffer's and Tippett's combinations is smallest is a test of its own, and on ten studies of nothing it rejects 9.66% of the time — not 5%, and nowhere near the 15% the three sizes add to, because the statistics are correlated at up to 0.903. Read at 2.448% each it is exact, and then it trails the best single combination by at most 7.45 points and leads the worst by at least 10.30.

testing · Uniformity
The difference in restricted mean survival at every horizon, in three worlds. Treatment minus control, in closed form, with dropout irrelevant to the truth. The proportional treatment's difference grows to 0.4766 at τ = 3 and the waning treatment's to 0.2675. The crossing treatment's rises to 0.1776 at τ = 2, near where the two survival curves cross, and falls back to 0.1366 at τ = 3. The ticks along the bottom are the eleven horizons, from 0.5 to 3 in quarters, at which a trial below reads its differences.

A horizon chosen after looking

A difference in restricted mean survival read at whichever of eleven horizons looks most convincing rejects 11.24% of trials in which the treatment does nothing, against 4.70% at a horizon fixed in advance. The correlation of the differences across horizons is closed, and the Gaussian process it defines prices the choice at a critical value of 2.317 — which brings the counted size back to 4.99% and keeps 96.92% of the power that a horizon nobody could have known to fix would have had.

survival · Censoring

Named alongside it

The objects these essays reach for when they reach for this one.

Statistical powerMonte CarloFamilywise error rateBonferroniError rateFalse positiveCorrelationp-valueFalse discovery rateHolmBenjamini–HochbergClosed form

All concepts