Concept

Bonferroni — where it appears

Dividing the significance level by the number of comparisons, which controls the chance of any false positive and prices comparisons that are independent. It is wrong in both directions on the same arithmetic: far too strong where the comparisons are dependent, and not strong enough where the excess is a shift in the statistic's mean rather than a maximum over many.

Named by 17 essays across 10 fields — each of them below, with the objects they name alongside it.

What each variant loses before anything has been searched for. The mean loss differential of each of the eight variants against the benchmark, over 600 tables of 60 origins, with every fit given 71 rows. The series is an AR(1) and every variant adds a lag whose coefficient is zero, so in population the two forecasts are the same forecast and the difference drawn here is estimation noise and nothing else. The marked line is σ²(q₁ − q₀)/n = -0.01408, which is an expression in how many coefficients each model has and how many rows it was fitted on — it knows nothing about the series, the persistence or which lag the variant added, and every bar is within a fifth of it. This is the amount a reference distribution recentred at each column's own sample mean believes the candidates are already behind by.

A table of nested models

A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.

search · Forecast
The familywise error rate with no correction, α = 0.05. Two routes: the curve is 1 − (1 − α)^m and the points are counted over 6,000 simulated families of true nulls. With twenty tests the chance of at least one false positive is 64.1%.

What the correction corrects

Twenty tests of true nulls produce at least one false positive 64% of the time, and the closed form and the count agree. Bonferroni holds it at 5% and Holm holds it at 5% while finding more. Nobody should still be using Bonferroni.

multiplicity · Multiplicity
One true null, one table, five readings. every subset of four, fifteen models, at a null where nothing any candidate holds is worth anything, over 500 draws. Each bar is the share of draws on which that reading declares a difference at a nominal 5%. The reading is the whole of the difference between the bars: the data is identical. An open search over all 210 ordered pairs rejects 76.2%; the table's own 5% point is 3.163 against the 1.671 a single comparison uses. Bonferroni takes the open reading to 0.6% — and on the nested ladder the same correction does not reach the nominal level at all, because there the excess is a shift in the mean rather than a maximum over many.

When the benchmark is a candidate

A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.

select · Forecast
What 20 analyses of one dataset are worth. The threshold giving a 5% family-wise error rate, read back as a number of independent analyses. At no correlation it is 20.05; at 0.6 it is 11.37; at 0.95 it is 2.58. Bonferroni divides by 20 throughout.

How many analyses there really were

Bonferroni divides by twenty because twenty analyses were run. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, and the threshold that controls exactly the same error rate is measurable rather than assumed.

paths · Forking
The distribution of the largest statistic in the table. Fit the benchmark to the whole series, resample its residuals, simulate 199 series in which the null is true by construction, re-run the entire eight-variant search on each, and keep the largest statistic. That is the distribution drawn here, and it is the distribution of the thing a specification search actually reports. It is centred at 1.045 — the maximum of eight statistics is not centred at zero however well each of them behaves — and its 5% point is 2.536. A table read against 1.671 is reading the distribution of one statistic; a Bonferroni correction reads it against 2.577 and is nearly right here, because eight variants that each add a different lag are nearly eight separate chances.

A null with a model in it

The distribution to read the winner of a table against cannot be resampled from the data, because the data does not contain the null. It has to be generated from a model — which is the assumption the resampling was chosen to avoid.

search · Bootstrap
Eight candidates, one of them exactly as good as the benchmark. The candidate set: moving averages of the last 1, 2, 3, 5, 8, 13, 21 and 34 observations, each drawn as its expected squared error divided by the benchmark's — the mean of all 60. The persistence is not chosen, it is solved for: at φ = 0.4895 the best candidate in the set, the average of 2, has exactly the benchmark's expected squared error, and every other candidate is worse by between 0.5% and 5.1%. So the null that no candidate beats the benchmark is true, with one candidate on its boundary. Everything a set comparison claims about its own error rate has to be measured here, because anywhere further inside the null every procedure flatters itself.

Eight forecasters and one benchmark

A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.

ranking · Multiplicity
Every candidate is behind by what its parameter count says. Each dot is one of the fifteen subsets of four predictors, fitted on a rolling window of 80 rows and scored against the benchmark out of sample over 60 origins, at a null where every one of them contains the truth. The line is σ²(q₀/(R − q₀ − 1) − q/(R − q − 1)), which is arithmetic on two integers and a window length. Most of these pairs are not nested — a subset of two predictors and a different subset of two share neither model — and the closed form does not care: the displacement is a statement about how many coefficients each side estimates. The candidates of the benchmark's own dimension sit at zero.

The displacement is a parameter count

A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.

select · Multiplicity
Twenty tests, 10 of them real — what each procedure holds. no correction: familywise 40.8%, false discovery 5.3%, power 85%. Bonferroni: familywise 2.8%, false discovery 0.5%, power 49%. Holm: familywise 3.6%, false discovery 0.6%, power 53%. Benjamini–Hochberg: familywise 20.0%, false discovery 2.6%, power 75%.

Two different promises

Bonferroni bounds the chance of any false positive. Benjamini–Hochberg bounds the share of the findings that are false. Both are called correcting for multiple comparisons, and one of them lets the familywise rate reach 20%.

multiplicity · Multiplicity
3 arms against one control, 360 units in all. Every control size, enumerated. The best is 132 on the control and 76 on each arm — a ratio of 1.74, against √3 = 1.73. Splitting the units evenly over all 4 groups costs 7.2%, which is small; what the larger control also does is lower the correlation between the comparisons, from 0.50 to 0.37, and that changes which multiplicity correction is right.

One control, many arms

The control appears in every comparison, so it is worth √k treatment arms — and the same sharing makes the k tests correlated at n/(n+n₀), which is the quantity Bonferroni ignores. Both facts come out of one design decision, and it is the size of the control.

allocation · Multiplicity
Power to find a real effect of 3 standard errors, 10 of 20 real. no correction finds 85.1%, Bonferroni finds 49.1%, Holm finds 52.5%, Benjamini–Hochberg finds 74.9%. The uncorrected procedure finds the most and controls nothing.

The price of control

Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.

multiplicity · Multiplicity
Twenty cells of an interval that is exactly 95%, 1,000 replications each. The t interval covers exactly 95% in every cell. Estimated at 1,000 replications its cells read 93.9% to 96.5%, and 2 of the twenty are flagged by their own ±1.96 standard errors.

A coverage table with its own error

Twenty cells estimating the coverage of an interval that is exactly 95%, at a thousand replications each, read from 93.9% to 96.5% — and a table like that flags at least one of its correct cells on 69.9% of honest runs. Ten times the replications does not repair it: at ten thousand the same table still flags one 63.3% of the time.

method · Seeds
What the family-wise correction does to the effect it lets through. At two standard errors the estimate that clears an uncorrected 5% threshold averages 1.35 times the truth, and the one that clears the family-wise threshold averages 1.69 times it. The correction fixes the error rate by demanding a larger estimate, and a larger estimate is a more selected one.

The correction that makes the estimate worse

Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.

paths · Forking
The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

testing · Forking
Eight encompassing nulls, all of them true at once. The forecast under test is the variance-minimising combination of the eight candidates, and the first-order condition that defines its weights is cov(e_c, e_j) = var(e_c) for every j — so every one of the eight nulls is exactly true simultaneously and every rejection counted here is false. 800 draws of 60 origins. The largest of the eight statistics rejects 25.0% of the time at a nominal 5% and one comparison stated in advance rejects 3.1%, which makes the set worth about 9.1 independent comparisons — nearly the eight it has. Bonferroni, which is far inside its level on a search over accuracy comparisons of the same eight forecasters, is at 4.8% here.

When every null is true

A reality check assumes that every candidate in the set is exactly as good as the benchmark, which is a configuration nobody's data is ever in. Test a combination against its own parts and that configuration is not assumed — it is what the arithmetic makes true.

search · Multiplicity
The factor that holds all of the next m observations with 95% probability, from a sample of 10. From 10 observations, the band for one future value has factor 2.371, for ten 3.716, for a hundred 4.942 and for a thousand 6.008. The 95%-content tolerance factor is 3.382 and is passed by m = 10; the Bonferroni stretch of the prediction factor reaches 7.567 at a thousand.

All of the next ten

A warranty, a batch release or a monitoring rule promises something about every one of the next ten observations, not about one. From a sample of ten, the band that holds all ten with 95% probability reaches 3.716 sample standard deviations either side of the mean — already wider than the 3.382 of a tolerance interval for 95% of the population — and it keeps widening: 4.942 for a hundred, 6.008 for a thousand, with no ceiling. A 95% prediction interval, read as the answer, holds all ten 67.9% of the time: more than 0.95 to the tenth power, because the ten succeed and fail together.

estimated · Bands
Two far rows, and the line with one of them deleted. Twenty clean points and two rows near x = 9. The slope is −0.511 with every row, −0.376 with one far row deleted, and 0.495 with both deleted. Deleting one of them barely moves the line, because the other is still there.

Two points that hide each other

One far observation among twenty-one has a Cook's distance of 24.1. Put a second beside it and the two read 0.966 and 0.772, neither crossing 1, while together they reverse the slope and deleting both moves the fit by 53.3.

regression · Leverage
Twenty hypotheses tested in a declared order, the ten real effects listed first. Effects of three standard errors, ten real, familywise 5%. fixed sequence: 85.3% at position 1, 45.0% at 5, 20.4% at 10; overall power 46.10%; fallback: 49.1% at position 1, 56.4% at 5, 57.9% at 10; overall power 55.87%; Holm: 52.5% at position 1, 52.2% at 5, 52.5% at 10; overall power 52.53%.

An order that spends the error rate

Test twenty hypotheses in a declared order, each at the full 5% and each only if every one before it was rejected, and the first is found 85.3% of the time where Holm finds it 52.5%. The tenth is found 20.4% of the time, the product of the powers before it. Move one true null to the head of the list and every real effect behind it is found no more than 4.3% of the time.

multiplicity · Multiplicity

Named alongside it

The objects these essays reach for when they reach for this one.

Multiple comparisonsFamilywise error rateFalse positiveBenchmark forecastError rateMean squared errorMonte CarloNull hypothesisStatistical powerClosed formHolmMultiplicity

All concepts