False positive — where it appears
Named by 14 essays across 6 fields — each of them below, with the objects they name alongside it.
What the correction corrects
Twenty tests of true nulls produce at least one false positive 64% of the time, and the closed form and the count agree. Bonferroni holds it at 5% and Holm holds it at 5% while finding more. Nobody should still be using Bonferroni.
How many analyses there really were
Bonferroni divides by twenty because twenty analyses were run. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, and the threshold that controls exactly the same error rate is measurable rather than assumed.
The test with no table
The statistic that separates a real long-run relation from a spurious one is computed as a t and is not a t. At two hundred observations its 5% point is −3.38 where the t table says −1.65, and reading it against the table calls two unrelated random walks cointegrated 70.5% of the time.
Two different promises
Bonferroni bounds the chance of any false positive. Benjamini–Hochberg bounds the share of the findings that are false. Both are called correcting for multiple comparisons, and one of them lets the familywise rate reach 20%.
What a p-value does not say
The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.
What a positive test is worth
A test that is 90% sensitive and 95% specific sounds accurate. For a condition affecting one person in a thousand, 98% of its positive results are wrong, and a worse test on a commoner condition beats a better test on a rare one.
What naming it in advance costs
Preregistration is argued for as free. Against an effect of two standard errors hiding in one of twenty analyses, naming the right one detects it 51.5% of the time and naming the wrong one detects it 4.7% of the time; correcting all twenty detects it 22.5% wherever it is. The two are worth the same when the chance of having named correctly is 38%.
The price of control
Every correction is paid for in power, and the exchange rate can be measured. Holm buys familywise control for 33 percentage points of power; Benjamini–Hochberg buys a weaker guarantee for 10. Neither is free and neither is a matter of taste.
What differencing costs
Differencing takes the false-positive rate between two unrelated walks from 76.7% to 4.9%, and takes a genuine relationship's R² from 0.91 to 0.33. Applied to a series that did not need it, it doubles the variance and installs a correlation of −0.5 that the data never had.
Twenty analyses of nothing
Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.
False discoveries that arrive together
Correlate twenty tests and Benjamini–Hochberg still holds its false discovery rate — 1.66% at a correlation of 0.9 with ten real effects, against 2.55% when the tests are independent. What changes is how the errors come. A family of true nulls reports anything 2.34% of the time instead of 5.08%, and when it does, it reports 16.56 false findings out of twenty.
The cliff that is a slope
A regression between two independent series is called significant 4.9% of the time at no persistence, 52.4% at a lag-one correlation of 0.9, and 83.4% at a unit root. The rule the field offers asks whether the last of those holds, and at 0.9 the unit-root test correctly refuses one 87.2% of the time.
An order that spends the error rate
Test twenty hypotheses in a declared order, each at the full 5% and each only if every one before it was rejected, and the first is found 85.3% of the time where Holm finds it 52.5%. The tenth is found 20.4% of the time, the product of the powers before it. Move one true null to the head of the list and every real effect behind it is found no more than 4.3% of the time.
The repair that keeps the question
A regression between two independent trending series is significant 82.9% of the time on random walks and 100.0% on trend-stationary ones. Subtracting a fitted line leaves 74.2% and 33.5%; differencing leaves 5.0% and 5.2% and throws away the trend the study was about.
Named alongside it
The objects these essays reach for when they reach for this one.
Multiple comparisonsStatistical powerBonferroniFamilywise error rateHolmCorrelationDependenceError rateMonte Carlop-valueAutocorrelationClosed form