Corrections, and what each controls

Estimating how many nulls are true

Benjamini–Hochberg at 5% delivers 2.55% when half of twenty nulls are false, because it cannot tell how many are. Storey's estimate of that share, read off the p-values above one half, spends the rest and finds 81.93% of the real effects instead of 74.70% on independent tests. Correlated at 0.9, the same procedure reports a finding in 19.29% of families in which every null is true.

Worth reading first: What the correction corrects.

Twenty tests, ten of them real effects of three standard errors, Benjamini–Hochberg at 5%: the false discovery rate delivered is 2.55%, and the same procedure on correlated tests fell short of 5% by the same margin or more at every correlation. The shortfall is not an accident of the setting. The procedure’s bound is q times the share of nulls that are true, and it spends q as though every null were true, because it has no way to know otherwise.

It has a way to estimate. A p-value from a true null is uniform, so it lands above one half half the time; a real effect of three standard errors almost never does. Count the p-values above one half and double the count, and the result estimates how many nulls are true. Storey’s procedure adds one to the count, divides by twenty, and runs Benjamini–Hochberg at q divided by that estimate. When half the nulls are false it should, roughly, run at twice the level and spend the half the fixed procedure leaves on the table.

Storey's estimate of the share of true nulls over twenty thousand families, independent and correlated at 0.6. The true share is 0.5. Independent tests: mean 0.610, spread 0.160, below half the truth in 0.92% of families. Correlated at 0.6: mean 0.609, spread 0.240, below half the truth in 9.33%.
Fig. 1 Storey’s estimate of the share of true nulls over twenty thousand families of twenty tests with ten real effects, with independent tests and correlated at 0.6. The true share is one half.

The estimate, family by family

With ten of twenty nulls true, the estimate averages 0.610 over twenty thousand families of independent tests, against a true share of 0.5. The excess is deliberate. The one added to the count pushes it up by a tenth, and an estimate that errs high spends less than it could rather than more than it may. Its spread across families is 0.160, and it falls below half the truth — below a quarter — in 0.92% of families.

That is the whole mechanism, and it depends on the twenty p-values being twenty separate readings of whether each null is true. Correlated at 0.6, the estimate still averages 0.609 — the centre does not move — while its spread widens to 0.240, and it falls below half the truth in 9.33% of families, ten times as often. The average is right, and the family in hand is less and less likely to be the average.

Everything below follows from those two histograms, and it is worth saying in advance which way. The procedure divides the level by the estimate. A family whose estimate is too low gets a threshold that is too generous, and under correlation the families whose estimates are too low are not a random tenth of families.

What the estimate buys

What estimating the share of true nulls buys, against how many effects are real. Storey: 55.36% with 2 real, 77.10% with 8 real, 89.59% with 14 real. Benjamini–Hochberg: 54.97% with 2 real, 71.61% with 8 real, 79.07% with 14 real.
Fig. 2 The share of real effects found by Storey’s procedure and by Benjamini–Hochberg, against how many of twenty independent tests have a real effect of three standard errors.

On independent tests, Storey’s procedure finds at least as many real effects as Benjamini–Hochberg at every count, and more the more real effects there are, because the share of true nulls is exactly what the fixed procedure wastes. With two of twenty real it finds 55.36% against 54.97%. With eight, 77.10% against 71.61%. With fourteen, 89.59% against 79.07% — ten and a half points. With ten, 81.93% against 74.70%, 7.22 points.

It buys that by spending what Benjamini–Hochberg left unused. Its false discovery rate on independent tests is 5.04% with every null true, 5.01% with six real and 4.89% with twelve, where Benjamini–Hochberg’s falls from 5.08% to 3.55% to 2.02%. Storey’s procedure holds the rate near the level it was set at, and Benjamini–Hochberg near the level times the share of true nulls. Nothing about the promise has been weakened; the fixed procedure was keeping it with a margin it did not need, and the adaptive one hands the margin back as findings.

That is the honest case for the procedure, and it is a strong one where it applies. The price of control put the cost of a correction in points of power; this is a correction that gives some of the points back at no cost to the rate it controls.

What smaller effects leave to buy

What estimating the share of true nulls buys, against the size of the real effects. Ten of twenty independent tests real. effect 1.5: Storey 12.77%, Benjamini–Hochberg 9.59%; effect 2: Storey 34.85%, Benjamini–Hochberg 26.59%; effect 2.5: Storey 61.88%, Benjamini–Hochberg 52.00%; effect 3: Storey 81.93%, Benjamini–Hochberg 74.70%; effect 3.5: Storey 92.75%, Benjamini–Hochberg 89.03%.
Fig. 3 The share of ten real effects found by Storey’s procedure and by Benjamini–Hochberg among twenty independent tests, against the size of the real effects.

The gain depends on how far the real effects’ p-values sit from the true nulls’, because the estimate can only count as a real effect a p-value that stays below one half. With ten real effects of two standard errors, Storey’s procedure finds 34.85% of them against Benjamini–Hochberg’s 26.59%. At one and a half standard errors it finds 12.77% against 9.59% — a larger share than at three standard errors, and a small absolute number, because at that size there is little either procedure can find.

Its false discovery rate falls away from 5% as the effects shrink: 4.31% at two standard errors and 3.74% at one and a half, where it was 4.93% at three. A real effect of one and a half standard errors puts its p-value above one half about one time in five, and every such p-value is counted as a true null — so the estimate is pushed up, the threshold down, and part of what the procedure was meant to hand back stays unspent. The estimate is conservative in exactly the regime where real effects are hardest to find.

In points of power the gain is largest in between. At two and a half standard errors Storey’s procedure finds 61.88% of the real effects against Benjamini–Hochberg’s 52.00%, a gain of 9.88 points; at three and a half it finds 92.75% against 89.03%, a gain of 3.72. Small effects leave little for either procedure to find, and large effects leave little for the adaptive one to add, because at that size Benjamini–Hochberg already finds most of what there is to find.

A cap that looks harmless

Storey's procedure with its estimate capped at one and uncapped, on independent tests. 0 of 20 real, Storey, λ = 0.5: 4.966%; 0 of 20 real, Storey, capped at one: 5.412%; 5 of 20 real, Storey, λ = 0.5: 4.979%; 5 of 20 real, Storey, capped at one: 5.021%; 10 of 20 real, Storey, λ = 0.5: 4.889%; 10 of 20 real, Storey, capped at one: 4.889%; 15 of 20 real, Storey, λ = 0.5: 4.620%; 15 of 20 real, Storey, capped at one: 4.620%. 200,000 families per row.
Fig. 4 Storey’s procedure on independent tests with its estimate of the share of true nulls capped at one and uncapped, counted over two hundred thousand families at each count of real effects, with two standard errors.

A share cannot exceed one, and the estimate is commonly written capped at one: the smaller of one and the count-based figure. On two hundred thousand families of independent true nulls, the capped procedure reports a finding 5.412% of the time where 5% was allowed. The uncapped estimate — the one the proof is written for — reports 4.966%. The gap is 0.446 points against a standard error of 0.049: nine standard errors.

The cap bites exactly where every finding is false. In a family of true nulls about half the p-values land above one half, so the uncapped estimate is near one or above it — 1.1 on average — and pulls the threshold below q. Capping holds the threshold at q in precisely those families, and in those families any finding is a false one. With five real effects the capped rate is 5.021% against 4.979%; with ten or fifteen the estimate is seldom above one and the two procedures agree to the third decimal place.

The guarantee of Storey, Taylor and Siegmund is for the estimate with one added and no cap. The cap is the kind of tidy change a reader of the code would approve — a proportion above one is visibly wrong — and it cannot be seen in one run and is plain in two hundred thousand. It is also a clean case of an estimate whose job is not to be a good estimate: what the procedure needs from it is not accuracy about the share but a particular distribution, and the “correction” moves the distribution.

The counts have an exact sum behind them. Given how many of twenty true-null p-values land above one half, the rest are uniform below it, and the chance that the step-up rule rejects anything among them is, by Simes’s identity, the smaller of one and a single ratio: 5% times the number below one half, divided by one half times twenty times the estimate. Summed over the binomial count, the uncapped procedure’s rate is 5% times one minus two to the minus twentieth — 4.999995% — and the capped procedure’s is 5.4405%. The two hundred thousand counted families sit 0.7 and 0.6 standard errors from those sums. The cap is not a quirk of a simulation. It is 0.44 points of error rate that the arithmetic says are there, and that the count merely finds.

Where to read the estimate from

Storey's procedure reading its estimate from the p-values above 0.2, 0.5 and 0.8. λ 0.2, independent: every null true 5.05%; ten real 4.75%, power 82.19%. λ 0.2, correlated at 0.6: every null true 6.37%; ten real 5.28%, power 78.66%. λ 0.5, independent: every null true 5.04%; ten real 4.93%, power 81.93%. λ 0.5, correlated at 0.6: every null true 8.90%; ten real 7.28%, power 80.40%. λ 0.8, independent: every null true 5.08%; ten real 4.45%, power 79.69%. λ 0.8, correlated at 0.6: every null true 8.79%; ten real 6.35%, power 77.29%.
Fig. 5 Storey’s procedure reading its estimate from the p-values above 0.2, 0.5 and 0.8, with independent tests and correlated at 0.6: the false discovery rate with every null true and with ten real effects, the exact rate where one exists, and the power with ten real.

One half is a convention, not a requirement. Reading the estimate from the p-values above 0.2 uses more of them — a true null lands above 0.2 four times in five — and lets more real effects into the count; reading it from those above 0.8 uses fewer and lets almost none in, at the cost of an estimate built from about four p-values in a family of twenty.

On independent tests with ten real effects, the three choices find 82.19%, 81.93% and 79.69% of the real effects, at false discovery rates of 4.75%, 4.93% and 4.45%. With every null true the exact rate is 5.0000% at 0.2 and at 0.5, and 4.9424% at 0.8, where the one added to a count of about four is a larger nudge. At a threshold of 0.8 a family of twenty true nulls expects four p-values above it, so the one added to the count is a quarter of what it is added to, and the estimate runs further above one than at a threshold of one half, where the count expects ten and the one is a tenth. The exact rate is below 5% for that reason and no other: the procedure is being told there are more true nulls than there are.

Capped, the three are 5.2182%, 5.4405% and 5.8152%: the higher the threshold, the more often an estimate from a family of true nulls lands above one, and the more a cap takes away.

Correlated at 0.6, the threshold of 0.2 holds the rate best — 6.37% with every null true, against 8.90% at one half and 8.79% at 0.8 — because an estimate built from sixteen expected p-values moves less with the shared draw than one built from ten or four. It is still above 5%, and with ten real effects it is 5.28% against 7.28%. A lower threshold softens the procedure’s failure under correlation; nothing in the choice removes it.

When the tests move together

Twenty true nulls: how often each procedure reports a finding, against the correlation. Storey: 5.04% at 0, 6.33% at 0.3, 8.90% at 0.6, 19.29% at 0.9. Storey, capped at one: 5.53% at 0, 6.54% at 0.3, 8.91% at 0.6, 19.29% at 0.9. Benjamini–Hochberg: 5.08% at 0, 4.86% at 0.3, 3.70% at 0.6, 2.34% at 0.9.
Fig. 6 Twenty true nulls: the chance that Storey’s procedure, the capped version and Benjamini–Hochberg report any finding, against the correlation between every pair of tests.

On twenty true nulls correlated with each other, Benjamini–Hochberg reports a finding less often as the correlation rises: 5.08%, 4.86%, 3.70% and 2.34% at 0, 0.3, 0.6 and 0.9. Storey’s procedure does the reverse — 5.04%, 6.33%, 8.90% and 19.29%. Capping stops mattering once the tests are correlated at 0.6 or more, because the estimate is then seldom near one.

The histogram of the estimate says why. A shared component that pushes every p-value down together also pushes the count above one half down together, so the families in which the estimate says most nulls are false are the families in which the shared draw made every null look real — and the estimate hands exactly those families a generous threshold. The procedure takes the single draw that produced the flood of false findings counted for Benjamini–Hochberg and reads it twice: once in the p-values, and again in the estimate that decides how many of them to believe.

The false discovery rate of Storey's procedure and Benjamini–Hochberg's, independent and correlated. Storey, independent: 5.04% with 0 real, 5.01% with 6 real, 4.89% with 12 real, 3.47% with 18 real. BH, independent: 5.08% with 0 real, 3.55% with 6 real, 2.02% with 12 real, 0.50% with 18 real. Storey, at 0.6: 8.90% with 0 real, 8.23% with 6 real, 6.76% with 12 real, 3.73% with 18 real. BH, at 0.6: 3.70% with 0 real, 2.94% with 6 real, 1.91% with 12 real, 0.51% with 18 real.
Fig. 7 The false discovery rate of Storey’s procedure and Benjamini–Hochberg against how many of twenty tests are real, with independent tests and correlated at 0.6.

With real effects present the same thing happens more gently. At a correlation of 0.6, Storey’s rate is 8.23% with six real effects and 6.76% with twelve, where Benjamini–Hochberg’s is 2.94% and 1.91%. Only when nearly every null is false does the excess fade — 3.73% against 0.51% with eighteen real — because there is almost nothing left to be wrong about, and the level times the share of true nulls is small for both.

Adaptive to what

The word “adaptive” invites a reading the counts rule out. The procedure adapts to the family: it reads the family’s p-values and sets a threshold for that family. Under independence that is a strength, because twenty separate readings of whether each null is true add up to a good estimate of how many are.

Under correlation, twenty p-values are closer to one reading than to twenty. The shared draw is the largest single thing that happened to the family, and an estimate built from the family’s p-values measures mostly that. So the procedure adapts, faithfully, to the one feature of the family that says nothing about how many nulls are true — and an estimate from correlated p-values is not a worse estimate of the same thing. It is an estimate of something else, a statistic whose meaning depends on what produced the data as much as on its value.

Read another way, Storey’s procedure is a plug-in. An estimate of a quantity nobody knows is substituted into a rule that was exact when the quantity was known, and a plug-in carries its estimate’s variability into the rule — which is the pattern an interval that averages over an estimated spread rather than substituting it measured elsewhere. Here the rule divides by the estimate. A reciprocal of a noisy number is skewed towards large values, so a low estimate moves the threshold much further than a high one, and that is why the tenfold rise in low estimates under correlation matters more than the average that did not move.

Which to use

Storey’s procedure where the tests are independent. It keeps the rate, and it finds between half a point and ten and a half points more of the real effects, most where most nulls are false — which is the setting, a screen with many real signals, where findings matter most.

Benjamini–Hochberg where they may not be. At a correlation of 0.3 Storey’s procedure already spends 6.33% on true nulls, and the correlation among real tests is seldom known to be smaller than that. Benjamini–Hochberg becomes more conservative under the same correlation; Storey’s becomes less.

The estimate uncapped, if it is used at all. The finite-sample guarantee is for the estimate with one added to the count and no cap. The capped version is wrong by nine standard errors in exactly the case that protects nothing but the rate.

A threshold nearer one fifth, if the procedure is used where correlation cannot be ruled out. On independent tests it finds as much — a quarter of a point more — and holds the rate at 6.37% rather than 8.90% at a correlation of 0.6. That is a mitigation to be named as one, not a guarantee.

Less, at small effects, than it appears to offer. At one and a half standard errors the gain is 3.18 points on a base of 9.59%. That is a third more findings, and it is still a tenth of the real effects found.

Which one, stated. At ten real effects on independent tests the two procedures differ by 7.22 points of power; at a correlation of 0.9 with every null true, by a factor of eight in the rate delivered. A list of findings “at a 5% false discovery rate” does not say which of those two worlds it came from.

What is proved here and what is counted

Proved, and quoted rather than derived. On independent tests the uncapped estimate with one added, with rejections confined to p-values at or below one half, keeps the false discovery rate at or below q (Storey, Taylor and Siegmund, 2004). Nothing comparable is proved for correlated tests, and nothing here suggests it could be.

Counted. Every rate and share above: twenty thousand families at each setting, and two hundred thousand in the comparison of the capped and uncapped estimates, where the difference to be seen is under half a point.

Particular to these families. Twenty tests; thresholds of 0.2, 0.5 and 0.8; real effects of one and a half to three and a half standard errors; and correlation from one component shared by every test. Storey’s procedure under block correlation was not counted, and the block results for Benjamini–Hochberg in the essay on correlated tests are not evidence about it, since its failure comes from the estimate rather than from the step-up rule.

Still open: an order that spends the error rate

Both procedures in this essay and the one before it treat the twenty hypotheses as a set with no order: any of them may be found, and the level is spread across whichever are. Trials rarely have that shape. They declare a primary hypothesis, then secondary ones, and the declaration is meant to say which finding the trial is about. A procedure can use the order directly — test the first at the full 5%, and the next only if the first is rejected — and that spends the error rate one hypothesis at a time rather than across the set. An order that spends the error rate counts what that buys the first hypothesis and what it costs every hypothesis after a true null.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Benjamini–HochbergCorrelationDependenceError rateFalse discovery rateMonte CarloMultiple comparisonsNull hypothesisp-valuePlug in estimateStatistical powerUniformity