False discoveries that arrive together
Worth reading first: What the correction corrects.
Two different promises set Benjamini–Hochberg beside Bonferroni and ended on a sentence worth taking literally: the false discovery rate is an expectation. It is the average, over the studies a procedure could be run on, of the share of each study’s findings that are false. Twenty independent tests, ten of them real effects of three standard errors, and that average comes to 2.55% where 5% was allowed.
The tests in a real family are seldom independent. Twenty outcomes measured on the same patients, twenty genes in one pathway, several arms compared against one shared control: every pair of statistics moves together to some degree. The proof that Benjamini–Hochberg holds its rate needs the p-values to be independent or positively dependent, so there are two questions. Does the rate still hold? And when it holds, is an average still a fair description of what one study is told?
The rate holds, and falls
Every pair of the twenty statistics was given the same correlation, from none to 0.9, by building each one from a draw shared by all twenty and a draw of its own. Twenty thousand families were counted at each setting.
With ten real effects the false discovery rate is 2.55% with independent tests, 2.53% at a correlation of 0.3, 2.26% at 0.6 and 1.66% at 0.9. With every null true it is 5.08%, 4.86%, 3.70% and 2.34%.
So the rate does not merely hold; it falls as the tests move together. A shared component makes twenty tests behave more like fewer, and a procedure calibrated for twenty independent chances to be wrong is conservative for fewer. Correlation of this kind — every pair alike and positive — is the kind the proof covers, and the counts sit where the proof says they must: at or below 5% times the share of nulls that are true, which is 2.5% with half the nulls false.
If the rate were the whole of the matter, this would be the end of the essay.
What the rate is an average of
The rate is the same kind of number at every correlation. What it averages is not.
Take the families in which every null is true. With independent tests, 5.08% of them report a finding, and a family that reports one reports 1.10 false findings on average — one test that crossed by chance, occasionally two. At a correlation of 0.9, 2.34% of families report anything, and each of those reports 16.56 false findings out of twenty.
When every null is true, the false discovery rate is exactly the chance of reporting anything, so the rate fell by half — and the number of false findings in each report rose fifteenfold. The expected number of false findings, over every family, is the product of the two: 0.056 a family with independent tests and 0.387 at a correlation of 0.9, seven times as many. A rate that fell and a count that rose are the same measurement read two ways.
The mechanism is the shared draw. When it is large it pushes all twenty statistics out together, and Benjamini–Hochberg’s step-up rule — which lets the k-th smallest p-value through at k times 5% over twenty — is most generous exactly when many p-values are small at once. A family that is unlucky is unlucky everywhere, and the rule that was built to be lenient with a long list of small p-values cannot tell a list produced by many real effects from a list produced by one large shared draw.
At the far end, the rate comes back
The fall in the rate is not the whole of its curve. Pushed past 0.9, the chance that a family of true nulls reports anything turns round: 2.55% at a correlation of 0.95, 3.45% at 0.99, and 5.10% at exactly one. At a correlation of one the twenty tests are one test written twenty times. Every p-value is the same p-value, Benjamini–Hochberg rejects all twenty whenever that one p-value is at or below 5% and none of them otherwise, and each family that reports anything reports exactly 20.00 false findings.
So the rate starts at 5%, dips to under half of it, and ends at 5%. The expected number of false findings does not dip. It rises the whole way: 0.056 a family with independent tests, 0.080 at a correlation of 0.3, 0.168 at 0.6, 0.387 at 0.9, 0.476 at 0.95, 0.682 at 0.99 and 1.019 at one, where the closed form — twenty findings on 5% of families — is exactly one false finding a family.
A procedure’s error rate and its expected number of errors are both honest summaries, and under dependence they move in opposite directions over most of the range. A study that reports its false discovery rate has reported the first. A reader deciding how many of a list’s findings to follow up, fund or retract needs the second, and nothing in the first lets the second be recovered.
One family at a time
The same comparison with ten real effects, drawn family by family. With independent tests, 398 of two thousand families make at least one false finding, and not one makes five or more. Correlated at 0.9, 75 make a false finding — five times fewer — and 41 of those 75 make five or more.
A reader of one study does not see the rate. They see one cell of this picture. With independent tests, the cells that contain a false finding contain one; at high correlation, the cells that contain any false findings mostly contain many. “A 5% false discovery rate” is a description of the whole picture, and it is equally true of both.
The share false, study by study
Written as the share of each study’s findings that are false, among the studies with at least one finding: with independent tests, 81.99% have fewer than a tenth of their findings false, and no study has every finding false. At a correlation of 0.9, 95.68% have fewer than a tenth false — the distribution is more concentrated on clean lists — and 0.808% have every finding false, about one study in 124.
Two more readings of the same counts, at a correlation of 0.6. The share of a study’s findings that are false exceeds a fifth in 3.03% of families, against 1.00% with independent tests, and exceeds a half in 0.88%, against none. The false discovery rate over those same families is 2.26%, lower than the independent 2.55%.
The average improves while the tail grows. That is not a paradox and it is not a failure of the procedure; it is what an expectation does when the thing being averaged becomes lumpier. It does mean that a study reporting ten findings from correlated tests, at a controlled false discovery rate of 5%, has been told something true about the procedure and much less than it seems about its own list. One run is an anecdote here in a particular sense: the promise was never about one run.
Correlation in blocks
Every pair of tests alike is the simplest correlation and not the usual one. Genes come in pathways, outcomes in domains, arms in groups sharing a control. So the twenty tests were also arranged in blocks — correlated at 0.9 inside a block and independent between blocks — of two, five, ten and twenty.
With every null true, blocks of two report anything 4.15% of the time, with 1.73 false findings in each such family; blocks of five, 3.36% and 3.80; blocks of ten, 2.87% and 7.74; and one block of twenty is the shared draw already counted, 2.34% and 16.56. With ten real effects among the twenty the false discovery rate is 2.41%, 2.18% and 2.00% for blocks of two, five and ten, and the share of families in which more than half of the findings are false is 0.01%, 0.14% and 0.50%, against 0.61% for one block of twenty. Power barely moves: the ten real effects are found 74.15%, 72.52% and 71.21% of the time with blocks of two, five and ten, against 74.70% with independent tests and 70.82% as one block.
The block size is the size of the flood. A shared draw can push out only the tests it is shared by, so a family built from blocks of five that is unlucky is unlucky, on average, in about four tests at once rather than in twenty. The rate holds under every arrangement counted; what the arrangement decides is the unit in which its failures arrive. For a list of genes that unit is a pathway, which is also the unit in which a list is usually followed up — and usually retracted.
The guarantee for any dependence, and what it costs
Benjamini and Yekutieli showed that dividing the level by 1 + 1/2 + … + 1/m keeps the false discovery rate under q whatever the dependence. For twenty tests the divisor is 3.598, so the procedure runs at 1.39% where Benjamini–Hochberg runs at 5%. And the divisor grows with the list: for a hundred tests it is 5.187, for a thousand 7.485, so the procedure runs at 0.96% and 0.67% respectively. The screens that produce a thousand p-values are exactly the settings where dependence is most likely and power scarcest, and the general guarantee’s price grows there with the logarithm of the list.
The divisor is not caution for its own sake. Under dependence arranged against it, Benjamini–Hochberg’s false discovery rate can climb to the level times that same sum — the bound was shown to be attained, not merely possible, by Guo and Rao — so the divisor is exactly the worst case. A procedure built to be safe against the worst case pays the worst case’s price on every family it is run on, including families like the ones counted here, whose rate under Benjamini–Hochberg never came near its bound.
Its false discovery rate with ten real effects is 0.72% with independent tests and 0.50% at a correlation of 0.9 — a guarantee kept with room to spare. Its price is in what it finds: 55.12% of the ten real effects with independent tests, against 74.70% for Benjamini–Hochberg, a loss of 19.59 points. Holm, which controls the far stricter familywise rate, finds 52.53%. At every correlation here Benjamini–Yekutieli is within three points of a familywise procedure and between 17.0 and 19.6 points short of the procedure it protects.
That price buys protection against the kinds of dependence under which Benjamini–Hochberg fails, and the families counted here include none. For the correlation a shared component produces, the guarantee is paid for and not used. The price of control set Holm’s thirty-three points of power against Benjamini–Hochberg’s ten as the cost of two different promises; Benjamini–Yekutieli charges nearly the stricter price for the weaker promise.
The arrangement the proof does not cover
Positive dependence is what the proof needs, and negatively correlated statistics are the standard case outside it: two-sided p-values from statistics that move in opposite directions are not positively dependent in the sense the proof requires. So the twenty tests were arranged in ten pairs correlated at −0.3, −0.6 and −0.9.
With every null true, Benjamini–Hochberg’s rate is 5.10%, 5.00% and 4.38%. With ten real effects it is 2.50%, 2.47% and 2.40%. The largest of those, 5.10% against 5% on twenty thousand families, is 0.7 standard errors from the level — nothing.
The arrangement that falls outside the proof does not, here, fall outside the rate. Constructions that do break Benjamini–Hochberg exist, and they are specific and extreme; this figure is not evidence that none does. It is evidence that the ordinary negative correlations a study might carry are not where the danger to the rate lies — and the danger to the list, counted above, lies with positive correlation, which the proof covers.
What a study can know about its own correlation
Nothing in a list of findings says how the tests behind it were correlated, and nothing in a false discovery rate does either. But the study that produced the list usually holds the evidence. Twenty outcomes measured on the same people have a sample correlation that can be computed before a single hypothesis is tested, and it is a property of the measurements rather than of the effects being looked for. Twenty genes drawn from named pathways have a block structure written into the database they were drawn from. The difference between the independent picture and the flooded one is not an unknowable feature of nature; it is a quantity sitting in the data set, usually unreported because the rate does not need it.
The counts above say what that quantity is for. A correlation near zero means a list’s false findings, if it has any, are strays. A correlation of 0.6 or more among the tests that produced a list means that, whatever its rate, the list is more likely than a reader of independent lists expects to be right as a whole or wrong as a whole — and a block structure says how large the whole is. None of that changes the decision about which findings to report. All of it changes how many of them to bet on together.
What a list of findings from correlated tests should carry
The correlation, or an honest statement that it is unknown. A false discovery rate means one thing for independent tests and another for tests that share a component, and the difference is not in the rate.
The expected number of false findings, not only the rate. Twenty true nulls produce 0.056 false findings a study with independent tests and 0.387 at a correlation of 0.9. The first is a stray result; the second is a whole list that will not replicate.
The unit a failure would come in. Correlated in blocks of five, a list that fails fails by about four findings at a time; correlated as a whole, by nearly all of them. Which of those a study’s tests resemble is usually known to the people who designed the measurements, and it says more about how to read the list than the rate does.
A second stage for any list that will be acted on as a whole. A finding selected for being far out is inflated one at a time; a list selected by a shared draw is wrong all at once, and only an independent second sample separates the two.
Benjamini–Hochberg rather than Benjamini–Yekutieli, unless the dependence is of a kind known to break it. The general guarantee costs nearly twenty points of power to insure against a failure that, for the correlations counted here, does not occur.
What is proved here and what is counted
Proved, and quoted rather than derived. Under positive dependence Benjamini–Hochberg’s false discovery rate is at most the share of true nulls times q; Benjamini–Yekutieli’s is at most q under any dependence (Benjamini and Yekutieli, 2001). What the correction corrects makes the same kind of argument for the familywise procedures, whose guarantees need no assumption about dependence at all.
Counted. Every rate, share and count above: twenty thousand families of twenty tests at each correlation, two thousand in the family-by-family figure, effects of three standard errors, and a level of 5%.
Particular to these families. Positive correlation shared by every pair, positive correlation in equal blocks of two to twenty at one within-block strength, and negative correlation in pairs. Blocks of unequal size, correlation that differs from block to block, and dependence between the real effects and the nulls are not measured, and nothing above says how they would fall between the cases that are.
Still open: estimating how many nulls are true
Every rate above sat below its bound by the share of real effects: 2.55% of a permitted 5% with half the nulls false. That unused half is not a margin anybody chose. It is the procedure spending q as though every null were true, because it has no way to know otherwise — and a p-value from a true null is flat, which gives a way to estimate how many are. Estimating how many nulls are true counts what spending the rest buys, why a cap on the estimate that looks harmless is not, and what the shared draw measured here does to an estimate built from p-values that move together.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An order that spends the error rate — both name error rate, false positive, familywise error rate, holm, monte carlo, multiple comparisons, statistical power
- The models that were never in the running — both name error rate, familywise error rate, monte carlo, multiple comparisons, statistical power
- Eight forecasters and one benchmark — both name error rate, familywise error rate, monte carlo, multiple comparisons
- How many analyses there really were — both name correlation, false positive, familywise error rate, multiple comparisons
- The rank is a decision — both name error rate, monte carlo, multiple comparisons, statistical power
- What naming it in advance costs — both name false positive, familywise error rate, multiple comparisons, statistical power
Named objects
A flat tag is an object no other essay names yet.
Benjamini–HochbergCorrelationDependenceError rateFalse discovery rateFalse positiveFamilywise error rateHolmMonte CarloMultiple comparisonsStatistical power