The smallest of three combinations
Worth reading first: A p-value that is not flat is not a p-value · What the correction corrects.
Two ways to combine p-values ended on a test nobody names. Ten studies report ten one-sided p-values, and a meta-analyst computes Fisher’s combination, Stouffer’s and Tippett’s, and reports whichever comes out smallest. Each of the three is exact under the global null, because each turns a flat p-value into a variable with a known law. The report is a fourth test — it rejects when any of the three does — and nothing about the three being exact makes the fourth exact.
Its size can be bounded before anything is computed. It is at least 5%, since it rejects whenever Fisher’s combination does, and at most 15%, since the probability of a union is at most the sum of its parts. Where it sits between those is a fact about how much the three rejection regions overlap, and the overlap is large: all three read the same ten numbers. This essay measures it exactly for two studies, encloses it for ten, follows it to fifty, and then prices the repair — reading all three at a stricter common level so that the smallest of them is an exact 5% test — against the power each combination has on its own.
Two studies, and a combination that adds nothing
For two studies the three rejection regions can be drawn, and the union is everything outside a single shaded region: below Tippett’s cut of 1.955 in both studies, below Stouffer’s line where the two z statistics sum to 2.326, and inside Fisher’s curve. Its probability is a one-dimensional integral over the first study’s z, with the second study’s normal distribution function evaluated at the lowest of the three boundaries. The integral gives 7.536%. Counted on a million pairs of null studies, the union rejects 7.54%, with a standard error of 0.03 points.
The two-study picture holds a surprise. Stouffer or Tippett alone reject 7.536% as well — the same to every digit the integral carries — so at two studies Fisher’s region lies entirely inside the other two, and computing Fisher’s combination adds nothing to a report that already reads Stouffer’s and Tippett’s. The reason is visible in the figure. Fisher’s curve leaves the Tippett square only where one study is past its cut, and inside the square the curve is closest to Stouffer’s line at its end, where one study sits at 1.955 and Fisher needs the other at 0.402. That pair sums to 2.357, which is still past 2.326. Every point Fisher rejects inside the square, Stouffer has already rejected.
The pairs that do overlap only partly say how much two combinations share. Fisher or Stouffer reject 6.073% of null pairs, Fisher or Tippett 6.464%. All three reject together on 2.46%, so nearly half of each combination’s 5% is spent on pairs every combination agrees about.
Ten studies, enclosed on a lattice in two sums
For ten studies the regions live in ten dimensions and cannot be drawn, but the event that nothing rejects has a structure that can be computed. It happens exactly when every study’s z is below Tippett’s cut, the sum of −2 log p is below Fisher’s critical value, and the sum of z is below Stouffer’s. Each study adds a point to two running sums, so the probability is a two-dimensional convolution of one study’s contribution with itself ten times.
That convolution is done on a lattice of cells, and the lattice error is not estimated but enclosed. Inside one cell of a study’s z, both coordinates it contributes are increasing in z, so moving every study to the lower corner of its cell can only shrink both sums, and moving it to the upper corner can only grow them. The first reading counts too many sets as silent and the second too few, whatever the cell width. On cells a twentieth wide, the union’s size lies between 8.76% and 10.57%; on cells a fortieth wide, between 9.18% and 10.09%. The midpoints of those two enclosures are 9.67% and 9.64%. Counted on a million sets of ten null studies — a route that shares nothing with the lattice — the union rejects 9.66%, with a standard error of 0.03 points.
So the smallest of the three rejects nearly twice as often as each combination alone, on nothing. The pairs fill in how. Fisher or Stouffer reject 6.63% of null sets, Fisher or Tippett 8.05%, Stouffer or Tippett 8.96%, and all three together 1.05%. At ten studies Fisher is no longer redundant: it adds 0.70 points to what Stouffer and Tippett already reject, where at two studies it added none.
Why the union stops well short of fifteen per cent
The 15% bound is what three tests with disjoint rejection regions would give, and these three are close to the opposite: their statistics are strongly correlated because each is a different summary of the same ten p-values. Counted over the null sets, Fisher’s statistic and Stouffer’s are correlated at 0.903, Fisher’s and the smallest p-value’s at 0.744, and Stouffer’s and the smallest p-value’s at 0.504.
The first of those is a closed form, and it does not depend on the number of studies. Fisher’s statistic and Stouffer’s are sums over the same independent pairs, minus twice the log of a study’s p-value and that study’s z, so the number of studies cancels from their correlation. By Stein’s lemma the covariance of the pair is twice the expected normal hazard, the density over the upper tail at a standard normal draw, and the variances are four and one. The correlation is therefore that expected hazard, and by quadrature it is 0.9032.
The ordering of the pairwise unions follows the correlations. The most correlated pair, Fisher and Stouffer, adds least to either: their union exceeds 5% by 1.63 points. The least correlated pair, Stouffer and Tippett, adds most, 3.96 points, because a set with one strong study and nine at nothing is exactly what Tippett rejects and exactly what Stouffer’s sum dilutes. Fisher sits between them in both senses, which is where its rejection region was drawn.
More studies, a larger fourth test
The union grows with the number of studies. Counted on null sets, it rejects 7.54% at two studies, 8.25% at three, 8.95% at five, 9.66% at ten, 10.12% at twenty and 10.66% at fifty. Fisher or Stouffer barely moves, from 6.08% to 6.71%, which is what a correlation fixed at 0.903 allows. What grows is Tippett’s share: its statistic is correlated with Fisher’s at 0.948 for two studies and 0.498 for fifty, and with Stouffer’s at 0.786 and 0.286.
That decoupling has an end, and the end can be computed. As the studies multiply, Fisher’s and Stouffer’s statistics become a bivariate normal pair with the correlation above, and the smallest p-value — the largest of the same exponential terms Fisher adds up — becomes independent of both sums. The union then tends to one minus 95% of the probability that neither sum rejects. For normal statistics correlated at 0.9032, either one rejects 6.78% of the time, and the smallest of three tends to 11.45%.
The limit says the problem does not go away with more studies, and does not become catastrophic either. A report that picks the smallest of three combinations roughly doubles its error rate at every number of studies a meta-analysis is likely to have, and more than doubles it past about twenty.
Reading all three at a stricter level
The repair is not to abandon the smallest of three but to calibrate it. Its combined p-value is itself a statistic with a null distribution, and the level at which that distribution reaches 5% is the level at which each of the three should be read.
For two studies the level is found by bisection on the integral: 3.221%. For ten studies it is located on the lattice, on cells a fortieth wide, at 2.448% — the level at which the midpoint of the enclosure is 5%, with the enclosure itself running from 4.72% to 5.27%. Counted on the same million null sets, the smallest of three read at 2.448% rejects 5.02%. Read off those sets directly, the level at which the counted smallest p-value reaches 5% is 3.216% at two studies, 2.439% at ten and 2.218% at fifty.
The familiar alternative is to divide by three. Reading each combination at a third of 5%, 1.667%, makes the union bound hold, and it holds with room to spare: the smallest of three then rejects 3.50% of null sets. That is the correlation left on the table. Dividing by three treats the three combinations as though they could disagree completely, and they mostly agree, so the division buys a test well inside its size and pays for it in power — the kind of cost the price of control put numbers on for corrections across several hypotheses.
What the honest version gains and gives up
The power comparison uses the alternative the earlier essay built: ten studies, m of them shifted by a common amount and the rest at nothing, with the shift set so that Stouffer’s power is exactly 50% at every m. Fisher’s combination is most powerful when the signal is concentrated, Tippett’s when it is in one study, Stouffer’s when it is shared. No single combination is best everywhere, and which one is best depends on an arrangement the analyst does not know in advance.
The calibrated smallest of three never beats the best single combination and never falls to the worst. With the signal in one study it has 99.31% power against Tippett’s 99.60%. In two it has 77.23% against Tippett’s 77.25%, which is the best combination to within its own lattice. From three studies on it trails whichever combination is best — by 4.60 points at three, where Fisher leads, and by 7.45 points at ten, where Stouffer does — and it leads the worst by at least 10.30 points, which is the margin at three studies over Stouffer’s 50%. Its power at ten studies is 42.55%. The counted power agrees with the lattice at every m to within two of the counts’ standard errors.
The uncalibrated version looks better and is not a test of 5%. Read at 5% each, the smallest of three has 55.85% power with the signal shared by all ten studies. Stouffer’s combination has 50% there, and for a shift common to every study no 5% test can do better than Stouffer’s, by the Neyman–Pearson lemma. A test that beats the most powerful test of its size is a test of a larger size, and this one is 9.66%.
Dividing by three instead of calibrating costs power at every m. At ten studies the three read at a third of 5% each reach 36.19%, six points under the calibrated test, and at three studies 53.82% against 60.30%.
The same trade with a stronger signal
With a signal strong enough to give Stouffer’s combination 80% power, every curve lifts and the trade keeps its shape. The calibrated smallest of three has 93.31% power with the signal in three studies, against Fisher’s 95.18% and Stouffer’s 80%, and 73.16% with it in all ten, against Stouffer’s 80%. It trails the best single combination by at most 6.84 points and leads the worst by at least 13.31. The uncalibrated version reaches 82.86% at ten studies, again past the most powerful 5% test, and the version read at a third of 5% reaches 67.50%.
A stronger signal narrows the premium a little and widens the margin over the worst choice, because every combination gets closer to certain rejection where the signal is concentrated and the calibrated test has less room to lose.
A choice made before the p-values are seen
The whole of the difference between the uncalibrated and the calibrated test is when the choice among combinations is made. An analyst who decides before seeing the ten p-values that the studies measure one shared effect, and computes only Stouffer’s combination, runs an exact 5% test with 50% power in the shared case. One who decides that the effect may live in a few populations and computes only Fisher’s runs a different exact test. One who computes all three and keeps the smallest, then reads it against 5%, has run an analysis chosen from the data, and the choice is part of the result in the same way the timing of a look is part of a p-value.
What distinguishes this case from twenty analyses of one dataset is the amount of inflation, which is small because the choices are few and highly correlated. Three readings of the same ten numbers inflate 5% to 9.66%, where twenty loosely related analyses inflate it far more. The calibrated test is the honest form of not knowing which arrangement the evidence will take: it pays up to 7.45 points of power, against an analyst who happened to guess right, for protection of at least 10.30 points against one who guessed wrong.
The trade is the same one multiplicity corrections make, one level up. There the question is which of several hypotheses to reject; here it is which of several statistics to believe about one hypothesis. In both, correlation between the pieces is what makes an exact correction cheaper than a union bound, and the calibrated level of 2.448% sits well above the 1.667% that division by three would impose, for the same reason that false discoveries that arrive together change what a procedure’s error rate means.
What a combined result should report
Which combination was chosen, and when. A combination named in a protocol is an exact test at its stated level. A combination selected after the p-values were computed is a draw from a larger test whose level is not 5%.
If several were computed, the calibrated level of their minimum. For ten independent studies it is 2.448% for the three standard combinations, and it is a computable number for any set of combinations and any number of studies — by integration for two, on a lattice or by counting beyond that.
The number of studies. The smallest of three rejects 7.54% of null sets at two studies and 10.66% at fifty, and the level that repairs it moves from 3.221% to 2.218%, so a calibration done for one number of studies does not carry over to another.
Exact, enclosed and counted
Exact. The two-study union and its calibrated level, by one-dimensional quadrature; the correlation between Fisher’s and Stouffer’s statistics, by Stein’s lemma and quadrature; the limit of the union as the studies multiply, from a bivariate normal integral.
Enclosed. The ten-study union at every level and every alternative, between a lower and an upper lattice reading that bound it whatever the cell. Powers are drawn at the midpoint on cells a twentieth wide and the calibrated level is located on cells a fortieth wide.
Counted. A million sets of null studies at two, three, five and ten studies and three hundred thousand at twenty and fifty; two hundred thousand sets at each alternative. Every counted union is held to the exact value or the enclosure where there is one.
Assumed and not tested. Independent studies, one-sided p-values from continuous tests, and signals of equal size in the studies that carry one.
Still open: studies that share a control group, and a choice made from the shape of the evidence
The first open question is dependence. Every exactness argument here, and every calibration, uses independent studies, and studies that share a control group, a population or an analyst do not produce independent p-values. A positive correlation between studies makes all three combinations anti-conservative on their own, by different amounts — Stouffer’s sum inflates its variance directly, while Tippett’s minimum barely notices — so the smallest of three under dependence has two inflations stacked, and the calibrated level for independent studies would no longer be exact. The same lattice cannot compute that case, since the studies no longer add independently, but a count can, and so can a normal approximation for Stouffer’s piece.
The second is a cheaper repair than calibration. An analyst could choose the combination from the shape of the p-values rather than their size — Tippett’s when one p-value is far smaller than the rest, Stouffer’s when they are all moderate — with the rule fixed in advance. The rule is a function of the data, so the resulting test has a size of its own, but the shape of ten p-values is only loosely related to whether any of the three rejects, and the inflation may be well below the 9.66% of keeping the smallest. Whether such a rule can recover most of the calibrated test’s protection while keeping more of each combination’s power is a measurement this lattice and this count can make, and have not made.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Estimating how many nulls are true — both name correlation, monte carlo, multiple comparisons, p-value, statistical power, uniformity
- A distribution drawn from the null — both name monte carlo, p-value, statistical power, uniformity
- The p-value a replication gets — both name monte carlo, p-value, statistical power, uniformity
- A coverage table with its own error — both name monte carlo, multiple comparisons, statistical power
- An order that spends the error rate — both name monte carlo, multiple comparisons, statistical power
- How many analyses there really were — both name correlation, multiple comparisons, p-value
Named objects
A flat tag is an object no other essay names yet.
CorrelationFisher's methodGlobal nullMonte CarloMultiple comparisonsp-valueStatistical powerStouffer's methodTippett's methodUniformity