Two ways to combine p-values
Worth reading first: A p-value that is not flat is not a p-value · What the correction corrects.
Ten studies of one question each report a p-value, and nothing else is available: no raw data, no effect estimates on a common scale, only ten numbers between zero and one. Two different questions can be put to them, and they are easy to confuse.
The first is which of the ten found something. That is the multiplicity question, and its corrections answer it by controlling how often a study is wrongly declared a finding — the familywise rate or the false discovery rate, two different promises that cost different amounts of power. The second is whether anything at all is going on: whether the ten together are evidence against the hypothesis that every one of their nulls is true. That is the global null, and a combination of p-values tests it. Rejecting it names no study, and a correction’s output cannot answer it.
There are three standard combinations and they are all exact, meaning each rejects a true global null at exactly its stated rate. Each is exact for the reason a p-value that is not flat is not a p-value spends its whole length on: under its own null every p-value is uniform. Under an alternative they stop agreeing, and the disagreement is not noise. It is a statement about which arrangement of evidence each one is built to find, and it can be computed rather than argued about.
Three transformations of a flat number
Every combination turns each uniform p-value into a variable with a known distribution, then adds. The p-values here are one-sided, since combining evidence about an effect needs a direction for the effect.
Fisher’s method uses . If is uniform, is exponential with mean two, which is a chi-squared on two degrees of freedom, so the sum over k studies is a chi-squared on 2k. At ten studies the 5% critical value is 31.410.
Stouffer’s method uses the normal score . If is uniform, is standard normal, so is standard normal and is compared with 1.645.
Tippett’s method uses the smallest p-value alone. The minimum of k independent uniforms falls below with probability , so the critical value is : 0.005116 at ten studies, which is a single study’s statistic beyond 2.568.
Nothing in any of the three depends on what the studies measured or how large they were. All three are exact because the inputs are uniform and independent, and for no other reason.
The figure holds the output to the check a single test is held to. The Kolmogorov–Smirnov distances from uniform are 0.0045 for Fisher and 0.0044 for Stouffer, with p-values of 0.81 and 0.84. The rejection rates are 4.92% and 4.90%, against a Monte Carlo standard error of 0.15 points. A combined p-value that is not flat is not a p-value either, and these two are flat.
Two studies, drawn where their statistics live
The figure below draws all three tests for two studies in the plane of their z statistics. The grey cloud is a sample of null pairs, and each test rejects in a different region.
Stouffer’s region lies above a straight line, . It cares about the sum and nothing else, so a pair at (2.3, 0) and a pair at (1.16, 1.16) are equally convincing to it.
Fisher’s region lies outside a curve on which , with . The curve runs parallel to both axes, and past one study rejects on its own, whatever the other reports — even a strongly negative z. Fisher is willing to be convinced by one result.
Tippett’s region is an L. It rejects when either statistic passes 1.955, the level 0.02532 of a single test, and it ignores the other study completely.
Each region holds exactly 5% of the standard bivariate normal. For Fisher that is closed form: the chi-squared on four degrees of freedom has upper tail , and at that is 0.05 to thirteen decimal places. Counted on a hundred thousand null pairs, the three regions catch 4.95%, 4.88% and 5.04%, with a standard error of 0.07 points. The three shapes are different because they spend the same 5% in different places, and where each spends it decides what it can find.
The two marked points carry the same Stouffer evidence. One is a single study at a shift of 2.326 with the other at nothing. The other is both studies at 1.163. Stouffer has exactly 50.0% power at both. At the single strong study Fisher has 62.7% and Tippett 65.4%. At the two weak ones Fisher has 47.7% and Tippett 38.3%.
For a common shift in every study, Stouffer’s test is not merely one choice among three. The likelihood ratio for a shared shift depends on the data only through , so by the Neyman–Pearson lemma no test of the same size is more powerful against that alternative. That is a theorem, and it holds only for equal shifts in a known direction. The question with content is how fast the advantage goes once the shifts are not equal.
One signal, spread over more and more of ten studies
Hold the total evidence fixed and move it around. Ten studies; m of them carry a shift μ and the rest carry nothing; and μ is set for each m so that Stouffer’s power is exactly 50%, which makes the same at every m. At m = 1 the one informative study is shifted by 5.2015 standard errors. At m = 10 all ten are shifted by 0.5201. Stouffer cannot tell these arrangements apart, by construction. Fisher and Tippett can.
Stouffer’s power is closed form, and so is Tippett’s, as a product of normal tails. Fisher’s is not, and it is computed two ways that share nothing.
The lattice. A shifted study contributes with , and has a closed-form distribution function, . The m shifted contributions are put on cells a hundredth wide and convolved. The ten minus m null contributions are the exact chi-squared. A sum that lands in cell s truly lies somewhere between s and s + m cells, so reading the chi-squared tail at both ends gives a lower and an upper bound. Fisher’s power is enclosed, not estimated, and at no signal the enclosure has to contain the size: it reads 4.982% to 5.018%.
The count. Twenty thousand sets of ten studies at each m, the same statistic computed on each.
Fisher’s power is 96.05% with the signal in one study, 77.13% in two, 64.90% in three, 57.95% in four and 53.64% in five. With the signal in six studies its enclosure is 50.62% to 50.90%, above Stouffer’s 50%. With the signal in seven it is 48.54% to 48.87%, below it. At ten it is 45.03%. The crossing sits between six and seven, and it is located rather than estimated: at no m does the enclosure straddle the line it is compared with.
Tippett is the same story pushed further. It reaches 99.60% when one study holds everything and falls to 18.54% when ten share it. It beats Fisher at one and at two studies (77.25% against 77.13%) and loses from three.
The counts agree with the enclosures everywhere, with one detail worth stating rather than smoothing over. At m = 1 the counted power is 96.13% with a standard error of 0.14 points. At m = 10 it is 45.91% against an enclosure whose top is 45.26%, which is 1.9 standard errors high. On the same draws Stouffer’s count is 50.73% against an exact 50%, also high, by 2.1 standard errors. Both statistics are read off the same twenty thousand sets of studies, so one lucky batch moves both counts the same way, which is what shared draws do.
A higher bar moves the crossing, and not by much
The crossing is not a property of the combinations alone. It depends on how strong the total signal is, which the next figure changes.
With Stouffer at 80%, Fisher is ahead through seven studies, at 80.30%, and behind from eight, at 78.09%. With the signal in one study it has 99.998%, and with it in all ten it has 74.75%; Tippett falls to 31.70%. A stronger total signal gives the concentrated arrangement a little more room. The direction of the trade does not change, and neither does the reason: Fisher’s region reaches out along the axes, and a stronger signal lets more of the spread-out arrangements get far enough along them.
Where Fisher stops leading as the studies multiply
At ten studies the crossing is at six. The obvious guesses are that it stays at a fixed fraction of the studies, or that it grows like the square root of their number, which is how the Stouffer statistic’s noise grows. Neither is right.
The crossing is 1 of 2, 2 of 3, 3 of 4, 3 of 5, 4 of 6, 5 of 8, 6 of 10, 7 of 12, 8 of 15, 10 of 20, 12 of 25, 14 of 30 and 18 of 40. As a share of the studies it drifts down, from 75% at four to 60% at ten and 45.0% at forty. Divided by the square root of the number of studies it rises, from 0.707 at two to 2.846 at forty. With Stouffer held at 80% instead, the crossing at forty studies is 21.
Each point is located by the same enclosure. At twenty studies and above, cells a fiftieth wide were too coarse at the rows nearest the crossing, so those rows alone were recomputed on cells five times finer until the enclosure cleared the comparison. There is no formula here. The points are measurements of one family of alternatives — equal shifts in m studies and nothing in the rest — at two strengths. They say nothing about signals of unequal size, and nothing about how the crossing behaves beyond forty studies.
The choice is a statement about the alternative
So the three combinations are not better and worse versions of one test. Each is a claim about where the evidence will be.
Stouffer’s is the right claim when the studies measure one shared effect, as replications of a single experiment do, and it is the most powerful test there is for that case. Fisher’s is the right claim when the effect might be present in some populations and absent in others. Tippett’s is the right claim when one study might carry the whole signal and the rest are noise. The regions in the two-study figure are those three claims drawn.
This is where the combination inherits a problem from the essay on analyses of nothing: the claim has to be made before the p-values are seen. Computing all three combinations and reporting the smallest is a fourth test. Its size is not 5% and has not been computed here. The choice of combination is part of the result, in the same way a stopping rule is part of a p-value.
Two-sided inputs are exact, and blind to direction
A two-sided p-value is uniform under its null too, so a combination of two-sided p-values is exactly as exact. Counted on eight thousand sets of ten null studies, Stouffer’s combination of them rejects 4.58% of the time, with a standard error of 0.24 points around 5%.
What it is exact about is a hypothesis with no direction in it. Shift every one of the ten studies by half a standard error upwards and it rejects 13.14% of the time. Shift every one downwards by the same amount and it rejects 12.35% — the same rate within the counting error, since the inputs cannot see the sign. Stouffer’s one-sided combination of the same upward shift has 47.46% power. A combination of two-sided p-values is legitimate, and it answers a question a study of one effect does not normally ask.
A small non-uniformity, combined ten times
The exactness depends on each input being flat, and the essay on flatness lists the ways an honest-looking p-value stops being flat. One of them is a variance divided by n instead of n − 1. Feed that to the combinations and they do not treat it the same way.
Take a one-sided t-test on five observations with that slip in it. On two hundred thousand null studies it rejects 6.46% of the time, with a standard error of 0.05 points. That is a real excess and a modest one. Combine ten of them by Fisher’s method and the size is 7.92%, standard error 0.19. Combine the same ten by Stouffer’s and it is 6.59%, standard error 0.18, which is the single test’s excess carried through, neither shed nor multiplied.
The difference is where each combination reads its inputs. The slip stretches every t statistic by one factor, √(5/4). A stretch by a constant factor moves a normal score by close to a constant fraction, which Stouffer adds up. It moves a far tail probability by a factor that grows the further out it is read, and Fisher’s statistic is dominated by the smallest p-values. Fisher reads exactly the region where a slightly wrong test is most wrong. The opposite claim — that a combination of these p-values keeps its 5% size — fails when it is computed.
The broader consequence reaches past this one slip. A combination has no way to see that its inputs are not flat, so it inherits every defect of every input. Selection is the extreme case. A set of p-values that is only the published significant ones is as far from uniform as a set can be, and the winner’s curse has already priced what that selection does to the estimates. Combining such a set tests nothing.
When the p-values cannot be flat
Some correct tests cannot produce a uniform p-value at all. The exact one-sided binomial test on ten trials at a null probability of one half can return only eleven p-values, and on its own it rejects 1.07% of the time at a nominal 5%, which is the conservatism of a discrete test that an interval’s coverage also shows. The mid-p value, which counts only half the probability of the observed outcome, rejects 5.47% of the time — closer to nominal, and on the wrong side of it.
Here the size of a combination does not have to be simulated. Each study’s p-value takes one of eleven values with a known probability, so every multiset of outcomes can be enumerated with its multinomial weight — 184,756 of them at ten studies — and the size is a sum. It is the move that counts an interval’s coverage over every possible sample rather than over a simulated few.
From ordinary p-values Fisher’s size is 1.07%, 1.59%, 1.67%, 1.33%, 1.12%, 0.94%, 0.71% and 0.55% at 1, 2, 3, 4, 5, 6, 8 and 10 studies. It rises before it falls, because the eleven possible values combine into a lattice of totals whose position against the critical value moves with k. From mid-p values it is 5.47%, 3.36%, 3.76%, 4.26%, 4.04%, 4.01%, 3.98% and 3.89%. The mid-p value is anti-conservative alone and conservative in every combination measured. The counted route agrees: 1.59% against an enumerated 1.59% at two studies, and 0.54% against 0.55% at ten. The mid-p counts are 3.34% against 3.36% at two studies and 3.74% against 3.89% at ten, the last 1.6 standard errors low.
Stouffer’s method does worse with the ordinary p-value. Its size is 1.07%, 2.07%, 1.98%, 0.88%, 0.81%, 0.71%, 0.51% and 0.37% at the same numbers of studies, and there is a structural reason for the collapse. A study that observes no successes has a p-value of exactly one, whose normal score is minus infinity, so one such study vetoes the whole combination whatever the other nine say. From mid-p values the size is 5.47%, 5.77%, 4.94%, 4.08%, 4.02%, 4.62%, 4.65% and 4.44%: above nominal at two studies, below it from three.
Neither choice of p-value gives a combination whose level is known in advance. For discrete inputs the level has to be computed, by enumeration as here or by an exact reference distribution for the combined statistic, and the conservative direction is not guaranteed by using the ordinary p-value either.
What is exact here, what is enclosed, and what is only counted
Exact: the null distribution of each combination, the two-study regions’ probabilities, and the size of every discrete combination, which is a finite sum over enumerated outcomes.
Enclosed: Fisher’s power under an alternative, between bounds that hold whatever the lattice’s cell width. Stouffer’s and Tippett’s powers are closed form.
Counted, and only as a second route: every one of those numbers is also checked against simulated studies, and the counts are held to their own standard errors. So are the measurements that have no other route: the two-sided rejection rates and the wrong-divisor sizes.
Assumed throughout and not tested: independence between studies. Studies that share a control group, a population or an analyst do not produce independent p-values, and all three exactness arguments use independence as heavily as they use uniformity. Nothing here measures what dependence does.
Where this goes next
The most direct continuation is the one named above and left uncomputed. A meta-analyst who computes Fisher’s, Stouffer’s and Tippett’s combinations and reports whichever is smallest has run a fourth test, and its size is a measurable quantity. It has to exceed 5%, and it is strictly less than three times that, since the three rejection regions overlap heavily. How much less is the number. It is distinct from twenty analyses of one data set because the choice here is among ways of reading the same ten p-values, so the three statistics are strongly correlated, and the inflation should be far smaller than the forking-paths figure.
The second continuation is the discrete one. A randomised p-value, which breaks ties with an auxiliary uniform, is exactly flat, so a combination of randomised p-values is exactly exact again. The price is power and a result that depends on a coin, and the size of that price has not been measured.
The third is weights. Stouffer’s statistic is usually weighted by each study’s size, which is the optimal weighting for a shared effect on a common scale, and the trade against Fisher traced above was drawn only for equal weights. Repeated replications of the same design are the case where equal weights are right, which is why that is where this one began.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A distribution drawn from the null — both name kolmogorov–smirnov, monte carlo, p-value, statistical power, uniformity
- Estimating how many nulls are true — both name monte carlo, multiple comparisons, p-value, statistical power, uniformity
- A coverage table with its own error — both name discreteness, monte carlo, multiple comparisons, statistical power
- A horizon chosen after looking — both name monte carlo, multiple comparisons, statistical power
- An order that spends the error rate — both name monte carlo, multiple comparisons, statistical power
- Draws that repeat each other — both name discreteness, monte carlo, p-value
Named objects
A flat tag is an object no other essay names yet.
Chi squaredConservative testDiscretenessExact enumerationFisher's methodGlobal nullKolmogorov–SmirnovMid-p valueMonte CarloMultiple comparisonsp-valueStatistical powerStouffer's methodTippett's methodUniformity