Tests, and the second number

Two ways to combine p-values

Fisher's and Stouffer's combinations are both exactly right when every null is true, for the single reason that each p-value is flat. Under a real effect they disagree about which evidence counts: with Stouffer held at 50% power across ten studies, Fisher is the more powerful while the signal sits in six or fewer of them and the less powerful from seven.

Worth reading first: A p-value that is not flat is not a p-value · What the correction corrects.

Ten studies of one question each report a p-value, and nothing else is available: no raw data, no effect estimates on a common scale, only ten numbers between zero and one. Two different questions can be put to them, and they are easy to confuse.

The first is which of the ten found something. That is the multiplicity question, and its corrections answer it by controlling how often a study is wrongly declared a finding — the familywise rate or the false discovery rate, two different promises that cost different amounts of power. The second is whether anything at all is going on: whether the ten together are evidence against the hypothesis that every one of their nulls is true. That is the global null, and a combination of p-values tests it. Rejecting it names no study, and a correction’s output cannot answer it.

There are three standard combinations and they are all exact, meaning each rejects a true global null at exactly its stated rate. Each is exact for the reason a p-value that is not flat is not a p-value spends its whole length on: under its own null every p-value is uniform. Under an alternative they stop agreeing, and the disagreement is not noise. It is a statement about which arrangement of evidence each one is built to find, and it can be computed rather than argued about.

Three transformations of a flat number

Every combination turns each uniform p-value into a variable with a known distribution, then adds. The p-values here are one-sided, since combining evidence about an effect needs a direction for the effect.

Fisher’s method uses 2logp-2\log p. If pp is uniform, 2logp-2\log p is exponential with mean two, which is a chi-squared on two degrees of freedom, so the sum over k studies is a chi-squared on 2k. At ten studies the 5% critical value is 31.410.

Stouffer’s method uses the normal score z=Q1(p)z = Q^{-1}(p). If pp is uniform, zz is standard normal, so zi/k\sum z_i/\sqrt{k} is standard normal and is compared with 1.645.

Tippett’s method uses the smallest p-value alone. The minimum of k independent uniforms falls below cc with probability 1(1c)k1-(1-c)^k, so the critical value is 10.951/k1 - 0.95^{1/k}: 0.005116 at ten studies, which is a single study’s statistic beyond 2.568.

Nothing in any of the three depends on what the studies measured or how large they were. All three are exact because the inputs are uniform and independent, and for no other reason.

20,000 combined p-values from 10 studies of nothing, combined two ways. Each trial draws 10 uniform p-values, as true nulls produce, and combines them by Fisher's method against a chi-square on 20 degrees of freedom and by Stouffer's against a standard normal. Left half of each bin Fisher, right half Stouffer. Both are flat: Kolmogorov–Smirnov distances 0.0045 and 0.0044, rejection rates 4.92% and 4.90%. Each is exact because its inputs are uniform, and for no other reason.
Fig. 1 Twenty thousand sets of ten uniform p-values, each combined both ways. The left bar of each pair is Fisher’s combined p-value and the right bar is Stouffer’s. Both are flat, which is what makes each of them a p-value.

The figure holds the output to the check a single test is held to. The Kolmogorov–Smirnov distances from uniform are 0.0045 for Fisher and 0.0044 for Stouffer, with p-values of 0.81 and 0.84. The rejection rates are 4.92% and 4.90%, against a Monte Carlo standard error of 0.15 points. A combined p-value that is not flat is not a p-value either, and these two are flat.

Two studies, drawn where their statistics live

The figure below draws all three tests for two studies in the plane of their z statistics. The grey cloud is a sample of null pairs, and each test rejects in a different region.

Three ways to reject with two studies, drawn where the two z statistics live. Two one-sided studies, each summarised by its z statistic. Fisher's combination rejects outside a curve that runs parallel to both axes, so one study past z = 2.378 decides it alone; Stouffer's rejects above the straight line z₁ + z₂ = 2.326; Tippett's rejects when either z passes 1.955. Each region holds exactly 5% of the standard bivariate normal — Fisher's in closed form, e^(−c/2)(1 + c/2) at c = 9.488 — and of 100,000 counted null pairs they catch 4.95%, 4.88% and 5.04%. Two alternatives carry the same Stouffer evidence: one study at 2.326 and the other at nothing, where the powers are 62.7%, 50.0% and 65.4%; and both at 1.163, where they are 47.7%, 50.0% and 38.3%.
Fig. 2 Two studies’ z statistics, with a sample of null pairs in grey. Stouffer rejects above the straight line, Fisher outside the curve and Tippett inside the L, and each region holds exactly 5% of the null. The two marked points carry the same Stouffer evidence, once as one strong study and once as two weak ones.

Stouffer’s region lies above a straight line, z1+z2=2.326z_1 + z_2 = 2.326. It cares about the sum and nothing else, so a pair at (2.3, 0) and a pair at (1.16, 1.16) are equally convincing to it.

Fisher’s region lies outside a curve on which p1p2=0.008705p_1 p_2 = 0.008705, ec/2e^{-c/2} with c=9.488c = 9.488. The curve runs parallel to both axes, and past z=2.378z = 2.378 one study rejects on its own, whatever the other reports — even a strongly negative z. Fisher is willing to be convinced by one result.

Tippett’s region is an L. It rejects when either statistic passes 1.955, the level 0.02532 of a single test, and it ignores the other study completely.

Each region holds exactly 5% of the standard bivariate normal. For Fisher that is closed form: the chi-squared on four degrees of freedom has upper tail ec/2(1+c/2)e^{-c/2}(1 + c/2), and at c=9.488c = 9.488 that is 0.05 to thirteen decimal places. Counted on a hundred thousand null pairs, the three regions catch 4.95%, 4.88% and 5.04%, with a standard error of 0.07 points. The three shapes are different because they spend the same 5% in different places, and where each spends it decides what it can find.

The two marked points carry the same Stouffer evidence. One is a single study at a shift of 2.326 with the other at nothing. The other is both studies at 1.163. Stouffer has exactly 50.0% power at both. At the single strong study Fisher has 62.7% and Tippett 65.4%. At the two weak ones Fisher has 47.7% and Tippett 38.3%.

For a common shift in every study, Stouffer’s test is not merely one choice among three. The likelihood ratio for a shared shift depends on the data only through zi\sum z_i, so by the Neyman–Pearson lemma no test of the same size is more powerful against that alternative. That is a theorem, and it holds only for equal shifts in a known direction. The question with content is how fast the advantage goes once the shifts are not equal.

One signal, spread over more and more of ten studies

Hold the total evidence fixed and move it around. Ten studies; m of them carry a shift μ and the rest carry nothing; and μ is set for each m so that Stouffer’s power is exactly 50%, which makes mμm\mu the same at every m. At m = 1 the one informative study is shifted by 5.2015 standard errors. At m = 10 all ten are shifted by 0.5201. Stouffer cannot tell these arrangements apart, by construction. Fisher and Tippett can.

Stouffer’s power is closed form, and so is Tippett’s, as a product of normal tails. Fisher’s is not, and it is computed two ways that share nothing.

The lattice. A shifted study contributes Y=2logQ(Z)Y = -2\log Q(Z) with ZN(μ,1)Z \sim N(\mu, 1), and YY has a closed-form distribution function, P(Yy)=Φ(Q1(ey/2)μ)P(Y \le y) = \Phi\big(Q^{-1}(e^{-y/2}) - \mu\big). The m shifted contributions are put on cells a hundredth wide and convolved. The ten minus m null contributions are the exact chi-squared. A sum that lands in cell s truly lies somewhere between s and s + m cells, so reading the chi-squared tail at both ends gives a lower and an upper bound. Fisher’s power is enclosed, not estimated, and at no signal the enclosure has to contain the size: it reads 4.982% to 5.018%.

The count. Twenty thousand sets of ten studies at each m, the same statistic computed on each.

One signal spread over 10 studies, and which combination finds it. The same total evidence — enough to give Stouffer's combination 50% power at every spread — placed in m of 10 one-sided studies: a shift of 5.201 standard errors in one study, down to 0.520 in each of all 10. Fisher's power, computed on a lattice whose error is enclosed rather than estimated, falls from 96.1% to 45.0%; Tippett's minimum-p test falls from 99.6% to 18.5%. Fisher beats Stouffer while the signal sits in 6 or fewer of the 10 studies and loses from 7 on. Open circles are Fisher's power counted on 20,000 trials a point.
Fig. 3 Power against the number of studies sharing one fixed total signal, with Stouffer’s power held at 50%. Fisher’s line is the midpoint of its enclosure and the open circles are its counted power. Tippett’s dashed line starts highest and falls fastest.

Fisher’s power is 96.05% with the signal in one study, 77.13% in two, 64.90% in three, 57.95% in four and 53.64% in five. With the signal in six studies its enclosure is 50.62% to 50.90%, above Stouffer’s 50%. With the signal in seven it is 48.54% to 48.87%, below it. At ten it is 45.03%. The crossing sits between six and seven, and it is located rather than estimated: at no m does the enclosure straddle the line it is compared with.

Tippett is the same story pushed further. It reaches 99.60% when one study holds everything and falls to 18.54% when ten share it. It beats Fisher at one and at two studies (77.25% against 77.13%) and loses from three.

The counts agree with the enclosures everywhere, with one detail worth stating rather than smoothing over. At m = 1 the counted power is 96.13% with a standard error of 0.14 points. At m = 10 it is 45.91% against an enclosure whose top is 45.26%, which is 1.9 standard errors high. On the same draws Stouffer’s count is 50.73% against an exact 50%, also high, by 2.1 standard errors. Both statistics are read off the same twenty thousand sets of studies, so one lucky batch moves both counts the same way, which is what shared draws do.

A higher bar moves the crossing, and not by much

The crossing is not a property of the combinations alone. It depends on how strong the total signal is, which the next figure changes.

One signal spread over 10 studies, and which combination finds it. The same total evidence — enough to give Stouffer's combination 80% power at every spread — placed in m of 10 one-sided studies: a shift of 7.863 standard errors in one study, down to 0.786 in each of all 10. Fisher's power, computed on a lattice whose error is enclosed rather than estimated, falls from 100.0% to 74.8%; Tippett's minimum-p test falls from 100.0% to 31.7%. Fisher beats Stouffer while the signal sits in 7 or fewer of the 10 studies and loses from 8 on. Open circles are Fisher's power counted on 20,000 trials a point.
Fig. 4 The same ten studies with Stouffer held at 80% power. Every curve lifts, and Fisher’s still crosses Stouffer’s, one study further along.

With Stouffer at 80%, Fisher is ahead through seven studies, at 80.30%, and behind from eight, at 78.09%. With the signal in one study it has 99.998%, and with it in all ten it has 74.75%; Tippett falls to 31.70%. A stronger total signal gives the concentrated arrangement a little more room. The direction of the trade does not change, and neither does the reason: Fisher’s region reaches out along the axes, and a stronger signal lets more of the spread-out arrangements get far enough along them.

Where Fisher stops leading as the studies multiply

At ten studies the crossing is at six. The obvious guesses are that it stays at a fixed fraction of the studies, or that it grows like the square root of their number, which is how the Stouffer statistic’s noise grows. Neither is right.

Where Fisher stops leading, as the number of studies grows. For each number of studies k, the largest number m of them the same total signal can be spread over while Fisher's combination is still more powerful than Stouffer's, with Stouffer held at 50% power. Every point is located by a bracketed power computation that does not straddle the comparison. It is 1 of 2, 6 of 10 and 18 of 40: a falling share of the studies — 75% at four, 60% at ten, 45.0% at 40 — and a count growing faster than the square root of k. Neither a fixed fraction nor √k describes it.
Fig. 5 For each number of studies, the most of them the fixed signal can be spread over with Fisher still the more powerful combination. The dashed reference is half the studies and the pale curve is the square root. The measured points climb faster than the square root and fall steadily behind the half.

The crossing is 1 of 2, 2 of 3, 3 of 4, 3 of 5, 4 of 6, 5 of 8, 6 of 10, 7 of 12, 8 of 15, 10 of 20, 12 of 25, 14 of 30 and 18 of 40. As a share of the studies it drifts down, from 75% at four to 60% at ten and 45.0% at forty. Divided by the square root of the number of studies it rises, from 0.707 at two to 2.846 at forty. With Stouffer held at 80% instead, the crossing at forty studies is 21.

Each point is located by the same enclosure. At twenty studies and above, cells a fiftieth wide were too coarse at the rows nearest the crossing, so those rows alone were recomputed on cells five times finer until the enclosure cleared the comparison. There is no formula here. The points are measurements of one family of alternatives — equal shifts in m studies and nothing in the rest — at two strengths. They say nothing about signals of unequal size, and nothing about how the crossing behaves beyond forty studies.

The choice is a statement about the alternative

So the three combinations are not better and worse versions of one test. Each is a claim about where the evidence will be.

Stouffer’s is the right claim when the studies measure one shared effect, as replications of a single experiment do, and it is the most powerful test there is for that case. Fisher’s is the right claim when the effect might be present in some populations and absent in others. Tippett’s is the right claim when one study might carry the whole signal and the rest are noise. The regions in the two-study figure are those three claims drawn.

This is where the combination inherits a problem from the essay on analyses of nothing: the claim has to be made before the p-values are seen. Computing all three combinations and reporting the smallest is a fourth test. Its size is not 5% and has not been computed here. The choice of combination is part of the result, in the same way a stopping rule is part of a p-value.

Two-sided inputs are exact, and blind to direction

A two-sided p-value is uniform under its null too, so a combination of two-sided p-values is exactly as exact. Counted on eight thousand sets of ten null studies, Stouffer’s combination of them rejects 4.58% of the time, with a standard error of 0.24 points around 5%.

What it is exact about is a hypothesis with no direction in it. Shift every one of the ten studies by half a standard error upwards and it rejects 13.14% of the time. Shift every one downwards by the same amount and it rejects 12.35% — the same rate within the counting error, since the inputs cannot see the sign. Stouffer’s one-sided combination of the same upward shift has 47.46% power. A combination of two-sided p-values is legitimate, and it answers a question a study of one effect does not normally ask.

A small non-uniformity, combined ten times

The exactness depends on each input being flat, and the essay on flatness lists the ways an honest-looking p-value stops being flat. One of them is a variance divided by n instead of n − 1. Feed that to the combinations and they do not treat it the same way.

Take a one-sided t-test on five observations with that slip in it. On two hundred thousand null studies it rejects 6.46% of the time, with a standard error of 0.05 points. That is a real excess and a modest one. Combine ten of them by Fisher’s method and the size is 7.92%, standard error 0.19. Combine the same ten by Stouffer’s and it is 6.59%, standard error 0.18, which is the single test’s excess carried through, neither shed nor multiplied.

The difference is where each combination reads its inputs. The slip stretches every t statistic by one factor, √(5/4). A stretch by a constant factor moves a normal score by close to a constant fraction, which Stouffer adds up. It moves a far tail probability by a factor that grows the further out it is read, and Fisher’s statistic is dominated by the smallest p-values. Fisher reads exactly the region where a slightly wrong test is most wrong. The opposite claim — that a combination of these p-values keeps its 5% size — fails when it is computed.

The broader consequence reaches past this one slip. A combination has no way to see that its inputs are not flat, so it inherits every defect of every input. Selection is the extreme case. A set of p-values that is only the published significant ones is as far from uniform as a set can be, and the winner’s curse has already priced what that selection does to the estimates. Combining such a set tests nothing.

When the p-values cannot be flat

Some correct tests cannot produce a uniform p-value at all. The exact one-sided binomial test on ten trials at a null probability of one half can return only eleven p-values, and on its own it rejects 1.07% of the time at a nominal 5%, which is the conservatism of a discrete test that an interval’s coverage also shows. The mid-p value, which counts only half the probability of the observed outcome, rejects 5.47% of the time — closer to nominal, and on the wrong side of it.

Here the size of a combination does not have to be simulated. Each study’s p-value takes one of eleven values with a known probability, so every multiset of outcomes can be enumerated with its multinomial weight — 184,756 of them at ten studies — and the size is a sum. It is the move that counts an interval’s coverage over every possible sample rather than over a simulated few.

The size of Fisher's combination when every p-value comes from a ten-trial binomial test. The exact one-sided binomial test on ten trials at a null probability of one half can return only eleven p-values, and alone it rejects 1.07% of the time at a nominal 5%; its mid-p version rejects 5.47%. Combining k of them by Fisher's method, with every multiset of outcomes enumerated, the true size is 1.07%, 1.59%, 1.67%, 1.33%, 1.12%, 0.94%, 0.71%, 0.55% at k = 1, 2, 3, 4, 5, 6, 8, 10 from exact p-values, and 5.47%, 3.36%, 3.76%, 4.26%, 4.04%, 4.01%, 3.98%, 3.89% from mid-p values. At ten studies: 0.55% and 3.89%. Open circles are sizes counted on 40,000 simulated sets, with two standard errors.
Fig. 6 The exact size of Fisher’s combination of ten-trial binomial tests against the number combined, for the ordinary p-value and for the mid-p value. The open circles are sizes counted on simulated sets of studies, with bars of two standard errors, beside the enumerated points at two, five and ten.

From ordinary p-values Fisher’s size is 1.07%, 1.59%, 1.67%, 1.33%, 1.12%, 0.94%, 0.71% and 0.55% at 1, 2, 3, 4, 5, 6, 8 and 10 studies. It rises before it falls, because the eleven possible values combine into a lattice of totals whose position against the critical value moves with k. From mid-p values it is 5.47%, 3.36%, 3.76%, 4.26%, 4.04%, 4.01%, 3.98% and 3.89%. The mid-p value is anti-conservative alone and conservative in every combination measured. The counted route agrees: 1.59% against an enumerated 1.59% at two studies, and 0.54% against 0.55% at ten. The mid-p counts are 3.34% against 3.36% at two studies and 3.74% against 3.89% at ten, the last 1.6 standard errors low.

The size of Stouffer's combination when every p-value comes from a ten-trial binomial test. The exact one-sided binomial test on ten trials at a null probability of one half can return only eleven p-values, and alone it rejects 1.07% of the time at a nominal 5%; its mid-p version rejects 5.47%. Combining k of them by Stouffer's method, with every multiset of outcomes enumerated, the true size is 1.07%, 2.07%, 1.98%, 0.88%, 0.81%, 0.71%, 0.51%, 0.37% at k = 1, 2, 3, 4, 5, 6, 8, 10 from exact p-values, and 5.47%, 5.77%, 4.94%, 4.08%, 4.02%, 4.62%, 4.65%, 4.44% from mid-p values. At ten studies: 0.37% and 4.44%.
Fig. 7 The same enumeration for Stouffer’s combination. The ordinary p-value is more conservative still, and the mid-p value crosses 5% at two studies.

Stouffer’s method does worse with the ordinary p-value. Its size is 1.07%, 2.07%, 1.98%, 0.88%, 0.81%, 0.71%, 0.51% and 0.37% at the same numbers of studies, and there is a structural reason for the collapse. A study that observes no successes has a p-value of exactly one, whose normal score is minus infinity, so one such study vetoes the whole combination whatever the other nine say. From mid-p values the size is 5.47%, 5.77%, 4.94%, 4.08%, 4.02%, 4.62%, 4.65% and 4.44%: above nominal at two studies, below it from three.

Neither choice of p-value gives a combination whose level is known in advance. For discrete inputs the level has to be computed, by enumeration as here or by an exact reference distribution for the combined statistic, and the conservative direction is not guaranteed by using the ordinary p-value either.

What is exact here, what is enclosed, and what is only counted

Exact: the null distribution of each combination, the two-study regions’ probabilities, and the size of every discrete combination, which is a finite sum over enumerated outcomes.

Enclosed: Fisher’s power under an alternative, between bounds that hold whatever the lattice’s cell width. Stouffer’s and Tippett’s powers are closed form.

Counted, and only as a second route: every one of those numbers is also checked against simulated studies, and the counts are held to their own standard errors. So are the measurements that have no other route: the two-sided rejection rates and the wrong-divisor sizes.

Assumed throughout and not tested: independence between studies. Studies that share a control group, a population or an analyst do not produce independent p-values, and all three exactness arguments use independence as heavily as they use uniformity. Nothing here measures what dependence does.

Where this goes next

The most direct continuation is the one named above and left uncomputed. A meta-analyst who computes Fisher’s, Stouffer’s and Tippett’s combinations and reports whichever is smallest has run a fourth test, and its size is a measurable quantity. It has to exceed 5%, and it is strictly less than three times that, since the three rejection regions overlap heavily. How much less is the number. It is distinct from twenty analyses of one data set because the choice here is among ways of reading the same ten p-values, so the three statistics are strongly correlated, and the inflation should be far smaller than the forking-paths figure.

The second continuation is the discrete one. A randomised p-value, which breaks ties with an auxiliary uniform, is exactly flat, so a combination of randomised p-values is exactly exact again. The price is power and a result that depends on a coin, and the size of that price has not been measured.

The third is weights. Stouffer’s statistic is usually weighted by each study’s size, which is the optimal weighting for a shared effect on a common scale, and the trade against Fisher traced above was drawn only for equal weights. Repeated replications of the same design are the case where equal weights are right, which is why that is where this one began.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Chi squaredConservative testDiscretenessExact enumerationFisher's methodGlobal nullKolmogorov–SmirnovMid-p valueMonte CarloMultiple comparisonsp-valueStatistical powerStouffer's methodTippett's methodUniformity