Tests, and the second number

Studies that share a control

Ten comparisons against one control arm of the same size are correlated at exactly one half, and combined as though they were independent, Stouffer's sum rejects 24.23% of sets in which nothing works, Fisher's 17.88%, while Tippett's smallest p-value falls to 3.72%. Repaired with the correlation the design states, each is exact again — and the smallest of the three is then inflated less than for independent studies, 7.05% against 9.66%, because sharing makes the three combinations agree. The two inflations the dependence was expected to stack do not stack.

Worth reading first: Two ways to combine p-values.

The smallest of three combinations measured what happens when a report keeps whichever of Fisher’s, Stouffer’s and Tippett’s combined p-values is smallest: on ten independent studies of nothing it rejects 9.66% of the time, and it can be repaired by reading each combination at a common level of 2.448%. Every number in it rested on the studies being independent, and it named the obvious way they are not. Studies that compare different treatments against one control group share that group’s noise.

The dependence has an exact size. A treatment arm and a control arm of the same size give a z statistic that is the difference of their means over the standard error of that difference; two such comparisons that share the control share half of that difference’s variance, so their statistics are correlated at exactly one half. One control, many arms derived that correlation, n/(n+n0)n/(n + n_0) for arms of nn and a control of n0n_0, and priced the control’s value to the design. Here the question is the one that comes after the trial is reported: what the correlation does to the combinations a meta-analyst or a reader would apply to its ten p-values.

Ten comparisons read as though they were independent

The lead into this expected the three combinations to be pushed in different directions — Stouffer’s sum strongly, Tippett’s minimum hardly at all — and expected the smallest of the three to stack the inflations. The first half of that is right, and more extreme than expected.

Ten studies every study shares the control: the three combinations read as though independent, by the correlationCounted on 400,000 null sets of ten one-sided studies at each correlation, every study shares the control. Fisher: 5.00%, 9.54%, 12.83%, 15.01%, 16.63%, 17.88%, 18.80%. Stouffer: 5.00%, 11.51%, 16.29%, 19.54%, 22.20%, 24.23%, 25.86%. Tippett: 4.98%, 4.86%, 4.68%, 4.45%, 4.17%, 3.72%, 3.27%. Smallest of three: 9.66%, 14.44%, 17.94%, 20.41%, 22.55%, 24.34%, 25.88% — at correlations 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6. Each combination is read at its critical value for independent studies.00.1000.20000.2000.4000.600correlation between any two studies' statisticsshare of null sets rejectedFisherStoufferTippettthe smallest of the three5%400,000 null sets of ten at each correlationa shared control at equal arms is 0.5
Fig. 1 The share of null sets of ten one-sided studies rejected by each combination read at its independent-study critical value, against the correlation between any two studies. At a shared control with equal arms the correlation is 0.5: Stouffer’s sum rejects 24.23%, Fisher’s 17.88%, Tippett’s 3.72%.

At a correlation of 0.5 Stouffer’s combination rejects 24.23% of null sets at a nominal 5%. That number is a closed form. The sum of ten z statistics correlated at ρ has variance 10 + 90ρ rather than 10, so dividing by 10\sqrt{10} leaves a statistic with standard deviation 1+9ρ\sqrt{1 + 9\rho}, which at one half is 2.345, and a one-sided test at 1.645 then rejects whenever a standard normal exceeds 1.645⁄2.345 — nearly a quarter of the time. The count of 400,000 null sets agrees with that closed form, and a much smaller correlation does damage already: at 0.1, which a shared population or a shared laboratory could produce without anybody noticing, Stouffer’s combination rejects 11.51%.

Fisher’s combination is pushed the same way, less: 9.54% at a correlation of 0.1 and 17.88% at 0.5. Its statistic adds −2 log p over the studies, and the covariance between two of those terms grows with the correlation of the underlying z’s, so the sum’s variance grows too; its mean does not change, since each p-value is still uniform on its own. Tippett’s combination goes the other way. Its rule rejects when the smallest of the ten p-values is below 1−0.951/101 - 0.95^{1/10}, a cut set so that ten independent chances share the 5%; correlated chances overlap, so the probability that any of them falls below the cut is smaller, and Tippett’s test rejects 3.72% at a correlation of 0.5 — conservative, as the error rates of correlated tests generally are for rules built on the union of independent events.

The smallest of the three rejects 24.34% at a correlation of 0.5, barely more than Stouffer’s alone. The stacking the lead feared is real in the sense that every combination’s error lands in the union, but it barely adds anything: once Stouffer’s sum is that badly calibrated it is the combination that rejects, and the other two mostly agree with it when they reject at all. The gap between the smallest of three and Stouffer’s alone is 4.65 points for independent studies, 2.93 at a correlation of 0.1, and 0.12 at the shared control.

What each combination is hearing

The three combinations hear the shared noise differently, and the difference is visible in what each adds up. Stouffer’s sum adds the z statistics themselves, and their correlation passes straight into the sum’s variance: ten terms correlated at one half have a sum whose variance is 5.5 times what independence would give. Fisher’s sum adds −2log⁡p-2\log p, a convex transformation of each z that stretches the upper tail and compresses the lower one, and the transformation weakens the correlation slightly: two such terms built from z’s correlated at one half are correlated at 0.453, and at 0.1 at 0.083. Fisher’s sum therefore inflates its variance by a little less than Stouffer’s, and its reference distribution, a χ2\chi^2 on twenty degrees of freedom, is already skewed towards large values, so a given inflation moves its tail less. Both effects point the same way, and Fisher’s 17.88% sits below Stouffer’s 24.23%.

Tippett’s rule hears the shared noise as a reason to reject less. The largest of ten correlated z statistics is smaller, in distribution, than the largest of ten independent ones, because the correlated ten tend to rise and fall together and so offer fewer separate chances to be large. A rule that set its cut for ten separate chances is too strict for fewer. None of this is specific to a shared control: any positive correlation between studies — a common population, a common assay, a common analyst choosing the same analysis — pushes the two sums up and the minimum down, by amounts set by the correlation and by which studies it links.

The ordering is the same one two ways to combine p-values drew for real effects, arriving here for shared noise. There, a signal spread across many studies favoured Stouffer’s sum and a signal concentrated in a few favoured Fisher’s and Tippett’s. A shared control is noise spread across every study at once — the most evenly spread “signal” there is — so the combination that is most powerful against an even spread is the one most fooled by it, and the combination built to notice one study standing out is the one least fooled. Each combination’s weakness under dependence is its strength under an alternative, read in a mirror, and a choice of combination made for power is therefore also a choice of which dependence the result will be most sensitive to.

Two studies at a time

The dependence need not be total. A meta-analysis of ten trials might contain five two-arm-plus-control trials each contributing two comparisons, so that studies share in pairs and are independent across pairs.

Ten studies studies share in pairs: the three combinations read as though independent, by the correlation. Counted on 400,000 null sets of ten one-sided studies at each correlation, studies share in pairs. Fisher: 4.99%, 5.60%, 6.22%, 6.82%, 7.45%, 8.08%, 8.60%. Stouffer: 5.00%, 5.84%, 6.63%, 7.50%, 8.19%, 8.98%, 9.64%. Tippett: 5.00%, 4.97%, 4.95%, 4.93%, 4.85%, 4.77%, 4.62%. Smallest of three: 9.65%, 10.26%, 10.90%, 11.57%, 12.10%, 12.70%, 13.19% — at correlations 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6. Each combination is read at its critical value for independent studies.
Fig. 2 The same counts with the ten studies sharing in pairs — five pairs, each correlated at the stated value, independent of one another. At 0.5 Stouffer’s combination rejects 8.98% and the smallest of three 12.70%.

Sharing in pairs does a fraction of the damage. At a correlation of 0.5 within pairs Stouffer’s sum has variance 15 rather than 10, its combination rejects 8.98%, Fisher’s 8.08%, Tippett’s 4.77%, and the smallest of the three 12.70% — still worse than for independent studies, and now the stacking is visible, because no single combination dominates the union the way Stouffer’s does when every study shares. Which structure a set of studies has is not a detail. The same nominal correlation does five times the damage to Stouffer’s combination when it links every pair of studies rather than five.

The repairs the design allows

A shared control is the friendly kind of dependence, because the design states its size. With equal arms the correlation is one half; with a control of a different size it is n/(n+n0)n/(n + n_0); either way a combination can be told it, and each of the three has a repair.

Stouffer’s is exact: divide the sum by its true standard deviation, k+2∑i<jρij\sqrt{k + 2\sum_{i<j}\rho_{ij}}, instead of k\sqrt k. Tippett’s is exact as well. The largest of ten z statistics correlated at ρ through a common component has the distribution P(max⁡≤c)=∫φ(u) Φ ⁣((c−ρ u)/1−ρ)10 duP(\max \le c) = \int \varphi(u)\,\Phi\!\big((c - \sqrt\rho\,u)/\sqrt{1-\rho}\big)^{10}\,du, a one-dimensional integral, and reading the largest z against its quantile is Dunnett’s test for many arms against one control. Fisher’s has no exact repair in closed form, and the standard one is Brown’s: keep the statistic, but read it as a scaled χ2\chi^2 whose mean and variance match the statistic’s true ones. The covariance of two −2 log p terms at a correlation of ρ is a two-dimensional integral; at one half it makes the ten-study sum behave like a χ2\chi^2 on 3.94 degrees of freedom rather than twenty.

Ten studies every study shares the control: the three combinations repaired for the sharing, by the correlation. Counted on 400,000 null sets of ten one-sided studies at each correlation, every study shares the control. Fisher: 5.00%, 4.93%, 4.98%, 4.95%, 5.05%, 5.01%, 5.00%. Stouffer: 5.00%, 4.92%, 4.96%, 4.92%, 5.05%, 4.99%, 4.98%. Tippett: 4.98%, 4.99%, 5.01%, 5.04%, 5.06%, 5.04%, 5.01%. Smallest of three: 9.66%, 8.77%, 8.24%, 7.79%, 7.43%, 7.05%, 6.69% — at correlations 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6. Stouffer's sum is divided by its true standard deviation, Tippett's largest z is read against the correlated normals' own quantile, and Fisher's statistic is scaled by Brown's moment match.
Fig. 3 The same null sets with each combination repaired for the correlation the design states: Stouffer’s sum divided by its true standard deviation, Tippett’s largest z read against Dunnett’s quantile, Fisher’s statistic scaled by Brown’s moment match. All three hold 5%, and the smallest of the three falls as the studies share more.

All three repairs hold their level. At every correlation from 0.1 to 0.6 the repaired Stouffer and Tippett tests reject between 4.92% and 5.06% of null sets, which is 5% to within the count’s error, and Brown’s Fisher stays within the same band — its approximation is good enough at ten studies that the count cannot tell it from exact.

And the smallest of the three repaired combinations moves the wrong way for the lead’s prediction. For independent studies it rejects 9.66%. With every study sharing it rejects 8.77% at a correlation of 0.1, 8.24% at 0.2 and 7.05% at the shared control’s 0.5. Dependence between the studies makes the three combinations more alike. When every study carries the same shared noise, all three summaries of the ten p-values are driven by it together: a set that has a large shared component looks significant to Stouffer’s sum, to Fisher’s sum and to the smallest p-value at once, and a set without one looks unremarkable to all three. The union of three tests that mostly agree is barely larger than any one of them.

The level the smallest of three needs

The repair for keeping the smallest of three is the calibration of the earlier essay, read off the counted null distribution of the smallest repaired p-value: the level at which it falls below the level 5% of the time.

The common level at which the smallest of three repaired combinations has size 5%, by the correlation between studies. The level, read off 400,000 counted null sets of ten, at which the smallest of the three repaired combined p-values falls below it 5% of the time. Every study sharing: 2.449%, 2.723%, 2.904%, 3.094%, 3.229%, 3.449%, 3.681%. Studies sharing in pairs: 2.457%, 2.471%, 2.527%, 2.526%, 2.543%, 2.572%, 2.614% — at correlations 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6. For independent studies the level is 2.449%; the further the studies share, the closer it moves to 5%, because the three repaired combinations agree more.
Fig. 4 The common level at which each repaired combination must be read for the smallest of the three to have size 5%, against the correlation, for every study sharing and for studies sharing in pairs. Both start at the independent studies’ 2.449%; sharing raises it, much more when every study shares.

For independent studies, counted on these sets, the level is 2.449%, the earlier essay’s 2.448% to within the count. With every study sharing it rises to 2.723% at 0.1, 2.904% at 0.2 and 3.449% at the shared control’s 0.5. In pairs it barely moves: 2.572% at 0.5. So the analyst who models the shared control and keeps the smallest of three can read each combination at nearly 3.5% rather than 2.4%, which is power returned for free by the same correlation that would otherwise have tripled the error rate.

Reading the repaired combinations at the independent-study level, 2.448%, is therefore safe and wasteful: the smallest of three then rejects 3.59% of null sets at the shared control. That is the mirror image of the earlier essay’s finding. Calibrating for independence when the studies share gives a test that is conservative; ignoring the sharing altogether gives one that rejects a quarter of the time. Both mistakes come from using one correlation’s arithmetic on data with another.

What the sharing costs

The repairs restore the error rate. They cannot restore the information, and for a common effect the loss has a closed form: a sum of ten correlated statistics estimates a shared mean with the precision of k2/(k+2∑ρ)k^2/(k + 2\sum\rho) independent studies.

How many independent studies ten sharing studies are worth to a combination of their effects. For a common effect, Stouffer's repaired statistic has the precision of a number of independent studies equal to the square of the count over the count plus twice the sum of the correlations between pairs. Ten studies all sharing at a correlation of 0.5 are worth 1.82; in pairs they are worth 6.67. At 0.2 the two are 3.57 and 8.33.
Fig. 5 How many independent studies ten sharing studies are worth to an estimate of a common effect, against the correlation, for every study sharing and for studies sharing in pairs. At the shared control’s 0.5, ten studies that all share are worth 1.82.

Ten comparisons against one shared control at equal arms are worth 1.82 independent studies for estimating a common effect, and ten studies sharing in pairs at the same correlation are worth 6.67. At a correlation of 0.2 the two figures are 3.57 and 8.33. The first number is the reason the naive combinations fail so badly: a combination that thinks it has ten studies’ worth of evidence and has under two will find significance that is not there. It is also the reason a multi-arm trial reported as ten comparisons is not ten confirmations of anything shared — a reader adding up ten significant arms against one control should know that their common part is one measurement of one control group.

It also says how large the evidence has to be. Read as though independent, Stouffer’s combination of ten one-sided studies rejects when the z statistics add up to more than 5.20; repaired for a shared control at equal arms, they have to add up to more than 12.20, nearly two and a half times as much. A platform trial that reports ten arms each with z near 0.6 has a naive combined result just significant at 5% and a repaired one nowhere near it, and the difference between the two is entirely the control group the ten arms have in common.

The number is about a common effect, and the arms of a trial usually test different treatments. For a question about any of them working — the question Tippett’s combination asks — the shared control costs much less, and in fact gives something back: Dunnett’s critical value for the largest of ten z statistics correlated at one half is 2.448, against 2.568 for ten independent ones, because correlated arms cannot all be lucky independently and so need a smaller margin to be believed one at a time. Borrowing a control from earlier trials meets the same arithmetic from the other side, where the shared component is the thing being borrowed. Which loss applies depends on the question, and so does which combination should be used. That is the earlier essay’s point about choosing a combination, now with a second reason the choice matters.

What a combination of shared studies should report

The sharing structure and its correlation. Which studies share a control, a population or an analyst, and the correlation that implies — n/(n+n0)n/(n + n_0) for a shared control, an estimate otherwise. Without it no combined p-value can be read, because the same ten numbers give 5% or 24% depending on it.

Which repair was used. Stouffer’s with the true variance and Tippett’s against the correlated quantile are exact; Brown’s Fisher is an approximation that holds well here. A combination read at its independent-study critical value on shared studies is not a 5% test.

The level, if the smallest of several is kept. It depends on the correlation and on the structure, and for shared studies it is higher than for independent ones.

Counted, and closed where it could be

Every rejection rate is counted on 400,000 null sets of ten one-sided studies at each correlation, one stated seed for each configuration, so the standard error of a rate near 5% is about 0.034 points and of a rate near 24% about 0.07. Stouffer’s naive size and its repair are closed forms and the counts agree with them — two routes to one number that share nothing but the question; Tippett’s repair is a one-dimensional integral; Brown’s moments come from a two-dimensional quadrature of the log p-values’ covariance. The calibrated levels are read off the counted distribution of the smallest repaired p-value, so they carry the count’s error, about a hundredth of a percentage point. Only two structures are measured, and only correlations that are positive and equal within a structure; a meta-analysis with some studies sharing and others not, or with correlations that differ by pair, needs its own count, and the repairs above extend to it directly because each uses only the matrix of correlations.

Still open: a correlation that has to be estimated

A shared control states its correlation. A shared population, a shared assay or a shared analyst does not, and then the correlation is an estimate — from the studies’ overlap if it is known, from the dispersion of their effects if it is not — and every repair above is computed with an estimated number.

The sensitivity is easy to read off the figures and uncomfortable: Stouffer’s combination is off by more than six points at a correlation of only 0.1, so a repair that underestimates a true 0.2 as 0.1 would still reject well above 5%. Whether a correlation estimated from ten studies is precise enough to repair anything, how the error of the estimate propagates into each combination’s size, and whether the smallest of three — whose inflation shrinks as sharing grows — is the more forgiving of an error in the correlation than any single combination, is the measurement this leaves. It joins the price of control on a familiar point: a correction is only as good as the quantity it corrects for, and a p-value is only uniform when the assumptions behind its reference distribution hold.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formCorrelationDependenceEffective sample sizeFisher's methodGlobal nullMeta analysisMonte CarloMultiple comparisonsp-valueShared controlStouffer's methodTippett's methodUniformity