Significant in one, not in the other
Worth reading first: Twenty intervals and one expected miss.
Two intervals that overlap priced the reading of two intervals against each other. The reading made more often in a results section compares each interval with zero instead: the effect is significant in one group and not in the other, and the sentence that follows is that it is present in the first and absent in the second — the drug works in younger patients and not in older ones, the intervention helped in one region and not in the other.
That sentence is a claim about a difference between two effects. What was compared was two verdicts, each against zero. The two are related only loosely, and the distance between them is large enough to measure.
Identical effects disagree half the time
Take two independent studies of exactly the same effect, each with power to detect it at 5%. Each is significant or not independently, so exactly one of the two is significant with probability
which is largest — 50% — when the studies have 50% power. At 80% power it is 32.1%; at 17% power, 28.2%. Two faithful studies of one effect disagreeing about significance is not an anomaly that needs explaining. At the moderate power most studies are run at it is the commonest outcome there is.
The hero figure shows it as geometry. Each pair of studies is a point in the plane of their two z statistics, scattered around the diagonal where both equal the true mean. The lines at 1.96 cut the plane into four quadrants, and the split pairs are the two off-diagonal quadrants — which, when the true mean sits at 1.96, is half of the scatter by symmetry.
When they disagree, the difference is rarely significant
The test of whether two effects differ looks at the difference of the two z statistics, divided by . In the figure that is the distance of a point from the diagonal, and the difference is significant only outside the two diagonal lines, at .
The split quadrants and the region outside the diagonal lines barely meet. Given a split, the difference is significant:
| power of each study | exactly one significant | difference significant, given a split |
|---|---|---|
| 17.0% | 28.2% | 14.95% |
| 50.0% | 50.0% | 9.75% |
| 80.0% | 32.1% | 13.45% |
| 93.8% | 11.6% | 24.03% |
At 50% power, nine splits in ten have a difference that is not remotely significant. The effects are identical by construction, so every significant difference is a false positive, and the 9.75% is not a detection rate but the false-positive rate of the difference test inside the split. What the table says is that the split and the difference test are close to unrelated: the split is common and the difference is rare, and knowing that the first occurred barely changes the chance of the second.
The conditional rate climbs at high power because a split there is rare and needs one study to fall far below its expected z, which is also what makes the gap between the two large. At 94% power a quarter of splits have a significant difference — and they are still, every one, false positives about identical effects.
A p-value beside another p-value
The same arithmetic applied to the numbers a reader actually sees. Two independent estimates on the same side, with two-sided p-values and , have z statistics and , and their difference has p-value :
| first result | second result | p-value of the difference |
|---|---|---|
| p = 0.01 | p = 0.20 | 0.360 |
| p = 0.001 | p = 0.10 | 0.245 |
| p = 0.04 | p = 0.06 | 0.903 |
A result at 0.01 beside one at 0.20 is a pair a results section would describe as a clear effect in one group and none in the other, and the evidence that the two groups differ is a p-value of 0.36. A result at 0.04 beside one at 0.06 falls on opposite sides of the line and is as close to identical as two noisy estimates get.
A p-value alone does not say how large an effect is. Two p-values do not say how different two effects are either, because the threshold that turns each into a verdict is applied separately, and the difference between two verdicts is not a verdict about a difference.
The gap the split manufactures
The split does more than fail to indicate a difference. It creates the appearance of one, through the selection that inflates a significant estimate.
A study that reached significance is, on average, one that landed above its true mean; a study that did not is, on average, one that landed below it. So among split pairs the significant estimate is pulled up and the non-significant one pulled down, and the pair shows a gap that neither the truth nor either study’s design contains:
| power of each study | significant study’s average z | non-significant study’s average z | apparent gap, as a share of the true effect |
|---|---|---|---|
| 17.0% | 2.49 | 0.70 | 179% |
| 50.0% | 2.76 | 1.16 | 81% |
| 80.0% | 3.15 | 1.40 | 62% |
At 50% power the average split pair shows the significant group’s effect 81% larger, relative to the true effect, than the non-significant group’s — with the true effects identical. At 17% power the manufactured gap is larger than the effect itself. The effect that looks present in one group and absent in the other is the same effect, seen once through a filter that inflated it and once through a filter that shrank it.
Reading the difference off a published table
The test of the difference rarely appears in print, but it can almost always be recovered from what does. A 95% interval’s width is standard errors, so each group’s standard error is its interval’s width divided by 3.92, and the difference has the root sum of squares of the two.
Suppose a paper reports an effect of 4.0 in one group with a 95% interval from 1.2 to 6.8, and an effect of 2.0 in another with an interval from −1.6 to 5.6. The first is significant at p = 0.0051 and the second is not, at p = 0.276, and the paper says the treatment works in the first group. The standard errors are 1.429 and 1.837; the difference is 2.0 with a standard error of 2.327, a 95% interval from −2.56 to 6.56, and a p-value of 0.390.
That interval is the honest summary of what the two groups say about each other. It runs from the second group’s effect being two and a half units larger to its being six and a half smaller, and a reader who had it in front of them would not have written the sentence about the treatment working in one group only. It takes one subtraction, one square root and one division, and it assumes the two groups are independent — which, for subgroups defined by a baseline characteristic in a randomised trial, they are.
A replication that “failed”
The same split has a more consequential name when the two studies are an original and its replication. The original was significant; the replication was not; the replication “failed”.
If the original was selected for significance and the replication has 50% power, the replication is non-significant half the time with the effect perfectly real — the p-value a replication gets makes the same point from the replication’s side. A failed replication is therefore not, by itself, evidence that the two studies disagree. The comparison that would be evidence is the one in this essay’s second table: the difference between the original’s estimate and the replication’s, against its own standard error. And because the original was selected, its estimate was inflated by exactly the gap tabulated above, so even that comparison leans towards finding a disagreement that the truth does not contain.
Five times in six found that a replication’s estimate lands outside the original’s interval one time in six with nothing wrong. Here the verdict “significant, then not” is reached half the time with nothing wrong. Both are readings of two estimates of one quantity through a device built for one estimate at a time.
Two non-significant results that agree
The split has a mirror image that is misread in the other direction. Two studies each report an effect with p = 0.134, neither significant, and the literature records two failures to find anything.
If the two estimates have the same size and the same precision, their combined evidence is the sum of their z statistics divided by — Stouffer’s combination, which the essay on combining p-values measures — and two z statistics of 1.5 combine to 2.12, a p-value of 0.034. Two non-significant studies that agree can be, together, a significant finding. A tally of verdicts records them as two negatives; the estimates record them as one positive.
So the verdict of each study against zero misleads in both directions at once: a significant and a non-significant result are read as a difference that is not there, and two non-significant results are read as an absence that the pair does not support. Both errors come from the same step, which is discarding the estimates and keeping the verdicts.
Many subgroups, and the certainty of a split
A trial rarely compares two groups. It reports its effect by sex, by age band, by site, by baseline severity, and the question becomes whether some subgroup is significant while some other is not.
With subgroups of the same effect, each with power , the chance that they are neither all significant nor all non-significant is :
| subgroups | at 50% power | at 80% power |
|---|---|---|
| 2 | 50.0% | 32.0% |
| 4 | 87.5% | 58.9% |
| 8 | 99.2% | 83.2% |
Among eight subgroups of one effect, a split is all but certain at any moderate power. Subgroups also have less power than the whole study — halving the sample halves the information — so a trial powered at 80% overall has subgroups at well under 80% and sits in the left-hand column rather than the right. The arithmetic is exact for two equal halves: a whole-trial z statistic with mean 2.80, which is 80% power, gives each half a z with mean 1.98 and 50.8% power — almost exactly the power at which two halves split most often. A trial sized conventionally for its overall effect is, by that sizing, a trial whose two halves disagree about significance half the time. Splitting into four quarters rather than two halves lowers each one’s power to 28.8% and leaves a 73.7% chance that at least one quarter is significant and at least one is not. A subgroup table in which the effect is significant in some rows and not in others is therefore the expected shape of a subgroup table for a treatment that works the same way in everyone, and it is not evidence of anything until the differences themselves have been tested.
That is also twenty analyses of nothing in another costume. Each subgroup’s test is fine on its own; the family of verdicts is what produces a finding, and the finding is a property of the family.
The test that answers the question, and what it costs
The question the split is standing in for has its own test: whether the effect differs between the groups, which is a test of the difference between two estimates and is often called a test of interaction. It is not an obscure procedure. It is expensive.
Split a study into two equal halves. Each half has half the information, so its standard error is times the whole study’s; the difference between the halves adds two such variances, so its standard error is twice the whole study’s. A difference between the halves that is as large as the overall effect is therefore half as many of its own standard errors as the overall effect is of its own.
In a study sized to give 80% power for the overall effect:
- a difference between the halves as large as the overall effect is detected 28.8% of the time, and needs four times the sample to be detected at 80%;
- a difference half as large is detected 10.8% of the time, and needs sixteen times the sample;
- only a difference twice the overall effect — one half with twice the overall effect and the other with none — reaches 80% at the original sample.
So the proper test is badly underpowered for any realistic difference, which is exactly why the improper reading is so tempting: it produces verdicts at the sample size the study has. How many subjects a study needs is a question about which question it is designed to answer, and a study designed to detect an effect is not designed to detect a difference in that effect and cannot be read as though it were.
The cost argues for a narrower practice rather than a hopeless one. A study that names, in advance, the one difference it expects and sizes itself for that difference can test it properly; a study that looks at every subgroup afterwards is running the family of verdicts tabulated above, with a split almost guaranteed and a difference test at a tenth of the power it would need. The difference between those two studies is not the analysis — the subtraction is the same — but whether the comparison was chosen before the split was seen.
What the formulas establish, and what they do not
At 50% power half of pairs of studies of one effect split, and fewer than one split in six has a significant difference. The split rate is the closed form , and the conditional rate is a one-dimensional integral whose value was also checked against four hundred thousand simulated pairs at 80% power, agreeing to within a tenth of a point.
A significant and a non-significant result do not differ significantly from each other in the examples tabulated — 0.01 against 0.20 gives a difference with p = 0.36.
A difference between two equal halves as large as the main effect needs four times the sample, and at the sample powered for the main effect the test of that difference is weak — 28.8%, and 10.8% for a difference half as large.
The difference recovered from two published intervals is the ordinary two-sample comparison, and the worked example is the same calculation as the p-value table with the standard errors read off interval widths rather than supplied: 2.0 with a standard error of 2.327 and a p-value of 0.390. Two non-significant results with z statistics of 1.5 combine to a significant one, at 0.034, by the sum that is exact for equally precise studies of one effect.
What does not survive is a split read as a difference. Every split in every table above occurred between identical effects.
Not claimed: that subgroup effects are never real, or that a split is uninformative in every design. A split between very large studies with high power, where the non-significant one is also precise, is some evidence of a difference, and the right tool for weighing it is still the test of the difference rather than the pair of verdicts. The apparent-gap table uses the truncated-normal means for a split on the side of the effect, which is where nearly all splits fall at the powers shown; the rare splits with one study significant in the wrong direction are left out of it.
Still open: the reverse reading
Every number here assumes the effects are identical and asks how often a difference is invented. The reverse question is how often a real difference is hidden — two groups with genuinely different effects, both significant or both not, read as consistent. That is a statement about the power of the verdict comparison against a real interaction, set beside the power of the difference test, and it would say whether the split reading is merely misleading or also blind.
There is also a middle case that is common in practice and is not computed here. When the second group is not a separate study but a subset of the first — the effect in everyone, and then the effect in women — the two estimates share data and are positively correlated, which narrows the standard error of their difference and changes both the split rate and the conditional rate. The arithmetic is the same integral with a correlated normal, and the direction of the change is not obvious in advance.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name multiple comparisons, sample size, statistical power
- Estimating how many nulls are true — both name multiple comparisons, p-value, statistical power
- The correction that makes the estimate worse — both name multiple comparisons, statistical power, the winner's curse
- The effect a stopped trial reports — both name sample size, statistical power, the winner's curse
- The price of control — both name multiple comparisons, sample size, statistical power
- The smallest of three combinations — both name multiple comparisons, p-value, statistical power
Named objects
A flat tag is an object no other essay names yet.
InteractionMultiple comparisonsp-valueSample sizeStatistical powerSubgroup analysisTwo-sample testThe winner's curse