The verdicts watch the threshold, not the gap
Worth reading first: Twenty intervals and one expected miss.
Significant in one, not in the other took two studies of exactly the same effect and counted how often their verdicts disagree: half the time, at the moderate power most studies run at, and in nine splits out of ten the difference between the two estimates is nowhere near significant. That is the reading that invents a difference. It left the reverse question open, and the reverse is the one a reviewer of a subgroup table actually needs answered: when two groups do have different effects, how often does the pair of verdicts say so — and when both are significant, or neither is, how often is a real difference sitting underneath the agreement?
The answer has a shape that no single number conveys. The chance that two verdicts disagree does not depend on how far apart the two effects are in any simple way. It depends on where they sit against the threshold of 1.96, and a difference of fixed size can be flagged three times in four or almost never, according to nothing but its height.
Two instruments, drawn on one plane
The hero figure puts both readings in the same picture. Each point of the plane is a pair of true effects, expressed as the z statistic each study would have on average, and the diagonal is where the effects are identical.
The test of the difference asks whether exceeds 1.96. Its power at a pair of true means is a function of alone, so it is constant along every line parallel to the diagonal. Two studies at means 3 and 1 and two at means 8 and 6 have the same power to be told apart, 29.30%, because the question is about their gap and the gap is the same.
The verdict comparison asks whether one study clears 1.96 and the other does not. Its chance at a pair of true means is , where each is that study’s power. That is a function of the two positions separately, and the shading shows the consequence: a cross of high disagreement centred on 1.96 in each direction, and a large pale corner where both effects are big. The two instruments are looking at different things. One measures distance from the diagonal; the other measures distance from a pair of lines that have nothing to do with the diagonal at all.
So the answer to “how often do the verdicts detect a real difference” is not a number but a map, and the map has a feature a reader would not guess: the verdicts are most sensitive to a difference exactly where the two effects straddle the threshold, and nearly blind to it where both are comfortably clear of it. A pair of large, precise studies with very different effects lands in the pale corner. The verdicts agree there, and they agree whatever the difference is.
A halving of the effect, with both studies significant
Take two large trials of one treatment, run in different populations. In the first the treatment’s z statistic has mean 6; in the second the effect is half as large and the mean is 3. The difference between them is real, it is a factor of two, and it is the kind of difference — a treatment that works half as well in one population — that decides whether the treatment is worth giving there.
The verdicts are “significant” and “significant” in 85.08% of such pairs. They split in 14.92%, and they are never both non-significant. A results section written from the verdicts says that the treatment works in both populations, and a reader takes that as consistency. The test of the difference, on the same pairs, finds the halving 56.41% of the time.
Push both studies further up. At true means of 8 and 4 — the same halving, in studies with more data — both are significant in 97.93% of pairs, the verdicts split in 2.07%, and the difference test’s power has risen to 80.74%. Adding data makes the verdict comparison less able to see the difference, because it drives both studies deeper into the region where both verdicts are certain, while it makes the difference test more able to see it, because the same ratio of effects becomes a larger gap in standard errors. The two readings move in opposite directions as the studies get better.
That is the reverse reading at its sharpest. The split invents differences between identical effects; the agreement conceals differences between unequal ones; and the concealment is worst in the studies a reader trusts most. It is also the reading most often applied to a replication: an original and a repeat that are both significant are reported as a successful replication, and a repeat whose effect is half the original’s is, by that account, a success.
One gap, read at every height
The plane can be cut along a line parallel to the diagonal, so that the gap between the two effects is fixed and only its height changes.
With the gap fixed at two z units the difference test has power 29.30% everywhere, which is modest and honest: a difference of two units in studies of this size is not easy to detect, and the test says so at every height. The split does something else entirely.
| midpoint of the two true means | verdicts split | both significant | neither | difference test |
|---|---|---|---|---|
| 0 | 28.23% | 2.89% | 68.88% | 29.30% |
| 1.97 | 73.18% | 13.74% | 13.08% | 29.30% |
| 4 | 15.00% | 84.98% | 0.02% | 29.30% |
| 5 | 2.07% | 97.93% | 0.00% | 29.30% |
| 6 | 0.12% | 99.88% | 0.00% | 29.30% |
The same difference is flagged by a split 73.18% of the time when the pair straddles the threshold and 0.12% of the time when both sit well above it. Near the threshold the split flags the difference more than twice as often as the test does, which looks like sensitivity and is not: it is the same position in the plane at which, in the essay on identical effects, the verdicts split half the time with no difference at all. The split’s high rate near 1.96 is mostly a property of 1.96.
Larger gaps keep the shape. At a gap of three units the split peaks at 86.89% at a midpoint of 2.00 and the difference test sits at 56.41%; at four units the peak is 93.47% at 2.15 against a test at 80.74%. The peak always sits near the threshold, and away from it the curve falls to nothing within three or four units whatever the gap.
Two opposite effects, both called null
The pale region of the plane is not only the corner where both effects are large. It is also the band near zero, where both are small — and small includes opposite in sign.
Two subgroups in which a treatment helps one and harms the other, each by an amount a study of that size has only a modest chance of detecting, produce “not significant” and “not significant” in 45.83% of pairs. The sentence that follows is that the treatment had no effect in either subgroup, which a reader hears as the subgroups agreeing. The test of the difference, whose gap here is three units, finds the reversal 56.41% of the time — as often as it found the halving between the large trials above, because the gap is the same.
The verdicts do occasionally register this pair: they split 43.74% of the time and are both significant, in opposite directions, 10.44% of the time. But a split here is read the way every split is read, as an effect present in one group and absent in the other, and so even the reading that notices something describes it wrongly. What the pair actually contains is an effect of one sign in one group and the other sign in the other, and only the estimates — not the verdicts — carry the sign. The mirror case, two non-significant results with the same sign that together make a significant one, is the arithmetic of combining two p-values; here the same two verdicts sit on estimates that should never be combined at all, and nothing in the verdicts distinguishes the two situations.
What each reading is worth as evidence
A map says where each reading responds. The question a reader is in a position to ask is narrower: having seen a split, or an agreement, how much more or less likely is it that the effects really differ? That is a ratio of two rates — how often the reading occurs when the effects differ, divided by how often it occurs when they do not — and it is the weight of evidence the reading carries, the number that multiplies the odds of a difference.
The natural place to compute it is the design the essay on identical effects ended with: a study sized for 80% power on its overall effect, split into two equal halves. Each half then has 50.84% power, and the halves split 49.99% of the time when their effects are the same. Let the halves’ effects differ by some multiple of the overall effect, keeping their average fixed.
| difference between the halves | verdicts split | split, times as likely | difference test significant | test, times as likely |
|---|---|---|---|---|
| none | 49.99% | 1 | 5.00% | 1 |
| half the effect | 57.18% | 1.144 | 10.78% | 2.156 |
| the whole effect | 72.87% | 1.458 | 28.84% | 5.768 |
| twice the effect | 92.96% | 1.860 | 80.00% | 16.000 |
A split is at most 1.86 times as likely under a real difference as under none, across every difference up to twice the overall effect — which is the largest difference that keeps both halves’ effects on the same side of zero. A reader who sees a split and concludes the halves differ is updating on evidence that, at best, not quite doubles the odds. A significant test of the difference, at the difference as large as the effect itself, multiplies them by 5.77; at twice the effect, by sixteen.
The agreement is weak evidence too, and in the direction a reader expects. When the halves’ verdicts agree, a difference the size of the whole effect is 0.542 times as likely as no difference — the agreement roughly halves the odds. A non-significant test of the difference at the same size is 0.749 times as likely, which is a weaker update: the test at 5% is conservative, and failing to reject tells less than rejecting does. So at this one design the agreement of verdicts is modestly more informative about sameness than a non-significant difference test, and it would be dishonest to leave that out.
It holds at this design and at no other in particular. The agreement’s weight comes entirely from where the halves sit. In the large trials above, a halving leaves both significant 97.93% of the time and equal effects at the same height leave both significant essentially always, so there the agreement is worth nothing at all — a likelihood ratio indistinguishable from one — while the difference test keeps the power that its gap gives it. The difference test’s evidence is a property of the difference; the verdicts’ evidence is a property of the design’s position against 1.96. A reader can compute the first from the published estimates. The second has to be worked out afresh for every table, and nobody does. A table of eight subgroups multiplies the problem rather than averaging it out, for the reason twenty analyses of nothing gives: every extra verdict is another chance for a position-driven split or agreement, and the family of verdicts tells a story that no single comparison in it supports.
The split as a test at a strange level
One more comparison makes the point without any appeal to evidence. The split is, formally, a test of “the effects differ”: it rejects when the verdicts disagree. At equal effects on this design its false-alarm rate is 49.99%, which would be an absurd significance level for any test anyone designed, but it is a level, and the test of the difference can be run at it.
At a two-sided level of 49.99% the difference test rejects when exceeds 0.675. Against a difference as large as the overall effect it then rejects 78.51% of the time, where the split reaches 72.87%; against half the effect, 59.48% against 57.18%; against twice the effect, 98.35% against 92.96%. At the level the split runs at, it is dominated by the test it stands in for. It is not a proper test at an unusual level, which would be defensible if someone wanted a very sensitive screen; it is a worse test than the proper one at that same level, because it throws away the estimates’ distance and keeps only which side of a line each one fell.
That also closes the escape a practised reader might try. Granting that the split has a high false-alarm rate, one might defend it as a sensitive screen for differences, to be followed up properly. As a screen it is beaten by the difference test run with a loose threshold, which costs the same subtraction and has a false-alarm rate that can be chosen rather than inherited from where the effects happen to sit.
Why position rather than gap
The mechanism is a threshold applied twice, and it is the same mechanism that makes a significant estimate an inflated one: a hard cut at 1.96 discards the estimate’s value and keeps only its side. Each verdict is a step function of its own estimate, flat on either side of 1.96, so it carries information about the estimate only when the estimate is near the step. Two estimates far above the step both return “significant” whatever their values; two far below both return “not significant”; the verdicts can disagree only when at least one estimate is close to the step. The pair of verdicts is therefore an instrument whose resolution is concentrated in a narrow band around one arbitrary value, and a difference is visible to it only when the difference happens to fall across that band.
The difference test has no such band. It is linear in both estimates, so every unit of distance between them counts the same wherever it occurs, and its resolution is uniform over the plane. That uniformity is exactly what “a test of the difference” means, and it is what the pair of verdicts, however they are combined, cannot supply: any function of two step functions is a step function of the plane, with its steps where the original ones were.
The same argument is what makes two overlapping intervals — and, one step earlier, a replication landing inside an interval — a poor test, in a milder form. Overlap uses the estimates’ distance from each other, so it is anchored to the diagonal and not to 1.96, and its defect is only that it runs at the wrong level — 0.56% rather than 5%. The verdict comparison is anchored to the wrong place altogether.
What a subgroup table can be made to say
The practical conclusion does not require abandoning subgroup tables. It requires reading them by their estimates.
A table that reports each subgroup’s effect with an interval already contains everything the difference test needs, as the essay on the split showed: each standard error is the interval’s width over 3.92, and the difference between two subgroups has the root sum of squares of the two. The resulting interval for the difference is the answer to the question both readings were trying to ask, in both directions at once — it excludes zero when the gap is large relative to its precision, and it is narrow and centred near zero when the subgroups genuinely agree. An interval for the difference that is narrow and includes zero is evidence of sameness that no pair of verdicts can supply, and it is the only such evidence the table contains.
What the table should not be asked to do is carry the verdicts as findings. A column of stars beside subgroup effects invites exactly the two misreadings measured here: stars that differ read as a difference, stars that match read as consistency, and neither reading tracks the thing it is taken to report. It is the same discipline that reading one interval as a statement about the procedure asks for, applied to two: the estimates and their precision are the evidence, and a verdict is a summary that has already thrown some of it away. What a p-value does not say is the size of an effect; what a pair of p-values does not say is the size of a difference, and what a matching pair does not say is that there is none.
There is also a design lesson, which the difference test’s power makes plain. A difference between two halves as large as the overall effect is found 28.84% of the time at the sample sized for the main effect, and the verdicts, read as a screen, flag it 72.87% of the time only by also flagging half of all identical pairs. Neither reading gets a study out of having been sized for a different question. A study that wants to know whether its effect differs between groups has to be sized for that difference, and at four times the sample for a difference the size of the effect, most studies cannot afford to ask — which is a reason to say so rather than to read the answer off the stars.
What the plane of two verdicts establishes, and what it assumes
The verdicts’ chance of disagreeing is and depends on where the two effects sit against 1.96; the difference test’s power depends only on their gap. Along a fixed gap of two z units the split runs from 73.18% near the threshold to 0.12% at a midpoint of six while the test holds 29.30% throughout.
A halving of the effect in two large studies is concealed by agreement: both significant in 85.08% of pairs at means 6 and 3 and in 97.93% at 8 and 4, where the difference test finds it 56.41% and 80.74% of the time.
On the halves of an 80%-power study, a split is at most 1.86 times as likely under a real difference as under none, across differences up to twice the overall effect, and a significant difference test is 5.77 times as likely at a difference the size of the effect. At the split’s own false-alarm rate of 49.99% the difference test has more power than the split against every difference computed.
An agreement of verdicts is modest evidence of sameness at this one design — a likelihood ratio of 0.542 against a difference the size of the effect, a little stronger than a non-significant difference test’s 0.749 — and that weight vanishes for studies whose effects are both well clear of the threshold.
Every rate is exact under normal sampling with known standard errors: products of normal tail areas for the verdicts, a single normal tail for the difference, and a one-dimensional integral for the joint probability of a split with a significant difference, which reduces at equal effects to the one checked against four hundred thousand counted pairs in the essay on identical effects. The scatters are five hundred seeded pairs and are illustrations; no rate is read off them.
Not claimed: that subgroups with agreeing verdicts differ, or that disagreeing ones do not. The claim is about what the verdicts can tell a reader, which is less than the estimates can at every point of the plane and nothing at all over much of it. Not considered: estimated standard errors, which add a t reference and change nothing in the shape, and subgroups that share data, which is the next question.
Still open: a subgroup that is part of the whole
Every pair here is two independent studies or two disjoint halves. The commonest pair in a paper is neither: the effect in everyone, reported beside the effect in a subgroup of those same people — all patients, then women; the whole trial, then its largest site. The subgroup’s estimate is built from part of the data that built the whole-sample estimate, so the two are positively correlated, with a correlation equal to the square root of the subgroup’s share of the sample.
That correlation changes every number above, and not all in one direction. It narrows the standard error of the difference between the two estimates, which should make the difference test sharper; but the difference between a whole and its part is not the quantity a reader cares about — the comparison that matters is the part against the rest — and the correlation also makes the two verdicts move together, which should make them split less. Which effect wins, how often “significant overall but not in women” appears when the women’s effect is identical to the men’s, and what the right test is when the reader has only the overall and the subgroup estimate in front of them, are the arithmetic of a correlated normal pair, and nothing here has done it.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An outcome cut in two — both name sample size, statistical power, two-sample test
- A boundary for giving up — both name sample size, statistical power
- A coverage table with its own error — both name sample size, statistical power
- A degrees of freedom that is not a count — both name sample size, two-sample test
- A distribution drawn from the null — both name p-value, statistical power
- Choosing n after looking — both name sample size, statistical power
Named objects
A flat tag is an object no other essay names yet.
InteractionLikelihood ratiop-valueSample sizeSignificance thresholdStatistical powerSubgroup analysisTwo-sample test