An interval read beside something else

The verdicts watch the threshold, not the gap

Two studies whose effects really differ are read through their verdicts as often as through their difference. Halves of an 80%-power study whose effects differ by the whole effect split 72.87% of the time against 49.99% when they do not — a split is never more than 1.86 times as likely under any difference up to twice the effect. And a halving of the effect in two large studies leaves both significant 97.93% of the time, while the test of the difference finds it 80.74%.

Worth reading first: Twenty intervals and one expected miss.

Significant in one, not in the other took two studies of exactly the same effect and counted how often their verdicts disagree: half the time, at the moderate power most studies run at, and in nine splits out of ten the difference between the two estimates is nowhere near significant. That is the reading that invents a difference. It left the reverse question open, and the reverse is the one a reviewer of a subgroup table actually needs answered: when two groups do have different effects, how often does the pair of verdicts say so — and when both are significant, or neither is, how often is a real difference sitting underneath the agreement?

The answer has a shape that no single number conveys. The chance that two verdicts disagree does not depend on how far apart the two effects are in any simple way. It depends on where they sit against the threshold of 1.96, and a difference of fixed size can be flagged three times in four or almost never, according to nothing but its height.

Where two verdicts disagree, across the true effects of two studies, beside where the test of their difference finds it. The shading is the chance that exactly one of two independent studies is significant, over the plane of their true z means. It is highest on a cross centred where either mean is 1.96 and lowest in the corner where both are large, whatever the difference between them. The difference test's power is constant along lines parallel to the diagonal: 50% on the inner pair and 80% on the outer. Two studies at means 8 and 4 — one effect half the other — split 2.1% of the time, and the difference test finds the halving 80.7% of the time.
Fig. 1 The plane of two studies’ true z means. The shading is the chance that exactly one of the two is significant; the dashed lines are 1.96 on each axis. The difference test’s power is constant along the diagonal lines — 50% on the dotted pair, 80% on the solid pair — and depends only on the distance from the diagonal.

Two instruments, drawn on one plane

The hero figure puts both readings in the same picture. Each point of the plane is a pair of true effects, expressed as the z statistic each study would have on average, and the diagonal is where the effects are identical.

The test of the difference asks whether (z1−z2)/2(z_1 - z_2)/\sqrt 2 exceeds 1.96. Its power at a pair of true means is a function of d1−d2d_1 - d_2 alone, so it is constant along every line parallel to the diagonal. Two studies at means 3 and 1 and two at means 8 and 6 have the same power to be told apart, 29.30%, because the question is about their gap and the gap is the same.

The verdict comparison asks whether one study clears 1.96 and the other does not. Its chance at a pair of true means is π1(1−π2)+π2(1−π1)\pi_1(1 - \pi_2) + \pi_2(1 - \pi_1), where each π\pi is that study’s power. That is a function of the two positions separately, and the shading shows the consequence: a cross of high disagreement centred on 1.96 in each direction, and a large pale corner where both effects are big. The two instruments are looking at different things. One measures distance from the diagonal; the other measures distance from a pair of lines that have nothing to do with the diagonal at all.

So the answer to “how often do the verdicts detect a real difference” is not a number but a map, and the map has a feature a reader would not guess: the verdicts are most sensitive to a difference exactly where the two effects straddle the threshold, and nearly blind to it where both are comfortably clear of it. A pair of large, precise studies with very different effects lands in the pale corner. The verdicts agree there, and they agree whatever the difference is.

A halving of the effect, with both studies significant

Take two large trials of one treatment, run in different populations. In the first the treatment’s z statistic has mean 6; in the second the effect is half as large and the mean is 3. The difference between them is real, it is a factor of two, and it is the kind of difference — a treatment that works half as well in one population — that decides whether the treatment is worth giving there.

Two studies whose effects differ, z statistics with means 6 and 3: where the verdicts split and where the difference shows. Five hundred pairs whose true effects differ. The test of the difference finds it in 56.4% of pairs; the two verdicts disagree in 14.9%, both are significant in 85.1% and neither in 0.0%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.
Fig. 2 Five hundred pairs of z statistics from two studies whose true means are 6 and 3. Blue points are pairs whose difference is significant; red points are pairs in which one study is significant and the other is not; grey points are neither. Most of the cloud sits above and to the right of both dashed lines.

The verdicts are “significant” and “significant” in 85.08% of such pairs. They split in 14.92%, and they are never both non-significant. A results section written from the verdicts says that the treatment works in both populations, and a reader takes that as consistency. The test of the difference, on the same pairs, finds the halving 56.41% of the time.

Push both studies further up. At true means of 8 and 4 — the same halving, in studies with more data — both are significant in 97.93% of pairs, the verdicts split in 2.07%, and the difference test’s power has risen to 80.74%. Adding data makes the verdict comparison less able to see the difference, because it drives both studies deeper into the region where both verdicts are certain, while it makes the difference test more able to see it, because the same ratio of effects becomes a larger gap in standard errors. The two readings move in opposite directions as the studies get better.

That is the reverse reading at its sharpest. The split invents differences between identical effects; the agreement conceals differences between unequal ones; and the concealment is worst in the studies a reader trusts most. It is also the reading most often applied to a replication: an original and a repeat that are both significant are reported as a successful replication, and a repeat whose effect is half the original’s is, by that account, a success.

One gap, read at every height

The plane can be cut along a line parallel to the diagonal, so that the gap between the two effects is fixed and only its height changes.

The same real difference, 2 z units, read by the verdicts at every height and by the test of the differenceThe two studies' true z means differ by 2 throughout and only their midpoint moves. The verdicts disagree 73.2% of the time at a midpoint of 2 and 0.00% of the time at 7; beyond a midpoint of about 4 both are nearly always significant. The test of the difference finds it 29.3% of the time at every midpoint.00.2500.5000.75010246midpoint of the two true z meansprobabilityone significant, one notboth significantneither significantdifference test significantindependent studies, exactone gap, read at every height
Fig. 3 Two studies whose true z means differ by 2, with their midpoint moved from 0 to 7. The solid curves are the chance the verdicts split, the chance both are significant and the chance neither is; the dashed line is the difference test’s power, which does not move. The vertical line marks a midpoint of 1.96. The slider sets the size of the difference.

With the gap fixed at two z units the difference test has power 29.30% everywhere, which is modest and honest: a difference of two units in studies of this size is not easy to detect, and the test says so at every height. The split does something else entirely.

midpoint of the two true means verdicts split both significant neither difference test
0 28.23% 2.89% 68.88% 29.30%
1.97 73.18% 13.74% 13.08% 29.30%
4 15.00% 84.98% 0.02% 29.30%
5 2.07% 97.93% 0.00% 29.30%
6 0.12% 99.88% 0.00% 29.30%

The same difference is flagged by a split 73.18% of the time when the pair straddles the threshold and 0.12% of the time when both sit well above it. Near the threshold the split flags the difference more than twice as often as the test does, which looks like sensitivity and is not: it is the same position in the plane at which, in the essay on identical effects, the verdicts split half the time with no difference at all. The split’s high rate near 1.96 is mostly a property of 1.96.

Larger gaps keep the shape. At a gap of three units the split peaks at 86.89% at a midpoint of 2.00 and the difference test sits at 56.41%; at four units the peak is 93.47% at 2.15 against a test at 80.74%. The peak always sits near the threshold, and away from it the curve falls to nothing within three or four units whatever the gap.

Two opposite effects, both called null

The pale region of the plane is not only the corner where both effects are large. It is also the band near zero, where both are small — and small includes opposite in sign.

Two studies whose effects differ, z statistics with means 1.5 and -1.5: where the verdicts split and where the difference shows. Five hundred pairs whose true effects differ. The test of the difference finds it in 56.4% of pairs; the two verdicts disagree in 43.7%, both are significant in 10.4% and neither in 45.8%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.
Fig. 4 Five hundred pairs from two studies whose true z means are 1.5 and −1.5: the treatment helps in one group and harms in the other by the same amount. Blue points have a significant difference; red points split; grey points are neither.

Two subgroups in which a treatment helps one and harms the other, each by an amount a study of that size has only a modest chance of detecting, produce “not significant” and “not significant” in 45.83% of pairs. The sentence that follows is that the treatment had no effect in either subgroup, which a reader hears as the subgroups agreeing. The test of the difference, whose gap here is three units, finds the reversal 56.41% of the time — as often as it found the halving between the large trials above, because the gap is the same.

The verdicts do occasionally register this pair: they split 43.74% of the time and are both significant, in opposite directions, 10.44% of the time. But a split here is read the way every split is read, as an effect present in one group and absent in the other, and so even the reading that notices something describes it wrongly. What the pair actually contains is an effect of one sign in one group and the other sign in the other, and only the estimates — not the verdicts — carry the sign. The mirror case, two non-significant results with the same sign that together make a significant one, is the arithmetic of combining two p-values; here the same two verdicts sit on estimates that should never be combined at all, and nothing in the verdicts distinguishes the two situations.

What each reading is worth as evidence

A map says where each reading responds. The question a reader is in a position to ask is narrower: having seen a split, or an agreement, how much more or less likely is it that the effects really differ? That is a ratio of two rates — how often the reading occurs when the effects differ, divided by how often it occurs when they do not — and it is the weight of evidence the reading carries, the number that multiplies the odds of a difference.

The natural place to compute it is the design the essay on identical effects ended with: a study sized for 80% power on its overall effect, split into two equal halves. Each half then has 50.84% power, and the halves split 49.99% of the time when their effects are the same. Let the halves’ effects differ by some multiple of the overall effect, keeping their average fixed.

How much more often each reading occurs when the halves of an 80%-power study really differ. A study sized for 80% power overall, split into two halves whose effects differ by a multiple of the overall effect. A split between the halves' verdicts is 1.46 times as likely when they differ by the whole effect as when they do not, and never more than 1.86 times as likely across the range. A significant test of the difference is 5.77 times as likely at the whole effect and 16.0 at twice it.
Fig. 5 For the two halves of a study sized for 80% power overall, how many times more often each reading occurs when the halves’ effects differ than when they are equal, on a doubling scale, against the size of the difference as a multiple of the overall effect.
difference between the halves verdicts split split, times as likely difference test significant test, times as likely
none 49.99% 1 5.00% 1
half the effect 57.18% 1.144 10.78% 2.156
the whole effect 72.87% 1.458 28.84% 5.768
twice the effect 92.96% 1.860 80.00% 16.000

A split is at most 1.86 times as likely under a real difference as under none, across every difference up to twice the overall effect — which is the largest difference that keeps both halves’ effects on the same side of zero. A reader who sees a split and concludes the halves differ is updating on evidence that, at best, not quite doubles the odds. A significant test of the difference, at the difference as large as the effect itself, multiplies them by 5.77; at twice the effect, by sixteen.

The agreement is weak evidence too, and in the direction a reader expects. When the halves’ verdicts agree, a difference the size of the whole effect is 0.542 times as likely as no difference — the agreement roughly halves the odds. A non-significant test of the difference at the same size is 0.749 times as likely, which is a weaker update: the test at 5% is conservative, and failing to reject tells less than rejecting does. So at this one design the agreement of verdicts is modestly more informative about sameness than a non-significant difference test, and it would be dishonest to leave that out.

It holds at this design and at no other in particular. The agreement’s weight comes entirely from where the halves sit. In the large trials above, a halving leaves both significant 97.93% of the time and equal effects at the same height leave both significant essentially always, so there the agreement is worth nothing at all — a likelihood ratio indistinguishable from one — while the difference test keeps the power that its gap gives it. The difference test’s evidence is a property of the difference; the verdicts’ evidence is a property of the design’s position against 1.96. A reader can compute the first from the published estimates. The second has to be worked out afresh for every table, and nobody does. A table of eight subgroups multiplies the problem rather than averaging it out, for the reason twenty analyses of nothing gives: every extra verdict is another chance for a position-driven split or agreement, and the family of verdicts tells a story that no single comparison in it supports.

The split as a test at a strange level

One more comparison makes the point without any appeal to evidence. The split is, formally, a test of “the effects differ”: it rejects when the verdicts disagree. At equal effects on this design its false-alarm rate is 49.99%, which would be an absurd significance level for any test anyone designed, but it is a level, and the test of the difference can be run at it.

At a two-sided level of 49.99% the difference test rejects when ∣z1−z2∣/2|z_1 - z_2|/\sqrt 2 exceeds 0.675. Against a difference as large as the overall effect it then rejects 78.51% of the time, where the split reaches 72.87%; against half the effect, 59.48% against 57.18%; against twice the effect, 98.35% against 92.96%. At the level the split runs at, it is dominated by the test it stands in for. It is not a proper test at an unusual level, which would be defensible if someone wanted a very sensitive screen; it is a worse test than the proper one at that same level, because it throws away the estimates’ distance and keeps only which side of a line each one fell.

That also closes the escape a practised reader might try. Granting that the split has a high false-alarm rate, one might defend it as a sensitive screen for differences, to be followed up properly. As a screen it is beaten by the difference test run with a loose threshold, which costs the same subtraction and has a false-alarm rate that can be chosen rather than inherited from where the effects happen to sit.

Why position rather than gap

The mechanism is a threshold applied twice, and it is the same mechanism that makes a significant estimate an inflated one: a hard cut at 1.96 discards the estimate’s value and keeps only its side. Each verdict is a step function of its own estimate, flat on either side of 1.96, so it carries information about the estimate only when the estimate is near the step. Two estimates far above the step both return “significant” whatever their values; two far below both return “not significant”; the verdicts can disagree only when at least one estimate is close to the step. The pair of verdicts is therefore an instrument whose resolution is concentrated in a narrow band around one arbitrary value, and a difference is visible to it only when the difference happens to fall across that band.

The difference test has no such band. It is linear in both estimates, so every unit of distance between them counts the same wherever it occurs, and its resolution is uniform over the plane. That uniformity is exactly what “a test of the difference” means, and it is what the pair of verdicts, however they are combined, cannot supply: any function of two step functions is a step function of the plane, with its steps where the original ones were.

The same argument is what makes two overlapping intervals — and, one step earlier, a replication landing inside an interval — a poor test, in a milder form. Overlap uses the estimates’ distance from each other, so it is anchored to the diagonal and not to 1.96, and its defect is only that it runs at the wrong level — 0.56% rather than 5%. The verdict comparison is anchored to the wrong place altogether.

What a subgroup table can be made to say

The practical conclusion does not require abandoning subgroup tables. It requires reading them by their estimates.

A table that reports each subgroup’s effect with an interval already contains everything the difference test needs, as the essay on the split showed: each standard error is the interval’s width over 3.92, and the difference between two subgroups has the root sum of squares of the two. The resulting interval for the difference is the answer to the question both readings were trying to ask, in both directions at once — it excludes zero when the gap is large relative to its precision, and it is narrow and centred near zero when the subgroups genuinely agree. An interval for the difference that is narrow and includes zero is evidence of sameness that no pair of verdicts can supply, and it is the only such evidence the table contains.

What the table should not be asked to do is carry the verdicts as findings. A column of stars beside subgroup effects invites exactly the two misreadings measured here: stars that differ read as a difference, stars that match read as consistency, and neither reading tracks the thing it is taken to report. It is the same discipline that reading one interval as a statement about the procedure asks for, applied to two: the estimates and their precision are the evidence, and a verdict is a summary that has already thrown some of it away. What a p-value does not say is the size of an effect; what a pair of p-values does not say is the size of a difference, and what a matching pair does not say is that there is none.

There is also a design lesson, which the difference test’s power makes plain. A difference between two halves as large as the overall effect is found 28.84% of the time at the sample sized for the main effect, and the verdicts, read as a screen, flag it 72.87% of the time only by also flagging half of all identical pairs. Neither reading gets a study out of having been sized for a different question. A study that wants to know whether its effect differs between groups has to be sized for that difference, and at four times the sample for a difference the size of the effect, most studies cannot afford to ask — which is a reason to say so rather than to read the answer off the stars.

What the plane of two verdicts establishes, and what it assumes

The verdicts’ chance of disagreeing is π1(1−π2)+π2(1−π1)\pi_1(1 - \pi_2) + \pi_2(1 - \pi_1) and depends on where the two effects sit against 1.96; the difference test’s power depends only on their gap. Along a fixed gap of two z units the split runs from 73.18% near the threshold to 0.12% at a midpoint of six while the test holds 29.30% throughout.

A halving of the effect in two large studies is concealed by agreement: both significant in 85.08% of pairs at means 6 and 3 and in 97.93% at 8 and 4, where the difference test finds it 56.41% and 80.74% of the time.

On the halves of an 80%-power study, a split is at most 1.86 times as likely under a real difference as under none, across differences up to twice the overall effect, and a significant difference test is 5.77 times as likely at a difference the size of the effect. At the split’s own false-alarm rate of 49.99% the difference test has more power than the split against every difference computed.

An agreement of verdicts is modest evidence of sameness at this one design — a likelihood ratio of 0.542 against a difference the size of the effect, a little stronger than a non-significant difference test’s 0.749 — and that weight vanishes for studies whose effects are both well clear of the threshold.

Every rate is exact under normal sampling with known standard errors: products of normal tail areas for the verdicts, a single normal tail for the difference, and a one-dimensional integral for the joint probability of a split with a significant difference, which reduces at equal effects to the one checked against four hundred thousand counted pairs in the essay on identical effects. The scatters are five hundred seeded pairs and are illustrations; no rate is read off them.

Not claimed: that subgroups with agreeing verdicts differ, or that disagreeing ones do not. The claim is about what the verdicts can tell a reader, which is less than the estimates can at every point of the plane and nothing at all over much of it. Not considered: estimated standard errors, which add a t reference and change nothing in the shape, and subgroups that share data, which is the next question.

Still open: a subgroup that is part of the whole

Every pair here is two independent studies or two disjoint halves. The commonest pair in a paper is neither: the effect in everyone, reported beside the effect in a subgroup of those same people — all patients, then women; the whole trial, then its largest site. The subgroup’s estimate is built from part of the data that built the whole-sample estimate, so the two are positively correlated, with a correlation equal to the square root of the subgroup’s share of the sample.

That correlation changes every number above, and not all in one direction. It narrows the standard error of the difference between the two estimates, which should make the difference test sharper; but the difference between a whole and its part is not the quantity a reader cares about — the comparison that matters is the part against the rest — and the correlation also makes the two verdicts move together, which should make them split less. Which effect wins, how often “significant overall but not in women” appears when the women’s effect is identical to the men’s, and what the right test is when the reader has only the overall and the subgroup estimate in front of them, are the arithmetic of a correlated normal pair, and nothing here has done it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

InteractionLikelihood ratiop-valueSample sizeSignificance thresholdStatistical powerSubgroup analysisTwo-sample test