An interval read beside something else

A subgroup inside its own trial

A trial with 80% power, significant overall and not in the fifth of it that is women, has told the reader almost nothing about women: that pattern turns up 57.45% of the time when women have the full effect and 58.26% when they have none. The subgroup is part of the whole, so the two verdicts are correlated, and the one test that answers the question — the subgroup against everyone else — can be recovered from the two printed intervals and nothing more.

Worth reading first: Twenty intervals and one expected miss.

The subgroup comparisons measured so far were between two groups that shared nothing: two studies, or two disjoint halves of one. Significant in one, not in the other found that such a pair splits about significance half the time with identical effects, and the verdicts watch the threshold, not the gap found that the split responds to where the effects sit rather than how far apart they are.

The pair that actually appears in results sections is usually different. It is the effect in everyone, reported beside the effect in a subgroup of those same people — all patients and then women, the whole trial and then its largest site, every participant and then those over seventy. The subgroup’s estimate is built from part of the data that built the overall one. The two are not independent, and the arithmetic that treats them as though they were is wrong in a direction that can be named.

A whole trial's z against a subgroup's, the subgroup 20% of it with the same effect as the restFive hundred pairs, correlated at 0.447 — the square root of its share — because the subgroup is part of the whole. The whole trial is significant and the subgroup is not in 57.5% of trials; the subgroup differs significantly from the rest in 5.0%. The dashed lines mark 1.96 on each axis and the slanted lines mark a significant difference between the subgroup and the rest.-20246-20246whole trial's zsubgroup'szsignificant overall,not in the subgroup 57.5% of trialssubgroup differs fromthe rest, significantly 5.0% of trialssubgroup's effect: sameshare of the trial: 20%whole trial at 80% power, 500 trials, seededthe part moves with the whole
Fig. 1 Five hundred trials’ z statistics: the whole trial’s on the horizontal axis and a subgroup’s, a fifth of the participants, on the vertical, with the effect the same in everyone. Red points are significant overall and not in the subgroup; blue points are those in which the subgroup differs significantly from the rest of the trial, which is the slanted pair of lines. The slider sets the subgroup’s share.

The correlation that sharing forces

Let a trial have nn participants and a subgroup a share ff of them. With a common spread σ\sigma per participant, the overall estimate has variance σ2/n\sigma^2/n and the subgroup’s σ2/(fn)\sigma^2/(fn). The overall estimate is a weighted average of the subgroup’s and the rest’s,

θ^whole=f θ^sub+(1−f) θ^rest,\hat\theta_{\text{whole}} = f\,\hat\theta_{\text{sub}} + (1 - f)\,\hat\theta_{\text{rest}},

and since the subgroup and the rest share no participants, the covariance of the whole with the subgroup is f⋅σ2/(fn)=σ2/nf \cdot \sigma^2/(fn) = \sigma^2/n — exactly the variance of the whole. Two exact facts follow.

The z statistics are correlated at f\sqrt f. A subgroup of a fifth of the trial moves with the whole at a correlation of 0.447; a half, at 0.707. The hero figure’s cloud is tilted accordingly: a trial whose overall z happened to come out high tends to have a high subgroup z too, because the subgroup’s participants are part of why the overall z came out high.

The variance of the difference between them is a subtraction, not a sum. Since the covariance equals the whole’s variance,

Var⁡(θ^sub−θ^whole)=Var⁡(θ^sub)−Var⁡(θ^whole),\operatorname{Var}(\hat\theta_{\text{sub}} - \hat\theta_{\text{whole}}) = \operatorname{Var}(\hat\theta_{\text{sub}}) - \operatorname{Var}(\hat\theta_{\text{whole}}),

where two independent estimates would add. And the difference itself is a fixed multiple of something familiar: θ^sub−θ^whole=(1−f)(θ^sub−θ^rest)\hat\theta_{\text{sub}} - \hat\theta_{\text{whole}} = (1 - f)(\hat\theta_{\text{sub}} - \hat\theta_{\text{rest}}). Divided by its correct standard error it is the same z statistic as the test of the subgroup against everyone else, not a similar one. Comparing a part with the whole, done correctly, is comparing the part with the rest.

Significant overall, not in the subgroup

Now read the pair the way results sections do. The trial is sized for 80% power on its overall effect, the effect is the same in everyone, and the question is how often the sentence “the effect was significant overall but not in women” can be written.

A subgroup of a fifth of the participants has a z statistic with mean 2.80×0.2=1.252.80 \times \sqrt{0.2} = 1.25, which is 24.04% power. So the subgroup is non-significant most of the time while the whole is significant most of the time, and the pair “significant overall, not in the subgroup” appears in 57.45% of trials. With the subgroup at half the trial its power is 50.84% and the pattern appears in 31.29%.

A whole trial's z against a subgroup's, the subgroup 50% of it with the same effect as the rest. Five hundred pairs, correlated at 0.707 — the square root of its share — because the subgroup is part of the whole. The whole trial is significant and the subgroup is not in 31.3% of trials; the subgroup differs significantly from the rest in 5.0%. The dashed lines mark 1.96 on each axis and the slanted lines mark a significant difference between the subgroup and the rest.
Fig. 2 The same five hundred trials with the subgroup at half the participants. The cloud is tilted more steeply, since the correlation is now 0.707, and fewer trials fall in the region that is significant overall and not in the subgroup.

The correlation does what the essay on the verdicts expected it to do: it makes the two verdicts move together. With the subgroup at half the trial, the verdicts disagree — one significant, the other not, in either order — in 33.39% of trials. Two disjoint halves of the same trial, each at the same 50.84% power, disagree in 49.99%. Sharing data removes a third of the splits, because a trial that did well overall has usually done well in its half too.

What it does not do is make the splits that remain meaningful. Among trials significant overall and not in a half-share subgroup, the subgroup differs significantly from the rest in 6.28% — and every one of those is a false alarm, since the effect is the same everywhere. At a fifth of the trial the rate is 3.46%. The reverse split, significant in the subgroup and not overall, is rare, 1.51% at a fifth and 2.10% at a half, since a subgroup can clear the threshold on its own only by an excursion the whole rarely fails to share.

The same pattern whatever the subgroup’s truth

The question a reader of “significant overall, not in women” is asking is whether women respond differently. So the pattern has to be priced against the alternatives: women with half the effect, women with none.

A whole trial's z against a subgroup's, the subgroup 20% of it with no effect at all. Five hundred pairs, correlated at 0.447 — the square root of its share — because the subgroup is part of the whole. The whole trial is significant and the subgroup is not in 58.3% of trials; the subgroup differs significantly from the rest in 20.2%. The dashed lines mark 1.96 on each axis and the slanted lines mark a significant difference between the subgroup and the rest.
Fig. 3 Five hundred trials in which the subgroup, a fifth of the participants, has no effect at all and the rest has the full effect. The cloud has moved down and left, and the region that is significant overall and not in the subgroup holds almost exactly the same share of it.

When the subgroup has no effect at all — the reading taken literally — the overall effect is diluted, the whole trial is significant less often, and the pattern appears in 58.26% of trials against 57.45% when the effect is the same. Its likelihood ratio for “women get nothing” against “women get the full effect” is 1.014. The sentence is written almost exactly as often in the world where it is true as in the world where it is false.

How often a trial is significant overall and not in a subgroup, by the subgroup's share, whatever the subgroup's true effect. The whole trial has 80% power when the effect is the same everywhere. At a subgroup of a fifth of the trial, "significant overall, not in the subgroup" happens 57.5% of the time when the subgroup has the same effect, 62.6% when it has half and 58.3% when it has none. The curves separate only at shares above a third, and there the reading becomes less likely when the subgroup has no effect.
Fig. 4 How often a trial is significant overall and not in a subgroup, against the subgroup’s share of the participants, for three truths about the subgroup: the same effect as the rest, half of it, and none. The trial has 80% power when the effect is uniform.
subgroup’s share same effect half the effect no effect none against same
0.2 57.45% 62.68% 58.26% 1.014
0.3 48.18% 55.35% 47.55% 0.987
0.5 31.29% 39.97% 26.48% 0.846
0.7 17.33% 24.27% 10.89% 0.628

Two things in the table are the opposite of what the reading assumes. Below about a third of the trial the pattern is indifferent to the subgroup’s truth: it is manufactured by the subgroup’s low power, and the subgroup’s power is set by its size, not by its effect. Above a third the pattern becomes less likely when the subgroup has no effect, not more, because a large subgroup with no effect drags the whole trial below significance and so removes the first half of the sentence. At a subgroup of half the trial, a reader who sees “significant overall, not in the subgroup” should, if anything, think slightly better of the subgroup’s effect than before — the reverse of the sentence.

The intermediate truth, half the effect, produces the pattern most often at every share. That is a combination of a subgroup weak enough to miss significance and a whole strong enough to reach it, and it is the one case the reading gets roughly right. Even there its likelihood ratio against the same effect is 1.091 at a fifth of the trial. A reader who multiplies their odds by that has barely moved them.

The test the two estimates can support

The sentence stands in for a test of whether the subgroup’s effect differs from the rest’s. A paper usually prints enough to run it, and the commonest way of running it is wrong.

The tempting calculation takes the overall estimate and the subgroup’s, each with its standard error, and tests their difference as if they were two independent studies — the calculation the essay on the split recommended for disjoint groups. For a part and its whole it divides by sesub2+sewhole2\sqrt{\text{se}_\text{sub}^2 + \text{se}_\text{whole}^2} where the correct denominator is sesub2−sewhole2\sqrt{\text{se}_\text{sub}^2 - \text{se}_\text{whole}^2}. The ratio of the two is (1−f)/(1+f)\sqrt{(1 - f)/(1 + f)}: 0.816 at a fifth of the trial and 0.577 at a half. Every statistic is shrunk by that factor, so the test is run at a much stricter level than it claims.

Detecting a subgroup with no effect from the whole trial's estimate and the subgroup's, computed two ways. Power against a subgroup that has no effect, when the rest has the effect that gives the whole trial 80% power if it were uniform. The right comparison is the subgroup against the rest. Treating the whole-trial and subgroup estimates as independent divides by too large a standard error: its size falls far below 5% and its power below the right test's at every share.
Fig. 5 Power to detect a subgroup with no effect, by the subgroup’s share, for the correct comparison of the subgroup with the rest and for the whole-and-subgroup comparison treated as independent. The dashed curve beneath is the second test’s actual false-alarm rate at a nominal 5%.

At a fifth of the trial the independent calculation has an actual size of 1.64% rather than 5%, and at a half 0.069%. Its power falls with it. Against a subgroup with no effect at all, the correct comparison detects the difference 20.17% of the time at a fifth of the trial and 28.84% at a half; the independent calculation detects it 10.05% and 2.31% of the time. At a half, the test a reader is most likely to run on a printed table has thrown away more than nine tenths of the power the data contain, which were not much to begin with.

So the shared data cut both ways, and the net is not in doubt. The correlation lowers the rate of spurious splits, which helps; it also invalidates the one test a reader can run from the two numbers in front of them, unless the reader knows to subtract.

The forest plot’s own reading

A subgroup table is usually drawn as a forest plot: each subgroup’s estimate and interval on its own row, the overall estimate as a vertical line down the page. The reading that goes with it is the reverse of the sentence priced above — if every subgroup’s interval crosses the overall line, the effect is “consistent across subgroups”, and the trial moves on.

For a subgroup that is part of the whole, that reading is nearly automatic. The gap between the subgroup’s estimate and the overall one has standard deviation sesub1−f\text{se}_\text{sub}\sqrt{1 - f}, smaller than the subgroup’s own standard error, because the overall estimate already contains the subgroup. So the subgroup’s 95% interval contains the overall estimate with probability 2Φ(1.96/1−f)−12\Phi(1.96/\sqrt{1 - f}) - 1 when the effect is the same everywhere: 97.16% for a fifth-share subgroup and 99.44% for a half.

Read as a test of “this subgroup differs”, crossing the line is a test run at 2.84% for a fifth and 0.557% for a half, and it is weaker than the subgroup-against-the-rest test for the same reason the independent calculation was: it measures the gap against a standard error that is too large. Against a subgroup with no effect at all it fails to cross the line 14.26% of the time at a fifth of the trial and 8.52% at a half, where the right comparison detects the difference 20.17% and 28.84% of the time. The forest plot’s reading, like the verdicts’ in the essay on the threshold, is an instrument set to find sameness, and at a half-share subgroup it finds it more than ninety-one times in a hundred when the subgroup gets nothing.

Both readings are therefore wrong in the direction that suits the page they appear on. The verdicts’ sentence invents a difference that the numbers do not support; the forest plot’s lines conceal a difference that the numbers could partly support. Neither is the comparison of two intervals read against each other, which at least uses the right pair; both compare a subgroup with a whole that contains it.

The subgroup that is significant when the trial is not

The reverse split — a subgroup significant while the whole trial is not — is rare, 1.51% of trials at a fifth-share subgroup, and it is the one most likely to be written up as a discovery: the drug failed overall but worked in the subgroup.

Its rarity is the warning. A subgroup reaches significance on its own, with a quarter of the trial’s power, only by an excursion well above its true effect, and a trial that did not share the excursion is one in which the rest of the data pulled the other way. So the subgroup’s estimate in that pattern is selected twice over — once by crossing 1.96 with a small sample, which is the winner’s curse at a quarter of the trial’s power, and once by the overall result’s failure, which says the rest did not agree. A subgroup found this way after the fact, among several, is the textbook case of a p-value that does not carry its own selection, and the effect estimate it reports is the least trustworthy number in the paper.

Recovering the comparison from a table

The correct test needs nothing that a table of whole-trial and subgroup intervals does not print, including — which is less obvious — the subgroup’s share of the trial. With a common per-participant spread, the share is the ratio of the two variances, f=sewhole2/sesub2f = \text{se}_\text{whole}^2 / \text{se}_\text{sub}^2.

Suppose a trial reports an overall effect of 4.0, 95% interval 1.6 to 6.4 (p=0.0011p = 0.0011), and an effect in women of 2.0, interval −1.8 to 5.8 (p=0.30p = 0.30), and says that the benefit was not seen in women. The standard errors are the widths over 3.92: 1.2245 and 1.9388. Then:

  • the women’s share is 1.22452/1.93882=1.2245^2/1.9388^2 = 0.399, about two fifths of the trial;
  • the effect in everyone else is (4.0−0.399×2.0)/(1−0.399)=(4.0 - 0.399 \times 2.0)/(1 - 0.399) = 5.33, with a standard error of 1.2245/0.601=1.5791.2245/\sqrt{0.601} = 1.579 and an interval from 2.23 to 8.42;
  • women against the rest differ by −3.33, standard error 2.501, interval −8.23 to 1.57, p = 0.183.

The same z statistic, −1.331, comes from dividing the women’s estimate minus the overall one by 1.93882−1.22452\sqrt{1.9388^2 - 1.2245^2}, which is the identity above written in the table’s own numbers. The calculation that treats the two as independent gives p=0.383p = 0.383 instead — more than twice as large, and a reader running it would conclude that the data are even less informative about women than they are.

The honest summary is the interval for the difference, and it is wide: the data are compatible with women benefiting 1.57 units more than men and with their benefiting 8.23 units less. That is a statement about how little a two-fifths subgroup of an 80%-power trial can say about a difference, and it is the statement the sentence “the benefit was not seen in women” replaces with something that sounds like a finding.

Why sharing does not rescue the reading

The effect of the correlation on the split was the question left open when two disjoint groups were compared, and it has an answer: sharing data makes the verdicts agree more often, a third fewer splits at half the trial. It would be natural to conclude that the nested pair is therefore a more trustworthy reading than the disjoint pair. It is not, for a reason the geometry of the hero figure shows.

The correlation tilts the cloud towards the line along which the whole and the part move together, and the test that matters — subgroup against rest — is measured across that line, in the direction the correlation does nothing to reduce. Fewer trials cross the verdict lines, but the ones that do cross them for the same reason as before: the subgroup is small and its estimate is noisy. Its power is set by its share alone — 24.04% at a fifth, whatever the rest is doing — and a verdict based on a quarter-powered estimate is mostly a statement about the quarter.

The same point in terms of evidence: the likelihood ratios in the table are close to one at every share below a third, because the thing that decides whether the subgroup’s verdict is “not significant” is how many people are in it. Sharing data with the whole makes the whole’s verdict partly predictable from the subgroup’s, which is why the rates change; it does not give the subgroup’s verdict any information about the subgroup’s effect that its sample size did not already fix.

What sharing data changes, exactly, and what it leaves alone

The whole trial and a subgroup of share ff have z statistics correlated at exactly f\sqrt f, and the difference between their estimates is (1−f)(1 - f) times the subgroup-against-the-rest difference, with a variance that is the subgroup’s minus the whole’s.

At 80% power overall, “significant overall, not in the subgroup” appears in 57.45% of trials with a fifth-share subgroup and the same effect everywhere, and in 58.26% when the subgroup has no effect, a likelihood ratio of 1.014. With a half-share subgroup it is 31.29% and 26.48%, so the pattern is less likely when the subgroup gets nothing.

Treating the whole and the subgroup as independent runs the test at 1.64% for a fifth-share subgroup and 0.069% for a half, with power 10.05% and 2.31% against a subgroup with no effect, where the subgroup-against-the-rest test has 20.17% and 28.84%.

The subgroup’s share and the rest’s estimate can both be recovered from two printed intervals, under a common spread, and the worked example gives the comparison a paper summarised as “not seen in women” a p-value of 0.183.

The rates are exact under normal sampling with known standard errors and a common per-participant spread: a one-dimensional integral over the subgroup’s z of the conditional normal for the whole’s, and closed forms for the difference tests. The scatters are five hundred seeded trials and are illustrations; their correlation is checked against f\sqrt f and no rate is read off them.

Not claimed: that a subgroup’s effect is never different, nor that the recovery from a table is exact when the subgroup’s spread differs from the rest’s — the share read off the standard errors is then a share of information rather than of participants, and the rest’s estimate is correct only if the paper’s overall estimate is the information-weighted average. Adjusted analyses, where the overall estimate is not a simple weighted average of subgroups, break the identity and need the interaction term the trial’s own model would give.

Still open: many subgroups of one trial

Every number here is one subgroup against its whole. A results table reports five or eight at once — sex, age band, site, severity, baseline risk — each a part of the same trial and overlapping one another as well as the whole. The pairs are then correlated in a pattern set by how the subgroups intersect, and the chance that some subgroup reads “not significant” while the whole does is the chance that the smallest of several correlated, underpowered estimates falls below the line.

That chance is close to one at ordinary sizes by the arithmetic of twenty analyses of nothing, and a table of eight subgroups is almost certain to contain the sentence this essay prices. What is not computed is the right family-wise treatment of the subgroup-against-the-rest tests when the subgroups overlap — how many effective comparisons a table of eight intersecting subgroups amounts to — which is the effective-count question asked of a structure a trial designs in advance rather than one an analyst wanders into.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CorrelationInteractionLikelihood ratiop-valueStandard errorStatistical powerSubgroup analysisTwo-sample test