A subgroup inside its own trial
Worth reading first: Twenty intervals and one expected miss.
The subgroup comparisons measured so far were between two groups that shared nothing: two studies, or two disjoint halves of one. Significant in one, not in the other found that such a pair splits about significance half the time with identical effects, and the verdicts watch the threshold, not the gap found that the split responds to where the effects sit rather than how far apart they are.
The pair that actually appears in results sections is usually different. It is the effect in everyone, reported beside the effect in a subgroup of those same people — all patients and then women, the whole trial and then its largest site, every participant and then those over seventy. The subgroup’s estimate is built from part of the data that built the overall one. The two are not independent, and the arithmetic that treats them as though they were is wrong in a direction that can be named.
The correlation that sharing forces
Let a trial have participants and a subgroup a share of them. With a common spread per participant, the overall estimate has variance and the subgroup’s . The overall estimate is a weighted average of the subgroup’s and the rest’s,
and since the subgroup and the rest share no participants, the covariance of the whole with the subgroup is — exactly the variance of the whole. Two exact facts follow.
The z statistics are correlated at . A subgroup of a fifth of the trial moves with the whole at a correlation of 0.447; a half, at 0.707. The hero figure’s cloud is tilted accordingly: a trial whose overall z happened to come out high tends to have a high subgroup z too, because the subgroup’s participants are part of why the overall z came out high.
The variance of the difference between them is a subtraction, not a sum. Since the covariance equals the whole’s variance,
where two independent estimates would add. And the difference itself is a fixed multiple of something familiar: . Divided by its correct standard error it is the same z statistic as the test of the subgroup against everyone else, not a similar one. Comparing a part with the whole, done correctly, is comparing the part with the rest.
Significant overall, not in the subgroup
Now read the pair the way results sections do. The trial is sized for 80% power on its overall effect, the effect is the same in everyone, and the question is how often the sentence “the effect was significant overall but not in women” can be written.
A subgroup of a fifth of the participants has a z statistic with mean , which is 24.04% power. So the subgroup is non-significant most of the time while the whole is significant most of the time, and the pair “significant overall, not in the subgroup” appears in 57.45% of trials. With the subgroup at half the trial its power is 50.84% and the pattern appears in 31.29%.
The correlation does what the essay on the verdicts expected it to do: it makes the two verdicts move together. With the subgroup at half the trial, the verdicts disagree — one significant, the other not, in either order — in 33.39% of trials. Two disjoint halves of the same trial, each at the same 50.84% power, disagree in 49.99%. Sharing data removes a third of the splits, because a trial that did well overall has usually done well in its half too.
What it does not do is make the splits that remain meaningful. Among trials significant overall and not in a half-share subgroup, the subgroup differs significantly from the rest in 6.28% — and every one of those is a false alarm, since the effect is the same everywhere. At a fifth of the trial the rate is 3.46%. The reverse split, significant in the subgroup and not overall, is rare, 1.51% at a fifth and 2.10% at a half, since a subgroup can clear the threshold on its own only by an excursion the whole rarely fails to share.
The same pattern whatever the subgroup’s truth
The question a reader of “significant overall, not in women” is asking is whether women respond differently. So the pattern has to be priced against the alternatives: women with half the effect, women with none.
When the subgroup has no effect at all — the reading taken literally — the overall effect is diluted, the whole trial is significant less often, and the pattern appears in 58.26% of trials against 57.45% when the effect is the same. Its likelihood ratio for “women get nothing” against “women get the full effect” is 1.014. The sentence is written almost exactly as often in the world where it is true as in the world where it is false.
| subgroup’s share | same effect | half the effect | no effect | none against same |
|---|---|---|---|---|
| 0.2 | 57.45% | 62.68% | 58.26% | 1.014 |
| 0.3 | 48.18% | 55.35% | 47.55% | 0.987 |
| 0.5 | 31.29% | 39.97% | 26.48% | 0.846 |
| 0.7 | 17.33% | 24.27% | 10.89% | 0.628 |
Two things in the table are the opposite of what the reading assumes. Below about a third of the trial the pattern is indifferent to the subgroup’s truth: it is manufactured by the subgroup’s low power, and the subgroup’s power is set by its size, not by its effect. Above a third the pattern becomes less likely when the subgroup has no effect, not more, because a large subgroup with no effect drags the whole trial below significance and so removes the first half of the sentence. At a subgroup of half the trial, a reader who sees “significant overall, not in the subgroup” should, if anything, think slightly better of the subgroup’s effect than before — the reverse of the sentence.
The intermediate truth, half the effect, produces the pattern most often at every share. That is a combination of a subgroup weak enough to miss significance and a whole strong enough to reach it, and it is the one case the reading gets roughly right. Even there its likelihood ratio against the same effect is 1.091 at a fifth of the trial. A reader who multiplies their odds by that has barely moved them.
The test the two estimates can support
The sentence stands in for a test of whether the subgroup’s effect differs from the rest’s. A paper usually prints enough to run it, and the commonest way of running it is wrong.
The tempting calculation takes the overall estimate and the subgroup’s, each with its standard error, and tests their difference as if they were two independent studies — the calculation the essay on the split recommended for disjoint groups. For a part and its whole it divides by where the correct denominator is . The ratio of the two is : 0.816 at a fifth of the trial and 0.577 at a half. Every statistic is shrunk by that factor, so the test is run at a much stricter level than it claims.
At a fifth of the trial the independent calculation has an actual size of 1.64% rather than 5%, and at a half 0.069%. Its power falls with it. Against a subgroup with no effect at all, the correct comparison detects the difference 20.17% of the time at a fifth of the trial and 28.84% at a half; the independent calculation detects it 10.05% and 2.31% of the time. At a half, the test a reader is most likely to run on a printed table has thrown away more than nine tenths of the power the data contain, which were not much to begin with.
So the shared data cut both ways, and the net is not in doubt. The correlation lowers the rate of spurious splits, which helps; it also invalidates the one test a reader can run from the two numbers in front of them, unless the reader knows to subtract.
The forest plot’s own reading
A subgroup table is usually drawn as a forest plot: each subgroup’s estimate and interval on its own row, the overall estimate as a vertical line down the page. The reading that goes with it is the reverse of the sentence priced above — if every subgroup’s interval crosses the overall line, the effect is “consistent across subgroups”, and the trial moves on.
For a subgroup that is part of the whole, that reading is nearly automatic. The gap between the subgroup’s estimate and the overall one has standard deviation , smaller than the subgroup’s own standard error, because the overall estimate already contains the subgroup. So the subgroup’s 95% interval contains the overall estimate with probability when the effect is the same everywhere: 97.16% for a fifth-share subgroup and 99.44% for a half.
Read as a test of “this subgroup differs”, crossing the line is a test run at 2.84% for a fifth and 0.557% for a half, and it is weaker than the subgroup-against-the-rest test for the same reason the independent calculation was: it measures the gap against a standard error that is too large. Against a subgroup with no effect at all it fails to cross the line 14.26% of the time at a fifth of the trial and 8.52% at a half, where the right comparison detects the difference 20.17% and 28.84% of the time. The forest plot’s reading, like the verdicts’ in the essay on the threshold, is an instrument set to find sameness, and at a half-share subgroup it finds it more than ninety-one times in a hundred when the subgroup gets nothing.
Both readings are therefore wrong in the direction that suits the page they appear on. The verdicts’ sentence invents a difference that the numbers do not support; the forest plot’s lines conceal a difference that the numbers could partly support. Neither is the comparison of two intervals read against each other, which at least uses the right pair; both compare a subgroup with a whole that contains it.
The subgroup that is significant when the trial is not
The reverse split — a subgroup significant while the whole trial is not — is rare, 1.51% of trials at a fifth-share subgroup, and it is the one most likely to be written up as a discovery: the drug failed overall but worked in the subgroup.
Its rarity is the warning. A subgroup reaches significance on its own, with a quarter of the trial’s power, only by an excursion well above its true effect, and a trial that did not share the excursion is one in which the rest of the data pulled the other way. So the subgroup’s estimate in that pattern is selected twice over — once by crossing 1.96 with a small sample, which is the winner’s curse at a quarter of the trial’s power, and once by the overall result’s failure, which says the rest did not agree. A subgroup found this way after the fact, among several, is the textbook case of a p-value that does not carry its own selection, and the effect estimate it reports is the least trustworthy number in the paper.
Recovering the comparison from a table
The correct test needs nothing that a table of whole-trial and subgroup intervals does not print, including — which is less obvious — the subgroup’s share of the trial. With a common per-participant spread, the share is the ratio of the two variances, .
Suppose a trial reports an overall effect of 4.0, 95% interval 1.6 to 6.4 (), and an effect in women of 2.0, interval −1.8 to 5.8 (), and says that the benefit was not seen in women. The standard errors are the widths over 3.92: 1.2245 and 1.9388. Then:
- the women’s share is 0.399, about two fifths of the trial;
- the effect in everyone else is 5.33, with a standard error of and an interval from 2.23 to 8.42;
- women against the rest differ by −3.33, standard error 2.501, interval −8.23 to 1.57, p = 0.183.
The same z statistic, −1.331, comes from dividing the women’s estimate minus the overall one by , which is the identity above written in the table’s own numbers. The calculation that treats the two as independent gives instead — more than twice as large, and a reader running it would conclude that the data are even less informative about women than they are.
The honest summary is the interval for the difference, and it is wide: the data are compatible with women benefiting 1.57 units more than men and with their benefiting 8.23 units less. That is a statement about how little a two-fifths subgroup of an 80%-power trial can say about a difference, and it is the statement the sentence “the benefit was not seen in women” replaces with something that sounds like a finding.
Why sharing does not rescue the reading
The effect of the correlation on the split was the question left open when two disjoint groups were compared, and it has an answer: sharing data makes the verdicts agree more often, a third fewer splits at half the trial. It would be natural to conclude that the nested pair is therefore a more trustworthy reading than the disjoint pair. It is not, for a reason the geometry of the hero figure shows.
The correlation tilts the cloud towards the line along which the whole and the part move together, and the test that matters — subgroup against rest — is measured across that line, in the direction the correlation does nothing to reduce. Fewer trials cross the verdict lines, but the ones that do cross them for the same reason as before: the subgroup is small and its estimate is noisy. Its power is set by its share alone — 24.04% at a fifth, whatever the rest is doing — and a verdict based on a quarter-powered estimate is mostly a statement about the quarter.
The same point in terms of evidence: the likelihood ratios in the table are close to one at every share below a third, because the thing that decides whether the subgroup’s verdict is “not significant” is how many people are in it. Sharing data with the whole makes the whole’s verdict partly predictable from the subgroup’s, which is why the rates change; it does not give the subgroup’s verdict any information about the subgroup’s effect that its sample size did not already fix.
What sharing data changes, exactly, and what it leaves alone
The whole trial and a subgroup of share have z statistics correlated at exactly , and the difference between their estimates is times the subgroup-against-the-rest difference, with a variance that is the subgroup’s minus the whole’s.
At 80% power overall, “significant overall, not in the subgroup” appears in 57.45% of trials with a fifth-share subgroup and the same effect everywhere, and in 58.26% when the subgroup has no effect, a likelihood ratio of 1.014. With a half-share subgroup it is 31.29% and 26.48%, so the pattern is less likely when the subgroup gets nothing.
Treating the whole and the subgroup as independent runs the test at 1.64% for a fifth-share subgroup and 0.069% for a half, with power 10.05% and 2.31% against a subgroup with no effect, where the subgroup-against-the-rest test has 20.17% and 28.84%.
The subgroup’s share and the rest’s estimate can both be recovered from two printed intervals, under a common spread, and the worked example gives the comparison a paper summarised as “not seen in women” a p-value of 0.183.
The rates are exact under normal sampling with known standard errors and a common per-participant spread: a one-dimensional integral over the subgroup’s z of the conditional normal for the whole’s, and closed forms for the difference tests. The scatters are five hundred seeded trials and are illustrations; their correlation is checked against and no rate is read off them.
Not claimed: that a subgroup’s effect is never different, nor that the recovery from a table is exact when the subgroup’s spread differs from the rest’s — the share read off the standard errors is then a share of information rather than of participants, and the rest’s estimate is correct only if the paper’s overall estimate is the information-weighted average. Adjusted analyses, where the overall estimate is not a simple weighted average of subgroups, break the identity and need the interaction term the trial’s own model would give.
Still open: many subgroups of one trial
Every number here is one subgroup against its whole. A results table reports five or eight at once — sex, age band, site, severity, baseline risk — each a part of the same trial and overlapping one another as well as the whole. The pairs are then correlated in a pattern set by how the subgroups intersect, and the chance that some subgroup reads “not significant” while the whole does is the chance that the smallest of several correlated, underpowered estimates falls below the line.
That chance is close to one at ordinary sizes by the arithmetic of twenty analyses of nothing, and a table of eight subgroups is almost certain to contain the sentence this essay prices. What is not computed is the right family-wise treatment of the subgroup-against-the-rest tests when the subgroups overlap — how many effective comparisons a table of eight intersecting subgroups amounts to — which is the effective-count question asked of a structure a trial designs in advance rather than one an analyst wanders into.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Estimating how many nulls are true — both name correlation, p-value, statistical power
- The smallest of three combinations — both name correlation, p-value, statistical power
- A copula that halves a marginal — both name correlation, interaction
- A coverage table with its own error — both name standard error, statistical power
- A cut is not a polynomial, and it does not have to be — both name correlation, interaction
- A dictionary that is a product — both name correlation, interaction
Named objects
A flat tag is an object no other essay names yet.
CorrelationInteractionLikelihood ratiop-valueStandard errorStatistical powerSubgroup analysisTwo-sample test