An interval read beside something else

Sixteen subgroups and one effect

A trial of four hundred with 80% power tabulates its result by eight binary characteristics — sixteen subgroups — and the treatment's effect is the same for every patient. Among trials significant overall, 93.5% report at least one subgroup where it is not, and the median one reports seven. Some subgroup-against-the-rest interaction reaches 5% in 34.6% of trials — eight independent chances — and when the characteristics overlap as two age thresholds do, 18.6%, about four. The effective count is set by the correlation between the characteristics' split indicators, (2/π) arcsin of their latent correlation, and a Bonferroni reading of the eight holds the rate at 5% whatever the overlap.

Worth reading first: What the 95% refers to.

A subgroup inside its own trial priced one subgroup read against the trial it belongs to. A subgroup of share ff is correlated f\sqrt f with its whole, the right test of “the subgroup differs” is the subgroup against the rest, and the sentence “significant overall, not in the subgroup” appears 57.45% of the time with the effect the same everywhere and 58.26% with no effect in the subgroup at all — a likelihood ratio of 1.014, which is to say none.

A results table does not report one subgroup. It reports five or eight characteristics — sex, age band, site, baseline severity, prior treatment — each splitting the trial in two, and every subgroup overlaps its whole and many of the other subgroups. The essay ended on that table: how many effective comparisons a table of intersecting subgroups amounts to, and what a family-wise reading of it should use.

How often a trial whose effect is the same for everyone shows a significant subgroup interaction, by the size of its subgroup tableFour hundred patients, 80% power overall, characteristics whose latents correlate 0 (their split indicators 0.000). Some interaction test reaches 5% in 5.10% at 1, 9.13% at 2, 17.43% at 4, 34.63% at 8, 47.13% at 12, 56.23% at 16; the equicorrelated-normal integral gives 5.00%, 9.75%, 18.55%, 33.66%, 45.96%, 55.99%. A Bonferroni reading of the eight holds 5.57%.00.2000.4000.60012481216binary characteristics tabulated (two subgroups each)trials with some interaction test at 5%some interaction at 5%, countedthe equicorrelated-normal integralBonferroni over the characteristics3,000 trials a point, effect the same everywheredashed red: 5%
Fig. 1 How often a trial of four hundred whose treatment effect is the same for every patient shows some subgroup-against-the-rest interaction at 5%, against the number of binary characteristics tabulated, counted and by an exact equicorrelated-normal integral; with a Bonferroni reading over the characteristics beside it. The slider sets how much the characteristics overlap.

The trial and its table

Four hundred patients, half to each arm, an outcome measured on a continuous scale, and a treatment effect that gives the whole trial 80% power and is identical for every patient. Each patient has eight binary characteristics. In the plainest case they are independent, as sex and enrolment site roughly are. In the others they overlap: each characteristic is the side of the median a latent measurement falls on, and the latents share a common factor, so that at a latent correlation of 0.9 two characteristics are as alike as “older than 60” and “older than 65”.

For each characteristic the table gives the effect in both of its subgroups, with a confidence interval, and the interaction test of one subgroup against the other. Nothing the table contains reflects a real difference: every subgroup has the same true effect as every other.

One trial's forest plot of sixteen subgroups, with the treatment effect the same for every patient. The whole trial estimates 0.215 with a 95% interval from 0.019 to 0.411. Of the sixteen subgroups, 13 have intervals that include zero; the true effect in every one of them is 0.280.
Fig. 2 One trial’s forest plot: sixteen subgroups from eight independent characteristics, each with its 95% interval, and the whole trial below. The dashed line is the true effect, the same in every subgroup.

The single trial drawn above is significant overall, and thirteen of its sixteen subgroups have intervals that include zero. Three are significant, scattered among characteristics with nothing in common. It is an ordinary trial.

What the subgroup verdicts say

With one characteristic, two subgroups of about two hundred each, a trial significant overall reports at least one of its two subgroups as not significant 69.7% of the time — each half has only about 50% power. With eight characteristics the figure is 93.5%; a significant trial reports on average 6.36 of its sixteen subgroups as not significant, and the median is seven.

How many of 16 subgroups read "not significant" when the whole trial is significant and the effect is the same in all of them. Over 2,411 trials significant overall, none of the 16 subgroups is non-significant in 6.51%; the median count is 7.
Fig. 3 How many of sixteen subgroups read “not significant” in trials that are significant overall, with the effect the same in all of them.

So a table in which every subgroup is significant is the unusual result, and a table with several “not significant” rows is what a homogeneous effect produces. Significant in one, not in the other found that a split between two verdicts is common and rarely real; a table of sixteen verdicts is sixteen chances to split, and at a subgroup’s power of about one half, it nearly always does.

Some subgroup’s point estimate also falls on the wrong side of zero — a treatment that helps everyone appearing to harm a subgroup — in 19.2% of trials, and with sixteen characteristics in 26.6%. That is the row a reader remembers.

The row a reader remembers

A subgroup whose estimate points the wrong way is the one that attracts explanations. It prompts a safety discussion, a mechanistic story about why those patients might respond differently, a recommendation to study them separately. In a trial of four hundred with a homogeneous effect, one trial in five supplies such a row by chance, and the row’s interval nearly always includes both zero and the true effect, which is what the explanation ignores.

The comparison a reader instinctively makes is between the wrong-way subgroup’s interval and the others’, and it is the comparison two intervals that overlap showed is miscalibrated: two 95% intervals can overlap substantially while their difference is significant, and can fail to overlap by little while it is not. The interaction test is the comparison with the right variance, and in a table of eight it reaches 5% for some characteristic a third of the time with nothing to find. The wrong-way row is usually not even that; it is a half-sized trial’s estimate landing below zero, which at 50% power happens to about one subgroup in forty on its own and to one table in five.

It is also the row a replication most often fails to reproduce, for the same reason a single study’s interval captures an equal-sized replication only five times in six: both the original subgroup and its replication are noisy, and the replication regresses towards the common effect the original strayed from.

Why every subgroup has about half the power

The numbers follow from one line of arithmetic. A trial designed for 80% power has an expected z statistic of 2.80. A subgroup holding half the patients has half the information, so its expected z is 2.80/2=1.982.80/\sqrt2 = 1.98 — almost exactly the significance threshold — and its power is 50.8%. Every subgroup in an evenly split table is a coin flip for significance, whatever the treatment does in it.

That is why the count of non-significant rows says nothing about heterogeneity. Sixteen coin flips produce eight tails on average, fewer here because the subgroups of a significant trial lean towards significance with it; a trial significant overall still reports six or seven “not significant” rows with the effect identical everywhere. A reader who takes those rows as evidence that the treatment “does not work” in those patients is reading the design’s power, not the patients.

The rows that are not coin flips are the ones a design would have to add patients for. To give each half of the trial 80% power on its own, the trial would need twice the patients; to give a subgroup of a quarter of the patients the same, four times. Subgroup tables are almost never designed that way, and their rows should be read as what they are: estimates from half-sized trials.

What the interaction tests say

The interaction tests are the right tests to read, and they hold their level one at a time: with one characteristic, the subgroup-against-the-rest test reaches 5% in 5.1% of trials. A table of eight holds eight of them. With the characteristics independent, some interaction reaches 5% in 34.6% of trials; with sixteen, 56.2%.

As a count of independent chances, eight independent characteristics are 8.29 tests and sixteen are 16.1 — exactly their number, within the noise of three thousand trials. Two characteristics that share no patients’ traits produce subgroup contrasts that share no noise, so each is a fresh chance at a false interaction.

How overlap thins the count

Overlapping characteristics are not fresh chances, and the reduction has an exact form. The interaction contrast for a characteristic is, in effect, a weighted sum of every patient’s outcome with a weight of plus or minus one by the side of the split they fall on. Two such contrasts are correlated by the correlation between the two characteristics’ indicators, and for median splits of latents correlated ρ\rho that is Sheppard’s (2/π)arcsin⁡ρ(2/\pi)\arcsin\rho — the formula the arcsine that closes it evaluated for cut points.

How many independent interaction tests 8 overlapping characteristics amount to, by how much they overlap. Eight independent characteristics are 8.00 independent tests; at an indicator correlation of 0.333 (latents at 0.5) the integral gives 6.74 and the count 7.07; at 0.713 (latents at 0.9), 4.03 and 4.01.
Fig. 4 The effective number of independent interaction tests eight characteristics amount to, against the correlation between their split indicators, by the equicorrelated-normal integral; the dots are the counts from three thousand trials at three overlaps.

With the interaction statistics treated as equicorrelated normals at that correlation, the chance that some one of KK exceeds 1.96 is a single integral over their shared factor, and it agrees with the count. At a latent correlation of 0.5 the indicators correlate 0.333, and eight characteristics amount to about 6.7 independent tests by the integral and 7.07 by the count; some interaction reaches 5% in 30.4% of trials. At a latent correlation of 0.9 the indicators correlate 0.713 and eight characteristics amount to about 4.0; some interaction reaches 5% in 18.6%.

How often a trial whose effect is the same for everyone shows a significant subgroup interaction, by the size of its subgroup table. Four hundred patients, 80% power overall, characteristics whose latents correlate 0.9 (their split indicators 0.713). Some interaction test reaches 5% in 5.40% at 1, 8.17% at 2, 12.47% at 4, 18.60% at 8, 23.47% at 12, 25.50% at 16; the equicorrelated-normal integral gives 5.00%, 8.28%, 12.84%, 18.68%, 22.63%, 25.65%. A Bonferroni reading of the eight holds 3.17%.
Fig. 5 The same table with characteristics as alike as two nested age thresholds, latents correlated 0.9. The family-wise rate grows far more slowly with the table’s size, and Bonferroni becomes conservative.

The curve is flat where most tables sit. Characteristics correlated at the levels baseline variables usually are — a third, a half — remove a chance or two from eight. Only characteristics that are nearly the same variable, such as nested age thresholds or two scores of one severity, thin the count substantially. A table of eight subgroups is worth close to eight tests.

What Bonferroni does with the table

Dividing the 5% among the characteristics — each interaction tested at 5% divided by eight — holds the chance of some false interaction at 5.6% with independent characteristics, 4.3% at a latent correlation of 0.5 and 3.2% at 0.9, in three thousand trials. It is right where the characteristics are independent and conservative where they overlap, and it never breaks: the correlation only makes it waste some of its level, which what the correction corrects found is the general behaviour of a Bonferroni correction on correlated tests.

The equicorrelated integral says how much is wasted and gives the exact threshold instead. At an indicator correlation of 0.713 the eight interactions are worth four independent tests, so a threshold set for four keeps the family at 5% and is less stringent than the one set for eight. That is a small gain, and it needs the overlap to be known — which, for characteristics measured at baseline, it is: the correlation between two split indicators can be computed from the trial’s own baseline data before any outcome is seen, the way the design a report could state can be.

Unequal subgroups are worse

Every characteristic here splits the trial in half, which is the most favourable case. Real characteristics rarely do. A subgroup holding a fifth of the patients has an expected z of 2.800.2=1.252.80\sqrt{0.2} = 1.25 in a trial designed for 80% power, and its power is 24.0%: three times in four it reads “not significant” while the whole trial and the other four fifths read significant. The interaction test against the rest is weaker too, because its variance is dominated by the small group, so a real difference there is harder to detect and a false one no rarer to produce.

So a table mixing a few balanced characteristics with several lopsided ones — a rare comorbidity, a small site, an unusual prior treatment — has even more “not significant” rows than the sixteen-subgroup table, and its wrong-way estimates concentrate in the small subgroups, where a half-dozen patients decide the sign. The subgroups a reader is most likely to find alarming are the ones the trial knows least about.

The subgroups a protocol names

A protocol can separate the table into two families. A characteristic named in advance with a stated reason to expect a difference — a mechanism, an earlier trial’s finding — is tested on its own at 5%, because the family it belongs to has one member. The rest are exploratory, and the effective count above is the right discount for them. Naming a handful in advance priced the choice of how many to name: each additional named analysis costs power on all of them, and the best list is a short one holding most of the belief about where a difference might be.

What the table should not do is present the named characteristic and the exploratory ones as rows of equal standing. A reader shown sixteen rows with sixteen intervals has no way to tell which one the trial was designed to answer, and the design’s protection of that one question is lost in the table’s layout.

What the family’s threshold costs a real difference

Holding the table’s family at 5% is not free, and the cost falls on the interactions that are real. An interaction test between two halves of a trial already needs four times the patients of the main comparison to detect a difference in effect as large as the effect itself, because it compares two half-sized estimates rather than one whole one; significant in one, not in the other priced that ratio, and the verdicts watch the threshold found the split between verdicts a far weaker guide to it. Testing it at 5% divided by eight rather than at 5% raises the threshold from 1.96 to 2.73 standard errors, and a real interaction that had a coin-flip’s chance at 1.96 has about one in five at 2.73.

That is the honest shape of the choice. A trial of a size designed for its main question cannot also answer eight heterogeneity questions; it can protect itself against false answers to all eight or give itself a modest chance at one of them, and a subgroup table read row by row does neither while appearing to do both.

Why a table is designed and a forking path is not

How many analyses there really were asked the effective-count question of analyses an analyst wandered into, where the family is unknown and has to be bounded. A subgroup table is the easier case. Its characteristics are chosen in the protocol, their overlap is measurable at baseline, and the family is on the page. Nothing prevents it being read correctly except the habit of reading it row by row.

That habit is expensive in a specific way. The row-by-row reading asks, of each subgroup, whether the treatment “works in” it, which is the verdict question the whole of this family of results has shown to be uninformative about differences. The table’s honest content is the interaction tests, read as a family: eight tests with a stated overlap, a threshold that holds the family’s error, and, in a trial whose effect is the same for everyone, nothing crossing it nineteen times in twenty.

What a reader of a subgroup table can do

Count the characteristics, not the rows. Sixteen subgroups from eight binary characteristics are eight interaction tests, not sixteen. A table with three-level characteristics has two contrasts for each, and those count separately.

Expect “not significant” rows in a significant trial. At a subgroup’s typical power, seven of sixteen is the median; none is the exception.

Discount a lone significant interaction by the table’s size. One interaction at p = 0.03 in a table of eight is what a homogeneous effect produces in a third of trials, and is not evidence of heterogeneity without a stated prior reason to look at that characteristic.

Ask for the overlap. Characteristics that are nearly the same variable — nested cut points, two measures of one thing — shrink the effective count, and a report that tabulates them should say so, because it changes the threshold a reader should apply.

What is counted here and what is exact

With eight binary characteristics and the effect the same everywhere, 93.5% of trials significant overall report a non-significant subgroup, and some interaction reaches 5% in 34.6% of trials, about 8.3 independent tests.

At a latent correlation of 0.9 between characteristics, some interaction reaches 5% in 18.6% of trials, about four tests; the equicorrelated-normal integral with the indicators’ correlation (2/π) arcsin ρ gives the same rates.

Bonferroni over the characteristics holds the family at 5.6%, 4.3% and 3.2% at latent correlations of 0, 0.5 and 0.9.

Every rate is counted over three thousand simulated trials of four hundred, with a normal outcome of known variance and characteristics split at their latents’ medians; the same trials are scanned with more characteristics, so the curves over table size are comparisons on identical data.

Not counted: characteristics with more than two levels, unequal splits — a subgroup of a tenth has far less power and makes a “not significant” row still more likely — and outcomes whose variance differs between subgroups, where the interaction test needs its own variance estimate.

Still open: a heterogeneity that is really there

Every table here is drawn with a homogeneous effect, so every interaction is false and every “not significant” row is noise. The question a trial actually faces is the mixed one: one characteristic genuinely modifies the effect and seven do not. A family-wise threshold then protects against the seven and costs power on the one.

What the table’s structure does to that trade — how much power a Bonferroni or equicorrelated threshold costs against a real interaction of a stated size, whether the real one is more likely than a false one to be the table’s largest, and how often a trial with one genuine modifier reports a different one — is the calculation that would turn a protocol’s subgroup list into a design with known properties, and it has not been made here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BonferroniCorrelationEffective number of testsFamilywise error rateInteractionMultiple comparisonsStatistical powerSubgroup analysis