Subgroup analysis — where it appears
Named by 4 essays across 2 fields — each of them below, with the objects they name alongside it.
Significant in one, not in the other
Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.
The verdicts watch the threshold, not the gap
Two studies whose effects really differ are read through their verdicts as often as through their difference. Halves of an 80%-power study whose effects differ by the whole effect split 72.87% of the time against 49.99% when they do not — a split is never more than 1.86 times as likely under any difference up to twice the effect. And a halving of the effect in two large studies leaves both significant 97.93% of the time, while the test of the difference finds it 80.74%.
A subgroup inside its own trial
A trial with 80% power, significant overall and not in the fifth of it that is women, has told the reader almost nothing about women: that pattern turns up 57.45% of the time when women have the full effect and 58.26% when they have none. The subgroup is part of the whole, so the two verdicts are correlated, and the one test that answers the question — the subgroup against everyone else — can be recovered from the two printed intervals and nothing more.
Twenty variables nobody stratified on
A randomised trial of eighty patients in which the treatment helps every patient by the same six points shows Simpson's reversal on one nominated, strongly prognostic baseline variable in 4.07% of trials. Tabulate twenty baseline variables of mixed prognostic value and some variable reverses in 16.07% of trials; a hundred, in 34.13%. Stratifying the randomisation on the strongest variable removes its reversal and leaves 12.17% and 31.93%: the protection is exactly as wide as the list it was given.
Named alongside it
The objects these essays reach for when they reach for this one.
Interactionp-valueStatistical powerTwo-sample testLikelihood ratioMultiple comparisonsSample sizeBaseline imbalanceCorrelationEffective number of testsPrognostic factorRandomisation