Concept

Subgroup analysis — where it appears

The estimation of a treatment's effect within parts of a trial defined by a baseline characteristic. Its verdicts are underpowered and correlated with the whole trial's, so a difference between subgroups needs its own test rather than a comparison of two significance verdicts.

Named by 4 essays across 2 fields — each of them below, with the objects they name alongside it.

Two studies of the same effect, z statistics with mean 1.96: where one is significant and the other is not. Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.

Significant in one, not in the other

Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.

alongside · Repetition
Where two verdicts disagree, across the true effects of two studies, beside where the test of their difference finds it. The shading is the chance that exactly one of two independent studies is significant, over the plane of their true z means. It is highest on a cross centred where either mean is 1.96 and lowest in the corner where both are large, whatever the difference between them. The difference test's power is constant along lines parallel to the diagonal: 50% on the inner pair and 80% on the outer. Two studies at means 8 and 4 — one effect half the other — split 2.1% of the time, and the difference test finds the halving 80.7% of the time.

The verdicts watch the threshold, not the gap

Two studies whose effects really differ are read through their verdicts as often as through their difference. Halves of an 80%-power study whose effects differ by the whole effect split 72.87% of the time against 49.99% when they do not — a split is never more than 1.86 times as likely under any difference up to twice the effect. And a halving of the effect in two large studies leaves both significant 97.93% of the time, while the test of the difference finds it 80.74%.

alongside · Repetition
A whole trial's z against a subgroup's, the subgroup 20% of it with the same effect as the rest. Five hundred pairs, correlated at 0.447 — the square root of its share — because the subgroup is part of the whole. The whole trial is significant and the subgroup is not in 57.5% of trials; the subgroup differs significantly from the rest in 5.0%. The dashed lines mark 1.96 on each axis and the slanted lines mark a significant difference between the subgroup and the rest.

A subgroup inside its own trial

A trial with 80% power, significant overall and not in the fifth of it that is women, has told the reader almost nothing about women: that pattern turns up 57.45% of the time when women have the full effect and 58.26% when they have none. The subgroup is part of the whole, so the two verdicts are correlated, and the one test that answers the question — the subgroup against everyone else — can be recovered from the two printed intervals and nothing more.

alongside · Repetition
How often a randomised trial of eighty shows a Simpson reversal on some baseline variable, against how many were tabulated. The treatment helps every patient equally. On one nominated, strongly prognostic variable the reversal appears in 4.1% of trials. Tabulating 2, 5, 10, 20, 50, 100 variables of mixed prognostic value, some variable reverses in 4.1%, 6.5%, 10.8%, 16.1%, 26.1%, 34.1% of trials; stratifying the randomisation on the most prognostic one gives 0.2%, 2.9%, 7.7%, 12.2%, 23.4%, 31.9%.

Twenty variables nobody stratified on

A randomised trial of eighty patients in which the treatment helps every patient by the same six points shows Simpson's reversal on one nominated, strongly prognostic baseline variable in 4.07% of trials. Tabulate twenty baseline variables of mixed prognostic value and some variable reverses in 16.07% of trials; a hundred, in 34.13%. Stratifying the randomisation on the strongest variable removes its reversal and leaves 12.17% and 31.93%: the protection is exactly as wide as the list it was given.

reversal · Simpson

Named alongside it

The objects these essays reach for when they reach for this one.

Interactionp-valueStatistical powerTwo-sample testLikelihood ratioMultiple comparisonsSample sizeBaseline imbalanceCorrelationEffective number of testsPrognostic factorRandomisation

All concepts