Twenty variables nobody stratified on
Worth reading first: Simpson's reversal is a region, not a table.
The reversal a coin cannot prevent found that correct randomisation removes Simpson’s reversal in expectation and not in any one trial: a trial of eighty, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials, and stratifying the randomisation on the group removes the reversal entirely. It ended with the part of the problem that stratification cannot touch. A trial stratifies on one or two variables. A patient carries thousands, a baseline table reports dozens, and the reversal is available on every variable that was not stratified on — and the rate that matters, when a subgroup table is scanned for exactly this pattern, is the rate of some variable reversing.
This essay counts it.
Twenty readings of one risk
A baseline table does not hold twenty independent variables. It holds age, disease stage, a biomarker, a comorbidity count, prior treatment — readings, of varying quality, of the same underlying thing: how sick the patient is. The model here makes that explicit. Each patient has an underlying risk , normal; the outcome’s probability rises with it, from about 17% on average for patients below the middle to about 76% above; and each of binary baseline variables is a dichotomised, noisy reading of , with its correlation with running evenly from 0.95 — nearly the risk itself — down to 0.10, nearly useless. The treatment adds six points to every patient’s chance of a good outcome, in every subgroup, always.
Each variable splits the trial into two subgroups, and the subgroup table for that variable reverses if the treatment is ahead in both subgroups and behind in the whole trial, or the other way round. With the treatment genuinely helping everyone, every reversal is an artefact of the particular allocation of patients to arms.
The mechanism, on one variable
The single-variable result is worth restating because every reversal in the tables below is an instance of it. A subgroup table reverses when the arms’ shares of the high-risk subgroup differ by , the subgroups’ risks differ by , and the product exceeds the treatment’s advantage within the subgroups. Randomisation makes zero on average and small in a large trial, but in a trial of eighty its standard deviation is about eleven percentage points, and a risk gap of fifty points turns an eleven-point imbalance into a five-point swing against a treatment worth six. On one prognostic variable that is enough about four times in a hundred.
With many variables the same inequality is checked many times, and each variable brings its own — the arms’ imbalance on that particular reading — and its own , which is large only for the prognostic ones. The tables below are the chance that the inequality holds for at least one of them.
One variable, and then many
On the single most prognostic variable, nominated in advance, the reversal appears in 4.07% of trials of eighty — close to the 3.40% the essay on a single variable found with two perfectly separated groups. That rate is what a reader who asks about one variable is exposed to.
A reader who scans a baseline table is exposed to much more:
| variables tabulated | some variable reverses | stratified on the strongest | effective independent variables |
|---|---|---|---|
| 1 | 4.07% | 0.00% | 1.0 |
| 5 | 6.50% | 2.90% | 4.5 |
| 10 | 10.80% | 7.73% | 8.6 |
| 20 | 16.07% | 12.17% | 13.0 |
| 100 | 34.13% | 31.93% | 31.2 |
With twenty variables tabulated, one trial in six shows the pattern somewhere; with a hundred, one in three. None of those trials contains a subgroup in which the treatment does anything but help by six points.
The last column is the number of independent variables, each reversing at the variables’ average rate, that would produce the same chance of some reversal. Twenty variables that are all readings of one risk behave like thirteen independent ones; a hundred, like thirty-one. The correlation through the shared risk makes the variables’ reversals overlap — a trial that happened to put sicker patients in one arm tends to show it on several readings of sickness at once — but not by much, because the noise in each reading is private and the reversals are driven as much by that noise as by the shared imbalance. It is the effective-count question for a family of subgroup tables, and the answer is that a correlated baseline table is worth most of its length.
Which variables can reverse
The reversal needs something specific from a variable: the two subgroups it defines must differ in their baseline risk, because the reversal is the arms’ different mixtures of high- and low-risk patients outweighing the treatment’s effect. A variable unrelated to the outcome defines two subgroups of the same risk, and no mixture of them can reverse anything.
The rate climbs steeply with the variable’s prognostic value: from 0.13% for the variable correlated 0.10 with the risk to 4.07% for the one correlated 0.95. So the reversals a scanned table finds are disproportionately on the variables a reader is most inclined to believe — age, stage, severity, the prognostic factors every trial reports first — which is part of why a reversal on one of them is so persuasive. The persuasiveness is an artefact of the same property that makes the reversal possible.
This is the region picture of Simpson’s reversal seen from inside a trial. The reversal occupies a region of the space of allocations, and the region is large only when the risk gap between the subgroups is large; randomisation places each variable’s allocation at a random point near the centre of that space, and the prognostic variables are the ones whose regions reach far enough in to be hit. A variable with no prognostic value has a region that does not reach the centre at all. And none of this is the reversal a change of effect measure produces, which needs no imbalance: the risk difference used here is collapsible, so every reversal counted is an allocation effect and nothing else.
It also sets a limit on what stratification can do. A trial can stratify on the one or two strongest prognostic variables and remove the largest single sources of reversal, but the next few in the list are nearly as prognostic and nearly as likely to reverse, and they are the ones left unprotected.
What stratifying on one variable buys
Stratified randomisation fills each arm alternately within the levels of the chosen variable, so the arms are balanced on it and its table cannot reverse. That is exactly what it does and exactly all it does.
With twenty variables tabulated, stratifying on the strongest takes the chance of some reversal from 16.07% to 12.17% — the strongest variable’s own four points, less a little overlap. With a hundred, from 34.13% to 31.93%. Balancing the arms on the best single reading of risk balances them partly on the other readings too, since they are correlated, and that shows in the modest reduction beyond the stratified variable’s own share; but the other variables’ private noise is untouched, and the reversals it drives remain.
A trial could stratify on more. Each stratification variable multiplies the number of strata, and with eighty patients a design stratified on four binary variables has sixteen strata of five patients each, in which the alternating allocation is barely balanced at all. Minimisation handles more variables at the price of predictability, and it too balances only what it is given. Every design that balances on a list leaves the reversal available on whatever is not on the list, and a subgroup table can always be longer than the list.
Larger trials, and the list that stays long
The single-variable reversal falls as trials grow, because a larger trial balances its arms more closely on every variable, and the imbalance that drives the reversal shrinks as the square root of the size. The scanned table falls too, and more slowly than a reader might expect.
With twenty variables, some reversal appears in 17.5% of trials of forty patients, 16.1% of eighty, 15.4% of 160, 12.6% of 320 and 6.2% of 640. The nominated variable’s rate is flat at about four per cent up to 160 patients and falls to 1.6% at 640. The ratio between the two — how many times more often the scanned table reverses than the nominated variable — stays between three and four and a half across the whole range. Growing the trial helps both readings in proportion; it does not make the list less dangerous relative to the single variable.
The flatness at small sizes has the same explanation as in the essay on a single variable. A reversal needs the treatment to be ahead in both subgroups, which a small trial’s noisy subgroup estimates often fail to deliver, and it needs the arms to differ in composition, which a large trial makes rare. The two pull against each other, and between forty and 160 patients they very nearly cancel. The practical range of trial sizes is exactly where the scanned table’s rate is highest.
The size a trial would need to make a scanned reversal rare is large. Halving the rate from its level at eighty takes between four and eight times the patients, and a trial of 640 still shows a reversal on some variable in one trial in sixteen. A trial large enough to make the subgroup table safe to scan is a trial whose subgroups are each large enough to be a trial, and few are.
Reading a subgroup table that reverses
The practical conclusion is not new, but it can now be stated with a number attached. A reversal on some baseline variable in a trial’s subgroup table is expected in one trial in six when twenty variables are reported, with no subgroup effect anywhere. It is not evidence of anything until it has been compared with that rate, and a single trial cannot make the comparison.
What a reader can do is what the essay on a single variable recommended, applied to the whole table: check the arms’ balance on the variable that reverses. Every reversal here is caused by the arms containing different mixtures of that variable’s subgroups, and the baseline table that reports the variable also reports the imbalance. A reversal on a variable that is visibly unbalanced between arms is the arithmetic of that imbalance; on a variable with exactly equal shares in both arms a reversal is arithmetically impossible, since the overall difference is then a weighted average of two positive subgroup differences with the same weights.
The test a reader might reach for, the test of whether the treatment effect differs between subgroups, is the wrong instrument. The subgroup effects are equal here by construction, so the interaction test is aimed at a quantity that is exactly zero, and it is as blind to the reversal as the verdicts are to a real difference: it asks a question the reversal does not raise. The reversal is a statement about the arms’ composition, not about the treatment, and the adjusted estimate — the treatment effect within subgroups, averaged — is what it should be read against.
What a trial report can do about it
None of this makes a subgroup table useless; it makes it a table whose surprises have a known rate. Three reporting habits turn that rate from a trap into a reading aid.
Report every variable tabulated, not only those that show something. The chance of a reversal somewhere depends on the length of the list, and a reader can apply the table above only if the list is printed. A report that mentions the one variable on which the treatment “appeared to fail” and not the nineteen on which it did not has hidden the denominator.
Report the adjusted estimate beside the unadjusted one. The reversal is a disagreement between the whole-trial comparison and the within-subgroup comparisons, and the within-subgroup comparisons, averaged, are what adjustment for the variable produces. In a randomised trial the adjusted estimate for a risk difference is the more precise of the two and is not biased by the imbalance, so it is the number to lead with when the table shows an imbalance on a strong prognostic variable.
Name in advance the variables the analysis will adjust for. Adjusting for whichever variable happens to be imbalanced is itself a search, and the adjusted estimate chosen that way inherits a selection of its own. A list fixed before the data are seen — ideally the strongest prognostic variables, which are also the likeliest to reverse — removes both problems at once, for the variables on it.
The same arithmetic as twenty analyses
Twenty analyses of nothing found 57% of datasets producing a significant result from twenty honest analyses of pure noise. The subgroup table is the same arithmetic in another costume: each variable’s table is an honest analysis, the reversal on any one has a small and correct probability, and the family of twenty has a large one. The difference is that the reversal is not usually tested. It is seen, in a table printed for another purpose, and it is described because it is striking — which is the selection that makes a family-wise rate the relevant one, whether or not anyone computed a p-value.
There is one respect in which the reversal is worse than a false positive in a list of outcomes. A significant outcome in a list of twenty is a claim a reader knows to discount. A reversal in a subgroup table looks like an explanation — the treatment “only seemed to fail” because of the mix of patients — and explanations are discounted less. The rate above is what that explanation should be discounted by.
How often some table reverses, and what one stratified variable removes
On one nominated, strongly prognostic variable, 4.07% of randomised trials of eighty show a Simpson reversal with the treatment helping every patient equally. With twenty variables of mixed prognostic value tabulated, some variable reverses in 16.07% of trials; with a hundred, 34.13%.
Stratifying on the strongest variable removes its reversal and lowers the chance of some reversal to 12.17% for twenty variables and 31.93% for a hundred.
Twenty correlated readings of one risk behave like thirteen independent variables and a hundred like thirty-one, and the per-variable rate rises from 0.13% for a variable correlated 0.10 with the risk to 4.07% for one correlated 0.95.
Every rate is a count over three thousand seeded trials at each setting. The trials share their patients, first variable, allocation and outcomes across every count of variables, so each row of the table scans the same three thousand trials with a longer list, and the rates can only rise down the table.
Not claimed: that real baseline variables are readings of a single risk. A table that holds several unrelated risk dimensions has less overlap between its variables and more effective independence, so the rates above are, if anything, conservative for such a table. Not claimed either that eighty is a typical trial; the essay on a single variable found the reversal rate falling at larger sizes, and at 320 patients some variable among twenty reverses in one trial in eight rather than one in six.
Still open: continuous variables cut at a chosen point
Every variable here is binary, cut at the middle. A continuous baseline variable — age, a biomarker, a score — has to be cut to make a subgroup table, and the cut is chosen. An outcome cut in two measured what cutting costs in information; here cutting adds a choice.
The reversal on a continuous variable depends on where it is cut, and an analyst who tries the median, the tertiles and a clinically conventional threshold is scanning a family within each variable as well as across them. The effective number of chances a continuous baseline table offers, when each variable may be cut in several places, is the product of two effective counts — across variables and across cuts — and the second has not been measured.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A covariate with no levels — both name randomisation, stratification
- A threshold in the tail — both name randomisation, stratification
- Balanced on the wrong function — both name randomisation, stratification
- Balancing more than one number — both name randomisation, stratification
- Conditioning on what the treatment caused — both name simpson's paradox, stratification
- Significant in one, not in the other — both name multiple comparisons, subgroup analysis
Named objects
A flat tag is an object no other essay names yet.
Baseline imbalanceEffective number of testsMultiple comparisonsPrognostic factorRandomisationSimpson's paradoxStratificationSubgroup analysis