When the stratified answer and the pooled one disagree

Every variable, every cut

Slide the cut along every one of twenty baseline variables in a randomised trial of eighty whose treatment helps every patient equally, and some variable shows a Simpson reversal at some cut in 49.0% of trials — a coin toss with nothing to find. The twenty scans do not add like independent chances: if they did they would reverse 92.4% of the time. They read one patient risk between them, and the weak variables seldom reverse at all, so the strongest five supply 41.0 of the 49.0 points. A trial that shows one reversal usually shows several: 4.8 variables on average.

Worth reading first: Simpson's reversal is a region, not a table.

A cut that slides took a trial of eighty patients whose treatment helps every one of them by the same six points, and slid the cut on its strongest baseline variable along every gap between observed values in the middle eighty per cent. Some cut showed a Simpson reversal — both subgroups favouring one arm, the whole trial the other — in 23.4% of trials, against 4.4% at the median and 13.2% at the nine deciles. Sixty-five cut positions were worth about nine and a half independent cuts.

Every scan in that essay ran along one variable, nominated in advance. The analysis it was approaching does not nominate. It takes the baseline table — age, weight, severity, every laboratory value — slides the cut along each, and reports the variable and the cut where the subgroups disagree most strikingly with the whole. That search runs over two dimensions at once, variables and positions, and its count of chances is neither the product of the two nor the larger, because variables that measure the same thing move together. What it amounts to is the subject here.

Twenty variables reading one risk

The trial is the one the earlier essays used. Each patient has a latent risk; the outcome’s probability rises with it; the treatment adds six points for everyone; eighty patients are randomised one to one. Twenty continuous baseline variables are measured, each a noisy reading of the same risk, tracking it with correlations running evenly from 0.95 down to 0.10 — a baseline table with a few strongly prognostic measures, several moderate ones, and a long tail of variables that know almost nothing.

Because they all read one risk, they are correlated with each other in proportion to their own strengths: the two strongest correlate about 0.86, a strong one and a weak one hardly at all. Nothing else connects them. The trial’s draws are exactly the earlier essay’s, patient by patient, so its numbers for the strongest variable are reproduced here rather than quoted.

Half of all trials

How often a trial with nothing to find shows a Simpson reversal somewhere, by how many variables are searched and how. Trials of eighty whose treatment helps every patient by six points, twenty baseline variables tracking the risk with correlations from 0.95 to 0.10; 4,000 trials. Searching the strongest 1, 2, 5, 10, 20: every cut, 23.4%, 30.7%, 41.0%, 46.9%, 49.0%; the deciles, 13.2%, 19.4%, 27.9%, 33.3%, 36.4%; the median, 4.4%, 6.9%, 11.3%, 14.4%, 16.2%.
Fig. 1 How often a trial whose treatment helps every patient equally shows a Simpson reversal somewhere, against how many of its twenty baseline variables are searched — strongest first — at the median, at the nine deciles, and at every cut in the middle eighty per cent.

Searched at its median, the strongest variable reverses in 4.4% of trials; the strongest twenty, each at its median, in 16.2%. Searched at nine deciles each, one variable reverses in 13.2% and twenty in 36.4%. Scanned over every cut, one reverses in 23.4% of trials, two in 30.7%, five in 41.0%, ten in 46.9% and all twenty in 49.0%.

So the full exploratory search, applied to a trial whose treatment helps every patient by exactly the same amount, finds a subgroup reversal in very nearly half of all trials. It is the most common analysis in subgroup reporting run on a trial with nothing to find, and its finding is a coin toss.

How the chances combine

If the twenty scans were independent, the chance that some one of them reverses would be one minus the product of the chances that each does not. Each variable’s own scan reverses between 23.4% of the time for the strongest and 3.4% for the weakest, and combined that way the twenty would reverse 92.4% of the time. They reverse 49.0%. The variables share the risk they read, so a trial whose risk happened to be unevenly split between the arms shows it on every variable that tracks the risk, and those variables’ scans reverse together or not at all.

The effective number of independent tables puts a number on that. Measured at the average table’s rate — one variable cut at one decile, which reverses about 1.1% of the time across the twenty — the scan of one variable is worth 9.5 independent tables, two variables 13.6, five 23.2, ten 34.8 and twenty 59.5. The search tries twenty variables at about sixty-five positions each, thirteen hundred tables, and is worth sixty.

The two dimensions do not multiply. Twenty variables at their medians are worth 12.9 independent medians, and one variable’s scan is worth 9.5 of its own cuts; a product would be more than a hundred, and the full search is worth well under that. Variables and positions overlap in the same way, through the risk: sliding the cut on one strong variable and switching to another strong variable at the same position both re-split the same patients by nearly the same risk.

Which variables carry it

Which of twenty baseline variables shows a reversal, by how closely it tracks the risk, scanned and at the median. 4,000 trials of eighty. Scanned over every cut, from the strongest variable to the weakest: 23.4%, 22.1%, 20.0%, 19.9%, 17.3%, 17.7%, 14.6%, 14.2%, 12.8%, 11.9%, 10.3%, 8.6%, 8.6%, 7.0%, 6.1%, 5.7%, 4.6%, 4.2%, 4.3%, 3.4%; at the median: 4.4%, 3.5%, 3.5%, 2.6%, 1.7%, 1.8%, 1.5%, 1.6%, 1.4%, 0.9%, 1.0%, 0.9%, 0.6%, 0.5%, 0.4%, 0.4%, 0.3%, 0.1%, 0.1%, 0.2%.
Fig. 2 Each of the twenty baseline variables’ own chance of showing a reversal at some cut and at its median, against how closely it tracks the patients’ risk.

The reversals come from the prognostic variables. The strongest, tracking the risk at 0.95, reverses somewhere in 23.4% of trials; the variable at 0.50 in 10.3%; the weakest, at 0.10, in 3.4%. At the median the same three reverse in 4.4%, 1.0% and 0.2%. A variable that knows nothing about the outcome cannot make the subgroups disagree with the whole, because splitting on it leaves each subgroup a random half of the trial; a variable that tracks the risk can, because splitting on it can put an unlucky excess of high-risk patients in one arm on each side.

That is why the curve in the first figure flattens. The strongest five variables take the search from nothing to 41.0%; the next fifteen add eight points between them. A baseline table’s long tail of weakly prognostic variables adds tables to the report and very little to the chance of a reversal, while its few strong variables carry nearly all of it. Twenty variables nobody stratified on found the same concentration with every variable split once; scanning sharpens it, since a strong variable’s scan finds the uneven split wherever it lies.

One trial, searched

Where a scan of every cut on twenty variables finds a Simpson reversal, in one trial with nothing to find. One seeded trial of eighty; the whole trial's difference is −2.5 points. 8 of the twenty variables reverse at some cut: variable 1 (correlation 0.95 with the risk) at 11 cuts, variable 2 (correlation 0.91 with the risk) at 2 cuts, variable 7 (correlation 0.68 with the risk) at 1 cut, variable 8 (correlation 0.64 with the risk) at 2 cuts, variable 10 (correlation 0.55 with the risk) at 6 cuts, variable 12 (correlation 0.46 with the risk) at 1 cut, variable 13 (correlation 0.41 with the risk) at 1 cut, variable 15 (correlation 0.32 with the risk) at 6 cuts.
Fig. 3 One trial of eighty with nothing to find, its twenty variables scanned: each row marks the cut positions, as a share of patients below the cut, at which both subgroups favour one arm and the whole trial the other.

The trial drawn above is one of the half that reverse. Its whole-trial difference is −2.5 points — an unlucky draw for a treatment that helps everyone by six — and eight of its twenty variables show a reversal somewhere: the strongest at eleven cut positions, the tenth at six, the fifteenth, tracking the risk at only 0.32, at six. A report of this trial could pick any of them, and the one it picked would be the most striking of eight, with a story to fit.

That is typical. Among trials that reverse anywhere, 4.8 variables reverse on average. A reversal found by the full search almost never comes alone, and the variable reported is chosen from among several by the analyst, which is a further search on top of the one being counted.

A reversal is a symptom of the whole

The reversals in these trials are not spread evenly over trials. A trial whose treatment helps everyone by six points still comes out with a negative whole-trial difference 30.8% of the time at eighty patients, and those trials are where the search finds most of its reversals: some variable reverses somewhere in 72.6% of them, against 39.4% of the trials whose whole difference came out positive. Across all trials the whole difference averages 5.6 points; across the trials with a reversal somewhere it averages 0.6.

The reason is the size of the trial against its effect. About 46% of control patients have the outcome, so the whole-trial difference between forty patients and forty has a standard error of about eleven points, and the six-point effect is about half a standard error. A trial of eighty cannot tell a six-point effect from none, and three trials in ten put it on the wrong side of zero. The subgroup search is not what makes these trials ambiguous; it is what makes their ambiguity look like a finding.

So a reversal found by the full search is, more than anything else, a report that the whole-trial estimate happened to land near zero. When the whole is near zero, a small uneven split of the risk on either side of almost any cut can push the two subgroup estimates the other way, and with twenty variables and sixty-five cuts each there is almost always a cut that does. The subgroups are not disagreeing with a clear overall result; they are fluctuating around a weak one, and the reversal is the fluctuation that happened to point both ways at once.

That reverses the way such a finding is usually read. A reported reversal invites the reading that the overall effect hides two opposite subgroup effects. In these trials it much more often marks an overall effect that came out too small by chance, with two subgroups that are both closer to the truth than the whole is — a reader who took either subgroup’s estimate over the whole’s would usually be nearer the six points every patient receives.

One reversal, or several

The distribution of how many variables reverse in a trial that has any is long. Exactly one variable reverses in 23.1% of such trials, two in 14.5%, three in 10.1%, and six or more in more than a third of them; the average is 4.8. A report that names a single reversing variable is, three times in four, a report of one variable chosen from several that would have served.

Restricting the cuts does not help much

A protocol might forbid cuts that leave few patients on one side. Requiring at least a fifth of the patients on each side of every cut, rather than a tenth, takes the full search from 49.0% to 47.2%; requiring three tenths, to 43.6%. The search is then worth 56.3 and 50.5 independent tables rather than 59.5. Most of the reversing cuts are in the middle of the variables already, where the subgroups are large enough to carry a stable estimate and still small enough to fluctuate. Trimming the range removes the edges of the search, and its chances were never mostly at the edges.

A larger effect, and none

The reversal rate depends on how far the whole effect sits from zero in units of its own noise. With no effect at all — the treatment helping nobody — the full search at eighty reverses in 56.3% of trials; with the six-point effect, 49.0%; with twelve points, 34.8%. Even a treatment that helps every patient by twelve points, a large effect for a trial of eighty, shows a reversal on some variable at some cut in a third of trials. The effect moves the rate less than the search does: the strongest variable at its median reverses 4.4% of the time at six points, and the full search at twelve points still reverses eight times as often.

A larger trial

The full search at two trial sizes: how often it finds a reversal, and how many independent tables it is worth. Every cut in the middle 80% of the strongest 1, 2, 5, 10, 20 variables. Eighty patients: 23.4%, worth 9.5 tables; 30.7%, worth 13.6 tables; 41.0%, worth 23.2 tables; 46.9%, worth 34.8 tables; 49.0%, worth 59.5 tables. 640 patients: 11.8%, worth 7.1; 14.5%, worth 8.8; 17.9%, worth 13.4; 20.5%, worth 21.5; 20.9%, worth 37.4.
Fig. 4 The full search — every cut on the strongest one, two, five, ten and twenty variables — at eighty and at 640 patients, with the effective number of independent tables each search is worth.

With eight times the patients and the same six-point effect, the whole trial’s difference is much less likely to come out negative, and a reversal needs both subgroups to come out against the effect while the whole does not. The full search reverses in 20.9% of trials of 640, against 49.0% at eighty. The strongest variable alone, scanned, reverses in 11.8%.

But the search’s count of chances barely shrinks. At 640 patients the strongest variable’s scan is worth 7.1 independent tables and the full search 37.4, against 9.5 and 59.5 at eighty. A cut that slides found the single scan’s count set by the range scanned rather than by the number of positions; the full search’s count is set by how many distinct ways of re-splitting the risk the baseline table offers, and a larger trial does not change the table. What a larger trial changes is the rate of each table, not how many tables the search is.

What a reader of a reported reversal can do with this

Treat a reversal from an unrestricted search as uninformative in a small trial. Half of all trials of eighty with a constant effect produce one. A paper reporting “the treatment helped overall but harmed patients on both sides of a threshold in the third baseline variable” has reported something that a trial with no heterogeneity at all would show by chance about as often as not.

Ask how many variables were searched, and how. At the median, twenty variables reverse 16.2% of the time; at nine fixed deciles, 36.4%; scanned, 49.0%. The same claim is three different strengths of evidence, and the only difference is a sentence in the methods.

Look at the variable’s prognostic strength. A reversal on a variable that tracks the outcome strongly is the expected kind of chance reversal, not a sign of a modifier: strong variables are the ones that reverse by chance. Five places to cut one variable found the same thing from the cuts’ side.

Read the whole estimate’s distance from zero first. Where the whole-trial difference is within a standard error or so of zero, a reversal somewhere is close to certain under a constant effect — 72.6% of trials whose whole came out negative showed one — and it says more about the whole than about any subgroup. A reversal beside a whole estimate several standard errors from zero is rarer and worth more attention, though at eighty patients such whole estimates are themselves rare.

Check how many variables reverse. A modifier would make one variable and its correlates reverse; chance makes several unrelated ones reverse together, because they all re-split the same uneven risk. A report of a single reversing variable from a full search is a report that left the others out.

How this extends the earlier counts

The reversal a coin cannot prevent found the reversal in a single nominated table and showed randomisation does not remove it; twenty variables nobody stratified on multiplied the tables by variables, five places to cut one variable by cuts, and a cut that slides by positions. Each multiplication added chances more slowly than its count, and the reason was the same each time: every table re-splits one underlying risk, and re-splitting a risk in a nearby way is not a new chance. The full search is the end of that sequence. Its thirteen hundred tables are worth sixty, and sixty chances at a table that reverses about one time in ninety make half of all trials.

Counted, and how

Every rate is over four thousand trials of eighty — fifteen hundred of 640 — with the treatment adding six points for every patient, the outcome’s probability a normal function of the latent risk, and twenty baseline variables tracking that risk at correlations from 0.95 to 0.10. Cuts are slid along every gap between observed values with at least a tenth of the patients on each side; the medians and deciles are those of the variable’s own distribution. The effective number of tables is log⁡(1−pany)/log⁡(1−pˉ)\log(1 - p_{\text{any}})/\log(1 - \bar p) with pˉ\bar p the average single-table rate, the convention the earlier essays used.

Not claimed: that a baseline table looks like this one. A table with many strongly prognostic variables reverses more often than one with a single strong variable and many weak ones, and the order of the variables here — strongest first — is the order a search would pick them in, which makes the curve as steep at the start as it can be. And not claimed that any of these reversals is real: every one is a trial in which the treatment helps every patient equally.

Still open: a reference distribution for the search a trial actually ran

The 49.0% is a property of this trial design and this baseline table, and a reader of a real trial has neither. What they could have is a reference built from the trial itself: re-draw the outcomes under the hypothesis that the effect is the same for everyone — from the trial’s own estimated risks and its estimated common effect — and run the full search on each re-draw. The share of re-draws that reverse somewhere is the chance the reported reversal would have turned up anyway.

Whether that reference can be built from eighty patients when the risks themselves must be estimated from the same baseline variables being searched, how close its rate comes to the true 49.0%, and whether a reversal can ever be rare under it in a trial this size, are measurable on the same trials and have not been measured here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Baseline imbalanceEffective number of testsMedian splitMultiple comparisonsRandomisationSimpson's paradoxSpecification searchSubgroup analysis