Every variable, every cut
Worth reading first: Simpson's reversal is a region, not a table.
A cut that slides took a trial of eighty patients whose treatment helps every one of them by the same six points, and slid the cut on its strongest baseline variable along every gap between observed values in the middle eighty per cent. Some cut showed a Simpson reversal — both subgroups favouring one arm, the whole trial the other — in 23.4% of trials, against 4.4% at the median and 13.2% at the nine deciles. Sixty-five cut positions were worth about nine and a half independent cuts.
Every scan in that essay ran along one variable, nominated in advance. The analysis it was approaching does not nominate. It takes the baseline table — age, weight, severity, every laboratory value — slides the cut along each, and reports the variable and the cut where the subgroups disagree most strikingly with the whole. That search runs over two dimensions at once, variables and positions, and its count of chances is neither the product of the two nor the larger, because variables that measure the same thing move together. What it amounts to is the subject here.
Twenty variables reading one risk
The trial is the one the earlier essays used. Each patient has a latent risk; the outcome’s probability rises with it; the treatment adds six points for everyone; eighty patients are randomised one to one. Twenty continuous baseline variables are measured, each a noisy reading of the same risk, tracking it with correlations running evenly from 0.95 down to 0.10 — a baseline table with a few strongly prognostic measures, several moderate ones, and a long tail of variables that know almost nothing.
Because they all read one risk, they are correlated with each other in proportion to their own strengths: the two strongest correlate about 0.86, a strong one and a weak one hardly at all. Nothing else connects them. The trial’s draws are exactly the earlier essay’s, patient by patient, so its numbers for the strongest variable are reproduced here rather than quoted.
Half of all trials
Searched at its median, the strongest variable reverses in 4.4% of trials; the strongest twenty, each at its median, in 16.2%. Searched at nine deciles each, one variable reverses in 13.2% and twenty in 36.4%. Scanned over every cut, one reverses in 23.4% of trials, two in 30.7%, five in 41.0%, ten in 46.9% and all twenty in 49.0%.
So the full exploratory search, applied to a trial whose treatment helps every patient by exactly the same amount, finds a subgroup reversal in very nearly half of all trials. It is the most common analysis in subgroup reporting run on a trial with nothing to find, and its finding is a coin toss.
How the chances combine
If the twenty scans were independent, the chance that some one of them reverses would be one minus the product of the chances that each does not. Each variable’s own scan reverses between 23.4% of the time for the strongest and 3.4% for the weakest, and combined that way the twenty would reverse 92.4% of the time. They reverse 49.0%. The variables share the risk they read, so a trial whose risk happened to be unevenly split between the arms shows it on every variable that tracks the risk, and those variables’ scans reverse together or not at all.
The effective number of independent tables puts a number on that. Measured at the average table’s rate — one variable cut at one decile, which reverses about 1.1% of the time across the twenty — the scan of one variable is worth 9.5 independent tables, two variables 13.6, five 23.2, ten 34.8 and twenty 59.5. The search tries twenty variables at about sixty-five positions each, thirteen hundred tables, and is worth sixty.
The two dimensions do not multiply. Twenty variables at their medians are worth 12.9 independent medians, and one variable’s scan is worth 9.5 of its own cuts; a product would be more than a hundred, and the full search is worth well under that. Variables and positions overlap in the same way, through the risk: sliding the cut on one strong variable and switching to another strong variable at the same position both re-split the same patients by nearly the same risk.
Which variables carry it
The reversals come from the prognostic variables. The strongest, tracking the risk at 0.95, reverses somewhere in 23.4% of trials; the variable at 0.50 in 10.3%; the weakest, at 0.10, in 3.4%. At the median the same three reverse in 4.4%, 1.0% and 0.2%. A variable that knows nothing about the outcome cannot make the subgroups disagree with the whole, because splitting on it leaves each subgroup a random half of the trial; a variable that tracks the risk can, because splitting on it can put an unlucky excess of high-risk patients in one arm on each side.
That is why the curve in the first figure flattens. The strongest five variables take the search from nothing to 41.0%; the next fifteen add eight points between them. A baseline table’s long tail of weakly prognostic variables adds tables to the report and very little to the chance of a reversal, while its few strong variables carry nearly all of it. Twenty variables nobody stratified on found the same concentration with every variable split once; scanning sharpens it, since a strong variable’s scan finds the uneven split wherever it lies.
One trial, searched
The trial drawn above is one of the half that reverse. Its whole-trial difference is −2.5 points — an unlucky draw for a treatment that helps everyone by six — and eight of its twenty variables show a reversal somewhere: the strongest at eleven cut positions, the tenth at six, the fifteenth, tracking the risk at only 0.32, at six. A report of this trial could pick any of them, and the one it picked would be the most striking of eight, with a story to fit.
That is typical. Among trials that reverse anywhere, 4.8 variables reverse on average. A reversal found by the full search almost never comes alone, and the variable reported is chosen from among several by the analyst, which is a further search on top of the one being counted.
A reversal is a symptom of the whole
The reversals in these trials are not spread evenly over trials. A trial whose treatment helps everyone by six points still comes out with a negative whole-trial difference 30.8% of the time at eighty patients, and those trials are where the search finds most of its reversals: some variable reverses somewhere in 72.6% of them, against 39.4% of the trials whose whole difference came out positive. Across all trials the whole difference averages 5.6 points; across the trials with a reversal somewhere it averages 0.6.
The reason is the size of the trial against its effect. About 46% of control patients have the outcome, so the whole-trial difference between forty patients and forty has a standard error of about eleven points, and the six-point effect is about half a standard error. A trial of eighty cannot tell a six-point effect from none, and three trials in ten put it on the wrong side of zero. The subgroup search is not what makes these trials ambiguous; it is what makes their ambiguity look like a finding.
So a reversal found by the full search is, more than anything else, a report that the whole-trial estimate happened to land near zero. When the whole is near zero, a small uneven split of the risk on either side of almost any cut can push the two subgroup estimates the other way, and with twenty variables and sixty-five cuts each there is almost always a cut that does. The subgroups are not disagreeing with a clear overall result; they are fluctuating around a weak one, and the reversal is the fluctuation that happened to point both ways at once.
That reverses the way such a finding is usually read. A reported reversal invites the reading that the overall effect hides two opposite subgroup effects. In these trials it much more often marks an overall effect that came out too small by chance, with two subgroups that are both closer to the truth than the whole is — a reader who took either subgroup’s estimate over the whole’s would usually be nearer the six points every patient receives.
One reversal, or several
The distribution of how many variables reverse in a trial that has any is long. Exactly one variable reverses in 23.1% of such trials, two in 14.5%, three in 10.1%, and six or more in more than a third of them; the average is 4.8. A report that names a single reversing variable is, three times in four, a report of one variable chosen from several that would have served.
Restricting the cuts does not help much
A protocol might forbid cuts that leave few patients on one side. Requiring at least a fifth of the patients on each side of every cut, rather than a tenth, takes the full search from 49.0% to 47.2%; requiring three tenths, to 43.6%. The search is then worth 56.3 and 50.5 independent tables rather than 59.5. Most of the reversing cuts are in the middle of the variables already, where the subgroups are large enough to carry a stable estimate and still small enough to fluctuate. Trimming the range removes the edges of the search, and its chances were never mostly at the edges.
A larger effect, and none
The reversal rate depends on how far the whole effect sits from zero in units of its own noise. With no effect at all — the treatment helping nobody — the full search at eighty reverses in 56.3% of trials; with the six-point effect, 49.0%; with twelve points, 34.8%. Even a treatment that helps every patient by twelve points, a large effect for a trial of eighty, shows a reversal on some variable at some cut in a third of trials. The effect moves the rate less than the search does: the strongest variable at its median reverses 4.4% of the time at six points, and the full search at twelve points still reverses eight times as often.
A larger trial
With eight times the patients and the same six-point effect, the whole trial’s difference is much less likely to come out negative, and a reversal needs both subgroups to come out against the effect while the whole does not. The full search reverses in 20.9% of trials of 640, against 49.0% at eighty. The strongest variable alone, scanned, reverses in 11.8%.
But the search’s count of chances barely shrinks. At 640 patients the strongest variable’s scan is worth 7.1 independent tables and the full search 37.4, against 9.5 and 59.5 at eighty. A cut that slides found the single scan’s count set by the range scanned rather than by the number of positions; the full search’s count is set by how many distinct ways of re-splitting the risk the baseline table offers, and a larger trial does not change the table. What a larger trial changes is the rate of each table, not how many tables the search is.
What a reader of a reported reversal can do with this
Treat a reversal from an unrestricted search as uninformative in a small trial. Half of all trials of eighty with a constant effect produce one. A paper reporting “the treatment helped overall but harmed patients on both sides of a threshold in the third baseline variable” has reported something that a trial with no heterogeneity at all would show by chance about as often as not.
Ask how many variables were searched, and how. At the median, twenty variables reverse 16.2% of the time; at nine fixed deciles, 36.4%; scanned, 49.0%. The same claim is three different strengths of evidence, and the only difference is a sentence in the methods.
Look at the variable’s prognostic strength. A reversal on a variable that tracks the outcome strongly is the expected kind of chance reversal, not a sign of a modifier: strong variables are the ones that reverse by chance. Five places to cut one variable found the same thing from the cuts’ side.
Read the whole estimate’s distance from zero first. Where the whole-trial difference is within a standard error or so of zero, a reversal somewhere is close to certain under a constant effect — 72.6% of trials whose whole came out negative showed one — and it says more about the whole than about any subgroup. A reversal beside a whole estimate several standard errors from zero is rarer and worth more attention, though at eighty patients such whole estimates are themselves rare.
Check how many variables reverse. A modifier would make one variable and its correlates reverse; chance makes several unrelated ones reverse together, because they all re-split the same uneven risk. A report of a single reversing variable from a full search is a report that left the others out.
How this extends the earlier counts
The reversal a coin cannot prevent found the reversal in a single nominated table and showed randomisation does not remove it; twenty variables nobody stratified on multiplied the tables by variables, five places to cut one variable by cuts, and a cut that slides by positions. Each multiplication added chances more slowly than its count, and the reason was the same each time: every table re-splits one underlying risk, and re-splitting a risk in a nearby way is not a new chance. The full search is the end of that sequence. Its thirteen hundred tables are worth sixty, and sixty chances at a table that reverses about one time in ninety make half of all trials.
Counted, and how
Every rate is over four thousand trials of eighty — fifteen hundred of 640 — with the treatment adding six points for every patient, the outcome’s probability a normal function of the latent risk, and twenty baseline variables tracking that risk at correlations from 0.95 to 0.10. Cuts are slid along every gap between observed values with at least a tenth of the patients on each side; the medians and deciles are those of the variable’s own distribution. The effective number of tables is with the average single-table rate, the convention the earlier essays used.
Not claimed: that a baseline table looks like this one. A table with many strongly prognostic variables reverses more often than one with a single strong variable and many weak ones, and the order of the variables here — strongest first — is the order a search would pick them in, which makes the curve as steep at the start as it can be. And not claimed that any of these reversals is real: every one is a trial in which the treatment helps every patient equally.
Still open: a reference distribution for the search a trial actually ran
The 49.0% is a property of this trial design and this baseline table, and a reader of a real trial has neither. What they could have is a reference built from the trial itself: re-draw the outcomes under the hypothesis that the effect is the same for everyone — from the trial’s own estimated risks and its estimated common effect — and run the full search on each re-draw. The share of re-draws that reverse somewhere is the chance the reported reversal would have turned up anyway.
Whether that reference can be built from eighty patients when the risks themselves must be estimated from the same baseline variables being searched, how close its rate comes to the true 49.0%, and whether a reversal can ever be rare under it in a trial this size, are measurable on the same trials and have not been measured here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- One characteristic that really matters — both name effective number of tests, multiple comparisons, subgroup analysis
- Sixteen subgroups and one effect — both name effective number of tests, multiple comparisons, subgroup analysis
- A contrast chosen after the data — both name multiple comparisons, specification search
- Significant in one, not in the other — both name multiple comparisons, subgroup analysis
- The cleanest of the significant — both name multiple comparisons, specification search
- The pattern a cause leaves on its neighbours — both name multiple comparisons, subgroup analysis
Named objects
A flat tag is an object no other essay names yet.
Baseline imbalanceEffective number of testsMedian splitMultiple comparisonsRandomisationSimpson's paradoxSpecification searchSubgroup analysis