A cut that slides
Worth reading first: Simpson's reversal is a region, not a table.
Five places to cut one variable counted how often a randomised trial’s subgroup table shows a Simpson reversal when its strongest continuous baseline variable is cut at a few chosen points. In a trial of eighty where the treatment helps every patient by the same six points, the median cut reverses in 4.4% of trials, some one of five candidate cuts in 10.8%, and some one of the nine deciles in 13.2% — nine correlated cuts worth about five independent chances.
Every cut there was fixed in advance, a family of tables an analyst might try. The essay ended on the version that matters more: the cut chosen by looking, slid along the variable until the subgroups differ most, which turns a family of nine into a continuum and a maximum over it. It named the quantity that should govern the count — how quickly the table changes as the cut moves — and noted that a sliding cut is the same object as a break that was looked for. Both turn out to be the right frame, and the numbers they give are not the ones a count of cut positions would suggest.
The scan
The trial is the one the earlier essays used, drawn on the same random numbers: each patient has an underlying risk, the treatment adds six percentage points to everyone’s chance of a good outcome, allocation is by coin, and the strongest baseline variable measures the risk with a correlation of 0.95. Kept continuous, the variable can be cut anywhere. The scan sorts the patients by it and moves the cut through every gap between consecutive values, from the point with a tenth of the patients below it to the point with a tenth above, and at each position asks the question the tables asked: do the two subgroups agree in sign and the whole trial disagree with both?
The median and the nine deciles are among the positions scanned, and the scan draws the same trials, so its count can only be higher. At eighty patients it is 23.4%: nearly one trial in four shows a reversal somewhere along its strongest variable, against 13.2% at some decile and 4.4% at the median. An analyst who slides the cut until the table looks interesting will find a reversal in about five times as many trials as one who cut at the median before looking.
A scan is not one chance per cut
At eighty patients the middle eighty per cent holds sixty-five cut positions. If they were independent chances at the median’s rate of 4.4% the scan would reverse in 94.6% of trials, nearly every one. It reverses in under a quarter, because moving the cut by one gap moves one patient from one subgroup to the other, and a table that changes by one patient is almost the same table.
Measured the way the earlier essays measured it — the number of independent cuts at the average decile’s reversal rate that would give the same chance of some reversal — the scan is worth 9.5 independent cuts at eighty patients, while the nine deciles are worth 5.1. At forty patients the scan is worth 9.0; at 160, 8.5; at 320, 7.5; at 640, 7.2 — while the number of positions it tries grows from 33 to 513.
The count does not grow with the positions because the positions are not what varies. As the trial grows, the gaps get finer, but the table at the thirtieth percentile and the table at the thirty-first stay nearly the same table whatever the trial’s size, and how quickly the table forgets itself as the cut moves is a property of the variable’s range, measured in quantiles, not of how many patients sit in each gap. A scan across the middle eighty per cent of a variable is worth something like seven to ten independent looks, at any size, because that is how many nearly separate tables the range holds.
This is the arithmetic a break that was looked for met in time rather than along a covariate. A change point searched for across every row of a series is a supremum over a process, and the charge for the search depends on the range searched and how smoothly the profile moves, not on the number of rows. A second break on a flat profile found the search for one change point in a hundred and twenty rows manufacturing five units of likelihood. The subgroup table’s sliding cut is the same search with a Simpson reversal in place of a likelihood.
One trial, scanned
The trial drawn above is one the fixed tables would have missed. The whole trial’s difference is −0.079: the treatment appears to harm. At no decile do both subgroups disagree with it. But with the cut at 28.75% or 30.00% of the patients below it, both subgroups show the treatment helping — a narrow window of two positions, a reversal an analyst sliding the cut would find and report as “in patients below the 29th percentile and above it alike, the treatment is beneficial”.
The picture also shows why the scan is worth only about ten chances. The two subgroup curves wander as the cut moves, and they wander slowly: each step changes one patient’s side, and the curves’ excursions last tens of positions. The shaded window is where both happen to lie above zero at the same moment, which a slowly moving pair of curves does rarely and a table of sixty-five independent pairs would do often.
The size of the trial
A reversal needs chance to overcome the treatment’s real effect, so it gets rarer as the trial grows and the effect becomes clearer. The scan follows that, with a rise first: some cut reverses in 18.8% of trials of forty, 23.4% of eighty, 24.3% of 160, 19.2% of 320 and 10.6% of 640. The median alone falls from 3.9% to 1.7% over the same range, and the deciles from 12.2% to 6.2%.
The ratio moves the other way. The scan finds a reversal 4.8 times as often as the median at forty patients, 5.3 at eighty, 5.7 at 160, 6.1 at 320 and 6.4 at 640. As a reversal at any one cut becomes rarer, more of the reversals that remain are found only by looking along the variable for them. In a large trial a subgroup reversal on a continuous variable is increasingly a product of where the cut was put.
Where the reversing cuts fall
A natural suspicion is that a scan’s reversals sit at its ends, where one subgroup is small and noisy, and that trimming the range would remove them. They do not sit there.
At eighty patients the reversing positions are spread across the whole scanned range, heaviest in the middle. That agrees with what the five-cut essay found for fixed cuts, where the median reversed more often than the first or ninth decile: the middle of the variable is where both subgroups are large enough to point clearly in one direction, and so where the two can agree against the whole. The scan does not change where reversals are likely; it changes how many of the likely places are tried. In only 5.3% of the trials whose scan reverses does every reversing position lie outside the middle sixty per cent. At 640 patients the spread is flatter and the outer share 10.4%. The reversal is a property of the whole table’s imbalance, and the whole table is out of balance at many cuts at once.
Nor does the range’s edge matter much. Widening the scan from the middle eighty to the middle ninety per cent leaves the rate at 23.4%; narrowing it to the middle sixty per cent lowers it to 22.1%, and to the middle forty per cent, 19.3%, with the effective count falling from 9.5 to 8.9 and 7.7. The positions near the extremes add almost nothing, because a subgroup of a tenth of eighty patients rarely has both arms well enough represented to point anywhere reliably, and the ones in the middle carry the scan.
A weaker variable, scanned
Everything so far is the strongest variable, the one whose table reverses most often because its subgroups differ most in risk. A baseline table is mostly weaker variables, and twenty variables nobody stratified on found that a variable’s chance of reversing falls with how well it tracks the risk. The scan changes how fast it falls.
At a correlation of 0.8 with the risk, the median cut reverses in 2.6% of trials and the scan in 18.3%. At 0.6, 1.3% and 14.0%. At 0.4, 0.7% and 8.5%. At 0.2 — a variable that barely tracks the risk at all — the median reverses in 0.3% of trials and the scan in 4.5%, about eighteen times as often. The deciles sit between: 10.1%, 7.2%, 3.8% and 1.7%.
So the scan is worth most on the variables where a fixed cut would show almost nothing. A reversal at the median of a weak variable is a rare event, and the scan, which gets to try every table the variable can make, finds the rare configurations in which its imbalance happens to line up. The effective count grows accordingly — from 9.5 independent cuts on the strongest variable to 14.4 at a correlation of 0.4 and 18.9 at 0.2 — which is the rare-event sensitivity of the convention showing itself, and also a warning in its own right: the less a variable should be able to produce a reversal, the more a reported reversal on it owes to the search.
For a reader, the upshot is that a reversal on a weak baseline variable — a characteristic with little reason to be related to the outcome — found at a cut nobody named in advance is almost entirely the product of the scan. Simpson’s reversal is usually told as a story about a strong confounder; the version a trial’s subgroup table produces by scanning is more often a story about a weak one, searched.
A threshold found by sliding is one of many
The reversing positions come in runs, and the runs multiply as the trial grows. At eighty patients a trial whose scan reverses does so in 2.62 separate runs of positions on average; at 640, in 8.36. A large trial scanned along its strongest variable does not show one threshold where the subgroups disagree with the whole; it shows several narrow windows, scattered along the variable, each of which could be reported as “the” cut-off.
That matters for how a scanned result is reported. A paper that names one threshold — “in patients whose score is below 42, and in those above it” — reports one window of several, chosen because it was found, and the width of the window, which is often a handful of patients, is not stated. The cut that fitted best found a baseline cut chosen for fit inflating the precision it reports; a cut chosen for a reversal is chosen for the most surprising table, and the precision it reports is that of a table nobody would have drawn in advance.
What the scan’s count is good for
The effective count gives a scanned reversal its discount, and the discount is modest in a way that is worth knowing. A reversal found by sliding the cut on one variable is about as surprising as one found among nine or ten independent pre-specified tables; not sixty-five, and not one. A scan over twenty variables of the kind twenty variables nobody stratified on and the five-cut essay counted would multiply that by a count of variables, and the five-cut essay found that product of families real: about thirty-one independent chances where a table of twenty variables was cut at five points each.
The honest reference for a scanned reversal is the scan itself, simulated: draw trials with the effect the same for everyone, slide the cut the same way, and read how often some position reverses. That is what the reversal a coin cannot prevent did for one fixed table, and the only change a scan requires is to take the maximum over positions inside each simulated trial. It is the same repair the break-point essays applied to a search over rows, and it is available wherever the scan can be written down.
What naming the cut in advance buys
The comparison can be read as the value of pre-specification. At eighty patients, a scan reverses in 23.4% of trials and the median alone in 4.4%, so fixing the median before looking removes 81% of the reversals the scan would have reported, all of them false. Fixing the nine deciles instead removes 43%. At 640 patients the median removes 84% and the deciles 42%: the share a single pre-specified cut removes rises with the trial’s size, because the scan’s advantage does.
None of this costs the trial anything it needs. The treatment effect is estimated from the whole trial, and a pre-specified cut on a continuous variable is only ever a way of drawing one table; the table drawn at the median says everything a subgroup table can honestly say about the variable, and the sixty-four others the scan adds say nothing new except by chance.
What a report of a cut on a continuous variable should state
How the cut was chosen. A median fixed in advance reverses in 4.4% of trials of eighty with nothing there; the same variable scanned, in 23.4%. The two are different claims and the report should say which it makes.
The range that was scanned. It, rather than the number of positions, sets the scan’s count: the middle eighty per cent of a variable is worth about nine and a half independent cuts at eighty patients, the middle forty per cent about seven and a half.
How many windows reverse, and how wide the one reported is. At 640 patients a reversing scan passes through eight separate runs on average. One of them reported as a threshold is a choice among them.
A reference simulated with the scan in it. Nothing else gives the scanned reversal its right rarity, and how many analyses there really were found that an analyst’s own count of what was tried is usually too low; a scan, at least, can be re-run exactly.
What is counted here
Counted over four thousand trials at each size, the effect the same for every patient, on the draws the fixed-cut essays used: the scan of the middle 80% reverses in 23.4% of trials of eighty, against 13.2% for the nine deciles and 4.4% for the median; it is worth 9.5 independent cuts at eighty patients and 7.2 at 640; and its reversing positions form 2.62 runs at eighty and 8.36 at 640.
Not claimed: that the effective count is exactly constant in the limit. It falls from 9.5 to 7.2 over the sizes measured, partly because reversals become rare events at large sizes, where the effective-count convention is sensitive to the rate it is measured at; the claim is only that it does not grow with the number of positions, which rise sixteenfold over the same range.
Still open: a cut chosen to look, on a variable chosen to look
Every scan here runs along one variable, the strongest, nominated in advance. The analysis these essays have been approaching slides the cut along every continuous variable in the baseline table and reports the variable and the cut where the table looks most striking. That is a scan over a two-dimensional family — variables and positions — whose count is neither the product of the two nor the larger of them, because variables that measure the same thing have scans that move together.
How often some variable of twenty, scanned, shows a reversal in a trial with nothing to find; how the count of variables and the count within each scan combine when the variables are correlated; and whether a simulated reference for the full search is practical for a trial’s own baseline table, are the measurements that would price the most common exploratory analysis in subgroup reporting. None of them has been made here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- One characteristic that really matters — both name effective number of tests, multiple comparisons, subgroup analysis
- Sixteen subgroups and one effect — both name effective number of tests, multiple comparisons, subgroup analysis
- A baseline cut in two — both name dichotomisation, randomisation
- An outcome cut in two — both name dichotomisation, the garden of forking paths
- Naming a handful in advance — both name the garden of forking paths, multiple comparisons
- Significant in one, not in the other — both name multiple comparisons, subgroup analysis
Named objects
A flat tag is an object no other essay names yet.
DichotomisationEffective number of testsThe garden of forking pathsMultiple comparisonsRandomisationSimpson's paradoxSubgroup analysisSupremum statistic