When the stratified answer and the pooled one disagree

Five places to cut one variable

A continuous baseline variable has to be cut before it makes a subgroup table, and the cut is chosen. In a randomised trial of eighty where the treatment helps every patient equally, the strongest baseline variable cut at its median shows a Simpson reversal in 4.4% of trials; tried at five candidate cuts, some cut reverses in 10.8%, and at nine, 13.2%. Nine correlated cuts are worth about five independent chances, not nine. Across twenty continuous variables cut at five points each, some table reverses in 32.2% of trials — twice the rate with each variable cut once, and about thirty-one independent chances where the table holds a hundred.

Worth reading first: Simpson's reversal is a region, not a table.

Twenty variables nobody stratified on counted how often a randomised trial whose treatment helps every patient shows a Simpson reversal on some baseline variable: 4.07% on one nominated variable at eighty patients, 16.07% on some variable of twenty, 34.13% on some variable of a hundred. Every variable there was binary — above or below its median. Most baseline variables are not. Age, a biomarker, a severity score, a blood pressure are continuous, and to appear in a subgroup table they have to be cut, somewhere, by someone.

The essay ended on what that choice adds. An analyst who tries the median, the tertiles and a conventional threshold is scanning a family of tables within each variable as well as across variables, and the effective number of chances a baseline table offers is then a product of two counts, the second of which had not been measured.

How often a trial of 80 shows a Simpson reversal on its strongest baseline variable, against how many cut points are triedThe treatment helps every patient equally. Cut at its median, the variable's table reverses in 4.4% of trials; trying 2, 3, 5, 9 cuts, in 6.4%, 8.7%, 10.8%, 13.2%. Independent chances would give 33.3% at nine.00.1000.2000.3000.40012359candidate cut points tried on one continuous variabletrials whose table reverses at some cutsome cut reversesas many independent chances4,000 trials of 80, cuts nested around the medianone variable is a family of tables
Fig. 1 How often a randomised trial of eighty, whose treatment helps every patient equally, shows a Simpson reversal on its strongest baseline variable at some cut, against how many candidate cut points are tried — from the median alone to all nine deciles — beside what the same number of independent chances would give. The slider sets the trial’s size.

The trial, and what a reversal is here

The trial is the one the twenty-variable essay used. Each patient has an underlying risk; the treatment raises every patient’s chance of a good outcome by six percentage points; allocation is by coin. The baseline variables are noisy measurements of the underlying risk, from one that tracks it closely to ones that barely track it at all. Here they are kept continuous, and a subgroup table is made by cutting a variable at a point fixed in advance — a quantile of its distribution, as a median, a tertile or a clinical threshold would be.

A reversal is the table’s two subgroups agreeing that the treatment helps and the whole trial saying it harms, or the reverse. Nothing about the treatment differs between subgroups, so every reversal is chance: an imbalance in the arms’ mix of high- and low-risk patients, which a coin produces and a subgroup table exposes.

One variable is a family of tables

Cut at its median, the strongest variable’s table reverses in 4.4% of trials of eighty — within sampling error of the 4.07% its binary version had, as it should be, since a median cut is that binary variable, here drawn on different random numbers. Tried at two cuts, the tertile points, some cut reverses in 6.4%; at three, 8.7%; at five — the twentieth, thirty-third, fiftieth, sixty-seventh and eightieth percentiles — 10.8%; at all nine deciles, 13.2%.

That rises with every cut tried and rises much less than the number of cuts. Nine independent chances at 4.4% each would give 33.3%; nine cuts of one variable give 13.2%. Measured as the number of independent chances at the average single cut’s rate that would produce the same total, five cuts amount to about 3.6 and nine to about 5.1.

The cuts are correlated for an obvious reason: a patient above the sixtieth percentile is above the fiftieth, and two cuts ten points apart split the patients into subgroups that share most of their members. A trial whose arms happen to be unbalanced in risk near the middle of the variable is unbalanced for every cut near the middle, and if one of those cuts shows a reversal the next one along usually does too.

How correlated two cuts of one variable are

The overlap has an exact size. Cutting one variable at the quantiles a<ba < b of its distribution gives two indicators — above the first cut, above the second — whose correlation is

a(1−b)(1−a) b,\sqrt{\frac{a(1-b)}{(1-a)\,b}},

whatever the variable’s distribution. The median and the sixty-seventh percentile correlate 0.707; the two tertile points, 0.5; the twentieth and eightieth percentiles, 0.25. Neighbouring cuts share most of their information, and the subgroup tables they produce share most of their chance imbalances, so a reversal on one tends to come with a reversal on the next.

That is the same geometry the arcsine that closes it measured between cuts of two correlated variables, with the correlation between the variables set to one. A family of cuts on one variable is a family of correlated binary variables, and the effective number of independent chances it offers is set by how quickly that correlation falls as the cuts move apart — which is why five cuts packed between the twentieth and eightieth percentiles amount to 3.6 chances rather than five.

Where on the variable a cut reverses

Not every cut is an equal chance.

How often one cut of the strongest baseline variable shows a reversal, by where it is cut, trials of 80. Cut at the median the table reverses in 4.4% of trials; at the thirtieth and seventieth percentiles 3.2% and 4.0%; at the first and ninth deciles 0.7% and 0.4%.
Fig. 2 How often a single cut of the strongest baseline variable shows a reversal, by where it is cut, from the first decile to the ninth, in trials of eighty.

The median is the most productive single cut, at 4.4%; the thirtieth and seventieth percentiles give 3.2% and 4.0%; the first and ninth deciles give 0.7% and 0.4%. A cut near the end of the distribution leaves one subgroup of eight or so patients, whose own treatment difference is so noisy that it agrees in sign with the other subgroup only by luck, and a reversal needs the two subgroups to agree. The middle cuts are where both subgroups are large enough to carry a stable sign and where an imbalance in risk between the arms is largest.

So the family that matters is the middle of the variable. An analyst who tries a “clinical threshold” at the ninetieth percentile adds almost nothing to the chance of a reversal; one who tries the median, the tertiles and the quartiles adds nearly everything a set of cuts can add, because those are the correlated, productive ones.

Cuts times variables

A baseline table holds many continuous variables, and each can be cut at several points.

How often a trial of 80 shows a reversal somewhere in its baseline table, cutting each variable once or at five points. With twenty continuous variables cut at their medians, some table reverses in 16.2% of trials; cut at five points each, in 32.2%. The hundred tables of the second amount to about 31.2 independent chances at the average single table's rate.
Fig. 3 How often a trial of eighty shows a reversal somewhere in its baseline table, against the number of continuous variables tabulated, with each variable cut once at its median or at five points.

A smaller table shows the same doubling. With five continuous variables — the handful a trial report tabulates by age, weight, severity, a biomarker and a score — some median-cut table reverses in 6.7% of trials, and with each cut at five points, 16.9%.

How often a randomised trial of eighty shows a Simpson reversal on some baseline variable, against how many were tabulated. The treatment helps every patient equally. On one nominated, strongly prognostic variable the reversal appears in 4.1% of trials. Tabulating 2, 5, 10, 20, 50, 100 variables of mixed prognostic value, some variable reverses in 4.1%, 6.5%, 10.8%, 16.1%, 26.1%, 34.1% of trials; stratifying the randomisation on the most prognostic one gives 0.2%, 2.9%, 7.7%, 12.2%, 23.4%, 31.9%.
Fig. 4 The binary version for comparison: how often some baseline variable’s median-split table reverses, against the number of variables tabulated, with and without stratifying the randomisation on the strongest. Each continuous variable cut at several points multiplies this curve’s chances.

With twenty continuous variables cut at their medians, some table reverses in 16.2% of trials — the binary result again, to within its sampling error. Cut at five points each, a hundred tables in all, some table reverses in 32.2%. The five-fold family on each variable doubles the chance of a reversal somewhere in the table.

As independent chances at the average single table’s rate, the hundred tables amount to about 31.2. The two families do not multiply cleanly. Twenty variables cut at their medians amount to about 13.1 independent variables at the variables’ average rate — the thirteen the twenty-variable table found — and five cuts of one variable to about 3.6 effective cuts; their product, about 47, overstates the whole table’s 31.2, because the variables are themselves correlated through the risk they all measure, and a trial unbalanced in risk is unbalanced on many of them at once, at many of their cuts.

Why the variables beyond the first add so little each

The curve over variables rises slowly at first because the variables are not alike. The second variable measured here tracks the underlying risk at a correlation of only 0.1, and a variable that carries almost no prognostic information cannot reverse a table: the subgroups it makes have almost the same mix of risk, so the arms’ imbalance does not separate them. The variables that reverse are the prognostic ones, and the strong ones are the few that reverse often — which is why twenty variables of mixed strength reverse at an average rate of 1.34% each against the strongest one’s 4.27%.

Cutting at several points multiplies the chances only on those prognostic variables. Adding cuts to a variable that predicts nothing adds tables that cannot reverse. The count that matters is the number of prognostic variables times the effective number of middle cuts on each, and in a typical baseline table both are small numbers whose product is not.

What a trial size changes

The reversal a coin cannot prevent found that a randomised trial’s reversal rate on a nominated variable falls slowly as the trial grows, because the arms’ imbalance and the treatment effect both shrink with it. The cut family behaves the same way. At forty patients the median alone reverses in 3.9% of trials and any of five cuts in 9.6%; at a hundred and sixty, 4.2% and 11.7%. The effective number of cuts moves the other way, from 3.8 at forty to 3.3 at a hundred and sixty, because a larger trial’s subgroup estimates are steadier and neighbouring cuts agree more often.

How often a trial of 160 shows a Simpson reversal on its strongest baseline variable, against how many cut points are tried. The treatment helps every patient equally. Cut at its median, the variable's table reverses in 4.2% of trials; trying 2, 3, 5, 9 cuts, in 6.9%, 8.8%, 11.7%, 14.1%. Independent chances would give 32.2% at nine.
Fig. 5 The same family of cuts in trials of a hundred and sixty. The median alone reverses about as often as at eighty; the five- and nine-cut families a little more often, because a larger trial’s subgroups are stable enough for more of the cuts to carry a sign.

None of these rates is small enough to ignore at any trial size a subgroup table is usually drawn from. A reversal on some cut of some continuous baseline variable is an event that one trial in three produces from a treatment that helps everyone.

Why the cut behaves like a p-value hunt

Trying cuts until a table looks interesting is the garden of forking paths in its most innocent form. Nobody computes a p-value, nobody tests anything; the analyst tabulates the variable at a few natural points and reports the one that tells a story. An outcome cut in two found that choosing the cut on the outcome after looking inflated the false-positive rate from 4.2% at one fixed cut to 17.7% at five chosen, and the cut that fitted best found that even a cut chosen without looking at the treatment difference flatters itself. A subgroup table is the same search pointed at a different target: the cut that makes the most striking table.

The reversal adds one feature the p-value hunt does not have. It needs no test to be striking. A table in which the treatment helps the young, helps the old and harms the trial as a whole reads as a paradox that demands an explanation, and the explanation it attracts — confounding, an interaction, a harmful effect in some hidden group — is wrong every time here, because none of those exists in the trial.

There is also an asymmetry in what gets written down. A cut that produced an ordinary table — both subgroups helped, the whole helped — is not a finding and is rarely reported as a cut at all; the variable is summarised and the reader never learns it was tabulated at five points. A cut that produced a reversal is a finding, and is reported with its threshold as though that threshold had been the only one. The family is visible only to the analyst, and the count of chances it contains is the one number that would let a reader discount the striking table correctly.

What a reader of a striking table can work out

A reader shown one reversed table does not know how many tables were looked at, but can often bound it. A trial report with a baseline table of five continuous variables, and a text mentioning subgroups by “age above and below 65” and “severity by tertile”, has looked at no fewer than two cuts on two variables and quite possibly the median and quartiles of everything. The counts here convert that into a chance: with five variables cut at five points each, one trial in six shows a reversal somewhere with nothing behind it.

The reversal’s own size is weak evidence against chance. The reversals here are not marginal: in a trial of eighty the two subgroup differences and the overall difference can each be several percentage points, because every one of them is a difference of two small proportions. Twenty analyses of nothing found that the most striking of twenty null analyses is usually striking; the most striking of a hundred subgroup tables is a reversal one time in three, and a reversal is the most striking thing a subgroup table can show.

The question worth asking of any reported reversal is therefore not whether it is surprising but how many tables were in the family it was drawn from, and how many analyses there really were is the general form of that question. For a continuous baseline table, the answer includes a second factor nobody writes down: the cuts.

The cut costs twice

Cutting a continuous baseline variable has a cost that has already been priced, and the reversal is a second cost of the same act. A baseline cut in two found that adjusting a trial’s analysis for a median-split baseline keeps only 2/π of the precision the variable as measured would give, and at a strong correlation with the outcome that costs more than twice the patients. The subgroup table built on the same cut throws away that precision and, as a by-product, manufactures the reversals counted here: each cut is a new way for a coin’s imbalance to be displayed as a paradox.

The two costs point to one repair. A variable used as measured — in a covariate-adjusted estimate, or in a plot of the treatment difference against the variable — keeps its full value for precision and offers no cut at which to reverse. A reader who wants to know whether the treatment’s benefit varies with age is better served by an estimate of how the benefit changes per decade, with its interval, than by a table at a cut chosen from several, and the estimate cannot show a Simpson reversal because it never compares a whole with two parts.

What a baseline table can report honestly

Report continuous variables as continuous. A mean and standard deviation by arm, or a covariate-adjusted estimate, uses the variable without choosing a cut, and cannot show a reversal. If subgroups are needed, fix the cut in the protocol.

If several cuts were looked at, say how many. The effective count is well below the nominal one, but it is not one: five middle cuts of a strong variable are worth three to four chances, and a reader judging a striking table needs to know the family it came from.

Treat a reversal in a randomised trial as a property of the allocation, not of the treatment. A coin does not balance a subgroup table, it balances the arms on average; the reversal a coin cannot prevent is the result that says so, and on continuous variables cut at several points it is roughly twice as common as on the same variables cut once.

Stratify on the strongest variable if a table of it will be reported. Stratifying the randomisation on the most prognostic variable removes its reversals and leaves the others; since the productive cuts are on the prognostic variables, that removes a large share of the table’s chances.

What is counted here and what is not

The strongest variable’s median cut reverses in 4.4% of trials of eighty; five cuts in 10.8% and nine in 13.2%, about 3.6 and 5.1 effective chances.

Twenty continuous variables cut at five points each show some reversal in 32.2% of trials, against 16.2% cut at their medians; about 31.2 effective chances in a hundred tables.

Every rate is counted over four thousand simulated trials — three thousand for the tables of many variables — with the treatment raising every patient’s chance of success by six percentage points, allocation by coin, and baseline variables that track the underlying risk at correlations spread evenly from 0.95 down to 0.1. Cuts are at population quantiles, fixed in advance.

Not counted: cuts chosen at the sample’s own quantiles, which follow the data and would behave slightly differently; variables whose relation to risk is not monotone, where cuts far from the median could become productive; and reversals of the treatment difference’s size rather than its sign, which a reader may find just as striking.

Still open: cuts chosen to look

Every cut here is fixed in advance, and the family is a list an analyst might try. The more consequential version chooses the cut by looking at the table — the threshold at which the subgroups differ most, found by sliding the cut along the variable — which turns a family of five into a continuum and a maximum over it.

For a continuous scan the effective number of cuts is not a count but a property of how quickly the table changes as the cut moves, the same quantity that governs a break that was looked for. How often a sliding cut produces a reversal somewhere along a strong variable, and how the answer depends on the trial’s size, is a supremum over a process that the counts here only sample at five and nine points, and it has not been computed.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

DichotomisationEffective number of testsThe garden of forking pathsMultiple comparisonsRandomisationSimpson's paradoxSubgroup analysisThreshold