Reversals that are not errors

Simpson's reversal is a region, not a table

The treatment wins in both groups and loses overall. That is normally shown with one famous table, which cannot answer the two questions a reader has — how often, and how large. Swept, it turns out to occupy 31% of the allocation space.

A treatment does better than the control in group A. It does better in group B. Put the groups together and it does worse. Every number is correct and no arithmetic has been abused.

Two groups, one treatment, and both readings of the same numbersThe treatment wins in group A (93.0% against 87.0%) and in group B (73.0% against 69.0%), and loses overall (78.1% against 82.6%). Nothing here is a trick; the allocation differs between the groups.treatedcontrolverdictgroup A93.0%87.0%treatmentgroup B73.0%69.0%treatmentboth together83.1%78.0%treatmentgroup A: 175 treated of 350 · group B: 175 of 350no reversal at this allocationthe rates never changemove the slider
Fig. 1 The table, with the allocation on a slider. The success rates never change as it moves; only how the groups are split between the two arms.

Watching it turn

The slider is the point of the figure. The four success rates are held absolutely fixed — treatment beats control in both groups, always, at every slider position. All that moves is how many people from each group ended up in each arm.

At some allocations the treatment wins overall. At others it loses. The data about the treatment’s effectiveness has not changed at all.

That is the whole of the paradox: an overall comparison is a weighted average, and the weights come from the allocation rather than from the treatment. If group A has a high success rate whatever is done, and most of the control arm is drawn from group A, then the control arm looks good for reasons that have nothing to do with the control.

How much of the space

The famous table is a single point, and a point cannot say whether this is a curiosity or a routine hazard.

Sweeping every allocation with the rates held fixed: 31% of allocations reverse the overall comparison, and the worst reverses it by 12.6 percentage points.

That is not a corner of the space. It is a third of it, and the region is connected rather than scattered — which means a study that is even moderately imbalanced between the groups is in it.

Which allocations reverse the overall comparisonThe per-group success rates are held fixed; only the split of each group between treatment and control changes. 32% of the allocations reverse, and the worst reverses by 13.1 percentage points.fraction of group A given the treatmentgroup B0.00.51.032% reverseshaded: the overall comparison reversesthe rates are identical everywhere here
Fig. 2 Which allocations reverse. The rates are identical at every point of this square.
A thousand people, a condition one in a hundred has, a test that is 90% and 95%9 people have it and test positive. 50 do not have it and test positive anyway. So of the 59 positive results, 15% are right — and that is with a test most people would call accurate.9 have it, test positive50 do not have it, test positive1 have it, test negative15% of positives are trueeach dot is one personthe arithmetic is not in dispute
Fig. 3 A different arithmetic trap in the same family: correct numbers, and an intuition that reads them backwards.

What cannot reverse

The gate checks the complementary claim, because a phenomenon that happens everywhere is as uninformative as one that happens nowhere.

With the groups balanced — each split evenly between treatment and control — the reversal is impossible. The weighted average has equal weights, so it cannot contradict its components.

That identifies the cause precisely. The reversal requires unequal allocation, and unequal allocation is what randomisation exists to prevent. A randomised trial is protected from this, in expectation, by design; an observational comparison is not protected at all, because the allocation was made by the world rather than by the experimenter.

What a positive test means, sensitivity 90%, specificity 95%At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in ten, 33 are. The test has not changed.00.2500.5000.7501how common the condition ischance a positive result is true1 in 10,0001 in 1,0001 in 1001 in 101 in 11.8%66.7%one test, every prevalencethe base rate outweighs the test
Fig. 4 How much work a base rate does compared with the quantity everyone attends to.

Which answer is right

The question people ask on meeting this is which comparison to believe, and the honest answer is that it depends on something the numbers do not contain.

If the grouping variable is a confounder — something that affects both the outcome and which arm a patient ended up in, like severity of illness — then the within-group comparison is the one to believe, and the aggregate is contaminated by the allocation.

If the grouping variable is a mediator — something the treatment itself causes, downstream of it — then conditioning on it removes part of the treatment’s effect, and the aggregate is the one to believe.

The table cannot say which. Both cases produce identical numbers. The distinction is causal, it comes from knowledge about the subject rather than from the data, and this is the reason causal inference exists as a separate discipline rather than being a branch of arithmetic.

Two measurements of the same thing, correlated 0.60Pick the worst 15% on the first measurement and their average rises by 0.55 on the second. Pick the best and theirs falls by 0.59. No treatment was given to anybody.first measurementsecond+0.55-0.59predicted 0.64 from the correlation alonenobody was treated
Fig. 5 Another effect that appears without any intervention, and that a poorly matched comparison group will not remove.

Where it shows up

The Berkeley admissions case is the standard example: an apparent bias against women in aggregate, absent or reversed within every department, because applications were not evenly distributed across departments with different admission rates.

The pattern recurs wherever aggregate rates are compared across groups with different compositions — hospital mortality rates with different case mixes, school results with different intakes, regional statistics with different age structures.

In every case the practical rule is the same and it is short: a comparison of aggregate rates between groups with different compositions is not a comparison of the thing in question. Whether the disaggregated version is the right one instead is a separate question, and a harder one.

Coverage of four nominal 95% intervals, n = 30Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.0.4000.6000.80010.2000.4000.6000.800the true proportionactual coverage of a nominal 95% intervalnominal 95%Wald — the textbook oneWilsonAgresti–CoullClopper–Pearsonsummed over all 31 outcomesno interval is 95% until it is counted
Fig. 6 The same move applied to an interval: parameterise the thing everyone recites, then measure across the range.

Why it belongs here

The connection to the rest of this site is the method rather than the subject.

The paradox is normally presented as a curiosity to be marvelled at. Parameterised, it becomes a region with a measurable extent, a stated cause, and a checkable claim about when it cannot happen. “How often” and “how large” have answers, and the answers are 31% and 12.6 points.

That is the same move as computing an interval’s coverage rather than quoting it: take the thing everyone recites, attach a parameter to it, and measure what happens across the range.

The weights, written out

The arithmetic behind the reversal is one line, and seeing it removes the sense of paradox entirely.

The overall rate for an arm is a weighted average of its group rates, with weights equal to the proportion of that arm drawn from each group:

rˉ=wArA+wBrB\bar{r} = w_A r_A + w_B r_B

The treatment and control arms have different weights, because they were filled differently. So the comparison of overall rates is a comparison of two different weighted averages of two different pairs of numbers, and there is no reason it should agree with a comparison made group by group.

Once stated like that the surprise dissolves and the right question appears: which weighting corresponds to the quantity of interest. That is not an arithmetic question.

Standardisation, and what it assumes

The standard fix is to compute both arms with the same weights — usually the population’s group proportions rather than the sample’s. This is direct standardisation, and it is how age-adjusted mortality rates and case-mix-adjusted hospital figures are produced.

It works, and it carries an assumption worth naming: it estimates what the rates would be if both arms had the reference composition, which is a counterfactual. If the treatment’s effect differs by group — if there is genuine interaction — then no single standardised number summarises it, and reporting one hides the thing that matters.

The check is simple. If the within-group effects point the same way and are of similar size, a standardised summary is reasonable. If they differ substantially, report the groups.

What randomisation does and does not fix

Randomisation makes the allocation independent of the group in expectation, so the weights in the two arms match and the reversal cannot systematically occur. That is the strongest practical argument for it and it is often stated too strongly.

In expectation does not mean in any particular trial. A small randomised trial can still land with imbalanced groups, and the reversal is available to it. Stratified randomisation and blocking exist to remove that residual chance rather than to fix a bias.

And randomisation only protects against the variables that exist at randomisation. A trial randomised properly can still show the reversal on a variable measured afterwards — which brings back the mediator problem, since a post-treatment variable may be caused by the treatment.

The causal question, stated properly

The essay has said the data cannot say whether to condition on the grouping variable. It is worth being precise about what would settle it.

The question is whether the grouping variable is a common cause of the treatment and the outcome — in which case conditioning removes confounding and the within-group comparison is right — or whether it lies on the causal path from treatment to outcome, in which case conditioning removes part of the effect being measured.

Those two situations produce identical tables. The distinction is in the causal structure, which is knowledge about the subject rather than about the numbers, and which is exactly what a causal diagram is for: it makes the assumption explicit and then says which adjustments are valid given it.

The uncomfortable consequence is that no amount of data settles it. Two analysts with the same table and different beliefs about the causal structure will correctly reach opposite conclusions, and the disagreement is about the structure rather than about the statistics.

Why this is the most consequential essay here

Of everything on this site, the reversal is the one with the most direct route to a bad decision.

Aggregate comparisons between groups with different compositions are made constantly — hospital league tables, school performance, regional health statistics, employment rates by demographic. In every case the composition differs, in every case the aggregate is a weighted average with mismatched weights, and in most cases nobody has checked whether the disaggregated comparison agrees.

The measurement on this page — 31% of allocations reverse — is the number to carry. It is not a rare configuration requiring adversarial construction. It is a third of the space, and any comparison drawn from observational data is somewhere in that space without knowing where.

Continuous versions of the same thing

The reversal is usually shown with two groups and a rate, and it is not confined to either.

The continuous analogue is a regression whose slope reverses when a covariate is included. The mechanism is identical: the overall slope is a weighted combination of the within-group slopes and the between-group differences, and the weights depend on how the covariate is distributed.

It appears in survival analysis, in time series aggregated over periods with different compositions, and in any comparison of averages across populations that differ in structure.

The name attaches to the two-by-two table because that is where it was first written down, and treating it as a curiosity about tables is what allows it to be missed everywhere else.

Testing whether the aggregate is trustworthy

A practical procedure, since the essay’s conclusion is otherwise “it depends”.

Check whether the grouping variable is associated with the arm. If treatment and control have similar compositions, the weights match and the aggregate is safe. This is a cross-tabulation and takes seconds.

Check whether the grouping variable is associated with the outcome. If it is not, it cannot confound, whatever the imbalance.

Both associations are required for the reversal. Either alone is harmless, which is why the check is cheap: a variable that fails one of the two can be set aside.

That is the operational definition of a confounder, and running it over the measured variables is the minimum diligence before reporting an aggregate comparison from observational data.

What it cannot do is say anything about the variables that were not measured, which is the permanent limitation of observational work and the reason randomisation is worth its cost.

Why the reversal is not a paradox

A closing note on the name, which does the phenomenon a disservice.

Nothing about it is contradictory. Two different weighted averages of the same numbers can order differently; that is a fact about weighted averages, available to anyone who writes down the arithmetic.

Calling it a paradox suggests a puzzle in the mathematics, and it invites the response that statistics is slippery or that numbers can be made to say anything. Neither is true here. The numbers say exactly one thing, precisely, and the difficulty is entirely in deciding which question the numbers were meant to answer.

That is a question about the causal structure of the situation, and it is the sort of question that data cannot settle and that the people who collected the data are usually best placed to answer. Framing it as a paradox obscures that, and framing it as a weighting problem makes it tractable.