Simpson's reversal is a region, not a table
A treatment does better than the control in group A. It does better in group B. Put the groups together and it does worse. Every number is correct and no arithmetic has been abused.
Watching it turn
The slider is the point of the figure. The four success rates are held absolutely fixed — treatment beats control in both groups, always, at every slider position. All that moves is how many people from each group ended up in each arm.
At some allocations the treatment wins overall. At others it loses. The data about the treatment’s effectiveness has not changed at all.
That is the whole of the paradox: an overall comparison is a weighted average, and the weights come from the allocation rather than from the treatment. If group A has a high success rate whatever is done, and most of the control arm is drawn from group A, then the control arm looks good for reasons that have nothing to do with the control.
How much of the space
The famous table is a single point, and a point cannot say whether this is a curiosity or a routine hazard.
Sweeping every allocation with the rates held fixed: 31% of allocations reverse the overall comparison, and the worst reverses it by 12.6 percentage points.
That is not a corner of the space. It is a third of it, and the region is connected rather than scattered — which means a study that is even moderately imbalanced between the groups is in it.
What cannot reverse
The gate checks the complementary claim, because a phenomenon that happens everywhere is as uninformative as one that happens nowhere.
With the groups balanced — each split evenly between treatment and control — the reversal is impossible. The weighted average has equal weights, so it cannot contradict its components.
That identifies the cause precisely. The reversal requires unequal allocation, and unequal allocation is what randomisation exists to prevent. A randomised trial is protected from this, in expectation, by design; an observational comparison is not protected at all, because the allocation was made by the world rather than by the experimenter.
Which answer is right
The question people ask on meeting this is which comparison to believe, and the honest answer is that it depends on something the numbers do not contain.
If the grouping variable is a confounder — something that affects both the outcome and which arm a patient ended up in, like severity of illness — then the within-group comparison is the one to believe, and the aggregate is contaminated by the allocation.
If the grouping variable is a mediator — something the treatment itself causes, downstream of it — then conditioning on it removes part of the treatment’s effect, and the aggregate is the one to believe.
The table cannot say which. Both cases produce identical numbers. The distinction is causal, it comes from knowledge about the subject rather than from the data, and this is the reason causal inference exists as a separate discipline rather than being a branch of arithmetic.
Where it shows up
The Berkeley admissions case is the standard example: an apparent bias against women in aggregate, absent or reversed within every department, because applications were not evenly distributed across departments with different admission rates.
The pattern recurs wherever aggregate rates are compared across groups with different compositions — hospital mortality rates with different case mixes, school results with different intakes, regional statistics with different age structures.
In every case the practical rule is the same and it is short: a comparison of aggregate rates between groups with different compositions is not a comparison of the thing in question. Whether the disaggregated version is the right one instead is a separate question, and a harder one.
Why it belongs here
The connection to the rest of this site is the method rather than the subject.
The paradox is normally presented as a curiosity to be marvelled at. Parameterised, it becomes a region with a measurable extent, a stated cause, and a checkable claim about when it cannot happen. “How often” and “how large” have answers, and the answers are 31% and 12.6 points.
That is the same move as computing an interval’s coverage rather than quoting it: take the thing everyone recites, attach a parameter to it, and measure what happens across the range.
The weights, written out
The arithmetic behind the reversal is one line, and seeing it removes the sense of paradox entirely.
The overall rate for an arm is a weighted average of its group rates, with weights equal to the proportion of that arm drawn from each group:
The treatment and control arms have different weights, because they were filled differently. So the comparison of overall rates is a comparison of two different weighted averages of two different pairs of numbers, and there is no reason it should agree with a comparison made group by group.
Once stated like that the surprise dissolves and the right question appears: which weighting corresponds to the quantity of interest. That is not an arithmetic question.
Standardisation, and what it assumes
The standard fix is to compute both arms with the same weights — usually the population’s group proportions rather than the sample’s. This is direct standardisation, and it is how age-adjusted mortality rates and case-mix-adjusted hospital figures are produced.
It works, and it carries an assumption worth naming: it estimates what the rates would be if both arms had the reference composition, which is a counterfactual. If the treatment’s effect differs by group — if there is genuine interaction — then no single standardised number summarises it, and reporting one hides the thing that matters.
The check is simple. If the within-group effects point the same way and are of similar size, a standardised summary is reasonable. If they differ substantially, report the groups.
What randomisation does and does not fix
Randomisation makes the allocation independent of the group in expectation, so the weights in the two arms match and the reversal cannot systematically occur. That is the strongest practical argument for it and it is often stated too strongly.
In expectation does not mean in any particular trial. A small randomised trial can still land with imbalanced groups, and the reversal is available to it. Stratified randomisation and blocking exist to remove that residual chance rather than to fix a bias.
And randomisation only protects against the variables that exist at randomisation. A trial randomised properly can still show the reversal on a variable measured afterwards — which brings back the mediator problem, since a post-treatment variable may be caused by the treatment.
The causal question, stated properly
The essay has said the data cannot say whether to condition on the grouping variable. It is worth being precise about what would settle it.
The question is whether the grouping variable is a common cause of the treatment and the outcome — in which case conditioning removes confounding and the within-group comparison is right — or whether it lies on the causal path from treatment to outcome, in which case conditioning removes part of the effect being measured.
Those two situations produce identical tables. The distinction is in the causal structure, which is knowledge about the subject rather than about the numbers, and which is exactly what a causal diagram is for: it makes the assumption explicit and then says which adjustments are valid given it.
The uncomfortable consequence is that no amount of data settles it. Two analysts with the same table and different beliefs about the causal structure will correctly reach opposite conclusions, and the disagreement is about the structure rather than about the statistics.
Why this is the most consequential essay here
Of everything on this site, the reversal is the one with the most direct route to a bad decision.
Aggregate comparisons between groups with different compositions are made constantly — hospital league tables, school performance, regional health statistics, employment rates by demographic. In every case the composition differs, in every case the aggregate is a weighted average with mismatched weights, and in most cases nobody has checked whether the disaggregated comparison agrees.
The measurement on this page — 31% of allocations reverse — is the number to carry. It is not a rare configuration requiring adversarial construction. It is a third of the space, and any comparison drawn from observational data is somewhere in that space without knowing where.
Continuous versions of the same thing
The reversal is usually shown with two groups and a rate, and it is not confined to either.
The continuous analogue is a regression whose slope reverses when a covariate is included. The mechanism is identical: the overall slope is a weighted combination of the within-group slopes and the between-group differences, and the weights depend on how the covariate is distributed.
It appears in survival analysis, in time series aggregated over periods with different compositions, and in any comparison of averages across populations that differ in structure.
The name attaches to the two-by-two table because that is where it was first written down, and treating it as a curiosity about tables is what allows it to be missed everywhere else.
Testing whether the aggregate is trustworthy
A practical procedure, since the essay’s conclusion is otherwise “it depends”.
Check whether the grouping variable is associated with the arm. If treatment and control have similar compositions, the weights match and the aggregate is safe. This is a cross-tabulation and takes seconds.
Check whether the grouping variable is associated with the outcome. If it is not, it cannot confound, whatever the imbalance.
Both associations are required for the reversal. Either alone is harmless, which is why the check is cheap: a variable that fails one of the two can be set aside.
That is the operational definition of a confounder, and running it over the measured variables is the minimum diligence before reporting an aggregate comparison from observational data.
What it cannot do is say anything about the variables that were not measured, which is the permanent limitation of observational work and the reason randomisation is worth its cost.
Why the reversal is not a paradox
A closing note on the name, which does the phenomenon a disservice.
Nothing about it is contradictory. Two different weighted averages of the same numbers can order differently; that is a fact about weighted averages, available to anyone who writes down the arithmetic.
Calling it a paradox suggests a puzzle in the mathematics, and it invites the response that statistics is slippery or that numbers can be made to say anything. Neither is true here. The numbers say exactly one thing, precisely, and the difficulty is entirely in deciding which question the numbers were meant to answer.
That is a question about the causal structure of the situation, and it is the sort of question that data cannot settle and that the people who collected the data are usually best placed to answer. Framing it as a paradox obscures that, and framing it as a weighting problem makes it tractable.