Reversals that are not errors

Simpson's reversal is a region, not a table

The treatment wins in both groups and loses overall. That is normally shown with one famous table, which cannot answer the two questions a reader has — how often, and how large. Swept, it turns out to occupy 31% of the allocation space.

A treatment does better than the control in group A. It does better in group B. Put the groups together and it does worse. Every number is correct and no arithmetic has been abused.

Two groups, one treatment, and both readings of the same numbersThe treatment wins in group A (93.0% against 87.0%) and in group B (73.0% against 69.0%), and loses overall (78.1% against 82.6%). Nothing here is a trick; the allocation differs between the groups.treatedcontrolverdictgroup A93.0%87.0%treatmentgroup B73.0%69.0%treatmentboth together78.1%82.6%controlgroup A: 88 treated of 351 · group B: 263 of 351wins in both groups, loses overallthe rates never changeonly the allocation does
Fig. 1 The table, with the allocation on a slider. The success rates never change as it moves; only how the groups are split between the two arms.

Watching it turn

The slider is the point of the figure. The four success rates are held absolutely fixed — treatment beats control in both groups, always, at every slider position. All that moves is how many people from each group ended up in each arm.

At some allocations the treatment wins overall. At others it loses. The data about the treatment’s effectiveness has not changed at all.

That is the whole of the paradox: an overall comparison is a weighted average, and the weights come from the allocation rather than from the treatment. If group A has a high success rate whatever is done, and most of the control arm is drawn from group A, then the control arm looks good for reasons that have nothing to do with the control.

How much of the space

The famous table is a single point, and a point cannot say whether this is a curiosity or a routine hazard.

Sweeping every allocation with the rates held fixed: 31% of allocations reverse the overall comparison, and the worst reverses it by 12.6 percentage points.

That is not a corner of the space. It is a third of it, and the region is connected rather than scattered — which means a study that is even moderately imbalanced between the groups is in it.

Which allocations reverse the overall comparison. The per-group success rates are held fixed; only the split of each group between treatment and control changes. 32% of the allocations reverse, and the worst reverses by 13.1 percentage points.
Fig. 2 Which allocations reverse. The rates are identical at every point of this square.
A thousand people, a condition one in 100 has, a test that is 90% and 95%. 9 people have it and test positive. 50 do not have it and test positive anyway. So of the 59 positive results, 15% are right — and that is with a test most people would call accurate.
Fig. 3 A different arithmetic trap in the same family: correct numbers, and an intuition that reads them backwards.

The inequality behind the slider

The reversal has a condition, and writing it down says exactly which three quantities are in a competition.

Each arm’s overall rate is a weighted average of the two group rates, with the weights being that arm’s group proportions. So the aggregate comparison is the within-group advantage plus a term made of the difference in weights.

Let δ be the treatment’s advantage inside each group, Δ the gap between the two groups’ own success rates, and w the difference between the two arms’ shares of the better group. Then the aggregate reverses when

wΔ  >  δ.w\,\Delta \;>\; \delta .

Three quantities, one inequality, and every case of the phenomenon is an instance of it.

Which says what the slider is doing. It moves w alone, holding δ and Δ fixed, so it moves the left-hand side through the right-hand side and the aggregate flips. It also says when the reversal is impossible: if the two arms have the same group composition, w = 0 and no value of Δ can produce one.

A confounder large enough to matter is not sufficient; it has to be unevenly allocated as well, which is the same structure as a positive test needing a base rate — a quantity computed inside the data being overturned by a quantity about how the data came to be.

What cannot reverse

The gate checks the complementary claim, because a phenomenon that happens everywhere is as uninformative as one that happens nowhere.

With the groups balanced — each split evenly between treatment and control — the reversal is impossible. The weighted average has equal weights, so it cannot contradict its components.

That identifies the cause precisely. The reversal requires unequal allocation, and unequal allocation is what randomisation exists to prevent. A randomised trial is protected from this, in expectation, by design; an observational comparison is not protected at all, because the allocation was made by the world rather than by the experimenter.

What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed.
Fig. 4 How much work a base rate does compared with the quantity everyone attends to.

Which answer is right

The question people ask on meeting this is which comparison to believe, and the honest answer is that it depends on something the numbers do not contain.

If the grouping variable is a confounder — something that affects both the outcome and which arm a patient ended up in, like severity of illness — then the within-group comparison is the one to believe, and the aggregate is contaminated by the allocation.

If the grouping variable is a mediator — something the treatment itself causes, downstream of it — then conditioning on it removes part of the treatment’s effect, and the aggregate is the one to believe.

The table cannot say which. Both cases produce identical numbers. The distinction is causal, it comes from knowledge about the subject rather than from the data, and this is the reason causal inference exists as a separate discipline rather than being a branch of arithmetic.

Two measurements of the same thing, correlated 0.60. Pick the worst 15% on the first measurement and their average rises by 0.78 on the second. Pick the best and theirs falls by 0.48. No treatment was given to anybody.
Fig. 5 Another effect that appears without any intervention, and that a poorly matched comparison group will not remove.

Where it shows up

The Berkeley admissions case is the standard example: an apparent bias against women in aggregate, absent or reversed within every department, because applications were not evenly distributed across departments with different admission rates.

The pattern recurs wherever aggregate rates are compared across groups with different compositions — hospital mortality rates with different case mixes, school results with different intakes, regional statistics with different age structures.

In every case the practical rule is the same and it is short: a comparison of aggregate rates between groups with different compositions is not a comparison of the thing in question. Whether the disaggregated version is the right one instead is a separate question, and a harder one.

Coverage of four nominal 95% intervals, n = 30. Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.
Fig. 6 The same move applied to an interval: parameterise the thing everyone recites, then measure across the range.

Why it belongs here

The connection to the rest of this site is the method rather than the subject.

The paradox is normally presented as a curiosity to be marvelled at. Parameterised, it becomes a region with a measurable extent, a stated cause, and a checkable claim about when it cannot happen. “How often” and “how large” have answers, and the answers are 31% and 12.6 points.

That is the same move as computing an interval’s coverage rather than quoting it: take the thing everyone recites, attach a parameter to it, and measure what happens across the range.

The weights, written out

The arithmetic behind the reversal is one line, and seeing it removes the sense of paradox entirely.

The overall rate for an arm is a weighted average of its group rates, with weights equal to the proportion of that arm drawn from each group:

rˉ=wArA+wBrB\bar{r} = w_A r_A + w_B r_B

The treatment and control arms have different weights, because they were filled differently. So the comparison of overall rates is a comparison of two different weighted averages of two different pairs of numbers, and there is no reason it should agree with a comparison made group by group.

Once stated like that the surprise dissolves and the right question appears: which weighting corresponds to the quantity of interest. That is not an arithmetic question.

Standardisation, and what it assumes

The standard fix is to compute both arms with the same weights — usually the population’s group proportions rather than the sample’s. This is direct standardisation, and it is how age-adjusted mortality rates and case-mix-adjusted hospital figures are produced.

It works, and it carries an assumption worth naming: it estimates what the rates would be if both arms had the reference composition, which is a counterfactual. If the treatment’s effect differs by group — if there is genuine interaction — then no single standardised number summarises it, and reporting one hides the thing that matters.

The check is simple. If the within-group effects point the same way and are of similar size, a standardised summary is reasonable. If they differ substantially, report the groups.

What randomisation does and does not fix

Randomisation makes the allocation independent of the group in expectation, so the weights in the two arms match and the reversal cannot systematically occur. That is the strongest practical argument for it and it is often stated too strongly.

In expectation does not mean in any particular trial. A small randomised trial can still land with imbalanced groups, and the reversal is available to it. Stratified randomisation and blocking exist to remove that residual chance rather than to fix a bias.

And randomisation only protects against the variables that exist at randomisation. A trial randomised properly can still show the reversal on a variable measured afterwards — which brings back the mediator problem, since a post-treatment variable may be caused by the treatment.

The causal question, stated properly

The essay has said the data cannot say whether to condition on the grouping variable. It is worth being precise about what would settle it.

The question is whether the grouping variable is a common cause of the treatment and the outcome — in which case conditioning removes confounding and the within-group comparison is right — or whether it lies on the causal path from treatment to outcome, in which case conditioning removes part of the effect being measured.

Those two situations produce identical tables. The distinction is in the causal structure, which is knowledge about the subject rather than about the numbers, and which is exactly what a causal diagram is for: it makes the assumption explicit and then says which adjustments are valid given it.

The uncomfortable consequence is that no amount of data settles it. Two analysts with the same table and different beliefs about the causal structure will correctly reach opposite conclusions, and the disagreement is about the structure rather than about the statistics.

Why this is the most consequential essay here

Of everything on this site, the reversal is the one with the most direct route to a bad decision.

Aggregate comparisons between groups with different compositions are made constantly — hospital league tables, school performance, regional health statistics, employment rates by demographic. In every case the composition differs, in every case the aggregate is a weighted average with mismatched weights, and in most cases nobody has checked whether the disaggregated comparison agrees.

The measurement on this page — 31% of allocations reverse — is the number to carry. It is not a rare configuration requiring adversarial construction. It is a third of the space, and any comparison drawn from observational data is somewhere in that space without knowing where.

Continuous versions of the same thing

The reversal is usually shown with two groups and a rate, and it is not confined to either.

The continuous analogue is a regression whose slope reverses when a covariate is included. The mechanism is identical: the overall slope is a weighted combination of the within-group slopes and the between-group differences, and the weights depend on how the covariate is distributed.

It appears in survival analysis, in time series aggregated over periods with different compositions, and in any comparison of averages across populations that differ in structure.

The name attaches to the two-by-two table because that is where it was first written down, and treating it as a curiosity about tables is what allows it to be missed everywhere else.

Testing whether the aggregate is trustworthy

A practical procedure, since the essay’s conclusion is otherwise “it depends”.

Check whether the grouping variable is associated with the arm. If treatment and control have similar compositions, the weights match and the aggregate is safe. This is a cross-tabulation and takes seconds.

Check whether the grouping variable is associated with the outcome. If it is not, it cannot confound, whatever the imbalance.

Both associations are required for the reversal. Either alone is harmless, which is why the check is cheap: a variable that fails one of the two can be set aside.

That is the operational definition of a confounder, and running it over the measured variables is the minimum diligence before reporting an aggregate comparison from observational data.

What it cannot do is say anything about the variables that were not measured, which is the permanent limitation of observational work and the reason randomisation is worth its cost.

Why the reversal is not a paradox

A closing note on the name, which does the phenomenon a disservice.

Nothing about it is contradictory. Two different weighted averages of the same numbers can order differently; that is a fact about weighted averages, available to anyone who writes down the arithmetic.

Calling it a paradox suggests a puzzle in the mathematics, and it invites the response that statistics is slippery or that numbers can be made to say anything. Neither is true here. The numbers say exactly one thing, precisely, and the difficulty is entirely in deciding which question the numbers were meant to answer.

That is a question about the causal structure of the situation, and it is the sort of question that data cannot settle and that the people who collected the data are usually best placed to answer. Framing it as a paradox obscures that, and framing it as a weighting problem makes it tractable.

The 31% is a measurement of a grid, not of nature

One caveat about the headline number, stated here because the essay has leaned on it throughout and because it is the kind of thing that quietly becomes false.

The allocation space is continuous: the two group sizes can take any values. The 31% is computed by laying a grid over that space and counting the cells that reverse — with the rates held fixed and a 31-point grid, 31% of the interior cells reverse, and the worst reverses the comparison by 12.6 percentage points.

Change the grid and the number moves. A 41-point grid over the same region gives 32%. That is not an error in either; it is what happens when a continuous region is measured by counting cells, and it means the third significant figure of the headline number is an artefact of a choice nobody would think to state.

The right reading is therefore the one the essay has been making — the reversal occupies a substantial fraction of the allocation space rather than a corner of it, and that qualitative statement is robust to the grid. The specific 31% should be read as “about a third”, and the figure’s grid is stated in its own caption so the number can be reproduced rather than merely believed.

Two things are exact and do not depend on the grid at all, and they are the ones carrying the argument. Balanced allocation cannot reverse, which is an algebraic fact about the weights and is asserted as one. And the worst case found is a genuine allocation with a genuine 12.6-point reversal, which is a witness rather than a proportion: it exists, it is constructed, and no grid refinement can remove it.

That division is worth generalising. A swept region reported as a fraction inherits the sweep’s resolution; a witness found inside it does not. Where an argument can be carried by the witness, it should be, because the witness survives every methodological objection that can be made to the sweep.

The same study read three ways. At time 2 the truth is 0.497. Kaplan–Meier gives 0.532; dropping the censored subjects gives 0.180; treating the censoring time as the event time gives 0.392. Both naive readings understate survival, because the subjects they mishandle are the ones doing well.
Fig. 7 Another aggregate that misleads when a subgroup is handled wrongly, and the reading that does not.

What has to be true for the aggregate to be trusted

The essay has argued that the disaggregated comparison is usually the right one and that the reason is causal rather than statistical. It is worth turning that around and asking the question a reader actually faces: given only an aggregate table, what would make it trustworthy?

Three conditions, and each is checkable from the margins of the table rather than from any theory about the subject.

Allocation is balanced across the groups. If each group splits the same way between the conditions being compared, the weights are identical and the reversal is arithmetically impossible. This is the exact result the gate asserts, and it is why randomisation with a fixed ratio buys so much: it does not make confounding unlikely, it makes this particular failure unavailable.

Or the groups are the same size in the relevant sense. Balance is about the split within each group, not the size of the groups themselves, and the two are constantly confused. A study with 1,000 subjects in one stratum and 50 in another can be perfectly balanced; a study with 500 in each can be badly unbalanced.

Or the grouping variable has no relation to the outcome. If the strata do not differ in their base rates, there is nothing for the weights to distort. This is the condition that cannot be checked from the aggregate table alone, and it is the one most often assumed.

The useful consequence is that the first two are visible in any table that reports its cell counts, and invisible in any table that reports only percentages. A results table giving rates without denominators has removed exactly the information needed to know whether its own aggregate is trustworthy — and that is the single most common way this failure is shipped.

So the practical demand is small and specific. Report the counts, not only the rates. Anyone can then check the balance in a few seconds, and the question of whether the aggregate can reverse stops being a matter of trust — as it does for the other reversal this field measures.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AggregationAllocationConfoundingRandomisationSimpson's paradox