The reversal a coin cannot prevent
Worth reading first: Simpson's reversal is a region, not a table.
Simpson’s reversal needs unequal allocation: the aggregate comparison is a weighted average and the weights come from how the groups were split between the arms. With the two arms drawn from the groups in the same proportions, the reversal is arithmetically impossible.
Randomisation makes those proportions equal in expectation. A trial does not get the expectation; it gets one draw, and the draw lands somewhere.
At eighty units the rate is 3.40%. One trial in thirty reports a Simpson’s reversal, with no confounding anywhere, no error in the analysis, and every number in the report correct.
The rate does not fall the way it should
The obvious expectation is that this is a small-sample problem, largest at forty units and shrinking steadily. It is not. The rate reads 3.23% at forty, 3.40% at eighty, 3.10% at a hundred and sixty, 2.48% at three hundred and twenty, 1.32% at six hundred and forty and 0.50% at twelve hundred and eighty.
It rises before it falls, and the peak is not where either half of the mechanism would put it.
Two conditions have to hold together for a trial to reverse, and they move in opposite directions with the trial’s size.
The treatment has to be ahead in both groups. Each group’s comparison is made on about a quarter of the trial, so at forty units each of the four cells holds about ten observations and a six-point difference is invisible in them. The two group comparisons agree in direction on 49.5% of trials at forty units, and on 92.7% at twelve hundred and eighty. This factor rises with the trial.
The two arms have to have different group compositions. That is the imbalance the reversal feeds on, and it shrinks like : the standard deviation of the difference in the arms’ group shares is 0.159 at forty units and 0.028 at twelve hundred and eighty. So a reversal conditional on the groups agreeing happens on 6.51% of trials at forty and 0.54% at twelve hundred and eighty. This factor falls with the trial.
The product of a rising factor and a falling one has an interior maximum, and it sits near eighty units. Neither factor alone would have located it, and neither is visible in a report — a paper stating that its subgroups agreed and its total did not has printed the numerator and nothing else.
A small true effect is barely protected at all
The size of the real effect is the other parameter, and it is the one that decides whether growing the trial helps.
At a true advantage of sixteen points the rate collapses quickly: 2.15% at forty units, 0.65% at a hundred and sixty, and nothing at all in four thousand trials at six hundred and forty or above. At six points it falls to 0.50% by twelve hundred and eighty. At two points it goes from 3.85% to 3.02% across the same thirty-two-fold range — which is to say that it does not usefully fall.
The inequality from the population table says why. A reversal needs w·Δ > δ, where Δ is the gap between the two groups’ own outcome rates, w is the difference in the arms’ group shares and δ is the treatment’s advantage. Δ is 0.50 here and fixed by the population. So the trial reverses when w exceeds δ/Δ, which is 0.32 at sixteen points, 0.12 at six and 0.04 at two — and 0.04 is still one and a half standard deviations of w at twelve hundred units.
The reading is uncomfortable and it is the point of the page. Randomisation’s protection against this is a function of how large the real effect is, and the trials where the effect is small are the trials where the protection is weakest. Those are also the trials whose subgroup tables get read most carefully.
The estimate is not biased, and that is the confusing part
Across every size and every effect drawn, the aggregate difference averages what it should: 0.060 against a true advantage of 0.060.
So nothing here is a bias. The aggregate estimator is unbiased, the two group estimators are unbiased, and the reversal is a statement about the joint behaviour of three unbiased estimates — about their ordering rather than about any one of their values. That is a different kind of object from everything else priced here, and it is why no correction to any single number addresses it.
It is also why the reversal is so persuasive when it appears. A reader checking each number individually finds nothing wrong, because there is nothing wrong with any number individually. The winner’s curse has the same shape from the other direction: there the selection is explicit and the individual estimates are honest, and here the ordering is the accident and the individual estimates are honest.
For scale: the observational version of the same population
The number to hold this against is the one from the population table, and the comparison is the argument for randomisation stated quantitatively rather than as a principle.
Sweeping the allocation with the rates fixed puts about a third of the allocation space in the reversing region. An observational comparison is a draw from that space made by the world, and nothing constrains where it lands: a study where the treated group is mostly drawn from the group with the high baseline rate sits deep inside the region, and no amount of data moves it out, because the region is a fact about the population rather than about the sample.
A randomised trial is also a draw from that space, but from a distribution concentrated hard on w = 0. Its reversal rate is 3.40% at eighty units and 0.50% at twelve hundred and eighty, and both of those are the tail of a distribution rather than a region of a space.
So randomisation converts a structural problem into a sampling one, and that is the whole of what it does here. The structural version does not improve with data and cannot be diagnosed from the table; the sampling version shrinks like and can be read off the arms’ margins. The difference between a third and a thirtieth is the measurement, and between “a property of the population” and “a property of this draw” is the reason it matters.
What randomisation does not do is make the number zero, which is the claim it usually gets. A design that stratifies does make it zero, and the gap between those two statements is this page.
Testing for heterogeneity finds nothing, because there is none
The reflex on meeting a reversing subgroup table is to test whether the treatment effect differs between the groups. It is the right instinct pointed at the wrong quantity, and the measurement says so unusually cleanly.
In every trial simulated here the treatment’s advantage is exactly six points in both groups. There is no interaction at all. The reversal arises entirely from the allocation, and the allocation is not a property of the treatment.
So the test finds nothing. Among the trials that reverse at eighty units, an interaction test declares the two group effects different on 0.28% of them. At three hundred and twenty units, on none of twenty thousand trials. Against a 5% test’s own false-positive rate — measured at 6.2% and 5.1% across all trials — the reversing trials are less likely than average to show a significant interaction.
The direction is not an accident. A reversal requires the two group differences to have the same sign, and an interaction test fires when they are far apart. Requiring them to agree in sign selects against exactly the configuration the test is looking for, so conditioning on a reversal makes the test quieter rather than louder.
The consequence is worth stating as a rule, because the two questions are asked with the same words. “Does the treatment work differently in the two groups?” and “why does the total disagree with the groups?” are different questions with different answers, and only the first has a test. The second is answered by looking at the allocation — the arms’ group shares, printed in the table’s own margins — and not by any comparison of effects. A trial that reports a reversal and a null interaction test has reported two facts that are entirely consistent, and a reader who takes the second as reassurance about the first has been reassured about something else.
The event that is not evidence
There is a weaker and much commoner observation that gets treated as a symptom of the same thing, and it is worth separating.
The aggregate sitting outside the range of the two group results happens on about 33% of trials, and — unlike the reversal — the share does not move with the trial’s size at all: 33.6% at forty units and 33.3% at twelve hundred and eighty.
That is because it is not a fact about aggregation. Three noisy estimates of the same quantity have an ordering, and one of the three is the middle one about a third of the time by symmetry; the aggregate is outside the other two whenever it is not the middle one, which is two thirds of the time for a random ordering and about a third here because the aggregate is a weighted average of the two and therefore pulled between them. A share that is constant in n is a share produced by the shape of the arithmetic rather than by anything in the data, and the constancy is the tell.
So the two observations look alike in a report and mean entirely different things. “The overall effect was smaller than in either subgroup” is worth nothing — it will happen to a third of honest trials whatever their size. “The overall effect had the opposite sign to both subgroups” is worth something, and how much depends on the size of the trial and the size of the effect, both of which have to be brought to the reading.
What stratifying does, and what it costs
The dashed line in the first figure is exactly zero, at every size, over twenty-four thousand trials. That is not a small rate rounded; stratified randomisation makes the reversal arithmetically unavailable, for the same reason the population version cannot reverse under balanced allocation.
Filling each group’s units alternately between the arms forces the two arms’ group shares to be equal up to one unit, so w is at most 1/n rather than a random variable with standard deviation 0.159. The inequality w·Δ > δ then fails by construction for any δ worth detecting.
The cost is the ordinary cost of blocking, and it is small: the trial has to know the group at the moment of allocation, and the analysis should then account for the stratification. What is bought is not a reduction in variance — though there is one — but the removal of a failure mode. Randomisation is not balance, and this is the cleanest demonstration of the difference on the site: the thing randomisation gives is a distribution over allocations, and the thing a trial needs is one good allocation.
The same argument covers every variable known before allocation, and it is the argument for balancing what is known in advance rather than adjusting for it afterwards. A trial that stratifies on the two or three variables it already knows matter has removed this entirely for those variables. It has removed nothing for the variables it did not think of, and a variable measured after allocation is a different problem again.
What is claimed here, and what is not
Three statements, each readable off the figures’ own frames.
A correctly randomised trial reverses at some size drawn. The rate is strictly positive somewhere on the curve at every value of the true effect the slider takes. It would not be if the arms had been balanced by accident — by allocating alternately, say, instead of by a coin — which is the easiest way to produce a picture that quietly measures the wrong design.
A stratified trial never does. Twenty-four thousand trials, no reversals, and the claim is an exact zero rather than a small number. An exact zero is strong and it is the right claim here, because it is arithmetic rather than statistical: with the arms’ compositions equal, the aggregate difference is a common-weighted average of two positive numbers and cannot be negative.
The share of trials whose aggregate is outside both group results barely moves with the trial’s size. That is what keeps the weak event from being read as the strong one. If the outside rate fell with n — which is what a reader expects — it would be a symptom of something, and it does not.
The reading that does not survive is the second claim taken for the first. A stratified design that reversed would mean either that the stratification was not doing what it says or that the reversal was not the event defined here, and either would make every other number on this page mean something different. The zero is stated as an equality for exactly that reason: a tolerance would let a genuine failure through as a rounding.
What to do with a trial that reports one
The rate on this page is small and non-zero, which is the least convenient possible answer, so it is worth saying what it licenses.
Check the allocation first, and it takes one line. The reversal needs w·Δ > δ, and every quantity in that inequality is printed in a standard subgroup table. w is the difference in the arms’ group shares, Δ is the gap between the groups’ own event rates, δ is the reported treatment effect. If w·Δ is smaller than δ, the reversal did not come from the allocation and something else is going on. If it is larger, the allocation is sufficient to explain it and nothing further is needed.
Prefer the stratified estimate, and say why. The aggregate in a trial with imbalanced arms is a comparison of two differently-composed populations. That is the same objection the population version makes, and it does not become weaker because the imbalance arrived by chance rather than by selection. Chance imbalance is still imbalance, and the standardised estimate — each group’s difference weighted by the trial’s overall group shares — is the one that compares like with like.
Do not read it as a finding about the treatment. Nothing in the reversal is a property of the treatment: it is a property of who ended up where. The temptation is to write that the treatment “works in subgroups but not overall”, which sounds like a mechanistic claim and is a statement about the randomisation’s luck.
And do not report the reversal as a result at all if the trial is small. At eighty units the pattern arrives on one trial in thirty from a population with no interaction and no confounding. A finding with a one-in-thirty chance of appearing from nothing is a finding that needs a second study rather than a paragraph of interpretation.
Still open: the variables that were not stratified on
Stratifying removes the reversal for the variable stratified on. A trial has one or two of those and a patient has thousands, and the reversal is available on every variable that was not.
What is not known here is how that scales. Measuring a hundred baseline variables in a randomised trial and tabulating the treatment effect within each gives a hundred chances for the pattern to appear, and the rate above is the rate for one variable chosen in advance. Whether the hundred behave like a hundred independent draws — they do not, since the variables are correlated with each other — and what the effective number is, is the same question a family of analyses raises in a different setting.
The practical form of the question is sharper than the theoretical one. A trial’s subgroup table is scanned for exactly this pattern, the scanning is not pre-specified, and the rate that matters is therefore the rate of some subgroup reversing rather than of a nominated one. That number is larger than 3.40% by an amount nobody on this page has counted.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A count that has to be estimated — both name allocation, blocking, covariate balance, randomisation
- A dictionary that is a product — both name allocation, blocking, covariate balance, randomisation
- How many subjects — both name allocation, blocking, sample size
- Randomising towards the winner — both name allocation, blocking, randomisation
- The variance removed before the data — both name allocation, blocking, randomisation
- Three arms and three scores — both name allocation, covariate balance, randomisation
Named objects
A flat tag is an object no other essay names yet.
AggregationAllocationBlockingCovariate balanceRandomisationSample sizeSampling variationSimpson's paradox