The change that is not confounding
Worth reading first: Simpson's reversal is a region, not a table.
A stratified analysis and a pooled analysis of the same data give different numbers. The standard reading is that the stratifying variable was a confounder, and which of the two to believe then depends on the causal structure.
Here is a table where that reading is simply wrong.
The allocation is a fair coin in every stratum. The strata are not associated with the arm, so they cannot confound anything, and the marginal and stratified comparisons are estimates of the same population. The odds ratio still moves, by 28% of the distance from 2.5 to no effect at all.
The odds ratio is not a weighted average of odds ratios
The risk difference and the risk ratio are. That is the whole of it, stated as arithmetic.
The pooled treated risk is a weighted average of the strata’s treated risks, and the pooled control risk is the same average of their control risks:
Subtracting gives the pooled risk difference as the weighted average of the strata’s own risk differences, immediately. Dividing gives the pooled risk ratio as an average of the strata’s risk ratios weighted by their control risks — which is a weighted average too, with different weights.
Neither move works for the odds ratio, because the odds ratio is a ratio of odds, and the odds of the average is not the average of the odds. p/(1 − p) is convex in p, so averaging risks and then taking odds gives something smaller than averaging odds, and the two averages in the numerator and denominator are shrunk by different amounts. The result is a pooled odds ratio strictly between 1 and the common conditional one, whenever the strata differ in baseline risk.
On this table: five strata at risks of 5%, 15%, 35%, 60% and 85%; a conditional odds ratio of 2.5 in each; pooled risks of 40.00% and 54.39%; a pooled odds ratio of 1.789.
The same table, read three ways
The three measures on the table above, stratum by stratum:
| stratum | baseline risk | odds ratio | risk ratio | risk difference |
|---|---|---|---|---|
| 1 | 5% | 2.500 | 2.326 | 0.066 |
| 2 | 15% | 2.500 | 2.041 | 0.156 |
| 3 | 35% | 2.500 | 1.639 | 0.224 |
| 4 | 60% | 2.500 | 1.316 | 0.189 |
| 5 | 85% | 2.500 | 1.099 | 0.084 |
| pooled | 40% | 1.789 | 1.360 | 0.144 |
The odds ratio column is constant and its pooled value is not in it. The other two columns vary wildly and their pooled values sit inside their own range, which is what a weighted average has to do.
Turn the construction around and hold the risk difference constant instead — 0.08 in every stratum, same five baseline risks, same coin:
| stratum | baseline risk | risk difference | odds ratio |
|---|---|---|---|
| 1 | 5% | 0.080 | 2.84 |
| 2 | 15% | 0.080 | 1.69 |
| 3 | 35% | 0.080 | 1.40 |
| 4 | 60% | 0.080 | 1.42 |
| 5 | 85% | 0.080 | 2.34 |
| pooled | 40% | 0.080 | 1.385 |
The pooled risk difference is 0.080000 — exactly, to every digit the arithmetic carries, because it is an identity rather than an approximation. And the odds ratios now range from 1.40 to 2.84 with no heterogeneity of any kind in the data, because a constant risk difference is a varying odds ratio by construction.
A constant risk ratio behaves the same way: with a risk ratio of 1.500 in four strata, the pooled risk ratio is 1.500000 and the strata’s odds ratios run from 1.54 to 6.00.
So “the effect is the same in every stratum” is not a statement about the data. It is a statement about the data and a choice of scale together, and the three choices give three different answers to whether there is heterogeneity, on identical counts.
How much it moves, and what decides it
The quantity that decides the size of the shift is how different the strata are in baseline risk.
With every stratum at the same risk there is nothing to collapse and the marginal odds ratio is 2.500, exactly. Pulling them apart by a half-width of 0.1 costs 3.0% of the distance to the null; 0.2 costs 11.8%; 0.3 costs 25.9%; 0.4 costs 44.5%.
Two consequences follow, and the second is the one that damages published work.
A more strongly prognostic covariate causes a larger shift. The better a variable predicts the outcome, the further apart it pulls the strata’s baseline risks, and the further the pooled odds ratio sits from the conditional one. So adjusting for the most useful covariate produces the largest apparent change — and a reader trained to read a large change as strong confounding reads it exactly backwards.
In a randomised trial, the adjusted odds ratio is systematically further from 1 than the unadjusted one. That is not an artefact and not a bias: both are consistent estimates, of different estimands. The unadjusted one estimates the population-averaged effect and the adjusted one estimates the effect conditional on the covariate, and non-collapsibility is precisely the statement that these differ. A trial reporting both and describing the gap as “adjustment for baseline imbalance” has attributed to imbalance something that would be there in a perfectly balanced trial.
What the gap is, in people
An odds ratio is hard to hold, so it is worth converting the whole construction into the counts it describes.
Inside the strata the treatment moves the risk from 5% to 11.6%, from 15% to 30.6%, from 35% to 57.4%, from 60% to 78.9% and from 85% to 93.4%. Those five pairs all have an odds ratio of 2.5 and risk differences from 6.6 to 22.4 points, which is the first thing a reader should take from the table: an odds ratio of 2.5 does not name an amount of benefit until a baseline risk is supplied.
Pooled, the treatment moves the risk from 40.00% to 54.39% — 14.4 points, or about one person in seven. That is the marginal effect and it is correct. The pooled odds ratio of 1.789 is also correct, and it is the number that will be quoted, and 1.789 against 2.5 will be read as a substantial difference between the adjusted and unadjusted analyses.
Nothing in the data changed between those two sentences. The risk difference is the same 14.4 points whichever way the table is cut, and the odds ratio is two different numbers because the odds ratio is measuring a shift on a scale whose spacing depends on where the shift starts. That is the same complaint a summary statistic usually invites, arriving from an unusual direction: here the summary is not hiding a feature of the data, it is reporting a feature of the scale.
A forest plot can be heterogeneous with no heterogeneity in it
The practical consequence for evidence synthesis is direct and is not a small correction.
Three studies of the same treatment, with the same conditional odds ratio of 2.5 in every stratum of every one, recruiting populations whose baseline risks are spread by half-widths of 0.1, 0.25 and 0.4 around the same centre. Their marginal odds ratios are 2.455, 2.226 and 1.832 — a range of 0.293 on the log-odds scale, and a ratio of 1.34 between the extremes.
A meta-analysis of those three reports heterogeneity. There is none: every individual in all three studies has the same conditional odds ratio, and the treatment behaves identically in all of them. What differs is the case mix, and the odds ratio is not invariant to case mix. The heterogeneity statistic is real, its p-value is honest, and it is measuring something other than what it is read as.
The same construction on the risk-difference scale produces three identical estimates and no heterogeneity at all, because the risk difference is collapsible. So the choice of scale decides whether the evidence base looks consistent, and the decision is usually made by convention rather than by argument — logistic regression is what was fitted, so an odds ratio is what is pooled.
The useful habit that follows is small. When a forest plot of odds ratios shows heterogeneity, plot the control-arm risks beside it. If the estimates line up against baseline risk, the heterogeneity is at least partly the scale, and the same analysis on a risk difference will say how much of it.
Why the intuition from linear regression is wrong here
Most readers arrive at this with an instinct trained on ordinary regression, and the instinct is correct there and false here.
In a linear model, omitting a covariate that is independent of the treatment changes the treatment’s coefficient by exactly nothing. That is the omitted-variable-bias formula: the bias is the omitted variable’s coefficient times the regression of the omitted variable on the treatment, and the second factor is zero. It is why a randomised trial analysed with ordinary regression gives the same treatment effect adjusted and unadjusted, up to noise.
In a logistic model it does not. Omitting an independent, prognostic covariate moves the treatment’s coefficient towards zero, and the amount is the shrinkage on this page. There is no bias in either estimate; the two coefficients are estimating different quantities, and only in the linear case do those quantities coincide.
The reason they coincide in the linear case is that the identity link is the one under which a conditional effect and a marginal effect are the same thing — averaging a linear function is the same as applying it to the average. Every non-linear link breaks that, and logistic, log and complementary-log-log links are all non-linear.
So “adjusting for a covariate that is balanced should not change the estimate” is a true statement about linear models being applied to a model where it is false. A reader carrying it will misdiagnose non-collapsibility as imbalance every time they meet it, and the misdiagnosis is invisible because the two look identical in the output.
Which of the two is right
Both, and the question as usually put has no answer, which is a different situation from the confounder-or-mediator case where one of them is right and the data cannot say which.
The conditional odds ratio answers: for a person in this stratum, what does the treatment multiply their odds by? The marginal odds ratio answers: if this whole population were treated rather than untreated, what would happen to the odds of the outcome in it? Those are different questions about the same treatment, and both have honest answers on the same table. The first is what a clinician wants and the second is what a health service wants.
So the check that distinguishes non-collapsibility from confounding is not a comparison of the two numbers, because that comparison cannot distinguish them — in the same way that counting an interval’s coverage is the only thing that separates an interval that works from one that merely claims to. It is a check on the allocation, and it is the same one-line check the randomised reversal needs: is the stratifying variable associated with the arm? If it is not, the gap is non-collapsibility and nothing else. If it is, the gap is both, in amounts nothing separates.
That last clause is the uncomfortable part. In an observational study the two effects are superimposed and no arithmetic on the table pulls them apart, so a stratified-versus-pooled comparison of odds ratios is never a clean measurement of confounding. The clean measurement uses a collapsible scale — compute the shift in the risk difference instead, where any gap is confounding by construction — and then converts back if an odds ratio is wanted.
Why anyone uses a measure with this property
The obvious conclusion is to abandon the odds ratio, and it is worth saying what that would cost, because the three measures cannot be made to share their virtues.
The odds ratio is invariant to which outcome is called the event. Recoding “survived” as “died” replaces the pooled odds ratio 1.789 by 0.5590, which is exactly its reciprocal. The risk ratio does not do this: the same recoding takes a pooled risk ratio of 1.360 to 0.7601, against a reciprocal of 0.7354. A measure whose value depends on an arbitrary labelling decision is a measure that can be chosen to look better.
The odds ratio is estimable from a case-control study and the other two are not. Sampling on the outcome destroys the risks and preserves their odds ratio, which is the reason case-control designs exist at all. A field that studies rare outcomes has no alternative, and the odds form is also what makes a screening calculation portable across prevalences — the same multiplicative structure, wanted there and resented here.
The odds ratio is what logistic regression produces, and logistic regression is what a model with several covariates and a binary outcome is fitted by, because it is the model whose fitted values cannot leave (0, 1). A risk-difference model can predict a negative risk.
So the three properties — collapsibility, invariance to outcome coding, and estimability under outcome-dependent sampling — are not available together, and every field has settled on a different two. That is the honest description, and it is a better basis for a choice than the reflex that non-collapsibility is a defect.
Two routes, and the reading that does not survive
The pooled odds ratio here is computed twice by arithmetic that shares nothing.
From the strata: average the risks with the strata’s weights, take the odds of each average, divide. Closed form, no random numbers, 1.7891.
From a cohort: draw four hundred thousand people, one stratum each, allocate by a coin, generate the outcome from that person’s own stratum risk, and read the pooled odds ratio off the two-by-two table of counts. 1.7828.
The two agree to three parts in a thousand, which is the counting error of four hundred thousand draws. Neither uses the other’s arithmetic: the first never samples and the second never touches the stratum structure after generating a person.
The reading this page refuses is a stratified and a marginal odds ratio disagreeing, taken as evidence of confounding. A reader who attributes the gap to confounding is claiming that a randomised table’s marginal odds ratio equals its conditional value, and on this table it does not — 1.789 against 2.5. If it did, either the allocation would have stopped being a coin or the odds ratio would have stopped being an odds ratio, and both would make every other number here mean something else.
The companion statement is the positive one, and it is an exact equality rather than a tolerance: with a constant risk difference, the pooled risk difference equals it to within 10⁻¹². That identity is what makes the odds ratio’s behaviour a property of the odds ratio rather than of stratification, and a tolerance on it would let a genuine arithmetic error through as a rounding.
Still open: the hazard ratio, which has the property and hides it better
Everything here is about a binary outcome at a fixed time. The same phenomenon appears in survival analysis, in a form that is harder to see and harder to argue about.
A hazard ratio is non-collapsible for the same reason an odds ratio is, and it has an additional difficulty the odds ratio does not: the population at risk at time t is a selected population — the ones who have not yet had the event — and the selection is stronger in the arm with the higher hazard. So even with a constant individual hazard ratio, the ratio of the observed hazards drifts towards 1 as follow-up lengthens, and two trials of the same treatment with different follow-up report different hazard ratios with nothing wrong in either.
Still open: telling non-collapsibility from depletion apart
What is not measured here is how large that drift is at the follow-up lengths trials actually use, whether it is comparable to the shrinkage on this page or an order of magnitude smaller, and whether it can be distinguished in practice from a genuinely time-varying effect — which is the interpretation it invariably receives, and which requires the hazard ratio to be a property of the treatment rather than of the treatment and the follow-up together.
The selection effect has a name in survival analysis and it is worth carrying across: the population that survives is not the population that started, and every quantity computed on the survivors inherits that. Non-collapsibility and depletion of susceptibles arrive at the same place from different directions, so a drifting hazard ratio has two candidate explanations before any biology is invoked, and separating them is a measurement nobody has made here.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A covariate with no levels — both name randomisation, stratification
- A threshold in the tail — both name randomisation, stratification
- Balanced on the wrong function — both name randomisation, stratification
- Balancing more than one number — both name randomisation, stratification
- Balancing what is known in advance — both name confounding, randomisation
- Randomising towards the winner — both name confounding, randomisation
Named objects
A flat tag is an object no other essay names yet.
AggregationCollapsibilityConfoundingOdds ratioRandomisationRisk differenceRisk ratioStratification