When the stratified answer and the pooled one disagree

Conditioning on what the treatment caused

When the grouping variable lies on the path from treatment to outcome, the stratified answer is the direct effect and the aggregate is the total effect. Both are correct. Over 15% of a sweep of the indirect path they have opposite signs, and no arithmetic on the table says which question was being asked.

Worth reading first: Simpson's reversal is a region, not a table.

The reversal’s causal question has two branches and only one of them has been priced. If the grouping variable is a confounder, the stratified comparison is the one to believe and the aggregate is contaminated by the allocation. If it is a mediator — something the treatment itself causes — the essay says conditioning on it “removes part of the treatment’s effect”, and stops.

That second branch is not a contamination. It is a different estimand, and it has a name, a value, and a sign that can disagree with the aggregate’s.

What stratifying on a mediator estimates, and what it does notThe direct effect stays at -0.109 throughout, because it is defined with the mediator held fixed. The total effect runs from -0.218 to 0.027 and crosses zero at 1.5. On 15% of the sweep the two have opposite signs, and both are correct answers.-0.200-0.1000-2-1012the treatment's effect on the mediator, in log oddseffect on the outcome, as a risk differencethe total effect changes signstratified on the mediator: the direct effectcomputed from the generating parameters15% of the sweep disagrees in sign
Fig. 1 A treatment with a direct effect on the outcome and an indirect one through a mediator. The dashed line is what stratifying on the mediator estimates; the solid line is what the aggregate estimates. The slider is how strongly the mediator affects the outcome.

Two quantities, two definitions

The construction is three arrows. The treatment moves the mediator; the mediator moves the outcome; and the treatment moves the outcome directly as well.

TMY,TYT \longrightarrow M \longrightarrow Y, \qquad T \longrightarrow Y

Everything below is computed from the generating parameters rather than simulated, so both effects are exact.

The total effect is what happens to the outcome when the treatment is given, with the mediator left to do whatever the treatment makes it do. It is the difference between the outcome’s probability under treatment and under control, averaging over the mediator’s own distribution in each arm — and those distributions differ, because the treatment changed them. Here the mediator is present in 42.6% of the control arm and 71.1% of the treated arm.

The direct effect is what happens when the treatment is given and the mediator is held where it was. That is what a stratified analysis estimates: inside each level of the mediator, compare the arms, then average with the mediator’s distribution held fixed.

On the default setting the two are −0.017 and −0.109. The treatment is harmful, and it is six times more harmful when the mediator is held still than when it is allowed to move, because moving the mediator is a benefit that partly cancels the harm.

Neither number is wrong

That is the whole difficulty, and it is a genuinely different situation from the confounded case.

With a confounder, one of the two comparisons is between differently-composed populations and the other is not. There is a right answer and the data cannot identify it, which is uncomfortable but familiar.

With a mediator, both comparisons are between comparable populations. Nothing is imbalanced; nothing is estimated with bias. The total effect answers should this treatment be given? and the direct effect answers how much of what the treatment does travels by a route other than the mediator? Those are different questions that a table cannot tell apart, and a stratified analysis answers the second whether or not it was asked.

The asymmetry between the two cases is worth stating plainly, because it changes what a reader should do. In the confounded case, an argument about the causal structure is an argument about which of two numbers is the estimate — one of them is contaminated and the disagreement is a problem to be resolved. In the mediated case there is nothing to resolve: the disagreement is information, both numbers are worth reporting, and the only error available is to report one of them under the other’s name.

So the same table supports a dispute in one case and a decomposition in the other, and which it is depends on a fact about timing that is not in the table.

What stratifying on a mediator estimates, and what it does not. The direct effect stays at -0.105 throughout, because it is defined with the mediator held fixed. The total effect runs from -0.105 to -0.105 and crosses zero at no point drawn. On 0% of the sweep the two have opposite signs, and both are correct answers.
Fig. 2 The same sweep with the mediator disconnected from the outcome. The two curves coincide everywhere: with nothing travelling by the indirect route, there is only one effect and stratifying costs nothing.

That figure is the control. With the mediator’s own effect on the outcome switched off, the total and direct effects agree at every setting of the indirect path, because a path with a zero in it carries nothing. The disagreement in the first figure is therefore the product of the two arrows and not an artefact of either.

The direct effect does not move, and that is what makes it the direct effect

Look again at the dashed line in the hero. It is flat: −0.109 at every one of the forty-one settings of the treatment’s effect on the mediator, to six decimal places.

That is not a coincidence, it is the definition. The direct effect holds the mediator fixed, so how strongly the treatment would have moved it is irrelevant — the arrow that is being swept is the one the definition has switched off.

The total effect does move, from −0.218 at the far negative end of the sweep to +0.027 at the far positive end, and it crosses zero at an indirect path strength of 1.5. So over the right-hand part of the sweep the treatment helps overall and harms with the mediator held fixed, which is the reversal in its causal form: two correct estimates with opposite signs from the same trial.

Over 15% of the swept range the two disagree in sign. Strengthening the mediator’s own effect on the outcome widens that: at an effect of 2.6 the two disagree over 34% of the range.

What stratifying on a mediator estimates, and what it does not. The direct effect stays at -0.087 throughout, because it is defined with the mediator held fixed. The total effect runs from -0.277 to 0.151 and crosses zero at 0.7. On 34% of the sweep the two have opposite signs, and both are correct answers.
Fig. 3 A mediator with a strong effect of its own. A third of the sweep now has the two estimates pointing in opposite directions, and the total effect crosses zero much earlier.

The collider that adjustment creates

There is a second, worse thing that adjusting for a post-treatment variable does, and it does not appear in the arithmetic above because the arithmetic above assumes the mediator is unconfounded.

Conditioning on a variable that two arrows point into opens a path between them. The mediator has an arrow from the treatment; if it also has an arrow from anything that affects the outcome — a prognostic factor, a measurement artefact, an unmeasured susceptibility — then conditioning on the mediator makes the treatment and that factor associated within each stratum, where they were independent before.

The consequence is that the stratified estimate stops being the direct effect and becomes the direct effect plus a bias, and the bias has no sign that can be argued in advance. Unlike confounding, which is removed by adjustment, this is created by adjustment: an analysis that adjusts for nothing has none of it, and every post-treatment variable added brings some.

That gives post-treatment adjustment a shape worth stating plainly. It answers a question nobody asked, and in doing so it opens a path nobody wanted, and both effects are invisible in the output. A model that adjusts for everything measured is not a cautious model; it is a model whose estimand is decided by which variables happened to be collected. The same structure, drawn on its own, is why a variable’s association with an outcome can change sign on conditioning with no causal relation between them at all.

What goes wrong when the second is reported as the first

The failure this describes is not exotic. It is the routine practice of adjusting for every measured covariate.

A trial measures a set of baseline variables and a set of variables collected during follow-up. Adjusting for the baseline ones is uncontroversial — they cannot have been caused by a treatment given later. Adjusting for the follow-up ones is the error, and it is easy to make because the variables arrive in the same dataset, look like the same kind of thing, and improve the model’s fit.

The consequence is stated exactly by the figure. A trial that adjusts for a post-treatment variable reports the direct effect, labels it the treatment effect, and can report it with the opposite sign to the effect of actually giving the treatment. The reported number is not wrong; the label is.

The tell is the timing and nothing else. A variable measured after allocation is a candidate mediator, and no property of the numbers distinguishes a mediator from a confounder — the two produce identical tables, which is the finding the treatment-and-control comparison ends on. The only thing that separates them is knowledge of what happened when, and that knowledge is not in the data.

Two groups, one treatment, and both readings of the same numbers. The treatment wins in group A (93.0% against 87.0%) and in group B (73.0% against 69.0%), and loses overall (78.1% against 82.6%). Nothing here is a trick; the allocation differs between the groups.
Fig. 4 The table both cases produce. The four rates and the allocation are all that a reader is given, and they are the same whether the grouping variable came before the treatment or after it.

Which level the mediator is held at changes the answer

“Held fixed” is not one instruction, and the figure’s dashed line quietly makes a choice.

Inside the two levels of the mediator the treatment’s effect is −0.1046 where the mediator is absent and −0.1155 where it is present. Those differ, because the model is logistic and a logistic model has an interaction on the risk scale even when it has none on the log-odds scale — the same non-linearity that moves a marginal odds ratio away from its conditional value.

So there is a controlled direct effect for each level of the mediator, and they are not equal. The number drawn is a weighted average of the two, with the weights being the mediator’s distribution in the control arm: 57.4% absent and 42.6% present, giving −0.109.

That weighting is a decision, and a different reasonable decision gives a different number. Weighting by the treated arm’s distribution gives a different average; using the population’s overall distribution gives a third; reporting the two controlled effects separately gives neither.

The consequence is that the disagreement between the stratified and the pooled estimate is not one number even after the question has been settled. It is a number plus a convention, and the convention is rarely stated. That is a weaker position than the confounded case, where the target is at least unambiguous once the causal structure is known — here the causal structure is known, and the estimand still needs a choice.

The practical version is short. Report the two controlled effects rather than their average, when there are two of them and they differ. An average of two numbers that a reader could have been given directly is a summary with nothing to recommend it, and the weights are precisely the part of the calculation nobody will check.

Adjusting for a mediator is not partial adjustment

A tempting halfway position is that adjusting for a mediator is over-adjustment, removing “too much”, so the truth lies between the two estimates. It does not, and the figure says why.

The direct effect is flat across the sweep. If it were a contaminated version of the total effect, contaminated by an amount that depends on the indirect path, it would move as that path moves. It does not move at all, because it is a clean estimate of something else.

So the two numbers do not bracket a third. They are two points, each exactly right about its own question, and a reader wanting a compromise between them is wanting an answer to a question they have not asked. The thing that is between them, and is a genuine quantity, is the indirect effect — the total minus the direct, which at the default setting is 0.092 and is the benefit that travels through the mediator.

That decomposition is worth having, and it is worth being honest about what it costs. It is only valid when the mediator itself is not confounded with the outcome — when nothing causes both the mediator and the outcome. Randomising the treatment guarantees nothing about that, because the mediator was not randomised: it was chosen by the world, in exactly the way an observational comparison’s allocation is. A randomised trial analysed by mediation is an observational study of its own mediator.

What is claimed here, and what is not

Two statements carry the figure and they are of different kinds, which is deliberate.

The direct effect is constant across the sweep, to six decimal places at every one of the forty-one settings. That is an exact property of the definition rather than a measurement, so it is stated exactly: a direct effect averaged over the mediator’s treated distribution instead of a fixed one would tilt the line, and a tolerance would let a small tilt through as noise.

The total effect does move. Stating only the first would be satisfied by a version in which both quantities were the direct effect, which is the plausible slip — the two come from the same three-arrow model and differ only in which distribution the mediator is averaged over.

The reading that does not survive is a mediator adjusted for as though it were a confounder. That reading requires the stratified and unstratified estimates to be equal, which is what a reader assumes when they call a post-treatment adjustment a correction, and here they are −0.109 and −0.017 with the indirect path switched on. They are equal when the path is switched off, which is the control figure above, and that is what makes the comparison a test rather than a statement that could not come out wrong.

One odds ratio in every stratum, and a different one marginally. Five strata with baseline risks from 5% to 85%, a treatment allocated by a coin in each, and a conditional odds ratio of exactly 2.5 throughout. The marginal odds ratio is 1.789. Nothing is confounded; the odds ratio is simply not a weighted average of odds ratios.
Fig. 5 The third reason a stratified and a pooled estimate differ, for completeness: no confounding, no mediator, and an odds ratio that moves because an odds ratio is not a weighted average of odds ratios.

How it looks in a trial that reports both

The abstract version above becomes concrete in a shape that appears often enough to recognise.

A trial randomises a treatment, measures an intermediate outcome at six weeks and the outcome of interest at a year. The pre-specified analysis compares the arms on the year outcome and finds a small harm. A secondary analysis adjusts for the six-week measurement — reasonably, on the grounds that it is strongly prognostic and improves precision — and finds a larger harm, or a benefit.

Everything in that paragraph is a correct calculation. What has happened is that the secondary analysis has estimated the effect of the treatment other than through the six-week measurement, and if the treatment’s benefit runs through exactly that channel, the secondary analysis has subtracted the benefit and reported the remainder.

The tell in the write-up is usually a sentence explaining that the adjusted analysis is “more precise” or “accounts for a baseline difference”. The first is often true and irrelevant — a more precise estimate of a different quantity is not an improvement — and the second is a category error, since a variable measured at six weeks is not a baseline.

Two questions separate the cases and neither needs the data. When was it measured, relative to allocation? And could the treatment have changed it? A yes to the second makes every adjustment for it a mediation analysis, whatever it is called, and a baseline measured twice is the case where the answer is genuinely ambiguous and the two analyses genuinely disagree.

Three reasons a stratified analysis disagrees with a pooled one

This ends with a list, because the three causes are routinely conflated and the response to each is different.

Confounding. The stratifying variable is associated with both the arm and the outcome. The stratified estimate is the one to believe, the disagreement is diagnosable from the table’s margins, and balancing the variable at allocation prevents it.

Non-collapsibility. The measure is an odds ratio or a hazard ratio and the strata differ in baseline risk. Both estimates are correct, of different estimands, and the disagreement appears in a perfectly randomised trial. Computing the same comparison on a risk difference makes it vanish.

Mediation. The stratifying variable was caused by the treatment. Both estimates are correct, of different estimands, and which one is wanted depends on the question. The disagreement is diagnosable only from the timing of the measurement, which is not in the numbers at all.

All three produce the same shape in a table: two numbers that disagree, with nothing obviously wrong in either. Only the first is an error, and it is the one all three tend to be read as.

Still open: how much of a published disagreement is which

The three causes above are separable in a constructed example, where the timing is known, the allocation is known and the scale is chosen. They are not separable in a paper.

What is not known here is the mixture in practice. A survey of published analyses reporting both an adjusted and an unadjusted effect could, in principle, be sorted: the measurement timing is usually reported, the effect measure always is, and the arm-covariate association is usually tabulated. The share of disagreements attributable to non-collapsibility alone — which requires no error, no mediator and no imbalance — is a number nobody appears to have.

It is worth having because it decides what the default reading should be. If most published adjusted-versus-unadjusted gaps in randomised trials are non-collapsibility, then the sentence “the adjusted analysis corrects for chance imbalance” is wrong most of the time it is written, and the correct default reading of such a gap is nothing happened.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AggregationCausal diagramConfoundingDirect effectMediatorSimpson's paradoxStratificationTotal effect