The variable the treatment caused
Worth reading first: One arithmetic, three decisions.
The rule of thumb that survives this whole ladder is about timing: do not adjust for anything the treatment could have caused. It is worth knowing what happens when it is broken, and the answer has two parts, of which only the first is usually told.
The first part is that adjusting for a mediator changes the question. The estimate stops being the total effect and becomes the direct one — here 1.1300 becomes 0.5000 — and the direct effect is a real quantity that somebody might have wanted. Told on its own, that reads as a warning about ambiguity rather than about error.
The second part is that it only changes the question when nothing unmeasured causes both the mediator and the outcome, and something usually does. Give the mediator an unmeasured cause it shares with the outcome and the adjusted regression returns 0.0500: a tenth of the direct effect, a twenty-third of the total, and not an estimate of anything anybody asked for. The unadjusted regression on the same data returns 1.1300 exactly.
Set the unmeasured cause to nothing and the same regression returns 0.5000 to fifteen places, so the failure is the confounding and not the timing. The timing is what makes the confounding unavoidable.
The structure, and what each regression is estimating
Three equations, with never measured:
The treatment causes the covariate at and the outcome directly at ; the covariate causes the outcome at ; and one unmeasured variable reaches the covariate at and the outcome at , both 1 in the base case. Every noise term is standard normal and independent.
Two quantities are on offer. The total effect is : the treatment reaches the outcome directly and again through the covariate, and moving the treatment moves both routes. The direct effect is : what the treatment would do with the covariate held where it was.
The regression that leaves the covariate out returns the total effect exactly, and the reason is worth stating because it is not the general case. The unmeasured variable does not reach the treatment — it causes the covariate and the outcome, and the treatment is upstream of everything — so there is no back-door route to block and the marginal regression of the outcome on the treatment recovers the sum of every causal route between them. Timing buys that: a treatment nothing causes is a treatment with no confounding to remove.
The regression that puts the covariate in is where the trouble is, and it is trouble of a specific shape. Holding the covariate fixed does two things at once. It blocks the causal route through the covariate, which is the intended change of question. And it conditions on a common effect of the treatment and the unmeasured variable, which opens a route that was closed: the covariate is caused by and by , so within a stratum of the two become dependent, and reaches the outcome directly.
What the adjustment would need in order to work is therefore a specific and rather strong absence: no unmeasured variable causing both the covariate and the outcome. That is a stronger requirement than the one the unadjusted estimate needs, which is only that nothing unmeasured causes both the treatment and the outcome — and the two are routinely confused, because both are stated as “no unmeasured confounding” and the words do not say confounding of what. A randomised treatment settles the first for free and does nothing at all for the second: randomisation deletes the arrows into the treatment, and the arrows this failure runs on go into the covariate and into the outcome. What randomisation buys is exactly one of the two, which is why a mediation analysis inside a randomised trial is an observational study with an experiment attached to its first step.
Neither effect, and no warning that it is neither
The adjusted coefficient has a closed form:
It sits −1.0800 below the total effect and −0.4500 below the direct one. As a share it is 0.100 of the direct effect and 0.044 of the total.
Those two gaps decompose, and the decomposition says which part is the change of question and which is the error. Against the direct effect the whole shortfall is , and every bit of it is bias. Against the total effect the shortfall is plus that same term: −0.6300 of intended change and −0.4500 of error, summing to −1.0800. So of the movement between the two regressions, rather more than half is the deliberate deletion of the mediated route and rather less than half is a route opened by accident, and no feature of the output separates the two contributions.
Two of those three numbers are quantities a study might have wanted and the third is what it gets. Nothing printed by the regression distinguishes them. The coefficient has a standard error, the standard error is correct for the model that was fitted, the residuals behave, and a reader who is told “adjusted for the intermediate variable, so this is the direct effect” has been told something false about a number that is 90% of the way to zero from the quantity named.
What makes the failure legible is that the two regressions disagree so much: 1.1300 against 0.0500. In practice that disagreement is precisely what gets read as evidence for the adjustment — the covariate obviously mattered, the estimate moved a long way, and something is being controlled for. The size of the movement carries no information about its direction, which is the arithmetic that is right in one world of three arriving with a fourth structure that none of the three covers.
The mediator’s own effect on the outcome is not in the answer
Read the closed form again and notice what is absent from it. The coefficient — the covariate’s own effect on the outcome, the edge whose route the adjustment is meant to be blocking — does not appear.
That is a cross-setting check rather than an observation about one expression, and it was run as one. At the adjusted coefficient is 0.0500 and the total effect is 0.6800. At it is 0.0500 against 1.1300. At , 0.0500 against 1.7600. At , 0.0500 against 2.3000. The total effect more than triples across that range, the adjusted regression’s answer does not move at all, and their ratio therefore runs from 0.074 down to 0.022.
The reason is that the adjusted coefficient is the coefficient on the treatment given the covariate, and once the covariate is in the model its own effect on the outcome is absorbed into its own coefficient. What is left on the treatment is the direct edge minus whatever the conditioning has opened, and what has been opened is the route through , which does not pass along at all.
The practical consequence runs against intuition. A covariate that carries almost none of the effect — a weak mediator, small — is exactly as dangerous to adjust for as one that carries most of it, provided it shares an unmeasured cause with the outcome. The size of the mediated share is a reason to care about the change of question and no reason at all to care about the bias, and the two are usually explained together as though they were one concern.
The worst unmeasured cause is not the strongest one
Sweep the strength of the unmeasured cause’s path into the covariate and the damage does not run the way a sensitivity argument would assume.
At the adjusted coefficient is 0.5000, exactly the direct effect: no unmeasured cause, no opened route, and the adjustment does what it is advertised to do. At it is 0.2882, at it is 0.1400, and at it is 0.0500. Then it turns. At it is 0.0846, at it is 0.1400 again, and at it is 0.2300.
The maximum damage is at — where the unmeasured cause contributes exactly as much to the covariate as the covariate’s own noise — and it is , from the same arithmetic–geometric mean inequality that bounds the covariate measured before the treatment. The two structures are different and the shape of their failure is the same: the bias is a ratio whose numerator is linear in a coefficient and whose denominator is quadratic in it, so it has an interior maximum rather than a limit.
The mechanism behind the turn is the same too. Past the covariate is mostly the unmeasured variable’s own contribution and carries progressively less about the treatment, so conditioning on it explains away while removing less and less of the treatment’s variation. A covariate that is nearly a pure measurement of the unmeasured confounder is nearly a safe thing to adjust for — which is a strange sentence and is what the arithmetic says.
So the sensitivity question is not “how strong could the unmeasured cause be”, and a reader who assumes the worst by taking large is assuming something that is not the worst. The worst is at a specific middling value, and it is computable.
Two routes, and what would have made them agree for the wrong reason
The forms above are algebra on a population covariance; beside them are five hundred least-squares fits of six hundred rows apiece, on rows simulated one variable at a time from the equations in causal order.
The adjusted coefficient counts at 0.0479 with a standard error of 0.00256 against a derived 0.0500 — under a standard error out. The unadjusted counts at 1.1277 with a standard error of 0.00368 against an exact 1.1300, which is 0.6 of a standard error. Both are the ordinary kind of agreement, and neither is the check that matters.
The check that matters is the one at . If the closed form had been derived on the wrong adjustment set — if the algebra had, say, conditioned on the wrong column — it would still have produced a number near 0.05 for some parameters and could have been talked into agreement. What it could not have survived is returning exactly the direct effect when the unmeasured cause is switched off, because that is a different structural claim about a different world, and the same expression has to satisfy both. It does: 0.5000, to fifteen places, from the same routine.
That is the same discipline as two routes to every number with a third route added — a limiting case whose answer is known independently of either. A closed form that matches its simulation and also collapses to a known value at a boundary is much harder to be accidentally right than one that only does the first.
A weaker unmeasured cause, and a stronger one that reverses the sign
Two more settings, because a single value of is a single value.
Weaker. At the adjusted coefficient is 0.1897 in closed form and 0.1879 counted, with a standard error of 0.00312. That is a third of the direct effect rather than a tenth — still not the direct effect, still reported as one, and now close enough to a plausible answer that nothing about it invites suspicion. A wildly wrong number is a number somebody investigates; 0.19 against a true direct effect of 0.50 is a number somebody publishes.
Stronger. Double the unmeasured cause’s path to the outcome, to , and the adjusted coefficient is −0.4000 in closed form and −0.4043 counted, on a standard error of 0.00362. The total effect is 1.1300 and the direct effect is 0.5000, both firmly positive; the adjusted regression reports a substantial negative effect. Its curve runs 0.5000 at , −0.2200 at 0.5, −0.4000 at 1, back to −0.2200 at 2 and −0.0400 at 3, and its worst case is −0.9000.
A sign reversal is the version of this failure that is visible, and it is worth noticing that it is the less dangerous case. An estimate whose sign contradicts everything known about the subject gets challenged. The reversal here needs a latent path twice the size of the base case; the merely-wrong estimates need nothing unusual at all.
The shape of that curve is the same shape as the base case scaled: it dips, turns at , and comes back towards zero, and the whole of it has been moved down by the doubled path to the outcome. The bias is linear in , which appears only in the numerator, and the turning point is at whatever is — so a sensitivity analysis over the strength of an unmeasured cause can be run in one dimension while the other is pinned to a value the data does supply, since is the residual spread of the mediator around its own regression on the treatment, and that is estimable.
Where this does not hold
The unmeasured variable does not reach the treatment. That is what makes the unadjusted regression exact, and it is the assumption a randomised treatment buys and an observed one does not. In an observational study both failures are live at once: leaving the covariate out is off by the ordinary confounding term and putting it in is off by this one, and neither route is clean. Nothing here measures the pair together, which is the case most real analyses of a mediator are actually in.
One mediator, measured perfectly. A chain of them, or one observed with error, is a different arithmetic. A noisy mediator is particularly worth naming: it fails to block the route it was included to block and opens the collider route, so it can carry both failures at once, and the closed form here does not cover it.
Everything is linear, and the direct effect is the same for every unit. Where the treatment’s effect on the outcome depends on the value of the mediator, no single number is the direct effect, and what a regression coefficient then reports is a weighted average over the population whose weights are a property of the fit rather than of the question. That is the same limitation the reversal that is a region records about standardisation, and it bites harder here because the mediator’s distribution is itself a function of the treatment.
And the structure is stipulated. The direct effect is not identified by the observed distribution here, and whether a covariate has an unmeasured cause in common with the outcome is not readable from the data — one covariance matrix holds two effects, and this world has a companion producing the same numbers with a different answer. The measurement here is of what the consequence is if the structure holds, not of how often it does.
What survives, and it is the timing rule
Three of this ladder’s structures ask for a diagram before an adjustment set can be chosen, and offer no shortcut. The one measured here does have a shortcut, and it is the reason to state it plainly at the end rather than to leave it buried under the failure.
A covariate recorded before the treatment was assigned cannot have been caused by it, so no version of this structure applies to it. That rules out the whole of this essay’s failure by a check on a date, which is cheaper than any of the alternatives — and it is why an adjustment set fixed at design time, as in balancing what is known in advance, is worth more than the same set chosen afterwards. It does not rule out the others: a pre-treatment covariate can still be a common effect of two unmeasured causes, and the sample the analysis is conditioned on has no timestamp at all.
What it also does not rule out is the change of question, and that is the honest close. A study that wants the direct effect has to condition on something post-treatment, so it has to make exactly the assumption this essay measures the cost of breaking: that nothing unmeasured causes both the mediator and the outcome. That assumption is not testable from the data, and there is no adjustment set that repairs it — which is why the total effect, which needs no such assumption, is the quantity a randomised trial is designed around, and why the essay that refuses “explained” as a causal word is refusing the move that leads here. The direct effect is a harder thing to know, and the regression that reports it does not say so.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Adjusting for a shadow — both name adjustment set, bias, causal diagram, covariate adjustment, latent variable, unmeasured confounding
- Adjusting for everything — both name adjustment set, bias, causal diagram, covariate adjustment, mediator
- Conditioning on what the treatment caused — both name causal diagram, direct effect, mediator, total effect
- A copula that halves a marginal — both name covariate adjustment, identification
- The bias that lands in the slope — both name bias, regression
Named objects
A flat tag is an object no other essay names yet.
Adjustment setBiasCausal diagramCollider biasCovariate adjustmentDirect effectIdentificationLatent variableMediatorPath coefficientPost-treatment variableRegressionTotal effectUnmeasured confounding