A collider before the treatment
Worth reading first: One arithmetic, three decisions.
The rule that rescues a sweep over random structures from being a curiosity is the one everybody already uses: only adjust for covariates recorded before the treatment was assigned, because the treatment cannot have caused them. It is a rule about timing rather than about structure, it needs no causal diagram, and it is checkable from a questionnaire’s date field.
Here is a covariate that satisfies it, satisfies every other rule of thumb in circulation, and biases the estimate by exactly −0.2000 against a treatment effect of 0.50 — 40.0% of the answer — while the regression that leaves it out is exact to fifteen places. It is recorded first. It is on no causal path from treatment to outcome. It is not a common cause of them. It is not an effect of the treatment, or of the outcome, or of anything either of them touches.
And the damage it can do has a ceiling. The bias is bounded above by 0.3536 at these latent strengths, and the two coefficients that make the covariate a common effect do not appear in that bound at all. A twenty-five by twenty-five sweep taking both of them to 3 reaches 0.3197, which is 90.4% of the bound and 63.9% of the treatment effect. The guess going in was that a bias like this grows without limit; it does not, and “how bad can it get” has an exact answer.
The structure, and why every rule of thumb passes it
Two variables are never measured. Call them and . The first causes the treatment; the second causes the outcome; and both of them cause the covariate.
Four coefficients name the four latent edges: for , for , for and for . Every one of them is 1 in the base case, the noise scales are all 1, and the treatment’s real effect on the outcome is .
Read the diagram as a route map. There is no unblocked back-door route from treatment to outcome: the route passes into the covariate from both sides, and a route through a common effect is blocked by default — the two causes of a common effect are independent until it is held fixed. So the estimate that ignores the covariate is unbiased, and the diagram says so before any arithmetic.
Conditioning on the covariate opens that route. The covariate is a collider on it, and holding a collider fixed makes its causes dependent, so and — which were independent — acquire an association inside every stratum of . That association runs from the treatment to the outcome through two latent variables nobody measured, and the regression reads it as an effect.
Every rule of thumb about which covariates are safe is a rule about the covariate’s relationship to the treatment and the outcome, and this covariate has none. It is not caused by either, it causes neither, and it is not a cause of both. The relationship that matters is a relationship to two variables that are not in the dataset, and no procedure reading the dataset can see it.
The bias, written down
The whole of it is one expression. Writing for the treatment’s variance,
which at unit paths is exactly. It is worth reading the numerator and the denominator separately, because they say different things.
The numerator is a product of all four latent edges. Set any one of them to zero and the bias vanishes. That is the sense in which this is a fragile structure rather than a common one: it needs a latent cause of the treatment that also causes the covariate, and a separate latent cause of the outcome that also causes the covariate, and all four edges to be real.
The denominator grows quadratically in the same two paths the numerator is linear in. That is where the ceiling comes from and it is the whole of the result. As and grow, the covariate becomes a better and better measurement of the two latents — but conditioning on it also removes more and more of the treatment’s own variance, and the second effect wins. The bias is a ratio in which the top is degree two in and the bottom is degree two as well, so it tends to a finite limit rather than to infinity.
The arithmetic–geometric mean inequality makes that limit exact. Since , the whole expression is bounded by
with equality approached only as both and grow without limit along the ray . Neither nor survives into that bound. How strongly the two latents reach the covariate cannot decide how bad the adjustment is; it can only decide how close to the bound a particular case sits.
Why there is a ceiling at all
A bound that falls out of an inequality is easy to check and easy to distrust, so it is worth saying in mechanism what the algebra is doing — and worth saying against the case where no ceiling exists.
Conditioning on the covariate does two things at once and they run in opposite directions. It opens the route between the two latents, which is where the bias comes from, and the strength of that opening rises with how much the covariate says about them — with , the product of the two paths into it. It also removes from the treatment exactly the part of its variation that the covariate can account for, and what the covariate accounts for also rises with , because is how much of — and therefore of — is written into . The first effect is the numerator and the second is in the denominator, and the second is the faster of the two.
Read as degrees, the numerator is degree two in jointly and the denominator is degree two in each of them separately. Scale both paths by a factor and the numerator grows by while and grow by as well, so the ratio tends to a constant. The only term that does not scale is , the covariate’s own noise, and dropping it is exactly what turns the bound into an equality — which is why the limit is approached rather than attained, and why the sweep reaches 90.4% of it rather than 100%.
The contrast that makes this worth a section is the ordinary confounder, which has no ceiling. Omitting a genuine common cause costs , and the edge from the covariate to the outcome appears in the numerator and nowhere else. Double it and the bias doubles; take it to a hundred and the bias goes to a hundred times what it was. There is no configuration of that structure in which the damage saturates, and the reason is structural rather than algebraic: a confounder’s route to the outcome is not something conditioning is also shutting. Here, the same coefficient that opens the route also closes the treatment’s variance around it, and a quantity that appears on both sides of a ratio cannot run away.
So the two failures this ladder prices are different in kind and not only in size. A missing adjustment is a bias with no upper limit and a knowable direction; an adjustment that should not have been made is a bias with an upper limit and no direction at all. Neither of those is the safer one to be in, and a reader who has only ever been warned about the first has been warned about the unbounded one.
Two routes, and one of them is 2.04 standard errors out
The forms above are algebra on a population covariance. Beside them: five hundred least-squares fits of eight hundred rows apiece, on data simulated one variable at a time from the structural equations in causal order.
The adjusted estimate counts at 0.3023 with a standard error of 0.00155, a bias of −0.19771 against a derived −0.2000 — a fifteenth of a standard error, and a bias more than a hundred and twenty standard errors from zero, so there is no reading on which this is noise. The unadjusted estimate counts at 0.5032 against an exact 0.5000, a bias of 0.00322 on a standard error of 0.00157.
That last one is 2.04 standard errors out, and it is worth not smoothing over. A claim of exactness that lands two standard errors high on five hundred fits is either an exact claim with an unremarkable draw behind it or a small real bias that the count is beginning to resolve. What decides it is that the closed form for the unadjusted estimate is identically, with the latent terms cancelling algebraically rather than approximately, and the population covariance computed by the same routine that feeds the adjusted case returns 0.5000 to fifteen places. A finite-sample bias of order in a least-squares slope is not zero at eight hundred rows, and 0.003 is about the size it would be. The honest statement is that the exactness is a population statement, and the count is consistent with it at a level nobody should describe as confirmation, which is what two routes to every number is for: the second route is there to catch an error in the first, and a two-standard-error reading is the point at which it is doing its job rather than failing.
How bad can it get, and where
A supremum reached only in a limit is a weak statement on its own, so the surface was swept: a twenty-five by twenty-five grid taking and from 0.12 to 3.
The largest bias anywhere on it is −0.3197, at and . That is 90.4% of the bound of 0.3536 and 63.9% of the treatment effect. The bound is therefore reachable rather than decorative: a structure with both latent paths at three is not exotic, and it gets nine tenths of the way to a limit that no structure of this shape can pass.
The location of that maximum is the second result. It is not at the corner of the grid. is interior in that direction while is against the edge, so the sweep found a genuine interior optimum in one variable and a monotone increase in the other. Differentiating the closed form in gives the ridge
which at is 2.3452 — and the grid’s maximum sits on the nearest grid point to it.
The worst second path is a middling one
Walk along the ridge and the arithmetic says something a reader would not guess: for each strength of the first latent path there is a particular strength of the second that does the most damage, and it is not the largest one.
At the worst second path is , doing 0.0299 of damage. At it is 1.0607 and 0.1179. At it is 1.2247 and 0.2041. At , 1.4577 and 0.2572; at 2.00, 1.7321 and 0.2887; at 2.50, 2.0310 and 0.3077; at 3.00, 2.3452 and 0.3198.
Two things are in that column. The damage rises with and rises more and more slowly — from 0.0299 to 0.2041 over the first unit and from 0.2887 to 0.3198 over the last — which is the saturation seen from the side. And the ratio climbs from 0.12 towards 1.4142, the ray the inequality names, reaching 1.279 at . The sweep’s own best ratio is 1.25, below the asymptotic one, and the gap between them is entirely the term in the denominator: the covariate’s own noise, which matters when the latent paths are small and becomes negligible when they are large.
Why the ridge exists at all is the same mechanism the bound is. Past the covariate is mostly the second latent’s own contribution, so conditioning on it explains away that latent while removing treatment variance that carries the signal — the bias falls back. It is a turning point, not a divergence.
What the bound does depend on
Since the two paths into the covariate do not decide the ceiling, something else does, and the expression names it: the ceiling is — the two latent paths that reach the treatment and the outcome, scaled by the treatment’s own noise.
At and the bound is 0.2236. At it is 0.3536. At , it is 0.7071; at , it is 0.4472; at it is 0.8944. It is linear in , which reaches only the outcome, and sublinear in , which reaches the treatment and therefore appears in the denominator too.
That asymmetry is the useful part. Doubling the latent cause of the outcome doubles the worst case. Doubling the latent cause of the treatment raises it by 26%, from 0.3536 to 0.4472, because a stronger also gives the treatment more variance for the bias to be a fraction of. So a reader worried about this structure should be asking how strongly something unmeasured drives the outcome, not how strongly something unmeasured drives assignment — which is the reverse of the ordinary confounding intuition, and the reverse for a reason: this is not a confounding route, it is one that adjustment builds.
Sweeping the grid at reaches −0.3586 at , — 80.2% of that world’s bound and 71.7% of the treatment effect, so the absolute damage is larger and the fraction of the bound smaller. Its bias at unit and is −0.1818, which is below the 0.2000 of the base case: a stronger latent cause of the treatment makes the worst case worse and the typical case better.
Where this does not hold
The bound is a one-covariate result. Nothing here measures several pre-treatment colliders at once. What is measured is that a population of structures gets monotonically worse as their share rises — from 65.5% of structures harmed by adjusting for everything at none, to 73.3% at three in ten, in the sweep that priced the crude rules. Whether such covariates give times the bias or something that saturates again is not measured, and the shape of the closed form suggests they compound. A suggestion is not a measurement, which is the whole reason it is stated here as one rather than as the other.
Everything is linear and Gaussian, so the covariate is a linear function of two latents plus noise. A covariate that is a threshold on such a function — a binary indicator of hospital admission, say — is a common effect too, and the arithmetic is not this arithmetic. The direction is the same and the bound is not.
The structure is stipulated rather than found. A reader cannot look at a dataset and discover that a covariate has this shape, because the shape is entirely about two variables that are not in it — which is the argument that one covariance holds two effects arriving in its most practical form. The claim is not that a given baseline covariate is dangerous. It is that no observed feature of a baseline covariate certifies that it is safe, and the timing rule is the certificate people use. That is the same uncomfortable place the reversal that no amount of data settles arrives at from the other direction, and the same one reached by a sample that was assembled by conditioning on something — being in the data as a condition — where the collider is the selection rule itself and is not a column at all.
What to do about it, since “do not adjust” is not the answer
The temptation this essay creates is to stop adjusting, and that would be the wrong lesson from it. The same covariate three ways round shows what leaving a genuine common cause out costs — 0.3481 against an effect of 0.5000, larger than anything here — and the sweep over random structures finds adjusting for nothing to be the best of three rules on only 18.4% of them. A structure that makes adjustment harmful does not make non-adjustment harmless.
What the M-structure argues for is narrower and more useful: the adjustment set is a decision about which covariates, argued from a diagram, rather than a decision about how many. That is a decision better made before the data exists than after, which is exactly why balancing what is known in advance and the variance blocking removes before any data are design-time operations: a set fixed at design time cannot be chosen to move the estimate, and a set fixed after the fact can. What randomisation buys helps here and does not solve it — randomising the treatment deletes , which sets the numerator of the bias to zero and the whole structure with it, but only for the treatment actually randomised, and a randomised trial that adjusts for a post-randomisation covariate is the covariate the treatment caused rather than this one.
The bound is the thing to carry. A bias with a ceiling is a bias somebody can reason about: at these latent strengths the worst an adjustment of this shape can do is 0.3536, whatever the covariate’s connection to the unmeasured causes, and a sensitivity analysis that assumes it could do anything is assuming something false.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A block weighted inside itself — both name bias, closed form
- A copula that halves a marginal — both name closed form, covariate adjustment
- A dictionary that is neither — both name closed form, covariate adjustment
- A space is not a relation — both name closed form, structural model
- A symmetry that was not enough — both name closed form, covariate adjustment
- A width rule on skewed outcomes — both name bias, closed form
Named objects
A flat tag is an object no other essay names yet.
Adjustment setBack door pathBiasCausal diagramClosed formColliderCollider biasCovariate adjustmentLatent variableM biasPath coefficientStructural modelStudy designUnmeasured confounding