What conditioning on a variable does

One arithmetic, three decisions

A covariate beside a treatment and an outcome can be a common cause of both, a step on the path between them, or an effect of both. The regression that includes it is the same arithmetic in all three, and it is right in one — returning 0.5000, deleting 0.6300 of the effect, and turning 0.5000 into −0.0872.

Worth reading first: Simpson's reversal is a region, not a table.

Three variables, and the same three numbers on the arrows between them. A treatment tt, an outcome yy, and a covariate cc that somebody has to decide whether to type into the regression. Point the two arrows that touch cc three different ways, leaving their strengths alone, and the effect the treatment has on the outcome is 0.5000, 1.1300 and 0.5000. Run the same regression of yy on tt and cc in each of the three and it returns 0.5000, 0.5000 and −0.0872.

The last of those is not an approximation of anything. The effect is a half, the estimate is negative, and every step of the arithmetic between them is correct. Nothing was measured badly, no assumption of least squares was violated, and the sample was not small — the number is the population value, and the counted estimate over eight hundred fits of six hundred rows lands on it at −0.0847.

Adjusting is right in exactly 1 of the three worlds and leaving the covariate out is right in exactly 2, and the regression output does not say which world it is in. That is the whole of this field’s subject, and this essay is the arithmetic of it.

Three arrows, and one estimator that cannot see them

Every world here is written as a set of equations, one per variable, each variable a weighted sum of its direct causes plus its own independent noise. Three coefficients name the three edges. The edge between the treatment and the covariate is a=0.90a = 0.90, the edge between the covariate and the outcome is d=0.70d = 0.70, and the direct edge from treatment to outcome is b=0.50b = 0.50. Only the arrowheads move.

A common cause of both:

c=ec,t=ac+et,y=bt+dc+ey.c = e_c, \qquad t = a\,c + e_t, \qquad y = b\,t + d\,c + e_y .

A step on the path from one to the other:

t=et,c=at+ec,y=bt+dc+ey.t = e_t, \qquad c = a\,t + e_c, \qquad y = b\,t + d\,c + e_y .

A common effect of both:

t=et,y=bt+ey,c=at+dy+ec.t = e_t, \qquad y = b\,t + e_y, \qquad c = a\,t + d\,y + e_c .

Every noise term is standard normal and independent of the others. The three equation sets differ in which variable is written on the left of which, and in nothing else.

What the treatment is worth is a different quantity in each. In the first it is b=0.5000b = 0.5000: the covariate is upstream of the treatment and none of the treatment’s influence goes through it. In the second it is b+ad=1.1300b + ad = 1.1300, because the treatment reaches the outcome directly and again around through the covariate. In the third it is b=0.5000b = 0.5000 again, since the covariate is downstream of both and carries no influence anywhere.

The estimator is the same object in all three. Write Σ\Sigma for the covariance matrix of the three variables; the regression of yy on tt alone returns

β^out=ΣtyΣtt,\hat\beta_{\text{out}} = \frac{\Sigma_{ty}}{\Sigma_{tt}},

and the regression of yy on tt and cc returns

β^in=ΣtyΣccΣtcΣcyΣttΣccΣtc2.\hat\beta_{\text{in}} = \frac{\Sigma_{ty}\Sigma_{cc} - \Sigma_{tc}\Sigma_{cy}}{\Sigma_{tt}\Sigma_{cc} - \Sigma_{tc}^{2}} .

Both are functions of Σ\Sigma and of nothing else. Neither takes a direction as an argument, because a covariance matrix does not hold one.

The same covariate, three ways round. Three worlds over a treatment, an outcome and a covariate, joined by the same three edges at the same three strengths — 0.90, 0.50 and 0.70 — differing only in which way the two edges touching the covariate point. In the first the covariate causes both and adjusting for it recovers the effect of 0.50 exactly. In the second the treatment causes the covariate, the effect is 1.13, and adjusting returns 0.50 — the direct edge alone, with the part that travels through the covariate deleted. In the third the treatment and the outcome both cause the covariate, the effect is 0.50, and adjusting returns -0.087. The regression that produces those three numbers is one formula, and nothing in the data says which panel it is being run in.
Fig. 1 The three arrangements of the same two edges, with the effect each holds and what each of the two regressions returns under it. The treatment sits bottom left, the outcome bottom right and the covariate at the apex in every panel, so the only thing that moves between panels is an arrowhead.

What one decision costs in each world

The two estimates and the effect are all closed forms in aa, bb, dd and the three noise scales, so the bias each decision carries can be written down rather than described.

Leaving the covariate out, in the common-cause world, is off by

adσc2a2σc2+σt2=0.3481,\frac{a\,d\,\sigma_c^{2}}{a^{2}\sigma_c^{2} + \sigma_t^{2}} = 0.3481 ,

which is the classical omitted-variable term: the product of the two edges of the back-door route, discounted by how much of the treatment’s variance that route accounts for. The estimate is 0.8481 against an effect of 0.5000, a ratio of 1.696. Putting the covariate in removes it exactly, to 0.0000.

Putting the covariate in, in the mediator world, is off by ad-a\,d, which is −0.6300 — the entire part of the effect that travels through the covariate — 55.8% of a total effect of 1.1300. Holding the covariate fixed while moving the treatment is a question about the direct edge alone, and the direct edge alone is 0.5000. That is a real quantity and it answers a real question; it is simply not the question the estimate is usually read as answering. Leaving the covariate out is off by nothing at all here, because the marginal regression of the outcome on the treatment recovers the sum of every route between them.

Putting the covariate in, in the common-effect world, is off by

d(a+db)σy2d2σy2+σc2=0.5872,-\,\frac{d\,(a + d\,b)\,\sigma_y^{2}}{d^{2}\sigma_y^{2} + \sigma_c^{2}} = -0.5872 ,

A bias of −0.5872 takes an effect of 0.5000 to an estimate of −0.0872. Nothing was omitted here; something was added. The covariate is a common effect of the treatment and the outcome, so holding it fixed forces its two causes to trade against each other — a high treatment value inside a stratum of fixed cc has to be paid for by a low outcome value, and the regression reads that trade as a negative effect. Leaving it out is exact.

One feature of that last expression is worth reading slowly, because it has no counterpart in the other two. The direct effect bb appears inside the bias, in the term a+dba + d\,b. The manufactured bias is a function of the very quantity it corrupts, which means there is no fixed amount to subtract off and no direction to correct in: a larger true effect produces a larger bias against it, and at these edges the two happen to cross. Nothing similar is true of the omitted-variable term in the common-cause world, which is adσc2/(a2σc2+σt2)a\,d\,\sigma_c^2/(a^2\sigma_c^2 + \sigma_t^2) and contains no bb at all — a confounded estimate is the effect plus a quantity that does not depend on the effect, and a collider-adjusted estimate is not the effect plus anything.

One regression, three answers, one of them right. What the treatment's estimated effect is off by in each of three worlds, with and without the covariate in the model, at edges of 0.90, 0.50 and 0.70. Where the covariate causes both, leaving it out is off by +0.348 and putting it in is off by nothing. Where the treatment causes the covariate, leaving it out gives the total effect exactly and putting it in deletes 0.6300 of it. Where the treatment and the outcome both cause the covariate, leaving it out is again exact and putting it in is off by −0.587 — enough to turn an effect of 0.50 into an estimate of -0.087. Each closed form is drawn beside 800 least-squares fits on 600 rows, the worst of which sits 1.76 standard errors from it.
Fig. 2 The bias each of the two decisions carries in each of the three worlds, closed form against the count from 800 least-squares fits of 600 rows. The three bars that sit at zero are the three decisions that are right.

Two routes, and the worst of six is 1.76 standard errors

The forms above are algebra on a population covariance. The bars they are drawn against are least squares run on simulated rows, and the two share no arithmetic — the method every number here is checked by is that a quantity with one route is a quantity nobody has checked.

Six readings, each 800 fits of 600 rows: the common-cause world at 0.8495 and 0.5024 against 0.8481 and 0.5000, the mediator world at 1.1293 and 0.4966 against 1.1300 and 0.5000, and the common-effect world at 0.5007 and −0.0847 against 0.5000 and −0.0872. In units of each count’s own standard error the worst disagreement anywhere in that table is 1.76, and it is the mediator world’s adjusted estimate.

That is the number a reader should hold against the alternative explanation of everything in the section above, which is that the closed forms are simply wrong — mis-transcribed, or derived for a matrix indexed the other way round. A transposed index in the covariance would have swapped the covariate’s role with the outcome’s and produced closed forms that miss their counts by a great deal rather than by a fifth of a standard error. The counts are not evidence that the structures are the right model of anything; they are evidence that the algebra describes the model it claims to describe, which is a smaller claim and the only one available.

The way that agreement could be hollow is worth naming, because it is the ordinary way this kind of check goes wrong. If the simulated rows were generated from the covariance matrix the closed forms are algebra on, rather than from the structural equations, the two routes would share their arithmetic and the agreement would be a tautology dressed as a measurement. They are not: the rows are drawn one variable at a time, each from its own equation and its own independent noise, in the causal order the world states, and the covariance is never formed on that side at all. The only object the two routes have in common is the coefficient set, which is the input to both rather than an intermediate in either.

The reversal is a region, not this arithmetic’s luck

A single set of edge strengths that produces a sign reversal is an anecdote, and the honest question is how much of the parameter space behaves like it. The edges were swept over a five-by-five grid of aa and dd running from 0.3 to 1.5, at a direct effect held at 0.5.

Adjusting in the common-cause world is exact on all 25 of those settings — the omitted-variable term is removed completely wherever it comes from. Adjusting in the common-effect world drives the estimate below zero on 15 of the 25, and the adjusted coefficient runs from −0.5385 to 0.3761 across the grid. The sign reversal is a region of the space rather than a coincidence of the three numbers this essay happened to pick.

The two ends of that grid say what the region is made of. The largest adjusted estimate, 0.3761, sits at the corner where both edges into the covariate are weakest, at a=d=0.3a = d = 0.3 — and it is still a quarter below the effect of 0.5000, so the mildest collider on the grid costs a quarter of the answer. The smallest, −0.5385, sits at the corner where both are strongest, at a=d=1.5a = d = 1.5, and is further below zero than the effect is above it. There is no setting on the grid at which adjusting is harmless, which is a different statement from the sign reversal and the more useful one: the reversal is where the failure becomes visible, not where it begins.

That is the same move as the essay that turned Simpson’s reversal into a region, which found 31% of an allocation space rather than one famous table, and it is the move that separates a curiosity from a hazard. A phenomenon that occupies one point of a parameter space is something to admire; one that occupies three-fifths of a grid is something a reader is standing in.

The same covariate, three ways round. Three worlds over a treatment, an outcome and a covariate, joined by the same three edges at the same three strengths — 0.90, 0.50 and 1.40 — differing only in which way the two edges touching the covariate point. In the first the covariate causes both and adjusting for it recovers the effect of 0.50 exactly. In the second the treatment causes the covariate, the effect is 1.76, and adjusting returns 0.50 — the direct edge alone, with the part that travels through the covariate deleted. In the third the treatment and the outcome both cause the covariate, the effect is 0.50, and adjusting returns -0.257. The regression that produces those three numbers is one formula, and nothing in the data says which panel it is being run in.
Fig. 3 The same three arrangements at a covariate–outcome edge of 1.4 rather than 0.7. The mediator world’s total effect rises to 1.7600 while the adjusted regression still returns 0.5000, and the common-effect world’s adjusted estimate falls to −0.2568.

What the data says about which world it is in

Nothing. That claim needs a measurement rather than an assertion, and the measurement is that the three worlds imply the same pattern of dependence and can be made to imply the same numbers.

All three are complete graphs on three nodes. Every pair of variables is joined by an edge, so no pair is independent, and every pair remains joined once the third is held fixed, so no pair is conditionally independent either. There is no independence anywhere in any of the three, which means there is no independence that distinguishes them.

That is stronger than saying the three worlds look similar. A test of conditional independence is the only feature of a joint distribution that a diagram of arrows constrains directly, so it is the whole of what any procedure reading structure off data has to work with. Three complete graphs impose no constraint at all, and a procedure with nothing to test is not a weak procedure — it is not a procedure.

The six correlations bear it out at the settings above. Marginally, the treatment and the covariate correlate at 0.6690, the treatment and the outcome at 0.7114, and the covariate and the outcome at 0.7170. Partialling each pair on the third gives 0.4472, 0.3244 and 0.4616. None of the six is zero and none is near zero, so a test of independence has nothing to reject and nothing to fail to reject that any of the three worlds would not predict.

Six correlations, three worlds, no difference. Every correlation and every partial correlation among the treatment, the covariate and the outcome, in three worlds parameterised to share one joint distribution. The six readings are 0.6690, 0.7114, 0.7170, 0.4472, 0.3244, 0.4616, and the largest disagreement between the three worlds on any of them is 5.0e-16 — machine noise. None is zero, so there is no conditional independence to test: all three are complete graphs on three nodes, and a complete graph has no missing edge for the data to notice. The effects those three worlds hold are 0.500, 0.848, 0.848.
Fig. 4 The three marginal correlations and the three partial correlations, in the three worlds that share one covariance matrix. Every bar is the same height in all three, and no bar is at zero.

One distribution, and it does not hold one effect

The stronger version of that claim is that the three worlds are not merely similar in their pattern of dependences but can be made identical in their numbers, and it is the subject of the argument that one covariance holds two effects.

Regressing each variable on its parents fits any causal order to a covariance matrix exactly and uniquely, so given the common-cause world at these edges there is a mediator world and a common-effect world whose joint distributions agree with it entry for entry, to 4.4·10⁻¹⁶. The adjusted regression returns 0.5000 in all three, to the same tolerance, because it is a function of the covariance and the covariance is the same. The effect those worlds hold is 0.5000 in one and 0.8481 in the other two.

Two of the numbers a reader wants are identical and the third differs by a third of its own size. Whatever settles which world the data came from, it is not more data — the distributions are the same and a larger sample estimates the same matrix more precisely. This is the point at which the subject stops being statistics and becomes something else, and it is exactly the thing the essay that refuses “explained” as a causal word is refusing on the way past.

One distribution, two effects. Three causal structures fitted to one covariance matrix over a treatment, a covariate and an outcome. Each reproduces it exactly — the largest entry-wise disagreement across all three is 4.4e-16 — so no sample of any size distinguishes them. The regression of the outcome on the treatment and the covariate returns 0.500 in all three, to within 4.4e-16, because that coefficient is a function of the covariance and of nothing else. The effect the three worlds hold is 0.500, 0.848 and 0.848: adjusting is exactly right in the first and off by −0.348 in the other two. The arithmetic cannot see the difference and the difference is the whole question.
Fig. 5 Three parameterisations of one covariance matrix, with the effect each of them holds. The largest disagreement between any two of the three matrices is 4.4·10⁻¹⁶, and the effects are 0.5000, 0.8481 and 0.8481.

Then leave everything out?

Two worlds of three reward leaving the covariate out and one rewards putting it in, so the arithmetic above appears to recommend a rule: adjust for nothing. It does not, and the reason is that the three worlds here are equally weighted because they were enumerated rather than sampled from anything.

Real covariates are not drawn uniformly from the three arrangements, and a set of six of them draws several roles at once. Over four thousand structures of six covariates each, “control for every covariate measured” leaves a larger bias than controlling for nothing on 65.5% of them and a smaller one on 33.8% — so the crude rule really is worse than the crude alternative, and the sweep that priced it also finds that adjusting for nothing has the lowest squared error on only 18.4% of the same structures. Neither crude rule is the answer; the structure is.

Two structures in three are made worse. What adjusting for every covariate measured does to the bias in the treatment's estimated effect, against adjusting for none, over 4000 randomly drawn structures of 6 covariates each. Each covariate is independently a common cause with probability 0.25, a cause of the treatment only, a cause of the outcome only, a cause of neither, a step on the causal path, or a common effect. The rule leaves a larger bias on 65.5% of structures, a smaller one on 33.8%, and the same on 0.7%. The share is a property of that population of structures rather than of adjustment, which is why the weights are stated; what does not depend on them is that the rule has no direction — it is not a conservative default that occasionally overcorrects, it is a rule whose error is whatever the structure happens to be.
Fig. 6 How often adjusting for every measured covariate leaves a larger bias than adjusting for none, over 4,000 randomly drawn structures of six covariates. It hurts on 65.5% and helps on 33.8%.

Where this does not hold

Everything above is linear and Gaussian, and that is what buys the closed forms rather than an incidental convenience.

The consequence worth naming is that the equivalence of the three worlds is an equivalence of second moments. A covariance matrix is the whole of a Gaussian joint law, so two structures that share one share everything; a non-Gaussian law carries information a covariance does not, and there are families in which the direction of an arrow is recoverable from observational data alone. The claim that the data cannot choose is therefore exactly as strong as the assumption that the data is Gaussian, and no stronger. Where that assumption fails, it fails in the direction of making the problem easier rather than harder, which is the rare direction for a limitation to run.

The second is that the covariate here is measured perfectly or not at all. A covariate observed with error is neither, and adjusting for a noisy version of a common cause removes some of the confounding while adding an attenuation of its own; nothing here measures where the two cross. The third is that all three worlds hold the same effect for every unit, so there is nothing here about an effect that differs across the population — the reversal that four datasets with one set of summaries is about, and the sensitivity that a fitted slope reversed by one point is about, are both about the fit rather than about the structure, and both are live in every world above.

What the regression output would have to say, and does not

The practical residue is short and it is unwelcome.

The estimate, its standard error, its tt statistic, its R2R^2, its residual plots and every diagnostic computable from the fit are identical in the common-cause world and in the two worlds that were built to share its covariance. The estimate is right in one and wrong by a third of its own size in the others. There is no number in the output that moves between them, because every number in the output is a function of Σ\Sigma and Σ\Sigma does not move.

This is why the decision is made before the data is read rather than after, and why what randomisation actually buys is worth its cost: randomising the treatment deletes every arrow into it, which removes the common-cause world from the list of candidates by construction rather than by test. It does not remove the other two — a randomised trial can still be ruined by adjusting for a variable measured afterwards, which is what the covariate the treatment caused measures, and it cannot protect against a sample that was assembled by conditioning on something, which is what being in the data as a condition is about.

And the covariate this essay is deciding about was a clean three-node picture with the causal question written on it. The harder case is a covariate that is measured before the treatment, is on no causal path, is not a common cause, and is still not safe: a collider before the treatment is that case, and its bias has a bound rather than a limit. The pattern recurs elsewhere on this site whenever two quantities share a cause — a positive test needing a base rate is the same structure counted in a different currency, and two mechanisms with one long-run relation is what it looks like when the shared cause is time.

What none of this measures is the frequency with which each of the three worlds is the true one, because that is not a statistical quantity at all. It is a fact about the subject the covariate is a covariate of, and the only honest thing an essay about the arithmetic can do is say which decision each world rewards and refuse to average over them.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adjustment setCausal diagramCausal effectColliderCollider biasConfoundingCovariate adjustmentDirect effectLeast squaresMediatorOmitted variable biasPath coefficientRegressionStructural modelTotal effect