One arithmetic, three decisions
Worth reading first: Simpson's reversal is a region, not a table.
Three variables, and the same three numbers on the arrows between them. A treatment , an outcome , and a covariate that somebody has to decide whether to type into the regression. Point the two arrows that touch three different ways, leaving their strengths alone, and the effect the treatment has on the outcome is 0.5000, 1.1300 and 0.5000. Run the same regression of on and in each of the three and it returns 0.5000, 0.5000 and −0.0872.
The last of those is not an approximation of anything. The effect is a half, the estimate is negative, and every step of the arithmetic between them is correct. Nothing was measured badly, no assumption of least squares was violated, and the sample was not small — the number is the population value, and the counted estimate over eight hundred fits of six hundred rows lands on it at −0.0847.
Adjusting is right in exactly 1 of the three worlds and leaving the covariate out is right in exactly 2, and the regression output does not say which world it is in. That is the whole of this field’s subject, and this essay is the arithmetic of it.
Three arrows, and one estimator that cannot see them
Every world here is written as a set of equations, one per variable, each variable a weighted sum of its direct causes plus its own independent noise. Three coefficients name the three edges. The edge between the treatment and the covariate is , the edge between the covariate and the outcome is , and the direct edge from treatment to outcome is . Only the arrowheads move.
A common cause of both:
A step on the path from one to the other:
A common effect of both:
Every noise term is standard normal and independent of the others. The three equation sets differ in which variable is written on the left of which, and in nothing else.
What the treatment is worth is a different quantity in each. In the first it is : the covariate is upstream of the treatment and none of the treatment’s influence goes through it. In the second it is , because the treatment reaches the outcome directly and again around through the covariate. In the third it is again, since the covariate is downstream of both and carries no influence anywhere.
The estimator is the same object in all three. Write for the covariance matrix of the three variables; the regression of on alone returns
and the regression of on and returns
Both are functions of and of nothing else. Neither takes a direction as an argument, because a covariance matrix does not hold one.
What one decision costs in each world
The two estimates and the effect are all closed forms in , , and the three noise scales, so the bias each decision carries can be written down rather than described.
Leaving the covariate out, in the common-cause world, is off by
which is the classical omitted-variable term: the product of the two edges of the back-door route, discounted by how much of the treatment’s variance that route accounts for. The estimate is 0.8481 against an effect of 0.5000, a ratio of 1.696. Putting the covariate in removes it exactly, to 0.0000.
Putting the covariate in, in the mediator world, is off by , which is −0.6300 — the entire part of the effect that travels through the covariate — 55.8% of a total effect of 1.1300. Holding the covariate fixed while moving the treatment is a question about the direct edge alone, and the direct edge alone is 0.5000. That is a real quantity and it answers a real question; it is simply not the question the estimate is usually read as answering. Leaving the covariate out is off by nothing at all here, because the marginal regression of the outcome on the treatment recovers the sum of every route between them.
Putting the covariate in, in the common-effect world, is off by
A bias of −0.5872 takes an effect of 0.5000 to an estimate of −0.0872. Nothing was omitted here; something was added. The covariate is a common effect of the treatment and the outcome, so holding it fixed forces its two causes to trade against each other — a high treatment value inside a stratum of fixed has to be paid for by a low outcome value, and the regression reads that trade as a negative effect. Leaving it out is exact.
One feature of that last expression is worth reading slowly, because it has no counterpart in the other two. The direct effect appears inside the bias, in the term . The manufactured bias is a function of the very quantity it corrupts, which means there is no fixed amount to subtract off and no direction to correct in: a larger true effect produces a larger bias against it, and at these edges the two happen to cross. Nothing similar is true of the omitted-variable term in the common-cause world, which is and contains no at all — a confounded estimate is the effect plus a quantity that does not depend on the effect, and a collider-adjusted estimate is not the effect plus anything.
Two routes, and the worst of six is 1.76 standard errors
The forms above are algebra on a population covariance. The bars they are drawn against are least squares run on simulated rows, and the two share no arithmetic — the method every number here is checked by is that a quantity with one route is a quantity nobody has checked.
Six readings, each 800 fits of 600 rows: the common-cause world at 0.8495 and 0.5024 against 0.8481 and 0.5000, the mediator world at 1.1293 and 0.4966 against 1.1300 and 0.5000, and the common-effect world at 0.5007 and −0.0847 against 0.5000 and −0.0872. In units of each count’s own standard error the worst disagreement anywhere in that table is 1.76, and it is the mediator world’s adjusted estimate.
That is the number a reader should hold against the alternative explanation of everything in the section above, which is that the closed forms are simply wrong — mis-transcribed, or derived for a matrix indexed the other way round. A transposed index in the covariance would have swapped the covariate’s role with the outcome’s and produced closed forms that miss their counts by a great deal rather than by a fifth of a standard error. The counts are not evidence that the structures are the right model of anything; they are evidence that the algebra describes the model it claims to describe, which is a smaller claim and the only one available.
The way that agreement could be hollow is worth naming, because it is the ordinary way this kind of check goes wrong. If the simulated rows were generated from the covariance matrix the closed forms are algebra on, rather than from the structural equations, the two routes would share their arithmetic and the agreement would be a tautology dressed as a measurement. They are not: the rows are drawn one variable at a time, each from its own equation and its own independent noise, in the causal order the world states, and the covariance is never formed on that side at all. The only object the two routes have in common is the coefficient set, which is the input to both rather than an intermediate in either.
The reversal is a region, not this arithmetic’s luck
A single set of edge strengths that produces a sign reversal is an anecdote, and the honest question is how much of the parameter space behaves like it. The edges were swept over a five-by-five grid of and running from 0.3 to 1.5, at a direct effect held at 0.5.
Adjusting in the common-cause world is exact on all 25 of those settings — the omitted-variable term is removed completely wherever it comes from. Adjusting in the common-effect world drives the estimate below zero on 15 of the 25, and the adjusted coefficient runs from −0.5385 to 0.3761 across the grid. The sign reversal is a region of the space rather than a coincidence of the three numbers this essay happened to pick.
The two ends of that grid say what the region is made of. The largest adjusted estimate, 0.3761, sits at the corner where both edges into the covariate are weakest, at — and it is still a quarter below the effect of 0.5000, so the mildest collider on the grid costs a quarter of the answer. The smallest, −0.5385, sits at the corner where both are strongest, at , and is further below zero than the effect is above it. There is no setting on the grid at which adjusting is harmless, which is a different statement from the sign reversal and the more useful one: the reversal is where the failure becomes visible, not where it begins.
That is the same move as the essay that turned Simpson’s reversal into a region, which found 31% of an allocation space rather than one famous table, and it is the move that separates a curiosity from a hazard. A phenomenon that occupies one point of a parameter space is something to admire; one that occupies three-fifths of a grid is something a reader is standing in.
What the data says about which world it is in
Nothing. That claim needs a measurement rather than an assertion, and the measurement is that the three worlds imply the same pattern of dependence and can be made to imply the same numbers.
All three are complete graphs on three nodes. Every pair of variables is joined by an edge, so no pair is independent, and every pair remains joined once the third is held fixed, so no pair is conditionally independent either. There is no independence anywhere in any of the three, which means there is no independence that distinguishes them.
That is stronger than saying the three worlds look similar. A test of conditional independence is the only feature of a joint distribution that a diagram of arrows constrains directly, so it is the whole of what any procedure reading structure off data has to work with. Three complete graphs impose no constraint at all, and a procedure with nothing to test is not a weak procedure — it is not a procedure.
The six correlations bear it out at the settings above. Marginally, the treatment and the covariate correlate at 0.6690, the treatment and the outcome at 0.7114, and the covariate and the outcome at 0.7170. Partialling each pair on the third gives 0.4472, 0.3244 and 0.4616. None of the six is zero and none is near zero, so a test of independence has nothing to reject and nothing to fail to reject that any of the three worlds would not predict.
One distribution, and it does not hold one effect
The stronger version of that claim is that the three worlds are not merely similar in their pattern of dependences but can be made identical in their numbers, and it is the subject of the argument that one covariance holds two effects.
Regressing each variable on its parents fits any causal order to a covariance matrix exactly and uniquely, so given the common-cause world at these edges there is a mediator world and a common-effect world whose joint distributions agree with it entry for entry, to 4.4·10⁻¹⁶. The adjusted regression returns 0.5000 in all three, to the same tolerance, because it is a function of the covariance and the covariance is the same. The effect those worlds hold is 0.5000 in one and 0.8481 in the other two.
Two of the numbers a reader wants are identical and the third differs by a third of its own size. Whatever settles which world the data came from, it is not more data — the distributions are the same and a larger sample estimates the same matrix more precisely. This is the point at which the subject stops being statistics and becomes something else, and it is exactly the thing the essay that refuses “explained” as a causal word is refusing on the way past.
Then leave everything out?
Two worlds of three reward leaving the covariate out and one rewards putting it in, so the arithmetic above appears to recommend a rule: adjust for nothing. It does not, and the reason is that the three worlds here are equally weighted because they were enumerated rather than sampled from anything.
Real covariates are not drawn uniformly from the three arrangements, and a set of six of them draws several roles at once. Over four thousand structures of six covariates each, “control for every covariate measured” leaves a larger bias than controlling for nothing on 65.5% of them and a smaller one on 33.8% — so the crude rule really is worse than the crude alternative, and the sweep that priced it also finds that adjusting for nothing has the lowest squared error on only 18.4% of the same structures. Neither crude rule is the answer; the structure is.
Where this does not hold
Everything above is linear and Gaussian, and that is what buys the closed forms rather than an incidental convenience.
The consequence worth naming is that the equivalence of the three worlds is an equivalence of second moments. A covariance matrix is the whole of a Gaussian joint law, so two structures that share one share everything; a non-Gaussian law carries information a covariance does not, and there are families in which the direction of an arrow is recoverable from observational data alone. The claim that the data cannot choose is therefore exactly as strong as the assumption that the data is Gaussian, and no stronger. Where that assumption fails, it fails in the direction of making the problem easier rather than harder, which is the rare direction for a limitation to run.
The second is that the covariate here is measured perfectly or not at all. A covariate observed with error is neither, and adjusting for a noisy version of a common cause removes some of the confounding while adding an attenuation of its own; nothing here measures where the two cross. The third is that all three worlds hold the same effect for every unit, so there is nothing here about an effect that differs across the population — the reversal that four datasets with one set of summaries is about, and the sensitivity that a fitted slope reversed by one point is about, are both about the fit rather than about the structure, and both are live in every world above.
What the regression output would have to say, and does not
The practical residue is short and it is unwelcome.
The estimate, its standard error, its statistic, its , its residual plots and every diagnostic computable from the fit are identical in the common-cause world and in the two worlds that were built to share its covariance. The estimate is right in one and wrong by a third of its own size in the others. There is no number in the output that moves between them, because every number in the output is a function of and does not move.
This is why the decision is made before the data is read rather than after, and why what randomisation actually buys is worth its cost: randomising the treatment deletes every arrow into it, which removes the common-cause world from the list of candidates by construction rather than by test. It does not remove the other two — a randomised trial can still be ruined by adjusting for a variable measured afterwards, which is what the covariate the treatment caused measures, and it cannot protect against a sample that was assembled by conditioning on something, which is what being in the data as a condition is about.
And the covariate this essay is deciding about was a clean three-node picture with the causal question written on it. The harder case is a covariate that is measured before the treatment, is on no causal path, is not a common cause, and is still not safe: a collider before the treatment is that case, and its bias has a bound rather than a limit. The pattern recurs elsewhere on this site whenever two quantities share a cause — a positive test needing a base rate is the same structure counted in a different currency, and two mechanisms with one long-run relation is what it looks like when the shared cause is time.
What none of this measures is the frequency with which each of the three worlds is the true one, because that is not a statistical quantity at all. It is a fact about the subject the covariate is a covariate of, and the only honest thing an essay about the arithmetic can do is say which decision each world rewards and refuse to average over them.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Conditioning on what the treatment caused — both name causal diagram, confounding, direct effect, mediator, total effect
- A penalty is a trace — both name least squares, regression
- Borrowing towards a line — both name confounding, least squares
- The repair that was exact and made it worse — both name least squares, regression
- The word a fraction costs — both name confounding, least squares
- Two analyses of one baseline — both name confounding, covariate adjustment
Named objects
A flat tag is an object no other essay names yet.
Adjustment setCausal diagramCausal effectColliderCollider biasConfoundingCovariate adjustmentDirect effectLeast squaresMediatorOmitted variable biasPath coefficientRegressionStructural modelTotal effect