Adjusting for everything
Worth reading first: One arithmetic, three decisions · Simpson's reversal is a region, not a table.
The commonest adjustment rule in observational work is not a rule about structure at all. It is: put every covariate that was measured into the model, on the reasoning that a covariate left out can only hurt and a covariate put in can only help. Priced over four thousand randomly drawn structures of six covariates each, it leaves a larger bias than putting none of them in on 65.5% of them and a smaller one on 33.8%.
Its mean squared error is 0.25974 against 0.06320 for using no covariate at all and 0.00356 for the back-door set — 4.110 times worse than doing nothing and 73.008 times worse than doing the right thing. Half of that error, 49.8%, comes from its worst tenth of structures.
The finding is not that adjusting for everything is a bad conservative default. It is that it is not a default in any direction: on a third of the structures it is the better of the two crude rules, on two thirds it is worse, and nothing computable from the fit says which third a given dataset is in. A rule whose error is whatever the structure happens to be is a rule that has delegated the decision to something nobody looked at.
The population of structures is part of the claim
A sweep over random structures reports properties of the process that drew them, so the process is stated first and in full rather than summarised.
Each structure has six covariates and one treatment effect of , and each fit runs on four hundred rows. Every edge weight is drawn uniform on — signed, so a route can cancel another rather than only add to it, and bounded away from zero, so a covariate assigned a role actually plays it. Each covariate is independently assigned one of six roles, at these weights: a cause of both treatment and outcome with probability 0.25, and with probability 0.15 each a cause of the treatment only, a cause of the outcome only, a cause of neither, a step on the causal path from treatment to outcome, or an effect of both.
Three of those six roles are harmless to the bias and three are not. A cause of neither does nothing. A cause of the treatment only is an instrumental variable: it adds no bias and removes variance from the treatment, which makes the estimate noisier rather than wrong. A cause of the outcome only adds no bias and removes variance from the outcome, which makes the estimate more precise. The three that matter are the common cause, which belongs in the adjustment set; the mediator, which deletes part of the effect; and the common effect, which manufactures a bias out of nothing.
Of the four thousand structures drawn, 0 were dropped as singular, so nothing in what follows is a survivorship statement about which structures could be fitted.
The number that does not depend on those weights is the one the essay is about. Change the mix and the shares move; what does not move is that the rule has no direction, because two of the three harmful roles are made worse by including the covariate and one is made worse by excluding it, and no rule expressed in terms of how many covariates to include can be right about all three. The arithmetic of that is the same covariate three ways round, and this is that arithmetic run over a population rather than over three enumerated cases.
It is worth being explicit about what the sweep inherits from that enumeration and what it does not. The three roles that matter are exactly the three worlds of the enumerated case, and their closed forms are the same closed forms; a structure of six covariates is a superposition of them, and the reason a superposition needs simulating rather than solving is that a covariate’s contribution depends on which other covariates are in the model with it. What the sweep adds is the arithmetic of that interaction over a stated population. What it cannot add is any information about which world a given covariate is in, since one covariance matrix holds two effects and every structure drawn here has a companion that would produce the same data with a different answer.
A fourth thing the sweep does not model is a treatment that was assigned rather than observed. Every structure here has arrows into the treatment, because the population of structures was drawn without a design; a real confounder created by a design is the case where the arrow into the treatment was put there by the experimenter, and it is the one case in which somebody knows the structure because somebody built it.
A rule with no direction
65.5% of structures are made worse by adjusting for everything and 33.8% are made better; 0.7% are left where they were, which is the share on which the drawn roles happened to make the two adjustment sets equivalent. Measured on squared error rather than on bias the picture is the same: 65.9% of structures have a higher squared error under the rule than under no adjustment at all.
Two thirds and one third is the number to carry, and the way to read it wrongly is as “the rule is right a third of the time, so it is a coin weighted against”. It is worse than that. A conservative default is a rule that errs in a known direction, so that somebody who knows it is conservative can reason about what the truth is likely to be relative to the estimate. This rule’s root-mean-square bias is 0.5064 against 0.2450 for adjusting for nothing, and its mean bias is a tenth of that — the errors point in every direction and cancel in the average while adding in the square. There is nothing to correct for and no direction to lean against.
The 0.7% that is left exactly where it was is worth a moment, because it is the only part of the distribution with a clean interpretation. Those are the structures on which the six drawn roles happened to make the two adjustment sets equivalent — a draw with no common cause, no mediator and no common effect, so that every covariate present is one of the three harmless kinds and including them changes the estimate’s precision without touching its bias. At the stated role weights that draw has probability 0.45 to the sixth, which is 0.8%, and the counted 0.7% over four thousand structures sits inside a standard error of it. It is a small check and it is the sort worth making: a share that has an arithmetic prediction and matches it is evidence that the role assignment is doing what the description says.
What the three rules cost
The back-door set — the covariates that are common causes of treatment and outcome, and no others — has a root-mean-square bias of 0.0000 over all four thousand structures. That is not a measurement that came out well. It is a property of the back door: the set blocks every route from treatment to outcome that starts with an arrow into the treatment, and blocks no causal route, so what remains is the effect. It would be exactly zero at any edge weights and any role mix, and its appearance here is a check that the machinery computes what it says it computes.
Its mean squared error is 0.00356, all of which is sampling variance on four hundred rows. Adjusting for nothing costs 0.06320 — 17.764 times as much — and adjusting for everything costs 0.25974, or 73.008 times as much. The ordering is the expected one and the gaps are larger than the ordering suggests: the difference between the best rule and the worst is not a factor of two or three but of seventy.
The mean is the tail
A mean squared error over a population of structures is a summary of a very skewed thing, and reporting it alone would hide what the rule actually does.
Adjusting for everything has a median squared error of 0.09087 — about a third of its mean. Its tenth percentile is 0.00250, which is lower than the do-nothing rule’s tenth percentile of 0.00309: on its best structures it is the better rule, and it is the better rule by a little. Its ninetieth percentile is 0.71164, its ninety-ninth is 2.13884, and its worst structure of four thousand costs 5.218 — against a treatment effect of 0.5, an error whose square is five is an estimate that has nothing to do with the answer.
49.8% of the rule’s total squared error comes from its worst tenth of structures. The do-nothing rule is far better behaved in the tail: its ninetieth percentile is 0.17242, its ninety-ninth 0.34592, and its worst case 0.567, which is a tenth of the other rule’s. The back-door set’s whole distribution fits in a band — tenth percentile 0.00163, ninetieth 0.00598, worst 0.011 — because it has no bias to have a distribution of and what is left is sampling noise on a fixed number of rows.
So the comparison of means understates the case against the rule rather than overstating it. A rule that is occasionally catastrophic and usually mediocre is worse to use than its mean suggests, because the structures on which it is catastrophic are not marked.
Unbiased is not best, and the gap is 26.8 points
Here is the result that came out against the guess, and it is the one this essay would most easily have slid over.
The back-door set is unbiased on every structure drawn. It has the lowest squared error on 73.2% of them. Those are different statements about different quantities, and the distance between them is 26.8 points.
Adjusting for nothing has the lowest squared error on 18.4% of structures, and adjusting for everything on 8.4%. The mechanism is not subtle once the role list above is read again. A covariate that causes the outcome and not the treatment adds no bias to either rule and removes variance from the outcome, so including it shrinks the estimator’s asymptotic variance; the back-door set, defined by common causes, leaves it out. On a structure with no common cause at all the back-door set is empty, the do-nothing rule and the correct rule coincide in bias, and the adjust-everything rule can beat both by picking up two or three outcome-only covariates that no bias argument would have included.
What this does not say is that the back-door set is the wrong target. It says that “correct” and “lowest error at this sample size” are two criteria, that the correct set answers the first outright and the second most of the time, and that the remaining quarter is a variance argument rather than a bias one. The rule a reader should leave with is the back-door set plus the covariates that cause only the outcome — a set that is still unbiased on every structure and is at least as precise as either. That set is not measured here, and naming it as the obvious next measurement is more honest than reporting 73.2% as though it were a defect in the concept.
The same distinction runs through every design essay on this site. What blocking removes, exactly is a variance argument made before any data exists, and balancing what is known in advance is the same operation moved to design time, where it costs nothing and cannot be chosen after seeing the answer.
The rule gets worse where the rules of thumb say it is safest
One of the six roles above is drawn from the structure that a covariate measured before the treatment is about: a covariate caused by two unmeasured variables, one reaching the treatment and one the outcome. It is prior to the treatment, on no causal path, and not a common cause, so every timing-based rule of thumb passes it.
Raising the share of covariates drawn from that structure from 0% to 10%, 20% and 30% moves the rule’s failure rate from 65.5% to 68.7%, 70.8% and 73.3%, and its helping rate from 33.8% down to 26.7%. Both directions are monotone across all four settings.
The other column moves too, and it moves the opposite way. The do-nothing rule’s mean squared error falls from 0.06320 to 0.05200, 0.04096 and 0.03332 as those covariates enter, because a covariate of that shape contributes nothing to an estimate that ignores it — it is not a confounder, so the rule that adjusts for nothing is already unbiased with respect to it, and replacing genuine confounders with harmless ones is a straight improvement for that rule. Adjusting for everything falls only from 0.25974 to 0.23289 over the same range, so the ratio between the two rules widens from 4.110 to nearly seven.
Is the population arithmetic what least squares actually finds?
Everything above is computed from population covariances rather than from fitted models, which is what makes a Monte Carlo sweep over four thousand structures affordable — the randomness is in which structure is drawn rather than in the estimate produced under it. The obvious way for it to be silently wrong is that the population formula is not what a regression on finite data converges to.
Six structures were taken and each of the three rules run on each over two hundred and fifty least-squares fits of three hundred rows apiece — eighteen structure-and-rule pairs. The worst disagreement between the population arithmetic and the mean of the fits is 2.65 standard errors, which on eighteen comparisons is what a normal sample of that size produces without anything being wrong.
The spread was checked too, because a formula can predict a mean correctly and be wrong about the precision of the thing it is predicting. The asymptotic variance sits within a log-ratio of 0.239 of the spread the fits actually show, with ratios running from 0.937 to 1.270 across the eighteen pairs — a quarter either way at three hundred rows, which is the ordinary finite-sample discrepancy for a variance and not a systematic offset.
The check that would have caught a genuine error is the second one rather than the first. A bias formula derived on the wrong adjustment set would still centre correctly on structures where the sets coincide, and there are many of those; a variance formula derived on the wrong set is wrong everywhere, because the set determines how many regressors there are. This is the same discipline as two routes to every number, applied to a quantity where the second route is expensive enough that it is run on six structures rather than four thousand.
Where this does not hold
Every covariate here is measured perfectly or not at all, and most real ones are neither. A confounder observed with error is the interesting case: adjusting for a noisy version removes some of the bias and adds an attenuation of its own, the two run in opposite directions, and there is a measurement error at which adjusting stops helping. Nothing here finds it. That is the single largest gap between this sweep and the practice it is about, and it is larger than the choice of role weights.
The effect is the same for every unit, so nothing here is about a treatment that works differently in different parts of the population. In that case no single adjusted number summarises anything, and the argument the reversal that is a region makes about standardisation applies with more force: choosing a weighting is choosing a question.
The sample size is fixed at four hundred rows. The bias comparisons do not depend on it, being properties of population covariances, but every squared-error comparison does, and the 73.2% in particular is a ratio between a bias term that does not shrink and a variance term that does. At four thousand rows the correct set would win more often, and at forty it would win less. The right way to read that number is as the share at a stated sample size, not as a property of the rules.
And the whole sweep prices bias against a known truth, which nobody has. What a practitioner has instead is the estimate and its standard error, and the standard error is computed under the assumption that the model is the right one — so the rule that adjusts for everything reports a narrower interval as it adds outcome-only covariates while the bias it is manufacturing does not appear anywhere in the arithmetic. That is the same failure as a curve that survives censoring: a description that is internally consistent, correctly computed, and about the wrong quantity.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The variable the treatment caused — both name adjustment set, bias, causal diagram, covariate adjustment, mediator
- Conditioning on what the treatment caused — both name causal diagram, confounding, mediator
- The gap a sample shows — both name bias, mean squared error, monte carlo
- A basis is a subspace — both name monte carlo, variance reduction
- A charge that reads the draw — both name mean squared error, monte carlo
- A criterion is a prediction of the hold-out — both name mean squared error, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Adjustment setAsymptotic varianceBack door pathBiasCausal diagramColliderConfoundingCovariate adjustmentInstrumental-variableMean squared errorMediatorMonte CarloPrecisionVariance reduction