A set of shadows
Worth reading first: Balancing what is known in advance · One arithmetic, three decisions.
Adjusting for a shadow measured one confounder through one proxy. A covariate that is 80% signal removes 68.85% of the confounding it stands for, not 80%, and the share is , with the proxy’s reliability and the squared correlation between the treatment and the confounder. That essay ended on a worry about sets. An applied adjustment set holds several covariates, most of them measured with error. The worry was that their residuals need not point the same way, and that what was already a bad rule — adjusting for everything measured, which left a larger bias than adjusting for nothing on 65.5% of four thousand random structures — could only get worse once every covariate was a shadow of itself.
That is a prediction, and it can be checked with the same four thousand structures. Each has six covariates, each drawn into a role — a cause of both treatment and outcome, a cause of one, a cause of neither, a step on the causal path, or an effect of both — with edge weights drawn away from zero and the treatment’s effect fixed at 0.5. Measurement error enters in the simplest way it can. Each observed covariate is the true one plus independent noise whose variance makes the reliability , and in covariance terms that changes nothing except each covariate’s own variance, which is divided by . Every structure is then solved on its inflated covariance exactly as before, under the three rules: adjust for nothing, adjust for everything measured, and adjust for the common causes, which is the set the causal diagram says is right.
The prediction fails on both counts. The right set does no harm at any common reliability, and adjusting for everything measured gets better as the covariates get worse.
Both adjusting rules close on doing nothing
Adjusting for nothing does not care how well the covariates are measured, because it uses none of them. Its mean absolute bias stays at 0.190 across the whole range. The other two rules move toward it from either side. Adjusting for the common causes starts at exactly zero with exact covariates, reaches 0.056 at a reliability of 0.8 and 0.118 at 0.5. Adjusting for everything measured starts at 0.377, nearly twice the bias of doing nothing, and falls to 0.279 at 0.8 and 0.203 at 0.5, where it is barely worse than leaving every covariate out. It beats adjusting for nothing on 33.8% of structures with exact covariates, 42.0% at 0.8 and 50.8% at 0.5.
Both movements have one cause. Measurement error weakens an adjustment, whatever the adjustment was doing. A noisy confounder removes less of the confounding, which is the shadow essay’s result again, and it is why the right set drifts upward. But a noisy collider also opens less of the path through it, and a noisy mediator blocks less of the effect running through it. Adjusting for everything measured did its damage almost entirely through those two roles. A collider manufactures a dependence between the treatment and the outcome when it is held fixed, as a sample selected on one does. Holding a noisy version fixed holds the collider only partly fixed, so it manufactures less. A mediator carries part of the effect, and adjusting for it deletes that part. Adjusting for a noisy version deletes less of it. 3,520 of the four thousand structures carry at least one variable the treatment causes. In those, adjusting for everything averages a bias of 0.428 with exact covariates, 0.306 at 0.8 and 0.211 at 0.5. Noise is a partial cure for adjusting for the wrong thing, for the same reason it is a partial failure in adjusting for the right one.
The same pattern shows up in mean squared error at 400 rows, which also counts the variance each adjustment costs.
Adjusting for everything measured has 4.11 times the squared error of adjusting for nothing with exact covariates, 2.17 times at 0.8 and 1.11 times at 0.5. The common causes have 0.056 times it with exact covariates, 0.145 at 0.8 and 0.434 at 0.5. Divide one by the other and the right set’s advantage over the everything rule falls from a factor of 73.0 to 14.9 at 0.8 and 2.6 at 0.5. Measurement error does not make the choice of adjustment set worse. It makes the choice matter less, and at the reliabilities common in survey and administrative data the three rules are closer together than any diagram drawn with exact variables would suggest.
That is not good news. It means the work of choosing the right set — eliciting the diagram, arguing about which variable caused which, refusing to adjust for a variable measured after treatment — buys much less than it appears to on paper whenever the covariates in hand are proxies. At reliability 0.5 the right set’s mean bias is 62% of what doing nothing leaves, and the median structure keeps 60.6% of its own unadjusted bias. No amount of care about the diagram recovers it.
Why the right set never does harm when every proxy is equally good
The worry said several proxies of different confounders would each leave a residual, and the residuals could cancel or accumulate. They can, but at a common reliability they cannot do the one thing that matters: make the right set worse than nothing. Across four thousand structures and six reliabilities, it never happens once.
The reason is a closed form that generalises the shadow essay’s. Take independent confounders , each reaching the treatment with weight and the outcome with weight , observed through proxies of reliability , with treatment noise of variance . Adjusted for all the proxies together, the bias in the treatment’s coefficient is
The numerator is the unadjusted bias’s own pieces, , each scaled by the share its proxy misses. The denominator is the treatment’s variance once the proxies are partialled out. With a single confounder the expression reduces to the shadow essay’s share removed. With every equal, the numerator is times the unadjusted numerator, while the denominator stays at least as large as . The bias left is therefore the unadjusted bias times a factor that is never above one. Pieces that cancelled unadjusted still cancel, and pieces that accumulated are each shrunk alike. The right set, measured equally badly throughout, is a weakened version of a correct adjustment, never a different one.
Two confounders that cancel
The closed form also says exactly when the guarantee fails. If the reliabilities differ, the pieces are scaled unequally, and a cancellation between them is undone. The cleanest case has two independent confounders whose biases cancel exactly. One reaches the treatment with weight 1.0 and the outcome with 0.5; the other reaches the treatment with 0.4 and the outcome with −1.25. Their pieces are +0.231 and −0.231, so the unadjusted estimate is exact, and adjusting for both true confounders is exact as well. Both analyses a careful reader would trust return 0.5.
Hold the first proxy at reliability 0.9 and the bias of adjusting for both proxies is zero only where the second is also at 0.9. At 0.7 it is −0.087. At 0.5, a typical reliability for a questionnaire scale, it is −0.169, a third of the effect, from an analysis whose unadjusted version had no bias at all. As the second proxy approaches pure noise, the bias tends to −0.357, which is the bias of adjusting for the first proxy alone. That is the right limit, since a proxy of reliability near zero adjusts for nothing. Reverse the qualities — the first proxy at 0.5, the second at 0.9 — and the bias is +0.132, the other way.
Drawn row by row rather than solved, four hundred studies of 2,000 rows with the proxies at 0.9 and 0.5 give a mean bias of −0.1699 against −0.1695 in closed form, and the unadjusted estimate a mean bias of −0.0018 within two standard errors of zero. The spread from study to study is 0.027, so a single study of that size reports the biased estimate about six standard errors from the truth. Nothing in that study’s output flags the problem. Its covariates are the right ones, they are measured as well as such covariates usually are, and the unadjusted estimate it might print as a robustness check differs from the adjusted one by about a third of the effect, which a reader would take as evidence that the adjustment was needed.
The arithmetic says something unwelcome about the usual habit of adjusting for whatever is measured best. If one confounder has a clinical measurement and another a self-report, adjusting for both does not give half the benefit of each. It removes most of the first confounder’s piece and little of the second’s, and when those pieces pointed opposite ways it manufactures a bias out of a balance. This is the same paradox of conditioning that made a pre-treatment collider dangerous, arriving through measurement rather than through the diagram: the adjustment is right, and incomplete in an uneven way.
How often unequal measurement does harm
The cancelling pair was built to show the mechanism. Whether it matters depends on how often confounders oppose each other and how unequal the measurements are. Back in the random structures, give each covariate its own reliability, drawn uniformly from , and widen from 0 to 0.25. The draws are made once per structure from a stream of their own, so every spread sees the same structures and the same relative qualities, only stretched.
With every covariate at 0.75 the right set never does harm. At a spread of 0.05 it does harm on 1.2% of structures, at 0.15 on 4.6% and at 0.25 — reliabilities anywhere from 0.5 to 1 — on 7.15%. Every one of those structures has confounders pushing opposite ways, and 28.4% of the four thousand do. Among them the share harmed reaches 25.1% at the widest spread: one structure in four whose confounders oppose is made worse by adjusting for exactly the right variables. The harm is usually modest, with a median of 0.028 against an effect of 0.5, but the largest is 0.160.
The aggregate picture hardly moves while this happens. Across all four thousand structures, the right set’s mean absolute bias goes from 0.068 at a common reliability to 0.072 at the widest spread, and its median share of the bias removed from 66.1% to 62.9%. A summary over structures shows a small, gradual deterioration. The structures underneath have split into a large majority helped as before and a minority hurt outright, and the minority is exactly the set of structures where the unadjusted estimate looked best.
Where adjusting for everything picks up a new cost
One more movement is hidden in the first figure. In the 480 structures with no mediator and no collider, adjusting for everything measured was exact with exact covariates, because it adjusted for the confounders plus variables that are harmless to include: causes of the treatment only, causes of the outcome only, causes of neither. Under noise those structures are no longer safe. At reliability 0.8, adjusting for everything in them averages a bias of 0.079 against 0.065 for the confounders alone, and it does worse than the confounders alone on 72.3% of them — 347 structures, which are exactly the 347 that carry both a confounder and a cause of the treatment only.
The culprit is the cause of the treatment only — an instrument, in the language of the essay on what conditioning means. Adjusting for an instrument removes variation from the treatment that had nothing to do with the confounders. What is left of the treatment is then more dominated by whatever confounding remains, so the remaining bias is a larger share of the remaining variance. With exact confounders there is no remaining confounding to amplify and the instrument is merely wasteful. With proxies there always is, and every instrument in the adjustment set enlarges it. The denominator in the closed form is the place to see it: an instrument in the set shrinks the treatment’s residual variance without touching the numerator.
The data cannot say which shadows are sharp
Everything above takes the reliabilities as known. A study usually does not know them, and its own data cannot supply them. A proxy of reliability 0.5 for a strong confounder and a proxy of reliability 1 for a weak confounder can produce the same covariance among treatment, covariate and outcome, because the covariance sees a confounder only through its proxy, and a strong confounder seen dimly and a weaker one seen sharply leave the same products behind. That is the observational equivalence that separated three causal structures arriving through a different door: there it was the direction of arrows the covariance could not see, here it is how much of each covariate is signal. The adjusted estimate is a function of the observed covariance, and the covariance is the same whichever reading is true, so the estimate is the same too. The bias differs between the two readings.
So the reliabilities have to come from outside the study: a test–retest sample, a validation subsample measured both ways, a published instrument’s manual. And the closed form says which outside number matters most. It is not the average reliability of the set, and not the reliability of the strongest confounder’s proxy, but the spread of reliabilities across confounders whose pieces have opposite signs. A set with every proxy at 0.6 is safer, in the specific sense of never doing harm, than a set with one proxy at 0.95 and another at 0.6 — even though the second set is better measured on average, and would be preferred by any rule that looks at reliabilities one covariate at a time.
What a study adjusting for proxies should report
The reliability of each covariate, not of the set. The guarantee that the right set does no harm rests on equal reliabilities. A study that reports one average reliability, or none, cannot say whether its adjustment shrank the bias or rearranged it.
Which confounders are thought to push in which direction. Opposing confounders are where unequal measurement does its damage, and where the unadjusted estimate is most misleadingly close to the truth. The direction of each is often arguable from subject knowledge even when its size is not.
The adjusted estimate with the best-measured covariates alone, and with the worst alone. For the cancelling pair those are −0.357 and +0.120 against an adjusted −0.169. The spread between them is a direct measure of how much the answer depends on which shadows were sharp.
Why each variable that causes only the treatment is in the set. With proxies for the confounders, every such variable amplifies what the proxies missed.
What is claimed and what is not
Proved. With independent confounders observed through independent proxies and adjusted for together, the bias is the sum of each confounder’s piece times one minus its proxy’s reliability, over the treatment’s residual variance. With a common reliability it is the unadjusted bias times a factor in , so the right set cannot do worse than nothing.
Computed exactly. Every bias and squared error on the four thousand structures, from each structure’s covariance with its covariate variances divided by their reliabilities, and the cancelling pair’s curve from its four-by-four covariance; the closed form matches the matrix solve to at every pair of reliabilities checked.
Counted. Four hundred studies of two thousand rows of the cancelling pair, with each proxy drawn as its confounder plus noise, landing on the closed form.
Not claimed. The random structures draw their confounders independently of each other and each covariate’s noise independently of everything else. Correlated confounders, noise that is correlated across covariates — two self-reports sharing a respondent’s habits — and structures with latent common causes of the covariates are not in the sweep.
Still open: noise that the covariates share
Every proxy here carries noise of its own. Real covariates often share it. Two questionnaire items answered by the same person carry the same mood and the same habit of agreeing. Two administrative fields filled in by the same clerk carry the same shortcuts. Shared noise is itself a common cause of the proxies, so it adds a latent variable to the diagram that reaches every covariate it touches, and adjusting for several of them conditions on that latent variable through all of them at once. Whether shared error makes the right set’s residual larger or smaller than independent error of the same reliability, and whether it can make the right set do harm even when every reliability is equal — which the closed form above forbids for independent error — is a measurement these covariances can make and have not yet made. A correction built on a retest estimated a reliability from two measurements of one person. A retest that shares the first measurement’s error overstates the reliability, which is the same problem seen from the measurement side.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two analyses of one baseline — both name closed form, confounding, covariate adjustment, measurement error, monte carlo
- The gap a sample shows — both name bias, closed form, mean squared error, monte carlo
- The measurement that got them enrolled — both name closed form, covariate adjustment, measurement error, monte carlo
- A lead that a heavy tail keeps — both name closed form, measurement error, monte carlo
- A standard error that knows about the instruments — both name closed form, instrumental-variable, monte carlo
- A width rule on skewed outcomes — both name bias, closed form, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Adjustment setBiasCausal diagramClosed formColliderConfoundingCovariate adjustmentInstrumental-variableMean squared errorMeasurement errorMediatorMonte CarloUnmeasured confounding