Adjusting for a shadow
Worth reading first: Balancing what is known in advance · One arithmetic, three decisions.
Every adjustment set on this field has been a list of variables the analysis has in hand exactly. The essay that opened this question put a covariate beside a treatment and an outcome and asked what role it plays; the third asked what happens when everything measured goes into the regression. Both took for granted that a variable named in a model is the variable it is named after.
Almost no applied covariate is. A test score stands in for ability, an income band for wealth, a diagnosis code for a condition, a single blood pressure reading for a person’s blood pressure. Each is the thing plus noise, and the question is not whether adjusting for one helps — it plainly does — but how much of the confounding it leaves, and whether anything in the output says so.
The answer is worse than the obvious guess, and the guess is obvious enough to be worth stating: a covariate that is 80% signal ought to remove about 80% of the problem.
The two agree only at the ends, and they disagree most in the middle, which is where every applied covariate sits.
The closed form, and what the second term is
The share removed is
where λ is the proxy’s reliability — the share of its variance that is signal — and is the squared correlation between the treatment and the confounder.
The second term is the part with no counterpart in the guess, and it says something a reader should find uncomfortable: the more strongly the confounder drives the treatment, the less of the confounding a proxy of given quality removes. At = 0 the formula collapses to λ, but at = 0 the confounder does not affect the treatment and there is no confounding to remove. The formula equals the guess exactly in the case where the question does not arise.
The mechanism is that adjustment works through the part of the proxy that predicts the treatment. The treatment carries the confounder’s signal too, so a regression that includes both has to separate the confounder’s contribution to the outcome from the treatment’s, and the noisier the proxy, the more of that separation is attributed to the treatment. Strong confounding means more to misattribute.
Here is 0.4475, which is a confounder explaining about 45% of the treatment’s variance — strong, and not implausibly so for the kind of variable people adjust for.
The formula is also the reason the defect is hardest to see where it is worst. A study with severe confounding has a large gap between its adjusted and unadjusted estimates, which reads as evidence that the adjustment did a lot of work — and by the formula, that is exactly the situation in which the proportion left behind is largest. The visible sign of a working adjustment and the invisible sign of an incomplete one are the same sign.
Where the guess comes from, and why it is wrong
The intuition behind the diagonal is worth reconstructing, because it is a reasonable intuition and its failure is specific.
A proxy of reliability λ can be written as the confounder plus noise, and regressing on it is — in the simplest possible reading — like regressing on a λ-weighted version of the confounder. The coefficient on a mismeasured regressor is attenuated towards zero by exactly λ, which is the classical result every applied course teaches, and the guess extends it: if the covariate’s own coefficient is attenuated by λ, the correction it applies to the treatment’s coefficient should be too.
That step does not hold, and the reason is that the treatment and the proxy are correlated. A multiple regression’s coefficient on one variable depends on what the others explain, so the confounder’s under-corrected contribution does not simply vanish — it is reallocated, and the variable it is reallocated to is the treatment, because the treatment is what correlates with what the proxy failed to capture. The stronger that correlation, the more is reallocated.
That reallocation is the in the denominator, and it turns a shortfall into a bigger shortfall. It is the same mechanism a mediator on the path exhibits from the other side: a regression distributes explanation among its regressors by correlation, and whichever regressor is best correlated with what is missing absorbs it.
What a reader actually sees
The numbers a study would print are 0.6084 against a truth of 0.5000, and the sentence that accompanies them is “adjusted for the confounder”. Everything about that analysis is correct as described. The variable is in the model, its coefficient is significant, the adjustment moved the estimate in the right direction by a visible amount, and the remaining error is 0.1084 — twenty-two per cent of the effect.
The direction is the awkward part. Adjusting for a proxy moves the estimate towards the truth, so the analysis looks like it is working, and the residual is a smaller version of the original bias rather than a new artefact. There is no sign change, no implausible magnitude, no diagnostic that fires. It is the shape of defect this site finds hardest to see: the symptom is the absence of the rest of a correction.
The simulation agrees with the closed form — 0.6071 against 0.6084 over 1,200 samples of two thousand rows — which is the check that the arithmetic above is the arithmetic of the world rather than a formula that happens to fit.
More data makes it worse
The residual bias is a property of the population, not of the sample. It does not shrink as data accumulates. The standard error does.
At a hundred rows the bias is 1.10 standard errors and the confidence interval covers the truth comfortably — not because the analysis is right but because it is imprecise. At 1,600 rows it is 4.40 standard errors and the interval excludes the truth. At 25,600 it is 17.61.
A larger study of a mismeasured confounder is a study that is confidently wrong. The coverage of its interval goes to zero in the sample size, which is the opposite of the property an interval is chosen for, and the reason is that every term in the standard error is about sampling and none is about the measurement.
This is the same arithmetic that makes a bias that does not average out worse than a variance, and it is the reason a residual confounding bias is not something a study can out-collect.
The size of the problem across the range
Collecting the closed form at eight reliabilities gives the table the rest of this essay is read against. At a reliability of 0.2 a proxy removes 12.14% of the confounding; at 0.4, 26.92%; at 0.6, 45.32%; at 0.8, 68.85%; at 0.9, 83.26%; at 0.95, 91.30%.
Two features of that column are worth naming. The gap between the curve and the diagonal is largest around a reliability of 0.6 to 0.7 — where it is about fifteen points — and it closes at both ends, so the guess is least wrong exactly where it matters least. And the curve is convex: improving a poor measurement buys less than improving a good one, which is the opposite of the usual returns to effort and means that a covariate measured badly is nearly worthless as an adjustment rather than partially useful.
A reliability of 0.2 is not a straw man. It is roughly what a single noisy reading of a variable that fluctuates day to day has as a measure of that variable’s usual level — one blood pressure reading for habitual blood pressure, one week’s income for permanent income — and such variables are routinely entered into adjustment sets as though they were the thing.
The one lever available
The confounder cannot be measured better by wishing. What can sometimes be done is measuring it more than once, and the arithmetic of averaging is generous at the start and stingy after.
Averaging k independent measurements of reliability λ gives one of reliability kλ/(1 + (k − 1)λ), so two measurements at λ = 0.5 are worth one at 0.667 and ten are worth one at 0.909. The share of bias removed runs 35.59%, 52.49%, 62.37%, 73.42%, 84.67%.
The second measurement is the valuable one: it removes seventeen points more of the bias than the first did, where the tenth removes about two. That is the usual shape of an averaging argument and it has a practical consequence worth stating — a study that can take a covariate twice on a subsample and nowhere else should take it twice on everybody instead, because the gain is front-loaded.
The arithmetic also says which repetitions count. The measurements have to be independent errors around the same quantity — two blood pressure readings on different days, two graders scoring the same essay — and not two readings that share whatever made the first one wrong. Two readings taken a minute apart repeat the instrument’s error rather than averaging it away, so they raise the apparent reliability of the average without raising the real one, and the adjustment gets worse while the diagnostic that would have caught it reports agreement.
And ten measurements still leave 15% of the bias. Reliability approaches one and does not reach it, so this lever narrows the problem rather than closing it.
What a defensible analysis looks like
Treat a measured confounder as a partly unmeasured one. The whole apparatus of sensitivity analysis exists for confounders nobody observed, and a proxy of reliability 0.8 is a confounder 80% observed. The residual is the same kind of object as an unmeasured confounder’s, with the advantage that its size is computable rather than postulated — which makes this the one sensitivity analysis that needs no assumed parameter.
Say what the covariate measures and how well. A reliability is a number many applied fields already have for their standard instruments, from test–retest studies or from internal consistency, and it is almost never carried into the analysis that adjusts for them. With it, the formula above turns “adjusted for X” into a statement with a residual attached.
Prefer two coarse measurements to one refined one, at equal cost. The averaging arithmetic says the second measurement removes seventeen points more of the bias at λ = 0.5, and no refinement of a single instrument is likely to move λ that far. Where a study can choose how to spend measurement effort, repetition beats precision until the reliability is already high.
Report the adjusted and the unadjusted estimate together. Their difference is the bias the adjustment removed, and the formula says what share of the total that was, so the two together bound the whole. Here 0.8481 and 0.6084 differ by 0.2397, which at a reliability of 0.8 is 68.85% of the bias — so the total was 0.3481 and 0.1084 remains. That calculation needs one number the analysis does not usually carry and two it already has.
Do not let a significant covariate coefficient stand in for a good covariate. A proxy of reliability 0.2 still has a coefficient many standard errors from zero at any realistic sample size, because a fifth of a strong confounder is still a strong predictor. The coefficient’s significance is evidence that the covariate is related to the outcome and no evidence at all that it has absorbed the confounding — the same distinction a collider exhibits when it enters a regression, where a large coefficient accompanies a larger bias.
And treat a large sample as a reason for more caution, not less. The usual reading of a tight interval is that the estimate is reliable. Where the adjustment set is measured with error, a tight interval around a biased estimate is the worst of the available outcomes, and the tightness is what makes it worst.
The three quantities a reader needs
Everything above reduces to three numbers, and a study that reports them turns an uninterpretable adjustment into a bounded one.
The unadjusted estimate. It is the starting point of the correction and it is almost always computed and often not shown.
The adjusted estimate. The difference between the two is the correction the proxy applied, and it is an observable quantity rather than an assumption.
The proxy’s reliability. With it, the formula says what share of the whole that correction was, and therefore what remains. Without it, the correction’s size is a number with no denominator: an adjustment that moves an estimate by 0.24 could have removed nearly all of the bias or a third of it, and nothing else in the output distinguishes those.
The third number is the one that is missing, and it is the only one that has to come from outside the study. That is why this defect persists in fields that have the number and do not carry it: the reliability lives in a methods literature about instruments and the adjustment lives in an analysis, and nothing puts them on the same page.
What is claimed here and what is not
Everything here is linear and Gaussian. The formula is exact for a linear structural model with normal disturbances and additive measurement error. With a binary proxy, a nonlinear outcome model or measurement error correlated with the confounder’s own level, the direction survives — a noisy adjustment under-adjusts — and the formula does not.
The measurement error is classical. The proxy is the confounder plus independent noise, which is the favourable case. Error correlated with the treatment or the outcome adds bias terms of its own that can point either way, and nothing here prices them.
One arrangement of the confounding. = 0.4475 and the effect is 0.5 against an unadjusted 0.8481, which is severe confounding. Weaker confounding gives a smaller residual in absolute terms and a larger share removed, since the formula’s second term matters less — so the share removed is not a constant of the method but a function of the world, and a study cannot read it off its own output.
Reliability is being used in its psychometric sense. λ is the share of the proxy’s variance attributable to the thing it measures, which is what a test–retest correlation estimates and what an internal-consistency coefficient approximates. Fields that report a reliability report this quantity; fields that report a measurement’s standard deviation report something that has to be divided by the construct’s variance to become it, and the conversion needs a quantity that is itself unobserved.
And the simulation is a check on the algebra, not a separate finding. The closed form is computed from the population covariance matrix by solving the normal equations exactly; the simulation draws finite samples and fits them. Their agreement to within 0.0013 says the matrix was assembled correctly, which is the one thing a derivation can get wrong invisibly.
Why this is worse than an omitted covariate
One comparison puts the finding in proportion, and it goes the uncomfortable way.
A confounder left out entirely produces a bias a careful reader is watching for. The analysis says “unadjusted”, the reader knows to discount it, and a sensitivity analysis can ask how strong an unmeasured confounder would have to be to explain the result away.
A confounder adjusted for badly produces a smaller bias and removes the watchfulness. The analysis says “adjusted”, the reader stops discounting, and the standard sensitivity analysis is not run because the confounder is not unmeasured — it is right there in the model with a significant coefficient.
So the defect converts a visible problem into an invisible smaller one, and whether that is a net improvement depends on how the result is used. The estimate is closer to the truth and the claim made about it is much stronger, and the second effect can be the larger. It is the same trade a pre-test estimator makes: a procedure that usually improves the estimate can worsen the inference by changing what the analyst believes about it.
Still open: what the residual does to an adjustment set
Every measurement here is of one proxy for one confounder. The situation that produced the third essay in this field is a set of covariates, several of them proxies, adjusted for together — and the arithmetic of that is not the arithmetic of this one repeated.
Two effects should compete. Several proxies of the same confounder behave like the averaging above, which helps. Several proxies of different confounders each leave their own residual, and the residuals need not point the same way, so they can cancel or accumulate. Which happens depends on how the confounders are correlated with each other and with the treatment, and that is a structure a study has no more access to than it has to the confounders themselves.
There is a specific reason to think the answer is unfavourable. “Adjust for everything measured” already leaves a larger bias than adjusting for nothing on two-thirds of randomly drawn structures, measured with every covariate observed exactly. Adding measurement error to each of them cannot make the set more informative, and it can make a covariate whose role was ambiguous look harmless. What that does to the two-thirds is the measurement this field has not made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The variable the treatment caused — both name adjustment set, bias, causal diagram, covariate adjustment, latent variable, unmeasured confounding
- The measurement that got them enrolled — both name closed form, covariate adjustment, measurement error, monte carlo, test–retest reliability
- Two analyses of one baseline — both name closed form, covariate adjustment, measurement error, monte carlo, test–retest reliability
- A lead that a heavy tail keeps — both name closed form, measurement error, monte carlo, test–retest reliability
- The slope of a density nobody can see — both name closed form, measurement error, monte carlo, test–retest reliability
- What a wrong model estimates — both name asymptotic bias, bias, closed form, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Adjustment setAsymptotic biasBiasCausal diagramClosed formCovariate adjustmentEstimation errorLatent variableMeasurement errorMonte CarloPartial correlationSensitivity analysisTest–retest reliabilityUnmeasured confounding