Either model, but not neither
Worth reading first: A score that balances.
Two nuisance models can be fitted on the way to a treatment effect: one for how the treatment was assigned, and one for how the outcome behaves. The augmented estimator combines them, and the theorem attached to it says it is consistent if either one is right. Counted over six hundred draws in each of the four cells made by getting each model right or wrong, that is exactly what happens: the bias is −0.0085 where both are right, −0.0083 where only the outcome model is, −0.0016 where only the propensity model is, and 0.8118 where neither is.
The two estimators it is built from do not manage that. The outcome regression alone is off by 0.8064 whenever it is the wrong model, and weighting alone by 0.8190 whenever the propensity model is the wrong one. Each of them is excellent in half the table and useless in the other half; the augmented estimator is excellent in three quarters of it.
The fourth cell is the one worth pausing on. Where neither model is right, the augmented estimator’s 0.8118 sits between its two components’ 0.8064 and 0.8190 rather than below both. The theorem promises nothing there and delivers nothing, and it delivers nothing in the most ordinary way available: an average of two wrong answers.
That is worth stating as a positive result rather than as an absence, because a reader who has met double robustness described as insurance may expect the fourth cell to be catastrophic — two wrong models compounding into something worse than either. It is not. Both components are off by about 0.81 in the same direction, for the same reason, since both omit the same covariate; the augmented estimator’s correction has nothing to correct with and lands between them. A pairing of misspecifications that pushed the two components in opposite directions would produce a different fourth cell, possibly a much better one and possibly a much worse one, and the table as measured says nothing about that case. What it does say is that in the ordinary case, where both models are wrong because the same variable is missing from both, the extra machinery is neither insurance nor a hazard: it is arithmetic on two answers that are already wrong.
What the correction actually is
The outcome-regression estimator fits each arm’s outcome surface and averages the fitted difference over the sample. The weighted estimator ignores the outcome model and reweights. The augmented estimator is the first with a weighted correction added:
The two extra terms are inverse-probability-weighted residuals from the outcome model. If the outcome model is right those residuals have mean zero in each arm whatever weights are applied, so the correction adds noise and nothing else. If the propensity model is right the weighting makes the correction repair whatever the outcome model got wrong. Only if both are wrong does the arithmetic have nothing to lean on.
There is a way of reading that correction which is more useful than “an addition”, and it came out of a defect in the machinery here rather than out of theory. The outcome-regression estimator’s standard error was first written as the spread of the fitted arm differences over the root of the sample size. That is a perfectly natural-looking formula and it is not a sampling error at all — it is the spread of the effect surface across the covariates, which is a property of the population and does not shrink towards anything. The interval built from it covered well under half the time, and it looked like a finding about outcome regression rather than a bug.
Written properly, the estimator is a function of two fitted coefficient vectors, and its influence function has to carry their estimation error. Doing so gives a standard error of 0.1011 against an actual spread of 0.1034, and an interval covering 94.5%. And the shape that appears when it is written out is the same shape as the augmentation term. The augmentation is not an addition to the outcome regression; it is the part of it that a naive standard error was already leaving out. That reading is worth more than the theorem’s usual statement, because it says why the correction exists rather than what it guarantees.
The check a four-cell table has no closed form for
Everywhere else in this field a counted number has an integral beside it, computed on a grid with no sample in it, so that neither route can confirm itself.
A bias table does not have one. There is no integral that returns “what a logistic fit omitting the second covariate, combined with a linear outcome fit omitting the same covariate, does to an augmented estimator at six hundred rows” — the quantity is defined by a fitting procedure at a finite sample and has to be counted.
So the second route here is a different one, and it is the reason the defect above was caught. Every estimator’s standard error is written out as an influence function and checked against the spread of the estimates themselves. Those are two independent statements: the first is a formula applied inside one sample, the second is the observed scatter of six hundred separate samples, and they share no arithmetic. Where both models are right the outcome regression reports 0.1011 against a spread of 0.1034, and the augmented estimator 0.1099 against 0.1165 — agreement to a few per cent, which is what a correct influence function looks like.
The weighting row is where the two routes disagree: 0.2674 reported against 0.1915 observed. That disagreement is not a bug, and the essay on estimated weights explains it; the point here is that the check reports it rather than absorbing it. A gate that only asked whether each interval covered near 95% would have passed the weighting row at 99.2% and called it success.
And it is the check that failed loudly on the broken standard error. A reported spread three times smaller than the observed one is not a subtle discrepancy — it is the kind of thing that shows up the moment two routes are required to agree, and the kind of thing that never shows up at all when a formula is trusted because it looks like a formula.
Which intervals cover, and which do not
Bias is not what a reader pays. An interval is, so every cell is also read as a coverage.
Where a model is wrong, the intervals do not merely widen — they collapse. The outcome regression covers 0.0% of the time once it is the wrong model, weighting 1.5%, and the augmented estimator 0.2% in the cell where neither is right. Zero per cent of six hundred draws is not a rounding of a small number; it means the bias is several interval widths, every time, and a reader looking at any one of those intervals sees a perfectly ordinary result.
That is the real content of “not neither”. The failure in the fourth cell is not that the estimate is off by 0.81; it is that the estimate is off by 0.81 and the interval says it is off by about 0.15. Nothing in the output announces which cell the analysis is in, because the cell is decided by a covariate nobody has. The same silence is what the essay that ran one regression through three causal structures found: the arithmetic is identical in all three, and only the structure says which answer it is producing.
One cell covers too well rather than too badly. Weighting alone, with both models right, covers 99.2% — and that is not the interval succeeding. Its average standard error is 0.2674 against an actual spread of 0.1915, so the interval is about forty per cent too wide. The reason is that the interval is computed from the formula for known weights and the weights were estimated, and estimated weights have a smaller variance than true ones. That sentence sounds backwards and it is the subject of the essay on where an estimate’s weights should come from; it arrives here as an over-coverage, which is the same fact seen from the other side.
What the theorem does not carry
Double robustness is a statement about bias in the limit. It says nothing about the spread at a finite sample, and the spread is where the augmentation term’s cost lives: those two correction terms are divided by the propensity, so their variance grows as the assignment approaches decidable.
Move the assignment rule from strength 1 to 2.5 and the table’s bias column barely changes — the augmented estimator is still at −0.0164 with both models right, and 0.0334 with only the propensity model right. The error column changes a great deal. With both models right the augmented estimator’s root mean square error goes from 0.1167 to 0.3331, against an outcome regression sitting at 0.1184, and its coverage holds at 94.2%. With only the propensity model right it goes to 0.9431 and its coverage falls to 76.0%.
So the estimator is doing what it promised at a setting where it has become expensive. A reader who chose it for its consistency has bought consistency and paid in a currency the choice was never priced in.
The comparison inside that cell is the awkward one. With both models right, the plain outcome regression sits at 0.1184 and the augmented estimator at 0.3331 — nearly three times the error, for a property neither of them needs, since the outcome model is right and its own bias is −0.0034. Double robustness is insurance, and the premium at strength 2.5 is about two hundred per cent of the loss it insures against. That is a defensible purchase when the buyer does not know which cell they are in, which is always; it is not a free one, and the usual account of the estimator does not mention a price at all.
The cell where having both is worse than either
Take one more step, to an assignment strength of 3.5, one past the end of the overlap sweep, and read the cell where the propensity model is right and the outcome model is wrong. This is the cell double robustness is for: the estimator’s whole justification is that a wrong outcome model does not sink it.
It does not sink it. The augmented estimator’s bias is 0.0857, against the outcome regression’s 1.6333 and weighting’s 0.5535. It is by a wide margin the least biased estimator on the table.
It is also the worst one on it. Its root mean square error is 1.9265, against the outcome regression’s 1.6402 and weighting’s 0.9383. A reader who took the nearly unbiased estimator over the badly biased one has doubled their expected squared error. Its interval covers 57.0%.
Being nearly unbiased and being any good are different properties, and this is the cleanest demonstration of it available on this population. Bias and spread are not comparable until they are squared, and once they are, an estimator that is right on average and enormous in a tenth of its draws loses to one that is wrong by half and steady.
The median absolute error is the number that says what “enormous in a tenth of its draws” means, and it turns the ordering over again. By median absolute error the augmented estimator is 0.6726, the weighted one 0.8143 and the outcome regression 1.6303 — so on the typical draw the augmented estimator is the best of the three, by a clear margin, and on the average of squares it is the worst. Both statements are true of the same six hundred draws. The distribution of its error is not one that a single summary describes: it is centred well and has a tail that owns the mean of its square.
There is a temptation to resolve the contradiction by picking the median and moving on, and it should be resisted. The root mean square is the right summary when the loss actually paid is squared error, which is what an interval width and a power calculation are both built on; the median is the right summary for the question “what will this estimate look like on the day”. They disagree here because the estimator has a tail, and reporting only the one that flatters it is how an estimator acquires a reputation it cannot support. The same discipline about which summary answers which question runs through the essay on what a residual is not, where a quantity that looks like an error and is measured like one turns out to be systematically the wrong size.
That is why the root mean square at this setting is not quoted alone. It moves between about 1.9 and 2.7 depending on which block of seeds is drawn, because it is an average of a heavy-tailed quantity; the median moves by about two per cent, and the cell where the augmented estimator is worse than both of its components fires in every block tried.
Where the cost comes from, and it is not the augmentation
The variance that ruins the augmented estimator at strength 3.5 is not a peculiarity of the correction term. It is the same quantity that ruins the plain weighted estimator, arriving through the same denominator.
The essay on the region with no comparison prices that failure directly, and its conclusion applies without modification here: past a certain thinness there is no estimator built on these weights that is not paying for a comparison the sample does not contain. Adding an outcome model does not conjure the missing units; it produces a correction term that is asked to reweight residuals in a region where the reweighting is a single observation carrying twenty-five million.
The reason the augmented estimator can be worse than the plain weighted one, rather than merely as bad, is that the weighted estimator here divides each arm’s total by its own summed weights and the augmentation term does not. The stabilisation that damps the weighted estimator’s tail is not available to the correction, so the correction’s tail is the one that survives into the root mean square. That is a fact about this particular pairing rather than about augmentation in general, and it is worth stating as such.
What “wrong” means here, and what it does not
Every wrong model in this table is wrong in exactly one way: it omits the second covariate, which both the assignment rule and the outcome depend on. That is a real misspecification and it is one kind.
A model with the right variables and the wrong functional form — a curved outcome fitted as a line, an assignment rule with an interaction fitted without one — biases the two components differently, and there is no reason to expect the four-cell table to keep its shape. In particular the cell where the augmented estimator is worse than both of its components was found by pushing one scalar until the augmentation’s spread exceeded the outcome regression’s bias; a different kind of wrongness moves both of those quantities and could move the crossing anywhere, including out of existence. Nothing here measures that, and it is the largest gap in this essay.
Two smaller things are worth naming. The intervals for the augmented estimator treat both nuisance fits as known, which is the usual reading and is conservative where the propensity model is right; in the cells where it is wrong the interval is not valid for anything and the coverage says so rather than the arithmetic. A 95% interval is a promise about a procedure repeated, and the essay on what that 95% refers to is the reason every cell here is counted over six hundred draws rather than argued about: a promise of that shape can only be checked by running the procedure. And the 99.2% over-coverage in the weighting row is a symptom of the same thing the decomposition above measures — that a fitted propensity is a different estimator from a known one, with a smaller variance, so a standard error derived for the known case is the wrong one. Nothing in this table adjusts for it.
Finally, none of the four cells is a cell anyone occupies knowingly. The whole table assumes unconfoundedness — that holding the recorded covariates fixed makes the assignment as good as random — and every cell in it is a statement about what happens when a recorded covariate is left out of a model. A covariate that was never recorded puts the analysis outside the table entirely, in a place where all three estimators agree with each other and none of them is right. The essay that priced adjusting for every covariate on hand and the essay on two worlds that produce identical data are where that boundary is measured; the balancing arithmetic this field opens with is exact and cannot see across it. Only assigning by a coin makes the question go away, and that is not available in the setting any of this is for.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- How many observations a weight leaves — both name inverse-probability weighting, propensity score, stabilised weights
- The draws aimed at the tail — both name confidence interval, heavy tail, standard error
- A coverage table with its own error — both name confidence interval, standard error
- A tenth as wide, and both of them right — both name confidence interval, standard error
- An imputation model the analysis does not contain — both name confidence interval, model misspecification
- An interval that carries its scale — both name confidence interval, standard error
Named objects
A flat tag is an object no other essay names yet.
Augmented estimatorAverage treatment effectConfidence intervalDoubly robustHeavy tailInfluence functionInverse-probability weightingModel misspecificationNuisance modelOutcome regressionPropensity scoreRoot mean square errorStabilised weightsStandard errorUnconfoundedness