Weighting one sample into another

Either model, but not neither

The augmented estimator's bias is −0.0085, −0.0083 and −0.0016 wherever one nuisance model is right, against components off by 0.8064 and 0.8190. One step past the overlap sweep it is the least biased estimator on the table at 0.0857 and the worst on it at 1.9265.

Worth reading first: A score that balances.

Two nuisance models can be fitted on the way to a treatment effect: one for how the treatment was assigned, and one for how the outcome behaves. The augmented estimator combines them, and the theorem attached to it says it is consistent if either one is right. Counted over six hundred draws in each of the four cells made by getting each model right or wrong, that is exactly what happens: the bias is −0.0085 where both are right, −0.0083 where only the outcome model is, −0.0016 where only the propensity model is, and 0.8118 where neither is.

The two estimators it is built from do not manage that. The outcome regression alone is off by 0.8064 whenever it is the wrong model, and weighting alone by 0.8190 whenever the propensity model is the wrong one. Each of them is excellent in half the table and useless in the other half; the augmented estimator is excellent in three quarters of it.

Either model is enough; neither is not. The bias of three estimators of an average effect of 1.0000, over 600 samples of 600 units with the assignment rule at strength 1, in each of the four cells made by getting each nuisance model right or wrong. The wrong model in both cases is one that omits the second covariate, which the outcome and the assignment both depend on. The outcome model alone is off by 0.8064 whenever it is the wrong one; weighting alone is off by 0.8190 whenever the propensity model is. The augmented estimator built from both is off by -0.0085, -0.0083 and -0.0016 in the three cells where at least one of them is right, and by 0.8118 in the fourth — which is between its two components rather than better than either.
Fig. 1 The bias of three estimators of an average effect of one, in each of the four cells made by getting each nuisance model right or wrong, over six hundred samples of six hundred units. The wrong model in every case is one that omits the second covariate, which both the assignment and the outcome depend on.

The fourth cell is the one worth pausing on. Where neither model is right, the augmented estimator’s 0.8118 sits between its two components’ 0.8064 and 0.8190 rather than below both. The theorem promises nothing there and delivers nothing, and it delivers nothing in the most ordinary way available: an average of two wrong answers.

That is worth stating as a positive result rather than as an absence, because a reader who has met double robustness described as insurance may expect the fourth cell to be catastrophic — two wrong models compounding into something worse than either. It is not. Both components are off by about 0.81 in the same direction, for the same reason, since both omit the same covariate; the augmented estimator’s correction has nothing to correct with and lands between them. A pairing of misspecifications that pushed the two components in opposite directions would produce a different fourth cell, possibly a much better one and possibly a much worse one, and the table as measured says nothing about that case. What it does say is that in the ordinary case, where both models are wrong because the same variable is missing from both, the extra machinery is neither insurance nor a hazard: it is arithmetic on two answers that are already wrong.

What the correction actually is

The outcome-regression estimator fits each arm’s outcome surface and averages the fitted difference over the sample. The weighted estimator ignores the outcome model and reweights. The augmented estimator is the first with a weighted correction added:

θ^=1ni[m^1(xi)m^0(xi)+Ti(yim^1(xi))e^(xi)(1Ti)(yim^0(xi))1e^(xi)].\begin{aligned} \hat\theta = \frac{1}{n}\sum_i \Big[\, &\hat m_1(x_i) - \hat m_0(x_i) \\ &+ \frac{T_i\,(y_i - \hat m_1(x_i))}{\hat e(x_i)} - \frac{(1-T_i)\,(y_i - \hat m_0(x_i))}{1 - \hat e(x_i)} \,\Big]. \end{aligned}

The two extra terms are inverse-probability-weighted residuals from the outcome model. If the outcome model is right those residuals have mean zero in each arm whatever weights are applied, so the correction adds noise and nothing else. If the propensity model is right the weighting makes the correction repair whatever the outcome model got wrong. Only if both are wrong does the arithmetic have nothing to lean on.

There is a way of reading that correction which is more useful than “an addition”, and it came out of a defect in the machinery here rather than out of theory. The outcome-regression estimator’s standard error was first written as the spread of the fitted arm differences over the root of the sample size. That is a perfectly natural-looking formula and it is not a sampling error at all — it is the spread of the effect surface across the covariates, which is a property of the population and does not shrink towards anything. The interval built from it covered well under half the time, and it looked like a finding about outcome regression rather than a bug.

Written properly, the estimator is a function of two fitted coefficient vectors, and its influence function has to carry their estimation error. Doing so gives a standard error of 0.1011 against an actual spread of 0.1034, and an interval covering 94.5%. And the shape that appears when it is written out is the same shape as the augmentation term. The augmentation is not an addition to the outcome regression; it is the part of it that a naive standard error was already leaving out. That reading is worth more than the theorem’s usual statement, because it says why the correction exists rather than what it guarantees.

The check a four-cell table has no closed form for

Everywhere else in this field a counted number has an integral beside it, computed on a grid with no sample in it, so that neither route can confirm itself.

A weight that balances, and one that unbalances. The standardised difference between the arms on each covariate, integrated over the population rather than counted in a sample. Unweighted, the arms differ by 0.8310 on the first covariate and 0.6015 on the second, which is what makes the raw difference of arm means 2.7102 against a true average effect of 1.0000. Weighting each unit by one over its own assignment probability removes both differences exactly — -2.78e-17 and -5.69e-19, which is machine precision and not a small number — because the weighted density of the treated arm is the population's own whatever the propensity is. Weighting by a score fitted without the second covariate balances the first to 0.0035 and pushes the second out to 0.7057, further apart than doing nothing.
Fig. 2 The one quantity in this field that does have a closed form: the standardised difference between the arms, 0.8310 and 0.6015 unweighted and machine zero under the true weights, integrated over the population. A four-cell bias table has no such route.

A bias table does not have one. There is no integral that returns “what a logistic fit omitting the second covariate, combined with a linear outcome fit omitting the same covariate, does to an augmented estimator at six hundred rows” — the quantity is defined by a fitting procedure at a finite sample and has to be counted.

So the second route here is a different one, and it is the reason the defect above was caught. Every estimator’s standard error is written out as an influence function and checked against the spread of the estimates themselves. Those are two independent statements: the first is a formula applied inside one sample, the second is the observed scatter of six hundred separate samples, and they share no arithmetic. Where both models are right the outcome regression reports 0.1011 against a spread of 0.1034, and the augmented estimator 0.1099 against 0.1165 — agreement to a few per cent, which is what a correct influence function looks like.

The weighting row is where the two routes disagree: 0.2674 reported against 0.1915 observed. That disagreement is not a bug, and the essay on estimated weights explains it; the point here is that the check reports it rather than absorbing it. A gate that only asked whether each interval covered near 95% would have passed the weighting row at 99.2% and called it success.

And it is the check that failed loudly on the broken standard error. A reported spread three times smaller than the observed one is not a subtle discrepancy — it is the kind of thing that shows up the moment two routes are required to agree, and the kind of thing that never shows up at all when a formula is trusted because it looks like a formula.

Which intervals cover, and which do not

Bias is not what a reader pays. An interval is, so every cell is also read as a coverage.

Which intervals cover, and which do not. How often a 95% interval contains the true average effect, over 600 samples of 600 units with the assignment rule at strength 1, for three estimators in each of the four cells made by getting each nuisance model right or wrong. The augmented estimator covers 93.7%, 93.3% and 99.2% in the three cells where a model is right and 0.2% in the one where neither is. The outcome model on its own covers 0.0% once it is wrong, and weighting on its own 1.5%. Weighting covers 99.2% when both are right, which is over-coverage rather than success: its interval is built from a formula that assumes the propensity was known, and it was estimated.
Fig. 3 How often a 95% interval contains the true average effect, for three estimators in each of the four cells, over six hundred samples. The augmented estimator covers 93.7%, 93.3% and 99.2% where a model is right and 0.2% where neither is.

Where a model is wrong, the intervals do not merely widen — they collapse. The outcome regression covers 0.0% of the time once it is the wrong model, weighting 1.5%, and the augmented estimator 0.2% in the cell where neither is right. Zero per cent of six hundred draws is not a rounding of a small number; it means the bias is several interval widths, every time, and a reader looking at any one of those intervals sees a perfectly ordinary result.

That is the real content of “not neither”. The failure in the fourth cell is not that the estimate is off by 0.81; it is that the estimate is off by 0.81 and the interval says it is off by about 0.15. Nothing in the output announces which cell the analysis is in, because the cell is decided by a covariate nobody has. The same silence is what the essay that ran one regression through three causal structures found: the arithmetic is identical in all three, and only the structure says which answer it is producing.

One cell covers too well rather than too badly. Weighting alone, with both models right, covers 99.2% — and that is not the interval succeeding. Its average standard error is 0.2674 against an actual spread of 0.1915, so the interval is about forty per cent too wide. The reason is that the interval is computed from the formula for known weights and the weights were estimated, and estimated weights have a smaller variance than true ones. That sentence sounds backwards and it is the subject of the essay on where an estimate’s weights should come from; it arrives here as an over-coverage, which is the same fact seen from the other side.

What the theorem does not carry

Double robustness is a statement about bias in the limit. It says nothing about the spread at a finite sample, and the spread is where the augmentation term’s cost lives: those two correction terms are divided by the propensity, so their variance grows as the assignment approaches decidable.

Move the assignment rule from strength 1 to 2.5 and the table’s bias column barely changes — the augmented estimator is still at −0.0164 with both models right, and 0.0334 with only the propensity model right. The error column changes a great deal. With both models right the augmented estimator’s root mean square error goes from 0.1167 to 0.3331, against an outcome regression sitting at 0.1184, and its coverage holds at 94.2%. With only the propensity model right it goes to 0.9431 and its coverage falls to 76.0%.

Either model is enough; neither is not. The bias of three estimators of an average effect of 1.0000, over 600 samples of 600 units with the assignment rule at strength 2.5, in each of the four cells made by getting each nuisance model right or wrong. The wrong model in both cases is one that omits the second covariate, which the outcome and the assignment both depend on. The outcome model alone is off by 1.4529 whenever it is the wrong one; weighting alone is off by 1.5962 whenever the propensity model is. The augmented estimator built from both is off by -0.0164, -0.0012 and 0.0334 in the three cells where at least one of them is right, and by 1.5422 in the fourth — which is between its two components rather than better than either.
Fig. 4 The same four cells with the assignment rule at strength 2.5. Every bias the theorem promises is still there; what has moved is the spread, which the theorem says nothing about.

So the estimator is doing what it promised at a setting where it has become expensive. A reader who chose it for its consistency has bought consistency and paid in a currency the choice was never priced in.

The comparison inside that cell is the awkward one. With both models right, the plain outcome regression sits at 0.1184 and the augmented estimator at 0.3331 — nearly three times the error, for a property neither of them needs, since the outcome model is right and its own bias is −0.0034. Double robustness is insurance, and the premium at strength 2.5 is about two hundred per cent of the loss it insures against. That is a defensible purchase when the buyer does not know which cell they are in, which is always; it is not a free one, and the usual account of the estimator does not mention a price at all.

The cell where having both is worse than either

Take one more step, to an assignment strength of 3.5, one past the end of the overlap sweep, and read the cell where the propensity model is right and the outcome model is wrong. This is the cell double robustness is for: the estimator’s whole justification is that a wrong outcome model does not sink it.

It does not sink it. The augmented estimator’s bias is 0.0857, against the outcome regression’s 1.6333 and weighting’s 0.5535. It is by a wide margin the least biased estimator on the table.

It is also the worst one on it. Its root mean square error is 1.9265, against the outcome regression’s 1.6402 and weighting’s 0.9383. A reader who took the nearly unbiased estimator over the badly biased one has doubled their expected squared error. Its interval covers 57.0%.

The only unbiased estimator on the table, and the worst. Root mean square error over 600 samples of 600 units with the assignment rule at strength 3.5, one step past the end of the overlap sweep, with the median absolute error beside each bar. In the cell where the propensity model is right and the outcome model is wrong, the augmented estimator carries a bias of 0.0857 against 1.6333 for the outcome model and 0.5535 for weighting — and a root mean square error of 1.9265 against their 1.6402 and 0.9383. It is the least biased estimator in that cell and the worst one in it, because double robustness is a statement about bias in the limit and the augmentation term carries a variance that grows with one over the propensity. Its interval covers 57.0%.
Fig. 5 Root mean square error at an assignment strength of 3.5, with the median absolute error beside each bar. In the cell where the propensity model is right, the augmented estimator carries a bias of 0.0857 against 1.6333 and 0.5535, and an error of 1.9265 against 1.6402 and 0.9383.

Being nearly unbiased and being any good are different properties, and this is the cleanest demonstration of it available on this population. Bias and spread are not comparable until they are squared, and once they are, an estimator that is right on average and enormous in a tenth of its draws loses to one that is wrong by half and steady.

The median absolute error is the number that says what “enormous in a tenth of its draws” means, and it turns the ordering over again. By median absolute error the augmented estimator is 0.6726, the weighted one 0.8143 and the outcome regression 1.6303 — so on the typical draw the augmented estimator is the best of the three, by a clear margin, and on the average of squares it is the worst. Both statements are true of the same six hundred draws. The distribution of its error is not one that a single summary describes: it is centred well and has a tail that owns the mean of its square.

There is a temptation to resolve the contradiction by picking the median and moving on, and it should be resisted. The root mean square is the right summary when the loss actually paid is squared error, which is what an interval width and a power calculation are both built on; the median is the right summary for the question “what will this estimate look like on the day”. They disagree here because the estimator has a tail, and reporting only the one that flatters it is how an estimator acquires a reputation it cannot support. The same discipline about which summary answers which question runs through the essay on what a residual is not, where a quantity that looks like an error and is measured like one turns out to be systematically the wrong size.

That is why the root mean square at this setting is not quoted alone. It moves between about 1.9 and 2.7 depending on which block of seeds is drawn, because it is an average of a heavy-tailed quantity; the median moves by about two per cent, and the cell where the augmented estimator is worse than both of its components fires in every block tried.

Where the cost comes from, and it is not the augmentation

The variance that ruins the augmented estimator at strength 3.5 is not a peculiarity of the correction term. It is the same quantity that ruins the plain weighted estimator, arriving through the same denominator.

The estimator has no upper bound on what it costs. What a thinning overlap does to a stabilised inverse-probability estimate of an average effect of 1.0000, over 600 samples of 600 at each of six settings. The spread rises from 0.1965 to 0.8122 and the root mean square error from 0.1964 to 0.9219, so at the thin end the error is very nearly the whole of the quantity being estimated. The lower line is the share of the arm's weighted total the single largest observation owns, averaged over the same draws: 0.69% to 9.78%, and in the worst single draw of the sweep 82.75%. Coverage of the 95% interval goes from 93.7% to 55.5%.
Fig. 6 What the same thinning overlap does to a plain weighted estimate: a spread rising from 0.1965 to 0.8122 and coverage falling from 93.7% to 55.5%. The augmentation term is divided by the same propensity.

The essay on the region with no comparison prices that failure directly, and its conclusion applies without modification here: past a certain thinness there is no estimator built on these weights that is not paying for a comparison the sample does not contain. Adding an outcome model does not conjure the missing units; it produces a correction term that is asked to reweight residuals in a region where the reweighting is a single observation carrying twenty-five million.

The reason the augmented estimator can be worse than the plain weighted one, rather than merely as bad, is that the weighted estimator here divides each arm’s total by its own summed weights and the augmentation term does not. The stabilisation that damps the weighted estimator’s tail is not available to the correction, so the correction’s tail is the one that survives into the root mean square. That is a fact about this particular pairing rather than about augmentation in general, and it is worth stating as such.

What “wrong” means here, and what it does not

Every wrong model in this table is wrong in exactly one way: it omits the second covariate, which both the assignment rule and the outcome depend on. That is a real misspecification and it is one kind.

A model with the right variables and the wrong functional form — a curved outcome fitted as a line, an assignment rule with an interaction fitted without one — biases the two components differently, and there is no reason to expect the four-cell table to keep its shape. In particular the cell where the augmented estimator is worse than both of its components was found by pushing one scalar until the augmentation’s spread exceeded the outcome regression’s bias; a different kind of wrongness moves both of those quantities and could move the crossing anywhere, including out of existence. Nothing here measures that, and it is the largest gap in this essay.

Where the difference between the two weightings goes. The variance of a stabilised inverse-probability estimate over 1200 samples of 600 units, weighted three ways. Weighting by the true propensity gives 0.072230. Regressing that estimator's error across draws on the draw's own imbalance — the sum over units of the treatment assigned less the probability it was assigned with, one component per covariate — explains 56.33% of it and leaves 0.031543. Weighting by a propensity fitted from the same sample gives 0.034446, which is 91.6% of the way to that floor and 47.7% of the true-weight variance. The same regression explains 0.05% of the fitted-weight estimator, because it has already used the imbalance up.
Fig. 7 Where the difference between weighting by a true propensity and weighting by a fitted one goes. The fitted weights leave 0.034446 against the true weights’ 0.072230, and a projection on the draw’s own imbalance accounts for the gap.

Two smaller things are worth naming. The intervals for the augmented estimator treat both nuisance fits as known, which is the usual reading and is conservative where the propensity model is right; in the cells where it is wrong the interval is not valid for anything and the coverage says so rather than the arithmetic. A 95% interval is a promise about a procedure repeated, and the essay on what that 95% refers to is the reason every cell here is counted over six hundred draws rather than argued about: a promise of that shape can only be checked by running the procedure. And the 99.2% over-coverage in the weighting row is a symptom of the same thing the decomposition above measures — that a fitted propensity is a different estimator from a known one, with a smaller variance, so a standard error derived for the known case is the wrong one. Nothing in this table adjusts for it.

Finally, none of the four cells is a cell anyone occupies knowingly. The whole table assumes unconfoundedness — that holding the recorded covariates fixed makes the assignment as good as random — and every cell in it is a statement about what happens when a recorded covariate is left out of a model. A covariate that was never recorded puts the analysis outside the table entirely, in a place where all three estimators agree with each other and none of them is right. The essay that priced adjusting for every covariate on hand and the essay on two worlds that produce identical data are where that boundary is measured; the balancing arithmetic this field opens with is exact and cannot see across it. Only assigning by a coin makes the question go away, and that is not available in the setting any of this is for.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Augmented estimatorAverage treatment effectConfidence intervalDoubly robustHeavy tailInfluence functionInverse-probability weightingModel misspecificationNuisance modelOutcome regressionPropensity scoreRoot mean square errorStabilised weightsStandard errorUnconfoundedness