Weighting one sample into another

The moments a balance is told

Weights fitted to balance the covariates' means were wrong by 0.6973 in the world where both the assignment and the outcome carry a square. Told the squares and the product as well, the same construction is off by −0.0045 there and its interval covers 91.0%. The failure moves up a moment rather than away: with a cube in both, the second-moment balance is off by 0.3099 and leaves the cube twice as far apart as no weighting. And where overlap is thin, 37.0% of samples have no such weights at all.

Worth reading first: A score that balances.

A weight fitted to balance fitted each arm’s weights so that its weighted covariate means equal the whole sample’s, exactly. The construction reached the semiparametric variance bound and was unbiased wherever either the assignment or the outcome was linear in the balanced covariates. In the world where both carried a square of the second covariate it was wrong by 0.6973, with an interval that covered 1.5% of the time, and its balance table read perfect, because the square was the one moment it had not been told about.

The obvious response is to tell it. Add the squares of both covariates and their product to the columns the weights must balance, and the fourth world’s failure should disappear. The obvious cost is that weights of this form have to satisfy six equations in each arm instead of three, and a sample whose treated units are sparse where the sample is dense may not be able to satisfy them. This essay measures both sides on the same draws — twelve hundred samples of six hundred units at the field’s standard overlap, six hundred in each world, five hundred at each strength of the assignment rule — and adds a fifth world to find where the failure goes next.

One weighting told the means and one told the second moments, in five worlds. The bias of the fit to balance over 600 samples of 600 units in each world, fitted to the covariates' means and fitted to their means, squares and product. Told the means it is off by -0.0004, -0.0020, 0.0103, 0.6973, 0.2698 in the worlds with no square, a square in the assignment, a square in the outcome, a square in both and a cube in both; told the second moments, by -0.0009, -0.0000, 0.0009, -0.0045, 0.3099. Its interval covers 94.0%, 94.7%, 95.0%, 1.5%, 51.0% and 93.7%, 89.8%, 94.0%, 91.0%, 48.3%.
Fig. 1 The bias of the fit to balance in five worlds — no square, a square in the assignment, a square in the outcome, a square in both, a cube in both — when the weights are told the covariates’ means, and when they are told the means, the squares and the product.

Six columns instead of three

The weights keep their form: one plus an exponential of a linear score, never below one, fitted by Newton’s method on a convex loss whose stationary point is exact balance. What changes is the list of columns. Beside the intercept and the two covariates the design now carries both centred squares and the product of the two, so each arm’s weighted means of all five must reproduce the sample’s. Because the intercept is a column the weights in each arm still sum to the sample size, and because every balanced moment is reproduced exactly the regression identity of the earlier essay still holds, now for a regression on all six columns.

That identity is the whole of the prediction. The estimate is the difference between two within-arm regressions, each evaluated at the sample’s means, and with squares and a product among the regressors the implied outcome model can represent an outcome that carries a square. So the protection that covered a linear outcome before should now cover a quadratic one, whatever the assignment rule does, and the protection that covered an assignment logit linear in the columns now covers one with a square in it.

The fourth world, repaired

Told only the means, the fit to balance was off by 0.6973 where both parts carry the square. Told the second moments, on the same six hundred samples, it is off by −0.0045, with a standard error of 0.0052. In the three worlds where the means were already enough, it stays unbiased: −0.0009 with no square anywhere, −0.0000 with the square in the assignment, and 0.0009 with the square in the outcome. Integrated over the population rather than counted, its limit in the fourth world is the true effect to the grid’s precision, where the means-only balance’s limit was off by 0.701.

The likelihood fit given the same six columns is off by 0.0940 in the fourth world, against 0.7903 when it was given three. The true weights are unbiased in every world, as they always were, with a standard deviation of 0.7008 in the fourth, which is why nobody uses them. The design, not the estimator, was what had failed.

An interval that stops covering where the bias does not

Weights told the squares, in the four worlds a square can be put into. The bias of three weighting estimators of an average effect of 1.0000, over 600 samples of 600 units in each of the four worlds, with the likelihood fit and the fit to balance both given the covariates, their squares and their product. The fit to balance is off by -0.0009, -0.0000, 0.0009, -0.0045 and its interval covers 93.7%, 89.8%, 94.0%, 91.0%. The likelihood fit on the same columns is off by 0.0452, 0.0468, 0.0504, 0.0940 and the true weights by 0.0184, 0.0405, 0.0237, 0.0485.
Fig. 2 The bias of the true weights, a likelihood fit and a fit to balance in the four square worlds, with both fits given the covariates, their squares and their product, and each estimator’s coverage beside it.

The repair is not free inside the worlds it repairs. The fit to balance’s interval, built from the same influence function as before, covers 93.7% with no square anywhere, 89.8% with the square in the assignment, 94.0% with it in the outcome and 91.0% with it in both. With the means alone it had covered 94.7% in the assignment world. The bias has not come back; the spread has grown and the standard error has not grown with it. With the square in the assignment the estimate’s standard deviation over the six hundred samples is 0.1265 told the second moments, against 0.1052 told the means.

The reason is the extra columns. A weight that must reproduce the square’s mean in the treated arm leans harder on the treated units with extreme values of the second covariate, and in the world whose assignment already over-selects those units the leaning is not needed to remove any bias — the means were enough there — but it is paid for in the variance. The influence-function standard error describes a regression with six columns fitted to six hundred units and understates what the reweighting adds. A balance extended to moments the data did not need is honest about the estimate and a little optimistic about its interval.

The failure moves one moment up

The moment each balance was not told about. Standardised differences between the arms, integrated over the population. In the world whose assignment carries a square of the second covariate, the square differs by 0.3766 unweighted, 0.4908 after balancing the means and 0.0e+0 after balancing the second moments. In the world whose assignment and outcome carry a cube, the cube differs by 0.1115 unweighted, 0.2402 after balancing the means and 0.2318 after balancing the second moments.
Fig. 3 Standardised differences between the arms, integrated over the population: on the square, in the world whose assignment carries a square, and on the cube, in the world whose assignment and outcome carry a cube — unweighted, after balancing the means, and after balancing the second moments.

The square was the moment the first design was not told about, and balancing its mean removed its imbalance exactly: in the world whose assignment carries a square, the arms differ on the square by 0.3766 standard deviations unweighted, by 0.4908 after balancing the means — further apart than no weighting, the earlier essay’s finding — and by exactly zero after balancing the second moments.

A second-moment balance has its own moment left out, and the fifth world puts the confounding there: a centred cube of the second covariate in both the assignment and the outcome, uncorrelated with the covariate and with its square. There the arms differ on the cube by 0.1115 unweighted. Balancing the means moves that to 0.2402, and balancing the second moments moves it to 0.2318 — more than twice as far apart as no weighting, the square’s pattern repeated exactly one moment up. The weights reproduce every moment they were given by leaning on the tails of the second covariate, and the tails are where a cube lives.

So the estimate is wrong there, told either design. With the means balanced it is off by 0.2698 and covers 51.0%; with the second moments balanced it is off by 0.3099 and covers 48.3%. Integrated over the population the two limits are 0.256 and 0.261. Extending the balance did not reach the cube, and it made the estimate a little worse there, because the extra columns pushed more weight into the tails that carry the moment nobody named.

That is the precise sense in which balancing more moments buys protection rather than robustness. Each column added protects against confounding that lives in that column and in nothing else. It does not approach protection against confounding in general, because there is always a next moment, and exact balance on the ones named is silent about it in the same way it was silent about the square — every balance diagnostic built from the design reads perfect on the samples where the estimate is furthest from the truth.

Why balancing more moves the cube further

The movement is the one a score that balances found for a covariate a likelihood fit was never given, and it shows in the counted samples as well as in the integral. Counted over the six hundred samples of the cube world, the arms differ on the cube by 0.1113 unweighted, by 0.2562 after balancing the means and by 0.2741 after balancing the second moments.

An assignment that rises with the cube over-represents the far positive tail of the second covariate among the treated units and the far negative tail among the controls. The weights are fitted to reproduce moments in which those two tails nearly cancel — a mean, a square, a product — and there are many sets of weights that reproduce such moments while leaving the tails’ own difference untouched or larger. The exponential form picks the set that leans on each arm’s extreme units, which is exactly where a cube’s difference lives, and every extra moment it is asked to reproduce gives it one more reason to lean there.

None of that shows in the weights’ spread. In the cube world the largest single weight in the median sample holds 1.6% of the treated arm told the means and 1.9% told the second moments, which the effective size a set of weights leaves would read as harmless. The imbalance is in where the mass sits rather than in how unevenly it is spread, so a diagnostic of weight spread cannot catch it and a diagnostic of the next moment can.

The variance of a longer design

Three weights and the floor none of them can go under. Told the squares and product as well, the likelihood fit's variance is 0.044425 and the fit to balance's 0.012993, 1.0304 of the bound, covering 93.6%. The variance of the estimated average effect over 1200 paired samples of 600 units, weighted three ways, beside the semiparametric bound — E[1/e + 1/(1 − e)] plus the variance of the effect over the population, divided by the sample size, which no estimator using these data can beat. Weighting by the true propensity gives 0.072230, by a likelihood fit 0.034446, and by a fit to balance 0.011883: 34.5% of the likelihood fit's and 16.5% of the true weights'. The bound is 0.012610; the counted variance is 0.9423 of it, inside the 4.1% relative error a variance over 1200 draws carries. Its bias is -0.0012 and its interval covers 94.0%.
Fig. 4 The variance of the estimated average effect over twelve hundred paired samples at the standard overlap, weighted by the true propensity, by likelihood fits and by fits to balance on the means and on the second moments, beside the semiparametric bound.

In the field’s own world, where the outcome and the assignment are both linear and none of the extra columns is needed, the price of carrying them can be read directly. Told the means the fit to balance had a variance of 0.011883, 0.9423 of the semiparametric bound. Told the second moments it has 0.012993, 1.0304 of the same bound of 0.012610. A variance counted over twelve hundred draws carries a relative standard error of about 4.1%, so both sit on the bound within their measurement, and the longer design costs about 9.3% in variance where it buys nothing.

The likelihood fit pays more for the same columns: its variance is 0.044425 told the second moments, against 0.034446 told the means, so the fit to balance’s advantage over it grows, to 0.2925 of its variance. In the median sample the largest single weight holds 1.9% of the treated arm’s total, against 1.7% told the means. At this overlap the extra columns are cheap, and a report that balanced them in a world that turned out not to need them has lost very little.

Where the overlap cannot support six moments

Where the overlap cannot support the moments a balance is asked for. Over 500 samples of 600 units at each setting of the assignment rule, the share of draws on which no weights of the form 1 + e^(±x'γ) reproduce the sample's moments: 0.0%, 0.0%, 0.0%, 0.4%, 3.8% when the means are balanced and 0.0%, 0.0%, 0.0%, 5.0%, 37.0% when the squares and product are balanced too. On the draws that have weights, the interval covers 96.4%, 93.6%, 93.0%, 89.0%, 82.1% and 96.4%, 93.6%, 88.4%, 78.3%, 70.8%.
Fig. 5 Over five hundred samples at each strength of the assignment rule, the share on which no weights of this form balance the sample, and the coverage of the interval on the samples that do have weights, for the means and for the second moments.

Thin overlap is where the second side of the trade arrives. A weight of the form one plus an exponential is never below one, so an arm that is sparse where the sample is dense cannot always reproduce the sample’s moments, and six moments are harder to reproduce than three. As the assignment rule strengthens from 0.5 to 2.5, the share of samples with no such weights is 0%, 0%, 0%, 0.4% and 3.8% when the means are balanced, and 0%, 0%, 0%, 5.0% and 37.0% when the second moments are. At the thinnest overlap more than a third of samples of six hundred units cannot be balanced on their squares at all, and the estimator is silent on them.

On the samples that do have weights the interval degrades faster too. Told the second moments it covers 96.4%, 93.6%, 88.4%, 78.3% and 70.8% across the sweep, against 96.4%, 93.6%, 93.0%, 89.0% and 82.1% told the means. In the median sample at the thinnest overlap the largest single weight holds 12.8% of the treated arm, against 10.4%. The reason is the one the region with no comparison measured for every weighting estimator, now sharpened: a balance that must be exact on six moments has to find its way into the region where the arms barely overlap, and when it cannot it fails outright rather than degrading.

The bound, read again at thin overlap

The bound, and how far from it thin overlap pushes, with the second moments balanced. The variance of the estimate weighted by a fit to balance, over 500 samples of 600 units at each of five settings of the assignment rule, divided two ways: by the variance of the estimate weighted by a likelihood fit on the same draws, and by the semiparametric bound for that setting, an integral with no sample in it. Against the likelihood fit it reads 0.7784, 0.2772, 0.2617, 0.2404, 0.3032. Against the bound it reads 0.8888, 0.9931, 0.9896, 0.5649, 0.1633. The coverage of its interval is 96.4%, 93.6%, 88.4%, 78.3%, 70.8%, and in the median draw the largest single weight holds 0.7%, 2.0%, 4.4%, 7.9%, 12.8% of the treated arm. On 0.0%, 0.0%, 0.0%, 5.0%, 37.0% of draws no weights of this form balance the sample at all, and those draws are left out of every reading on this figure.
Fig. 6 The variance of the second-moment balance across the overlap sweep, divided by the likelihood fit’s variance on the same draws and by the semiparametric bound for each strength.

Against the likelihood fit the second-moment balance’s variance reads 0.7784, 0.2772, 0.2617, 0.2404 and 0.3032 across the sweep: better everywhere, as the means-only balance was. Against the bound it reads 0.8888, 0.9931, 0.9896, 0.5649 and 0.1633. The means-only balance had read 0.7544 at a strength of 1.5; told the second moments the estimator sits on the bound there too, where the shorter design was already dipping below it.

The readings below the bound at the two thinnest settings are the earlier essay’s warning in a new form. The bound is an integral over the whole population, most of it at thin overlap in units a sample of six hundred almost never contains, and a sample that has not reached where an integral is computed reports a variance below it. Told the second moments the estimator is measured on even fewer samples there — 315 of five hundred at the thinnest setting — so the counted variance describes a population of samples that could be balanced, which is a narrower and better-behaved population than the one the bound describes.

How many moments, then

The measurements lay the choice out as a trade with three terms. Each moment added protects exactly the worlds whose confounding lives in that moment: the square was repaired completely, and the cube not at all. Each moment added costs a little variance and a little coverage where it was not needed — 9.3% of variance and five points of coverage in the square-in-assignment world at standard overlap. And each moment added removes samples at thin overlap, by an amount that grows much faster than the number of columns: from 3.8% to 37.0% of samples at the thinnest setting for two squares and a product.

So the design is a claim about which moments the confounding could live in, and it is a claim that has to be made before the outcomes are seen, the way arranging units before any outcome exists is. A study with good overlap can afford to balance second moments it may not need. A study with thin overlap cannot, and has to choose between a design that is silent on the square and a design that is silent on a third of its samples. What no design can do is what the mechanism the data cannot see already established: balance on what was measured says nothing about what was not.

What a balancing analysis should report

The columns the weights were fitted to balance. A balance table that lists them reads perfect by construction; the columns themselves are the claim.

The balance on the next moment up. For a means-only balance, the squares; for a second-moment balance, the cubes and the interactions of the squares. That is the one diagnostic a balance table can print that is not zero by construction, and in the worlds measured here it is the one that moved the wrong way on the samples where the estimate failed.

How many samples, resamples or subgroups had no feasible weights. An estimator that is exact when it exists and silent otherwise has to report its silence, and at thin overlap a second-moment balance is silent more than a third of the time.

The broader shape is the double robustness of an augmented estimator, made visible as a list: every column added to the balance is a term added to both the implied propensity model and the implied outcome model at once, and the estimate is right whenever either of those models is.

Proved, computed and counted

Proved. Exact balance on every column in the design, and the identity that makes the estimate a within-arm regression on those columns evaluated at the sample means; the protection in a world whose assignment or outcome is linear in the columns follows from it.

Computed without a sample. The population limits and the integrated standardised differences on the square and the cube in each world, and the semiparametric bound at each overlap.

Counted. Every bias, coverage, variance, weight share and infeasible share, over twelve hundred samples at the standard overlap, six hundred in each world and five hundred at each strength, on the seeds the means-only balance used, so the two designs differ on a sample only in the columns they were told.

Particular to this population. Two normal covariates, one square and one cube as the unnamed moments, and weights of one exponential form. Other forms of weight — ones allowed below one, or fitted to approximate rather than exact balance — trade the infeasible samples differently.

Still open: balance up to a tolerance

The infeasible samples are the cost of asking for exact equality, and exact equality is a choice. Weights fitted to make each moment’s imbalance no larger than a stated tolerance, with the smallest spread of weights that achieves it, exist on every sample and approach the exact balance as the tolerance shrinks. That turns the design question into a dial: how many moments to balance, and how tightly each. Whether a tolerance of a tenth of a standard deviation on the second moments keeps the fourth world’s repair, recovers the 37.0% of samples the exact balance lost, and at what bias from the moments it no longer balances exactly, is a measurement the same four worlds and the same sweep can make.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formCovariate balanceDoubly robustEfficiency boundEstimandInverse-probability weightingModel misspecificationOverlapPositivityPropensity scoreStandardised difference