A weight fitted to balance
Worth reading first: A score that balances.
The estimated weight is the better one found that weighting by a propensity fitted by maximum likelihood halves the variance of the estimate compared with weighting by the true propensity, and traced the gain to one line: the likelihood’s score equation sets the draw’s own covariate imbalance to zero as a side effect. It ended with the estimator that follows from reading that line the other way round — a propensity fitted to balance, whose defining equation is the balance itself rather than the likelihood — and did not fit one.
This essay fits it, on the same twelve hundred draws of six hundred units, beside the true weights and the likelihood weights, so every comparison is paired. The result is better than the likelihood fit by more than the likelihood fit was better than the truth, and it comes with a precise description of what it cannot see.
Weights defined by the balance they produce
A propensity weight for a treated unit is and for a control unit . With a logistic model for e, a treated unit’s weight is and a control unit’s . Maximum likelihood chooses γ to predict who was treated. The alternative chooses it so that the treated units, weighted, have exactly the sample’s covariate means — one equation per column of the design, solved by Newton’s method from the likelihood coefficients — and separately so that the control units, weighted, do too.
Because the intercept is one of the columns, each arm’s weights sum to the sample size exactly, so the two usual forms of the weighted estimator coincide and there is one estimate rather than two. And because the balance is imposed rather than hoped for, it can be read off every draw.
The true propensity leaves a root-mean-square standardised difference of 0.1317 on the first covariate and 0.1186 on the second — a sampling error, since the true propensity is right about the population and knows nothing about this draw. The likelihood fit leaves 0.0770 and 0.0657, having absorbed part of the draw’s imbalance as the earlier essay showed. The fit to balance leaves 1.4×10⁻¹⁴ and 1.2×10⁻¹⁴, and the largest gap between a weighted arm mean and the sample mean in any of the twelve hundred draws is 9.8×10⁻¹⁴. That is not a small imbalance. It is the arithmetic’s floor, and a balance table built from these weights reads zeros.
Exact balance makes the estimate a regression
Exact balance has a consequence that is easy to state and that decides everything below. Fit a least-squares line of the outcome on the balanced covariates within one arm, using the balance weights. Its prediction at the sample’s covariate means is, by algebra, the arm’s weighted mean outcome — because the weighted covariate means are the sample’s means, and a weighted least-squares line passes through the weighted means. Across the twelve hundred draws the largest gap between those two quantities is 2.0×10⁻¹³.
So the estimate from weights fitted to balance is simultaneously a weighting estimator and a regression estimator: the difference between two arms’ regression lines, each evaluated at the whole sample’s covariate means. That is the structure the augmented estimator had to build by adding an outcome model to a weighting model. Here it arrives inside a single set of weights, and it means the estimate inherits a regression’s protection whenever the outcome really is linear in the covariates that were balanced, whatever the weights got wrong about the assignment.
A variance at the bound
Weighted by the true propensity the estimate has a variance of 0.072230; by the likelihood fit, 0.034446; by the fit to balance, 0.011883. That is 34.5% of the likelihood fit’s variance and 16.5% of the true weights’. The likelihood fit took half the variance out of the truth; the balance fit takes two thirds out of the likelihood fit.
There is a floor under that sequence, and it can be computed without drawing anything. For any estimator of an average effect that uses only the covariates, the assignment and the outcome, the asymptotic variance cannot be smaller than the semiparametric efficiency bound, which here is plus the variance of the effect across the population, divided by the sample size — 0.012610 in this world. The counted variance of the balance fit is 0.9423 of that bound, and a variance counted over twelve hundred draws carries a relative standard error of about 4.1%, so the estimator sits on the floor to within its own measurement. Two routes that share nothing — a grid integral over the population and a count over samples — agree that this estimator has taken all of the precision there is to take.
It pays nothing for it. Its bias is −0.0012 against an effect of one, and the interval built from its influence function covers 94.0%. The likelihood fit, whose standard error is derived as though its weights were known, covers 99.0% with a variance three times larger, the over-coverage the earlier essay explained.
Why this fit reaches the floor and the likelihood fit does not
The sequence 0.072230, 0.034446, 0.011883 is not three points on one curve, and the regression identity says why the last step is different in kind from the first.
The likelihood fit’s gain over the truth was, as the earlier essay measured, the draw’s chance imbalance absorbed as a location: the score equation sets one weighted combination of the covariate differences to zero, through a nonlinear likelihood, and the estimator’s variance falls by the share of its error that combination could predict. The balance fit sets every covariate difference to zero exactly, which removes that same share — and then, through the identity, it also does what an outcome regression does. The estimate is the difference of two within-arm regression lines evaluated at the sample’s means, so the part of each unit’s outcome that is predictable from its covariates no longer contributes noise to the comparison. Only the residual around those lines does.
The efficiency bound is the variance of exactly that kind of estimator: one that uses the propensity to correct the assignment and a regression to strip the predictable part of the outcome. In this world the outcome really is linear in the two covariates, so the regression the balance implies is the right one, and the estimator lands on the bound. That is also the limit of the result. Where the outcome depends on something the balanced columns do not span, the implied regression is the wrong one, the variance stays above the bound, and — as the next section shows — the bias can return with it.
The weights themselves are unremarkable at this overlap. In the median draw the largest single balance weight carries 1.73% of the treated arm’s total, so no unit dominates the estimate and the effective size a set of weights leaves is not what bought the precision. The gain came from where the weights put their mass, not from how evenly they spread it.
Four worlds a square can be put into
Every number so far is from a world where the assignment and the outcome both depend linearly on the two covariates, which is the world the balance was told about. The protection it claims has to be measured where that is false, and the cleanest falsehood is a square. Put a centred square of the second covariate into the assignment rule, into the outcome, into both, or into neither; keep the balance fitted to the two covariates’ means in every case; and count six hundred samples in each world.
With no square anywhere the balance fit’s bias is −0.0004, with a standard error of 0.0045. With the square only in the assignment it is −0.0020 — the propensity model is now wrong, and the estimate does not notice, because the outcome is still linear in what was balanced and the regression identity above leaves no room for a bias. With the square only in the outcome it is 0.0103, inside its standard error of 0.0075 — the outcome model the identity implies is now wrong, and the estimate does not notice that either, because the propensity model is right and weights fitted to balance are consistent for the true propensity when the model contains it. Integrated over the population rather than counted, the balance fit’s limit is the true effect to 4.1×10⁻¹² in all three worlds, where the likelihood fit’s limit with the square in the assignment is off by 0.0075.
With the square in both, neither protection is left, and the same exactly balanced weights are wrong by 0.6973, with an interval that covers 1.5% of the time. The likelihood fit is wrong by 0.7903 and covers 8.5%. The true weights are unbiased in every world — 0.0485 in the last, covering 83.0% — and are by far the noisiest estimator on the table in every one.
That is the double robustness either model but not neither measured, arriving in one set of weights: right if the assignment is linear in what was balanced, right if the outcome is, and wrong by most of the effect if both are not. And in the fourth cell the balance is still exact. Every balance diagnostic the analysis could print reads perfect on the draws where the estimate is furthest from the truth.
Exact on the covariate, worse on its square
Why the fourth cell fails can be read directly, because a balance table can be extended to the moment the weights were not given.
In the world whose assignment carries the square, the arms differ unweighted by 0.4751 standard deviations on the second covariate and 0.3766 on its square. Weights fitted to balance remove the first to 4.3×10⁻¹⁷. They move the second to 0.4908 — further apart than no weighting at all. The likelihood fit with the same two covariates does the same, to 0.5281. Counted over six hundred samples rather than integrated, the square’s difference goes from 0.3731 unweighted to 0.4957 after balancing.
The weights are not merely failing to fix the square; they are making it worse, and the reason is the construction. Assignment depends on the square, so units with large squares — both tails of the second covariate — are over-represented among the treated. Balancing the mean of the second covariate reweights those treated units so the tails cancel in the mean, and a reweighting that makes a symmetric excess in both tails cancel in the mean does nothing to cancel it in the square, and here pushes more weight into it. Exact balance on the means is not balance on the distribution, and the moment it was not told about is the moment a balancing estimator can move in the wrong direction while reporting nothing.
A covariate never named
The last world is the first field’s. A score that balances showed that a likelihood fit given only the first covariate leaves the second, which both the assignment and the outcome depend on, further apart than no weighting — 0.7057 against 0.6015. The question is whether fitting to balance, with its exact balance and its bound-level variance, rescues that.
It does not, and it could not. The fit to balance balances the first covariate to 1.3×10⁻¹³ on the counted draws and moves the second from 0.6015 unweighted to 0.7062 — to within 0.0005 of where the likelihood fit left it. The two estimators’ limits are off the effect by 0.8158 and 0.8130, and counted over the twelve hundred draws the balance fit is off by 0.8091. Balancing is a statement about the columns in the design. A column that is not in the design is not balanced by any method that works only on the design, and its imbalance is exactly the unmeasured confounding a covariate the data cannot see turns into a bias no estimator on these data removes.
What thin overlap does, and a variance below its own bound
The variance result was at the population’s standard overlap. Where overlap thins, every weighting estimator gets worse, and the balance fit has to be read there too.
Against the likelihood fit the balance fit’s variance reads 0.8032, 0.3475, 0.1433, 0.1169 and 0.1545 as the assignment rule strengthens from 0.5 to 2.5 — better everywhere, and by a factor of six to eight once the overlap is poor, where the likelihood fit’s estimate is also biased by 0.1433 and 0.2497. Against the bound it reads 0.8847, 0.9344, 0.7544, 0.3687 and 0.1307.
A variance a seventh of an efficiency bound needs explaining rather than celebrating, because an estimator cannot beat that bound in the limit it is stated for. Three things are happening, and the last is the one that matters.
Some draws have no answer. A weight of the form is never below one, so an arm that is sparse where the sample is dense cannot always reproduce the sample’s means. On 0.4% of draws at strength 2 and 3.8% at 2.5, no such weights exist, and those draws are left out of every reading. An estimator that is exact when it exists and silent otherwise has a variance measured on the draws where it exists.
Its interval knows something is wrong. Coverage falls from 96.4% at the widest overlap to 89.0% and 82.1% at the two thinnest, with the bias still near zero, and in the median draw at 2.5 the largest single weight carries 10.4% of the treated arm. The influence-function standard error is describing a narrower spread than the estimator has.
And the bound is computed where no sample goes. It is an integral of over the population, and at thin overlap that integral lives in the tails of the propensity. At strength 1 the units with a propensity within 0.01 of zero or one contribute 1.2% of it. At strength 2.5, units within contribute 59.5% of it and units within contribute 31.3% — and those last are 0.027% of the population, so a sample of six hundred holds even one of them on only 15.1% of draws. The bound is the variance of an estimator that has seen the whole population, extreme units included; a sample of six hundred has not seen most of what the bound is made of. It is the situation the draws aimed at the tail measured for an importance sampler whose variance was infinite and invisible: what a finite sample’s variance estimates is the part of the integral its draws reach, and a figure below the bound is a sample that has not yet reached where the bound is computed.
So the honest reading of the overlap sweep is two-sided. The balance fit is much better than the likelihood fit at every overlap, and where overlap is standard it is at the floor. Where overlap is thin, its variance is not evidence of super-efficiency but a measurement taken inside the region a sample of this size can see, with a failure rate, an undercovering interval and a largest weight each saying so.
What is proved here and what is counted
Proved. Exact balance on the means is imposed by construction and read to machine precision. The regression identity — weighted arm mean equals the weighted least-squares line at the sample means — is algebra once the balance is exact, and the double robustness in three of the four worlds follows from it: consistency when the outcome is linear in the balanced columns, and when the assignment’s logit is. The semiparametric bound is a theorem, and here it is an integral computed on a grid.
Computed without a sample. The population limits in the four worlds, the integrated balance on the square, the first field’s short design, the bound at every overlap, and the share of the bound’s integral in the propensity tails.
Counted. Every variance, bias, coverage, balance and weight share on draws from stated seeds shared by all three estimators: twelve hundred at the standard overlap, six hundred in each of the four worlds, five hundred at each strength of the sweep.
Not examined. The balance here is on means only. Balancing further moments — the squares, the interactions — is the obvious extension, and it moves the failure of the fourth world to whatever moment is still left out, at a cost in the weights’ spread that grows with every moment added.
Where this goes next: how many moments a balance should be told
The square was not in the design, and the design decided what the weights could protect. The next question is the design itself. Add the squares and the product of the two covariates to the columns the weights must balance, and the fourth world’s failure should disappear; add enough columns and the weights must satisfy more constraints than the overlap can support, and the draws with no feasible weights — already 3.8% at the thinnest overlap with two columns — should multiply.
The measurement that would settle how far to go is the same four-world table and the same overlap sweep with the balance extended to second moments: the bias in the fourth world, the variance against the bound, and the share of draws on which no weights exist, at each overlap. It is distinct from this essay because it trades one kind of silence for another — a moment left out against a sample that cannot satisfy every moment put in — and because the answer is the size of a design, which arranging the units before any outcome exists chooses in advance and a weighting estimator has to choose afterwards.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A dictionary that is neither — both name closed form, covariate balance, model misspecification
- A threshold in the tail — both name closed form, covariate balance, model misspecification
- An imputation model the analysis does not contain — both name closed form, estimand, model misspecification
- Dropping the incomplete rows — both name closed form, effective sample size, estimand
- What a wrong model estimates — both name closed form, estimand, model misspecification
- What the rule blocks — both name closed form, covariate balance, standardised difference
Named objects
A flat tag is an object no other essay names yet.
Closed formCovariate balanceDoubly robustEffective sample sizeEfficiency boundEstimandInverse-probability weightingModel misspecificationOverlapPositivityPropensity scoreStandardised differenceVariance ratio