Weighting one sample into another

A balance allowed a tolerance

Weights that balance the means exactly and the second moments only to within a tenth of a standard deviation were supposed to keep the second-moment repair and recover the samples exact balance loses. Where a square drives both the assignment and the outcome they keep neither half of the repair: the bias is 0.2782 against −0.0074 with exact balance, and the interval covers 26.5%, because the weights spend the whole tolerance — every second-moment gap sits at the edge of its band — and the outcome converts a gap of a tenth into bias at its own slope, 2·δ·√2 = 0.283. At thin overlap the same tolerance does recover samples: 89.0% have weights at all, against 62.0% with exact balance.

Worth reading first: A score that balances.

The moments a balance is told gave weights fitted to balance the treatment arms the squares and the product of two covariates as well as their means, and the repair held: in the world where a square drives both who is treated and what happens to them, weights that balanced only the means were wrong by 0.6973, and weights that balanced the second moments too were off by −0.0045 with an interval covering 91.0%. The repair had a price where overlap was thin. Six columns balanced exactly is a demand some samples cannot meet with positive weights, and at the thinnest overlap more than a third of samples had no such weights at all.

That essay named the obvious relaxation. Exact equality is a choice; weights could instead be fitted so that each second moment’s imbalance is no larger than a stated tolerance — a tenth of a standard deviation, say — with the means still balanced exactly. That turns the design into a dial, and the dial was expected to keep the repair, recover the lost samples, and give up a little bias on the moments no longer balanced exactly. How much bias is the part that turns out to be the whole story.

Exact on the means, within a band on the squares

The population is the earlier essays’. Two standard normal covariates, an assignment whose log-odds are linear in them with a centred square of the second added with weight 0.5, an outcome linear in them with the same square added with weight 1, and a treatment effect averaging one. The weights are the balancing tilt the earlier essays used — each arm’s weights the logistic form 1+e∓x′γ1 + e^{\mp x'\gamma} — fitted so that each arm’s weighted means of the intercept and both covariates equal the whole sample’s exactly, and its weighted means of both squares and the product lie within δ\delta of the sample’s, in units of each column’s own standard deviation.

The fit is the exact tilt with a penalty: the dual of the problem is the same convex loss with δ\delta times each tolerated column’s standard deviation charged on the absolute value of its coefficient. Where the penalty does not bind, the column’s coefficient is zero and its gap lies somewhere inside the band; where it binds, the gap sits exactly at the band’s edge. A tolerance of zero is the exact second-moment balance; a tolerance large enough that nothing binds is the balance on the means alone.

The tolerance is spent

The bias a tolerance on the second moments leaves where a square drives both the assignment and the outcome. 600 samples a point. Bias at δ = 0, 0.02, 0.05, 0.1, 0.2, 0.5: −0.0074, 0.0500, 0.1353, 0.2782, 0.5473, 0.6983; coverage 89.7%, 89.5%, 72.4%, 26.5%, 1.0%, 1.0%. The tolerance spent in full at the outcome's slope on the square would give 2·δ·√2: 0.0566, 0.1414, 0.2828, 0.5657 at the four smaller values; balancing the means alone gives 0.6973.
Fig. 1 The bias of the estimate against the tolerance δ on each second moment, in the world with a square in both the assignment and the outcome, beside the bias a second-moment gap at the edge of its band produces at the outcome’s slope on the square (dashed), and the bias of balancing the means alone (red).

In the world with a square in both, exact balance on the second moments leaves a bias of −0.0074 over six hundred samples. A tolerance of two hundredths of a standard deviation leaves 0.0500; five hundredths, 0.1353; a tenth, 0.2782; two tenths, 0.5473; half a standard deviation, 0.6983, which is the bias of balancing the means alone. The interval covers 89.7% with exact balance and 26.5% at a tenth.

The bias is not a gradual leak. It is arithmetic. The assignment depends on the square, so the treated arm, unweighted, has more units with large squares than the sample does, and the weights must pull its weighted mean of the square towards the sample’s. Given a band, they pull exactly to the edge of it and no further: every unit pulled further would cost weight variance the penalty does not repay, and at every tolerance up to two tenths the median sample’s largest second-moment gap is the tolerance itself — 0.020, 0.050, 0.100, 0.200. The control arm is pulled to the opposite edge. The two arms therefore differ in their weighted mean of the square by twice the tolerance, in the square’s units.

And the outcome converts that difference into bias at its own slope. The centred square has a standard deviation of 2\sqrt 2, so a gap of δ\delta standard deviations is δ2\delta\sqrt2 in its own units, twice that between the arms, and with the outcome rising one unit per unit of the square the estimate is off by 2δ22\delta\sqrt2: 0.057, 0.141, 0.283 and 0.566 at two hundredths, five hundredths, a tenth and two tenths, against counted biases of 0.050, 0.135, 0.278 and 0.547. The small shortfall is the samples in which the band does not quite bind.

So a tolerance on a moment that drives the assignment is not “approximately balanced”. It is unbalanced by exactly the tolerance, in the direction the assignment pushes, on every sample, and the bias is the tolerance times whatever the outcome does with that moment.

The conventional tenth

A tenth of a standard deviation is not an arbitrary setting. It is the threshold the balance diagnostics most studies report are read against: a standardised difference below 0.1 is conventionally called balanced, and a table of covariates all under it is the usual evidence that weighting has done its job. A score that balances began from that diagnostic. Weights fitted with a tolerance of exactly 0.1 pass it by construction, on every column, on every sample.

They also put every binding gap at 0.1 exactly, in the direction the assignment pushes, which is the least favourable place inside the band for the estimate. The diagnostic table looks the same as a table of gaps scattered around zero with none above a tenth; the estimate is different by 0.28 in this world. A threshold used to read imbalance after the fact was turned into a target, and a target that the fit is allowed to reach is one it will reach.

That is the general hazard of fitting to a diagnostic. A standardised difference is informative about a weighting that was not chosen to make it small: a gap of 0.08 left by weights fitted to the treatment indicator is a sampling fluctuation, as likely to point one way as the other. The same 0.08 left by weights fitted to be within 0.1 is a decision, made in the direction that costs least in weight variance, which is the direction the assignment was pushing — and that direction is exactly the one the outcome’s dependence on the moment turns into bias.

Why “balanced to within sampling error” does not save it

A tolerance is sometimes justified as matching the noise: in a sample of six hundred, a weighted mean is uncertain by roughly a tenth of a standard deviation anyway, so balancing more tightly than that is chasing noise. The counts say otherwise. At five hundredths — half the noise — the bias in the fourth world is already 0.1353 and the interval covers 72.4%; at two hundredths, 0.0500 and 89.5%. The bias scales with the tolerance at the outcome’s slope whatever the noise is, because it is not noise: it is the same gap, in the same direction, in every sample.

The distinction is the one the estimated weight is the better one drew between a sample’s own chance imbalance and the population’s. Exact balance removes the sample’s chance imbalance along with the population’s; a tolerance leaves the population’s imbalance in place up to the band, and a population imbalance is a bias, not an error that averages out.

Where it costs, and where it does not

Bias and coverage with the second moments balanced exactly and within a tenth of a standard deviation, in four worlds. 600 samples a world. no square anywhere: exact, bias −0.0071 and coverage 93.3%; within a tenth, −0.0060 and 93.7%; a square in the assignment: exact, bias −0.0072 and coverage 90.2%; within a tenth, −0.0049 and 93.3%; a square in the outcome: exact, bias −0.0072 and coverage 93.3%; within a tenth, −0.0006 and 80.1%; a square in both: exact, bias −0.0074 and coverage 89.7%; within a tenth, 0.2782 and 26.5%.
Fig. 2 Bias and coverage with the second moments balanced exactly and within a tenth of a standard deviation, in the four worlds: a square in neither the assignment nor the outcome, in the assignment only, in the outcome only, and in both.

The arithmetic says where the tolerance is harmless, and the four worlds confirm it. With no square anywhere the bias is −0.0071 exact and −0.0060 with a tenth’s tolerance, coverage 93.3% and 93.7%. With the square only in the assignment the band binds just as it did — the assignment still pushes the square apart — but the outcome does nothing with the square, so the gap costs nothing: −0.0072 and −0.0049, coverage 90.2% and 93.3%. With the square only in the outcome the assignment does not push the square apart, so the band does not bind on average and the bias stays at −0.0006.

Only in the fourth world, where the assignment pushes the square apart and the outcome responds to it, does the tolerance become bias. That is the world the second-moment balance was introduced for, so the tolerance gives up the repair exactly where the repair was needed. In the three worlds where exact second-moment balance was never necessary, the tolerance is harmless or better; in the one where it was, it is close to useless.

The same reasoning holds at any tolerance. A weight fitted to balance found the balancing estimator doubly robust: right if either the assignment or the outcome is linear in the balanced columns. A tolerance moves a column partly out of the balanced set, and the double robustness goes with it in proportion — the estimator is right if the outcome is linear in the column or the assignment does not push it apart, and wrong by the tolerance times the outcome’s slope if both fail.

What the slack does to the spread

What a tolerance on the second moments does to the estimate's spread, in four worlds. 600 samples a point. Standard deviation at δ = 0, 0.02, 0.05, 0.1, 0.2, 0.5: no square anywhere, 0.1116, 0.1102, 0.1090, 0.1085, 0.1085, 0.1086; a square in the assignment, 0.1243, 0.1213, 0.1176, 0.1136, 0.1091, 0.1076; a square in the outcome, 0.1119, 0.1144, 0.1329, 0.1594, 0.1734, 0.1755; a square in both, 0.1253, 0.1225, 0.1195, 0.1175, 0.1257, 0.1790.
Fig. 3 The standard deviation of the estimate against the tolerance δ in each of the four worlds.

Relaxing a constraint should lower variance, and where the assignment depends on the square it does: the estimate’s standard deviation falls from 0.1243 with exact balance to 0.1136 at a tenth and 0.1091 at two tenths in the assignment-only world, the effective size of the treated arm rising from 249 to 271 and 276. That is the variance the exact balance was paying to force the square’s means together, given back.

Where the outcome depends on the square and the assignment does not, the opposite happens. The standard deviation rises from 0.1119 exact to 0.1594 at a tenth and 0.1734 at two tenths, and the interval’s coverage falls from 93.3% to 80.1% and 76.8%. The band does not bind on average here, but in any one sample the weights are free to leave the square’s gap anywhere inside it, and the gap they leave is whatever the sample’s own chance imbalance was. That gap, multiplied by the outcome’s slope, is a noise term the exact balance removed and the tolerance puts back — and the standard error, built from the balanced columns, does not know it is there, which is why the coverage falls while the bias stays at zero.

So the tolerance trades in a direction set by the world. It saves variance where the assignment depends on the moment, costs variance and coverage where the outcome depends on it, and costs bias where both do. Only the first is what a tolerance is usually adopted for.

Where it does help: samples with no exact weights

How many samples have balancing weights at all, as the tolerance on the second moments grows, at two thin overlaps. 600 samples a point. At δ = 0, 0.02, 0.05, 0.1, 0.2, 0.5: assignment strength 2, 94.2%, 95.7%, 97.2%, 98.7%, 100.0%, 100.0%; assignment strength 2.5, 62.0%, 68.8%, 76.5%, 89.0%, 94.0%, 95.3%.
Fig. 4 The share of samples with balancing weights at all — means exact, second moments within δ — against δ, with no square anywhere and the assignment’s dependence on the covariates multiplied by two and by two and a half, where the overlap between the arms is thin.

The tolerance does what it was proposed for where overlap is thin. With the assignment’s dependence on the covariates two and a half times the earlier essays’ — a treated arm and a control arm that barely share a region — exact balance on six columns has weights in 62.0% of samples, about the share the earlier count found. A band of two hundredths raises that to 68.8%, five hundredths to 76.5%, a tenth to 89.0%, two tenths to 94.0% and half a standard deviation to 95.3%. The samples a tolerance cannot rescue are the ones where even the means cannot be balanced with weights of this form; no tolerance on the squares makes room the means still need.

What the rescued samples deliver is less encouraging. At that overlap the estimate’s standard deviation is 0.29 or so whatever the tolerance, the interval covers between 68% and 75% of the time, and the treated arm’s weights concentrate on a few units — the largest carrying about a tenth of the arm, an effective size of thirty-seven to fifty-one of three hundred. A tolerance makes weights exist where exact balance had none; it does not make the comparison they support informative. How many observations a weight leaves found Kish’s count exact only for outcomes that do not move with the covariates; here the outcome moves with them, and fifty effective units of three hundred is, if anything, an optimistic reading of what the rescued samples hold. The region with no comparison is still a region with no comparison.

Choosing the dial

Do not loosen a moment that the outcome may depend on. The tolerance is spent in full wherever the assignment pushes that moment apart, so its bias is the tolerance times the outcome’s slope on the moment, and a reader cannot tell from the weights whether that slope is zero. A tolerance of a tenth on the square cost a quarter of the effect here.

Loosen moments only where the outcome is known not to need them. Then the tolerance saves variance, as it did in the assignment-only world, and costs nothing. That is a claim about the outcome, and it has to be made on substantive grounds; the balance diagnostics will report every gap within its band and say nothing about whether the band was safe.

Prefer exact balance on the moments a world might need, and pay for it in samples. The moments a balance is told found the exact second-moment balance repairing the fourth world at a cost in variance and, at thin overlap, in samples with no weights at all. The tolerance does not trade between those two costs; it trades the repair itself for them.

Report the gaps, not the tolerance. Every gap here sat at the edge of its band in the worlds where it mattered. A table that reports “within 0.1 standard deviations” invites the reading that the gaps are small and scattered; they are equal to the tolerance and all point the same way.

Use a tolerance to make weights exist, and then ask what they say. At thin overlap a tenth’s band rescued a third of the samples exact balance lost, and the estimates those samples gave covered in seven trials out of ten. The rescue is real and the inference is not.

Computed and counted

Exact: the tolerant fit is the dual of the exact tilt with an L1 penalty on the tolerated coefficients, so a tolerated column’s gap lies strictly inside its band only where its coefficient is zero and sits at the edge wherever the coefficient is not; a gap of δ\delta standard deviations in the centred square, opposite in the two arms, is 2δ22\delta\sqrt2 in the square’s units.

Counted, over six hundred samples of six hundred units in each world and tolerance: the biases of −0.0074, 0.0500, 0.1353, 0.2782, 0.5473 and 0.6983 in the world with a square in both, with coverage from 89.7% to 26.5% at a tenth; the spreads and coverages in the other three worlds; and the share of samples with weights at the thin overlap, from 62.0% exact to 89.0% at a tenth.

Not claimed: anything about tolerances on the means, which are held exact throughout, or about other relaxations — a penalty on the sum of the gaps rather than a band on each, or weights of a different form, which would spend the tolerance differently. The arithmetic of a binding band depends on the band being an edge the weights can sit on.

Still open: a tolerance set by the outcome

A tolerance on a moment costs its gap times the outcome’s slope on that moment, and the slope is the one thing the weights do not use. An outcome model fitted to the control arm estimates it, and the tolerance could be set moment by moment so that each moment’s possible contribution to the bias — its band times its estimated slope — is held below a stated amount, tight where the outcome is curved and loose where it is flat.

Whether a tolerance chosen that way keeps the fourth world’s repair at the variance of a loose band, how much the slope’s own estimation error costs, and whether it is any different in the end from fitting the outcome model and using it directly — which is what either model, but not neither found doubly robust estimators doing — are measurable on the same four worlds and have not been measured here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Covariate balanceDoubly robustEffective sample sizeInverse-probability weightingModel misspecificationOverlapPropensity scoreStandardised difference