A weight that balances, and one that unbalances
The standardised difference between the arms on each covariate, integrated over the population rather than counted in a sample. Unweighted, the arms differ by 0.8310 on the first covariate and 0.6015 on the second, which is what makes the raw difference of arm means 2.7102 against a true average effect of 1.0000. Weighting each unit by one over its own assignment probability removes both differences exactly — -2.78e-17 and -5.69e-19, which is machine precision and not a small number — because the weighted density of the treated arm is the population's own whatever the propensity is. Weighting by a score fitted without the second covariate balances the first to 0.0035 and pushes the second out to 0.7057, further apart than doing nothing.
Weighting one sample into anotherwide6 views
What else it draws
The same object, drawn to answer the other questions the essays put to it.
What a set of inverse-probability weights leaves of the treated arm, by three routes, at six settings of the assignment rule. The integral 1/(π∫φ/e) reads the whole covariate space and falls from 0.9392 to 2.655e-3. Kish's effective size counted in samples of 600 falls only to 0.2861, because almost all of the integral's fall is in a region a sample of six hundred never draws from. And the fraction the variance of the weighted mean actually delivers is lower again — 0.1155 — because the variance is the average of one over the effective size and the effective size averaged is not the same number. At the widest overlap all three agree to 0.05%.
What a thinning overlap does to a stabilised inverse-probability estimate of an average effect of 1.0000, over 600 samples of 600 at each of six settings. The spread rises from 0.1965 to 0.8122 and the root mean square error from 0.1964 to 0.9219, so at the thin end the error is very nearly the whole of the quantity being estimated. The lower line is the share of the arm's weighted total the single largest observation owns, averaged over the same draws: 0.69% to 9.78%, and in the worst single draw of the sweep 82.75%. Coverage of the 95% interval goes from 93.7% to 55.5%.
The bias of three estimators of an average effect of 1.0000, over 600 samples of 600 units with the assignment rule at strength 1, in each of the four cells made by getting each nuisance model right or wrong. The wrong model in both cases is one that omits the second covariate, which the outcome and the assignment both depend on. The outcome model alone is off by 0.8064 whenever it is the wrong one; weighting alone is off by 0.8190 whenever the propensity model is. The augmented estimator built from both is off by -0.0085, -0.0083 and -0.0016 in the three cells where at least one of them is right, and by 0.8118 in the fourth — which is between its two components rather than better than either.
The variance of an inverse-probability estimate weighted by a propensity fitted from the sample, over the variance of the same estimate weighted by the true propensity, paired on the same 500 samples of 600 units at each of five settings. Every reading is below one: the stabilised estimator keeps 27.8% of its true-weight variance where the assignment is nearly a coin toss and 72.0% where it is nearly decidable, and the unstabilised one 30.0% and 49.3%. Neither estimator is materially biased, so this is a variance rather than a trade. The true weights are right about the population and know nothing about the draw; the fitted weights are the value that sets this draw's own imbalance to zero, and that imbalance was what the variance was made of.
What three sets of weights leave of the standardised difference between the arms on each covariate, as a root mean square over 1200 samples of 600 units. The true propensity leaves 0.1317 and 0.1186 — a sampling error, since it is right about the population and knows nothing of the draw. A likelihood fit leaves 0.0770 and 0.0657, having absorbed part of the draw's imbalance as a side effect of fitting the treatment. Weights fitted so that each arm's weighted means are the sample's leave 1.4e-14 and 1.2e-14, which is the arithmetic's floor rather than a small number: the largest gap between a weighted arm mean and the sample mean in any draw is 9.8e-14.
Where it is used
7 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 7 different questions.
- A score that balances Weighting one sample into another
- How many observations a weight leaves Weighting one sample into another
- The region with no comparison Weighting one sample into another
- Either model, but not neither Weighting one sample into another
- The estimated weight is the better one Weighting one sample into another
- A weight fitted to balance Weighting one sample into another
- The moments a balance is told Weighting one sample into another