Weighting one sample into another

How many observations a weight leaves

Kish's effective sample size is exact — for an outcome whose mean does not move with the covariates the weights are built from, the studentised variance reads 1.0680 where the formula says one. For the population's own outcome the same reading is 6.769, rising to 52.497.

Worth reading first: A score that balances.

Kish’s effective sample size is quoted as a rule of thumb, and it is not one. Studentise the weighted mean of an outcome that does not move with the covariates the weights are built from — divide each draw’s weighted mean by the root of what the formula predicts for that same draw — and the variance of what comes out reads 1.0680, 1.1106, 1.1134, 0.9603, 1.0704 and 1.1186 across six settings of the assignment rule. The formula says one. Those readings are one, within a standard error of about 0.058.

Run the identical arithmetic on the population’s own outcome and the same quantity reads 6.769 at the widest overlap and 52.497 at the thinnest, rising at every step in between. Nothing degraded. The formula is exact where it applies and wrong by a factor of fifty where it does not, and the distance between those two columns is what this essay is about.

Two further partings turn out to matter more than the formula itself. A closed form for how much of the treated arm the weights leave falls to 0.0027 across the sweep while the same fraction counted in samples of six hundred only falls to 0.2861 — two orders of magnitude apart, with neither of them wrong. And the fraction the variance actually delivers is lower again, 0.1155, because a variance is the average of one over the effective size and the effective size averaged is a different number.

What the number counts

The weights here are the ones the essay on the balancing score showed remove a covariate imbalance exactly: one over the probability each unit had of receiving the treatment it received. Kish’s quantity is

ESS=(iwi)2iwi2,\mathrm{ESS} = \frac{\left(\sum_i w_i\right)^2}{\sum_i w_i^2},

which equals nn when every weight is the same and falls towards one as a single weight comes to dominate. It has a second form that says what it is rather than how to compute it. Writing cv(w)\mathrm{cv}(w) for the weights’ coefficient of variation — their standard deviation over their mean —

nESS=1+cv2(w),\frac{n}{\mathrm{ESS}} = 1 + \mathrm{cv}^2(w),

exactly, for any set of weights at all. The two expressions share no arithmetic: one squares a sum, the other centres around a mean. They agree to 101210^{-12} on four hand-written weight sets, including one running from 10310^{-3} to 10310^{3}, which is the second route to every design-effect number below.

The second form is the useful one, because it says the cost is the spread of the weights and nothing else. A weighted mean does not care that a weight is large; it cares that the weights differ. Where the assignment rule is nearly a coin toss every unit has a propensity near a half, every weight is near two, the spread is small, and the arm keeps almost all of itself. Where the assignment is nearly decidable a few units have propensities near zero and their weights are enormous, and the arm is a handful of rows in a trench coat.

What one observation owns of a weighted arm. The forty largest inverse-probability weights in one treated arm of 341 units, each as a share of the arm's whole weighted total, at narrow overlap — the assignment rule at strength 2. If every unit were assigned with the same probability each would own 0.29%; here the largest owns 6.30% and the ten largest own 18.08% between them. The open marks are the largest share in each of 20 further draws, because one sample is one sample: they run from 2.25% to 24.11%. Kish's effective size for this arm is 122.9 of its 341 rows.
Fig. 1 The forty largest weights in one treated arm, each as a share of the arm’s whole weighted total, with the assignment rule at strength 2. If every unit were assigned with the same probability each would own about a third of a per cent; the open marks are the largest share in each of twenty further draws.

That picture is one draw, with twenty more beside it because one draw is one draw. What it shows is the shape of the thing being averaged: not a gentle tilt across the arm, but a short cliff at one end and a long flat plain. The effective size is the number that turns that shape into a count.

Exact, for an outcome nobody has

The claim usually attached to the effective size is that the variance of a weighted mean is σ2/ESS\sigma^2/\mathrm{ESS}. Whether that is true depends on something the formula contains no term for, which is the outcome.

Two outcomes are carried through the same draws here. The flat one is the population’s noise with its mean surface removed — an outcome whose conditional mean does not vary over the covariates. There the weights are fixed relative to the noise, the variance of the weighted mean is σ2w2/(w)2\sigma^2 \sum w^2 / (\sum w)^2 identically, and the formula is not an approximation but an equality. That is the column reading 1.0680 through 1.1186.

The shaped outcome is the population’s actual one, whose conditional mean rises with both covariates — the same covariates the weights are built from. There a draw that happens to give a large weight to a unit with high covariates gets a large weighted mean twice over, once through the weight and once through the outcome, and the two reinforce. The formula has no term for that, and the studentised variance goes to 6.769, 16.575, 27.061, 43.399, 48.135 and 52.497.

The effective size is exact, for an outcome nobody has. The variance of a weighted mean over 600 samples, divided by what Kish's effective sample size predicts for that same draw, at six settings of the assignment rule. For an outcome that does not move with the covariates the weights are built from — the population's own noise with its mean surface removed — the reading is 1.0680, 1.1106, 1.1134, 0.9603, 1.0704 and 1.1186: the formula is not an approximation there, it is the variance. For the population's actual outcome the same reading is 6.77 at the widest overlap and 52.50 at the thinnest, and it rises at every step. The formula counts what the weights cost and knows nothing about the outcome, so an outcome whose conditional mean varies over the same covariates carries a second source of variance it has no term for.
Fig. 2 The variance of a weighted mean over six hundred samples, divided by what Kish’s formula predicts for that same draw, at six settings of the assignment rule. For an outcome flat in the covariates the reading is one; for the population’s own outcome it is 6.769 at the widest overlap and 52.497 at the thinnest.

Two things about the exactness column deserve to be said rather than left for a reader to assume. It rises monotonically in neither direction — 0.9603 at strength 2 sits below one — which is what a set of readings scattering around a true value of one looks like and is not what a degrading approximation looks like. And the six readings are not six independent readings: the same noise draws are used at every setting so that the sweep’s comparisons are paired, so this is one reading of the exactness claim taken six ways, at a standard error of about 0.058, rather than six confirmations of it.

The shaped column, by contrast, rises at every step without exception, which is why it is quoted studentised rather than as a raw ratio of variances. A ratio of two variances of a heavy-tailed sum is a noisy object and moves around; the studentised version is monotone and says the same thing.

What could have produced that column without the formula being exact

An exactness claim that comes out at 1.07 when it predicts 1.00 is the kind of result worth attacking before believing, because several uninteresting things produce it.

The first is a flat outcome constructed to be flat. The flat column is the population’s noise with its mean surface subtracted — the true surface, known here because the population is stated. That is not a cheat, but it is a strong condition, and it is worth being explicit that the exactness is being demonstrated on an outcome nobody has access to rather than on a plausible one. In an application the analyst has yy, not yy minus a surface they would have to estimate; the shaped column is the honest picture of what a real outcome does, and it is the one reading 52.497. The flat column exists to establish that the formula is not merely a bound, which matters because a bound that is loose by a factor of fifty and a formula that is exact under a condition nobody meets are different objects and call for different warnings.

The second is a variance estimated on too few draws. Six hundred draws give a variance whose own standard error is about 2/600\sqrt{2/600}, near 6% — so a reading of 1.07 sits about one standard error above one, and a set of six such readings scattering on both sides of one is exactly what a true value of one produces. Had the readings been 1.07, 1.11, 1.11, 1.10, 1.07, 1.12 with none below one, the honest conclusion would have been a small positive bias rather than exactness. The 0.9603 at strength 2 is what makes the column a scatter rather than a shift, and it is the single most informative number in it. Counting the sign of the deviation is the check that the essay on what a 95% interval refers to makes into a habit: a property advertised as exact should miss in both directions.

The third is a studentisation that hides its own denominator. Each draw’s weighted mean is divided by the root of what the formula predicts for that draw, not by the average prediction — so a draw with unusually spread weights is divided by an unusually large number, and the operation could in principle flatten any variance towards one. It cannot flatten the shaped column, which is computed the same way and comes out at fifty. That is the control: one arithmetic, two outcomes, and only one of them lands on the value the formula names.

Three answers to how much of the arm is left

The effective size has a closed form on this population, and it is worth having because it contains no sample at all. Among treated units the covariates have density eφe\varphi and the weight is 1/e1/e, so the mean weight is 1/π1/\pi with π\pi the treated share and the mean squared weight is (1/π) ⁣ ⁣φ/e(1/\pi)\!\int\!\varphi/e. Kish’s fraction is therefore

ESSn1=1πφ/e,\frac{\mathrm{ESS}}{n_1} = \frac{1}{\pi \int \varphi/e},

a pure integral. Beside it sit two counted quantities: the mean of Kish’s formula over six hundred drawn samples, and the fraction the variance of the weighted mean actually delivers, measured as the ratio of an unweighted variance to a weighted one.

At the open end of the sweep all three agree. The integral says the treated arm keeps 0.9392 of itself, the counted Kish fraction says 0.9397, and the variance delivers 0.9101. One step in, they read 0.7234, 0.7377 and 0.6501. Those are three routes to one quantity, sharing no arithmetic, agreeing to within a per cent at the first setting — which is the check that the integral, the sampler and the variance are all describing the same population.

Three answers to how much sample is left. What a set of inverse-probability weights leaves of the treated arm, by three routes, at six settings of the assignment rule. The integral 1/(π∫φ/e) reads the whole covariate space and falls from 0.9392 to 2.655e-3. Kish's effective size counted in samples of 600 falls only to 0.2861, because almost all of the integral's fall is in a region a sample of six hundred never draws from. And the fraction the variance of the weighted mean actually delivers is lower again — 0.1155 — because the variance is the average of one over the effective size and the effective size averaged is not the same number. At the widest overlap all three agree to 0.05%.
Fig. 3 What a set of inverse-probability weights leaves of the treated arm, by three routes, across six settings of the assignment rule. The integral falls to 0.0027, the counted effective fraction only to 0.2861, and the fraction the variance delivers to 0.1155.

At the thin end they are three different numbers: 0.0027, 0.2861 and 0.1155. Two orders of magnitude separate the first from the second. That is the reading this essay was not expecting, and it is not a defect in any of the three.

Why a closed form that reads 687.24 is not wrong

The integral  ⁣φ/e\int\!\varphi/e is the thing the effective fraction is one over, and across the sweep it reads 1.74, 2.32, 4.52, 14.80, 81.06 and 687.24. That last number is where the disagreement comes from, and it is real arithmetic: as the assignment rule steepens, the propensity approaches zero exponentially fast in the covariates while the normal density falls off only like a Gaussian, so φ/e\varphi/e diverges and the integral is dominated by a region far out in the tail.

Far out in the tail is a place a sample of six hundred does not go. The integral is reading the whole covariate space, including the corner where a unit would be assigned with probability 4×1084\times10^{-8} and would carry a weight of twenty-five million. No draw of six hundred rows from a standard normal pair reaches that corner. Kish’s formula computed on the draw prices only the weights the draw actually contains, and those weights, while large, are nothing like twenty-five million.

So the closed form and the count are answering different questions and both are answering theirs correctly. The closed form says what a set of weights would cost if the sample were the population. The count says what a set of weights did cost in a sample of six hundred. A quantity computed before the data is a prediction about the data only where the data goes, and this is the cleanest instance of that in the field: an asymptotic quantity whose value is dominated by a region no finite sample visits.

The practical residue is uncomfortable rather than reassuring. The count is the smaller worry only because the sample is small: at six hundred rows the sample cannot see how bad the overlap is, so the diagnostic computed on it reports a milder failure than the population has. A larger sample would reach further into the corner, find the weight of twenty-five million, and report an effective size that had fallen. This is one of the few diagnostics on this site that gets worse as more data arrives, and it gets worse because it is getting more honest.

What the variance delivers is a harmonic mean

The third route sits below the second at every setting, and the gap is not noise. It runs 1.033, 1.135, 1.341, 1.727, 2.169 and 2.477 — negligible where the weights are near equal, a factor of two and a half where they are not.

The reason is Jensen’s inequality applied the way round nobody applies it. The effective size is a property of a draw: each sample has its own weights and therefore its own ESS\mathrm{ESS}. The variance of the weighted mean across draws is the average of σ2/ESS\sigma^2/\mathrm{ESS}, so the effective size the variance delivers is the harmonic mean of the per-draw effective sizes. Quoting the arithmetic mean of ESS\mathrm{ESS} and then dividing σ2\sigma^2 by it is the wrong average, and it is wrong in the optimistic direction — the harmonic mean is always the smaller.

Where the weights are near equal every draw has nearly the same effective size, the two means coincide, and the gap is 1.033. Where the weights are wild the per-draw effective size is itself a wildly varying quantity, and the draws that happen to have a very small effective size contribute very large variances that the arithmetic mean of the effective sizes never sees. The distribution of the effective size across draws is part of the answer, and a single reported number has thrown it away.

The estimator has no upper bound on what it costs. What a thinning overlap does to a stabilised inverse-probability estimate of an average effect of 1.0000, over 600 samples of 600 at each of six settings. The spread rises from 0.1965 to 0.8122 and the root mean square error from 0.1964 to 0.9219, so at the thin end the error is very nearly the whole of the quantity being estimated. The lower line is the share of the arm's weighted total the single largest observation owns, averaged over the same draws: 0.69% to 9.78%, and in the worst single draw of the sweep 82.75%. Coverage of the 95% interval goes from 93.7% to 55.5%.
Fig. 4 What a thinning overlap does to the estimate itself: the spread rises from 0.1965 to 0.8122 and the coverage of a 95% interval falls from 93.7% to 55.5%, on the same draws every effective-size reading above is taken from.

That figure is the effective size cashed out. The three fractions above are statements about what the weights cost; the spread and the coverage are what the cost buys, and the essay on the region with no comparison is where they are read.

Two design effects, and one word between them

A sample worth fewer observations than it has rows is a thing this collection has priced before, from a completely different picture. The essay on three-level grouping computes a design effect of 1+(m1)ρ1 + (m-1)\rho from a correlation between observations inside a cluster, and the essay on autocorrelated series finds a fifty-point series with a lag-one correlation of 0.8 worth about six independent observations. This field computes 1+cv2(w)1 + \mathrm{cv}^2(w) from the spread of a set of weights. Both are called the design effect.

They are not the same number, and on this population that is measurable rather than arguable. At an assignment strength of 2, the weighting design effect is 4.6721. Split the same population into ten bands of the propensity, read the intraclass correlation of the outcome surface across those bands — 0.9076 — and the clustering design effect is 54.5505. Same population, same day, a factor of eleven between them.

Two design effects, one word. How many rows a sample of 600 is worth one over, at six settings of the assignment rule, over 600 samples apiece. Read from the spread of the weights — one plus their squared coefficient of variation, which is Kish's effective size written the other way up — the design effect runs from 1.0642 where the assignment is nearly a coin toss to 8.2568 where it is nearly decidable. The bottom bar is the design effect the grouped-data field computes for the same population, one plus the number in a band times the intraclass correlation across ten bands of the propensity: 54.5505, at an intraclass correlation of 0.9076. Both say how many independent observations a sample is worth; they are functions of different things and they are not the same number.
Fig. 5 The design effect computed from the spread of the weights, across six settings of the assignment rule: 1.0642 where the assignment is nearly a coin toss, 8.2568 where it is nearly decidable. The bottom bar is the clustering design effect on the same population, 54.5505 at an intraclass correlation of 0.9076.

Neither is wrong, because they are functions of different things. The weighting version knows only the weights and nothing about the outcome; the clustering version knows only the outcome’s correlation structure and nothing about the weights. A reader who takes “the design effect of this study is 4.67” and divides a sample size by it has made a claim about the weights that is true, and a claim about how many independent observations the study contains that is not supported by anything measured. Say which one. The same care is why the essay that pooled a proportion across groups reports its intraclass correlation beside its design effect rather than the design effect alone.

Where the weights came from is not in the formula

Kish’s number is a function of a set of weights, so two different sets of weights with the same spread get the same answer — and that is a real limitation rather than a pedantic one, because the weights actually used in practice are fitted rather than known.

The weights that balance best are not the true ones. What each set of weights leaves of the standardised difference between the arms, as a root mean square over 600 samples of 600 units. Weighting by the true propensity leaves 0.1396 and 0.1310 — the size of a sampling error, since the true weights balance the population exactly and say nothing about the draw. Weighting by a propensity fitted from the same sample leaves 0.0837 and 0.0733, because the fitted score is by construction the value that sets the draw's own imbalance to zero. A score fitted without the second covariate balances the first to 0.0487 and leaves the second at 0.7143, which is worse than the 0.6038 it started at.
Fig. 6 What each set of weights leaves of the standardised difference between the arms, as a root mean square over six hundred samples. True weights leave 0.1396 and 0.1310; weights fitted from the same sample leave 0.0837 and 0.0733.

Weights fitted from the sample leave a little over half the residual imbalance the true weights leave, on both covariates. Their spread is not materially different, so Kish’s formula prices the two sets almost identically — and the estimator built on them has a variance not quite half as large. The effective size cannot see that difference at all, because the difference is not in the weights’ spread; it is in what the weights are correlated with, which is the particular draw they were fitted on.

Estimating a weight you already know is worth doing. The variance of an inverse-probability estimate weighted by a propensity fitted from the sample, over the variance of the same estimate weighted by the true propensity, paired on the same 500 samples of 600 units at each of five settings. Every reading is below one: the stabilised estimator keeps 27.8% of its true-weight variance where the assignment is nearly a coin toss and 72.0% where it is nearly decidable, and the unstabilised one 30.0% and 49.3%. Neither estimator is materially biased, so this is a variance rather than a trade. The true weights are right about the population and know nothing about the draw; the fitted weights are the value that sets this draw's own imbalance to zero, and that imbalance was what the variance was made of.
Fig. 7 The variance of an estimate weighted by a fitted propensity over the variance of the same estimate weighted by the true one, paired on the same samples at five settings. Every reading is below one.

That is the essay on why an estimated weight beats a known one, and it is the clearest evidence that the effective size is a statement about weights rather than about estimators. A reader who wants to know how precise an estimate is has to measure the estimate.

Two things this essay does not settle are worth naming. The first is that every covariate here is standard normal, so every band of the score is a band of a normal linear predictor and the integral  ⁣φ/e\int\!\varphi/e takes the particular shape it does because of that. The balancing identity is exact whatever the covariate law, but a skewed or discrete second covariate would move this integral a long way and might move the parting from the count with it; nothing here tests that. The second is that all of this prices a single weighted mean. The essay on inverse-variance weights is about the case where the weights are chosen to minimise a variance rather than forced by an assignment rule, and the two problems share a formula and not a purpose. What the weights cost here was never optional, and what a coin buys instead is the alternative that has no weights in it at all — as is arranging the units in pairs before assigning them, which removes a share of the variance that is knowable before a single row is drawn.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formCoefficient of variationDesign effectEffective sample sizeExtreme weightHarmonic meanIntraclass correlationInverse-probability weightingJensens inequalityKish effective sizeOverlapPositivityPropensity scoreStabilised weightsVariance ratio