The region with no comparison
Worth reading first: A score that balances.
Positivity is usually written down as a condition: every unit must have had some chance of either treatment. Written that way it is a box to tick, and a reader ticks it and moves on. It is better read as a quantity, because it fails by degrees, and every degree of the failure is priced.
Across a single sweep in how strongly the covariates decide the assignment — the same six settings the effective size is read at — the estimator’s spread rises from 0.1965 to 0.8122, a factor of 4.133, rising at every step. The share of the weighted total owned by the single largest observation rises from 0.69% to 9.78%. The coverage of a 95% interval falls from 93.7% to 55.5%. And a procedure that is exactly unbiased in the population carries a counted bias of 0.4374 on an effect of one — nearly half the quantity it is estimating, arriving from nowhere except the sample being finite.
Nothing in that list is a failure of the identity the essay on the balancing score established. The weights still turn each arm into the population exactly. What has changed is that the sample no longer contains the units the identity is asking about.
What is left where nothing is comparable
Take the thinnest setting in the sweep and read the population rather than the estimator. 43.85% of the covariate mass is assigned with probability outside 0.05 to 0.95, and 22.47% outside 0.01 to 0.99. Nearly a quarter of the population is, for practical purposes, decided: it is going to be treated, or it is going to be untreated, and whichever of those it is there is almost nothing in the other arm to compare it to.
The estimator does not decline to answer. It answers by finding the few units that fell the other way and giving them enormous weights. The smallest propensity seen in any draw of the sweep is — one unit, standing in for twenty-five million.
Those two mass figures are the useful way to read a failure of positivity, because they are properties of the population rather than of any sample and they move smoothly. At the widest setting the mass outside 0.05 to 0.95 is zero to two decimal places: there is no region without a comparison, and nothing in the arithmetic is under strain. One step in it is 2.82%, then 12.97%, 24.84%, 35.06% and finally 43.85%. There is no threshold anywhere in that sequence, no setting at which positivity stops holding and starts failing — which is why treating it as a condition to be checked rather than a quantity to be reported loses the whole of the information.
What makes the failure hard to see from inside one dataset is that the units in the region contribute nothing visible. A unit assigned with probability is, almost always, simply absent from the arm it would have had to appear in. The sample does not contain a gap where the comparison should be; it contains no trace of the comparison having been asked for. Every diagnostic that reads what is present reports a healthy dataset, and the failure is that something is not there.
The mean of the largest share is a mild-sounding 6.36% at strength 2, and the mean is the wrong statistic for it. In the worst single draw at that setting one observation owns 82.75% of its arm’s entire weighted total. That estimate is not a weighted average of six hundred rows; it is one row, with a rounding error attached. Twenty draws are shown beside the first for exactly that reason — the quantity being illustrated has a distribution, and its upper tail is where the damage is. It is the same shape the essay on a regression line drawn by one point measures for leverage, arriving through weights rather than through a design matrix.
Unbiased in the population, and biased in the sample
A counted bias of 0.4374 from an estimator with no bias in it needs an account, because “finite-sample bias” is a phrase that explains nothing on its own.
The estimator used here divides each arm’s weighted total by its own summed weights rather than by the sample size. That stabilisation is worth a great deal — an unstabilised version reports an effect scaled by however much the realised weights happened to overshoot — and it makes the estimator a ratio of two random quantities rather than a mean. A ratio of two random quantities is biased whenever the denominator varies, and the denominator here is a sum of weights whose largest term can be a tenth of the total. So the bias is not an error; it is the price of the repair, and it grows with exactly the quantity the sweep is varying.
The two routes agree that this is about the sample rather than about the population. The population overlap has its own integral, computed on a grid with no sample in it, and it says the arms can be balanced exactly at every setting; the counted estimator says the balance is unreachable at the thin end. Both are right, and the disagreement is the whole of what positivity means.
Trimming buys spread and sells the question
The standard repair is to drop units whose propensity is outside some band and weight what is left. It works, on the quantity it is aimed at. At an assignment strength of 2.5, over six hundred draws, the untrimmed estimate has a spread of 0.7418 and a threshold of 0.2 takes it to 0.1755 — a factor of four, which is more than the whole overlap sweep did in the other direction.
And it changes what is being estimated. The average effect over the units a threshold keeps is not the average effect over everybody, because the effect varies with the first covariate and the threshold selects on a function of both. That is a population fact with no sample in it, computable by integration: the untrimmed estimand is 1.0000, and at thresholds of 0.01, 0.05 and 0.2 the trimmed estimand is 0.9531, 0.9258 and 0.9056. The shift at the last of those is −0.0944, about a tenth of the effect, and by then 33.9% of the population is what remains.
The intermediate thresholds are worth reading rather than skipping, because the estimand moves fastest where the trimming is mildest. Between no trimming and a threshold of 0.01 the estimand falls by 0.0469 while the kept population is still 85.4% of the original; between thresholds of 0.15 and 0.2 it barely moves at all, while the kept population goes from 41.6% to 33.9%. That is the shape a reader should expect and rarely pictures: the units nearest the edges of the propensity range are the ones whose effect is furthest from the average, so the first units dropped are the ones that move the target most, and a threshold small enough to look harmless has already done most of the damage.
Read against the target it actually has — the average effect over the units it kept — the trimmed estimator is excellent: 0.2825 untrimmed, 0.0025 at a threshold of 0.2. Read against the target the question was asked about it goes 0.2825, 0.0247, −0.0106, and out to −0.0919 in the other direction. The two columns are the whole finding, and no threshold makes both small.
And at six hundred rows it does reduce the bias, for the wrong reason
The sentence a reader expects at this point — that trimming trades bias for variance, buying spread at the cost of accuracy against the original target — is false as stated on this population at six hundred rows, and the numbers above say so. The distance to the average effect over everybody falls from 0.2825 to 0.0247 at the first threshold and to 0.0106 at the second before it starts growing. For two thresholds, trimming makes the estimate nearer the thing a reader believes is being estimated.
That is not the estimand being hit. It is the previous section’s finite-sample bias being removed. At six hundred rows and strength 2.5 the untrimmed estimator carries an instability of about 0.28 of its own, from the handful of enormous weights, and the first thresholds delete exactly those weights. Two effects are running in opposite directions and cancelling near a threshold of 0.02, where the distance passes through zero and changes sign.
The coverage column at that setting shows the same cancellation from the other side. Untrimmed, the interval covers 65.7% — a catastrophe driven by the instability rather than by any substitution. At a threshold of 0.05 it covers 94.2% against the original target, which is the two errors passing each other, and at 0.2 it is back down to 91.3% as the estimand’s movement takes over. A reader looking only at that column would conclude that a moderate threshold is optimal, and the optimum would be an artefact of two unrelated quantities crossing.
A cancellation is not a demonstration of anything. Both effects have to be separated before the claim can be made, and the way to separate them is that one of them goes away with the sample size and the other cannot. The estimand’s movement is an integral: 0.9056 at a threshold of 0.2 whether the sample has six hundred rows or six million. The estimator’s instability is a finite-sample quantity and shrinks.
An interval that covers less the more data it is given
So the demonstration is a sweep in , at three sizes with the threshold held fixed.
The untrimmed estimator’s own instability behaves as promised: its distance to the average effect over everybody is 0.2950 at six hundred rows, 0.1339 at two thousand four hundred and 0.0681 at nine thousand six hundred. Sixteen times the data, a quarter of the instability.
The trimmed estimator’s distance does not move at all: −0.0925, −0.0952 and −0.0966 at the same three sizes. It cannot move, because it is not an error — it is the gap between two population quantities, and a gap between two population quantities has nothing to shrink with.
What that does to an interval is the sharpest reading in the field. Against the average effect over the units the rule kept, the trimmed interval covers 94.3% at six hundred rows, 94.5% at two thousand four hundred and 96.0% at nine thousand six hundred — it works, and it goes on working. Against the average effect over everybody it covers 90.8%, 81.5% and 41.0%.
The mechanism is arithmetic rather than subtle: the distance stays at about a tenth while the interval’s half-width shrinks like one over the root of the sample, so a fixed offset that was inside the interval at six hundred rows is outside it at nine thousand six hundred. The same trap catches any adjustment that quietly redefines its target, which is why the essay that priced adjusting for every covariate on hand reports what each rule is estimating before it reports how well. An estimator whose interval covers less often the more data it is given is not a noisy estimator of the right thing; it is a precise estimator of something else. More data is not monotonically better — the essay that found an interval getting worse with an extra observation found the same shape in a much smaller setting — and here more data is the instrument that reveals the substitution rather than the thing that causes it.
At six hundred rows alone, all of this is invisible. The trimmed interval covers 91.3% against the original target and 94.0% against its own — two numbers a reader would call close enough, on a sample size a reader would call ordinary.
What would have produced that fall without a substitution
A coverage that collapses from 90.8% to 41.0% is a large enough number to be a bug, and three ordinary explanations have to be closed off before it is read as a finding.
The standard error could be wrong. An interval covers badly if its half-width is computed from a formula that does not describe the estimator’s actual spread, and the influence function for a stabilised weighted estimate on a trimmed sample is not an obvious thing to get right. But the same interval, on the same draws, with the same half-width, covers its own target 94.3%, 94.5% and 96.0% at the three sizes. A half-width that were too small would fail against both targets, not one; a half-width that were too large would over-cover against both. One interval cannot be simultaneously right-sized for one number and wrong-sized for another number, so what differs between the two columns is the number, not the interval.
The trimmed estimator could be biased for its own target too, and merely less so. It is not, and this is what the second column measures: 0.0019 at six hundred rows, −0.0008 at two thousand four hundred, −0.0022 at nine thousand six hundred. Those are zero to the precision the counts support, and they stay zero as the sample grows, which is what consistency looks like.
The three sample sizes could differ in something other than size. They are drawn from the same population with the same rule, the same threshold, and independent seed blocks; the draw counts fall from four hundred to one hundred as the size grows, because the quantity being read is a coverage and a coverage needs the same absolute precision at each. That choice costs precision on the largest cell — a coverage counted over a hundred draws carries a standard error of about two points at 94% and five points at 41% — and neither of those is anywhere near the twenty-point and fifty-point gaps being read.
What survives all three is the arithmetic already stated: a fixed offset of about a tenth, and a half-width shrinking like . At six hundred rows the interval is wide enough to swallow an offset of a tenth several times over; at nine thousand six hundred its half-width is about 0.08, and a tenth is outside it. The interval has not become wrong — it has become narrow enough for the wrongness to show, which is what the essay that watched twenty intervals and expected one miss is about from the healthy side.
The same failure reaches every estimator built on the same weights
Trimming is one response to a thin overlap. Another is to stop using weights alone and add an outcome model, which is the augmented estimator. It does not escape.
At an assignment strength one step past the sweep, the augmented estimator is the least biased thing on the table at 0.0857 and the worst thing on the table at 1.9265, against a plain weighted estimate’s 0.9383. Its interval covers 57.0%. The augmentation term carries a variance that grows with one over the propensity, which is the same quantity that is destroying the weighted estimator, and the essay on what having two models buys is where that is priced. There is no estimator in this field that is not paying for the same missing comparison.
Where a trimmed population stops being one anybody asked about
The honest description of the trimmed estimator is that it estimates the average effect over the units it kept, and that its interval for that quantity is correct at every sample size measured. That is a real answer to a real question, and it is the answer a reader should take when the overlap is thin — provided the switch is reported rather than left for somebody to discover, which is the same discipline the essay on what a 95% interval refers to applies to the word “confidence”.
What has not been measured is the point at which that answer stops being worth reporting. At a threshold of 0.2 the kept population is 33.9% of the original — a third of a population, selected on its own assignment probability. Whether that is a group with a name, or an artefact of a threshold nobody would have chosen for any reason except that the weights were misbehaving, is not something any number in this field addresses. The estimand is well defined and the interval is honest; the question of whether anyone asked about it is not a statistical question, and this field does not pretend to answer it.
Two smaller things are left open. Every setting in the sweep thins the overlap in the same direction, by steepening one scalar in the assignment rule, so “thin overlap” here always means the same shape of failure — a rule with an interaction, or a covariate with a long tail, would thin it differently. And the effect is linear in one covariate, which is what makes the trimmed estimand move by a tenth rather than by a half; with a constant effect it could not move at all, and the whole of this essay would be a section about spread. The essay on how many observations a weight leaves prices the spread; the essay on where an estimate’s weights should come from finds a reduction in it that costs nothing. Neither of them touches the region with no comparison, because nothing does. Assigning by a coin is the one procedure in this collection that makes the region impossible rather than expensive, and that is the whole of what a design buys over an analysis.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Whose effect it is — both name average treatment effect, estimand, treatment effect heterogeneity
- An imputation model the analysis does not contain — both name confidence interval, estimand
- Dropping the incomplete rows — both name confidence interval, estimand
- The count that is not the rows — both name confidence interval, kish effective size
- The draws aimed at the tail — both name confidence interval, heavy tail
- The mechanism the data cannot see — both name confidence interval, estimand
Named objects
A flat tag is an object no other essay names yet.
Average treatment effectConfidence intervalConsistencyEstimandExtreme weightHeavy tailInverse-probability weightingKish effective sizeOverlapPositivityPropensity scoreRoot mean square errorStabilised weightsTreatment effect heterogeneityTrimming