A variable that moves one thing only

Whose effect it is

With a perfectly valid instrument and no violation of anything, the estimate converges on 1.1000 where the population average effect is 0.5000. The gap is exactly θ(1 − p_c), the always-takers and never-takers cancel out of both halves of the ratio, and five per cent defiers move the answer to 1.2667.

Worth reading first: The assumption nothing tests · Which series does the moving.

Everything up to here has been about an instrument that is invalid, or weak, or both. Take one that is neither. The exclusion restriction holds exactly, the first stage is strong, the sample is large, and there is no confounding the method fails to remove. The estimate converges on 1.1000. The average causal effect in the population is 0.5000.

Nothing has gone wrong. The estimator is doing precisely what it is designed to do, and what it is designed to do is answer a question about a subpopulation nobody specified and nobody can identify: the units whose treatment the instrument actually changed. If the effect is the same for everybody, that subpopulation’s average is the population’s and the distinction is empty. If it is not — and the entire reason for wanting a causal effect is usually that it is not — then the number reported and the number wanted differ by an amount the data has nothing to say about.

The population here has three kinds of unit in fixed proportions: 40% compliers, 25% always-takers and 35% never-takers. The effects are arranged so that the population average is held at 0.5000 whatever else moves, which means every number below is a statement about who carries the effect and never about how large the effect is on average.

Four numbers, and only one of them is the question. Four quantities in a population where the effect is not the same for everybody: 40.0% compliers, 25.0% always-takers and 35.0% never-takers, with the compliers carrying an effect 1.0000 larger than everybody else's. The population average effect is 0.5000. The compliers' average effect is 1.1000. A valid instrument converges on 1.1000 — the second of those, not the first, and the two differ by 0.6000. Comparing the treated with the untreated as they stand gives 1.4182, wrong by 0.9182 in the same direction, because always-takers start 1.5000 above never-takers before any treatment happens. The instrument removes the selection and changes the question at the same time, and only one of those is reported.
Fig. 1 Four quantities in a population with 40% compliers, 25% always-takers and 35% never-takers. The population average effect is 0.5000, the compliers’ average is 1.1000, the instrument converges on 1.1000, and comparing the treated with the untreated as they stand gives 1.4182.

Two thirds of the population cancel out of the ratio

The arithmetic is short enough to be worth doing rather than describing.

The Wald ratio is the instrument’s effect on the outcome divided by its effect on the treatment. Consider an always-taker: treated whether the instrument pushes or not, so their treatment is identical in both arms and their outcome is identical in both arms. They contribute exactly nothing to the numerator and exactly nothing to the denominator. A never-taker contributes nothing for the same reason. Only compliers respond, so only compliers appear in either half, and the ratio is their average effect.

Sixty per cent of this population cancels out of both halves of the fraction, and the instrument says nothing about any of them however large that share is. If the always-takers were 80% of the population the estimate would be unchanged; it would still be the compliers’ effect, now computed on a fifth of the units.

It is worth noticing that this cancellation is exact rather than approximate, and that it is a feature of the design rather than an artefact of the world being simple. Nothing about it depends on the effects being drawn from any particular distribution, on the baselines, or on the sample size: an always-taker’s treatment is a constant across the two arms, so their contribution to a difference of arm means is a difference of two identical quantities. The estimator is not ignoring them or downweighting them. They are algebraically absent, and the same sentence applies whatever their effects happen to be — which is why no amount of data, and no diagnostic run on the sample, can report anything about those units at all.

The gap between what is reported and the population average has a closed form. With the compliers carrying an effect θ\theta above everybody else’s, and the scores centred so the population average never moves, the gap is exactly θ(1pc)\theta(1 - p_c). At θ=1\theta = 1 and pc=0.4p_c = 0.4 that is 0.6000, which is 1.1000 against 0.5000, and it is a gap that grows without bound as the effect and compliance become more strongly related.

That cancellation is the same shape as the zero that turned out to be made of two effects pointing opposite ways: a quantity reads a particular value not because the contributions are small but because they are absent from the arithmetic, and nothing in the output distinguishes “these units contributed nothing” from “these units were not there”.

Counted twice, including once by a route no analyst has

A closed form for an estimand is a claim about a limit, and the limit is checked here two ways rather than asserted.

Six hundred samples of four thousand rows, with the instrument assigned by a fair coin and each unit’s treatment fixed by its type, give a counted Wald ratio of 1.1055 ± 0.0040. Against the closed form’s 1.1000 that is a t of 1.39 — agreement, on a comparison that was given every opportunity to fail, since two thousand four hundred thousand simulated units resolve the ratio to four decimal places.

The second route is the one worth having. Inside the simulation every unit’s compliance type is known, so the average effect over the units that are compliers can be computed directly — a quantity no analyst has ever had, because compliance type is never observed for any individual. It comes out at 1.1004. The population average effect over the same draws is 0.4993 against a designed 0.5000.

So three numbers are on the table and they say the same thing three ways: the estimator converges on 1.1055, the actual compliers average 1.1004, and the population averages 0.4993. The first two agree and the third does not, and the disagreement is not error. This is the discipline of computing every number by arithmetic that shares nothing applied where it does the most work — a simulation can be interrogated about quantities the world cannot, and the identity of the estimand is exactly such a quantity.

The target does not move, so only the identified quantity can

A sweep in which two things change at once settles nothing, so this one is constructed to change exactly one.

The effects are centred on the profile at every setting, which fixes the population average at 0.5000 throughout. What moves is how much of the effect sits with the compliers. Raise the relationship by half and the compliers average 1.4000 and the instrument converges on 1.4000; double it and they average 1.7000 and the instrument converges on 1.7000. The population average is 0.5000 in all three.

Four numbers, and only one of them is the question. Four quantities in a population where the effect is not the same for everybody: 40.0% compliers, 25.0% always-takers and 35.0% never-takers, with the compliers carrying an effect 1.5000 larger than everybody else's. The population average effect is 0.5000. The compliers' average effect is 1.4000. A valid instrument converges on 1.4000 — the second of those, not the first, and the two differ by 0.9000. Comparing the treated with the untreated as they stand gives 1.4404, wrong by 0.9404 in the same direction, because always-takers start 1.5000 above never-takers before any treatment happens. The instrument removes the selection and changes the question at the same time, and only one of those is reported.
Fig. 2 The same four quantities with the compliers one and a half units above everybody else. The population average effect is still 0.5000 and the instrument converges on 1.4000.

That construction is what makes the finding a finding. A sweep in which the identified quantity rose because the average effect rose would say nothing at all; here the average effect is pinned, and every movement in the estimand is movement in who is being described.

Four numbers, and only one of them is the question. Four quantities in a population where the effect is not the same for everybody: 40.0% compliers, 25.0% always-takers and 35.0% never-takers, with the compliers carrying an effect 2.0000 larger than everybody else's. The population average effect is 0.5000. The compliers' average effect is 1.7000. A valid instrument converges on 1.7000 — the second of those, not the first, and the two differ by 1.2000. Comparing the treated with the untreated as they stand gives 1.4626, wrong by 0.9626 in the same direction, because always-takers start 1.5000 above never-takers before any treatment happens. The instrument removes the selection and changes the question at the same time, and only one of those is reported.
Fig. 3 And with the compliers two units above everybody else: the instrument converges on 1.7000, the population average is still 0.5000, and the comparison of treated against untreated has moved by less than either.

The third reading in each of those pictures is the naive comparison — the difference between the treated and the untreated as they stand, with no instrument at all. It is 1.4182, 1.4404 and 1.4626 across the three settings, moving by less than a twentieth while the identified quantity moves by six tenths. It barely moves, because most of what it is measuring is not an effect: the always-takers start a full unit and a half above the never-takers before any treatment happens, and that baseline difference dominates. The selection bias is nearly constant while the identified quantity moves by half again, which is why a naive comparison that happens to land near the right answer at one setting is not evidence of anything.

The instrument is still much better than not having one

An essay that stops at “the instrument answers the wrong question” has produced a complaint rather than a measurement, and the measurement cuts the other way.

The comparison available without an instrument is the difference between the treated and the untreated as they stand, and at the standing setting it is 1.4182 against a population average of 0.5000 — wrong by 0.9182. The instrument’s answer of 1.1000 is wrong by 0.6000. So against the quantity a reader probably wants, the instrument removes about a third of the error and leaves the rest, and it does that while removing all of the confounding rather than some of it.

The two errors are also different in kind, and the difference matters more than their sizes. The naive comparison’s 0.9182 is selection bias: it is made of always-takers starting a unit and a half above never-takers, it depends on baselines that have nothing to do with any treatment effect, and it would be there if the effect were exactly zero for everybody. The instrument’s 0.6000 is not bias at all. It is the estimator correctly reporting a different estimand, and it would vanish if the effect were the same for everybody while the naive comparison’s error would not.

That distinction is what makes the honest description of the instrument’s failure hard to state and easy to misunderstand. It is not doing anything wrong; it is doing something else. The equivalent case in a purely observational setting is the sweep that priced controlling for every covariate available, where the failure is genuinely a bias and can be compared against doing nothing on the same scale. Here the two quantities are not on one scale, and averaging them or choosing between them by size is the mistake this section exists to prevent.

What the gap actually depends on

Calling the driver “heterogeneity” is too loose, because heterogeneity alone does not produce a gap. Effects that vary wildly across units but vary independently of compliance leave the compliers’ average equal to everybody’s.

What produces the gap is the correlation between an individual’s effect and their being a complier. At the standing setting that correlation is 0.4399 and the gap is 0.6000. At three times the setting the correlation is 0.8268, the estimand is 2.3000, and the gap is 1.8000 — more than three times the population average effect it is meant to be about.

The instrument answers a question about the compliers. A binary instrument, a binary treatment and four compliance types in fixed proportions — 40.0% compliers, 25.0% always-takers, 35.0% never-takers — with the individual effect made progressively more related to being a complier. The population average effect is held at 0.5000 by construction at every point, so nothing on the horizontal axis moves the target. What the Wald ratio converges on moves from 0.5000 to 2.3000 — a gap of 1.8000 — because it is the average effect among compliers and nothing else. The identified quantity is a fact about who responds to the instrument, and it is unbounded in this gap: it grows exactly as θ(1 − p_c) with the strength of the relationship.
Fig. 4 The identified quantity against the correlation between an individual’s effect and being a complier. The population average is held at 0.5000 throughout; the estimand runs to 2.3000 at a correlation of 0.8268.

A correlation of 0.44 is not exotic, and the substantive story behind it is the one that motivates most instruments anyone uses. People who respond to an incentive to take a treatment are disproportionately people for whom the treatment is worth taking. The design’s own logic — find something that nudges people into treatment — selects for units whose effect is large, and the estimand inherits that selection. The instrument does not merely fail to average over everybody; it averages over exactly the units the design recruited, and the recruitment was on the effect.

That makes this a case of the rule being part of the result rather than a defect in the estimator. The identified quantity depends on what the instrument moved, so two valid instruments for the same treatment identify two different quantities, and there is no sense in which either is wrong. It is the same lesson as one regression standing for three different causal claims: the arithmetic is fixed and the object it refers to is decided outside the arithmetic.

Monotonicity is the second assumption nothing tests

Everything above assumes there are no defiers — no units that take the treatment when pushed away from it and refuse it when pushed towards it. That assumption has the same character as the exclusion restriction: it is substantive, it is usually defended in a sentence, and nothing in the data can check it.

It is also, unlike the exclusion restriction, an assumption whose failure is amplified through the denominator. With defiers present the Wald ratio is not the compliers’ average; it is (pcmcpdmd)/(pcpd)(p_c m_c - p_d m_d)/(p_c - p_d), and the denominator is the share of compliers less the share of defiers.

Replace two per cent of the compliers with defiers and the estimand moves from 1.1000 to 1.1556. Replace five per cent and it moves to 1.2667, an error of 0.1667, with the first stage falling from 0.400 to 0.300. Replace fifteen per cent and it is 2.6000, an error of 1.5000, on a first stage of 0.100.

A subpopulation that does the opposite, priced. What the Wald ratio converges on when some of the compliers are replaced by defiers — units the instrument pushes the other way. The compliers' own average effect is 1.1000 throughout and never moves. What moves is the denominator, which is the share of compliers less the share of defiers rather than the share of compliers: 5.0% defiers take the first stage from 0.400 to 0.300 and the estimand from 1.1000 to 1.2667, an error of 0.1667. At 15.0% it is 1.5000. The error is p_d(m_c − m_d)/(p_c − p_d) exactly, so it grows faster than the share of defiers does, and monotonicity is the second assumption here that no amount of data can check.
Fig. 5 What the Wald ratio converges on as compliers are replaced by defiers. The compliers’ own average effect stays at 1.1000 throughout; the estimand runs to 2.6000 at fifteen per cent, because the denominator has fallen to 0.100.

The compliers’ own average effect is 1.1000 at every point on that sweep. Nothing about them changed. What changed is the divisor, and the error grows faster than the share of defiers does — tripling the share from five per cent to fifteen multiplies the error by nine — because the numerator loses while the denominator shrinks. A subpopulation of one unit in twenty, invisible in every diagnostic, moves the answer by a seventh; one in seven moves it past twice the truth.

The observable symptom of a defier problem is a small first stage, which is indistinguishable from a weak instrument. So the two failures this collection has priced separately — an estimator collapsing back towards least squares and an estimand drifting because the population has units pointing both ways — present identically in the one number a reader is trained to check.

The reversal a subpopulation can produce

There is a sharper version of the defier arithmetic worth stating, because it connects to something already measured elsewhere in this collection.

The Wald ratio with defiers is a weighted difference in which one weight is negative. A negative weight is what allows an aggregate to sit outside the range of the quantities it aggregates, which is exactly the mechanism behind a reversal that is a region rather than a table: the aggregate and every subgroup can disagree in sign, and no amount of data resolves it, because both readings are correct about different quantities.

Here the compliers’ effect is 1.1000, the defiers’ is 0.1000, and the estimand at fifteen per cent defiers is 2.6000 — outside the range of both. An estimate lying outside every subpopulation’s effect is not a symptom of noise or of a bad instrument. It is what a difference of weighted averages with a small denominator does, and it is the only visible trace monotonicity leaves when it fails.

The subpopulation cannot be described, only counted

The natural response to all of this is to report the estimate together with a description of who it is about. That response is not available, and the reason is worth being exact about.

Compliance type is defined by two potential decisions — what a unit would do under each value of the instrument — and only one of the two is ever observed for any unit. So no individual can be classified. What is identified is the proportion: the first stage estimates the share of compliers, at 0.400 here, and that single number is the whole of what a study can say about the population its estimate describes. Their average outcome, their average covariates, whether they are the units a policy would reach — none of it is recoverable, because the set is not a set anybody can list.

This is the sense in which the estimand is not merely local but anonymous. A summary that refers to a subpopulation nobody can exhibit is a summary that cannot be checked against anything, and it has the same defect as four datasets agreeing on every summary statistic while being four different things: the number is right and the number does not pin down what produced it.

The proportion at least moves observably. Adding defiers drops it from 0.400 to 0.300 and then to 0.100 on the sweep drawn above, and a first stage collapsing is the one symptom the failure has. What the proportion cannot do is say which direction the units are pointing, since a compliance rate of 0.300 is produced identically by thirty per cent compliers and by forty per cent compliers alongside ten per cent defiers. The observable quantity is a net, and the two worlds behind it identify quantities that differ by a seventh.

Where this arithmetic stops

The four-type enumeration above is exact, and it is exact because both the instrument and the treatment are binary. That is what makes the population a finite list of types with fixed proportions, and it is what lets every quantity here be computed rather than simulated. It is also the boundary.

With a continuous instrument the four types have nowhere to go. What is identified becomes a weighted average of effects at margins — the units moved from untreated to treated as the instrument passes each of its values — and the weights are the object worth measuring, since a weighted average whose weights nobody has seen is a summary of an unknown population. That is a different construction rather than a parameter change, and it is not measured here. It is the obvious next question against this argument and the honest statement is that this essay does not answer it.

Two smaller limits. The baselines that make the naive comparison wrong — always-takers starting above never-takers — are fixed rather than swept, so 1.4182 is one world’s selection bias and not a general figure. And the correlation reported between effect and compliance depends on an individual-level noise term whose scale is a free parameter of the construction, so the correlation is a legible index of the relationship rather than a quantity to be compared against one measured in a real population. What does not depend on any of that is the gap itself, which is θ(1pc)\theta(1 - p_c) exactly, at every setting on the sweep, to twelve digits.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Always-takerAverage treatment effectCausal effectCompliance typeComplierDefierEstimandFirst stageInstrumental-variableLocal average treatment effectMonotonicityNever takerSelection biasTreatment effect heterogeneityWald ratio