The assumption nothing tests
Worth reading first: Which series does the moving · Correcting the persistence.
A regression of an outcome on a treatment that shares a cause with it converges on the wrong number, and no sample size removes the error. In the world every measurement here is read at — a confounder pushing the treatment with a coefficient of 0.6 and the outcome with a coefficient of 0.6 — least squares converges on 1.3303 where the causal effect is 1.0000. The inconsistency is 0.3303 and it is permanent.
An instrument is the standard escape. Find a variable that moves the treatment, argue that it reaches the outcome by no other route, and the ratio of its effect on the outcome to its effect on the treatment converges on the causal effect exactly. Run it in this world and it answers 1.0000. That is not an approximation and it is not a large-sample courtesy: with the second route closed, the answer is right in the limit whatever the confounder is doing.
The argument’s whole weight sits on the second sentence, and the second sentence is an assumption. It is a different kind of assumption from the ones a regression makes about which covariates to include — those can at least be varied and their consequences counted, as the sweep that priced controlling for every covariate available does over four thousand randomly drawn structures. This one cannot be varied and counted, and the second half of this essay is why. What this essay measures is what that assumption is worth when it is slightly false — how much slightly, in units the same world prices least squares in — and then whether any amount of data could have told the difference. The two answers are that the tolerance is much narrower than the assumption’s usual defence suggests, and that no, nothing could.
The error is the reciprocal of the thing that made it work
Write the world down. An instrument , a treatment , an outcome , an unobserved confounder , and everything standard normal and independent:
The exclusion restriction is the claim that . Suppose instead it is not, by some amount too small for anybody to have argued about. The instrument’s estimate converges on
and the shape of that expression is the entire subject. The violation is not added to the answer; it is divided by the first stage. The quantity that makes an instrument worth using is the quantity that magnifies any fault in it.
At a first stage of 0.30 — which nobody would call weak; the population first-stage F is 19 — a direct effect of 0.05 costs 0.1667. That is a twentieth of the causal effect entering as a sixth of it. At a first stage of 0.12 the same violation costs 0.4167, and at 0.02 it costs 2.5000, which is two and a half times the effect being estimated.
The arithmetic that cancels, and why the rule is a straight line
The comparison that decides anything is not between the instrument and a world without confounding. It is between the two estimators available to somebody standing in one world, both of them wrong. Least squares in that same world carries
where is the residual variance of the treatment given the instrument. Note the term: once the instrument has a direct effect, that effect also reaches the ordinary regression, and dropping it is the mistake that makes this question look harder than it is.
Set the two errors equal. The left side is , so multiplying up gives , and the terms cancel on both sides. What is left is , which is linear:
The first stage an instrument needs is proportional to the violation it is assumed not to have, with a constant of proportionality that depends on the confounding and on nothing else. Here that constant is 2.7778. A direct effect of 0.02 demands a first stage of 0.0556; one of 0.05 demands 0.1389; one of 0.10 demands 0.2778. Below the line, the method being corrected is the better estimator.
That cancellation is worth pausing on, because a quadratic was the first answer and it was wrong. Comparing the instrument’s error against the confounding the instrument was hired for — the inconsistency in a world where , rather than the one still present in the world where it is not — gives , which is a genuine quadratic with no real root at all above . Both quantities are kept and named apart here, because they were conflated once and an assertion comparing the two errors at the crossing is what caught it. They answer different questions: one asks whether the instrument beats the problem as it now stands, the other whether it beats the problem it was brought in to solve.
The instrument is the worse estimator on half the grid
A crossing formula is a line, and a line invites the reply that the region on the wrong side of it is small. It is not. Across a grid of nine first stages from 0.02 to 0.60 and five live violations from 0.01 to 0.20, the instrument has the larger asymptotic error in 23 of 45 cells.
The two redrawings above and below are the whole of the sensitivity analysis. At the crossing is at 0.0556 and an instrument would have to be nearly worthless to lose. At it is at 0.2778, which is a first stage many published instruments do not reach, and losing is the ordinary case rather than the pathological one.
What is worth extracting from the three pictures is not any one crossing but that the crossing moves proportionally. There is no first stage that is safe in the abstract. A first stage is safe only relative to a bound on the violation, and the bound on the violation is the thing nobody has.
Two worlds with different causal effects and the same observable law
The usual defence of the exclusion restriction is that it is untested rather than untestable — that the argument for it is substantive, that a clever check might be found, and that in the meantime it is one assumption among many. The first half of that is right and the second is not, and the difference is worth making concrete rather than asserting.
In this world the joint law of is Gaussian with mean zero, so the six second moments are literally everything a sample can contain. Nothing else exists to look at. So take the world above, pick a different causal effect, and solve for the parameters that reproduce the same six numbers. Three adjustments are forced and the algebra is short: matching fixes , matching fixes , and matching fixes what is left of the outcome’s own spread.
Run it at a causal effect of 1.5000 against a truth of 1.0000. The instrument’s direct effect moves from zero to −0.1500, the confounder’s path into the outcome from 0.6000 to −0.2333, and the outcome’s residual spread from 0.8000 to 0.9141. Four parameters move, three of them a great deal. Every one of the six observable second moments moves by zero — exactly zero in closed form, with the largest gap over all six at , which is floating point rather than arithmetic.
The closed form is one route and it could be an algebraic accident, so the moments are also counted. Three hundred paired samples of two thousand rows each, drawn from the two worlds on the same seed, differenced moment by moment and compared against their own Monte Carlo standard errors: the largest discrepancy anywhere is 1.20 standard errors. The two laws are not close, they are the same, and a simulation given every chance to refuse the claim does not.
The estimator’s behaviour is the point. The instrument answers 1.0000 in both worlds — correctly in the first and wrongly in the second — and reports nothing that distinguishes them, because there is nothing to report. This is the same shape as the argument in the essay that fitted three causal structures to one covariance matrix, where the regression returns the same coefficient under three diagrams holding three different effects, and the same shape as the finding that one regression is the same arithmetic whether its covariate is a confounder, a mediator or a collider. The arithmetic never knew which world it was in. What decides is a claim about structure, and a claim about structure is not in the data.
The whole set of effects the moments allow
If two effects fit the same law, a natural next question is how many do. The construction has exactly one limit: the leftover variance of the outcome cannot be negative. That leftover is a quadratic in the candidate effect — every term in it is at most quadratic and none is higher — so three evaluations determine it exactly, and its roots are the endpoints of the set.
The answer is 0.6603 to 2.0597, a span of 1.3994 around a true effect of 1.0000. Every causal effect in that interval has a world behind it producing the identical joint distribution of instrument, treatment and outcome, and the instrument answers 1.0000 in all of them.
The endpoints are computed twice — from the quadratic’s roots, and by bisecting the same function outward from its peak, which shares none of the roots’ arithmetic — and the two agree to twelve digits. That is the discipline the essay on computing every number by two routes argues for, and it earns its keep here: a set derived from a single algebraic manipulation is exactly the kind of result that is wrong in a way nobody notices.
An interval more than one unit wide, containing a true effect of one, is not a confidence interval and should not be read as one. A confidence interval shrinks with the sample and makes a checkable statement about a procedure; this one does not shrink at all, because it is the set that survives having infinite data. Reporting a set rather than a point is what is left when the assumptions on hand pin a parameter only to a range, and it is the honest answer here — but it is an answer nobody wants, since a range that runs from two thirds of an effect to twice it supports no decision the analysis was commissioned to inform.
It is also the quantity that gives the crossing above its meaning. A reader may reasonably object that a violation of 0.05 is a number somebody made up, and the objection is correct: nothing in the data suggests 0.05 rather than 0.20 or zero. What the equivalent set adds is that the data does not suggest anything at all. The violation is not merely unmeasured; it is a direction in parameter space along which the likelihood is exactly flat, and a flat likelihood is the reason a sensitivity analysis has to be run over a range rather than tested at a point.
The bound is what a variance can be, which is a claim about Gaussianity
That interval has an edge, and naming it is the measurement that shows it is about something.
What limits the set is a variance having to stay non-negative — nothing else. That is a weak constraint, and it is weak precisely because the world is Gaussian: with the joint law fixed by second moments alone, the second moments are all the data has to spend and there is nothing to buy identification with. Loosen that and the picture changes. Under heteroskedasticity the six moments are no longer the data, because the way the outcome’s variance moves with the instrument carries information the second moments do not, and there is a body of work identifying a causal effect from exactly that.
So the strict claim proved here is narrower than “an instrument’s exclusion restriction is untestable”. It is that under joint Gaussianity the restriction leaves no observable trace at all, and that the set of effects consistent with what remains is wider than the effect itself. How much of the untestability is a fact about instruments and how much is a fact about Gaussianity is a quantity this essay does not have, and getting it would mean matching four moments rather than two — a different construction rather than a parameter change. It is named here because a claim about what nothing can see should say what “nothing” was allowed to look at.
What a first stage is evidence of, and what it is not
The practical residue is a distinction between two things a first stage does.
A large first stage makes an instrumental estimate precise, and that is what the reported F is about. It also divides the violation, and that is what nothing reports. The two are the same number doing two jobs, and only one of them is ever printed — which is why “the first stage is strong” is heard as a general reassurance when it is a statement about one of the two failures.
The rule this essay measures makes the second job quantitative, and it is easiest read backwards from the crossings already computed. A first stage of 0.1389 is exactly the point at which a violation of 0.05 becomes as expensive as the confounding; a first stage of 0.2778 is that point for a violation of 0.10. So an instrument reporting a first stage of 0.30 is buying protection against direct effects up to about a tenth of the causal effect and no further, and one reporting 0.12 is buying protection against about a twentieth.
Those first stages do not look weak by the usual diagnostic, which is the other half of the problem. The population first-stage F is 19.0000 at and 7.4800 at 0.18, against the conventional threshold of ten — so a first stage passing that threshold comfortably still tolerates only a violation of about a tenth, and one sitting just below it tolerates rather less. The F is a statement about precision. It has been read for two decades as a statement about validity, and the two quantities coincide only in the sense that the same appears in both.
The same structure appears wherever a correction is scaled by an estimated quantity: a repair that moves the wrong number and a bias that lands in the slope rather than the intercept are both cases where the fix is priced by something the fitter chose rather than by the fault.
What none of this says is that instruments are a bad idea, and the grid says why. At the crossing is at a first stage of 0.0278, so anything usable wins comfortably; the confounding is 0.3303 and the instrument’s error is a thirtieth of that. The claim is narrower and less comfortable: the method’s advantage is contingent on a bound nobody supplies, the bound is tighter than the reassurance suggests, and the quantity that makes the estimate look good is the quantity that makes the bound tight.
What this measured and what it assumed
Three things were counted here and one was assumed, and separating them is the point of writing it down.
Counted: the two estimators’ asymptotic errors across a grid, the crossing between them, the six moments of two worlds in closed form and in simulation, and the interval of causal effects those moments allow. Every limit has a closed form beside a count, and the counts were given the chance to refuse — the worst Monte Carlo discrepancy over six moments and three hundred paired samples would have had to reach three and a half standard errors to fail, and reached 1.20.
Assumed: linearity, one instrument, homoskedastic Gaussian errors, and a treatment effect that is the same for everybody. The last of those is not innocuous — with heterogeneous effects a valid instrument converges on something other than the population average effect, which is a second way of answering a question nobody asked, and it needs its own measurement rather than a sentence here.
And one thing is deliberately absent. Nothing in this essay is a test. There is no statistic whose distribution differs between the two worlds, because the two worlds have the same distribution; a test would be an instrument reading on a quantity that does not vary. That absence is the finding, and it is the shape that a rule which removes exactly none of an interaction and a walk that covers half a reference distribution for ever have in common with it: every check that asks whether what is there is right passes, because the defect is that something is not there to check. What can be done instead is to say what the assumption is worth if it fails by a stated amount — which is the crossing — and to report the set that survives it, which is 0.6603 to 2.0597.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a wrong model estimates — both name asymptotic bias, closed form, consistency, least squares
- A dropout the data cannot see — both name non-identifiability, observational equivalence, partial identification
- Adjusting for a shadow — both name asymptotic bias, closed form, unmeasured confounding
- Three mechanisms and one dataset — both name closed form, least squares, non-identifiability
- A collider before the treatment — both name closed form, unmeasured confounding
- A flat point with more than one direction — both name asymptotic bias, closed form
Named objects
A flat tag is an object no other essay names yet.
Asymptotic biasCausal effectClosed formConsistencyEndogeneityExclusion restrictionFirst stageInstrumental-variableLeast squaresNon-identifiabilityObservational equivalencePartial identificationReduced formUnmeasured confoundingWald ratio