A variable that moves one thing only

Two instruments that disagree

The overidentification test keeps its size at 5.0% and reaches 86.4% power against a violation carried by one instrument. Against the same error carried by both in proportion to their first stages it rejects on 4.6% of draws — its own size — while the estimate is wrong by 0.3000, which is 94.2% of the confounding the instruments were brought in to remove.

Worth reading first: The assumption nothing tests · What a p-value does not say.

One instrument leaves the exclusion restriction with nothing to check. Two leave something over, and the something is the standard reply to everything in this field: if the instruments are both valid they should agree, so test whether they do. The test exists, it is routine, and it works — size 5.0% under the null, rising to 86.4% power against the largest violation swept here.

Then arrange the same violation differently. Give each instrument a direct effect on the outcome standing in the same ratio to its own first stage, so that the two are equally corrupted rather than one of them being wholly so. The estimate is now wrong by 0.3000, which is 94.2% of the least-squares inconsistency the instruments were brought in to remove. The test rejects on 4.6% of draws.

That is not low power. It is the test’s own size: it is behaving exactly as it does when nothing is wrong, because from where it stands nothing is. A test with no power against the case that matters most is the whole of this essay, and the reason is not subtle once it is stated — an overidentification test compares the instruments with each other, and two instruments can agree about the wrong number.

The same error, caught or invisible, by how it is arranged. The overidentification test's rejection rate against the error the violation actually puts into the estimate, so the two rows are the same estimate being equally wrong. With the whole violation on one instrument the test keeps its size at 5.0% under the null and reaches 86.4% by an error of 0.800. With both instruments violating in the same ratio as their first stages the two Wald ratios are identical, the test has nothing to compare, and it rejects at 5.8% at that same error — its own size. Over 1000 draws of 300 rows at each setting, at a nominal 5.0%. The test is a comparison between instruments and it was never a check on either.
Fig. 1 The test’s rejection rate against the error the violation puts into the estimate, under two arrangements of the same violation. With it loaded onto one instrument the rate reaches 86.4%; spread in proportion to the first stages it stays at 4.6% where the estimate is out by 0.3000.

What the statistic is a statistic about

The two-instrument world here has first stages of 0.30 and 0.20 and three hundred rows. Each instrument, used alone, gives a Wald ratio converging on β+δj/πj\beta + \delta_j/\pi_j — the reciprocal amplification priced against the first stage, one instrument at a time. Two-stage least squares combines them, landing at a strength-weighted average of the two.

The overidentification statistic is, up to arithmetic, the difference between what the two say. Load the whole violation onto the first instrument at an error of 0.3 and the first converges on 1.4333 while the second stays at 1.0000; the gap the test measures is 0.4333, and the direct effects producing it are 0.130 and 0.000.

What the test is actually comparing. Two instruments with first stages of 0.30 and 0.20, the whole exclusion violation loaded onto the first. Each instrument used alone converges on β + δ/π, so the first one's answer walks away from the truth at 1.156 by the end of the sweep while the second stays at 0.000. Two-stage least squares lands between them, at a strength-weighted average, and is out by 0.800. It is the vertical gap between the two lines — 1.156 at the far end — that the overidentification test measures, and that gap is what closes to exactly zero when both instruments are wrong in proportion to their strengths.
Fig. 2 What each instrument converges on when the whole violation is loaded onto the first. The first walks away from the truth while the second stays at 1.0000, and the vertical gap between them — 0.4333 at an error of 0.3 — is what the test measures.

Now put the same total damage on both. Set δj=bπj\delta_j = b\pi_j, so the direct effects are 0.090 and 0.060 rather than 0.130 and 0.000. Each instrument’s Wald limit is then β+bπj/πj=β+b\beta + b\pi_j/\pi_j = \beta + b — the first stage cancels, both instruments converge on the same number, and the gap is 0.0000 exactly. Not small: zero, at every point on the sweep, to the last bit the arithmetic carries.

The estimate is wrong by the same bb it was before. Two-stage least squares converges on β+(πδ)/(ππ)\beta + (\pi\cdot\delta)/ (\pi\cdot\pi), which under this arrangement is β+b\beta + b — the identical error, produced by a violation the test cannot see rather than one it can.

Why the sweep is indexed by the error rather than by the violation

A comparison between two arrangements is worthless if they are not doing equal damage, and the obvious way to make them comparable is the wrong one.

Setting both arrangements to the same δ\delta would compare a violation of 0.130 on one instrument against violations of 0.090 and 0.060 on two, and those produce different errors in the estimate. The test would then be being asked to catch two different faults, and a difference in rejection rates would be a difference in what it was looking for rather than in what it could see.

So the sweep is indexed by the error the violation actually puts into the estimate. At every point on the horizontal axis the two arrangements produce the same asymptotic error — 0.3000 at 0.3, 0.8000 at 0.8 — and are therefore the same estimate being equally wrong. The counted errors confirm it: 0.3101 ± 0.0049 against a closed-form 0.3000 at the middle of the sweep, and 0.8121 against 0.8000 at the end.

The violation an overidentification test cannot see. Two instruments, both violating the exclusion restriction, with each one's direct effect in the same ratio to its own first stage. The two Wald ratios are then equal to the last bit — there is nothing for a comparison of them to find — so the test rejects at 5.0% under the null and at 5.8% when the estimate is wrong by 0.800, a span of 1.4% across the whole sweep. The counted estimate tracks the closed-form error throughout: 0.812 ± 0.0055 against 0.800. Over 1000 draws of 300 rows at each setting. A passing overidentification test says the instruments agree; it does not say they are right.
Fig. 3 The proportional arrangement alone: the rejection rate against the error in the estimate. It never leaves its own size across the whole sweep — a span of 1.4 points — while the estimate goes wrong by up to 0.8000.

Across the whole proportional sweep the rejection rate moves by 1.4% — from its size at the null to 5.8% at the largest violation, which is the kind of movement two thousand draws produce from nothing. The test is not weak against this arrangement; it is indifferent to it.

What the pair was hired for, and how much of it survives

The rejection rate is one half of the finding and the other half is what the estimate is worth while the test is passing, because a blind spot over a harmless fault would be a curiosity rather than a defect.

In this two-instrument world with the exclusion restriction intact, least squares converges 0.3186 away from the causal effect. That is the confounding the instruments were brought in to remove, and it is the whole reason for the design. Under the proportional violation at the middle of the sweep the instrumental estimate converges 0.3000 away — leaving 0.3000 of the 0.3186 in place, or 94.2% of it.

So the design has been run, the diagnostic has been passed, and what has been bought is the removal of about a sixteenth of the problem. The remaining error is not confounding any more — it is a direct effect of the instruments on the outcome, arriving through a different route — but a reader has no way to tell those apart and no reason to care, since the number in front of them is wrong by nearly as much either way.

That is what makes the blindness a defect rather than a limitation. A test that missed a fault costing a hundredth of the effect would be a test with a sensible threshold. This one misses a fault that undoes the entire purpose of the analysis, and it misses it not marginally but completely. The situation is the mirror image of a randomisation test paying nineteen points of power for exactness: there a procedure gives up detectable power for a guarantee it can state; here a procedure keeps its power, states a guarantee, and the guarantee is about something else.

Three ways this could have been an artefact

A result this convenient for the essay’s thesis deserves the same scrutiny as one that contradicted it, so here is what was checked.

The test could simply be underpowered at this sample size. Three hundred rows is not many, and a test that rejected at 4.6% against everything would be a test with nothing in it rather than a test with a blind spot. It is refuted on the same draws, at the same sample size, by the same statistic: loaded onto one instrument the identical error is caught on 24.5% of draws and the largest violation on 86.4%. The power exists. It is directional.

The proportional arrangement could be doing less damage than the other. That is exactly what indexing the sweep by the error rules out, and the counted estimates confirm it rather than the design merely promising it — 0.3101 ± 0.0049 under the proportional arrangement against a closed-form 0.3000, and 0.8121 against 0.8000 at the end of the sweep. Both arrangements produce the same error to within a fiftieth at every point.

The exact zero could be floating-point luck. A gap between two Wald limits that comes out as 0.0000 to four decimal places might be a small number rounded, and a small number would mean a test with a little power rather than none. It is not: the identity δj/πj=b\delta_j/\pi_j = b holds at any first stages whatever, the assertion behind it requires the gap to fall below 101210^{-12} rather than merely to look small, and it is the algebra rather than the arithmetic that makes the test’s target vanish.

What none of that rules out is the arrangement being a knife edge — a measure-zero case in a family of violations that are mostly detectable. That objection is answered in part by the substantive argument above and not at all by measurement, and it is the first thing the deferral at the end of this essay names.

The statistic’s whole distribution is the null’s

“Fails to reject” is a weak claim, because a test can fail to reject while its statistic drifts steadily towards the critical value. That would be a test with some information in it, and would make the finding a matter of degree.

It is not. The critical value here is 3.8415, the 95th percentile of χ2(1)\chi^2(1). Under the null the statistic’s median is 0.411; with the estimate wrong by 0.3000 under the proportional arrangement it is 0.406. Both sit essentially at the median of the reference distribution, which is 0.4549. The statistic has not moved towards the critical value at all — its whole distribution is where it would be if nothing were wrong.

Compare the same error under the arrangement the test can see, where the median statistic rises to 1.731 and the rejection rate to 24.5%. That is the contrast worth carrying: one arrangement moves the statistic and the other does not touch it, and the two do identical damage to the answer.

This is the shape a rule that removes exactly none of an interaction has, and the shape a walk that reaches half its reference distribution for ever has: the defect leaves nothing behind, so every check that asks whether the output is right passes, and the failure is that something is not there to be checked. A test statistic sitting at its null median is not evidence of validity. It is the absence of evidence of anything, reported in a format that looks like evidence.

What a passing test actually licenses

The direction the test can see is worth pricing properly, because the essay is not that overidentification tests are useless.

Loaded onto one instrument, the rejection rate runs 24.5% at an error of 0.3, 41.5% at 0.4, 71.7% at 0.6 and 86.4% at 0.8. Against a large violation carried by one instrument the test is genuinely informative. Against a moderate one it is not: at an error of 0.3 — nearly the whole of the confounding the design was meant to remove — it catches one case in four, and three cases in four pass with the estimate badly wrong in the direction the test was built to find.

So even in its own direction the test’s honest reading is that a rejection is informative and a pass is close to uninformative, and outside its direction a pass carries exactly no information at all. What a passing test licenses is the statement these instruments agree, and nothing beyond it. The step from there to these instruments are valid requires the assumption that a violation would have been arranged unevenly, and nothing supplies that assumption.

The shape of the power curve is worth reading as well as its endpoints. It rises slowly at first — 4.9% at an error of 0.05 and 7.1% at 0.1, barely above size — and then steeply. That is what a test whose statistic grows with the square of the discrepancy does, and it means the region where the test is most needed is precisely the region where it is weakest: a violation large enough to matter and small enough to be arguable sits in the flat part of the curve. A diagnostic that only fires once the answer is obviously ruined is a diagnostic doing very little work.

It is worth noticing which arrangement is more plausible substantively. Instruments in the same study are usually variants of one design — two distances, two policy discontinuities, two components of the same eligibility rule — and a mechanism that reaches the outcome directly is likely to reach it through both, roughly in proportion to how strongly each moves the treatment. The invisible arrangement is the natural one, and the visible one requires two instruments whose faults are unrelated in a way their construction argues against.

Two instruments do not narrow the set of effects the data allows

The clearest way to see why the test cannot help is to ask what it would have to be testing.

The first stage an instrument needs is set by the violation nobody can see. The error each estimator converges on when the instrument has a direct effect of 0.05 on the outcome — a path the exclusion restriction asserts is zero and no sample can check. The instrument's error is δ/π exactly, so it is the reciprocal of the very quantity that made the method work: 1.0000 at a first stage of 0.05 and 0.0833 at 0.60. Least squares carries the confounding instead, at 0.3440 at a first stage of 0.30. The two cross at π = 0.1389, and the crossing is exactly δ times 2.7778 — the first stage an instrument needs is proportional to the violation it is assumed not to have, and below that line the method being corrected is the better estimator.
Fig. 4 The single-instrument comparison the pair inherits: the instrument’s asymptotic error is the direct effect divided by the first stage, so the same violation costs more the less the instrument moves the treatment.

Under the proportional arrangement each instrument’s Wald limit is β+b\beta + b, so the pair is observationally identical to a single instrument with the same violation — the second instrument adds precision and no information about validity. Everything measured where two worlds were constructed with identical observable moments therefore carries over unchanged, including the width of the set of causal effects the data cannot rule out.

Every effect in this interval fits the data equally well. Matching a world's six observable second moments fixes the instrument's direct effect, the confounder's path and the outcome's residual variance for any candidate causal effect. The last of those cannot be negative, and that is the only thing limiting how far the causal effect may move: the curve is what is left of the outcome's variance, and it is positive from 0.6603 to 2.0597 — a span of 1.3994 around a true effect of 1.0000. Every world inside that span produces the identical joint law of (z, t, y), and the instrument answers 1.0000 in all of them. The endpoints are read twice, from the quadratic's roots and from a bisection sharing none of its arithmetic, and agree to twelve digits.
Fig. 5 The set of causal effects consistent with the observable moments of the one-instrument world, running from 0.6603 to 2.0597. Adding a second instrument that is wrong in proportion to its first stage does not narrow it.

It is the same argument the equivalent set makes, one level up. A second instrument is another observable, so the natural expectation is that it buys something — more moments, more restrictions, a narrower answer. Under this arrangement it buys precision and nothing else, and precision around a wrong number is what the interval that got shorter as it got wrong is about from the other direction.

The set of effects consistent with the data runs 0.6603 to 2.0597 in the single-instrument world, a span of 1.3994 around a true effect of 1. A second instrument narrows that set only if it constrains the violation, and under the proportional arrangement it constrains nothing — the two instruments impose one restriction between them, and the restriction is satisfied. The test is measuring a quantity that is exactly zero under a fault, which is the definition of a check with no power rather than a check that failed.

Which promise the test was making

The confusion this essay is about is a confusion between two promises, and separating them makes the result read as arithmetic rather than as a paradox.

The promise the test makes: if these instruments identify different quantities, the difference will be detected at a stated rate. That promise is kept, at 86.4% against the largest violation swept.

The promise it is read as making: if the instruments are invalid, this will show. That promise is not made, cannot be made, and is refuted by 4.6% at an error of 0.3000.

The distinction is the same one two corrections for multiple comparisons turn on, and the same one that separates the two intervals measured where the conventional one over-covered while the exact one stayed flat: procedures advertised at one nominal rate control different quantities, and a procedure keeping its own promise perfectly can be useless for the question it is put to. The overidentification test controls a rate about disagreement. Validity is not disagreement, and no amount of power against the first buys any against the second.

There is a second, quieter reason the test reads as more than it is. Its statistic is a specification test, and specification tests are usually reported as passing or failing a model. What is being specified here is not the model but the restrictions the model happens to impose, and one instrument imposes none — which is why the same worry, in the just-identified case, produces no test at all rather than a test that passes. The move from no test to a passing test feels like acquiring evidence, and what has actually been acquired is one comparison between two things that were never independent checks on anything.

What this measured, and what it did not

The sweep is one pair of first stages, 0.30 and 0.20, at three hundred rows and a thousand draws a setting. Both size and power depend on those, and a pair with more unequal strengths would change the arithmetic of what “proportional” means without changing that the proportional gap is exactly zero — that identity is algebraic and holds at any first stages, which is why the assertion behind it demands the gap be under 101210^{-12} rather than merely small.

Three things are left open. The test’s power against partial proportionality — violations that are neither loaded onto one instrument nor exactly proportional — is not swept, and it is the case a real design would actually be in; the two arrangements here are the endpoints of a family whose interior is unmeasured. The rejection rates are read at a single sample size, so nothing here says how the blind spot behaves as the sample grows, though the gap being exactly zero rather than merely small suggests it does not close. And with more than two instruments the test has more than one degree of freedom and more directions to look in, which raises a question this world cannot answer: whether the invisible set of violations shrinks as the count rises, or merely becomes harder to describe. Given what spreading the strength over many instruments does to the interval, the cost of finding out that way would be high.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Asymptotic biasChi squaredCritical valueExclusion restrictionFirst stageInstrumental-variableNull hypothesisObservational equivalenceOveridentificationSargan testSpecification testStatistical powerTwo-stage least squaresUnmeasured confoundingWald ratio