A variable that moves one thing only

What the first stage does not know

A single weak instrument does not make the conventional interval undercover — it makes it cover 99.1% at a width of 7.320. Where the promise actually breaks is many instruments — coverage falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583.

Worth reading first: The assumption nothing tests · What the 95% refers to.

The received account of a weak instrument is that it makes the confidence interval lie: the estimate is unreliable, the standard error does not know it, and the stated 95% is really something much smaller. Counted here, at nine first stages from 0.02 to 0.60 over two thousand draws of two hundred rows each, the nominal 95% interval covers 99.1% at the weakest setting and 95.6% at the strongest. It does not undercover. It over-covers, monotonically, and by four points at the setting where the estimator is worthless.

That is not a lucky world. It is a mechanical consequence of how the interval is built, it can be read off the arithmetic in one line, and it means the received account has the sign wrong.

What follows takes the reversal apart, prices what the over-coverage costs, and then finds where the promise genuinely does fail — which is many instruments rather than one weak one, and which fails in the one direction that leaves no trace in the output: the interval gets shorter as it gets wrong.

The interval that over-covers when the instrument fails. Counted coverage of two nominal 95.0% intervals for the same causal effect, read off the same 2000 draws of 200 rows at each first stage. The exact Anderson–Rubin set covers 95.3% at every setting — flat, because the statistic it inverts is built from y − tβ, which contains no π at all, and is therefore the same number on the same draw whatever the instrument is worth. The conventional interval covers 99.1% at π = 0.02 and 95.6% at π = 0.6: it goes wrong at the weak end by covering too MUCH, at a median width of 7.320, because its standard error is computed from residuals taken at an estimate that has itself gone wrong. A weak instrument does not make this interval lie about its coverage; it makes it useless while telling the truth.
Fig. 1 Counted coverage of two nominal 95% intervals for the same causal effect, read off the same draws at each first stage. The conventional interval covers 99.1% at π = 0.02 and 95.6% at π = 0.60; the exact set covers 95.3% at every setting.

The instrument touches only the denominator

The conventional t-statistic for the causal effect can be written, after the centring an intercept performs,

zεσ^2(β^)zz,\frac{z'\varepsilon}{\sqrt{\hat\sigma^2(\hat\beta)\, z'z}},

where ε\varepsilon is the structural error at the true effect. The numerator is a covariance between the instrument and an error that, under the exclusion restriction, is independent of the instrument by construction. It does not contain the first stage at all. Whatever π\pi is worth, zεz'\varepsilon is the same random variable.

So the first stage reaches the statistic through exactly one channel: the denominator. And the denominator’s variance is not σ2\sigma^2; it is σ^2\hat\sigma^2, computed from residuals taken at β^\hat\beta rather than at β\beta, because β\beta is the unknown. An estimate that has gone a long way wrong leaves large residuals behind it, and large residuals make a large standard error, and a large standard error makes a wide interval.

The estimator’s own failure pays for the interval’s coverage. The inflation factor σ^2(β^)/σ^2(β)\hat\sigma^2(\hat\beta)/\hat\sigma^2(\beta), taken as a median over the draws, is 1.634 at a first stage of 0.02 and 1.002 at 0.60, and it falls monotonically across the sweep in step with the coverage.

The interval is paid for by the estimator's own failure. The residual variance the conventional 2SLS standard error is computed from, divided by the same quantity computed at the true effect, median over 2000 draws of 200 rows. Software estimates that variance from y − t β̂, so an estimate that has gone a long way wrong makes its own residuals large and its own interval wide: the factor is 1.634 at π = 0.02 and 1.002 at π = 0.6. That is the whole of why the conventional interval covers 99.1% rather than ninety-five at the weakest instrument here, at a median width of 7.320 against 0.466 at the strongest. The protection is real and it is not a promise: nothing in the construction arranges it.
Fig. 2 The residual variance the conventional standard error is computed from, divided by the same quantity computed at the true effect. The factor is 1.634 at π = 0.02 and 1.002 at π = 0.60.

Nothing in the construction arranged that protection, which is why it should not be relied on. It is a side effect of estimating a variance at an estimate, and the arithmetic that produces it is the same arithmetic that produces a residual smoother than the errors it came from — a fit’s leftovers are not the quantity the formula was derived for, and here that discrepancy happens to point in the helpful direction.

Useless while telling the truth

An interval covering 99.1% of the time is not good news, and the width says why. At the weakest setting the median width is 7.320, for a causal effect of 1. At the strongest it is 0.466.

A statement that the effect lies somewhere in a range seven units wide, around a truth of one, is technically correct and carries no information. This is the same trade the essay that measured four intervals’ widths and coverages together is about, read from the other end: there the shortest interval was the one that failed its stated level, and here the interval that keeps its promise most comfortably is the one that has stopped saying anything. Coverage alone is not a property worth checking, and a procedure that keeps its stated rate by widening indefinitely is not a procedure a reader should be reassured by.

The width surcharge for the exact alternative is also worth reading the other way round. At a first stage of 0.30 the exact set is 1.157 times as wide as the conventional interval and covers 95.3% against 96.0% — so at moderate strength the conventional interval is already buying its excess coverage with width, in the same currency, and the comparison between the two procedures at that setting is close to a wash. What separates them is not what they cost at a well-behaved setting; it is what each does when the setting stops being well behaved, and that is a property no single reading can show.

The honest summary of a single weak instrument is therefore not “the interval lies”. It is that the interval tells the truth and the truth is useless, and that nothing in the reported output distinguishes that case from a wide interval that is wide because the effect is genuinely uncertain by a small amount of data.

A set whose coverage does not depend on the first stage at all

There is a construction that does not have this problem, and it does not have it by design rather than by side effect.

Test the hypothesis that the effect equals some stated bb by forming ytby - tb and regressing it on the instrument. Under the exclusion restriction, at the true bb, that residual is λu+σee\lambda u + \sigma_e e — it contains no π\pi whatever — and it is independent of the instrument, so the F statistic for the regression is exactly F(1,n2)F(1, n-2). No approximation, no large-sample appeal, no dependence on instrument strength. Collect every bb the test does not reject and the result is a set with exact coverage.

Counted, the set covers 95.3% at every one of the nine first stages, with a spread across the sweep of exactly 0. Not approximately flat — the same draws are covered at every setting, because on one coupling of the random streams the statistic is the same number at every π\pi to fourteen digits. The residue is arithmetic rather than probability: yy and tt are each centred before the statistic is formed, and centring two sums separately then subtracting is not the same rounding as centring their difference.

The point the test is inverted at is 3.888853, which is the 95th percentile of F(1,198)F(1, 198) computed by bisection on its own tail, and it agrees with the square of the two-sided t quantile at 198 degrees of freedom — two routes to a critical value that share no arithmetic. That agreement matters more than it looks, because the whole claim about this set is that its distribution is exact, and an exact claim resting on one implementation of one special function is a claim about the implementation.

What exactness costs, and it is not width

The set is not free, and the price is not the one a reader would guess.

Where both are bounded, the exact set costs 1.042 times the conventional interval’s width at a first stage of 0.60 and 1.157 at 0.30. A four per cent surcharge for a guarantee that holds at every instrument strength is a bargain, and if that were the whole story nobody would use anything else.

The real price is that the set is sometimes not an interval. Inverting the test is a quadratic inequality, and the sign of its leading coefficient decides the shape of the answer: positive gives a bounded interval, negative gives the complement of one — two rays running to infinity in both directions — and zero gives a single ray. That coefficient is positive exactly when the sample first stage clears a threshold the critical value sets. A weak instrument does not make this set wrong; it makes it infinite.

What an honest interval says when it has nothing to say. The share of 2000 draws on which the exact Anderson–Rubin set is unbounded, at each first stage, 200 rows. Inverting the exact test is a quadratic inequality whose leading coefficient changes sign exactly when the sample first stage falls below a threshold the critical value sets, and below it the answer is two rays rather than an interval. That is 94.4% of draws at π = 0.02 and 0.0% at π = 0.6. Where it is bounded it is not much dearer than the conventional interval — a median 0.485 against 0.466 at the strongest setting, a factor of 1.042. The price of never lying about coverage is being allowed to say nothing.
Fig. 3 The share of draws on which the exact set is unbounded, at each first stage. It is 94.4% at π = 0.02, 60.2% at 0.12, 1.2% at 0.30 and 0.0% at 0.60.

At a first stage of 0.02 the set is unbounded on 94.4% of draws; at 0.12 on 60.2%; at 0.30 on 1.2%; at 0.60 on 0.0%. So the honest procedure and the conventional one deliver the same message at the weak end by different means — one says “somewhere in seven units”, the other says “somewhere on the real line” — and only one of them is telling the reader that the data has not identified anything. The price of never lying about coverage is being allowed to say nothing, and being allowed to say nothing is what makes the guarantee worth having.

Three ways this could have been an artefact

A result that reverses the received account earns more scepticism than one that confirms it, so it is worth saying what was checked before it was believed.

It could be simulation noise. A coverage rate counted over two thousand draws has a sampling error of its own, and a four-point excess would be unremarkable if that error were large. It is many times smaller than the gap, and — more to the point — the excess is not one reading. It falls monotonically across nine settings, from 99.1% at the weakest through 98.1% at 0.12 and 96.0% at 0.30 to 95.6% at the strongest, and the inflation factor falls in step with it. A single anomalous cell is noise; nine cells walking in order beside a second quantity that has an arithmetic reason to move with them is a mechanism.

It could be the wrong critical value. The conventional interval here uses a normal quantile, which is what the asymptotic derivation and most software give. A t quantile at 198 degrees of freedom would widen it slightly and raise the coverage slightly at every setting, so it would move both ends of the sweep in the same direction and could not produce the gradient. The gradient is what the claim rests on.

It could be a coincidence of the seeds. The two intervals are read off the same draws at every setting, so the comparison between them is within-draw rather than between-runs, and the exact set’s flatness is checked by a much stronger statement than equal rates: on one coupling of the streams the statistic is the same number at every first stage to fourteen digits, and the inverted quadratic agrees with the statistic it inverts on every draw at every setting. That is the discipline the essay on what a single simulated figure is worth asks for, and the reason a picture of coverage here is a picture of a rate over draws rather than twenty intervals from twenty samples, which is what makes a rate legible and cannot make a gradient legible.

Where the promise genuinely fails, and it gets shorter as it goes

Everything above is about one instrument getting weaker. Hold the strength fixed and change how many instruments carry it, and the reversal reverses.

At a total concentration parameter of 8 — the instruments together explaining the same amount of the treatment at every count — the conventional interval covers 97.2% at one instrument, 94.0% at four, 86.7% at eight, 73.2% at sixteen and 51.5% at thirty-two. A nominal 95% procedure covering barely half the time is a genuine failure of the promise, and it is the failure the received account was reaching for and attributing to the wrong cause.

Where the conventional interval really does fail. Coverage of a nominal 95.0% interval when the same total first-stage strength — a concentration parameter of 8 at 200 rows — is spread across more and more instruments, over 1000 draws each. At one instrument the conventional interval covers 97.2%; at 32 it covers 51.5%, and it does so while getting SHORTER, from a median 1.454 to 0.583. Nothing in the output looks wrong: the estimate has simply gone 80.5% of the way back to least squares, so it stops moving, so its residuals stop being large, so its standard error stops being generous. The exact set covers 95.1%, 94.1%, 94.8%, 95.4%, 94.2%, 95.2% across the same row.
Fig. 4 Coverage of the nominal 95% interval as the same total first-stage strength is spread over more instruments. It falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583; the exact set covers 95.2% at thirty-two.

And the width falls while the coverage does — from a median 1.454 at one instrument to 0.583 at thirty-two. That pairing is the strongest thing in this field, and it is what makes the failure invisible in practice. The protection at the weak end came from an estimate wandering far enough to inflate its own residuals; here the estimate has gone 80.5% of the way back to least squares and stopped moving, so the residuals stop being large, so the standard error stops being generous, so the interval tightens. Everything on the printout improves. A reader comparing two analyses, one with four instruments and one with thirty-two, sees a shorter interval and a larger sample of instruments and concludes the second is better identified.

The exact set is unmoved: 95.2% at thirty-two instruments, on the same draws, against 51.5%.

It is worth being precise about what the count sweep holds fixed, because the finding depends on it entirely. The total concentration parameter is 8 at every point, so the instruments jointly explain the same share of the treatment at one as at thirty-two; what changes is how many coefficients that explanation is spread over, and therefore how much of each one is estimation noise. Nothing here says that adding a genuinely informative instrument to a weak one is harmful. It says that dividing a fixed amount of information into more pieces is, and that the diagnostic every reader is trained to look at — the interval’s width — moves the wrong way while it happens.

The tail underneath all of this

Coverage is a statement about a rate, and rates conceal what the individual draws are doing. The distribution the weak-instrument intervals are built around is the one measured where the just-identified estimator was shown to have no mean, and it is worth putting the two beside each other.

Where the estimate lands ten times the truth away. The share of 2000 draws whose just-identified instrumental estimate falls more than 10 times the true effect away from it, at each first stage, 200 rows. At π = 0.02 it is 5.5% of draws and the central ninety per cent of estimates runs from -4.436 to 6.425 for a true effect of 1; at π = 0.6 it is 0.0% and the same band is 0.786 to 1.188. This is the part of the distribution that has no mean to average away, and it is why every summary of this estimator here is a median.
Fig. 5 The share of draws landing more than ten times the true effect away from it, at each first stage — 5.55% at π = 0.02 and 0.00% at π = 0.30, with the central ninety per cent of estimates running from −4.436 to 6.425 at the weak end.

At a first stage of 0.02, 5.55% of estimates land more than ten times the true effect away from it. An interval 7.320 wide, centred on an estimate that is out by a factor of ten on one draw in eighteen, covers 99.1% of the time — which is another way of saying that most of the coverage is bought by the draws where the estimate is fine and the interval is enormous anyway, not by the draws where the estimate has failed. Coverage is a marginal rate over draws, and a marginal rate can be kept while the conditional rate on the draws a reader would care about is anything at all. That distinction is the whole of the argument that the 95% belongs to the procedure and not to the interval in front of anybody, and it bites harder here than in the ordinary case, because here the two kinds of draw differ by a factor of ten in the estimate.

The exactness is conditional on the assumption nothing tests

The exact set’s guarantee is stated carefully above and it is worth restating as a limitation, because the phrase “exact whatever the first stage is worth” invites a reading it will not support.

The statistic is exactly FF-distributed at the true effect because the residual ytβy - t\beta is independent of the instrument, and that independence is the exclusion restriction. Give the instrument a direct effect on the outcome and the residual at the true β\beta contains δz\delta z, which is a function of the instrument; the statistic is then non-central, the test rejects the true effect more often than its nominal rate, and the set it inverts covers less than it claims. The exact set is exact conditional on the assumption that nothing in the data can check — and it is worth noticing that the same violation is not divided by the first stage here, because the statistic never divides by anything the first stage appears in.

So the two procedures fail under different things, and a reader choosing between them is choosing which failure to be exposed to rather than buying safety. The conventional interval is exposed to instrument strength and to the number of instruments. The exact set is not exposed to either, and both are exposed to the exclusion restriction, which neither can see.

That is a distinction about which promise is being kept, and it is the same distinction two corrections for multiple comparisons turn on: procedures advertised at the same nominal five per cent are controlling different quantities, and keeping one can mean breaking another without anything going wrong. Coverage at every first stage and coverage under a violated exclusion restriction are two different promises, and no procedure in this essay makes the second.

What this leaves undone

The many-instrument failure is shown here and its standard repair is not measured here, and that has to be said plainly because a reader may well know the repair exists.

The usual fix removes each observation’s own contribution to its fitted treatment before the second stage, which breaks the correlation that drives the estimate back towards least squares. It is the reason this failure is normally described as fixable. Leaving each row out of its own first stage runs it on the sweep drawn above, on the same draws, and answers the question this essay leaves open — whether the collapse from 97.2% to 51.5% is a fact about many instruments or a fact about one estimator applied to many instruments. The answer is both: the repair holds coverage at every count, and at this strength it does so with an interval several times as wide around an estimate that often misses by more than the whole effect.

Two smaller things are also left. The exact set is inverted in closed form only at one instrument, and at higher counts the coverage is read from the statistic rather than from a set anybody could report — so “the exact set covers 95.2% at thirty-two instruments” is a statement about a test, and turning it into a reportable region at that count is a numerical problem this does not solve. And every world here is homoskedastic and Gaussian, which is what makes the F(1,n2)F(1, n-2) claim exact rather than asymptotic; under heteroskedasticity the same construction is approximate and its flatness across first stages becomes a claim requiring its own count.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Anderson rubinConcentration parameterConditional coverageConservative intervalCritical valueExact testFirst stageInterval widthMarginal coverageNominal levelOvercoveragePartial identificationStandard errorTwo-stage least squaresWeak instrument