Leaving each row out of its own first stage
Worth reading first: The assumption nothing tests · What the 95% refers to.
What the first stage does not know found the place where the conventional instrumental-variable interval really does break. Holding the instruments’ total strength fixed at a concentration parameter of 8 and spreading it over more of them, two-stage least squares covered 97.2% at one instrument and 51.5% at thirty-two, while its interval narrowed. That essay named the standard repair — remove each observation’s own contribution to its fitted treatment — and did not measure it, and it left a question open: whether the collapse is a fact about many instruments or a fact about one estimator applied to them.
This essay runs the repair on exactly those draws: the same world, the same seeds, the same thousand datasets of two hundred rows at each count, so every difference in a row is a difference between methods. The answer to the open question turns out to be both, in proportions set by how strong the instruments are.
Where the bias comes from: each row’s own leverage
Two-stage least squares replaces the treatment with its fitted value from a regression on the instruments, then regresses the outcome on that fitted value. Its error, given the fitted values, is a ratio: the covariance of the fitted treatment with the structural error, over the covariance of the fitted treatment with the treatment. The instruments are independent of the structural error, so the numerator ought to average zero. It does not, and the reason is that each row’s fitted treatment was computed from a regression that included that row.
The fitted value for row i is a weighted sum of every row’s treatment, and the weight on its own treatment is its first-stage leverage, — the same hat-matrix diagonal that decided which rows could hide from a deletion diagnostic. Row i’s own treatment carries row i’s own structural error through the confounding the instrument was brought in to remove. So the numerator splits exactly into two parts: the part carried by each row’s own leverage, and the part carried by every other row. The first has an expectation that is not zero:
where is the covariance the confounding puts between treatment and error, K is the number of instruments, and the leverages sum to K exactly because they are the diagonal of a projection.
Counted, the own-row part averages 0.354, 0.722, 1.434, 2.890, 5.777 and 11.375 at one to thirty-two instruments, against closed forms of 0.358, 0.716, 1.433, 2.866, 5.731 and 11.462. The rest averages 0.062 at thirty-two against a closed form of 0.058 — noise around a number two hundred times smaller than the own-row part. And the denominator, which the fitted values’ own variance inflates by the same mechanism, averages 40.465 at thirty-two against = 39.960. On every one of the six thousand draws the leverages sum to the number of instruments to within .
That is the whole mechanism of the many-instrument bias, and two routes that share nothing agree about it. The ratio of the two expectations at thirty-two instruments, 11.462 over 39.960, is 0.287; the counted median error of two-stage least squares there is 0.279. Each instrument adds 0.36 to the numerator and one to the denominator, and with a concentration of 8 shared among thirty-two instruments the ones in the denominator overwhelm the signal. The estimate goes most of the way back to least squares, whose inconsistency in this world is 0.346, not because the instruments are invalid but because each row is partly its own instrument.
The bias as a rule of thumb, and where it stops being one
The two expectations give a prediction for two-stage least squares’ error at every count, over , and it is worth setting against the counted medians row by row, because the agreement is not uniform and the way it fails is informative. At one instrument the prediction is 0.040 and the median error is −0.016; at two, 0.072 against 0.045; at four, 0.120 against 0.089; at eight, 0.180 against 0.183; at sixteen, 0.239 against 0.261; at thirty-two, 0.287 against 0.279.
The prediction is a ratio of expectations and the median is the centre of a ratio, and the two agree only when the ratio’s denominator is steady enough for its own noise not to move the centre. With one weak instrument the denominator is the concentration of 8 plus one, and it varies from draw to draw by about as much as its size; the ratio’s distribution is skewed and its median sits on the other side of zero from the expectation argument. With thirty-two instruments the denominator is 8 plus 32, dominated by the part the instruments’ noise contributes, and that part is nearly constant — so the prediction and the median meet. The many-instrument bias is most predictable exactly where it is largest, which is also why it is so consistent from study to study: a large count of instruments turns the bias from a random quantity into an almost fixed one.
The quantity that governs all of it is the concentration per instrument, . At a total of 8 over thirty-two instruments it is a quarter, and each instrument’s first-stage coefficient is estimated with far more noise than signal. In the familiar diagnostic, the first-stage F statistic averages roughly for samples this size, so a study reporting an F of about 1.25 on thirty-two instruments is reporting this sweep’s last column — a relation stated here from the standard large-sample approximation, not counted.
Deleting the row from its own fit
The repair follows directly. Rebuild each row’s fitted treatment from a first stage that never saw that row. The leave-one-out fitted value has a closed form that needs no refitting —
with including the intercept’s — and using as a single instrument gives the jackknife instrumental-variable estimator. The own-row term disappears from the numerator by construction, and the K disappears from the denominator’s expectation. Nothing is estimated that was not already on the page; it is one pass over the rows.
On the draws where two-stage least squares covered 97.2%, 96.4%, 94.0%, 86.7%, 73.2% and 51.5%, the jackknife’s conventional interval covers 97.1%, 97.2%, 98.3%, 97.8%, 97.9% and 98.7%. It never falls below its promise at any count. So the collapse was a fact about the estimator: an estimator that does not use each row as its own instrument does not lose its coverage as the instruments multiply.
What the repair costs in width
It is also a fact about many instruments, and the width says so.
Two-stage least squares’ interval narrows from 1.454 at one instrument to 0.583 at thirty-two while its coverage collapses — the pattern the earlier essay found most dangerous, because the printout improves as the answer worsens. The jackknife’s interval moves the other way, from 2.012 to 3.492. At thirty-two instruments it is 5.99 times as wide as the interval it replaces. It covers by being honest about how little a concentration of 8 divided into thirty-two pieces can say, which is the trade the shortest interval on a table always makes in the other direction.
An interval for an effect of one that runs three and a half units wide is correct and nearly uninformative, and it is worth being exact about why it is so wide. The next two figures are the reason.
An estimate that misses by more than the whole effect
Coverage and width describe the interval. The estimate inside it has a distribution of its own, and at this strength it is not a distribution a reader would want to be one draw from.
At thirty-two instruments, the middle ninety per cent of two-stage least squares’ estimates runs from 1.026 to 1.544 — a band that no longer contains the truth, and the picture of an interval that has become confidently wrong. The jackknife’s runs from −3.296 to 4.690. It misses the truth by more than the whole effect — an estimate below zero or above two — on 34.7% of draws, where two-stage least squares does so on 0.0%, and even at one instrument the jackknife misses that badly on 19.4% against two-stage least squares’ 5.5%.
And removing the own-row term does not remove the bias from the median. At thirty-two instruments two-stage least squares’ median sits 0.279 above the truth, with an order-statistic interval from 0.270 to 0.292. The jackknife’s median sits 0.201 above it, with an interval from 0.143 to 0.256 — clearly still on the least-squares side, and 58.0% of the way back to least squares where two-stage least squares is 80.5%. The expectation argument above was about a numerator and a denominator separately; the median of their ratio is a different quantity, and when the denominator is noisy the ratio’s centre is pulled by the noise however unbiased its pieces are.
A denominator that crosses zero
The noise in that denominator is the source of both the width and the misses, and it has a specific form.
The jackknife’s denominator averages close to the signal it estimates at every count, between 6.7 and 7.4 against a concentration of 8. Its spread across draws grows from 5.78 at one instrument to 10.85 at thirty-two, and the share of draws on which it is negative grows from 10.2% to 25.5%. On one draw in four the estimate’s denominator has the wrong sign, and the estimate is on the far side of zero from the effect.
That is the situation a ratio whose interval has to be the whole line is about: a ratio whose denominator has real probability near zero has no finite mean and no bounded interval that covers at every parameter value, and the just-identified estimator with no mean was the first appearance of it in this field. Two-stage least squares’ denominator never crosses zero here, because the K it adds keeps it away — which is exactly the bias. The jackknife takes the K out and puts the ratio back where a ratio of weak signals lives. Every summary of it in this essay is a median or a percentile for that reason; a mean of its estimates is not a quantity with a value.
The likelihood estimator beside it
Limited-information maximum likelihood is the other standard answer to many instruments, and it costs nothing extra to compute on the same draws: it is a k-class estimator with κ the smaller root of a two-by-two determinant, and at one instrument κ is exactly one and it is two-stage least squares, which the first row of every figure shows to the last draw.
Its median is the least biased of the three: 0.074 above the truth at thirty-two instruments, 21.3% of the way to least squares. Its conventional interval stays near the width of two-stage least squares’, 1.250 at thirty-two, and covers 97.2%, 96.7%, 95.7%, 90.9%, 85.2% and 79.0% across the counts. So it undercovers by less than two-stage least squares and more than its promise, and its estimate misses by more than the whole effect on 25.5% of draws at thirty-two. The interval reported beside it here is the conventional one; the correction to its standard error that the many-instrument literature derives for exactly this case is not computed, and the undercoverage should be read as what the uncorrected interval does rather than as what the estimator can do.
At eight times the strength
Everything above is at a total concentration of 8, the strength at which the earlier essay found the collapse. Multiply the strength by eight and run the whole comparison again.
At a concentration of 64 two-stage least squares still undercovers as the instruments multiply — 96.0% at one instrument, 92.5% at eight, 88.0% at sixteen and 75.7% at thirty-two — with its median 0.117 above the truth at thirty-two, 42.9% of the way to least squares. The jackknife covers 96.0%, 96.3%, 95.6%, 95.9%, 95.5% and 95.3%, with its median within 0.019 of the truth at every count and its denominator negative on no draw at all. Its interval at thirty-two instruments is 0.644 wide against two-stage least squares’ 0.388, a factor of 1.66. The likelihood estimator covers 89.6% at thirty-two.
At this strength the repair is a repair: correct coverage, a centred estimate and a modest price in width. At the strength of the earlier essay it is a correct statement that the data say very little. The same algebra does both, and what separates them is whether the signal left after deleting the own-row term is large enough relative to its noise for a ratio to behave.
What is proved here and what is counted
Proved. The split of two-stage least squares’ numerator into own-row and other-row parts is an identity. Its two expectations, and the denominator’s, are closed forms in λγ, K, n and μ², and the leverages summing to K is the trace of a projection. The leave-one-out fitted value’s closed form is algebra, and at one instrument limited-information maximum likelihood equals two-stage least squares exactly.
Counted. Every coverage, width, percentile, median, miss rate and denominator share is over a thousand draws of two hundred rows at each count, from stated seeds shared by all three estimators. A coverage over a thousand draws carries a standard error of about 0.7 points near 95%, and the medians carry distribution-free intervals from the order statistics because two of the estimators have tails too heavy for a standard error of a mean to mean anything.
Particular to this world. One confounding strength, homoskedastic normal errors, equal instrument strengths and two concentrations. The direction of the own-row bias is general, since it follows from the identity; its size relative to the signal, and therefore whether the jackknife’s cure is worse than the disease at a given strength, is not.
Assumed away entirely. Every instrument here moves the treatment by the same amount in every row, so all thirty-two identify one effect. When the effect differs between people, each instrument estimates the effect for the people it moves, and thirty-two instruments are thirty-two weighted averages over thirty-two overlapping groups of compliers. An estimator that pools them — any of the three here — reports a weighted average of those averages, with weights set by first-stage strengths the analyst did not choose. Leaving each row out does nothing about that, and a sweep that gave the instruments different complier populations would be measuring a different question from this one.
Many weak instruments in practice
The setting is not a laboratory curiosity. It is what an analysis produces when a treatment is instrumented by a set of indicators — a birth quarter interacted with state and year, a judge among dozens assigned at random, a set of genetic variants each with a tiny effect on an exposure. Each indicator carries a sliver of the first stage, their total may be respectable, and the concentration per instrument is well under one. Those are the analyses in which two-stage least squares is most confidently wrong in exactly the way measured above: a narrow interval, close to least squares, that looks better identified the more instruments are added.
The two estimators that remove the own-row term answer differently in that setting and the difference is worth carrying. The jackknife reports the truth about the data’s information and pays for it in an interval that may be too wide to use and an estimate that may point the wrong way. The likelihood estimator keeps its estimate close to the truth and, with the conventional standard error, overstates how close. Neither is a free lunch, and the choice between them is a choice between a wide honest interval and a narrow one whose correction has to be supplied separately.
What an analysis with many instruments can report
The concentration parameter, or its estimate, divided by the number of instruments. It is the quantity that decided every contrast above, and a first-stage F statistic reports it directly.
More than one estimator on the same data. Two-stage least squares and the jackknife pulling apart — a narrow interval and a wide one that barely overlap — is itself a finding: it says the own-row term is large relative to the signal, which is the condition under which the narrow interval is wrong.
Medians and percentiles rather than means for any estimator that deletes the own-row term, and the share of resamples or draws on which its denominator changes sign, since that share is the probability the estimate points the wrong way.
Where this goes next: the likelihood estimator with the right standard error
The estimator that came out best on bias here was limited-information maximum likelihood, and the one thing withheld from it was a standard error built for many instruments. The many-instrument asymptotics that derive that correction keep the ratio of instruments to observations fixed as both grow, which is exactly the sweep this field has been running, and they predict an interval that is wider than the conventional one but nowhere near the jackknife’s.
The measurement that would settle it is the corrected interval’s coverage and width on these same draws at both strengths. If it covers near 95% at a width close to two-stage least squares’ at a concentration of 8, it dominates both estimators measured here, and the jackknife’s honesty turns out to have been more expensive than it needed to be. If it undercovers at the weak end, the jackknife’s wide interval was the price the data actually charged. It is distinct from this essay because the question is no longer which estimator removes the own-row bias, but whether a standard error can be made to know about it without giving up the estimator’s precision. Either way, the result still rests on the assumption no instrument count can test.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Three corrections and a leverage — both name closed form, leave-one-out, leverage, monte carlo
- A lead that a heavy tail keeps — both name closed form, heavy tail, monte carlo
- Intervals for the findings — both name closed form, interval width, monte carlo
- Robust is not free — both name interval width, leave-one-out, monte carlo
- The bread and the filling — both name closed form, leverage, monte carlo
- The check worth more than the check — both name closed form, interval width, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Closed formConcentration parameterFirst stageHeavy tailInstrumental-variableInterval widthJackknife instrumental-variableLeave-one-outLeverageLimited-information maximum likelihoodMonte CarloTwo-stage least squaresWeak instrument