A standard error that knows about the instruments
Worth reading first: The assumption nothing tests · What the 95% refers to.
Leaving each row out of its own first stage measured three estimators on the draws where the conventional instrumental-variable interval collapses: a concentration parameter of 8, spread over one to thirty-two instruments, a thousand samples of two hundred rows at each count. Two-stage least squares fell to 51.5% coverage. The jackknife estimator, which builds each row’s fitted treatment without that row, kept its coverage at 98.7% by widening its interval six-fold. Limited-information maximum likelihood had the least biased median of the three and a conventional interval that covered 79.0%.
That essay withheld one thing from the likelihood estimator: a standard error built for many instruments. The large-sample theory that keeps the ratio of instruments to observations fixed as both grow — exactly the sweep the field runs — derives one, due to Bekker, and predicts an interval wider than the conventional one and narrower than the jackknife’s. Whether that interval covers in samples of two hundred, and at what width, decides which of the two repairs the data actually required. This essay computes it on the same draws.
What the conventional interval leaves out
The likelihood estimator is a k-class estimator, and its conventional variance is the residual variance divided by the treatment’s variation net of κ times the part the instruments cannot explain. That is the variance the estimate would have if the first-stage coefficients were known. With one instrument they nearly are, relative to the signal. With thirty-two, each coefficient carries a sliver of a concentration of 8 and is estimated with far more noise than signal, and the estimate inherits that noise through every one of them.
Bekker’s correction adds that noise back. In the form Hansen, Hausman and Newey give for errors of constant variance, it takes the share of the residual the instruments appear to explain, α̃ = u’Pu/u’u at the estimate, forms H as the treatment’s projected variation minus α̃ times its total variation, and replaces the conventional middle term with the residual variance times a weighted sum of the treatment’s projected and unprojected variation, after the part explained by the residual has been removed. Every piece is a sum the conventional interval already computed or one triangular solve away.
At one instrument the correction vanishes. The just-identified estimator’s residual is exactly orthogonal to the instrument, so u’Pu is zero, α̃ is zero, and Bekker’s standard error is the conventional one to the last digit — 0.371 both, as the median over a thousand draws. As the instruments multiply, α̃ grows: its median is 0.028 at eight instruments and 0.138 at thirty-two, which is roughly the share of rows the instruments’ count takes up, and the correction grows with it.
The large-sample variance the correction estimates has a closed form in this world, where the structural and first-stage errors each have variance one and covariance : one over , plus times one minus 0.36 squared, over one minus times . The first term is the conventional variance. The second is what it leaves out, and it grows with the number of instruments and falls with the square of the concentration, so at a concentration of 8 and thirty-two instruments it is more than four times the first. Like a robust standard error, the correction leaves the estimate exactly where it was and changes only the claim made about its spread.
How much of the spread the conventional interval misses
Read as a share, the closed form says how far the conventional interval is from the truth before a single sample is drawn. At a concentration of 8, the term the conventional variance leaves out is 18.0% of the whole at two instruments, 47.6% at eight and 80.6% at thirty-two. At a concentration of 64 and thirty-two instruments it is 34.1%, and there the two counted standard errors say the same thing independently: the conventional median squared is 70.0% of Bekker’s median squared, so the conventional interval is built on about seven-tenths of the variance the correction finds.
A share missing from the variance translates into a coverage, if the estimates are normal. An interval nominally at 1.96 standard errors, built on only part of the variance, is really an interval at 1.96 times the square root of that part, in units of the true spread. At a concentration of 64 and thirty-two instruments that reading predicts the conventional likelihood interval should cover 88.8%, and it covers 89.6%. The closed form, the normal distribution and the count agree within a point.
At a concentration of 8 the same reading predicts 61.2% at thirty-two instruments, and the interval covers 79.0%. The gap is not an error in the closed form, which matched the counted spread there to the third decimal. It is the shape of the estimates. At the weak end the likelihood estimate’s distribution has a narrow centre and long tails, so a good part of its spread lives in a minority of draws far from the truth; an interval too narrow for the spread still catches the many draws in the centre, and misses mostly in the tails. That is also why the corrected interval’s median standard error can fall short of the spread and its coverage still come close to the promise — the variance is set by draws that a standard error computed from the same sample can see coming.
Coverage at a concentration of 8
On the draws where the conventional likelihood interval covered 97.2%, 96.7%, 95.7%, 90.9%, 85.2% and 79.0% at one to thirty-two instruments, the same estimates with Bekker’s standard error cover 97.2%, 97.2%, 97.2%, 95.0%, 94.3% and 93.8%. The correction does almost all of what it claims to: from eight instruments on, where the conventional interval has begun to fail, it holds the likelihood interval within a point and a quarter of its promise.
It does not quite reach 95% at the weak end, and the shortfall at thirty-two instruments is 1.2 points on a count whose standard error is about 0.7. The jackknife covered 98.7% there and never fell below its promise at any count. On coverage alone, then, the jackknife is the safer interval and the corrected likelihood interval the nearer one.
What each repair costs
The width is where the two repairs part. Bekker’s interval widens from 1.454 at one instrument to 2.198 at thirty-two, 1.76 times the conventional likelihood interval’s 1.250 there. The jackknife’s widens to 3.492. At thirty-two instruments the corrected likelihood interval is 63% of the jackknife’s width, covering 93.8% against 98.7%.
That trade is the one the shortest interval on a table always poses, and here it can be read off both axes at once. The jackknife pays 1.29 units of width to buy 4.9 points of coverage, most of them above the promise it only needed to meet. And the width is not the whole of the difference, because the intervals sit around different estimates: the likelihood estimate’s median is 0.074 above the truth at thirty-two instruments, the jackknife’s 0.201, and the likelihood estimate misses by more than the whole effect on 25.5% of draws against the jackknife’s 34.7%. A narrower interval around a better-centred estimate is not a cheaper version of the jackknife’s; it is a different and better answer that happens to cover a point and a bit less.
Both are still wide. On 53.7% of draws at thirty-two instruments Bekker’s interval is more than two units wide for an effect of one. A concentration of 8 divided among thirty-two instruments says little, and no standard error can make it say more; the correction only stops the interval from pretending otherwise.
A standard error too small in the median that still covers
The closed form and the count agree about how spread out the likelihood estimates are, and they agree by two routes that share nothing. Half the distance between the 16th and 84th percentiles of the thousand estimates is 0.379 at one instrument and 0.797 at thirty-two; the closed-form standard deviation is 0.372 and 0.802. The conventional standard error’s median stays near a third of a unit throughout, 0.319 at thirty-two, less than half the spread.
Bekker’s median is 0.561 at thirty-two instruments. That is seventy per cent of the spread it estimates, and yet the interval built from it covers 93.8%. The two facts are consistent because a standard error is not one number but one per sample, and this one is largest on the samples where the estimate is furthest out. A draw whose first stage came out weak by chance gives a likelihood estimate far from the truth and, through the same weak first stage, a small H and a large corrected variance. The median standard error describes the typical draw, which is not the draw coverage is decided on.
The conventional standard error has no such term, which is why its median and its coverage fail together, and why the ratio of the two medians, Bekker’s over the conventional, is the quickest diagnostic a single study has: 1.76 at thirty-two instruments here, against one at one instrument.
A t-ratio that leans towards least squares
Coverage counts misses without saying which side they fall on, and the standardised error does. Divided by Bekker’s standard error, the likelihood estimate’s error has its 2.5th percentile at −1.305 and its 97.5th at 2.461 at thirty-two instruments, where a standard normal would put both at 1.96. The ratio is lopsided in the direction least squares lies: the estimate overshoots the truth by more than 1.96 standard errors on more than 2.5% of draws, and undershoots by that much on far fewer.
So the corrected interval’s 93.8% is a sum of two errors that partly offset. It misses high too often and low too rarely, and the net is a point and a quarter short. A test of whether the effect is below some value would be anti-conservative and a test of whether it is above would be conservative, which a symmetric interval cannot show. Even at one instrument the percentiles are −1.266 and 2.019, the asymmetry a just-identified estimator with no mean carries at a concentration of 8 and the reason that interval’s 97.2% is not the whole story either.
At eight times the strength
Multiply the strength by eight and the correction becomes what the theory promises. The conventional likelihood interval still undercovers as the instruments multiply, 96.0% at one instrument, 92.4% at sixteen and 89.6% at thirty-two. With Bekker’s standard error the same estimates cover 96.0%, 95.9%, 95.6%, 95.6%, 95.0% and 94.9%. The median estimate is within 0.010 of the truth at every count.
At this strength the standard error and the spread agree in the median as well. At thirty-two instruments Bekker’s median standard error is 0.150, the counted spread 0.162 and the closed form 0.154, while the conventional median is 0.125. The interval is 0.587 wide, 91% of the jackknife’s 0.644, which covers 95.3%, and wider than two-stage least squares’ 0.388, which covers 75.7%. Its standardised error runs from −1.706 to 2.173, still leaning towards least squares but by much less.
Nothing measured in this field does better at a concentration of 64. At thirty-two instruments the corrected likelihood interval covers within a tenth of a point of its promise, around the best-centred estimate, at the narrowest width of any interval that covers.
Which repair the data required
The question the jackknife essay left was whether its honesty had been more expensive than it needed to be, and the answer depends on strength in the way that essay’s own results did. At a concentration of 64 it was: the corrected likelihood interval covers as well and is narrower around a better estimate. At a concentration of 8 it was partly: the corrected interval is 37% narrower, around an estimate that misses by the whole effect less often, and falls a point and a quarter short of its promise on one side. A reader choosing between them at the weak end is choosing between an interval that covers and one that is more useful, and the counts above say what each is worth.
What neither repair touches is the assumption beneath all of them. Every interval here covers the effect the instruments identify only if no instrument reaches the outcome except through the treatment, and with thirty-two instruments that is thirty-two assumptions, each untestable on its own. And every instrument here moves the same effect in every row; when effects differ, each instrument identifies the effect for the people it moves, and a standard error, however well corrected, is a standard error for a weighted average the analyst did not choose.
A check a study can make before trusting either
The shares above are closed forms in the concentration and the count of instruments, and a single study can estimate both. The first-stage F statistic averages roughly one plus the concentration per instrument, so a study with thirty-two instruments and an F near 1.25 is reading this sweep’s weakest column, and the closed form then says the conventional interval leaves out 80.6% of the variance before any interval has been printed. A study with the same thirty-two instruments and an F near 3 has a concentration near 64 and is missing about a third. The check needs no simulation and no second estimator, only the first stage every analysis already reports and the number of rows, and it says in advance which of the two strengths measured here a study resembles — and so whether the conventional interval is merely optimistic or not worth reading at all.
What a many-instrument analysis should report
The likelihood estimate with Bekker’s standard error, beside the conventional one. The ratio of the two is the size of the many-instrument problem in this sample — 1.76 at thirty-two weak instruments here, one at a single instrument — and a ratio well above one says the conventional interval is not to be read.
The concentration per instrument. It decided every comparison: at a total of 8 over thirty-two instruments the corrected interval falls slightly short and is more than two units wide on half the draws; at 64 it covers and is precise.
Which side the interval is likely to miss on. At weak strength the standardised error leans towards least squares, so a one-sided conclusion drawn from a symmetric interval can be wrong in one direction and safe in the other. A ratio whose interval has to be the whole line is the limit this shape is approaching.
Proved, computed and counted
Proved. Bekker’s standard error equals the conventional one at one instrument, because the just-identified residual is orthogonal to its instrument. The closed-form variance is the many-instrument limit in a world whose error variances and covariance are known.
Computed for this world. The closed-form spread at each count and strength, from , , and .
Counted. Every coverage, width, median standard error, spread, percentile and median bias, over a thousand draws of two hundred rows at each count and strength, on the seeds the three earlier estimators used, so the new interval’s rows differ from theirs only in the method.
Particular to this world. Normal errors with constant variance, equal instrument strengths, two concentrations and one confounding strength.
Still open: errors whose variance changes from row to row
Bekker’s standard error, and the likelihood estimator itself, lean on the errors having one variance. When the structural error’s variance differs from row to row, the likelihood estimator stops being consistent as the instruments multiply, for the same reason two-stage least squares was not: a term that averages to zero under constant variance picks up each row’s own leverage when it does not. The standard repair combines both ideas measured in this field — a likelihood estimator whose objective leaves each row out of its own projection, with a standard error built for many instruments and unequal variances. Whether it keeps the corrected likelihood interval’s coverage and width when the error variance moves with an instrument, and how much the jackknife’s wider interval regains its advantage there, is the next measurement.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Intervals for the findings — both name closed form, coverage, interval width, monte carlo
- The check worth more than the check — both name closed form, coverage, interval width, monte carlo
- A coverage table with its own error — both name closed form, coverage, monte carlo
- A flat point with more than one direction — both name closed form, coverage, monte carlo
- A simulation that stops when it looks settled — both name closed form, coverage, monte carlo
- A width rule on skewed outcomes — both name closed form, coverage, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Closed formConcentration parameterCoverageFirst stageInstrumental-variableInterval widthJackknife instrumental-variableLimited-information maximum likelihoodMonte CarloTwo-stage least squaresWeak instrument