Concept

Two-stage least squares — where it appears

Fitting the treatment on the instruments, then the outcome on the fitted treatment. With one instrument it reduces to the ratio of two covariances, and with several it is a strength-weighted average of what each instrument says on its own — which is why instruments that disagree can be averaged into an answer none of them gives.

Named by 5 essays across one field — each of them below, with the objects they name alongside it.

A weak instrument gives back the problem it was hired for. The counted mean bias of two-stage least squares at 4 instruments and 200 rows, over 2000 draws a setting, against the standard approximation and against the least-squares inconsistency the instrument was brought in to remove. At π = 0.02 the counted bias is 0.3220 ± 0.0142 where least squares is out by 0.3594 — 89.6% of the way back. At π = 0.3 it is 0.0118 against 0.2647. The approximation, the inconsistency over the population first-stage F, tracks the count at the weak end and sits above it in the middle: 0.968, 0.971, 0.918, 0.810, 0.740, 0.722, 0.846 as the ratio of counted to approximated bias.

Weak, and back where it started

A consistent instrumental estimate at two hundred rows and a concentration parameter of 0.32 is biased by 0.3220 ± 0.0142 against a least-squares inconsistency of 0.3594 — 89.6% of the way back to the problem it was hired to solve. Just identified, it has no mean at all, and that is measured as a rate rather than assumed.

instrument · Exclusion
The interval that over-covers when the instrument fails. Counted coverage of two nominal 95.0% intervals for the same causal effect, read off the same 2000 draws of 200 rows at each first stage. The exact Anderson–Rubin set covers 95.3% at every setting — flat, because the statistic it inverts is built from y − tβ, which contains no π at all, and is therefore the same number on the same draw whatever the instrument is worth. The conventional interval covers 99.1% at π = 0.02 and 95.6% at π = 0.6: it goes wrong at the weak end by covering too MUCH, at a median width of 7.320, because its standard error is computed from residuals taken at an estimate that has itself gone wrong. A weak instrument does not make this interval lie about its coverage; it makes it useless while telling the truth.

What the first stage does not know

A single weak instrument does not make the conventional interval undercover — it makes it cover 99.1% at a width of 7.320. Where the promise actually breaks is many instruments — coverage falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583.

instrument · Exclusion
The same error, caught or invisible, by how it is arranged. The overidentification test's rejection rate against the error the violation actually puts into the estimate, so the two rows are the same estimate being equally wrong. With the whole violation on one instrument the test keeps its size at 5.0% under the null and reaches 86.4% by an error of 0.800. With both instruments violating in the same ratio as their first stages the two Wald ratios are identical, the test has nothing to compare, and it rejects at 5.8% at that same error — its own size. Over 1000 draws of 300 rows at each setting, at a nominal 5.0%. The test is a comparison between instruments and it was never a check on either.

Two instruments that disagree

The overidentification test keeps its size at 5.0% and reaches 86.4% power against a violation carried by one instrument. Against the same error carried by both in proportion to their first stages it rejects on 4.6% of draws — its own size — while the estimate is wrong by 0.3000, which is 94.2% of the confounding the instruments were brought in to remove.

instrument · Exclusion
Three intervals as one strength is spread thinner, at a concentration of 8. Coverage of three nominal 95.0% intervals on the same 1000 draws of 200 rows at each count, when a total concentration parameter of 8 is spread over 1 to 32 instruments. Two-stage least squares covers 97.2%, 96.4%, 94.0%, 86.7%, 73.2%, 51.5%. Building each row's fitted treatment from a first stage that never saw that row covers 97.1%, 97.2%, 98.3%, 97.8%, 97.9%, 98.7%. Limited-information maximum likelihood with its conventional standard error covers 97.2%, 96.7%, 95.7%, 90.9%, 85.2%, 79.0%. At one instrument the likelihood estimator is two-stage least squares exactly, which is why the first readings of those two agree to the last draw.

Leaving each row out of its own first stage

Spread a fixed first-stage strength over thirty-two instruments and two-stage least squares covers 51.5%. Build each row's fitted treatment from a first stage that never saw that row and the same draws cover 98.7% — through an interval 5.99 times as wide, around an estimate that misses by more than the whole effect on 34.7% of draws. At eight times the strength the same repair covers 95.3% and costs a width factor of 1.66.

instrument · Exclusion
Three intervals as one strength is spread thinner, at a concentration of 8. Coverage of four nominal 95.0% intervals on the same 1000 draws of 200 rows at each count, when a total concentration parameter of 8 is spread over 1 to 32 instruments. Two-stage least squares covers 97.2%, 96.4%, 94.0%, 86.7%, 73.2%, 51.5%. Building each row's fitted treatment from a first stage that never saw that row covers 97.1%, 97.2%, 98.3%, 97.8%, 97.9%, 98.7%. Limited-information maximum likelihood with its conventional standard error covers 97.2%, 96.7%, 95.7%, 90.9%, 85.2%, 79.0%. The same estimate with Bekker's many-instrument standard error covers 97.2%, 97.2%, 97.2%, 95.0%, 94.3%, 93.8%. At one instrument the likelihood estimator is two-stage least squares exactly, which is why the first readings of those two agree to the last draw.

A standard error that knows about the instruments

Limited-information maximum likelihood came out least biased when a concentration parameter of 8 was spread over thirty-two instruments, and its conventional interval covered 79.0%. Bekker's many-instrument standard error covers 93.8% on the same draws, at 63% of the jackknife's width — and it gets there with a median standard error of 0.561 against a true spread of 0.797, because it is large on the draws that need it. At eight times the strength it covers 94.9% at 91% of the jackknife's width, and nothing measured here beats it.

instrument · Exclusion

Named alongside it

The objects these essays reach for when they reach for this one.

First stageConcentration parameterInstrumental-variableWeak instrumentInterval widthMonte CarloAsymptotic biasClosed formCritical valueHeavy tailJackknife instrumental-variableLimited-information maximum likelihood

All concepts