Series

Exclusion — the series

7 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. The first stage an instrument needs is set by the violation nobody can see. The error each estimator converges on when the instrument has a direct effect of 0.05 on the outcome — a path the exclusion restriction asserts is zero and no sample can check. The instrument's error is δ/π exactly, so it is the reciprocal of the very quantity that made the method work: 1.0000 at a first stage of 0.05 and 0.0833 at 0.60. Least squares carries the confounding instead, at 0.3440 at a first stage of 0.30. The two cross at π = 0.1389, and the crossing is exactly δ times 2.7778 — the first stage an instrument needs is proportional to the violation it is assumed not to have, and below that line the method being corrected is the better estimator.

    The assumption nothing tests

    An instrument buys a causal effect with an assumption no sample can check, and the price is set by the same quantity that made the method work. The first stage it needs is 2.7778 times the violation it is assumed not to have, so a direct effect of 0.05 demands a first stage of 0.1389 and least squares wins below it.

    part 1 · instrument
  2. A weak instrument gives back the problem it was hired for. The counted mean bias of two-stage least squares at 4 instruments and 200 rows, over 2000 draws a setting, against the standard approximation and against the least-squares inconsistency the instrument was brought in to remove. At π = 0.02 the counted bias is 0.3220 ± 0.0142 where least squares is out by 0.3594 — 89.6% of the way back. At π = 0.3 it is 0.0118 against 0.2647. The approximation, the inconsistency over the population first-stage F, tracks the count at the weak end and sits above it in the middle: 0.968, 0.971, 0.918, 0.810, 0.740, 0.722, 0.846 as the ratio of counted to approximated bias.

    Weak, and back where it started

    A consistent instrumental estimate at two hundred rows and a concentration parameter of 0.32 is biased by 0.3220 ± 0.0142 against a least-squares inconsistency of 0.3594 — 89.6% of the way back to the problem it was hired to solve. Just identified, it has no mean at all, and that is measured as a rate rather than assumed.

    part 2 · instrument
  3. The interval that over-covers when the instrument fails. Counted coverage of two nominal 95.0% intervals for the same causal effect, read off the same 2000 draws of 200 rows at each first stage. The exact Anderson–Rubin set covers 95.3% at every setting — flat, because the statistic it inverts is built from y − tβ, which contains no π at all, and is therefore the same number on the same draw whatever the instrument is worth. The conventional interval covers 99.1% at π = 0.02 and 95.6% at π = 0.6: it goes wrong at the weak end by covering too MUCH, at a median width of 7.320, because its standard error is computed from residuals taken at an estimate that has itself gone wrong. A weak instrument does not make this interval lie about its coverage; it makes it useless while telling the truth.

    What the first stage does not know

    A single weak instrument does not make the conventional interval undercover — it makes it cover 99.1% at a width of 7.320. Where the promise actually breaks is many instruments — coverage falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583.

    part 3 · instrument
  4. Four numbers, and only one of them is the question. Four quantities in a population where the effect is not the same for everybody: 40.0% compliers, 25.0% always-takers and 35.0% never-takers, with the compliers carrying an effect 1.0000 larger than everybody else's. The population average effect is 0.5000. The compliers' average effect is 1.1000. A valid instrument converges on 1.1000 — the second of those, not the first, and the two differ by 0.6000. Comparing the treated with the untreated as they stand gives 1.4182, wrong by 0.9182 in the same direction, because always-takers start 1.5000 above never-takers before any treatment happens. The instrument removes the selection and changes the question at the same time, and only one of those is reported.

    Whose effect it is

    With a perfectly valid instrument and no violation of anything, the estimate converges on 1.1000 where the population average effect is 0.5000. The gap is exactly θ(1 − p_c), the always-takers and never-takers cancel out of both halves of the ratio, and five per cent defiers move the answer to 1.2667.

    part 4 · instrument
  5. The same error, caught or invisible, by how it is arranged. The overidentification test's rejection rate against the error the violation actually puts into the estimate, so the two rows are the same estimate being equally wrong. With the whole violation on one instrument the test keeps its size at 5.0% under the null and reaches 86.4% by an error of 0.800. With both instruments violating in the same ratio as their first stages the two Wald ratios are identical, the test has nothing to compare, and it rejects at 5.8% at that same error — its own size. Over 1000 draws of 300 rows at each setting, at a nominal 5.0%. The test is a comparison between instruments and it was never a check on either.

    Two instruments that disagree

    The overidentification test keeps its size at 5.0% and reaches 86.4% power against a violation carried by one instrument. Against the same error carried by both in proportion to their first stages it rejects on 4.6% of draws — its own size — while the estimate is wrong by 0.3000, which is 94.2% of the confounding the instruments were brought in to remove.

    part 5 · instrument
  6. Three intervals as one strength is spread thinner, at a concentration of 8. Coverage of three nominal 95.0% intervals on the same 1000 draws of 200 rows at each count, when a total concentration parameter of 8 is spread over 1 to 32 instruments. Two-stage least squares covers 97.2%, 96.4%, 94.0%, 86.7%, 73.2%, 51.5%. Building each row's fitted treatment from a first stage that never saw that row covers 97.1%, 97.2%, 98.3%, 97.8%, 97.9%, 98.7%. Limited-information maximum likelihood with its conventional standard error covers 97.2%, 96.7%, 95.7%, 90.9%, 85.2%, 79.0%. At one instrument the likelihood estimator is two-stage least squares exactly, which is why the first readings of those two agree to the last draw.

    Leaving each row out of its own first stage

    Spread a fixed first-stage strength over thirty-two instruments and two-stage least squares covers 51.5%. Build each row's fitted treatment from a first stage that never saw that row and the same draws cover 98.7% — through an interval 5.99 times as wide, around an estimate that misses by more than the whole effect on 34.7% of draws. At eight times the strength the same repair covers 95.3% and costs a width factor of 1.66.

    part 6 · instrument
  7. Three intervals as one strength is spread thinner, at a concentration of 8. Coverage of four nominal 95.0% intervals on the same 1000 draws of 200 rows at each count, when a total concentration parameter of 8 is spread over 1 to 32 instruments. Two-stage least squares covers 97.2%, 96.4%, 94.0%, 86.7%, 73.2%, 51.5%. Building each row's fitted treatment from a first stage that never saw that row covers 97.1%, 97.2%, 98.3%, 97.8%, 97.9%, 98.7%. Limited-information maximum likelihood with its conventional standard error covers 97.2%, 96.7%, 95.7%, 90.9%, 85.2%, 79.0%. The same estimate with Bekker's many-instrument standard error covers 97.2%, 97.2%, 97.2%, 95.0%, 94.3%, 93.8%. At one instrument the likelihood estimator is two-stage least squares exactly, which is why the first readings of those two agree to the last draw.

    A standard error that knows about the instruments

    Limited-information maximum likelihood came out least biased when a concentration parameter of 8 was spread over thirty-two instruments, and its conventional interval covered 79.0%. Bekker's many-instrument standard error covers 93.8% on the same draws, at 63% of the jackknife's width — and it gets there with a median standard error of 0.561 against a true spread of 0.797, because it is large on the draws that need it. At eight times the strength it covers 94.9% at 91% of the jackknife's width, and nothing measured here beats it.

    part 7 · instrument

All series