Cells of unequal size and an error that follows them
Worth reading first: The assumption nothing tests.
A standard error that knows about the instruments spread a weak first stage over thirty-two instruments and found the combination that worked: limited-information maximum likelihood, the least biased of the estimators tried, with Bekker’s many-instrument standard error, which covered 93.8% where the conventional interval covered 79.0%. Every draw in that essay had errors of one variance, and the essay ended on why that mattered. Bekker’s standard error and the likelihood estimator itself both lean on it. When the structural error’s variance differs from row to row, the likelihood estimator stops being consistent as instruments multiply, for the same reason two-stage least squares was not.
The instruments in that essay were thirty-two independent normal columns. That is a convenient world and an unusual one. Many instruments in practice are indicators — quarter of birth crossed with state, the judge a case was assigned to, the school a lottery placed a child in — and the cells they mark are of very different sizes. That shape turns out to be what decides whether unequal variances matter at all.
Thirty-three cells and their leverage
The world is the earlier essays’ standing one: a treatment confounded with the outcome through a shared unobserved factor, a true effect of one, two hundred rows. The instruments are now the indicators of thirty-three cells, all but one, so there are thirty-two of them. The cells run in size from two units to fifteen, geometrically, and each cell moves the treatment up or down by the same amount, set so that the concentration parameter is 8 — the weak case — or 64.
With cell indicators as instruments, a row’s first-stage leverage — the weight its own treatment carries in its own fitted value — is exactly one over its cell’s size. A unit in a cell of two has leverage 0.500; a unit in a cell of fifteen, 0.067. Two-stage least squares’ fitted treatment for a unit in a small cell is half that unit’s own treatment, and the bias leaving each row out of its own first stage measured — each row’s own confounding reaching its own fitted value — is concentrated in those rows.
The heteroskedasticity is put where it is most plausible and most damaging. The confounding’s variance in cell is times what it is in a cell of average size: at every cell is alike; at a cell of two carries 3.03 times the average-sized cell’s variance and a cell of fifteen 0.40 of it. Small cells being noisier is the ordinary state of affairs — a small court has a narrower mix of cases, a small state a less averaged population — and it means the covariance between the structural error and the first-stage error is largest exactly where the leverage is.
What the likelihood’s correction assumes
Limited-information maximum likelihood is a k-class estimator. It takes two-stage least squares’ quadratic forms in the projection and subtracts times the same forms in the residual space, with found from the data. With errors of one variance that subtraction removes the own-row bias in expectation, because the bias is times the average covariance and the residual-space forms estimate that same average.
With cell indicators the two averages have a plain reading. Each unit in cell has leverage , so the own-row bias sums, over cells, units times times the cell’s covariance: every cell counts once, whatever its size. The residual-space forms the likelihood’s is computed from sum over units, so every unit counts once, and a cell of fifteen counts seven and a half times as much as a cell of two. The bias weights cells equally and the correction weights units equally. With every cell alike the two weightings agree. With the small cells noisier, the bias is driven by the many small cells and the correction by the few large ones.
With the covariance moving across rows the bias is no longer times the average. It is the sum, over rows, of each row’s leverage times its own covariance — a leverage-weighted average. In this world at the leverage-weighted covariance is 1.583 times the plain one, and at it is 2.217 times. The likelihood’s correction removes the plain average’s worth and leaves the rest, and the estimator that was consistent with many instruments under equal variances is no longer consistent under these.
How far each estimator errs
With equal variances the cells behave like the earlier essay’s normal instruments. Two-stage least squares errs by a median 0.286, the likelihood by 0.049, the jackknife likelihood by 0.066 and the jackknife instrumental estimator by 0.146.
At the likelihood’s median error is 0.534 — more than half the effect, and further from the truth than two-stage least squares’ 0.387. The estimator chosen because it was the least biased is now the most biased of the four. At it errs by 0.830 against two-stage least squares’ 0.593. Its correction is computed as though every row carried the average covariance, and on these draws that leaves it on the least-squares side by more than the estimator it was meant to correct.
The jackknife likelihood errs by 0.066 at , the same as with equal variances. It is the likelihood with every quadratic form in the projection taken with its diagonal removed — rather than — so no row’s own leverage ever reaches its own fitted value, and the leverage-weighted covariance has nothing to act on. At , with a weak first stage, it too drifts, to 0.195; there the heteroskedasticity also inflates the confounding’s total variance and the first stage is too weak to separate it.
The jackknife instrumental estimator, which removes the diagonal from two-stage least squares rather than from the likelihood, errs by 0.325 at . Its trouble is not the heteroskedasticity but the weakness, which the earlier essays found: at a concentration of 8 its denominator is barely larger than its own noise.
The medians hide how spread the estimates are, and at a concentration of 8 they are very spread. The likelihood’s estimate misses by more than the whole effect on 25.0% of draws with equal variances and 28.4% at ; the jackknife likelihood’s on 22.2% and 24.5%. Neither is a precise estimator at this strength. The difference between them is where the spread is centred: the likelihood’s centre has moved half the effect towards least squares, and the jackknife likelihood’s has not moved at all. A wide interval around the right centre is a weak-instrument problem, which what the first stage does not know showed an honest interval can at least report. A narrow interval around a moved centre is a failure no width can report.
How often each interval covers
Bekker’s interval for the likelihood covers 94.2% with equal variances, 74.8% at and 29.6% at . It is an interval built for many instruments and one variance, centred on an estimate that is no longer consistent, and both halves fail together.
The jackknife likelihood’s interval uses a standard error built for many instruments and unequal variances — a sandwich whose middle sums each row’s squared error times its squared leave-one-out fitted value, plus the cross terms through the off-diagonal projection weights. It covers 92.4% with equal variances, 90.6% at and 85.0% at , at median widths of 2.06 to 2.20 against Bekker’s 1.09 to 2.20. With equal variances it gives up two points of coverage to Bekker’s, the price of estimating a variance it did not need to estimate; at it gains sixteen.
The jackknife instrumental estimator covers 98.0%, 93.8% and 75.5%, by being nearly twice as wide as anything else — 3.86 at — which is the price leaving each row out found at a concentration of 8. Two-stage least squares covers 52.9%, 21.4% and 1.7%.
At eight times the strength
With a first stage eight times as strong, every estimator’s weak-instrument trouble shrinks and what is left is the heteroskedasticity. Bekker’s interval covers 96.5% with equal variances, 87.3% at and 62.2% at ; the likelihood’s median error is −0.006, 0.080 and 0.260. The jackknife likelihood covers 95.6%, 94.8% and 93.4%, with median errors of −0.003, −0.002 and −0.000 — on the truth at every setting, at widths of 0.57 to 0.68 against Bekker’s 0.59 to 0.65.
The likelihood’s loss at this strength can be split into its two causes. Bekker’s interval at is 0.60 wide, so its standard error is about 0.154, and a centre moved by 0.080 is half a standard error from the truth. A correct standard error around a centre half a standard error off would cover about 92%; Bekker’s covers 87.3%. Roughly three of the eight lost points are the moved centre and the other five are the standard error itself, computed for one variance and applied to many. The jackknife likelihood repairs both at once, because its estimator removes the moved centre and its sandwich removes the second.
That is the clean version of the result. Once the first stage is strong enough for any of these estimators to work, the likelihood’s error grows with the heteroskedasticity and the jackknife likelihood’s does not, and its robust interval holds its level while Bekker’s loses a third of it. The jackknife instrumental estimator is also unbiased here, −0.016 at , but its conventional standard error does not know about the unequal variances, and it covers 94.0% and 85.2% at the two heteroskedastic settings.
Normal instruments do not show it
The same construction on the earlier essay’s thirty-two normal instruments — the confounding’s variance rising with the squared length of a row’s instrument vector, which is that row’s leverage up to a constant — barely moves the likelihood. Bekker’s interval covers 94.1%, 92.7% and 91.7% at , 1 and 2, and the likelihood’s median error goes from 0.079 to 0.122 and 0.165.
The difference is the spread of the leverages. Thirty-two independent normal instruments give every row a leverage close to ; the heteroskedasticity can only correlate with a quantity that barely varies, and the leverage-weighted covariance is barely different from the average. The ratio can be written down. On normal instruments a row’s leverage is close to its squared length over , and with the confounding’s variance going as that length over raised to , the leverage-weighted covariance over the plain one is for a chi-square on degrees of freedom divided by . That is at and at , against the cells’ 1.583 and 2.217. A sixteenth extra against nearly three fifths: the normal design leaves the likelihood a sixteenth of a bias to miss, and the cells leave it most of one.
Cells of two to fifteen give leverages that vary sevenfold, and the same heteroskedasticity, attached to them, produces a bias that grows with the number of small cells. The failure is not a property of many instruments or of unequal variances. It is a property of unequal variances attached to unequal leverage, which is the shape a design with small cells has by construction.
What the estimate is an estimate of
Every draw here gives every unit the same treatment effect, so a consistent estimator has one thing to be consistent for. Designs with cells are also the designs in which whose effect it is bites: if the effect differs across cells, an instrumental estimate is a weighted average of cell effects, with weights set by how strongly each cell moves the treatment, and the jackknife likelihood’s target is that weighted average too. Removing the own-row bias does not change which average is estimated. A study with cells of unequal size and an effect that varies with cell size has two questions to answer, and this essay has answered the one about bias, not the one about the estimand.
The exclusion restriction is untouched as well. Every cell’s indicator is a valid instrument here by construction; the assumption nothing tests still holds for each of them, and with thirty-two cells it has thirty-two chances to fail. The overidentification test that two instruments that disagree priced is the natural check, and its usual form assumes the one variance this world does not have; under unequal variances it needs a robust version of its own, which has not been counted here.
What a design with small cells should report
Check whether leverage varies. The diagonal of the first-stage projection is computable before any outcome is seen. If it is nearly constant, as with many continuous instruments of similar scale, the likelihood with Bekker’s interval is safe even under unequal variances. If it varies — cells of very different sizes, a few instruments with heavy tails — the likelihood’s correction is calibrated for an average that the bias does not use.
Look at the residual variance against leverage. The covariance that matters is the one between the two errors, which a study cannot see, but the structural residual’s variance by cell is visible, and a trend in it across cell sizes is the warning sign this world was built from.
Use the jackknife likelihood with the robust standard error when cells are small. It costs two points of coverage and a little width when variances are equal and gains sixteen points at with a weak first stage, and holds 93–96% at every setting once the first stage is moderate. Weak and back where it started showed what a weak first stage does to every estimator; the jackknife likelihood is the one that does not add a second failure to it.
Do not read the likelihood’s small bias under equal variances as a property it keeps. The ordering of the earlier essays — the likelihood least biased, two-stage least squares most — reversed here with nothing changed but which cells were noisy.
Proved, computed and counted
Proved: with cell indicators as instruments, a row’s first-stage leverage is one over its cell’s size; the bias of two-stage least squares from a row’s own leverage is the leverage-weighted covariance between the errors, which the likelihood’s κ removes only as the plain average. Computed: the ratio of the two averages in this design, 1.583 at and 2.217 at .
Counted, over a thousand draws of two hundred rows at each strength and setting: every median error, coverage and width above, the same draws for all estimators, and the normal-instrument comparison on that world’s own draws.
Not claimed: that small cells are always the noisy ones. The direction matters — confounding that is smaller in small cells would make the leverage-weighted covariance smaller than the average, and the likelihood would over-correct instead — and it is a property of the application. Nor that the robust interval is exact: at a concentration of 8 it covers 90.6% and 85.0% under heteroskedasticity, and a weak first stage still costs it.
Still open: cells too small to leave out
A cell of two units gives each of its rows a leverage of one half, and a jackknife that removes a row’s own contribution leaves that row’s fitted value resting on a single other unit. A cell of one — a judge who heard one case — gives leverage one, and the leave-one-out fitted value does not exist. Designs with many instruments routinely contain such cells, and the usual practice is to drop them or pool them into a residual category.
Whether dropping singleton cells biases the jackknife likelihood when the dropped cells are the noisy ones, how the robust interval behaves as the smallest cell shrinks to two and then one, and whether pooling the smallest cells into one recovers the information without bringing back the leverage-weighted bias, are measurable in this same world and have not been measured here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Right for the wrong reason — both name coverage, heteroskedasticity, leave-one-out
- Robust is not free — both name coverage, heteroskedasticity, leave-one-out
- A ratio whose interval has to be the whole line — both name coverage, weak instrument
- An effective number of clusters — both name coverage, leverage
- Marginal is not conditional — both name coverage, heteroskedasticity
- The bread and the filling — both name heteroskedasticity, leverage
Named objects
A flat tag is an object no other essay names yet.
Concentration parameterCoverageHeteroskedasticityInstrumental-variableJackknife instrumental-variableLeave-one-outLeverageLimited-information maximum likelihoodTwo-stage least squaresWeak instrument