Concept

Nominal level — where it appears

The error rate or coverage a procedure claims, as distinct from the one it delivers. Every claim of this kind here is counted rather than accepted, and a gap in either direction is a defect: an interval that covers too often is buying its coverage with width.

Named by 5 essays across 3 fields — each of them below, with the objects they name alongside it.

Robust, at the sample sizes it is reached for. Counted coverage of five 95% intervals for a slope, at six sample sizes, under an error variance leaning towards the edges of the design (γ = 0.8), over 20000 draws at the small end. The model-based interval sits at about 87.06% everywhere and does not improve with the sample, because it is a claim about a variance it is not estimating. The robust ones do improve: HC0 covers 88.73% at 20 rows, 92.70% at 50 and 94.93% at 1,000. Its promise is asymptotic and its use is not, and the gap between those two facts is this picture. The leave-one-out correction read against a t on n − 2 is the only line that is near its promise at the small end: 94.55% at 20 rows.

Robust is not free

A robust standard error's promise is asymptotic and its use is not. Its 95% interval covers 88.73% at twenty rows, and under mild heteroskedasticity it is the worse of the two intervals until a hundred.

sandwich · Misspecification
The interval that over-covers when the instrument fails. Counted coverage of two nominal 95.0% intervals for the same causal effect, read off the same 2000 draws of 200 rows at each first stage. The exact Anderson–Rubin set covers 95.3% at every setting — flat, because the statistic it inverts is built from y − tβ, which contains no π at all, and is therefore the same number on the same draw whatever the instrument is worth. The conventional interval covers 99.1% at π = 0.02 and 95.6% at π = 0.6: it goes wrong at the weak end by covering too MUCH, at a median width of 7.320, because its standard error is computed from residuals taken at an estimate that has itself gone wrong. A weak instrument does not make this interval lie about its coverage; it makes it useless while telling the truth.

What the first stage does not know

A single weak instrument does not make the conventional interval undercover — it makes it cover 99.1% at a width of 7.320. Where the promise actually breaks is many instruments — coverage falls from 97.2% to 51.5% while the median width falls from 1.454 to 0.583.

instrument · Exclusion
The price of insurance is noise, not coverage. What a robust standard error costs under a constant error variance — the case where the model-based one is exactly right — at five sample sizes over 20000 draws at the small end. It is not coverage: the leave-one-out interval read against a t on n − 2 covers 95.52% at 20 rows against the model-based 95.06%. It is a wider interval, by a factor of 1.0689 at 20 rows falling to 1.0049 at 250, and it is a variance estimate 2.57 times as variable at 20 rows and 1.85 times at 250 — against a denominator that is exactly V²·2/(n − 2), so only the numerator is counted. The uncorrected estimate is the cheaper of the two at small samples and the dearer at large: 1.31 against 1.76.

Right for the wrong reason

A robust standard error costs no coverage where the risk is absent — 95.52% against 95.06% at twenty rows. It costs a 6.89% wider interval and a variance estimate 2.572 times as variable, and the pre-test that would avoid paying recovers 15.9% of what the insurance is worth.

sandwich · Misspecification
One estimator, three answers, and only the reference changes. Coverage of the cluster-robust 95% interval for the slope against the number of clusters, at 30 rows in each. The estimator is identical in all three curves; what differs is the number it is compared against. At 5 clusters it covers 75.05% against a normal, 85.30% against a t on 4 degrees of freedom and 87.95% against a t on 3. At 80 clusters the three agree to within a point. The correction costs nothing: the same standard error, a different table.

The reference the sandwich is read against

The cluster-robust interval covers 75.05% at five clusters and 93.58% at eighty. The same estimate read against a t on G − 2 covers 87.95% at five, and the estimator is unchanged — three hundred rows grouped into five clusters cover 74.28% where the same three hundred grouped into seventy-five cover 94.63%.

sandwich · Misspecification
Three promises, and no procedure keeps all three. Average coverage and worst-case coverage for four 95% intervals for a proportion at n = 40, computed exactly. Their expected widths are 0.2418, 0.2417, 0.2472, 0.2641 in the same order. The textbook interval and the score interval have the same expected width to four digits — 0.2418 and 0.2417 — and worst-case coverages of 55.31% and 92.21%. The exact interval never breaks its promise and is 9.3% wider than the score interval to do it. Each of the three columns orders the four procedures differently.

An interval that covers and says nothing

A procedure returning the whole line 95% of the time and the empty set otherwise has coverage exactly 95% at every parameter value. Two real intervals at forty observations have expected widths of 0.2418 and 0.2417 and worst-case coverages of 55.31% and 92.21%.

intervals · Coverage

Named alongside it

The objects these essays reach for when they reach for this one.

CoverageInterval widthConfidence intervalConservative intervalEstimated varianceFinite-sample correctionHeteroskedasticityLeave-one-outMonte CarloOvercoverageReference distributionRobust standard error

All concepts