A standard error for a model that is wrong

Robust is not free

A robust standard error's promise is asymptotic and its use is not. Its 95% interval covers 88.73% at twenty rows, and under mild heteroskedasticity it is the worse of the two intervals until a hundred.

Worth reading first: What a wrong model estimates.

Under an error variance leaning hard towards the edges of the design, a robust 95% interval for a slope covers 88.73% at twenty rows. The model-based interval covers 87.06% there, and unlike the robust one it does not improve with the sample: it is still at 87.40% at a thousand rows, because it is a precise interval for a quantity it is not estimating.

So the robust error wins, and the margin at twenty rows is under two points. That is the number this essay is about. The correction whose whole justification is that the model-based error is estimating the wrong variance recovers, at the sample size where anybody actually worries, about a quarter of the gap it exists to close. Everything else it promises is delivered at a thousand rows, where 94.93% against 87.40% is exactly the picture the textbook draws.

Two facts turn that from a caveat into a decision. The first is that most of the small-sample shortfall is repairable and it is not repaired by default. The second is that under a mild departure the ordering reverses: at twenty rows with the variance leaning gently, the robust interval covers 89.63% against the model-based interval’s 93.52%, so the robust error is the worse of the two at the sample sizes it is most often reached for, under exactly the departure it exists for.

The promise is asymptotic; the use is not

Six sample sizes, an error variance leaning towards the edges of the design, and twenty thousand draws at the small end. HC0 — the uncorrected sandwich — covers 88.73% at twenty rows, 91.16% at thirty, 92.70% at fifty, 93.73% at a hundred and 94.93% at a thousand. Every one of those is a count, and the binomial standard error on the first is 0.224 points, so the shortfall at twenty rows is twenty-eight standard errors from nominal rather than a fluctuation.

The model-based line is flat: 87.06%, 87.37%, 87.43%, 86.98%, 86.55%, 87.40%. That flatness is the signature of the defect the field’s arithmetic computes — the estimator is converging perfectly well on a number that is 0.6081 of the variance the slope has, so more data buys precision about the wrong quantity and nothing else. A coverage that does not improve with the sample is not a small-sample problem, and no correction to the reference distribution will touch it.

The two lines therefore have different shapes as well as different levels, and reading them together is what the picture is for. One is asymptotically right and finite-sample poor; the other is finite-sample stable and asymptotically wrong. Choosing between them at a given sample size is a real decision rather than an application of a rule.

Robust, at the sample sizes it is reached for. Counted coverage of five 95% intervals for a slope, at six sample sizes, under an error variance leaning towards the edges of the design (γ = 0.8), over 20000 draws at the small end. The model-based interval sits at about 87.06% everywhere and does not improve with the sample, because it is a claim about a variance it is not estimating. The robust ones do improve: HC0 covers 88.73% at 20 rows, 92.70% at 50 and 94.93% at 1,000. Its promise is asymptotic and its use is not, and the gap between those two facts is this picture. The leave-one-out correction read against a t on n − 2 is the only line that is near its promise at the small end: 94.55% at 20 rows.
Fig. 1 Counted coverage of five 95% intervals across six sample sizes, under an error variance leaning towards the edges of the design. HC0 climbs from 88.73% to 94.93%; the model-based interval sits at about 87% throughout.

Two defects, and only one of them is the estimate

A confidence interval is a variance estimate and a quantile, and it is worth asking which of the two is failing.

The same HC0 variance read against a tt on n2n - 2 degrees of freedom rather than against a normal covers 90.68% at twenty rows instead of 88.73%. That is 1.95 points bought by changing a multiplier from 1.959964 to 2.1009 — a number that costs nothing to change and that every regression package already computes for the model-based interval sitting next to it. Roughly two of the six missing points, at the sample size where the miss is largest, are the reference distribution rather than the estimate.

The rest of the repair is in the estimator. HC3, the leave-one-out weighting, covers 92.91% against a normal and 94.55% against a tt on n2n-2 — the only interval on the table near its promise at twenty rows, and it is within half a point of nominal there. The four robust corrections and two references make eight intervals, and the ordering is stable across the whole sweep: the dearer the correction and the fatter the reference, the better the small-sample coverage.

Half the shortfall is the quantile, not the estimate. What each 95% interval covers at 20 rows over 20000 draws, read against two references. The uncorrected robust interval covers 88.73% against a normal quantile and 90.68% against a t on n − 2 degrees of freedom — the same variance estimate, 1.95 points apart, bought by changing a number that costs nothing to change. The leave-one-out correction with a t reference reaches 94.55%. So a robust interval's small-sample shortfall is two defects rather than one: a variance estimate that is short, and a reference distribution that is wrong for the same reason the model-based one needs a t.
Fig. 2 What each 95% interval covers at twenty rows, read against a normal quantile and against a t on n − 2. HC0 covers 88.73% and 90.68% — the same variance estimate, 1.95 points apart — and HC3 with a t reference reaches 94.55%.

The justification for a tt reference here is not the usual one. The tt distribution exists because s2s^2 is σ2χn22/(n2)\sigma^2\chi^2_{n-2}/(n-2), and that derivation says nothing whatever about a sandwich, whose variance estimate is not a scaled chi-square and has no exact distribution. The essay on why the normal is not good enough when the spread is estimated gives the argument in the case where it is exact. Using tt on n2n-2 for a robust interval is a heuristic, and its justification here is the count in the picture rather than a theorem. It is worth being explicit about that, because the number of degrees of freedom to use is genuinely unsettled and the sweep does not settle it — it shows only that n2n-2 is better than infinity at every cell of this table.

What the estimate recovers, and what it never will

Coverage mixes the variance estimate’s bias with its noise and with the reference. The bias on its own is readable directly, as each estimate’s average divided by the variance the slope actually has across the same draws.

HC0 reads 0.8151 at twenty rows, 0.8936 at thirty, 0.9335 at fifty, 0.9493 at a hundred and 0.9935 at a thousand. That climb is the whole of the small-sample defect and the whole of the asymptotic promise, on one axis. HC3 approaches from the other side — 1.1373 at twenty rows, falling to 0.9996 at a thousand — which is why its interval is the widest on the table and why its coverage is the best at the small end. The model-based estimate reads 0.5909 at twenty and 0.6078 at a thousand: a flat line at three-fifths of the truth.

Two things follow. The 0.8151 is not noise, it is structure, and the field’s exact expectations compute it from the hat matrix without a simulation: under a constant error variance HC0’s expectation is 11/nui4/(ui2)21 - 1/n - \sum u_i^4/(\sum u_i^2)^2, which at twenty rows on this design is 0.8603 before the variance shape does anything. And the corrections are not competing to be unbiased — HC2 is exactly unbiased under a constant error variance and covers 90.88% here, worse than HC3’s 92.91%, because coverage at a small sample is bought by over-statement rather than by being right on average.

What the estimate recovers, and what it never will. Each variance estimate's average across 20000 draws, divided by the variance the slope actually has, at six sample sizes with the error variance leaning towards the edges of the design (γ = 0.8). The model-based estimate sits at about 0.6078 at every size: it is not converging on anything, because it is estimating a different quantity. The uncorrected robust estimate climbs from 0.8151 at 20 rows to 0.9935 at 1,000, which is the whole of its small-sample defect and the whole of its asymptotic promise on one axis. The leave-one-out correction overshoots at the small end — 1.1373 at 20 rows — which is what buys its coverage there and what makes its interval the widest on the table.
Fig. 3 Each variance estimate’s average divided by the variance the slope has, at six sample sizes. HC0 climbs from 0.8151 to 0.9935; HC3 falls from 1.1373 to 0.9996; the model-based estimate sits at about 0.60 throughout.

The crossing is the number nobody quotes

The comparison so far has held the departure fixed and strong. The interesting quantity is the sample size at which the robust interval is nearer 95% than the model-based one — read as nearer, since both can miss in either direction — and it depends on how strong the departure is.

At the strong setting it is twenty: the robust interval wins at the smallest size on the sweep. Halve the departure and it is fifty. Halve it again and it is a hundred. Those three numbers are the practical content of this field, and they are the ones that never appear beside the recommendation. Under the mild setting at twenty rows the model-based interval covers 93.52% and the robust one 89.63% — the default is nearer its promise by nearly four points, and it is nearer at thirty and at fifty as well.

The mechanism is a race between two errors that shrink at different rates. The model-based interval carries a fixed shortfall proportional to how far 1+γ(κ1)1 + \gamma(\kappa - 1) sits from one, and at a mild γ\gamma that shortfall is small. The robust interval carries a shortfall that goes to nought like one over the sample size but starts large. A small fixed error beats a large shrinking one until the sample is big enough, and the crossing is where the two curves meet.

Robust, at the sample sizes it is reached for. Counted coverage of five 95% intervals for a slope, at six sample sizes, under an error variance leaning towards the edges of the design (γ = 0.3), over 20000 draws at the small end. The model-based interval sits at about 92.09% everywhere and does not improve with the sample, because it is a claim about a variance it is not estimating. The robust ones do improve: HC0 covers 89.30% at 20 rows, 93.09% at 50 and 94.60% at 1,000. Its promise is asymptotic and its use is not, and the gap between those two facts is this picture. The leave-one-out correction read against a t on n − 2 is the only line that is near its promise at the small end: 95.05% at 20 rows.
Fig. 4 The same five intervals under a variance leaning half as hard. The model-based interval now covers 92.09% at twenty rows against HC0’s 89.30%, and does not lose its advantage until fifty.

There is a circularity in using it, and naming it is more useful than pretending otherwise. To know which side of the crossing a study is on, an analyst needs to know how strong the departure is — which is the quantity the robust error was reached for because it was unknown. So the crossing cannot be a decision rule on its own. What it can do is bound the damage in the direction nobody checks: a study of twenty or thirty rows that switches to a robust error is very likely making its interval worse unless the heteroskedasticity is severe, and the severity that would justify the switch is severe enough that the corresponding correction is large and visible. A robust error that barely moves the number at twenty rows is evidence that switching was not worth its own small-sample cost.

What the sweep does not do is say what function of the departure the crossing is. Both halves are available in closed form — the model-based interval’s fixed shortfall is a coverage under a stated variance ratio, and the robust estimate’s bias has an exact expectation — so the crossing is a solvable equation and nobody here solved it. It is reported as three measured points and not as a rule, which is the honest state of it and is also a whole argument somebody could make.

What the correction costs in width

Coverage is only half a comparison; an interval that covers by being enormous has not helped. At twenty rows under the strong departure the mean widths are 1.5706 for the model-based interval, 1.6984 for HC0 against a normal, 2.0022 for HC3 against a normal and 2.1462 for HC3 against a tt.

So the interval that reaches 94.55% is 37% wider than the one that covers 87.06%, and a reader who wanted to believe the two were comparable should stop. This is not a case of one procedure dominating another: the model-based interval is shorter and misses more often, in the exact proportion that makes both statements true simultaneously. The choice is between a stated coverage and a stated width, which is the trade the essay on what studentising costs prices for a resampled interval and reaches the same shape of answer.

The width also explains why the reference correction is nearly free and the estimator correction is not. Going from a normal to a tt at twenty rows multiplies every interval by 2.1009/1.959964, about 7%, and buys 1.95 points of coverage on HC0. Going from HC0 to HC3 multiplies by about 18% and buys 4.18 points. Both are worth having and only one of them is a rounding error.

Half the shortfall is the quantile, not the estimate. What each 95% interval covers at 50 rows over 20000 draws, read against two references. The uncorrected robust interval covers 92.70% against a normal quantile and 93.44% against a t on n − 2 degrees of freedom — the same variance estimate, 0.74 points apart, bought by changing a number that costs nothing to change. The leave-one-out correction with a t reference reaches 94.89%. So a robust interval's small-sample shortfall is two defects rather than one: a variance estimate that is short, and a reference distribution that is wrong for the same reason the model-based one needs a t.
Fig. 5 The same eight intervals at fifty rows. HC0 now covers 92.70% against a normal and 93.44% against a t — the reference is worth 0.74 points here against 1.95 at twenty rows, because the two quantiles have converged.

An unbiased variance estimate does not give a calibrated interval

One row of the table is worth isolating, because it breaks the reflex that says the right estimator is the unbiased one. HC2’s weighting 1/(1hii)1/(1 - h_{ii}) is constructed so that its expectation is exactly the slope’s variance under a constant error variance, and at twenty rows it duly reads 0.9621 of the truth here — much the closest of the four to one. Its interval covers 90.88%. HC3, which reads 1.1373 and is therefore wrong by a seventh in the direction of caution, covers 92.91%.

The estimate that is right on average gives the worse interval, and the reason is that an interval is not a linear function of a variance. It takes a square root, which is concave, so an unbiased variance is already a slightly short standard error; and it is a tail statement, so what matters is not the average of the estimate but how often it lands low. A variance estimate with a coefficient of variation above a quarter — HC2’s is 0.2615 at eighty rows, against the model-based estimate’s 0.1976 — lands low often enough that an interval built from it misses more than its nominal rate even when the estimate is centred correctly. Over-statement is what buys coverage back, and HC3’s weighting is over-statement by construction rather than by accident.

That is why unbiased is the wrong criterion for choosing between the corrections, and why the choice cannot be made without a coverage count. It also explains a pattern that would otherwise look like a coincidence: the ordering of the eight intervals by coverage at twenty rows is exactly their ordering by width, and both are exactly their ordering by how much each correction inflates a high-leverage point’s squared residual.

What each error is a claim about, and what the claim comes out as. Each variance estimate's average over 20000 draws, divided by the variance the slope actually has across those same draws, at 80 rows with the error variance leaning towards the edges of the design (γ = 0.8). One is a standard error that is right. The model-based estimate reads 0.6081 of the spread, so its standard error is 77.98% of the one it should report; the four robust corrections read 0.9576, 0.9821, 0.9961, 1.0362. Two further routes agree with the count and share none of its arithmetic: n times the counted variance is 4.8905 against a population sandwich of 4.9200, and the counted ratio of the two standard errors is 1.2799 against a closed form of 1.2806.
Fig. 6 Each estimate’s average over twenty thousand draws of eighty rows, against the variance the slope actually has. HC2 reads 0.9961 and HC3 1.0362 — a spread of a twenty-fifth between the corrections, against the model-based estimate’s 0.6081.

The same reading gives the honest version of “use a robust standard error”. At eighty rows the four corrections read 0.9576, 0.9821, 0.9961 and 1.0362 of the truth, and the choice between them is worth a twenty-fifth of a variance. At twenty rows they read 0.8151, 0.9056, 0.9621 and 1.1373, and the choice is worth forty per cent. The advice is sample-size dependent in a way its usual statement is not, and the sample sizes where it matters are the ones the advice is most often given about.

What could have produced these counts without the claim being true

Simulation noise. Each cell reports its own binomial standard error, and the sweep spends its budget where the answer is interesting: twenty thousand draws at twenty, thirty and fifty rows, ten thousand at a hundred, four thousand at two hundred and fifty and fifteen hundred at a thousand. The standard error is 0.224 points at the small end and 0.566 at the large one, so the thousand-row cell is the least certain number on the table and the one carrying the least weight in the argument. The claims that matter — 88.73% against a promise of 95%, and 89.63% against 93.52% — are between four and twenty-eight standard errors wide.

The estimand could be moving. It is not: the mean model in this sweep is correct, a straight line fitted to a straight-line truth, so the target is the truth’s own slope and every miss is a variance failure rather than the projection shift a wrong mean produces. Mixing the two would make a coverage table uninterpretable, since an interval can miss because it is the wrong width or because it is centred somewhere else.

The design could be doing the work. The covariate is a fixed even grid at every sample size, so the leverage structure is as benign as it gets and the maximum leverage at twenty rows is 0.1857. On a design with one far-out point the same corrections separate by a factor of sixteen rather than by a fifth, which is the argument about which correction to use and is a different essay’s subject. The numbers here are therefore the best case for the four corrections agreeing.

The reference could be the whole story. It is not: HC0 on a tt still covers only 90.68% at twenty rows, four points short, and the remaining shortfall tracks the estimate’s own bias of 0.8151 exactly as a first-order argument says it should. Two defects, and the sweep separates them by holding one fixed while the other moves — the same discipline the essay that separates two resampling defects uses on a different pair.

Where the shortfall is not the point

Everything above prices a robust standard error where a robust standard error is needed. It says nothing about what the correction costs where the model-based error was right all along, which is a separate sweep and a separate result: under a constant error variance the leave-one-out interval read against a tt covers 95.52% at twenty rows against the model-based 95.06%, and the price is width and variability rather than coverage. The essay that prices insurance where the risk is absent is where that count lives, and it changes the reading of this one: the case against defaulting to robust errors is not that they cost coverage when the assumption holds, but that they cost coverage when it fails mildly and the sample is small.

And the third repair is not measured here. The small-sample shortfall has two components, the estimate and the reference, and both have been treated by assuming a better reference rather than by generating one. Resampling the reference distribution is the standard third answer, and the machinery for the version that keeps a non-constant error variance is already measured in the essay on which residual resampling preserves it — a multiplier that leaves each residual where it is. Whether that closes the gap between 88.73% and 95% at twenty rows better than HC3 and a tt close it to 94.55% is one sweep, and it is not run here. It would have been a seventh figure in an essay already carrying six, which is an honest reason and not a good one.

What is settled is narrower than the recommendation it is usually attached to. A robust standard error delivers its promise asymptotically, delivers about a quarter of it at twenty rows uncorrected, and delivers nearly all of it at twenty rows if the leave-one-out weighting and a tt reference are both used. The default in most software is the uncorrected version against a normal, which is the worst cell of that table and the one every quoted justification is written about.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Asymptotic varianceConfidence intervalCoverageFinite-sample correctionHeteroskedasticityHeteroskedasticity-consistentInterval widthLeave-one-outMonte CarloNominal levelReference distributionRobust standard errorSample sizeT intervalWald interval