A standard error for a model that is wrong

Three corrections and a leverage

On an even design of twenty rows the four robust corrections read 0.8603, 0.9559, 1.0000 and 1.1647 of the truth and the choice barely matters. Add one point at x = 8 and they read 0.3191, 0.3419, 1.0000 and 5.1127.

Worth reading first: What a wrong model estimates · The line that one point drew.

The four robust variance estimators differ in one number: what each multiplies a point’s squared residual by, given that point’s leverage. On an even design of twenty rows their exact expectations, as shares of the variance the slope actually has, are 0.8603, 0.9559, 1.0000 and 1.1647. The whole span is a fifth, and any statistical argument that turns on which of the four was used is an argument about nothing.

Replace one of those twenty rows with a point far outside the design — twenty-nine even points on the unit interval doubled and one at x=8x = 8 — and the same four expectations read 0.3191, 0.3419, 1.0000 and 5.1127. The span is a factor of 16.02, the uncorrected estimate is under a third of the truth and the leave-one-out one is five times it, and the choice between them decides whether a 95% interval is a third of its proper width or five times it.

Neither number is a simulation result. Under a constant error variance the expectation of every one of these estimators is closed, reads nothing but the covariates, and can be computed before a single response is looked at. Which correction to use is therefore a question about the design, answerable in advance, and the reason it is usually answered by a package default is that nobody computes the quantity that decides it.

The residual is short by exactly the leverage

The whole of the arithmetic comes from one exact fact about least squares. The residual vector is (IH)ε(I - H)\varepsilon, so under a constant error variance

E[ei2]=σ2(1hii)E[e_i^2] = \sigma^2(1 - h_{ii})

exactly, with hiih_{ii} the ii-th diagonal of the hat matrix. A residual is not an error, it is an error with part of itself projected away, and the part removed is larger where the point has more say in where the line goes. That is the statement the essay separating residuals from errors makes at length; here it is used as an identity.

Feed it through each estimator’s own weighting and every expectation follows without a random number:

E[V^HC0]V=11nui4(ui2)2,E[V^HC2]V=1,E[V^HC3]V=ui2/(1hii)ui2,\frac{E[\hat{V}_{\text{HC0}}]}{V} = 1 - \frac{1}{n} - \frac{\sum u_i^4}{\left(\sum u_i^2\right)^2}, \qquad \frac{E[\hat{V}_{\text{HC2}}]}{V} = 1, \qquad \frac{E[\hat{V}_{\text{HC3}}]}{V} = \frac{\sum u_i^2/(1 - h_{ii})}{\sum u_i^2},

with uiu_i the covariate less its mean, and HC1 the first of those times n/(n2)n/(n-2). On the even design of twenty rows the fourth-moment share is 0.089699 and 11/200.089699=1 - 1/20 - 0.089699 = 0.860301, which is the first table entry to six places. The maximum leverage there is 0.185714.

This is a trace computation, and it is worth saying which one it is not. The trace in the essay on what a selection penalty is estimating is tr(HΩ)\mathrm{tr}(H\Omega), the optimism a criterion charges for having fitted; the trace here is the exact expectation of a variance estimator under a stated design. The two read the same hat matrix and answer different questions, and conflating them would suggest that a penalty and a correction are the same repair.

What each correction is worth, exactly. Each variance estimate's expectation under a constant error variance, divided by the variance the slope actually has, at 20 rows on an even design, by two routes: the closed form E[eᵢ²] = σ²(1 − hᵢᵢ) carried through each correction's own weight, and the mean of 20000 counted estimates. The maximum leverage here is 0.1857 and the design's fourth-moment share Σu⁴/(Σu²)² is 0.0897, which is the only thing the closed form reads. HC0 comes out at 0.8603 — short by construction, since its factor is exactly 1 − 1/n − Σu⁴/(Σu²)². HC1 reaches 0.9559, HC2 is exactly 1.0000 at every design and every sample size, and HC3 overshoots to 1.1647. On an even design the four are within a fifth of each other and the choice barely matters.
Fig. 1 Each estimator’s expectation on an even design of twenty rows, in closed form and counted over twenty thousand draws. HC0 reads 0.8603 against a count of 0.8636, HC2 is exactly 1.0000, and HC3 reads 1.1647.

Four corrections, one curve

Everything that separates the four estimators is a single function of leverage, and the function reads only the covariates.

HC0 multiplies every squared residual by one, which is to say it does not correct at all. HC1 multiplies by n/(n2)n/(n-2)1.1111 at twenty rows — the same factor for every point, borrowed from the degrees-of-freedom adjustment that makes s2s^2 unbiased in the model-based estimate and applied here to an estimator with a different defect. HC2 multiplies by 1/(1hii)1/(1 - h_{ii}), which is exactly the reciprocal of the shortfall the identity above states and is therefore the unique weight making the expectation right. HC3 multiplies by 1/(1hii)21/(1 - h_{ii})^2, an approximation to what the residual would have been had the fit not seen that point.

At the largest leverage on an even design of twenty rows, 0.185714, those four are 1.0000, 1.1111, 1.2281 and 1.5082. At the far point’s leverage of 0.836320 they are 1.0000, 1.1111, 6.1095 and 37.3258. The first list is why the choice does not matter on a well-spread design and the second is why it decides everything on a design with a point outside it.

The four corrections differ in one number. What each of the four heteroskedasticity-consistent corrections multiplies a point's squared residual by, as a function of that point's leverage, at 20 rows. HC0 multiplies by one; HC1 by n/(n − 2) = 1.1111, the same for every point; HC2 by 1/(1 − h), which is exactly the factor that makes the expectation right under a constant error variance; HC3 by 1/(1 − h)², which is the leave-one-out approximation and is deliberately dearer. At an average point of an even design the four are 1.0000, 1.1111, 1.0526 and 1.1080 and the choice is nearly arithmetic. At a leverage of 0.86 they are 1.00, 1.11, 7.14 and 51.02. Every difference between the four estimators is this curve, and the curve reads only the covariates.
Fig. 2 What each correction multiplies a point’s squared residual by, as a function of that point’s leverage. HC0 and HC1 are flat; HC2 and HC3 climb, and at a leverage of 0.86 they stand at about 7 and 50.

Two properties of that picture are worth naming. The weights are known before any outcome is looked at, since leverage is a function of the covariates alone — so a reader can compute what a correction will do to a dataset without having the response, which is exactly what makes the choice answerable in advance. And HC1’s flat factor is aimed at the wrong quantity: it corrects a whole-sample average by a whole-sample constant, when the defect it is repairing is distributed across points in proportion to their leverage. At twenty rows on an even design that mismatch is invisible, because every leverage is near 1/n1/n; on a design with one dominant point HC1 corrects 0.3191 to 0.3419 and leaves the estimate two thirds short.

One point, and the corrections come apart

The far-point design is twenty-nine even points on the doubled unit interval with one observation at x=8x = 8. It is not exotic — a single subject with an unusual exposure, one country with a much larger population, one day with an outlying price — and it is the case the essay on the line one point drew is entirely about.

At thirty rows that point’s leverage is 0.836320: it alone accounts for five sixths of the slope. The design’s fourth-moment share is 0.647561, so HC0’s expectation is 11/300.647561=1 - 1/30 - 0.647561 = 0.319106. HC1 reaches 0.341899. HC2 is exactly 1.000000. HC3 reads 5.112664. Counted against twenty thousand draws, those four come back as 0.3184, 0.3411, 0.9935 and 5.0709 — agreement to within the sampling error at every cell, from two routes that share no arithmetic.

The span between the extremes is ×16.02, against ×1.217 on the even design of the same size. So the difference between the corrections is not a property of the corrections; it is a property of the design, amplified by them. A reader who wants a single number for how much the choice matters can compute the design’s maximum leverage and read it off the weight curve.

One far-out point, and the corrections separate. Each variance estimate's expectation under a constant error variance, divided by the variance the slope actually has, at 30 rows on a design of 29 even points with one placed at 8, by two routes: the closed form E[eᵢ²] = σ²(1 − hᵢᵢ) carried through each correction's own weight, and the mean of 20000 counted estimates. The maximum leverage here is 0.8363 and the design's fourth-moment share Σu⁴/(Σu²)² is 0.6476, which is the only thing the closed form reads. HC0 comes out at 0.3191 — short by construction, since its factor is exactly 1 − 1/n − Σu⁴/(Σu²)². HC1 reaches 0.3419, HC2 is exactly 1.0000 at every design and every sample size, and HC3 overshoots to 5.1127. A single observation has taken the uncorrected estimate to a third of the truth and the leave-one-out one to five times it.
Fig. 3 The same four expectations on twenty-nine even points plus one at x = 8, where the far point’s leverage is 0.8363. HC0 reads 0.3191 of the truth and HC3 reads 5.1127, a span of a factor of sixteen.

HC2 is exactly one on both designs, and at every sample size on the sweep. That is not a numerical coincidence to be verified; it is the construction, since 1/(1hii)1/(1 - h_{ii}) was chosen to cancel the (1hii)(1 - h_{ii}) the identity puts there. The sweep checks it anyway, at sixteen cells, and the check is worth its cost: an assertion that HC2’s expectation is one would fail immediately if the hat matrix were being computed for the wrong design, if the leverage were being read off the wrong diagonal, or if the sampler and the closed form disagreed about which points the design contains. It is the cell that cannot be right by accident.

The shortfall is the kurtosis over the sample size

The closed form for HC0 has two terms in it, 1/n1/n and the design’s fourth-moment share ui4/(ui2)2\sum u_i^4/(\sum u_i^2)^2, and the second turns out to be the first in disguise on any well-spread design.

Multiply that share by the sample size on the even grid and it reads 1.7940 at twenty rows, 1.7973 at thirty, 1.7990 at fifty, 1.7998 at a hundred and 1.8000 at four hundred. It is converging on the covariate’s own kurtosis, which for an even spread is exactly 1.8. So on a design whose points are spread the way a population is, ui4/(ui2)2κ/n\sum u_i^4/(\sum u_i^2)^2 \to \kappa/n and

E[V^HC0]V11+κn,\frac{E[\hat{V}_{\text{HC0}}]}{V} \approx 1 - \frac{1 + \kappa}{n},

which at κ=1.8\kappa = 1.8 is 12.8/n1 - 2.8/n: 0.860000 against the exact 0.860301 at twenty rows, and 0.993000 against 0.993000 at four hundred. The uncorrected sandwich is short by about three observations’ worth, and by how many is set by the covariate’s kurtosis — the same moment that the field’s ratio arithmetic says decides how far the robust error can sit from the model-based one. One design quantity governs both the size of the correction and the cost of not correcting it.

The far-point design is the case where that approximation fails completely, and the failure is diagnostic. The same product there reads 14.06, 19.43, 26.89 and 35.03 at the four sample sizes — growing rather than settling, because the share is not κ/n\kappa/n for any fixed κ\kappa. A single point carrying most of the design’s fourth moment makes the fourth-moment share an O(1)O(1) quantity, so HC0’s shortfall stops being a small-sample correction and becomes a permanent factor. That is the arithmetic behind the next section’s picture, and it says the right question is never “is the sample large” but “is ui4/(ui2)2\sum u_i^4/(\sum u_i^2)^2 small”.

The far point does not wash out

An obvious hope is that all of this is a small-sample story and that enough rows fix it. On the even design that hope is correct: HC0’s expectation climbs from 0.8603 at twenty rows to 0.9068 at thirty, 0.9440 at fifty and 0.9720 at a hundred, while HC3 falls from 1.1647 to 1.0289, and the four estimators converge on each other as the maximum leverage falls from 0.1857 to 0.0394.

On the far-point design it is wrong. At a hundred rows — ninety-nine even points and the one at x=8x = 8 — the far point’s leverage is still 0.5992, HC0’s expectation is 0.6397 and HC3’s is 1.8883. The span has fallen from sixteen to about three and it has not gone away, because the point’s leverage is not shrinking like 1/n1/n: it is shrinking towards the share of the design’s total spread that the point itself contributes, which stays large as long as the point stays far out. Adding rows near the middle of a design does not dilute a point outside it.

What each correction is worth, exactly. Each variance estimate's expectation under a constant error variance, divided by the variance the slope actually has, at 100 rows on an even design, by two routes: the closed form E[eᵢ²] = σ²(1 − hᵢᵢ) carried through each correction's own weight, and the mean of 20000 counted estimates. The maximum leverage here is 0.0394 and the design's fourth-moment share Σu⁴/(Σu²)² is 0.0180, which is the only thing the closed form reads. HC0 comes out at 0.9720 — short by construction, since its factor is exactly 1 − 1/n − Σu⁴/(Σu²)². HC1 reaches 0.9918, HC2 is exactly 1.0000 at every design and every sample size, and HC3 overshoots to 1.0289. On an even design the four are within a fifth of each other and the choice barely matters.
Fig. 4 The same four expectations at a hundred rows on the even design. The maximum leverage is 0.0394 and the four corrections have converged to within a twentieth of each other.

That is the practical rule this essay produces, and it is the opposite of the usual one. The usual advice is to use HC3 at small samples and stop worrying at large ones. The measurement says the thing to compute is the leverage, not the sample size: a thousand well-spread rows need no correction at all and a hundred rows with one point at leverage 0.6 need the same care as twenty.

One far-out point, and the corrections separate. Each variance estimate's expectation under a constant error variance, divided by the variance the slope actually has, at 100 rows on a design of 99 even points with one placed at 8, by two routes: the closed form E[eᵢ²] = σ²(1 − hᵢᵢ) carried through each correction's own weight, and the mean of 20000 counted estimates. The maximum leverage here is 0.5992 and the design's fourth-moment share Σu⁴/(Σu²)² is 0.3503, which is the only thing the closed form reads. HC0 comes out at 0.6397 — short by construction, since its factor is exactly 1 − 1/n − Σu⁴/(Σu²)². HC1 reaches 0.6528, HC2 is exactly 1.0000 at every design and every sample size, and HC3 overshoots to 1.8883. A single observation has taken the uncorrected estimate to a third of the truth and the leave-one-out one to five times it.
Fig. 5 The same at a hundred rows with the far point still present. Its leverage is 0.5992, HC0 still reads 0.6397 of the truth and HC3 1.8883, so a hundred rows have not dissolved a single observation’s influence.

What could have made these numbers without the claim being true

The closed form could be wrong and the count could be following it. It could not: the two share no code path. The expectations are sums over the hat matrix’s diagonal on a stated design; the counts are averages of twenty thousand variance estimates from simulated responses, computed by the same routine an analyst would use. They agree at all thirty-two cells across two designs and four sample sizes, and the residual disagreements — 0.8603 against 0.8636, 5.1127 against 5.0709 — are within the sampling error of the count.

The constant error variance could be doing the work. It is, and that is the scope rather than a flaw. Every expectation here is derived under σ2(xi)=σ2\sigma^2(x_i) = \sigma^2, which is the case in which the model-based estimate is already right and none of these corrections is needed for its own sake. That choice is what makes the comparison clean: it isolates each correction’s finite-sample behaviour from the heteroskedasticity it is designed for. Under a non-constant variance HC2 is no longer exactly unbiased, and the coverage sweep under a leaning variance reads it at 0.9621 rather than 1.0000 at twenty rows.

What the estimate recovers, and what it never will. Each variance estimate's average across 20000 draws, divided by the variance the slope actually has, at six sample sizes with the error variance leaning towards the edges of the design (γ = 0.8). The model-based estimate sits at about 0.6078 at every size: it is not converging on anything, because it is estimating a different quantity. The uncorrected robust estimate climbs from 0.8151 at 20 rows to 0.9935 at 1,000, which is the whole of its small-sample defect and the whole of its asymptotic promise on one axis. The leave-one-out correction overshoots at the small end — 1.1373 at 20 rows — which is what buys its coverage there and what makes its interval the widest on the table.
Fig. 6 The same estimators under an error variance leaning towards the edges of the design, where HC2’s exactness no longer holds: it reads 0.9621 at twenty rows and 0.9965 at a thousand.

The far point could be an outlier in the response rather than in the design. It is not, and the distinction is the whole of why this essay is about leverage rather than about outliers. The point at x=8x = 8 is generated by exactly the same model as the other twenty-nine, with the same error variance; nothing about its response is unusual. What is unusual is where it sits on the covariate, which is a property of the design a reader can see before collecting anything. An outlier in the response would be a different problem with a different repair, and the essay on four datasets sharing one summary contains both cases side by side.

The expectation could be the wrong summary. This is the real limitation and it is not resolvable inside this essay. An expectation says where an estimator sits on average, and the coverage measurement shows that the estimator with the right average is not the one with the right coverage: HC2 is exactly unbiased and covers 90.88% at twenty rows, while HC3 overshoots by a seventh and covers 92.91%. So the four exact numbers here rank the estimators by a criterion that is not the one a reader cares about. They are still the right thing to compute, because they say precisely how far apart the four can be on a given design — but “which is best” needs the count and not the closed form.

An expectation carried by one observation

The two routes agree least well on the far-point design, and the pattern of the disagreement is worth reading rather than absorbing into the sampling error. At thirty rows HC3’s exact expectation is 5.112664 and the count over twenty thousand draws is 5.0709; at twenty rows the exact value is 7.5461 and the count is 7.6505. Both gaps are around one per cent, and the even design’s cells agree to a few parts in a thousand on the same number of draws.

That is not a defect in either route; it is the estimator telling a reader what it has become. On a design whose maximum leverage is 0.84, the sum ui2ei2/(1hii)2\sum u_i^2 e_i^2/(1 - h_{ii})^2 is dominated by the single term belonging to the far point, because that point supplies most of ui2u_i^2 and the largest weight. A sum dominated by one term is an estimate with about one degree of freedom, whose own coefficient of variation is near 2\sqrt{2}, so twenty thousand draws resolve its mean about as well as a few hundred draws resolve a well-conditioned one.

The consequence for an analyst is sharper than the arithmetic suggests. On such a design the robust variance estimate is right on average and nearly uninformative in any single sample, because the average is over a distribution whose spread is comparable to its mean. An estimator can be unbiased, exactly, and still be the wrong thing to hand somebody holding one dataset — which is the same distinction the coverage measurement draws between HC2’s exactness and its 90.88%, arrived at from a design where no single point dominates. Here it arrives from the design instead, and it is much larger.

It also explains why nothing in this field recommends a correction on the strength of its expectation alone. The exact values rank the estimators and bound how far apart they can be; what they cannot do is say which produces a usable interval, and on a design carrying a leverage of 0.84 the honest answer may be that none of them does. That is a statement about the configuration rather than about the arithmetic, and it is the sort of thing a plot of the residuals is calibrated to show or to miss long before a variance estimate is chosen.

Where the arithmetic stops

The weight 1/(1hii)21/(1 - h_{ii})^2 has a singularity, and a design can approach it. At twenty rows with the far point present its leverage is 0.8865 and HC3’s expectation is 7.5461; push the point further out and the leverage approaches one, the weight approaches infinity, and the estimator’s expectation does too. That is not a defect of the formula so much as an accurate report: at a leverage of one the fit passes exactly through that point, its residual is identically nought, and the data contains no information at all about the error at that position. An estimator that returns an enormous variance there is telling the truth loudly.

What is not measured here is the sequence a practitioner actually faces, which is deciding whether such a point belongs in the analysis at all. Every number in this essay is conditional on the design as given. A design with a leverage of 0.84 is one where a single observation has been given five sixths of the say, and the correction to use is a smaller question than whether to fit a line through that configuration — the essay on what one point can do to a slope is the argument that it usually is not. The corrections repair a variance estimate; nothing repairs a design.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formDegrees of freedomEstimated varianceFinite-sample correctionHat matrixHeteroskedasticity-consistentLeast squaresLeave-one-outLeverageMonte CarloOutliersResidualRobust standard errorSandwich estimatorStudentised residual