Robust is not free
Worth reading first: What a wrong model estimates.
Under an error variance leaning hard towards the edges of the design, a robust 95% interval for a slope covers 88.73% at twenty rows. The model-based interval covers 87.06% there, and unlike the robust one it does not improve with the sample: it is still at 87.40% at a thousand rows, because it is a precise interval for a quantity it is not estimating.
So the robust error wins, and the margin at twenty rows is under two points. That is the number this essay is about. The correction whose whole justification is that the model-based error is estimating the wrong variance recovers, at the sample size where anybody actually worries, about a quarter of the gap it exists to close. Everything else it promises is delivered at a thousand rows, where 94.93% against 87.40% is exactly the picture the textbook draws.
Two facts turn that from a caveat into a decision. The first is that most of the small-sample shortfall is repairable and it is not repaired by default. The second is that under a mild departure the ordering reverses: at twenty rows with the variance leaning gently, the robust interval covers 89.63% against the model-based interval’s 93.52%, so the robust error is the worse of the two at the sample sizes it is most often reached for, under exactly the departure it exists for.
The promise is asymptotic; the use is not
Six sample sizes, an error variance leaning towards the edges of the design, and twenty thousand draws at the small end. HC0 — the uncorrected sandwich — covers 88.73% at twenty rows, 91.16% at thirty, 92.70% at fifty, 93.73% at a hundred and 94.93% at a thousand. Every one of those is a count, and the binomial standard error on the first is 0.224 points, so the shortfall at twenty rows is twenty-eight standard errors from nominal rather than a fluctuation.
The model-based line is flat: 87.06%, 87.37%, 87.43%, 86.98%, 86.55%, 87.40%. That flatness is the signature of the defect the field’s arithmetic computes — the estimator is converging perfectly well on a number that is 0.6081 of the variance the slope has, so more data buys precision about the wrong quantity and nothing else. A coverage that does not improve with the sample is not a small-sample problem, and no correction to the reference distribution will touch it.
The two lines therefore have different shapes as well as different levels, and reading them together is what the picture is for. One is asymptotically right and finite-sample poor; the other is finite-sample stable and asymptotically wrong. Choosing between them at a given sample size is a real decision rather than an application of a rule.
Two defects, and only one of them is the estimate
A confidence interval is a variance estimate and a quantile, and it is worth asking which of the two is failing.
The same HC0 variance read against a on degrees of freedom rather than against a normal covers 90.68% at twenty rows instead of 88.73%. That is 1.95 points bought by changing a multiplier from 1.959964 to 2.1009 — a number that costs nothing to change and that every regression package already computes for the model-based interval sitting next to it. Roughly two of the six missing points, at the sample size where the miss is largest, are the reference distribution rather than the estimate.
The rest of the repair is in the estimator. HC3, the leave-one-out weighting, covers 92.91% against a normal and 94.55% against a on — the only interval on the table near its promise at twenty rows, and it is within half a point of nominal there. The four robust corrections and two references make eight intervals, and the ordering is stable across the whole sweep: the dearer the correction and the fatter the reference, the better the small-sample coverage.
The justification for a reference here is not the usual one. The distribution exists because is , and that derivation says nothing whatever about a sandwich, whose variance estimate is not a scaled chi-square and has no exact distribution. The essay on why the normal is not good enough when the spread is estimated gives the argument in the case where it is exact. Using on for a robust interval is a heuristic, and its justification here is the count in the picture rather than a theorem. It is worth being explicit about that, because the number of degrees of freedom to use is genuinely unsettled and the sweep does not settle it — it shows only that is better than infinity at every cell of this table.
What the estimate recovers, and what it never will
Coverage mixes the variance estimate’s bias with its noise and with the reference. The bias on its own is readable directly, as each estimate’s average divided by the variance the slope actually has across the same draws.
HC0 reads 0.8151 at twenty rows, 0.8936 at thirty, 0.9335 at fifty, 0.9493 at a hundred and 0.9935 at a thousand. That climb is the whole of the small-sample defect and the whole of the asymptotic promise, on one axis. HC3 approaches from the other side — 1.1373 at twenty rows, falling to 0.9996 at a thousand — which is why its interval is the widest on the table and why its coverage is the best at the small end. The model-based estimate reads 0.5909 at twenty and 0.6078 at a thousand: a flat line at three-fifths of the truth.
Two things follow. The 0.8151 is not noise, it is structure, and the field’s exact expectations compute it from the hat matrix without a simulation: under a constant error variance HC0’s expectation is , which at twenty rows on this design is 0.8603 before the variance shape does anything. And the corrections are not competing to be unbiased — HC2 is exactly unbiased under a constant error variance and covers 90.88% here, worse than HC3’s 92.91%, because coverage at a small sample is bought by over-statement rather than by being right on average.
The crossing is the number nobody quotes
The comparison so far has held the departure fixed and strong. The interesting quantity is the sample size at which the robust interval is nearer 95% than the model-based one — read as nearer, since both can miss in either direction — and it depends on how strong the departure is.
At the strong setting it is twenty: the robust interval wins at the smallest size on the sweep. Halve the departure and it is fifty. Halve it again and it is a hundred. Those three numbers are the practical content of this field, and they are the ones that never appear beside the recommendation. Under the mild setting at twenty rows the model-based interval covers 93.52% and the robust one 89.63% — the default is nearer its promise by nearly four points, and it is nearer at thirty and at fifty as well.
The mechanism is a race between two errors that shrink at different rates. The model-based interval carries a fixed shortfall proportional to how far sits from one, and at a mild that shortfall is small. The robust interval carries a shortfall that goes to nought like one over the sample size but starts large. A small fixed error beats a large shrinking one until the sample is big enough, and the crossing is where the two curves meet.
There is a circularity in using it, and naming it is more useful than pretending otherwise. To know which side of the crossing a study is on, an analyst needs to know how strong the departure is — which is the quantity the robust error was reached for because it was unknown. So the crossing cannot be a decision rule on its own. What it can do is bound the damage in the direction nobody checks: a study of twenty or thirty rows that switches to a robust error is very likely making its interval worse unless the heteroskedasticity is severe, and the severity that would justify the switch is severe enough that the corresponding correction is large and visible. A robust error that barely moves the number at twenty rows is evidence that switching was not worth its own small-sample cost.
What the sweep does not do is say what function of the departure the crossing is. Both halves are available in closed form — the model-based interval’s fixed shortfall is a coverage under a stated variance ratio, and the robust estimate’s bias has an exact expectation — so the crossing is a solvable equation and nobody here solved it. It is reported as three measured points and not as a rule, which is the honest state of it and is also a whole argument somebody could make.
What the correction costs in width
Coverage is only half a comparison; an interval that covers by being enormous has not helped. At twenty rows under the strong departure the mean widths are 1.5706 for the model-based interval, 1.6984 for HC0 against a normal, 2.0022 for HC3 against a normal and 2.1462 for HC3 against a .
So the interval that reaches 94.55% is 37% wider than the one that covers 87.06%, and a reader who wanted to believe the two were comparable should stop. This is not a case of one procedure dominating another: the model-based interval is shorter and misses more often, in the exact proportion that makes both statements true simultaneously. The choice is between a stated coverage and a stated width, which is the trade the essay on what studentising costs prices for a resampled interval and reaches the same shape of answer.
The width also explains why the reference correction is nearly free and the estimator correction is not. Going from a normal to a at twenty rows multiplies every interval by 2.1009/1.959964, about 7%, and buys 1.95 points of coverage on HC0. Going from HC0 to HC3 multiplies by about 18% and buys 4.18 points. Both are worth having and only one of them is a rounding error.
An unbiased variance estimate does not give a calibrated interval
One row of the table is worth isolating, because it breaks the reflex that says the right estimator is the unbiased one. HC2’s weighting is constructed so that its expectation is exactly the slope’s variance under a constant error variance, and at twenty rows it duly reads 0.9621 of the truth here — much the closest of the four to one. Its interval covers 90.88%. HC3, which reads 1.1373 and is therefore wrong by a seventh in the direction of caution, covers 92.91%.
The estimate that is right on average gives the worse interval, and the reason is that an interval is not a linear function of a variance. It takes a square root, which is concave, so an unbiased variance is already a slightly short standard error; and it is a tail statement, so what matters is not the average of the estimate but how often it lands low. A variance estimate with a coefficient of variation above a quarter — HC2’s is 0.2615 at eighty rows, against the model-based estimate’s 0.1976 — lands low often enough that an interval built from it misses more than its nominal rate even when the estimate is centred correctly. Over-statement is what buys coverage back, and HC3’s weighting is over-statement by construction rather than by accident.
That is why unbiased is the wrong criterion for choosing between the corrections, and why the choice cannot be made without a coverage count. It also explains a pattern that would otherwise look like a coincidence: the ordering of the eight intervals by coverage at twenty rows is exactly their ordering by width, and both are exactly their ordering by how much each correction inflates a high-leverage point’s squared residual.
The same reading gives the honest version of “use a robust standard error”. At eighty rows the four corrections read 0.9576, 0.9821, 0.9961 and 1.0362 of the truth, and the choice between them is worth a twenty-fifth of a variance. At twenty rows they read 0.8151, 0.9056, 0.9621 and 1.1373, and the choice is worth forty per cent. The advice is sample-size dependent in a way its usual statement is not, and the sample sizes where it matters are the ones the advice is most often given about.
What could have produced these counts without the claim being true
Simulation noise. Each cell reports its own binomial standard error, and the sweep spends its budget where the answer is interesting: twenty thousand draws at twenty, thirty and fifty rows, ten thousand at a hundred, four thousand at two hundred and fifty and fifteen hundred at a thousand. The standard error is 0.224 points at the small end and 0.566 at the large one, so the thousand-row cell is the least certain number on the table and the one carrying the least weight in the argument. The claims that matter — 88.73% against a promise of 95%, and 89.63% against 93.52% — are between four and twenty-eight standard errors wide.
The estimand could be moving. It is not: the mean model in this sweep is correct, a straight line fitted to a straight-line truth, so the target is the truth’s own slope and every miss is a variance failure rather than the projection shift a wrong mean produces. Mixing the two would make a coverage table uninterpretable, since an interval can miss because it is the wrong width or because it is centred somewhere else.
The design could be doing the work. The covariate is a fixed even grid at every sample size, so the leverage structure is as benign as it gets and the maximum leverage at twenty rows is 0.1857. On a design with one far-out point the same corrections separate by a factor of sixteen rather than by a fifth, which is the argument about which correction to use and is a different essay’s subject. The numbers here are therefore the best case for the four corrections agreeing.
The reference could be the whole story. It is not: HC0 on a still covers only 90.68% at twenty rows, four points short, and the remaining shortfall tracks the estimate’s own bias of 0.8151 exactly as a first-order argument says it should. Two defects, and the sweep separates them by holding one fixed while the other moves — the same discipline the essay that separates two resampling defects uses on a different pair.
Where the shortfall is not the point
Everything above prices a robust standard error where a robust standard error is needed. It says nothing about what the correction costs where the model-based error was right all along, which is a separate sweep and a separate result: under a constant error variance the leave-one-out interval read against a covers 95.52% at twenty rows against the model-based 95.06%, and the price is width and variability rather than coverage. The essay that prices insurance where the risk is absent is where that count lives, and it changes the reading of this one: the case against defaulting to robust errors is not that they cost coverage when the assumption holds, but that they cost coverage when it fails mildly and the sample is small.
And the third repair is not measured here. The small-sample shortfall has two components, the estimate and the reference, and both have been treated by assuming a better reference rather than by generating one. Resampling the reference distribution is the standard third answer, and the machinery for the version that keeps a non-constant error variance is already measured in the essay on which residual resampling preserves it — a multiplier that leaves each residual where it is. Whether that closes the gap between 88.73% and 95% at twenty rows better than HC3 and a close it to 94.55% is one sweep, and it is not run here. It would have been a seventh figure in an essay already carrying six, which is an honest reason and not a good one.
What is settled is narrower than the recommendation it is usually attached to. A robust standard error delivers its promise asymptotically, delivers about a quarter of it at twenty rows uncorrected, and delivers nearly all of it at twenty rows if the leave-one-out weighting and a reference are both used. The default in most software is the uncorrected version against a normal, which is the worst cell of that table and the one every quoted justification is written about.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name confidence interval, coverage, monte carlo, sample size, t interval
- An interval that covers and says nothing — both name confidence interval, coverage, interval width, nominal level, wald interval
- The same draws for both methods — both name confidence interval, coverage, monte carlo, sample size, wald interval
- What the 95% refers to — both name confidence interval, coverage, monte carlo, sample size, wald interval
- A block size that changes — both name confidence interval, coverage, monte carlo, sample size
- A schedule that reads the mean — both name confidence interval, coverage, monte carlo, sample size
Named objects
A flat tag is an object no other essay names yet.
Asymptotic varianceConfidence intervalCoverageFinite-sample correctionHeteroskedasticityHeteroskedasticity-consistentInterval widthLeave-one-outMonte CarloNominal levelReference distributionRobust standard errorSample sizeT intervalWald interval