A standard error for a model that is wrong

Right for the wrong reason

A robust standard error costs no coverage where the risk is absent — 95.52% against 95.06% at twenty rows. It costs a 6.89% wider interval and a variance estimate 2.572 times as variable, and the pre-test that would avoid paying recovers 15.9% of what the insurance is worth.

Worth reading first: What a wrong model estimates.

The sweep behind this essay was built to price a cost that turns out not to exist. Under a constant error variance — the case where the model-based standard error is exactly right and a robust one is insurance against nothing — the leave-one-out interval read against a tt on n2n - 2 covers 95.52% at twenty rows. The model-based interval covers 95.06%. The binomial standard error on the second is 0.153 points, so the robust interval is, if anything, on the wrong side of nominal rather than short of it.

That is worth stating plainly because the standard argument against defaulting to robust errors is a coverage argument, and it is not supported here. What insurance costs is two other things: an interval 6.89% wider at twenty rows, and a variance estimate 2.572 times as variable as the one it replaces. Neither of those appears in a coverage table, which is why the price is usually described as if it were free or as if it were a miss.

The obvious way to avoid paying is to test first and switch only if the test rejects. That procedure was expected to cover worse than both of the errors it chooses between, and it does not do that either: it lands between them at every cell of every sweep here. What it does is land almost entirely on the wrong one, recovering 15.9% of what the robust error is worth at twenty rows and 13.8% at two hundred and fifty — a share that does not improve with the sample, for a reason that turns out to be an exact identity rather than a simulation result.

What insurance costs where there is nothing to insure

Five sample sizes, a constant error variance, twenty thousand draws at the small end. The three coverages at twenty rows are 95.06% for the model-based interval, 95.52% for the robust one and 95.22% for the pre-test procedure; at two hundred and fifty they are 95.06%, 95.08% and 95.09%. Nothing on that table is short of its promise, and the differences are within a few standard errors of each other.

The width ratio is where the cost shows: 1.0689 at twenty rows, 1.0446 at thirty, 1.0248 at fifty, 1.0127 at a hundred and 1.0049 at two hundred and fifty. Seven per cent at twenty rows and half a per cent at two hundred and fifty, falling like 1/n1/n, which is the shape a finite-sample penalty has.

The second cost is larger and much less often quoted. The variance of the robust variance estimate, divided by the variance of the model-based one, is 2.572 at twenty rows for the leave-one-out weighting and 1.853 at two hundred and fifty. The denominator there is not counted: s2s^2 is σ2χn22/(n2)\sigma^2\chi^2_{n-2}/(n-2) exactly, so

Var(V^ols)=V22n2\mathrm{Var}(\hat{V}_{\text{ols}}) = V^2\cdot\frac{2}{n-2}

with no simulation at all — 2.5125e-3 at twenty rows against a counted 2.5137e-3, and 1.1613e-6 against 1.1886e-6 at two hundred and fifty. Only the numerator of the ratio is a count, which is what makes the ratio worth quoting to three figures.

The price of insurance is noise, not coverage. What a robust standard error costs under a constant error variance — the case where the model-based one is exactly right — at five sample sizes over 20000 draws at the small end. It is not coverage: the leave-one-out interval read against a t on n − 2 covers 95.52% at 20 rows against the model-based 95.06%. It is a wider interval, by a factor of 1.0689 at 20 rows falling to 1.0049 at 250, and it is a variance estimate 2.57 times as variable at 20 rows and 1.85 times at 250 — against a denominator that is exactly V²·2/(n − 2), so only the numerator is counted. The uncorrected estimate is the cheaper of the two at small samples and the dearer at large: 1.31 against 1.76.
Fig. 1 What a robust standard error costs under a constant error variance: an interval 1.0689 times as wide at twenty rows falling to 1.0049 at two hundred and fifty, and a variance estimate 2.572 times as variable falling to 1.853.

A variance estimate 2.572 times as variable is worth translating, since nobody reads a variance. Its standard deviation is 2.572=1.60\sqrt{2.572} = 1.60 times the other’s, and the interval’s half-width is the square root of the variance, so the width an analyst reports is itself about sixty per cent more variable under the robust error at twenty rows. Two analysts with the same design and the same true variance, drawing different samples, will disagree about how wide their intervals should be by half as much again as they would have under the default. That is a cost paid on every dataset, including the great majority in which the assumption was fine, and it is invisible to anybody comparing the two procedures on average coverage — which is the only comparison usually made.

There is a crossing inside that column worth noticing, because it inverts the usual ordering. At twenty rows the uncorrected estimate is the less variable of the two robust ones — 1.305 against HC3’s 2.572 — and at two hundred and fifty it is the more variable trend, 1.761 against 1.853. The two approach about 1.8 from opposite sides. HC3’s extra variability at small samples is the price of the leverage weighting that the field’s account of the four corrections shows buying its coverage; once the leverages are all near 1/n1/n that weighting does nothing and the two estimators become the same object with a constant between them.

The procedure that tests first is not worse than both

Turn the variance dial up and the picture the pre-test is meant to improve appears. At a variance leaning hard towards the edges of the design, the model-based interval covers 87.17% at twenty rows and the robust one 94.50%. The procedure that runs a heteroskedasticity test at 5% and reports whichever error the test points to covers 88.33%.

It is between them, which was the first thing the sweep was built to check and which it refuses to disprove: at every sample size, both variance shapes and both auxiliary regressions, the pre-test’s coverage lands inside the interval spanned by its two branches. So the standard warning — that a pre-test procedure can be worse than either of the choices it selects between — is not what happens to coverage here, whatever it does to other properties.

What it is instead is nearly useless. The share of the gap between the two branches that the procedure recovers is 0.159 at twenty rows, 0.148 at thirty, 0.140 at fifty, 0.138 at a hundred and 0.138 at two hundred and fifty. The share does not grow. A reader expecting a procedure that behaves like the model-based error at small samples and like the robust one at large samples gets a procedure that behaves like the model-based error throughout.

The procedure is not either of its branches. What three 95% intervals cover at five sample sizes over 20000 draws at the small end, under an error variance leaning towards the edges of the design (γ = 0.8), with the heteroskedasticity test run on the auxiliary regression everybody runs. The model-based interval covers 87.17% at 20 rows; the robust one covers 94.50%; the procedure that tests at 5% and takes whichever the test points to covers 88.33%. It is not worse than both — which is what this sweep was built to price and did not find — and it is not much better than the worse one either, because it is the test's power that decides which branch is taken and the test rejects 13.2% of the time here. At 250 rows the three are 86.51%, 87.65% and 94.76%.
Fig. 2 Three 95% intervals at five sample sizes under a leaning error variance, with the usual auxiliary regression. The pre-test covers 88.33% at twenty rows against the model-based 87.17% and the robust 94.50%, and is still at 87.65% at two hundred and fifty.

A pre-test’s coverage is its test’s power

The reason the share is flat is arithmetic, and stating it as arithmetic is what turns a simulation result into something a reader can check without running anything.

Every draw either rejects or does not. On the draws that reject, the reported interval is the robust one; on the rest it is the model-based one. So by the law of total probability the procedure’s coverage is

Cpre=πCreject+(1π)Ckeep,C_{\text{pre}} = \pi\, C_{\text{reject}} + (1 - \pi)\, C_{\text{keep}},

with π\pi the test’s power and CrejectC_{\text{reject}}, CkeepC_{\text{keep}} the coverages of the branches on the samples that took them. At twenty rows those three numbers are π=\pi = 13.2%, Creject=C_{\text{reject}} = 95.80% and Ckeep=C_{\text{keep}} = 87.19%, which give 88.33% — the counted coverage, to the digit. The identity is checked at every cell of every sweep here and agrees to nine decimal places, which it must, since both sides count the same events.

So a pre-test inherits its branches in the proportion its test rejects, and nothing else. With a power of 13.2%, seven draws in eight report the model-based interval, and the procedure is the model-based interval with an eighth of the robust one mixed in. There is no adaptivity in it beyond the power, and a reader who wants to know what a diagnostic-then-decide procedure will deliver need only ask how often the diagnostic fires — the same question the essay on the check before a standard error asks about a different diagnostic and answers with the same shape of number.

The weakness is the auxiliary regression, not the procedure

A power of 13.2% at twenty rows against a departure severe enough to cost eight points of coverage is a bad test, and it stays bad: 13.8%, 13.9%, 14.0%, 13.7% across a factor of twelve in sample size. A test whose rejection rate does not move with the sample is not underpowered, it is aimed at the wrong thing — its own size is 5%, and it is barely clearing it.

The error variance here is σ02(1+γ(u2/v1))\sigma_0^2(1 + \gamma(u^2/v - 1)), which is symmetric about the design’s centre: large at both ends and small in the middle. The auxiliary regression everybody runs regresses the squared residuals on the covariate itself, a linear term, and a linear function has no way to represent a symmetric bowl. The test is looking for a component the departure does not have, and it finds it at almost exactly the rate chance provides. This is the same failure as the balancing rule pointed at the wrong function of a covariate, where a rule that halves a variance when the covariate enters linearly is worth a fifth of that when it enters as a threshold.

Add a squared term to the auxiliary regression and the same procedure, with the same branches and the same 5% threshold, recovers 42.9% of the gap at twenty rows and 100% by two hundred and fifty. Its power goes 37.5%, 53.2%, 76.9%, 97.8%, 100.0%, and its coverage goes 90.31%, 91.75%, 93.29%, 94.67%, 94.76% — which is the robust interval’s own coverage, because by then every draw rejects. Under the milder setting the same change takes the power from 8.7%, 9.3%, 9.3%, 9.6%, 9.9% to 12.2%, 15.3%, 21.0%, 36.2%, 71.9%.

The procedure is not either of its branches. What three 95% intervals cover at five sample sizes over 20000 draws at the small end, under an error variance leaning towards the edges of the design (γ = 0.8), with the heteroskedasticity test run on an auxiliary regression carrying a squared term. The model-based interval covers 87.17% at 20 rows; the robust one covers 94.50%; the procedure that tests at 5% and takes whichever the test points to covers 90.31%. It is not worse than both — which is what this sweep was built to price and did not find — and it is not much better than the worse one either, because it is the test's power that decides which branch is taken and the test rejects 37.5% of the time here. At 250 rows the three are 86.51%, 94.76% and 94.76%.
Fig. 3 The same three intervals with a squared term in the auxiliary regression. The pre-test now covers 90.31% at twenty rows and 94.76% at two hundred and fifty, where it has become the robust interval outright.

Reporting both is the point. Quoting the procedure under a test aimed at the wrong function would be a straw man, and the honest summary is a conditional one: a pre-test delivers what its test can detect, and the usual auxiliary regression cannot detect the departure used here at all. The corollary is uncomfortable, because a practitioner choosing an auxiliary regression is choosing which departures to be protected against, before seeing which one is present.

A pre-test keeps what its test rejects. The share of the gap between the model-based interval and the robust one that a pre-test procedure recovers, against how often its test rejects, at five sample sizes and two auxiliary regressions, under an error variance leaning towards the edges of the design (γ = 0.8). With the auxiliary regression everybody runs — the squared residuals on the covariate — the test rejects 13.2% of the time at 20 rows and 13.7% at 250, and the procedure keeps 15.9% and 13.8% of what the robust error is worth. The share does not grow with the sample because the power does not: this error variance is symmetric in the covariate and a linear auxiliary regression has nothing to find. Add a squared term and the same procedure keeps 42.9% at 20 rows and 100.0% at 250. A pre-test's coverage is its test's power and nothing else.
Fig. 4 The share of the robust error’s worth a pre-test recovers, against how often its test rejects, at five sizes and two auxiliary regressions. The points lie near the diagonal on which the share is the power.

The setting where a pre-test would be most use is the one it is worst at

The mild departure is the interesting case for a procedure like this, because it is the one an analyst cannot see by eye and the one where the two branches genuinely differ. At a variance leaning half as hard, twenty rows, the model-based interval covers 92.13% and the robust one 95.00% — nearly three points at stake — and the pre-test covers 92.59%, recovering 0.160 of the gap.

Its power there is 8.7%, against a size of 5%. Halving the departure has nearly halved the excess rejection rate while leaving most of the coverage shortfall in place, which is the general shape of the problem: the coverage cost of heteroskedasticity falls roughly with γ\gamma and the detectability of it falls much faster, so the ratio of what is at stake to what can be found gets worse as the departure gets milder. That is the same relation between an effect and the probability of noticing it that the essay on what a p-value does not say states for a mean, and here it decides whether a diagnostic is worth running at all.

Even with the squared term the mild setting takes until two hundred and fifty rows to reach 71.9% power and 0.773 of the gap. So the procedure is reliable exactly where the departure is severe enough that a robust error would have been used anyway, and it is unreliable across the whole region where the question is live — which is what a diagnostic aimed at a decision must not be, and what the essay measuring a diagnostic at a size no enumeration reaches finds when it pushes a different test into the region where its answers stop agreeing.

The branch it keeps is worse than the interval it kept

There is a second effect inside the procedure, and it is the one that survives however good the test is.

At fifty rows with the aimed auxiliary regression, the model-based interval covers 87.41% across all draws. Across only the draws whose test did not reject — the draws on which the procedure actually reports it — it covers 84.32%. That is 3.09 points worse, and at a hundred rows the same pair reads 87.26% against 84.47%.

The mechanism is selection and it is not subtle. Passing a heteroskedasticity test means the squared residuals looked flat, and samples whose residuals look flat are disproportionately samples in which s2s^2 came out small. A small s2s^2 is a narrow model-based interval, so the procedure hands a reader the model-based interval precisely on the draws where it is narrowest. The test does not merely fail to switch; it switches away from the samples where switching was least needed and keeps the ones where the interval it keeps is worst. This is the structure the essay on an estimate made after a choice measures for a selected arm and the winner’s curse measures for a selected effect, arriving here through a variance rather than through a mean.

It also explains why the conditional coverage gets worse as the test gets better. With the blind auxiliary regression the kept branch covers 87.19% at twenty rows, almost exactly the unconditional 87.17%, because a test that rejects at random selects nothing. With the aimed one it covers 85.69% at twenty rows and 84.32% at fifty: a test with real power against the variance shape also has real correlation with the size of s2s^2, and the selection sharpens as the power grows.

What could have produced these numbers without the claim being true

The test could be broken rather than blind. It is not: under a constant error variance its rejection rate is 5.1%, 5.1%, 4.9%, 4.9% and 5.4% across the five sample sizes, which is its nominal size. A test holding its size and failing to gain power against a specific alternative is a test aimed elsewhere, which is the claim; a test that had lost its size would be a bug.

The identity could be a rearrangement of the count. It is two counts of different things — the coverage flag over all draws on one side, and the power together with two conditional coverages on the other — combined by an argument that is exact rather than approximate. It would fail immediately if the branch labels and the interval actually reported ever disagreed, which is the sort of off-by-one this arrangement is built to catch.

The choice of robust estimator could be carrying the result. The robust branch here is the leave-one-out correction on a tt reference, which the coverage measurement shows is the best cell available at small samples. Choosing a weaker robust branch would narrow the gap between the branches and make the pre-test look better by making its alternative worse, which is the wrong direction for a favourable result.

What each error is a claim about, and what the claim comes out as. Each variance estimate's average over 20000 draws, divided by the variance the slope actually has across those same draws, at 80 rows with the error variance leaning towards the edges of the design (γ = 0). One is a standard error that is right. The model-based estimate reads 1.0069 of the spread, so its standard error is 100.34% of the one it should report; the four robust corrections read 0.9714, 0.9963, 1.0066, 1.0432. Two further routes agree with the count and share none of its arithmetic: n times the counted variance is 2.9783 against a population sandwich of 3.0000, and the counted ratio of the two standard errors is 0.9999 against a closed form of 1.0000.
Fig. 5 Each estimate’s average over twenty thousand draws of eighty rows under a constant error variance, against the spread the slope actually has. The model-based estimate reads 1.0069 and HC2 1.0066 — equally right, and not equally precise.

The width comparison could be hiding a coverage failure. It is not, because both are reported: the robust interval at twenty rows is 6.89% wider and covers 95.52% against 95.06%. An interval that is wider and covers more is paying for what it gets, and the whole content of the insurance framing is that the payment is in width and estimator noise rather than in the promise.

Half the shortfall is the quantile, not the estimate. What each 95% interval covers at 20 rows over 20000 draws, read against two references. The uncorrected robust interval covers 88.73% against a normal quantile and 90.68% against a t on n − 2 degrees of freedom — the same variance estimate, 1.95 points apart, bought by changing a number that costs nothing to change. The leave-one-out correction with a t reference reaches 94.55%. So a robust interval's small-sample shortfall is two defects rather than one: a variance estimate that is short, and a reference distribution that is wrong for the same reason the model-based one needs a t.
Fig. 6 The eight intervals at twenty rows under a leaning variance, which is what the pre-test’s robust branch is choosing between. The leave-one-out correction on a t reference covers 94.55%; the uncorrected one on a normal covers 88.73%.

Where the procedure is harmless, and what was not measured

Under a constant error variance the pre-test is fine: it covers 95.22% at twenty rows, between its branches’ 95.06% and 95.52%, and its width sits between theirs. A procedure that does nothing when there is nothing to do is not a failure, and if the departure a study faces is one its auxiliary regression can see, the procedure converges on the robust branch quickly — 97.8% power at a hundred rows with the squared term.

The case against it is narrower and harder to escape: its coverage is its test’s power, its test’s power is decided by a functional form chosen before the data, and the branch it keeps is selected to be the worst version of itself. Those three facts do not depend on the departure used here, and the third is the only one that cannot be repaired by a better auxiliary regression.

What is not measured is the version of the question a practitioner would ask next. This sweep prices the interval, and a pre-test also changes the point estimate’s distribution — nothing here counts the estimator’s own bias conditional on the test, which is a different sweep and a more familiar objection to pre-testing. Nor is there a continuum: the procedure is a hard switch at a 5% threshold, and the natural alternative is a weighted average of the two variance estimates whose weight is a smooth function of the test statistic. That would land somewhere between the two branches with a smaller variance than either, and whether the somewhere is better than 13.8% of the way across is one sweep and it is not run. The identity above says what to expect of it: whatever weight the data ends up putting on the robust branch is what the procedure’s coverage will be a mixture in, and a smooth weight is only an improvement if it is on average larger than the power of the test it replaces.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Auxiliary regressionBreusch–Pagan testConditional inferenceConfidence intervalCoverageEstimated varianceHeteroskedasticityInterval widthLeave-one-outNominal levelPre-test estimatorRobust standard errorSelection effectSpecification testStatistical power