Right for the wrong reason
Worth reading first: What a wrong model estimates.
The sweep behind this essay was built to price a cost that turns out not to exist. Under a constant error variance — the case where the model-based standard error is exactly right and a robust one is insurance against nothing — the leave-one-out interval read against a on covers 95.52% at twenty rows. The model-based interval covers 95.06%. The binomial standard error on the second is 0.153 points, so the robust interval is, if anything, on the wrong side of nominal rather than short of it.
That is worth stating plainly because the standard argument against defaulting to robust errors is a coverage argument, and it is not supported here. What insurance costs is two other things: an interval 6.89% wider at twenty rows, and a variance estimate 2.572 times as variable as the one it replaces. Neither of those appears in a coverage table, which is why the price is usually described as if it were free or as if it were a miss.
The obvious way to avoid paying is to test first and switch only if the test rejects. That procedure was expected to cover worse than both of the errors it chooses between, and it does not do that either: it lands between them at every cell of every sweep here. What it does is land almost entirely on the wrong one, recovering 15.9% of what the robust error is worth at twenty rows and 13.8% at two hundred and fifty — a share that does not improve with the sample, for a reason that turns out to be an exact identity rather than a simulation result.
What insurance costs where there is nothing to insure
Five sample sizes, a constant error variance, twenty thousand draws at the small end. The three coverages at twenty rows are 95.06% for the model-based interval, 95.52% for the robust one and 95.22% for the pre-test procedure; at two hundred and fifty they are 95.06%, 95.08% and 95.09%. Nothing on that table is short of its promise, and the differences are within a few standard errors of each other.
The width ratio is where the cost shows: 1.0689 at twenty rows, 1.0446 at thirty, 1.0248 at fifty, 1.0127 at a hundred and 1.0049 at two hundred and fifty. Seven per cent at twenty rows and half a per cent at two hundred and fifty, falling like , which is the shape a finite-sample penalty has.
The second cost is larger and much less often quoted. The variance of the robust variance estimate, divided by the variance of the model-based one, is 2.572 at twenty rows for the leave-one-out weighting and 1.853 at two hundred and fifty. The denominator there is not counted: is exactly, so
with no simulation at all — 2.5125e-3 at twenty rows against a counted 2.5137e-3, and 1.1613e-6 against 1.1886e-6 at two hundred and fifty. Only the numerator of the ratio is a count, which is what makes the ratio worth quoting to three figures.
A variance estimate 2.572 times as variable is worth translating, since nobody reads a variance. Its standard deviation is times the other’s, and the interval’s half-width is the square root of the variance, so the width an analyst reports is itself about sixty per cent more variable under the robust error at twenty rows. Two analysts with the same design and the same true variance, drawing different samples, will disagree about how wide their intervals should be by half as much again as they would have under the default. That is a cost paid on every dataset, including the great majority in which the assumption was fine, and it is invisible to anybody comparing the two procedures on average coverage — which is the only comparison usually made.
There is a crossing inside that column worth noticing, because it inverts the usual ordering. At twenty rows the uncorrected estimate is the less variable of the two robust ones — 1.305 against HC3’s 2.572 — and at two hundred and fifty it is the more variable trend, 1.761 against 1.853. The two approach about 1.8 from opposite sides. HC3’s extra variability at small samples is the price of the leverage weighting that the field’s account of the four corrections shows buying its coverage; once the leverages are all near that weighting does nothing and the two estimators become the same object with a constant between them.
The procedure that tests first is not worse than both
Turn the variance dial up and the picture the pre-test is meant to improve appears. At a variance leaning hard towards the edges of the design, the model-based interval covers 87.17% at twenty rows and the robust one 94.50%. The procedure that runs a heteroskedasticity test at 5% and reports whichever error the test points to covers 88.33%.
It is between them, which was the first thing the sweep was built to check and which it refuses to disprove: at every sample size, both variance shapes and both auxiliary regressions, the pre-test’s coverage lands inside the interval spanned by its two branches. So the standard warning — that a pre-test procedure can be worse than either of the choices it selects between — is not what happens to coverage here, whatever it does to other properties.
What it is instead is nearly useless. The share of the gap between the two branches that the procedure recovers is 0.159 at twenty rows, 0.148 at thirty, 0.140 at fifty, 0.138 at a hundred and 0.138 at two hundred and fifty. The share does not grow. A reader expecting a procedure that behaves like the model-based error at small samples and like the robust one at large samples gets a procedure that behaves like the model-based error throughout.
A pre-test’s coverage is its test’s power
The reason the share is flat is arithmetic, and stating it as arithmetic is what turns a simulation result into something a reader can check without running anything.
Every draw either rejects or does not. On the draws that reject, the reported interval is the robust one; on the rest it is the model-based one. So by the law of total probability the procedure’s coverage is
with the test’s power and , the coverages of the branches on the samples that took them. At twenty rows those three numbers are 13.2%, 95.80% and 87.19%, which give 88.33% — the counted coverage, to the digit. The identity is checked at every cell of every sweep here and agrees to nine decimal places, which it must, since both sides count the same events.
So a pre-test inherits its branches in the proportion its test rejects, and nothing else. With a power of 13.2%, seven draws in eight report the model-based interval, and the procedure is the model-based interval with an eighth of the robust one mixed in. There is no adaptivity in it beyond the power, and a reader who wants to know what a diagnostic-then-decide procedure will deliver need only ask how often the diagnostic fires — the same question the essay on the check before a standard error asks about a different diagnostic and answers with the same shape of number.
The weakness is the auxiliary regression, not the procedure
A power of 13.2% at twenty rows against a departure severe enough to cost eight points of coverage is a bad test, and it stays bad: 13.8%, 13.9%, 14.0%, 13.7% across a factor of twelve in sample size. A test whose rejection rate does not move with the sample is not underpowered, it is aimed at the wrong thing — its own size is 5%, and it is barely clearing it.
The error variance here is , which is symmetric about the design’s centre: large at both ends and small in the middle. The auxiliary regression everybody runs regresses the squared residuals on the covariate itself, a linear term, and a linear function has no way to represent a symmetric bowl. The test is looking for a component the departure does not have, and it finds it at almost exactly the rate chance provides. This is the same failure as the balancing rule pointed at the wrong function of a covariate, where a rule that halves a variance when the covariate enters linearly is worth a fifth of that when it enters as a threshold.
Add a squared term to the auxiliary regression and the same procedure, with the same branches and the same 5% threshold, recovers 42.9% of the gap at twenty rows and 100% by two hundred and fifty. Its power goes 37.5%, 53.2%, 76.9%, 97.8%, 100.0%, and its coverage goes 90.31%, 91.75%, 93.29%, 94.67%, 94.76% — which is the robust interval’s own coverage, because by then every draw rejects. Under the milder setting the same change takes the power from 8.7%, 9.3%, 9.3%, 9.6%, 9.9% to 12.2%, 15.3%, 21.0%, 36.2%, 71.9%.
Reporting both is the point. Quoting the procedure under a test aimed at the wrong function would be a straw man, and the honest summary is a conditional one: a pre-test delivers what its test can detect, and the usual auxiliary regression cannot detect the departure used here at all. The corollary is uncomfortable, because a practitioner choosing an auxiliary regression is choosing which departures to be protected against, before seeing which one is present.
The setting where a pre-test would be most use is the one it is worst at
The mild departure is the interesting case for a procedure like this, because it is the one an analyst cannot see by eye and the one where the two branches genuinely differ. At a variance leaning half as hard, twenty rows, the model-based interval covers 92.13% and the robust one 95.00% — nearly three points at stake — and the pre-test covers 92.59%, recovering 0.160 of the gap.
Its power there is 8.7%, against a size of 5%. Halving the departure has nearly halved the excess rejection rate while leaving most of the coverage shortfall in place, which is the general shape of the problem: the coverage cost of heteroskedasticity falls roughly with and the detectability of it falls much faster, so the ratio of what is at stake to what can be found gets worse as the departure gets milder. That is the same relation between an effect and the probability of noticing it that the essay on what a p-value does not say states for a mean, and here it decides whether a diagnostic is worth running at all.
Even with the squared term the mild setting takes until two hundred and fifty rows to reach 71.9% power and 0.773 of the gap. So the procedure is reliable exactly where the departure is severe enough that a robust error would have been used anyway, and it is unreliable across the whole region where the question is live — which is what a diagnostic aimed at a decision must not be, and what the essay measuring a diagnostic at a size no enumeration reaches finds when it pushes a different test into the region where its answers stop agreeing.
The branch it keeps is worse than the interval it kept
There is a second effect inside the procedure, and it is the one that survives however good the test is.
At fifty rows with the aimed auxiliary regression, the model-based interval covers 87.41% across all draws. Across only the draws whose test did not reject — the draws on which the procedure actually reports it — it covers 84.32%. That is 3.09 points worse, and at a hundred rows the same pair reads 87.26% against 84.47%.
The mechanism is selection and it is not subtle. Passing a heteroskedasticity test means the squared residuals looked flat, and samples whose residuals look flat are disproportionately samples in which came out small. A small is a narrow model-based interval, so the procedure hands a reader the model-based interval precisely on the draws where it is narrowest. The test does not merely fail to switch; it switches away from the samples where switching was least needed and keeps the ones where the interval it keeps is worst. This is the structure the essay on an estimate made after a choice measures for a selected arm and the winner’s curse measures for a selected effect, arriving here through a variance rather than through a mean.
It also explains why the conditional coverage gets worse as the test gets better. With the blind auxiliary regression the kept branch covers 87.19% at twenty rows, almost exactly the unconditional 87.17%, because a test that rejects at random selects nothing. With the aimed one it covers 85.69% at twenty rows and 84.32% at fifty: a test with real power against the variance shape also has real correlation with the size of , and the selection sharpens as the power grows.
What could have produced these numbers without the claim being true
The test could be broken rather than blind. It is not: under a constant error variance its rejection rate is 5.1%, 5.1%, 4.9%, 4.9% and 5.4% across the five sample sizes, which is its nominal size. A test holding its size and failing to gain power against a specific alternative is a test aimed elsewhere, which is the claim; a test that had lost its size would be a bug.
The identity could be a rearrangement of the count. It is two counts of different things — the coverage flag over all draws on one side, and the power together with two conditional coverages on the other — combined by an argument that is exact rather than approximate. It would fail immediately if the branch labels and the interval actually reported ever disagreed, which is the sort of off-by-one this arrangement is built to catch.
The choice of robust estimator could be carrying the result. The robust branch here is the leave-one-out correction on a reference, which the coverage measurement shows is the best cell available at small samples. Choosing a weaker robust branch would narrow the gap between the branches and make the pre-test look better by making its alternative worse, which is the wrong direction for a favourable result.
The width comparison could be hiding a coverage failure. It is not, because both are reported: the robust interval at twenty rows is 6.89% wider and covers 95.52% against 95.06%. An interval that is wider and covers more is paying for what it gets, and the whole content of the insurance framing is that the payment is in width and estimator noise rather than in the promise.
Where the procedure is harmless, and what was not measured
Under a constant error variance the pre-test is fine: it covers 95.22% at twenty rows, between its branches’ 95.06% and 95.52%, and its width sits between theirs. A procedure that does nothing when there is nothing to do is not a failure, and if the departure a study faces is one its auxiliary regression can see, the procedure converges on the robust branch quickly — 97.8% power at a hundred rows with the squared term.
The case against it is narrower and harder to escape: its coverage is its test’s power, its test’s power is decided by a functional form chosen before the data, and the branch it keeps is selected to be the worst version of itself. Those three facts do not depend on the departure used here, and the third is the only one that cannot be repaired by a better auxiliary regression.
What is not measured is the version of the question a practitioner would ask next. This sweep prices the interval, and a pre-test also changes the point estimate’s distribution — nothing here counts the estimator’s own bias conditional on the test, which is a different sweep and a more familiar objection to pre-testing. Nor is there a continuum: the procedure is a hard switch at a 5% threshold, and the natural alternative is a weighted average of the two variance estimates whose weight is a smooth function of the test statistic. That would land somewhere between the two branches with a smaller variance than either, and whether the somewhere is better than 13.8% of the way across is one sweep and it is not run. The identity above says what to expect of it: whatever weight the data ends up putting on the robust branch is what the procedure’s coverage will be a mixture in, and a smooth weight is only an improvement if it is on average larger than the power of the test it replaces.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An interval that covers and says nothing — both name confidence interval, coverage, interval width, nominal level
- Intervals for the findings — both name confidence interval, coverage, interval width, selection effect
- The condition that cannot be dropped — both name coverage, estimated variance, interval width, selection effect
- The count or the length — both name confidence interval, coverage, estimated variance, interval width
- The interval with no resampling in it — both name confidence interval, coverage, estimated variance, interval width
- What studentising costs — both name confidence interval, coverage, estimated variance, interval width
Named objects
A flat tag is an object no other essay names yet.
Auxiliary regressionBreusch–Pagan testConditional inferenceConfidence intervalCoverageEstimated varianceHeteroskedasticityInterval widthLeave-one-outNominal levelPre-test estimatorRobust standard errorSelection effectSpecification testStatistical power