When the observations repeat each other

The interval the joint fit reports

Maximising a regression's coefficients and an eight-lag covariance band together gives a better estimate than plugging the band in. It also prints an interval that covers 71% to 85% of the time where 95% was promised — on the white predictor too, where every other rule is honest — because it reports almost the precision of the law's own covariance and delivers less than half of that advantage. Counting the eight lags it estimated buys one to two points.

Worth reading first: The observations that repeat each other · A design is a number.

The essay that priced a joint fit over a general covariance measured what practitioners experience as a risk averaged over draws, and set aside what they actually read off the output: the standard error printed beside each coefficient. It named the omission as a decision. A risk and an interval are different kinds of claim, and a fit whose covariance was maximised rather than plugged in might report its uncertainty correctly or might not; putting an unchecked interval next to a measured risk would have mixed the two.

This essay makes the check. The regression is the one the whole field runs: four predictors with lag-one correlations of 0.9, 0.6, 0.3 and nothing, true coefficients of 0.30, 0.20, 0.12 and zero, a hundred and twenty rows, errors drawn from each of the four laws in turn. Six rules fit it and each prints the usual interval — the coefficient plus or minus 1.96 standard errors, the standard error being what generalised least squares reports when the covariance it whitened by is taken as known. The question is how often that interval contains the true coefficient, over a thousand draws under each law.

The answer splits the rules cleanly. Five of them are about as honest as anyone would expect. The sixth is the joint fit, which wins on risk where it wins at all, and whose interval is the least honest of any rule that whitens.

Six intervals, one of them honest by right

Begin with the law the joint fit was found to help: a five-period moving average, whose dependence stops dead after four lags and is therefore exactly a band.

Six intervals, one of them honest by right. How often each rule's reported 95% interval covers the true coefficient under a five-period moving average, over 1000 draws of 120 rows, for the most persistent predictor (its own lag-one correlation 0.9) and the white one. The law's own covariance covers 94.9% and 94.1%. Least squares covers 65.0% on the persistent predictor and 94.4% on the white one; a fitted ρ̂ 93.0% and 96.3%; the two-step band 86.5% and 97.3%. The joint fit covers 73.2% and 71.4% — short on both, and short on the white predictor where every other rule is honest — and counting its eight band coefficients gives 76.1% and 73.5%.
Fig. 1 Coverage of each rule’s reported 95% interval under a five-period moving average, for the most persistent predictor and for the white one.

The law’s own covariance is the reference, and it does what the theory says it should: 94.9% on the persistent predictor and 94.1% on the white one. That is the oracle rule — it whitens by the true covariance and so is entitled to treat it as known — and its coverage is the evidence that the computation is right. Everything else is read against it.

Least squares, which treats the errors as independent, covers 65.0% on the persistent predictor and 94.4% on the white one. That is the oldest result in this field: a dependence in the errors matters to the standard error only through a dependence in the predictor, so it ruins the interval for the column that persists and leaves the column that does not almost untouched.

A fitted first-order correlation covers 93.0% and 96.3%, and the two-step band plug-in 86.5% and 97.3%. Both whiten, so both repair most of what least squares got wrong on the persistent predictor, and neither damages the white one.

The joint fit covers 73.2% on the persistent predictor and 71.4% on the white one. It is short on both, and on the white predictor it is short by twenty-four points where every other rule is within three of nominal. That is the result, and the rest of the essay is about where it comes from.

It is worth saying what that coverage means in the form a reader meets it. The white predictor’s true coefficient is zero, so its interval missing the truth is the same event as its t-test rejecting a true null. An analyst testing whether that column matters, at the conventional 5%, would be told it does on 28.6% of samples under the moving average if they used the joint fit — and on 2.7% if they used the two-step band, whose interval over-covers. One of those is a test at about its stated level and erring on the safe side; the other finds an effect that is not there more than one time in four.

It is not one law’s

The moving average is the law on which the joint fit has the most to offer, so it might be the law on which it overreaches. Run the same comparison under the first-order autoregression these regressions are mostly written under.

Six intervals, one of them honest by right. How often each rule's reported 95% interval covers the true coefficient under AR(1) at 0.8, over 1000 draws of 120 rows, for the most persistent predictor (its own lag-one correlation 0.9) and the white one. The law's own covariance covers 94.1% and 95.5%. Least squares covers 56.9% on the persistent predictor and 94.6% on the white one; a fitted ρ̂ 90.6% and 96.0%; the two-step band 83.7% and 96.5%. The joint fit covers 79.4% and 83.5% — short on both, and short on the white predictor where every other rule is honest — and counting its eight band coefficients gives 81.7% and 84.9%.
Fig. 2 The same six intervals under a first-order autoregression at 0.8.

The oracle covers 94.1% and 95.5%. Least squares covers 56.9% on the persistent predictor and 94.6% on the white one; a fitted ρ̂ 90.6% and 96.0%; the two-step band 83.7% and 96.5%. The joint fit covers 79.4% and 83.5%.

The shortfall is smaller here and has the same shape: the joint fit is the only rule short on the white predictor. Under all four laws the pattern repeats, and the figure below reads the white predictor alone, because that is the coefficient on which no other rule has anything to hide.

Short under every law. How often the reported 95% interval covers the true coefficient of the white predictor — the one column with no persistence of its own — for four rules under each of the four laws, over 1000 draws of 120 rows apiece. The law's own covariance covers 95.5%, 94.1%, 95.3%, 95.4%, and the two-step band 96.5%, 97.3%, 94.9%, 96.4%. The joint fit covers 83.5%, 71.4%, 85.3%, 79.5%, and charging its eight band coefficients to the residual degrees of freedom moves that to 84.9%, 73.5%, 86.1%, 81.0%. The shortfall is not a law's; it is the joint fit's.
Fig. 3 Coverage of the white predictor’s interval under each law, for the two-step band, the joint fit, the joint fit with its band counted, and the law’s own covariance.

The oracle covers 95.5%, 94.1%, 95.3% and 95.4% under the autoregression, the moving average, long memory and the break; the two-step band covers 96.5%, 97.3%, 94.9% and 96.4%. The joint fit covers 83.5%, 71.4%, 85.3% and 79.5%. Under no law does it come within nine points of what it promises, and it falls furthest under the moving average — the one law on which it was measured to beat the two-step on risk.

What it reports and what it delivers

A coverage figure is a symptom. Its cause is always the same: the standard error printed and the actual spread of the estimate across samples disagree. So set the two side by side, for the white predictor under the moving average.

What each rule reports and what it delivers. For the white predictor under a five-period moving average, over 1000 draws: the mean standard error each rule reports beside the actual standard deviation of its estimate across draws. Least squares reports 0.0884 and delivers 0.0885; the two-step band reports 0.0436 and delivers 0.0376; the law's own covariance reports 0.0173 and delivers 0.0177. The joint fit delivers 0.0289 — better than the two-step, worse than the law's own — and reports 0.0167, which is 97% of what the law's own covariance reports. It states nearly the precision of the oracle and delivers about 44% of the oracle's advantage over the two-step.
Fig. 4 For the white predictor under a five-period moving average: the mean standard error each rule reports beside the actual standard deviation of its estimate across draws.

Least squares reports 0.0884 and delivers 0.0885, which is why its interval covers on this column. The two-step band reports 0.0436 and delivers 0.0376 — a little pessimistic, which is why it over-covers. The oracle reports 0.0173 and delivers 0.0177.

The joint fit delivers 0.0289. That is a genuine improvement on the two-step, and it is the improvement the earlier essay measured as a risk: the joint estimate of a white predictor’s coefficient really is tighter. But the standard error it reports is 0.0167 — 97% of what the oracle reports. The joint fit prints nearly the precision of a rule that knows the covariance exactly, and it delivers about 44% of the oracle’s advantage over the two-step.

Put the other way round: its actual standard deviation is 1.585 times the root-mean-square of the standard errors it reports. Every rule that whitens by an estimate misstates its precision a little; the two-step on this column misstates it in the safe direction, at 0.858. The joint fit misstates it by more than half in the unsafe one.

Not always the oracle’s precision

“Nearly the oracle’s precision” is a description of the moving average, and it would be easy to promote it into a mechanism it is not. Under the break it fails.

The law with a break in its persistence is the one law here whose covariance no band contains: the dependence changes half way through the sample, so the true covariance is not Toeplitz and the oracle — whitening by it exactly — has an advantage nothing banded can approach. On the white predictor the oracle reports 0.0286 and delivers 0.0279. The two-step band reports 0.0472 and delivers 0.0441. The joint fit reports 0.0355 and delivers 0.0523.

So under the break the joint fit does not claim the oracle’s precision; it claims about a quarter more standard error than the oracle does. What it still does is claim more precision than it has: it reports a standard error two thirds the size of its spread, while the two-step reports one a little larger than its own. The joint fit’s spread here is no better than the two-step’s — 0.0523 against 0.0441, worse in fact — and its interval is narrower. That is the general statement, and the moving average is its sharpest instance rather than its definition: the joint fit’s reported precision runs ahead of its delivered precision under every law, by a ratio of 1.316, 1.585, 1.291 and 1.421 on the white predictor, against the two-step’s 0.920, 0.858, 0.982 and 0.929.

The persistent predictor, where everybody is short

The white predictor isolates the joint fit’s problem because nothing else is wrong there. The persistent predictor is the harder column, and every feasible rule is short on it.

Under the autoregression the actual spread of the persistent predictor’s estimate is 1.126 times the root-mean-square reported standard error for a fitted ρ̂, 1.332 for the two-step band and 1.430 for the joint fit; the oracle reads 1.007. Under long memory the fitted ρ̂ and the two-step cover 85.3% and 84.9%, and the joint fit 83.0%.

Two different shortfalls are stacked on this column. The first is shared by every rule that estimates the dependence: a predictor with a lag-one correlation of 0.9 regressed against errors with their own persistence is the setting in which an effective sample size collapses furthest, and an estimate of the covariance that is slightly wrong moves the variance of the coefficient a great deal. Under long memory, where no band of eight lags contains the law, the two-step and a fitted ρ̂ are short by the same ten points for the same reason. The second shortfall is the joint fit’s own, and it is the one already measured on the white predictor. On the persistent column it adds a few points to a gap that was already there, which is why the white predictor is the cleaner place to see it.

Why the maximum overstates what it knows

The reported standard error is σ^2(X~⊤X~)−1\hat\sigma^2(\tilde X^\top \tilde X)^{-1} on the whitened design — the generalised least-squares formula with the covariance treated as fixed. Each rule plugs in its own covariance, and the formula has no term for the fact that the covariance was estimated from the same hundred and twenty rows. That omission is the same for every rule. What differs is what the omission is hiding.

The two-step band’s covariance is tapered, and a taper shrinks every sample autocovariance towards zero. So the two-step believes the errors are less dependent than they are — a fact the field has measured from the likelihood’s side, where the tapered plug-in sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. For a white predictor, believing in less dependence means believing the predictor is less informative relative to the noise, and so a larger standard error. The two-step’s bias in its covariance becomes a cushion in its interval.

The joint fit removes much of the bias. Under the autoregression its first band coefficient averages 0.714, against the two-step’s 0.635 and a true value of 0.8, and its log-determinant averages −103.5 against −64.5. That is what it was built to do, and the essay that compared iteration with maximisation showed that re-reading the correlation from the generalised residuals recovers most of that kind of bias and the likelihood’s own determinant term the rest. But removing the bias removes the cushion as well, and what is left is the estimation noise in eight band coefficients fitted to a hundred and twenty rows, which the formula never priced for any rule.

The estimate it settles on is not only less biased; it is also selected. A maximum over a band of eight free numbers is drawn towards the covariance that makes this sample’s residuals look most independent after whitening, and a sample’s residuals looking more independent than its errors are is exactly the condition under which a standard error comes out too small. The fit that takes the memory out found the same thing one level down, with one correlation and residuals instead of errors; here it is eight coefficients and a likelihood.

The further the maximum goes, the less it covers

That reading makes a prediction: the draws on which the maximum moves furthest from the plug-in should be the draws on which the interval covers least. Sort the thousand draws under the autoregression by the log-determinant of the joint fit’s band — the more negative, the closer the band has been pushed towards a singular matrix — and read coverage of the white predictor in each quarter.

The further the maximum goes, the less it covers. The joint fit's coverage for the white predictor under AR(1) at 0.8, over 1000 draws sorted into quarters by the log-determinant of the band the maximum settles on — the more negative, the nearer the band sits to a singular matrix. The joint fit's band averages −103.5 against the two-step's −64.5, so on average the maximum moves towards the edge, and its first lag rises from 0.635 to 0.714. Coverage runs 75.6%, 84.0%, 84.4%, 90.0% from the quarter nearest the edge to the quarter furthest from it, and in that first quarter the reported standard error is 77% of the oracle's own.
Fig. 5 The joint fit’s coverage for the white predictor under a first-order autoregression, by quarter of how near its band sits to the edge of the positive definite cone.

Coverage runs 75.6%, 84.0%, 84.4% and 90.0% from the quarter nearest the edge to the quarter furthest from it. In the nearest quarter the reported standard error is 77% of the oracle’s own — smaller than what a rule that knew the covariance would report — and in the furthest it is 94%. The prediction holds: where the maximum goes furthest, the interval it reports is narrowest and wrongest.

Even the furthest quarter covers only 90.0%, so this is not a story about a few degenerate draws. It is a gradient across all of them.

Counting the lags is not the repair

The cheapest repair available is the one any textbook would reach for first: the joint fit estimated eight band coefficients, so charge them to the residual degrees of freedom and read the interval against Student’s t on what remains. That moves the critical value from 1.96 to 1.982 and the variance estimate up by eight rows’ worth.

It buys a point or two. On the white predictor it moves coverage from 83.5% to 84.9% under the autoregression, 71.4% to 73.5% under the moving average, 85.3% to 86.1% under long memory and 79.5% to 81.0% under the break. Every one is still short by more than eight points.

The reason is visible in the critical value that would actually have given 95% coverage — the 95th percentile of the absolute t-ratio across draws. For the oracle it is 1.92 on the white predictor under the autoregression, close to 1.96 as it should be. For the joint fit it is 3.73, and under the moving average 4.90. A rule whose t-ratio needs a critical value of nearly five is not a rule that miscounted its degrees of freedom by eight. Its t-ratio has heavy tails, because its denominator moves from draw to draw with the covariance it maximised, and the draws where the denominator is smallest are the draws where the numerator is not.

What stays honest

None of this makes the joint fit a bad estimator. On the moving average it delivers the tightest feasible estimate of every coefficient on the table, and on risk it is the one rule that beats the two-step on a law where it should. What it cannot do is report its own precision with the formula every other rule uses.

The rules that do report honestly are worth naming. The oracle is honest by right, and no analyst has it. A fitted first-order correlation covers within two points of nominal on three of four predictors under every law here and falls short on the persistent one by between two and ten points — its one estimated number is cheap to be wrong about. The two-step band is honest on everything but the persistent predictor, partly by design and partly by the accident of its taper.

The joint fit is the one rule here whose estimate improves and whose interval worsens, and the two are the same event. It moves the covariance towards the truth and towards this sample at once, and the formula reads both movements as knowledge.

Still open: an interval that knows the band was maximised

The next measurement is the repair the degrees-of-freedom count could not supply. Two are standard and neither is measured here. A sandwich estimate replaces σ^2(X~⊤X~)−1\hat\sigma^2(\tilde X^\top \tilde X)^{-1} with an expression that reads the whitened residuals’ own dependence, so a band that whitened too eagerly is caught by what it left behind; it costs nothing to compute and its honesty on a hundred and twenty rows is an open question, since the check that reads a residual’s dependence misses a lag-one correlation of 0.2 four times in five. A parametric bootstrap redraws from the fitted band, refits the joint maximum on each redraw, and reads the spread directly; it prices the maximisation by repeating it, and costs thirty joint fits per interval at least. Which of the two restores 95% under all four laws — and whether either does so without giving back the risk the joint fit earned — is a distinct question with a measurable answer, and it is the one this field should take next.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Confidence intervalCovariance matrixDegrees of freedomDependenceGeneralised least squaresLong memoryMonte CarloNuisance parameterParametric bootstrapProfile likelihoodSandwich estimatorStandard errorStructural breakTaperingWhitening