Estimating the dependence, not naming it

The window that has to be chosen, and the term that was dropped

An estimated covariance has a bandwidth in it, and both ends of the dial are wrong for different reasons. The rule a practitioner would reach for is two thirds worse than the best window there is.

Worth reading first: Choosing the order · The observations that repeat each other.

Estimating a dependence rather than naming it moves the problem rather than solving it. The named version needs one number and the estimated version needs a window — how many lags of the sample autocovariance sequence to keep — and a window is a choice with a cost on either side of it.

This is not a new shape. The essay about a block length has the same dial with the same two wrong ends, and so does choosing the order of an autoregression. What is worth measuring here is where the optimum sits, how far the rule a practitioner would actually use sits from it, and what happens to a comparison when one of the rules being compared does not exist on most of the draws.

Both ends, and neither near

A window with two wrong ends. The regret of a rule whitened by a Bartlett-tapered Ω̂, as the window widens, against three rules that need no window at all. At L = 0 the estimate is the identity and the rule is exactly least squares — 0.08556, the same number to five places. It falls to 0.01754 at L = 20 and rises again by L = 30, because a quarter of the sample's lags are then being estimated from it. The automatic bandwidth a practitioner would reach for, 4(n/100) to the power 2/9, which at this sample size is 4, gives 0.02963 — 69% above the best available window. The rule told the dependence is an AR(1) sits at 0.01440 throughout, which is the price of not knowing the form.
Fig. 1 The regret of a rule whitened by a Bartlett-tapered estimate, as the window widens, against three rules that need no window at all. At L = 0 the estimate is the identity, so the rule is exactly least squares — 0.08556, the same number to five places, which is the check that the sweep starts where it should.

The curve falls to 0.01754 at L = 20 and rises again to 0.01977 by L = 30. Both ends are wrong and they are wrong for unrelated reasons.

At the narrow end the estimate is nearly the identity: a window of one keeps a single autocorrelation and gets 0.05312, which is barely a third of the way from least squares to the rule that knows ρ. The whitening cannot remove a dependence it is not allowed to see.

At the wide end the estimate is over-parameterised. Thirty lags estimated from a hundred and twenty rows means the longest of them is a product of ninety terms, and the taper is putting weight 1/31 on it. Those estimates are mostly noise, the noise goes into a matrix that then gets inverted, and the whitening starts removing structure that is not there.

What sits between them is not a plateau. The best window beats the wide end by 13%, which is more than a standard error, and beats the narrow end by a factor of three.

The window a practitioner would use

There is a standard automatic bandwidth for exactly this problem — 4(n/100)2/94(n/100)^{2/9}, which at n = 120 gives L = 4 — and it is not the minimiser of anything measured here. At four lags the rule gives 0.02963, which is 68.9% above the best window available.

That is worth stating carefully, because the automatic rule is not wrong. It is derived for a different loss: it minimises the mean squared error of a long-run variance estimate, which is what a standard error needs. Here the estimate is not being reported, it is being used to whiten a fit, and a whitening cares about the shape of the dependence across the whole band rather than about the sum of it. The window that estimates γ̂ well and the window that whitens well are answers to different questions, and the fact that everybody uses the first for the second is the kind of substitution this site exists to count.

Even so, the automatic window is inside the useful range: at 0.02963 it recovers 65.4% of what counting rows gives up, against 79.5% for the best window and 90.3% for the rule that is told ρ. A rule chosen for the wrong loss is still much better than no rule, which is the honest version of the finding and the less quotable one.

There is a second reason the two losses pull apart, and it is specific to what a criterion is. A long-run variance is a sum over the band — Σ γ(k) — so an error at one lag can be cancelled by an error of the other sign at another, and a window is chosen to trade the bias of the lags it drops against the variance of the lags it keeps. A whitening is not a sum. It inverts the matrix those autocovariances build, and an inverse is sensitive to the shape of the sequence rather than to its total: two sequences with the same sum and different profiles whiten differently. That is why the useful window here is five times the automatic one — the extra lags contribute almost nothing to the sum and a good deal to the shape.

A rule that looks best where it barely exists

The truncated estimate — the one that is not always a covariance matrix — has to be handled with more care than a row in a table allows, and the reason is a selection effect that the table cannot show.

A rule that looks best where it barely exists. The share of draws on which a truncated Ω̂ has a Cholesky factor, by window. At L = 1 no draw admits a factorisation at all, so the rule has no regret to report there. At L = 2 it is 1.0%, and averaged over that 1% the rule's regret reads 0.00718 — below the rule that is told the true ρ. It is not a better rule; it is the same rule measured on the draws whose residual sequences happened to be benign enough to factor. An availability rate is not a footnote to a comparison like this one, it is the comparison, and a table without it reads as though every row were computed on the same draws.
Fig. 2 The share of draws on which the truncated estimate has a Cholesky factor at all, by window, with the regret each of those averages is taken over. At two lags it exists on 1.0% of draws and its regret reads 0.00718 — better than the rule that is told the true ρ.

A rule cannot beat the rule that knows the answer. What is happening is that at L = 2 the truncated sequence factors only when the residuals happen to look benign, and a sample whose residuals look benign is a sample the selection problem is easy on. The 0.00718 is a number about which draws survived, not about the rule.

This is the same defect as reporting a search’s winner at its own score, arriving somewhere nobody would look for it. There the selection is over models and here it is over draws; in both cases a number is quoted for a set that was chosen using the quantity being reported.

An availability rate is not a footnote to a comparison like this one. It is the comparison. A table of regrets with one row computed on 45% of the draws and the rest on all of them is a table that reads as though every row were the same experiment.

The dial is steeper on the side the standard rule errs towards

Both ends of the sweep are wrong and they are not equally wrong, and the asymmetry decides which direction an unsure practitioner should lean.

Between the automatic window’s four lags and the best twenty, the regret falls by 0.01209 — about 0.00076 a lag. Between twenty and thirty it rises by 0.00223, about 0.00022 a lag.

The narrow side is three and a half times steeper. So a window that is too wide by ten lags costs a fifth of what one too narrow by sixteen costs, and the safe direction is unambiguously outward.

Which is the direction the standard rule does not go. 4(n/100)2/94(n/100)^{2/9} lands at four lags, sixteen short of the optimum, on the steep side — so the rule everybody uses is not merely aimed at a different loss but errs in the expensive direction of the loss it is being used for. A practitioner who simply doubled the automatic window would recover about half of the 68.9% gap for nothing, and one who quadrupled it would land at the optimum.

The essay’s own reading of why the two losses differ predicts that direction. A long-run variance is a sum, so the automatic rule stops adding lags once they contribute nothing to the total; a whitening reads the shape, and the lags that contribute nothing to a sum still contribute to an inverse.

The truncated rule’s one per cent beats the oracle

The availability effect can be given a floor, and the floor is stark.

At two lags the truncated estimate exists on 1.0% of draws and posts a regret of 0.00718. The rule that is told the true dependence, averaged over all draws, posts about 0.0083.

A rule cannot beat the oracle on the same draws, so the one per cent that survive must be draws on which the oracle itself would score below 0.00718 — at least 13% easier than the average draw, and almost certainly far easier than that, since the truncated rule is nowhere near the oracle on any draw it does exist on.

So the availability column does not merely qualify the regret column; it inverts it. A row reading better than the oracle is not evidence of a good rule, it is arithmetic proof that the row was computed on a selected subset — and the selection is on exactly the quantity being reported, since a residual sequence benign enough to factor is a residual sequence the whitening has little to repair.

That is the cleanest available statement of why an availability rate belongs in the table. Without it the row reads as a rule that outperforms an oracle, which is impossible, and nothing else in the table says so.

The term the likelihood has and the criterion drops

There is a second thing in this field that only becomes visible once a whitening stops being a single transform, and it was there all along.

A criterion of the form n·log(RSS/n) + 2q is a linearisation of the Gaussian log-likelihood with the variance concentrated out. For y ~ N(Xβ, σ²Ω) that likelihood is

2logL=nlog(2πσ2)+logΩ+(yXβ^)Ω1(yXβ^)σ2-2\log L = n\log(2\pi\sigma^2) + \log|\Omega| + \frac{(y - X\hat\beta)^\top \Omega^{-1} (y - X\hat\beta)}{\sigma^2}

and concentrating σ² leaves n·log(RSS_w/n) + log|Ω| up to constants, where RSS_w is the whitened residual sum. So the honest whitened criterion is

AIC=nlog ⁣(RSSwn)+logΩ+2q.\mathrm{AIC} = n\log\!\left(\frac{\mathrm{RSS}_w}{n}\right) + \log|\Omega| + 2q .

The rule that whitens and then applies the ordinary penalty drops the middle term. For the first-order transform that term is exactly −log(1 − ρ̂²) — one term, not n of them, because the transform’s Jacobian and the determinant of Ω very nearly cancel.

Dropping a constant changes no comparison, and if ρ̂ were shared across the table the term would be a constant. It is not: each candidate reads its own ρ̂ from its own residuals, so each carries its own determinant.

At a persistence of 0.85 the dropped term runs from 0.8784 to 1.0115 across the fifteen candidates — a spread of 0.1331, where a penalty difference between candidates of adjacent size is 2. So the missing term is worth about a fifteenth of one coefficient, and the honest prediction is that it should barely matter.

It barely matters, and it is not nothing. Restoring it moves the rule from 0.01442 to 0.01366, a paired difference of 0.00076 at 2.3 standard errors of the pair. Nine tenths of a point of the eighty-three the rule recovers.

The price of each thing the rule is not told. What each rule gives up against the best model available, at a persistence of 0.85 on a fifteen-candidate table, over 400 draws. Reading down: least squares with the ordinary penalty; the whitening at the true ρ; the same at a ρ̂ estimated per candidate; that rule with the term the Gaussian likelihood carries and it omits; a Bartlett-tapered Ω̂ estimated once from the fullest candidate at L = 8; the same estimated per candidate; and the truncated Ω̂, which exists on only 45.0% of draws and is averaged over those. Knowing ρ recovers 89.9% of what counting rows gives up, estimating it 83.4%, and estimating a whole covariance 75.6%.
Fig. 3 The two rules side by side, third and fourth from the top: the whitened criterion as the field before this one wrote it, and the same rule with its determinant restored.

This is a refusal that barely bites, and those are worth as much as the ones that do. The claim being refused — a whitened criterion may use the ordinary penalty — is exactly right when the whitening is shared and wrong when it is not, and the size of the error is a fact about how much ρ̂ varies across the table rather than about the principle. Had the candidates disagreed about ρ̂ by half instead of by three hundredths, the same omission would have been worth a coefficient and a half.

What a window is, next to what a parameter is

Two dials, and it is worth putting them side by side because they look alike and are not.

The parameter ρ̂ is estimated from a hundred and twenty residuals and its sampling error is a few hundredths. Its systematic part — a fit smooths its own residuals — is of the same order. A criterion reads differences and the systematic part is nearly shared, so most of the error cancels.

The window L is not estimated at all. It is chosen, before any data is seen, by a rule that is derived for a different loss, and its error does not cancel between candidates because every candidate is whitened by the same badly chosen matrix. What a shared nuisance buys is cancellation of the sampling error; a shared choice is shared exactly, and so is the damage it does.

That is why the window’s cost shows up as a level shift in the whole ladder rather than as noise, and why the sweep has a shape at all.

The distinction matters beyond this field. Every essay in this collection about an estimated nuisance — the ratio a two-arm interval needs, the spread a hierarchical model borrows against, the dependence a resampling has to reproduce — has been about sampling error, and the standing result is that a nuisance shared by everything being compared is cheap. A choice shared by everything being compared is not cheap in the same way and is not expensive in the same way either: it moves every score by the same amount, so it cannot change a comparison, and it changes what is being compared. The window is the first object in this collection that behaves like that, and it is worth naming as a category rather than as a special case.

A rule derived for a sum, applied to an inverse

Newey–West’s bandwidth is not an arbitrary rule of thumb and it is not wrong. It is derived, carefully, to minimise the mean squared error of a long-run variance — the scalar Σ_k γ(k), summed across the band with whatever weights the window supplies. Against that target the derivation is right and the answer it gives at n = 120 is 4.

A whitening is not that target. It needs Ω1/2\Omega^{-1/2}, and an inverse square root is not a functional of the sum of the autocovariances; it is a functional of the whole matrix they build, and it is dominated by the matrix’s smallest eigenvalues rather than by its largest. A sum is forgiving of shape — errors of opposite sign in different lags cancel inside it, which is precisely why a short band can estimate it well. An inverse is the opposite: an eigenvalue estimated near zero when it is not near zero produces an enormous multiplier in the whitened design, and no amount of accuracy elsewhere in the band compensates for it.

That is why a bandwidth chosen for one is 68.9% above the best available for the other, and why the direction of the error is the one it is. Four lags do not underestimate the sum badly. They misstate the decay, and a persistent process whose decay is cut off at four lags looks, to an inverse, like a process with far less long-range structure than it has — so the whitening under-corrects, and 0.02963 is what under-correcting costs against 0.01754 at the right window.

The general form is worth more than the number. An automatic tuning rule carries with it the loss function it was derived under, and that loss function is almost never stated at the call site. It travels as a default, gets applied to whatever quantity is nearby, and is right exactly as often as the two quantities happen to share a minimiser. Here they do not, and nothing in the interface says so: the same autocovariances, the same window family, the same code, one number that means something different in each place.

And there is no escaping the choice by declining to make it, because both ends are wrong: L = 0 reproduces least squares to five places and gives up the whole repair, thirty lags costs 0.01977 by fitting noise into the band. The optimum at 20 is interior, which is what makes this a selection problem rather than a limit.

What is claimed here, and what is not

This essay takes the window an estimated dependence has to be read over. The claims are that the sweep has an interior optimum at L = 20 which beats both ends; that the standard automatic bandwidth of 4 is 68.9% above it, because it is derived to estimate a long-run variance well rather than to whiten well; that the truncated rule’s regret is unreadable without its availability rate, since at the narrowest windows it exists only on the draws it finds easy; and that the log determinant a whitened criterion drops is worth 0.00076 of regret at 2.3 paired standard errors, which is small because the candidates’ ρ̂ agree to within three hundredths.

What stays out and is named as a decision: choosing the window from the data, by any of the plug-in rules that exist for it. That is a third selection problem on top of the two this field already has, it would have to be priced against the same table it is being chosen on, and every measurement here holds L fixed so that the sweep is a sweep over one thing.

The boundary against the first essay in this field is that it is about what an estimated covariance costs at a fixed window and this one is about the window.

The checks, and the refusals that make them mean something

Two claims are gated. The sweep is required to have an interior optimum that beats both ends by more than a standard error, which is the only reason a window is a choice rather than a formality. And L = 0 is required to reproduce least squares exactly, because a whitening by the identity is not a whitening and a sweep whose zero point disagreed with the rule it should equal would be measuring something other than the window.

Two refusals bite. A truncated estimate quoted without its availability rate is rejected — at L = 2 that rate is 1.0% and the rule reads better than the one that knows the truth. And a whitened criterion at a per-candidate Ω̂ without its determinant is rejected, on the ground that the term runs 0.8784 to 1.0115 across the table rather than taking one value, so it is not a constant and does not cancel.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • One number for a table of candidates — both name autocorrelation, degrees of freedom, dependence, generalised least squares, information criterion, least squares, model selection, persistence
  • The charge nobody derived — both name covariance matrix, degrees of freedom, information criterion, long-run variance, model selection, nuisance parameter, persistence, tapering
  • A family before a fit — both name autocorrelation, covariance matrix, dependence, generalised least squares, model selection, nuisance parameter, tapering
  • A lag the sample has less of — both name degrees of freedom, dependence, estimation error, information criterion, long-run variance, model selection, tapering
  • A line that beats two curves — both name degrees of freedom, dependence, estimation error, information criterion, least squares, model selection, tapering
  • The comparison that was not made — both name covariance matrix, degrees of freedom, dependence, information criterion, model selection, nuisance parameter, tapering

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationCovariance matrixDegrees of freedomDependenceEstimation errorGeneralised least squaresInformation criterionLeast squaresLikelihoodLong-run varianceModel selectionNuisance parameterPersistenceSpecification searchTapering