The shape a dependence has

The window a whitening wants

Every law here is best whitened by a window several times longer than its own memory, including the one whose memory ends at the fourth lag. The three ways of choosing it from the sample all land in the same place, and it is the wrong one.

Worth reading first: Choosing the order · The observations that repeat each other.

An estimated covariance has a window in it: how many lags the estimate carries before it stops. The rule that repairs a criterion by whitening its sample holds that window fixed and says so, and the obvious reading of what it should be fixed at is the memory of the errors — carry the lags that have something in them and stop.

That reading is wrong in every world measured here, and it is wrong by a factor of between three and eight.

The best window is not the memory

The window a whitening wants is not the memory of the errorsRegret under a five-period moving average as the tapered estimate is given more lags, over 120 draws at n = 120. The best window is L = 30; the automatic bandwidth is 4 and the error model's own likelihood chooses 9.7 on average. Both land in the same place and both are short, and the reason is the taper: a Bartlett weight at lag k is 1 − k/(L + 1), so a window of 8 keeps 0.556 of whatever the fourth lag carries and a window of 30 keeps 0.871. A window has to be several times the memory before it stops removing the memory. The dashed line is the rule told the errors are a first-order autoregression, which needs no window at all.0.0200.0300.0400.05000.3010.6020.9031.081.301.481.601.78log₁₀ of the window — how many lags the estimate carriesregret against the best model availabletold the form, no windowbest at L = 30the automatic bandwidth120 draws, trianglethe choices all land left of the optimum
Fig. 1 Regret as the tapered estimate is given more lags, with the law on the slider. The dashed line is the rule told the errors are a first-order autoregression, which needs no window at all.

The clearest case is the error that is a five-period moving average. Its autocorrelation is 0.8, 0.6, 0.4, 0.2 and then exactly zero for ever: there is nothing at the fifth lag or at any lag after it, and this is not an approximation. The best window for whitening it is L = 30.

Thirty lags, of which twenty-six are known in advance to be empty. At L = 8 — twice the memory — the rule reads 0.02374 against 0.01740 at thirty, so carrying nothing but noise in twenty-two extra lags is worth a third of the regret.

The other three worlds say the same thing more quietly. Under the geometric law the best window is 12; under long memory 30; under the break 40. Every one of them is between a tenth and a third of the sample, and the automatic bandwidth a practitioner would reach for is 4.

Because the taper removes what it is estimating

The mechanism is the taper, and it is one line of arithmetic. A tapered estimate does not use γ̂(k); it uses γ̂(k) w(k, L), where the Bartlett weight is

w(k,L)=1kL+1,w(k, L) = 1 - \frac{k}{L + 1},

which is what makes the result a covariance matrix rather than a table of numbers that happens to be symmetric. The weight is not a detail at the edge of the window. At L = 8 the fourth lag is kept at 0.556 of what the sample reported, and at L = 30 it is kept at 0.871.

So under the moving average, a window of 8 throws away nearly half of the dependence at the very lag where the dependence is, and a window of 30 throws away an eighth. The extra lags are not carrying information; they are paying for the taper, by pushing the weights at the lags that matter back towards one.

There is a second attenuation underneath it and it points the same way. A sample autocovariance at lag k averages n − k pairs and subtracts a sample mean, so what a sample reports is short before any taper touches it — under long memory the first lag arrives at 0.538 against a truth of 0.800. The estimate is shrunk twice, and only one of the two shrinkages is under the analyst’s control.

What a window set at the memory actually keeps

The Bartlett weight has a fixed point in it, and the fixed point settles the naive reading in one line. At the outermost lag a window carries, k = L, so

w(L,L)=1LL+1=1L+1.w(L, L) = 1 - \frac{L}{L + 1} = \frac{1}{L + 1}.

A window set exactly at the memory of the errors keeps one part in L + 1 of the lag it was set to capture. Under the moving average, L = 4 keeps a fifth of the fourth lag — the one lag the law’s memory is defined by. The reading “carry the lags that have something in them and stop” is therefore the reading that discards four fifths of the last thing it carries, and it does so at every memory length: a memory of nine keeps a tenth, a memory of nineteen a twentieth. It is not nearly right with a problem at the edge. The edge is the whole of it.

The automatic bandwidth is 4 at n = 120, so under the moving average it happens to be exactly that rule, reached from a different derivation. Its weights across the four live lags are 0.8, 0.6, 0.4 and 0.2, against true autocorrelations of 0.8, 0.6, 0.4 and 0.2, so what reaches the whitening is 0.64 + 0.36 + 0.16 + 0.04 = 1.20 of a true 2.00 — sixty per cent of the dependence there is. At the best window of 30 the weights are 0.968, 0.935, 0.903 and 0.871, and the delivered total is 1.871, or 93.5%. The extra twenty-six lags, every one of them known to be empty, are what buys the last third.

Ten times the memory, not one

Inverting the weight gives the window that any stated fidelity requires. Keeping a fraction f of lag k takes

L=k1f1,L = \frac{k}{1 - f} - 1,

so keeping nine tenths of the last live lag of a memory-q law takes L = 10q − 1. For the moving average that is 39, against a measured optimum of 30, which keeps 87.1%. Keeping four fifths takes 5q − 1, which is 19.

A rule of thumb falls out of the taper alone, with no sweep behind it: multiply the memory by something between five and ten, not by one. The factor of three to eight this field measures is what that multiplication looks like once variance has trimmed it back, and the two other stationary optima sit inside the same band once each law’s effective memory is read off rather than its nominal one — the geometric law’s autocorrelation is under a tenth past the tenth lag, which puts 10q − 1 near a hundred and the sweep’s answer of 12 well short of it, and that shortfall is the variance term doing its work rather than the taper doing less of it.

The bias half, with no sampling in it

Both statements above are about bias, and bias can be computed with no draws at all. Build the covariance the estimate converges to — the taper’s weight times the (1 − k/n) that comes from averaging n − k pairs, times the law’s own autocovariance — factor it, and apply the factor to the true covariance from both sides. If the window were right the result would be the identity, and how far it is from the identity is the whole of the bias.

Where a wider window stops helping, and where it never does. The window's bias, alone, with no sampling anywhere in it: Ω_L is built from the sequence a sample of 120 rows reports on average — the taper's weight times the (1 − k/n) that comes from having n − k pairs at lag k — and its factor is applied to the true covariance from both sides. If the window were right the result would be the identity, and what is plotted is how far it is from one. Under all three stationary laws it falls the whole way, so a wider window would keep improving the whitening if variance were free, and the interior optimum in the sweep beside this one is therefore a fact about variance rather than about fit. Under the break it stops at L = 20 and turns back up: the object being approximated is not a function of the gap between two rows, and no matrix that is can get closer.
Fig. 2 The bias half of the window’s trade under each law, computed in closed form. Three curves fall the whole way; the fourth has a floor.

Under all three stationary laws that quantity falls monotonically, all the way to the end of the range: at sixty lags it is 0.0900 under the geometric law and still falling. So if variance were free the best window would be as wide as the sample allows, and every interior optimum in the sweep above is bought by variance rather than by fit.

That is worth stating plainly because it inverts the usual intuition about bandwidth. The received account of a window is that too many lags fit noise and too few miss the signal — a symmetric trade with a middle. Here one side of the trade is monotone: more lags always describe the dependence better, and what stops it is that each extra lag is a number estimated from n − k pairs.

Under the break the bias curve bottoms at twenty and turns back up, which is the wall that essay is about and is a different failure entirely.

Three ways to choose it, all landing short

Nobody knows the best window. There are three feasible ways to choose one and they are the subject of the rest of this essay.

The automatic bandwidth. The plug-in rule, 4(n/100)2/94(n/100)^{2/9}, gives 4 at n = 120. It costs 50.2% more than the best window under the geometric law, 89.5% more under the moving average, 11.9% more under long memory and 9.9% more under the break.

The reason it is short is not that it is a bad rule. It is derived for a different quantity: a long-run variance is a sum over the band, Σ γ(k) w(k, L), and in a sum the taper’s attenuation is part of what makes the estimator consistent — the weights are chosen so that the bias in the sum vanishes as the sample grows. A whitening does not sum the sequence; it inverts the matrix the sequence builds, and an inverse is sensitive to the shape of what it inverts in a way a sum is not.

The error model’s own likelihood. A tapered Ω̂ at window L is a Gaussian model for the residual series, so it has a likelihood, and charging two per lag makes choosing L an order-selection problem of the ordinary kind. It lands at 5.20 on average under the geometric law, 9.73 under the moving average, 3.53 under long memory and 5.75 under the break — better than the automatic rule everywhere, and still short of the optimum everywhere: 34.9%, 30.7%, 11.5% and 6.2% above the best available.

Three ways to choose a window, and the one that would have been best. Regret under a five-period moving average, over 120 draws. The automatic bandwidth is the plug-in rule a practitioner reaches for, and it is derived to estimate a long-run variance — a sum over the band — where a whitening inverts the matrix that band builds; at n = 120 it is 4, and it costs 89% more than the best window available. Choosing the window by the error model's own likelihood, two per lag, lands at 9.7 on average and costs 31% more. The rule that reads what a whitening actually promises — how flat it leaves the residuals — is not on this chart because it is not a rule: it is monotone in the window and picks the widest one it is offered, on 100% of draws under three of the four laws.
Fig. 3 What each way of choosing is worth, with the law on the slider. The best fixed window is at the top and is not available to anybody; the rule told the form is on the same axis for scale.

And the rule this site’s own premise suggests, which does not work at all, and is worth its own section.

A rule that is monotone in what it is choosing

A whitening makes a promise: it claims the transformed series has no dependence left in it. This collection’s habit is to count what a procedure claims, so the obvious selection rule is to count it — choose the window that leaves the whitened residuals flattest, on a portmanteau statistic over the first dozen lags.

It selects the widest window it is offered, on every draw, under three of the four laws, and on 119 draws of 120 under the fourth.

It is not a close call and it is not noise. Ω̂_L is fitted to the autocorrelations the flatness statistic then scores it on removing, so the statistic falls with L by construction: a window of thirty has thirty numbers tuned to flatten the first thirty lags of this particular sample. Handed candidates up to 8 it picks 8; handed candidates up to 30 it picks 30. A rule whose answer is the last element of the list it was given is not choosing anything.

This is in-sample fit choosing the most complex model, arriving in a place where nobody looks for it — not in the model being selected, but in the nuisance model underneath it, where there is no penalty term because nobody thought of the window as a fitted object. It is refused here, and the refusal is the reason the likelihood rule above carries a charge of two per lag.

Choosing costs nothing; choosing that value costs a third

There are two ways for a data-driven window to lose, and they are worth separating because only one of them is usually named. It can lose by variance — the window wobbles from draw to draw, and a rule with a moving tuning parameter is worse than the same rule with a fixed one at the same average. Or it can lose by location — it settles reliably on the wrong number.

Under the geometric law the likelihood rule chooses 5.20 on average and reads 0.02097. Fixed windows at 4 and 8 read 0.02336 and 0.01672, and a straight interpolation between them at 5.20 gives 0.0212. Within the noise, a window that moves from draw to draw is worth exactly what a fixed window at its own average is worth.

So the whole of the loss is location. Choosing the window on the sample costs nothing for the act of choosing; it costs 34.9% because it chooses about five when twelve was available. A rule that does not look at the data at all — take L = 12 and stop — beats every rule here that does. That is not an argument against data-driven tuning in general; it is a measurement of what these particular criteria are aimed at, which is fit rather than whitening.

How much the window is worth deciding

A tuning parameter deserves attention in proportion to what it moves, and this one moves a different amount in each world. Measured as the span of the sweep — the worst window in the range against the best — as a share of everything there was to recover, the window is worth 49.8% under the geometric law, 45.2% under the moving average, 35.6% under long memory and 20.4% under the break.

So the window is roughly half of the whole question where the dependence has an edge to be captured, and a fifth of it where the problem is that the covariance is not stationary at all. Under the break, an analyst who spent a week on the bandwidth and none on whether the persistence changes has optimised the smaller of the two by a wide margin.

What the measurement supports as a working rule is narrow and worth stating narrowly: at n = 120, on this table, a window between a tenth and a third of the sample beat every data-driven choice in all four worlds. It is not a theorem and it is not measured at other sample sizes — the automatic rule’s own n2/9n^{2/9} says the right answer grows slowly with n, and a fixed fraction of the sample grows linearly, so the two must cross somewhere this field has not looked. What can be said is that at the sample size a hundred and twenty rows describes, the standard choice is between three and eight times too small, and the error is in the same direction every time.

That direction is the useful part. A window that is too wide costs variance and the cost is shallow — the sweeps above rise by a tenth between their optimum and sixty lags. A window that is too narrow removes the dependence before the whitening sees it, and that cost is steep: at four lags the rule has given up between a fifth and a half of the repair. The two errors are not symmetric, and a rule derived to be optimal for a different target should be expected to sit on the wrong side of an asymmetric loss rather than in the middle of it.

What is claimed here, and what is not

This essay takes what window a whitening wants and how to choose it. The claims are that the best fixed window is 12, 30, 30 and 40 in the four worlds, all far past the memory of the errors and including a world with no dependence past its fourth lag; that the mechanism is the taper, which keeps 0.556 of the fourth lag at L = 8 and 0.871 at L = 30; that the bias half of the trade falls monotonically in every stationary world, so the interior optimum is bought by variance; that the automatic bandwidth is 4 and costs between 9.9% and 89.5% more than the best available; that the likelihood rule lands between 3.53 and 9.73 and costs between 6.2% and 34.9%; and that a window chosen by the flatness of the whitened residuals is monotone in the window and selects the end of whatever list it is shown.

What stays out and is named as a decision: a window chosen for each candidate rather than once. The estimate here is made once, from the fullest candidate’s residuals, before any selection, because a nuisance that is shared cancels out of the differences a criterion reads and a nuisance estimated per candidate does not. That argument was made for the estimate and it is assumed here for the window, which is a different object; a window chosen per candidate is a selection problem inside a selection problem and it is not measured.

The other omission is a penalty for having chosen. The criterion reads a window selected from the same hundred and twenty rows, and nothing in the criterion charges for it. The measurement above says the charge would be small — choosing costs nothing beyond where it chooses — but small is not zero and this field does not price it.

The boundary against the essay that introduced the window is that it measured one world and asked what the window costs at both ends of its own dial; this one asks what the window is, across four worlds, and finds it is not the memory.

The checks, and the refusals

Two claims are gated. The bias-only mismatch is required to fall monotonically under every stationary law, which is the assertion that the interior optimum is a variance effect; and the best window found by the sweep is required to be at least eight lags in every world, which is the assertion this essay is named for.

Two refusals. The flatness rule is rejected for being monotone: it is required to pick the top of the list on more than nine draws in ten and to change its answer when the list is extended, and it does both. And the automatic bandwidth is rejected as a window for a whitening, at a measured 89.5% above the best available under the moving average — not because the formula is wrong, but because it is a correct answer to the question of how to estimate a sum.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A family before a fit — both name autocorrelation, covariance matrix, generalised least squares, long memory, model selection, nuisance parameter, whitening
  • How often it matters — both name bandwidth selection, information criterion, long memory, model selection, regret, whitening
  • The comparison that was not made — both name covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
  • The volume a whitening moves — both name covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
  • What fitting them together buys — both name covariance matrix, generalised least squares, long memory, model selection, nuisance parameter, whitening
  • A charge that reads the draw — both name bandwidth selection, information criterion, model selection, regret, sample autocovariance

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationBandwidth selectionBias-varianceCovariance matrixGeneralised least squaresInformation criterionLong memoryModel diagnosticsModel selectionMoving averageNuisance parameterRegretSample autocovarianceWhitening