The shape a dependence has

The order the tail is drawn at

A fitted autoregression reproduces the sample exactly at the lags it was fitted on, so everything it says past them is extrapolation — and the order is the dial that decides how much of it there is.

Worth reading first: Choosing the order · The observations that repeat each other.

A window carries the lags it was given and zero everywhere else. A fitted autoregression carries an autocovariance at every lag, including the ones past the end of its own order, because a model is a rule for continuing rather than a table of numbers.

That is the whole of what a sieve has over a window, and it is why the construction exists in this collection at all: a resampling built from a fitted model is the one construction not bounded by what the residuals themselves report. The order is the dial that decides how far the continuation is extrapolated from, and this essay is about what it is worth and what choosing it costs.

Everything a fitted model says is past its own order

The Yule–Walker equations are solved so that the fitted model reproduces the sample’s own autocorrelations out to lag p. That is not an approximation and it is not a coincidence: it is what the equations are.

So an AR(4) agrees with the series it was fitted to at the first four lags, exactly, and disagrees with it everywhere else — which means a comparison drawn inside the order is a comparison of a thing with itself, and the only content in the fitted sequence is the tail.

Four sequences, and the rule only ever sees the last oneUnder long memory at d = 4/9, four things that are all called the dependence. The law itself is the top line. What a sample of 120 rows reports on average is the second, computed exactly: subtracting a sample mean takes the first lag from 0.800 to 0.538. What a candidate's residuals report is the third, lower again at 0.472, because a fit removes dependence along with signal. The autoregressions are fitted to that third sequence and reproduce it exactly out to their own order — the Yule–Walker equations are solved to make it so — so everything they say past that is extrapolation. At the twentieth lag the law has 0.576, the residuals report 0.006, and an AR(8) extrapolates 0.028.00.50011248122030lagautocorrelationthe lawwhat 120 rows reportwhat the residuals reportAR(1)AR(2)AR(4)AR(8)200 draws, n = 120, 2 lags inside the fitpast the order it is all extrapolation
Fig. 1 Four sequences that all get called the dependence, with the law on the slider. Each fitted order agrees with the residuals up to its own p and continues on its own from there.

Four sequences, and the fit sees the last one

Under long memory there are four different quantities in that picture and the distance between them is the subject.

The law says 0.800 at the first lag. What a sample of a hundred and twenty rows reports, computed exactly, is 0.538 — the arithmetic of subtracting a sample mean from a series that barely has one. What a candidate’s residuals report is 0.472, lower again, because a fit removes dependence along with signal and what remains has had the persistent part of the design taken out of it.

The autoregression is fitted to that fourth sequence. So the extrapolation is anchored at a point that is already 40% short of the truth before any modelling decision has been made, and the modelling decision can only make the shape wrong, not the level right.

At the twentieth lag the law has 0.5758, the residuals report 0.006, and an AR(8) extrapolates 0.0284 — wrong by a factor of 20. It is worth being precise about which part of that is the model’s fault: the residuals had already lost the tail, and the AR(8) is putting a little of it back.

And under a moving average it invents one

The failure reverses when the truth stops. A five-period moving average has exactly nothing past its fourth lag, and an AR(1) fitted to it extrapolates 0.301 at the fourth lag and 0.097 at the eighth, where the truth is zero and stays zero. A first-order model has no way to represent an autocorrelation that ends; the only sequence it can continue with is a geometric one.

Raising the order fixes it, and the fix is visible in the same picture: an AR(4) reads 0.101 at the fourth lag — the residuals’ own value, because four lags is where it was fitted — and −0.042 at the eighth, which is the sample’s noise rather than an invented tail.

So the order is not a smoothing parameter with a bias in one direction. Too low invents a tail where there is none and too low also loses one where there is — under long memory an AR(1) reads 0.007 at the eighth lag where the residuals themselves report 0.108. Which error a low order makes depends on the world, and the two look identical from inside the fit.

Which attenuation dominates depends on the law

The first lag loses 0.262 to the sample mean and a further 0.066 to the fit — so under long memory the centring is four times the design.

That is the reverse of what the same decomposition gives under a first-order autoregression, where the exact identity puts the centring at 0.0227 and the design’s four extra columns at 0.0435: the design costs about twice the intercept there, and about a quarter of it here.

As shares it is starker. Centring costs 33% of the law’s first lag under long memory and 3% under the autoregression — a factor of eleven — while the design costs 12% and 6%, which are within a factor of two of each other.

The reason is the whole of what long memory is. A process whose autocorrelation is still 0.58 at the twentieth lag has an enormous amount of its variation at low frequency, and a sample mean is the lowest-frequency thing there is; subtracting it removes a third of the first lag before any model is fitted. An autoregression at 0.8 has far less down there, so its mean is a much smaller thing to subtract.

So the same two attenuations swap ranks between two laws, and a rule tuned to the one where the design dominates is tuned for the wrong half of the problem in the other.

Four dependences a single parameter cannot tell apart. Every law here is standardised to a lag-one autocorrelation of 0.8, so a rule told the errors are a first-order autoregression finds the same number in all four and has no way of seeing what separates them. The geometric decay is the world in which estimating a covariance rather than naming it was priced, and found to cost. The five-period moving average has 0.200 at the fourth lag and exactly nothing past it, where the geometric law says 0.328 at the fifth. Long memory at d = 4/9 is still at 0.576 by the twentieth lag, where the geometric law has reached 0.012. The break has no autocorrelation function at all: what is drawn for it is the average over the pairs at each gap, which is what a stationary estimate converges to.
Fig. 2 The four laws, standardised to a lag-one autocorrelation of 0.8, so a rule told the errors are a first-order autoregression finds the same number in all four. The five-period moving average has 0.200 at the fourth lag and exactly nothing past it where the geometric law says 0.328 at the fifth; long memory is still at 0.576 by the twentieth.

The order buys five orders of magnitude and is still twenty short

The twentieth lag is where the extrapolation is doing all of its work, and splitting the factor of twenty from what it started with says how much the model actually delivers.

The truth is 0.5758. The residuals report 0.006 — ninety-six times short. The AR(8) extrapolates 0.0284, which is 4.7 times the residuals’ reading and 20 times short of the truth. So the ninety-six factors as 4.7 × 20.4: the model recovers a fifth of the distance in logarithms and the anchor keeps the rest.

The extra orders matter more than that ratio suggests. An AR(1) fitted at the residuals’ first lag of 0.472 would extrapolate 0.47220=3×1070.472^{20} = 3\times10^{-7} at the twentieth lag. So the seven extra parameters take the extrapolation from three ten-millionths to 0.0284 — a factor of about ninety thousand — and it is still twenty times too small.

That is the honest reading of what an order buys. It is not a modest improvement on a low-order fit; it is nearly the whole distance, in a problem where the whole distance is five orders of magnitude and the remaining factor of twenty is set by a number the model was never allowed to choose.

What the order is worth

What the order is worth, and what choosing it costs. Regret under a five-period moving average at each fixed order, over 200 draws. At p = 1 the rule is exactly the first-order autoregression the parameterised rule fits, which is why the two agree there to five places; the best order available is p = 12 at 0.01586. Choosing the order on the sample by the criterion that is standard for it — the residuals' own likelihood with two per lag — gives 0.01762, and picks 7.50 on average. The gap between the chosen order and the best one is the price of the third selection problem this rule carries: the criterion selects a model, and the criterion's error model selects an order, and both read the same hundred and twenty rows.
Fig. 3 Regret at each fixed order, with the law on the slider. At p = 1 the rule is the parameterised one, exactly.

Under the geometric law the best order is 1, at 0.01294, and every higher order is worse: 0.01406 at two, 0.01560 at four, 0.01937 at twelve. Each extra coefficient is a number estimated from a hundred and twenty rows and paid for in the whitening.

Under the moving average the ordering reverses completely. The first order reads 0.02935 and the twelfth 0.01586 — a difference larger than anything else in this field — because an autocorrelation that stops abruptly is an infinite autoregression, and the fitted order is how much of that infinity is kept.

Under long memory the best available order is 12 at 0.04753 against 0.05319 at one, and under the break the whole dial is worth 0.00299 from end to end, which is nothing: the order cannot repair a covariance that changes, and the flatness of that sweep is the same statement the window’s sweep makes there.

The order chosen from the sample

The standard way to choose it is the residuals’ own likelihood with two per coefficient, minimised over p — the criterion that is standard for order selection, applied to the nuisance model rather than to the model of interest.

Under the geometric law it picks 1 on 152 of 200 draws and reads 0.01444 against the best fixed order’s 0.01294. Under the moving average it averages 7.50 and reads 0.01762 against 0.01586, giving up 13.1% of what the order had to give. Both are the criterion working.

Under long memory it picks one or two on 136 of 200 draws, where twelve is the best available, and reads 0.05023 against 0.04753: it gives up 47.7% of what the order had to give.

The reason it lands short is the same reason the automatic bandwidth does. A likelihood is a sum over pairs of rows, and the lags with the most pairs are the short ones, so a criterion built on it is dominated by a part of the sequence every candidate order already gets right. The tail — the only thing the order actually controls — enters the likelihood with almost no weight at all. The order is chosen for the part of the fit that does not depend on it.

Why an order is cheaper than a window

The two general constructions in this field estimate the same object and pay for it differently. A tapered window at twenty lags carries twenty numbers, and the number at lag k is an average over the n − k pairs at that gap — so the last of them is estimated from a hundred pairs and the estimate at every lag is a separate quantity with its own noise. An autoregression of order six carries six numbers, each of them a solution of a linear system that reads every row of the sample, and the whole tail follows from those six by a recursion.

The measurement agrees. Against the tapered window at twenty lags the fitted model is ahead by 0.00474 under the geometric law, 0.00342 under the moving average and 0.00116 under the break, at 3.4, 3.6 and 0.8 paired standard errors, and the two are tied to five places under long memory. Three wins and a tie, with no world in which the window is ahead.

It is the same economy that makes a penalty count parameters rather than rows: what is expensive is not how much structure a model describes but how many independent numbers had to be estimated to describe it. A model that continues is cheaper than a table that stops, provided the continuation is roughly the right shape — and the rest of this essay is about what happens when it is not.

The window a whitening wants is not the memory of the errors. Regret under AR(1) at 0.8 as the tapered estimate is given more lags, over 120 draws at n = 120. The best window is L = 12; the automatic bandwidth is 4 and the error model's own likelihood chooses 5.2 on average. Both land in the same place and both are short, and the reason is the taper: a Bartlett weight at lag k is 1 − k/(L + 1), so a window of 8 keeps 0.556 of whatever the fourth lag carries and a window of 30 keeps 0.871. A window has to be several times the memory before it stops removing the memory. The dashed line is the rule told the errors are a first-order autoregression, which needs no window at all.
Fig. 4 Regret against the number of lags a tapered estimate is given, at ρ = 0.8 over 120 draws. The best window is 12 and the automatic bandwidth is 4: a Bartlett weight at lag k is 1 − k/(L + 1), so a window of eight keeps 0.556 of whatever the fourth lag carries and a window of thirty keeps 0.871.

Two ways to land short, separated by one measurement

There are two explanations for a criterion that picks two when twelve was available, and they have different repairs. The criterion may be aimed at the wrong part of the sequence. Or it may be reading a series that has already lost the dependence — the residuals, which are not the errors and report 0.472 at the first lag where the law says 0.800.

The world is simulated, so the errors themselves are available and the two explanations can be told apart by running the same criterion on both.

Under long memory the criterion picks 2.46 on the errors and 2.33 on the residuals, a difference of 0.13 against standard errors of about 0.11. Handed a perfect series it makes the same choice. The residuals are not the problem there; the criterion is doing exactly what it is built to do, on a sequence whose short lags are informative and whose tail is not, and the tail is the only thing the order controls.

Under the moving average the same comparison reads 9.38 on the errors and 7.59 on the residuals, a difference of 1.79 at ten standard errors. There the attenuation costs nearly two orders, and it costs them in the world where the order is worth the most.

So the two failures are different failures. Under long memory the repair would have to be a different criterion — one that weights the tail, which is to say one that already knows the tail matters. Under the moving average the repair is upstream: fit the error model to something that has not had a regression taken out of it, or fit both at once.

The check that says the general rule contains the special one

One number in the table above is not a measurement but an identity. At p = 1 the sieve reads 0.01294018 and the rule told the errors are a first-order autoregression reads 0.01294018: the paired difference across two hundred draws is 6·10⁻¹⁷, which is zero.

They share no arithmetic. One applies the Prais–Winsten transform at ρ^=etet1/et2\hat\rho = \sum e_t e_{t-1} / \sum e_t^2; the other fits an AR(1) by Yule–Walker, builds the model’s whole autocovariance sequence from the coefficient, forms the dense covariance matrix and factors it. That they agree to machine precision is a check on both, and the reason they must is worth having: a fit with an intercept has residuals whose mean is exactly zero, so the centred autocovariance the Yule–Walker recursion uses and the uncentred ratio the parameterised rule uses are the same number.

A generalisation that does not reduce to the case it generalises is not a generalisation. Nothing else in this field would have noticed if it had stopped doing so.

The extrapolation is wrong and the whitening barely notices

Here is the awkward part, and it is the reason this essay does not end with a warning.

The sieve’s extrapolation under long memory is wrong by a factor of twenty at the twentieth lag. The sieve is also, in that world, tied with the tapered window for the best feasible rule in the table — 0.05023 against 0.05024 — and it beats the window under the geometric law and the moving average. A construction whose tail is wrong by twenty times is not being punished for it.

The explanation is that a whitening reads a covariance matrix, and the entries of that matrix near the diagonal outnumber the ones far from it. A hundred and twenty rows have 119 pairs at lag one and 100 at lag twenty; the transform is dominated by the short lags, which is exactly where the fitted model is reproducing the sample rather than extrapolating. The same counting is why the window’s own dial is worth what it is: both tuning parameters are read through a matrix whose mass is near its diagonal.

So the tail’s wrongness costs little in this use, and that is a statement about the use rather than about the tail. The use where it would bite is the one the sieve was introduced for: a resampling that generates errors from the fitted model carries the extrapolated sequence into every draw, so a reference distribution built that way inherits a tail that is wrong by a factor of twenty rather than a matrix that is mostly right. The same object, read for a different purpose, has a completely different error.

What is claimed here, and what is not

This essay takes what an autoregression’s order is worth to a whitening, and what choosing it costs. The claims are that a Yule–Walker fit reproduces the sample’s autocorrelations exactly out to its own order, so the only content past that is extrapolation; that the extrapolation invents a tail under a moving average — 0.097 at the eighth lag where the truth is zero — and loses one under long memory, at 0.0284 against 0.5758 at the twentieth; that the best fixed order is 1, 12, 12 and 6 in the four worlds; that the residuals’ own likelihood picks one or two on 136 of 200 long-memory draws and gives up 47.7% of what the order had to give; and that the sieve at order one is the parameterised rule to machine precision.

What stays out and is named as a decision: the order chosen per candidate. As with the window, one order is chosen once from the fullest candidate’s residuals, before any selection, and the case for that was made about an estimated covariance rather than about a tuning parameter. The order is also chosen with no charge for having chosen it, which is the same omission the window carries and is stated there.

And the extrapolation is priced here only as an input to a whitening. What it is worth as a reference distribution is measured elsewhere, on a construction where the tail is not a small part of the answer, and the two measurements should not be read as one.

The checks, and the refusals

Two claims are gated. Every fitted order is required to reproduce the residuals’ own autocorrelations out to its own p, to 10⁻⁸, which is the identity the whole essay rests on; and the sieve at order one is required to equal the parameterised rule to 10⁻¹², which is the check that the general machinery contains the special case.

Two refusals. A fitted autoregression’s extrapolation is rejected as an estimate of the dependence: at the twentieth lag under long memory it reports 0.0284 where the law has 0.5758, and a construction that carries that sequence has confirmed a model with the model’s own assumption. And an order chosen by the residuals’ likelihood is rejected where the tail is the target, at a measured 47.7% of the available gain — not because the criterion is wrong, but because it is aimed at a part of the sequence the order does not control.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A family before a fit — both name autocorrelation, generalised least squares, long memory, model selection, nuisance parameter, whitening
  • A list is not a rule — both name autoregression, information criterion, model selection, nuisance parameter, regret, whitening
  • The comparison that was not made — both name autoregression, information criterion, model selection, nuisance parameter, regret, whitening
  • The volume a whitening moves — both name autoregression, information criterion, model selection, nuisance parameter, regret, whitening
  • A covariance with no parameter in it — both name autocorrelation, generalised least squares, information criterion, model selection, nuisance parameter
  • How often it matters — both name information criterion, long memory, model selection, regret, whitening

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationAutoregressionExtrapolationGeneralised least squaresInformation criterionLong memoryModel selectionMoving averageNuisance parameterOrder-selectionRegretSample autocovarianceWhitening