The order the tail is drawn at
Worth reading first: Choosing the order · The observations that repeat each other.
A window carries the lags it was given and zero everywhere else. A fitted autoregression carries an autocovariance at every lag, including the ones past the end of its own order, because a model is a rule for continuing rather than a table of numbers.
That is the whole of what a sieve has over a window, and it is why the construction exists in this collection at all: a resampling built from a fitted model is the one construction not bounded by what the residuals themselves report. The order is the dial that decides how far the continuation is extrapolated from, and this essay is about what it is worth and what choosing it costs.
Everything a fitted model says is past its own order
The Yule–Walker equations are solved so that the fitted model reproduces the sample’s own autocorrelations out to lag p. That is not an approximation and it is not a coincidence: it is what the equations are.
So an AR(4) agrees with the series it was fitted to at the first four lags, exactly, and disagrees with it everywhere else — which means a comparison drawn inside the order is a comparison of a thing with itself, and the only content in the fitted sequence is the tail.
Four sequences, and the fit sees the last one
Under long memory there are four different quantities in that picture and the distance between them is the subject.
The law says 0.800 at the first lag. What a sample of a hundred and twenty rows reports, computed exactly, is 0.538 — the arithmetic of subtracting a sample mean from a series that barely has one. What a candidate’s residuals report is 0.472, lower again, because a fit removes dependence along with signal and what remains has had the persistent part of the design taken out of it.
The autoregression is fitted to that fourth sequence. So the extrapolation is anchored at a point that is already 40% short of the truth before any modelling decision has been made, and the modelling decision can only make the shape wrong, not the level right.
At the twentieth lag the law has 0.5758, the residuals report 0.006, and an AR(8) extrapolates 0.0284 — wrong by a factor of 20. It is worth being precise about which part of that is the model’s fault: the residuals had already lost the tail, and the AR(8) is putting a little of it back.
And under a moving average it invents one
The failure reverses when the truth stops. A five-period moving average has exactly nothing past its fourth lag, and an AR(1) fitted to it extrapolates 0.301 at the fourth lag and 0.097 at the eighth, where the truth is zero and stays zero. A first-order model has no way to represent an autocorrelation that ends; the only sequence it can continue with is a geometric one.
Raising the order fixes it, and the fix is visible in the same picture: an AR(4) reads 0.101 at the fourth lag — the residuals’ own value, because four lags is where it was fitted — and −0.042 at the eighth, which is the sample’s noise rather than an invented tail.
So the order is not a smoothing parameter with a bias in one direction. Too low invents a tail where there is none and too low also loses one where there is — under long memory an AR(1) reads 0.007 at the eighth lag where the residuals themselves report 0.108. Which error a low order makes depends on the world, and the two look identical from inside the fit.
Which attenuation dominates depends on the law
The first lag loses 0.262 to the sample mean and a further 0.066 to the fit — so under long memory the centring is four times the design.
That is the reverse of what the same decomposition gives under a first-order autoregression, where the exact identity puts the centring at 0.0227 and the design’s four extra columns at 0.0435: the design costs about twice the intercept there, and about a quarter of it here.
As shares it is starker. Centring costs 33% of the law’s first lag under long memory and 3% under the autoregression — a factor of eleven — while the design costs 12% and 6%, which are within a factor of two of each other.
The reason is the whole of what long memory is. A process whose autocorrelation is still 0.58 at the twentieth lag has an enormous amount of its variation at low frequency, and a sample mean is the lowest-frequency thing there is; subtracting it removes a third of the first lag before any model is fitted. An autoregression at 0.8 has far less down there, so its mean is a much smaller thing to subtract.
So the same two attenuations swap ranks between two laws, and a rule tuned to the one where the design dominates is tuned for the wrong half of the problem in the other.
The order buys five orders of magnitude and is still twenty short
The twentieth lag is where the extrapolation is doing all of its work, and splitting the factor of twenty from what it started with says how much the model actually delivers.
The truth is 0.5758. The residuals report 0.006 — ninety-six times short. The AR(8) extrapolates 0.0284, which is 4.7 times the residuals’ reading and 20 times short of the truth. So the ninety-six factors as 4.7 × 20.4: the model recovers a fifth of the distance in logarithms and the anchor keeps the rest.
The extra orders matter more than that ratio suggests. An AR(1) fitted at the residuals’ first lag of 0.472 would extrapolate at the twentieth lag. So the seven extra parameters take the extrapolation from three ten-millionths to 0.0284 — a factor of about ninety thousand — and it is still twenty times too small.
That is the honest reading of what an order buys. It is not a modest improvement on a low-order fit; it is nearly the whole distance, in a problem where the whole distance is five orders of magnitude and the remaining factor of twenty is set by a number the model was never allowed to choose.
What the order is worth
Under the geometric law the best order is 1, at 0.01294, and every higher order is worse: 0.01406 at two, 0.01560 at four, 0.01937 at twelve. Each extra coefficient is a number estimated from a hundred and twenty rows and paid for in the whitening.
Under the moving average the ordering reverses completely. The first order reads 0.02935 and the twelfth 0.01586 — a difference larger than anything else in this field — because an autocorrelation that stops abruptly is an infinite autoregression, and the fitted order is how much of that infinity is kept.
Under long memory the best available order is 12 at 0.04753 against 0.05319 at one, and under the break the whole dial is worth 0.00299 from end to end, which is nothing: the order cannot repair a covariance that changes, and the flatness of that sweep is the same statement the window’s sweep makes there.
The order chosen from the sample
The standard way to choose it is the residuals’ own likelihood with two per coefficient, minimised over p — the criterion that is standard for order selection, applied to the nuisance model rather than to the model of interest.
Under the geometric law it picks 1 on 152 of 200 draws and reads 0.01444 against the best fixed order’s 0.01294. Under the moving average it averages 7.50 and reads 0.01762 against 0.01586, giving up 13.1% of what the order had to give. Both are the criterion working.
Under long memory it picks one or two on 136 of 200 draws, where twelve is the best available, and reads 0.05023 against 0.04753: it gives up 47.7% of what the order had to give.
The reason it lands short is the same reason the automatic bandwidth does. A likelihood is a sum over pairs of rows, and the lags with the most pairs are the short ones, so a criterion built on it is dominated by a part of the sequence every candidate order already gets right. The tail — the only thing the order actually controls — enters the likelihood with almost no weight at all. The order is chosen for the part of the fit that does not depend on it.
Why an order is cheaper than a window
The two general constructions in this field estimate the same object and pay for it differently. A tapered window at twenty lags carries twenty numbers, and the number at lag k is an average over the n − k pairs at that gap — so the last of them is estimated from a hundred pairs and the estimate at every lag is a separate quantity with its own noise. An autoregression of order six carries six numbers, each of them a solution of a linear system that reads every row of the sample, and the whole tail follows from those six by a recursion.
The measurement agrees. Against the tapered window at twenty lags the fitted model is ahead by 0.00474 under the geometric law, 0.00342 under the moving average and 0.00116 under the break, at 3.4, 3.6 and 0.8 paired standard errors, and the two are tied to five places under long memory. Three wins and a tie, with no world in which the window is ahead.
It is the same economy that makes a penalty count parameters rather than rows: what is expensive is not how much structure a model describes but how many independent numbers had to be estimated to describe it. A model that continues is cheaper than a table that stops, provided the continuation is roughly the right shape — and the rest of this essay is about what happens when it is not.
Two ways to land short, separated by one measurement
There are two explanations for a criterion that picks two when twelve was available, and they have different repairs. The criterion may be aimed at the wrong part of the sequence. Or it may be reading a series that has already lost the dependence — the residuals, which are not the errors and report 0.472 at the first lag where the law says 0.800.
The world is simulated, so the errors themselves are available and the two explanations can be told apart by running the same criterion on both.
Under long memory the criterion picks 2.46 on the errors and 2.33 on the residuals, a difference of 0.13 against standard errors of about 0.11. Handed a perfect series it makes the same choice. The residuals are not the problem there; the criterion is doing exactly what it is built to do, on a sequence whose short lags are informative and whose tail is not, and the tail is the only thing the order controls.
Under the moving average the same comparison reads 9.38 on the errors and 7.59 on the residuals, a difference of 1.79 at ten standard errors. There the attenuation costs nearly two orders, and it costs them in the world where the order is worth the most.
So the two failures are different failures. Under long memory the repair would have to be a different criterion — one that weights the tail, which is to say one that already knows the tail matters. Under the moving average the repair is upstream: fit the error model to something that has not had a regression taken out of it, or fit both at once.
The check that says the general rule contains the special one
One number in the table above is not a measurement but an identity. At p = 1 the sieve reads 0.01294018 and the rule told the errors are a first-order autoregression reads 0.01294018: the paired difference across two hundred draws is 6·10⁻¹⁷, which is zero.
They share no arithmetic. One applies the Prais–Winsten transform at ; the other fits an AR(1) by Yule–Walker, builds the model’s whole autocovariance sequence from the coefficient, forms the dense covariance matrix and factors it. That they agree to machine precision is a check on both, and the reason they must is worth having: a fit with an intercept has residuals whose mean is exactly zero, so the centred autocovariance the Yule–Walker recursion uses and the uncentred ratio the parameterised rule uses are the same number.
A generalisation that does not reduce to the case it generalises is not a generalisation. Nothing else in this field would have noticed if it had stopped doing so.
The extrapolation is wrong and the whitening barely notices
Here is the awkward part, and it is the reason this essay does not end with a warning.
The sieve’s extrapolation under long memory is wrong by a factor of twenty at the twentieth lag. The sieve is also, in that world, tied with the tapered window for the best feasible rule in the table — 0.05023 against 0.05024 — and it beats the window under the geometric law and the moving average. A construction whose tail is wrong by twenty times is not being punished for it.
The explanation is that a whitening reads a covariance matrix, and the entries of that matrix near the diagonal outnumber the ones far from it. A hundred and twenty rows have 119 pairs at lag one and 100 at lag twenty; the transform is dominated by the short lags, which is exactly where the fitted model is reproducing the sample rather than extrapolating. The same counting is why the window’s own dial is worth what it is: both tuning parameters are read through a matrix whose mass is near its diagonal.
So the tail’s wrongness costs little in this use, and that is a statement about the use rather than about the tail. The use where it would bite is the one the sieve was introduced for: a resampling that generates errors from the fitted model carries the extrapolated sequence into every draw, so a reference distribution built that way inherits a tail that is wrong by a factor of twenty rather than a matrix that is mostly right. The same object, read for a different purpose, has a completely different error.
What is claimed here, and what is not
This essay takes what an autoregression’s order is worth to a whitening, and what choosing it costs. The claims are that a Yule–Walker fit reproduces the sample’s autocorrelations exactly out to its own order, so the only content past that is extrapolation; that the extrapolation invents a tail under a moving average — 0.097 at the eighth lag where the truth is zero — and loses one under long memory, at 0.0284 against 0.5758 at the twentieth; that the best fixed order is 1, 12, 12 and 6 in the four worlds; that the residuals’ own likelihood picks one or two on 136 of 200 long-memory draws and gives up 47.7% of what the order had to give; and that the sieve at order one is the parameterised rule to machine precision.
What stays out and is named as a decision: the order chosen per candidate. As with the window, one order is chosen once from the fullest candidate’s residuals, before any selection, and the case for that was made about an estimated covariance rather than about a tuning parameter. The order is also chosen with no charge for having chosen it, which is the same omission the window carries and is stated there.
And the extrapolation is priced here only as an input to a whitening. What it is worth as a reference distribution is measured elsewhere, on a construction where the tail is not a small part of the answer, and the two measurements should not be read as one.
The checks, and the refusals
Two claims are gated. Every fitted order is required to reproduce the residuals’ own autocorrelations out to its own p, to 10⁻⁸, which is the identity the whole essay rests on; and the sieve at order one is required to equal the parameterised rule to 10⁻¹², which is the check that the general machinery contains the special case.
Two refusals. A fitted autoregression’s extrapolation is rejected as an estimate of the dependence: at the twentieth lag under long memory it reports 0.0284 where the law has 0.5758, and a construction that carries that sequence has confirmed a model with the model’s own assumption. And an order chosen by the residuals’ likelihood is rejected where the tail is the target, at a measured 47.7% of the available gain — not because the criterion is wrong, but because it is aimed at a part of the sequence the order does not control.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A family before a fit — both name autocorrelation, generalised least squares, long memory, model selection, nuisance parameter, whitening
- A list is not a rule — both name autoregression, information criterion, model selection, nuisance parameter, regret, whitening
- The comparison that was not made — both name autoregression, information criterion, model selection, nuisance parameter, regret, whitening
- The volume a whitening moves — both name autoregression, information criterion, model selection, nuisance parameter, regret, whitening
- A covariance with no parameter in it — both name autocorrelation, generalised least squares, information criterion, model selection, nuisance parameter
- How often it matters — both name information criterion, long memory, model selection, regret, whitening
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationAutoregressionExtrapolationGeneralised least squaresInformation criterionLong memoryModel selectionMoving averageNuisance parameterOrder-selectionRegretSample autocovarianceWhitening