A dependence with a shape
Worth reading first: The observations that repeat each other · A design is a number.
The whitening that repairs an information criterion can be told the dependence is a first-order autoregression and left to find one number, or it can be left to estimate the whole covariance. On errors that really are a first-order autoregression, being left to estimate it costs 0.00672 of regret — the price of a generality the design does not reward.
That number is a cost with no benefit beside it, and it was reported as one. Every world it was measured in was a world the parameterisation was true in, so the comparison could only run one way. What is missing is the other half: a dependence the parameter cannot represent, and the same table, and the same measurement.
Here are four.
Four laws with the same first lag
Everything below is standardised twice: to unit variance, and to a lag-one autocorrelation of 0.8. The second is what makes the comparison mean anything. A rule told the errors are a first-order autoregression has one number to find, it finds it from the first lag, and it therefore finds the same number in all four of these worlds. Whatever separates them is invisible to it by construction.
The geometric law is ρ^k at ρ = 0.8: the world in which estimating a covariance rather than naming it was priced, and the only one of the four in which the parameterisation is true.
The moving average is an error that is a five-period average of white noise, so its autocorrelation is the triangle (5 − k)/5: 0.8, 0.6, 0.4, 0.2, and then exactly nothing. At the fifth lag the geometric law says 0.328 and this one says zero, and no ρ makes a geometric sequence stop.
Long memory is ARFIMA(0, d, 0) at d = 4/9, whose autocorrelation is Γ(k + d)Γ(1 − d)/(Γ(k + 1 − d)Γ(d)) and decays like . At the twentieth lag it is still 0.576, where the geometric law has reached 0.012. It is the opposite failure of the same parameter: a tail that will not end rather than one that ends abruptly.
The break is not a shape at all. The persistence is 0.95 for the first sixty rows and 0.65 for the rest, so the covariance of a pair depends on where the pair sits and not only on the gap between them. Averaged over the pairs at each gap it reports 0.7987 at the first lag, which is what a lag-one estimate reads — and the average of 0.95^k and 0.65^k is not 0.8^k and is not any ρ^k at all.
How fast two laws with one lag in common come apart
Matching at the first lag sounds like a strong constraint and it is worth seeing how weak it is by the third.
The geometric law runs 0.8, 0.64, 0.512, 0.410, 0.328. The moving average runs 0.8, 0.6, 0.4, 0.2, 0. As a share of the geometric’s own value, the gap between them is
0% at the first lag, 6.3% at the second, 21.9% at the third, 51.2% at the fourth and 100% at the fifth.
So the agreement the standardisation buys survives one lag. By the third the two laws disagree about a fifth of the dependence and by the fifth one of them has none left.
And what that does to the quantity most rules are after
Summed, the difference is larger than any single lag suggests.
The geometric law’s long-run variance is , which is 9.0. The moving average’s is , which is 5.0.
Two laws with an identical first lag, and long-run variances a factor of 1.8 apart. Since a whitening, a block resample and a criterion’s penalty are all ultimately statements about that sum, a rule that reads only the first lag is not making a small error on the quantity it needs — it is out by eighty per cent while being exactly right on the number it looked at.
Two different best first-order fits for the same law
That gap has a consequence worth stating, because it says the parameterisation’s failure is not simply that it is too rigid.
Fitting a first-order autoregression to the moving average by matching the first lag gives ρ = 0.8. Fitting it by matching the long-run variance — solving — gives .
Two answers, twenty per cent apart, for the same law and the same family, differing only in which feature the fit is asked to reproduce.
The rule under test matches the first lag, because that is what a lag-one estimate does. Nothing about the family forces that choice, and the alternative would be right about the sum and wrong about the first lag instead. A one-parameter fit to a law outside its family has to choose what to be right about, and the diagnostic a practitioner would run — check the first lag — is the one that confirms the choice already made rather than the one that would have questioned it.
The estimate starts short
Before any rule is compared with any other there is a fact about all of them, and it is arithmetic rather than noise. Every estimated whitening in this collection is built from sample autocovariances, γ̂(k) averages , and ē is not zero. Subtracting a sample mean from a series that barely has one removes a large part of the series.
With the covariance known the whole expectation is available in closed form:
where R_t is the average of row t of Σ and G is the average of all of it. So what a sample of a given length will report is computable before any sample is drawn.
At d = 4/9 the first lag arrives as 0.538 against a truth of 0.800, and by the twelfth it is 0.089 against 0.609. Under the geometric law the shortfall is much smaller — 0.777 against 0.800 — and in the same direction.
This is not an estimation error anybody could fix by estimating better. It is what the estimator is of, at this sample size, and it means the sequence every window in this field tapers was already short before the taper touched it. It also means the comparison to make is between rules that all read the same short sequence, which is what the rest of this essay does.
What each rule is worth
Six rules on the same fifteen-candidate table, at n = 120, over two hundred draws. Two of them are fixed points: least squares with the ordinary penalty, which is told nothing, and a whitening told the entire covariance, which is infeasible and is here to say how much there was to win. In between are the four answers to what is the analyst told: the form is an autoregression of order one, find ρ; a window, take the sample autocovariances out to L; the same window, tapered so that the result is a covariance matrix; or the form is an autoregression of order p, find p coefficients.
Under the geometric law the parameterised rule reads 0.01294 and the tapered estimate 0.01919: a paired difference of 0.00625 at 3.9 standard errors in favour of knowing the form. That reproduces the earlier finding on new machinery, at a slightly milder persistence, and it is the whole of the case against generality.
Under the moving average the same two rules read 0.02935 and 0.02104, and the paired difference is 0.00831 at 4.5 standard errors the other way.
Those two numbers are the point of the field. The cost of generality where it is not needed and the benefit of it where it is are the same size. Nobody choosing between the two rules can decide by the magnitudes, because the magnitudes are equal and opposite; the decision is a statement about what shape the dependence has, which is exactly the thing neither rule can see.
Under long memory the general estimate wins by 0.00295 at 2.8 standard errors, which is real and small. Under the break the difference is 0.00066 at 0.4 standard errors, which is nothing at all — and that is the subject of its own essay, because the reason is not that the estimate is bad.
An order beats a window
The two general rules are not one rule. Estimating a covariance from a window of lags and estimating it from a fitted autoregression are different constructions, and the fitted model is ahead in three of the four worlds: by 0.00474 under the geometric law, 0.00342 under the moving average, and 0.00116 under the break, at 3.4, 3.6 and 0.8 standard errors respectively. Under long memory the two agree to 0.00001, which is a tie by any standard.
The mechanism is parameter count rather than shape. A tapered window at L = 20 carries twenty numbers estimated from a hundred and twenty rows; an autoregression of order six carries six, and each of them is estimated from every row rather than from the n − k pairs at one lag. The model is also the only construction here that extrapolates: a fitted autoregression has an autocovariance at every lag, including the ones past the end of its own order, where a window has exactly zero.
That extrapolation is not free and is not always right, which is what the order is for.
How much there was to win
The rules differ by less than the worlds do. Read the oracle row across the four laws and the range is enormous: being told the whole covariance recovers 98.5% of what least squares gives up under the moving average, 90.7% under the break, 82.2% under the geometric law, and 51.5% under long memory.
So the moving average is the world where knowing the dependence is worth almost everything, and long memory is the world where knowing it is worth half. That ordering is not the ordering of how hard the laws look. It is worth stating why the moving average is so generous: an error that is a five-period average of white noise has linear combinations with very little noise in them — the differences that cancel the shared innovations — and a whitening told the law puts its weight exactly there. Nothing estimated from a hundred and twenty rows finds those directions, which is why the best feasible rule recovers 77.7% where the oracle recovers 98.5%.
Under long memory the gap runs the other way: the oracle recovers half, and every feasible rule recovers about a quarter. The dependence is real, it is enormous, and most of it is at lags a sample of this length reports as nothing.
What the parameterisation is worth when it is right
One row of the table deserves its own sentence. Under the geometric law the feasible parameterised rule reads 0.01294 and the oracle reads 0.01362, a paired difference of 0.00068 at 0.5 standard errors. The rule that estimates one number from a hundred and twenty residuals is, within the noise, as good as the rule that is handed the covariance.
That is what a correct parameterisation buys, and it is the strongest argument for using one. It is also the argument’s own limit: the same rule recovers 62.9% under the moving average where the oracle recovers 98.5%, and 35.1% under the break where the oracle recovers 90.7%. A parameterisation that is right costs nothing and a parameterisation that is wrong costs most of what was available, and the sample cannot be asked which case it is in — not, at least, by fitting one number to it.
The truncated estimate is still not a covariance
The obvious general estimate is the sample autocovariances cut off at L, and it fails in a way that has nothing to do with accuracy: the truncated sequence is not always a covariance matrix, so a whitening built from it does not exist. The rate at which it fails is a fact about the world it is estimated in, and the four laws here span it — the rule exists on 55% of draws under long memory, 30% under the geometric law, 24% under the break, and 6% under the moving average.
The moving-average world is the extreme case and it is not a coincidence. That law’s spectral density has zeros in it, so its covariance is nearly singular and a truncated estimate of it is on the wrong side of nearly every time. Where the rule does exist its regret is reported over the draws it survived, which is an average of two rules under one name and is why it is quoted with its availability attached rather than ranked against anything.
The diagnostic cannot see what the choice depends on
The two halves of this essay meet in an awkward place. Which rule to use depends on the shape of the dependence, and the shape is what the sample reports worst.
A practitioner deciding whether to name the dependence or estimate it has one diagnostic to hand: the autocorrelations of the residuals. The first of them is the same in all four worlds by construction, so it carries no information about the decision at all. The lags that do carry it — the fourth, where the moving average stops; the twentieth, where long memory has not — are exactly the ones a hundred and twenty rows report as nearly zero. Under long memory the truth at the twentieth lag is 0.576 and the expected sample value is 0.016, which is indistinguishable from a world with no memory past the first few lags.
So the decision this table prices is not a decision the data will make for anyone. What the table can do is say what each answer is worth if it is right and what it costs if it is wrong, and those two numbers are equal here, which is a reason to make the choice on grounds outside the sample — how the series was generated, what a five-period average of anything would mean in the subject, whether the recording changed half way through.
What is claimed here, and what is not
This essay takes what estimating a dependence is worth when the parameterisation is wrong. The claims are that four laws standardised to the same lag-one autocorrelation are indistinguishable to a rule with one parameter; that the general estimate costs 0.00625 where the parameterisation is true and earns 0.00831 where it is not, at 3.9 and 4.5 paired standard errors, so the two prices are the same size; that a fitted autoregression beats a tapered window in three worlds of four; that a correct parameterisation is within half a standard error of being told the whole covariance; and that what a sample of a hundred and twenty rows reports at the first lag is 0.538 under long memory where the law says 0.800, computed exactly and confirmed by counting.
What stays out, and is named as a decision: which law a real sample came from. Nothing here is a test, and the four worlds are not a menu to choose from after looking at the data — choosing one of them on the same rows that the criterion then reads is a third selection problem, and the two that this collection does measure are the window and the order. What the table says is what each rule is worth given a world, which is the quantity a person who knows something about their errors can use and the quantity a person who knows nothing cannot.
The boundary against the essay that priced generality is that it measured the cost of estimating a covariance in the world where naming it is right, and this one measures the benefit in the worlds where it is not. Neither is a claim about which case is commoner.
The checks, and the refusals that make them mean something
Four claims are gated. Every law is required to have the same lag-one autocorrelation to machine precision, with the break — which has no autocorrelation function — required to average 0.7987 over its own pairs, because a drift in that number would silently turn every comparison here into a comparison between four different amounts of dependence. The series drawn from each law are required to match the exactly computed expectation of a sample that length rather than the law itself, which is the only comparison that is fair to a correct generator. The dense factorisation these laws need is checked against the banded one used elsewhere, on the moving average, which is the one law that is genuinely banded. And a series whitened by its own law’s factor is required to be flat on a portmanteau statistic over twelve lags, against the same statistic on the unwhitened series, which is what says the check has teeth.
The refusal is the one this essay’s own construction invites. A rule that exists on a fraction of draws is refused as a comparison unless the fraction is reported with it: under the moving average the truncated estimate survives on 6% of draws, and the draws it survives are the ones whose residuals happened to look benign.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The comparison that was not made — both name autoregression, covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
- The volume a whitening moves — both name autoregression, covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
- Iterating is not maximising — both name autocorrelation, bias, generalised least squares, nuisance parameter, regret, whitening
- Nothing in the fit picks the width — both name covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
- A window for every candidate — both name information criterion, model selection, nuisance parameter, regret, whitening
- The plug-in and the maximum — both name bias, covariance matrix, long memory, sample autocovariance, whitening
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationAutoregressionBiasCovariance matrixGeneralised least squaresInformation criterionLong memoryModel misspecificationModel selectionMoving averageNuisance parameterRegretSample autocovarianceStructural breakWhitening