Concept

Long memory — where it appears

Dependence whose autocorrelation decays like a power of the lag rather than geometrically, so that it is still substantial hundreds of steps out. A fractionally integrated series is the standard example, and its sums and averages converge more slowly than a short-memory series of the same variance.

Named by 10 essays across 5 fields — each of them below, with the objects they name alongside it.

The dependence, at four removes. Under AR(1) at 0.8, four different sequences all called the dependence. The top line is the law. The middle line is what a sample of 120 errors reports on average — computable exactly, because the expectation of a sample autocovariance is arithmetic once the covariance is known. The lower line is what a candidate's residuals report, which is what every two-step rule in this collection actually reads: a fit removes variance, and it removes more of the persistent part than of the rest. At the first lag the three are 0.800, 0.7773 and 0.7338. The dots are counted from draws and share no arithmetic with the line they sit on; the worst departure is 1.2 standard errors.

A dependence fitted with the line

Every whitening in this collection reads the dependence off a set of residuals, and residuals are not errors. Fitting the two together recovers most of what that costs, and changes almost nothing about the decision it feeds.

together · Dependence
Four dependences a single parameter cannot tell apart. Every law here is standardised to a lag-one autocorrelation of 0.8, so a rule told the errors are a first-order autoregression finds the same number in all four and has no way of seeing what separates them. The geometric decay is the world in which estimating a covariance rather than naming it was priced, and found to cost. The five-period moving average has 0.200 at the fourth lag and exactly nothing past it, where the geometric law says 0.328 at the fifth. Long memory at d = 4/9 is still at 0.576 by the twentieth lag, where the geometric law has reached 0.012. The break has no autocorrelation function at all: what is drawn for it is the average over the pairs at each gap, which is what a stationary estimate converges to.

A dependence with a shape

Four ways for errors to repeat, all with the same first lag and nothing else in common. A rule told the errors are a first-order autoregression finds the same number in all four, and is right about one of them.

general · Dependence
Where the general fit becomes the parametric one. The band family's objective at the autoregression's own geometric sequence, cut off at each width, on one sample of 60 rows. The horizontal line is the profile likelihood the parametric fit maximises, written independently through a different whitening. At the full width the two are the same number to 3e-14, which is what says the general construction contains the parametric one rather than resembling it. Below 10 lags there is no line at all: the geometric sequence cut off short is not a covariance matrix, so the objective has nothing to evaluate. Between the two the truncation is briefly above the parametric likelihood — a wrong covariance can fit one sample better than the right one, which is the whole reason a width has to be charged for rather than chosen.

A family before a fit

A regression's coefficients and one correlation can be maximised together. Replace the correlation with an estimated covariance and there is nothing left for "jointly" to mean — until a set of covariances is named, and the set turns out not to contain the truth.

family · Dependence
How much memory a fit takes out, candidate by candidate. Under AR(1) at 0.8, the lag-one autocorrelation a candidate's residuals report, computed exactly for each candidate on 200 draws. The upper line is the law at 0.8000. A candidate that is an intercept alone reports 0.7773 — which is exactly what a sample of 120 errors reports, because an intercept annihilates the sample mean and nothing else, and the two arithmetics agree to the last bit. Every predictor after that takes more out, down to 0.7341 at the fullest candidate. That is the collision this field is about: the rule every whitening here uses estimates its nuisance once, from the fullest candidate, so that the criteria stay comparable — and the fullest candidate is the one whose residuals report the least.

The fit that takes the memory out

A candidate's residuals report less dependence than its errors do, and how much less is arithmetic rather than noise. The rule used for a good reason reads the series that has lost the most.

together · Dependence
Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.

The plug-in and the maximum

A tapered covariance estimate sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. Four fifths of that is what the optimiser would have found if nothing were missing.

family · Dependence
The quantity that does not depend on the list. The probability that letting each candidate choose its own tuning parameter changes which candidate the table selects — the product of the two moving shares — against the length of the list, over 1200 draws apiece. It is 14.2%, 11.9%, 12.3%: a spread of 2.2% across a list length that moves the disagreement rate by a factor of 1.52. This is the invariant the whole field turns on. Everything downstream of the winner — the coefficients, the regret, whatever a reader is going to quote — is a function of whether the winner changed, and how often that happens is not something the list controls. A longer list changes how often the candidates quarrel and not how often the quarrel matters.

How often it matters

The disagreement rate rises by half across the list and the share of disagreements that decide anything falls by nearly the same factor. Their product — how often the tuning list changes which candidate wins — sits at an eighth and does not move.

apiece · Order-selection
The window a whitening wants is not the memory of the errors. Regret under a five-period moving average as the tapered estimate is given more lags, over 120 draws at n = 120. The best window is L = 30; the automatic bandwidth is 4 and the error model's own likelihood chooses 9.7 on average. Both land in the same place and both are short, and the reason is the taper: a Bartlett weight at lag k is 1 − k/(L + 1), so a window of 8 keeps 0.556 of whatever the fourth lag carries and a window of 30 keeps 0.871. A window has to be several times the memory before it stops removing the memory. The dashed line is the rule told the errors are a first-order autoregression, which needs no window at all.

The window a whitening wants

Every law here is best whitened by a window several times longer than its own memory, including the one whose memory ends at the fourth lag. The three ways of choosing it from the sample all land in the same place, and it is the wrong one.

general · Order-selection
The ladder is the same ladder under every law. Each pair's overlap under each of the four laws, over 400 draws apiece. The scale is fixed at both ends by construction: two searches over disjoint sets of independent columns read -0.001, 0.013, -0.002, 0.004, and a break search paired with a step column it already contains reads exactly one under every law. Between them the pair that reads one residual series twice runs 0.763, 0.805, 0.577, 0.752 — lowest under long memory, where a whitening has most to do and the break search has least left to find that the whitening has not taken. The rung that moves most is the pair of step dictionaries, from -0.148 under a moving average to 0.419 under a break; and the pair that is negative is negative under all four.

Three quarters of the way to one search

The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.

apart · Criterion
Four sequences, and the rule only ever sees the last one. Under long memory at d = 4/9, four things that are all called the dependence. The law itself is the top line. What a sample of 120 rows reports on average is the second, computed exactly: subtracting a sample mean takes the first lag from 0.800 to 0.538. What a candidate's residuals report is the third, lower again at 0.472, because a fit removes dependence along with signal. The autoregressions are fitted to that third sequence and reproduce it exactly out to their own order — the Yule–Walker equations are solved to make it so — so everything they say past that is extrapolation. At the twentieth lag the law has 0.576, the residuals report 0.006, and an AR(8) extrapolates 0.028.

The order the tail is drawn at

A fitted autoregression reproduces the sample exactly at the lags it was fitted on, so everything it says past them is extrapolation — and the order is the dial that decides how much of it there is.

general · Order-selection
Fitting them together is worth something under one law. Four fits of the same regression under four dependences: least squares, the two-step plug-in every whitened rule in this collection runs, the coefficients and the band maximised together, and a whitening at the law's own covariance that nobody has. Under the moving average — the one law the band family contains — the joint fit beats the two-step by 0.0077 at 3.3 paired standard errors. Under the autoregression, long memory and the break it is a tie: 0.4, 1.0, 0.3 standard errors. That is the same ordering the likelihood gap gave, arrived at through the coefficients rather than through the objective.

What fitting them together buys

Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.

family · Dependence

Named alongside it

The objects these essays reach for when they reach for this one.

WhiteningModel selectionNuisance parameterAutocorrelationCovariance matrixGeneralised least squaresInformation criterionRegretSample autocovarianceBiasDependenceStructural break

All concepts