Fitted together, or fitted after

The fit that takes the memory out

A candidate's residuals report less dependence than its errors do, and how much less is arithmetic rather than noise. The rule used for a good reason reads the series that has lost the most.

Worth reading first: The observations that repeat each other · A design is a number.

The whitening every rule in this collection builds is built from a sequence of sample autocovariances, and that sequence is read off a set of residuals. The essay that fitted the dependence with the line quotes three numbers for the first lag under a first-order autoregression at 0.8 — the law at 0.800, what a hundred and twenty errors report at 0.7773, and what a candidate’s residuals report at 0.7338 — and moves on to what can be done about it.

This essay is about the third number. It is not estimated, it is not simulated, and it depends on which candidate is doing the fitting.

The identity

A least-squares fit leaves ê = Me, where M = I − X(X′X)⁻¹X′ annihilates the candidate’s own design. So

E ⁣[γ^(k)]=1ntk(MΣM)t,tk\operatorname{E}\!\left[\hat\gamma(k)\right] = \frac{1}{n}\sum_{t \ge k}\left(M\Sigma M'\right)_{t,\,t-k}

with Σ the error covariance. Everything on the right is known before a sample exists: the design is drawn but the expectation is taken given it, and averaging over designs is then an average of exact numbers rather than an estimate of an unknown one.

There is no mean to subtract. A fit with an intercept leaves residuals summing to exactly zero, so the usual centring correction is already in M. That is what makes this a generalisation rather than a second effect: the sample-mean attenuation the laws field computes for the errors is this identity at the design that is an intercept and nothing else.

How much memory a fit takes out, candidate by candidate. Under AR(1) at 0.8, the lag-one autocorrelation a candidate's residuals report, computed exactly for each candidate on 200 draws. The upper line is the law at 0.8000. A candidate that is an intercept alone reports 0.7773 — which is exactly what a sample of 120 errors reports, because an intercept annihilates the sample mean and nothing else, and the two arithmetics agree to the last bit. Every predictor after that takes more out, down to 0.7341 at the fullest candidate. That is the collision this field is about: the rule every whitening here uses estimates its nuisance once, from the fullest candidate, so that the criteria stay comparable — and the fullest candidate is the one whose residuals report the least.
Fig. 1 The first-lag autocorrelation each candidate’s residuals report, exactly.

Neither M nor MΣM′ is ever formed. With P = X(X′X)⁻¹ and A = ΣX, each entry of the product is

(MΣM)ts=ΣtsPt ⁣ ⁣AsAt ⁣ ⁣Ps+Pt(XΣX)Ps,\left(M\Sigma M'\right)_{ts} = \Sigma_{ts} - P_t\!\cdot\!A_s - A_t\!\cdot\!P_s + P_t\left(X'\Sigma X\right)P_s,

which is a handful of multiplications per entry against a matrix product per lag. At a hundred and twenty rows and five columns that is the difference between a table and a wait, and the saving matters because the quantity is wanted per candidate, per law and per draw.

At the empty design the two arithmetics agree to the last bit — 0.7773 by row-and-grand averages of Σ, 0.7773 through a hat matrix — and they are written independently. That is the strongest check in the field, and it is the reason the rest of the curve can be believed: an identity that reproduces a known special case exactly is an identity, and one that comes close to it is a bug with a tolerance.

The intercept is worth about two of the other columns

The three numbers split the attenuation into two parts and the parts are not the same size.

From the law’s 0.800 to what the errors report at 0.7773 is 0.0227, and that whole step is the centring — the empty design is an intercept and nothing else. From 0.7773 to the candidate’s 0.7338 is a further 0.0435, and that step is four more columns.

So the four non-constant columns together take 1.9 times what the intercept takes, which is about 0.0109 apiece against the intercept’s 0.0227. One constant column costs what two ordinary predictors cost.

That is the same fact this collection meets from the other direction when it asks what a single scalar can say about a whole table of candidates: a constant column is the zero-frequency direction, and a persistent error process has more of its variation there than anywhere else. Every other column has variation the errors do not share, so its projection removes less. The ordering is not a property of these particular predictors; it is a property of what the errors look like.

What eight per cent at the first lag is worth downstream

The first-lag figure is not what any rule uses. What a whitening implies is a long-run variance, and the map from one to the other amplifies.

For a first-order process the inflation is (1+ρ)/(1ρ)(1+\rho)/(1-\rho), which gives 9.00 at the law’s 0.800, 7.98 at the errors’ 0.7773 and 6.51 at the residuals’ 0.7338.

So a shortfall of 8.3% in the correlation becomes a shortfall of 27.6% in the quantity every rule in the collection actually reads — an amplification of 3.3. The local elasticity is larger still: 2ρ/(1ρ2)2\rho/(1-\rho^2) is 4.4 at ρ = 0.8, so the amplification would be worse for a smaller misstatement and worse again at a higher persistence.

Split as before, 11.4 points of the 27.6 are the centring and 16.2 are the rest of the design.

Which is why this identity is worth computing exactly rather than bounding. A correction of two hundredths at the first lag looks like something a wide window would absorb, and by the time it reaches the object the rule is built from it is more than a quarter of the answer.

It is monotone, and it points the wrong way

Every predictor added to a candidate takes more dependence out. Under the first-order autoregression the five candidates on the ladder report 0.7773, 0.7619, 0.7497, 0.7406 and 0.7341; under the moving average 0.7846 down to 0.7479; under long memory 0.5377 down to 0.4894; and under the break 0.7598 down to 0.7116.

The direction is not a surprise once it is stated. A design column that is itself persistent partly tracks a persistent error, so the fit removes some of the error along with the signal, and it removes the part that is smooth in time — which is the part the dependence lives in. What is worth stating is the size: from an intercept to four predictors the first lag drops by 0.0433 under the geometric law, which is nearly twice the 0.0227 the sample had already lost before any fit touched it.

And it points the wrong way for the rule this collection uses. Every field here that estimates a nuisance estimates it once, from the fullest candidate, so that fifteen criteria are comparable — a decision the estimated-covariance field made deliberately and the shared-against-own comparison priced. The fullest candidate is the one at the bottom of this curve.

Both arguments are right. A nuisance estimated per candidate makes the criteria incommensurable, which is worse than an attenuation. But the cost of the choice was never a number until now, and it is a fifth of a standard error’s worth of correlation on every rule in the collection.

It is worth being precise about what would happen if the rule were reversed. A whitening built from a one-predictor candidate’s residuals is built from a sequence that is 0.0279 less attenuated at the first lag, and it is then applied to all fifteen candidates alike — so nothing about comparability breaks, and the only objection is that the choice of which candidate to read is arbitrary. That is a real objection and it is a different one from the objection to reading each candidate’s own. Estimating the nuisance from the smallest candidate is admissible and estimating it from each is not, and the collection has never distinguished the two because it never had a reason to prefer either.

The reason is here and it is small. A fifth of a standard error of correlation buys, at the sizes the joint fit measures, something under a tenth of what estimating the coefficient costs at all — which is itself a fiftieth of what not knowing the law costs. The honest conclusion is that the choice of source candidate does not matter, and the honest way to reach it is to have computed the thing rather than assumed it.

What the whole sequence looks like

The first lag is where the effect is easiest to state and not where it is largest.

The dependence, at four removes. Under AR(1) at 0.8, four different sequences all called the dependence. The top line is the law. The middle line is what a sample of 120 errors reports on average — computable exactly, because the expectation of a sample autocovariance is arithmetic once the covariance is known. The lower line is what a candidate's residuals report, which is what every two-step rule in this collection actually reads: a fit removes variance, and it removes more of the persistent part than of the rest. At the first lag the three are 0.800, 0.7773 and 0.7338. The dots are counted from draws and share no arithmetic with the line they sit on; the worst departure is 1.2 standard errors.
Fig. 2 The law, what the errors report, and what a candidate’s residuals report, at six lags.

Under the geometric law the residual sequence crosses zero somewhere past the eighth lag and is −0.0239 at the twelfth, where the law says 0.0687 and is positive. Under the moving average, whose autocorrelation is exactly zero past the fourth lag, the residuals report −0.0559 at the eighth and −0.0438 at the twelfth: a series with no memory at all past four lags is reported as having negative memory at eight.

The mechanism is arithmetic and has nothing to do with the law. A residual sequence sums to zero by construction, so its autocovariances at lag zero and above cannot all be positive: whatever the errors do, the sequence a fit leaves has to give some of it back somewhere. The only question is where, and the answer is the tail — because the head is where the dependence is and the fit could not remove all of it without removing the signal too.

That negative tail is not an artefact to be corrected. It is what a whitening built from these numbers is built from, and it is why a tapered estimate exists at all: a sequence that goes negative in its tail is not always a covariance sequence, and a window that keeps it is not always invertible.

What the size of it depends on

Three things move the attenuation and it is worth separating them, because two of them are under nobody’s control and the third is a design decision.

The sample length. Everything here is at a hundred and twenty rows, which is the length this collection’s predictive regressions run at. The shortfall is roughly proportional to the number of coefficients over the number of rows, so at four hundred and eighty rows the same fit would cost a quarter as much. That is not a recommendation to use more rows; it is the statement that the effect is a small-sample effect and that nothing about it is a bias in the usual sense — it does not survive to any limit.

The persistence of the design columns. The four predictors here are drawn at persistences of 0.9, 0.6, 0.3 and 0, which is the tiered design the effective-size field introduced precisely so that candidates of the same dimension are not interchangeable. A persistent column can track a persistent error and a white one cannot, so most of the 0.043 comes from the first column. A table whose regressors were all white would show almost none of this, which is why a comparison made on such a table would have found nothing.

The persistence of the errors. Under long memory the fit removes 0.0483 of the first lag against 0.0433 under the geometric law, and the sequence it removes it from was already 0.2623 short. The effect is largest exactly where the sequence can least afford it.

The law that has no autocorrelation

The break is the one law here whose covariance depends on where a pair sits rather than only on the gap between the pair. It has no autocorrelation function, so what a residual sequence reports for it is an average over positions — and the identity above computes exactly that, because it never assumed stationarity in the first place. M Σ M′ is a matrix, and the sum over t is a sum over positions whatever Σ is.

The dependence, at four removes. Under a break in the persistence, four different sequences all called the dependence. The top line is the law. The middle line is what a sample of 120 errors reports on average — computable exactly, because the expectation of a sample autocovariance is arithmetic once the covariance is known. The lower line is what a candidate's residuals report, which is what every two-step rule in this collection actually reads: a fit removes variance, and it removes more of the persistent part than of the rest. At the first lag the three are 0.000, 0.7598 and 0.7114. The dots are counted from draws and share no arithmetic with the line they sit on; the worst departure is 0.7 standard errors.
Fig. 3 The same four removes on a covariance that changes with position rather than with gap.

Averaged over its own pairs the law reads 0.7987 at the first lag; the errors report 0.7598 and the fullest candidate’s residuals 0.7116. So a stationary reading of the residuals of a non-stationary error is short twice over — once because the sample is finite, once because the fit removed some of it — before it is short a third time for being an average of two regimes that no single ρ describes.

Two routes, and the one that is not a route

The exact expectation is checked against counted residuals, and the comparison has a trap in it that this collection has walked into before.

γ̂(k)/γ̂(0) is a ratio, and the expectation above is linear in γ̂ rather than in the ratio. So the counted figure has to be formed as a ratio of averages — mean γ̂(k) over draws, divided by mean γ̂(0) — and not as an average of ratios. Formed the wrong way the two routes disagree at the first lag by 0.016 against a standard error of 0.006, which is nearly three standard errors and looks exactly like a defect in the identity. It is not: it is a comparison between two different quantities, and it is the same refusal the laws table records one level up.

Formed the right way, the exact and counted first lags are 0.7338 and 0.7340 against a standard error of 0.0181, and every lag out to the twentieth agrees inside a standard error.

What this changes about the fields that read it

Three constructions in this collection read a residual autocovariance sequence and none of them was built knowing what the sequence is.

The window a whitening wants is chosen by a criterion on residuals, and the criterion is choosing over a sequence whose tail is already negative. The order the tail is drawn at is chosen the same way, and its extrapolation is an extrapolation of an attenuated sequence. The tapered estimate exists because the untapered one is not always a covariance matrix, and how often it fails is a function of how negative the tail is.

None of those three is wrong. What changes is that the quantity each of them is a function of is now computable, so a claim about any of them can be checked against what its input actually is rather than against what the law is.

The one that changes most is the window. A criterion choosing a window is choosing how many of these lags to keep, and the argument for a long window is that a long dependence needs one. Under the moving average the law has nothing past the fourth lag, the residuals report −0.056 at the eighth, and the best fixed window is thirty — a number that made no sense against the law and makes a different kind of sense against this sequence: a window of thirty is keeping thirty lags of something that is not zero, and what it is keeping is the fit’s own signature rather than the error’s memory. That does not make the window wrong, because the whitening it builds is judged on what it whitens rather than on what it is an estimate of. It does mean the usual reading of the number — the window should be several times the memory — is a reading of the wrong sequence.

What is claimed here, and what is not

This essay takes what a fit does to a dependence before any rule reads it. The claims are that the expectation of a residual autocovariance is exact given the design and computable in the cost of one small matrix per candidate; that the identity reduces at the empty design to the sample-mean attenuation the laws field computes, to machine precision, by independent arithmetic; that the attenuation is monotone in the number of coefficients a candidate fits, dropping the first lag from 0.7773 to 0.7341 under the geometric law; that the residual sequence is negative past the eighth lag under a law whose autocorrelation is positive everywhere; and that the counted route agrees with the exact one when the ratio is formed the way the expectation is linear in.

What stays out, and is named as a decision: a correction. Nothing here proposes to add the attenuation back. A bias-corrected autocovariance sequence is not guaranteed to be a covariance sequence, which is the same wall the truncated estimate runs into, and correcting a sequence that a rule then inverts is a repair whose failure mode is a matrix that does not exist rather than a number that is wrong. The route taken instead is fitting the dependence with the line, which changes what is estimated rather than adjusting what was.

The boundary against the essay on residuals and errors is that it establishes that a resampling of residuals cannot carry more dependence than the residuals have, and this one computes how much they have. The first is a ceiling and the second is the number the ceiling is at.

The checks, and the refusals that make them mean something

Three claims are gated. The exact expectation is required to agree with counted residuals at every lag, as a ratio of averages. The identity is required to reproduce the sample-mean expectation exactly at the empty design, which is the reduction that makes it an identity rather than an approximation. And the attenuation is required to be monotone in the candidate’s dimension, on every law, because a direction that held on one law and not on another would be a fact about that law rather than about fitting.

The refusal is the one that produced the check: an average of ratios compared against a ratio of expectations is refused, with the 0.016 discrepancy it manufactures printed beside the standard error it would have been read against.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AttenuationAutocorrelationBiasClosed formCovariance matrixDependenceHat matrixLong memoryModel selectionNuisance parameterProjectionResidualSample autocovarianceStructural breakWhitening