Fitted together, or fitted after

A dependence fitted with the line

Every whitening in this collection reads the dependence off a set of residuals, and residuals are not errors. Fitting the two together recovers most of what that costs, and changes almost nothing about the decision it feeds.

Worth reading first: The observations that repeat each other · A design is a number.

Every rule in this collection that repairs a criterion by whitening does it in two steps. A candidate is fitted by least squares, the dependence is estimated from what the fit leaves behind, the sample is whitened with that estimate, and the candidate is fitted again. The rule that reads one number does it, the rule that reads a window does it, and the rule that reads an order does it.

The step nobody has examined is the first one. The residuals a fit leaves behind are not the errors, and the difference between them is not noise: it is arithmetic, it has a sign, and it is computable before any sample is drawn.

What the fit takes out

A candidate’s residuals are ê = Me, where M is the annihilator I − X(X′X)⁻¹X′ of that candidate’s own design. So the expectation of the sample autocovariance at lag k is

E ⁣[γ^(k)]=1ntk(MΣM)t,tk,\operatorname{E}\!\left[\hat\gamma(k)\right] = \frac{1}{n}\sum_{t \ge k} \left(M\Sigma M'\right)_{t,\,t-k},

which is exact given the design, with no simulation anywhere in it.

The dependence, at four removes. Under AR(1) at 0.8, four different sequences all called the dependence. The top line is the law. The middle line is what a sample of 120 errors reports on average — computable exactly, because the expectation of a sample autocovariance is arithmetic once the covariance is known. The lower line is what a candidate's residuals report, which is what every two-step rule in this collection actually reads: a fit removes variance, and it removes more of the persistent part than of the rest. At the first lag the three are 0.800, 0.7773 and 0.7338. The dots are counted from draws and share no arithmetic with the line they sit on; the worst departure is 1.2 standard errors.
Fig. 1 The dependence at four removes, under a first-order autoregression. The dots are counted from draws and share no arithmetic with the line they sit on.

Three quantities get called the dependence and a rule only ever sees the third. The law is 0.800 at the first lag. What a sample of a hundred and twenty errors reports is 0.7773, because subtracting a sample mean from a persistent series removes part of the series. And what a candidate’s residuals report is 0.7338, because a fit removes more than a mean does — and removes more of the persistent part than of the rest, since a persistent error is exactly what a smooth design column can partly track.

Under long memory the same three numbers are 0.800, 0.5377 and 0.4889. The law is unchanged, the sample’s shortfall is enormous, and the fit takes another five points off the top of it.

That third row is what every two-step rule in this collection has actually been reading, and until now it was not computed anywhere.

The size of the effect is a fact about the design rather than about the dependence. A candidate holding one predictor removes one degree of freedom and a little of the low-frequency part of the error; a candidate holding four removes five, and the fullest candidate is the one every rule in this field estimates its nuisance from — which is the choice the shared-against-own comparison settled for a different reason and which turns out to cost the most attenuation of any candidate on the table. The two arguments point opposite ways, and nothing before this could tell how far — which is the arithmetic underneath this one.

How much memory a fit takes out, candidate by candidate. Under AR(1) at 0.8, the lag-one autocorrelation a candidate's residuals report, computed exactly for each candidate on 200 draws. The upper line is the law at 0.8000. A candidate that is an intercept alone reports 0.7773 — which is exactly what a sample of 120 errors reports, because an intercept annihilates the sample mean and nothing else, and the two arithmetics agree to the last bit. Every predictor after that takes more out, down to 0.7341 at the fullest candidate. That is the collision this field is about: the rule every whitening here uses estimates its nuisance once, from the fullest candidate, so that the criteria stay comparable — and the fullest candidate is the one whose residuals report the least.
Fig. 2 How much of the dependence each candidate’s own fit removes, computed exactly.

A likelihood that has both in it

There is a rule that does not read residuals at all. Write the model with its dependence in it — a regression whose errors are a first-order autoregression — and maximise the Gaussian likelihood over the coefficients and the correlation together. The Prais–Winsten transform makes the errors independent at any given ρ, so the concentrated log-likelihood is

(ρ)=n2logσ^2(ρ)+12log(1ρ2),\ell(\rho) = -\tfrac{n}{2}\log \hat\sigma^2(\rho) + \tfrac12 \log(1 - \rho^2),

where σ̂²(ρ) is the residual variance of a least-squares fit on the transformed sample. The second term is the Jacobian of the transform: the first row is scaled by √(1 − ρ²) and every other row is a difference, whose Jacobian is one.

A two-step rule never sees that term. It reads ρ̂ off a set of residuals and stops.

It is worth saying why the term is there rather than treating it as bookkeeping. Whitening is a change of variables, and a likelihood written in transformed coordinates has to carry the determinant of the transform or it is not a likelihood at all — it is a sum of squares wearing one. The determinant is √(1 − ρ²), from the single scaled row, and it rises as ρ falls. So the term pulls the maximiser towards zero and the sum of squares alone does not, which is the one place a first-order condition on residuals and a first-order condition on the likelihood can disagree in a stated direction.

Three routes to one number. The lag-one coefficient of the errors, estimated three ways on the same 200 draws under AR(1) at 0.8. The law is 0.800. Reading it off the least-squares residuals gives 0.7237; iterating that rule to its fixed point gives 0.7673; maximising the Gaussian likelihood over the coefficients and the correlation together gives 0.7762. All three are attenuated, because a hundred and twenty rows are a hundred and twenty rows, and the paired gap between the first and the last is 0.0525 at 21.2 standard errors. The spread hardly moves — 0.068 against 0.063 — so what separates them is bias and not noise.
Fig. 3 The same coefficient, estimated three ways on the same two hundred draws.

On two hundred draws under a first-order autoregression at 0.8, reading the coefficient off the least-squares residuals gives 0.7237. Maximising the likelihood gives 0.7762, a paired difference of 0.0525 at 21.2 standard errors. Both are still short of the law — a hundred and twenty rows are a hundred and twenty rows — but the joint estimate gives back two thirds of what the two-step estimate loses, and it does it without changing the spread: 0.0678 against 0.0628.

The middle row is the one that is interesting. Iterating the two-step rule — re-reading the coefficient from the generalised least-squares residuals and refitting until nothing moves — gives 0.7673. That is most of the way, and it is not all of the way, and the reason it is not is the Jacobian term.

What a better estimate is worth

The coefficient is not the answer to anything. What the coefficient is for is a criterion, and the criterion picks a model out of fifteen candidates. So the question this field has to ask is not how much closer the joint estimate is but how much better the choice gets.

What a better estimate is worth to the decisionThe regret each rule gives up against the best model available, under AR(1) at 0.8 over 300 draws. The two ends are fixed points: least squares at 0.07848, told nothing, and a whitening told the whole covariance at 0.01551. Between them the three estimates of one number sit 0.01771, 0.01554 and 0.01549 — a paired gain of 0.00222 at 1.8 standard errors for the joint fit over the two-step one, where the whole of what estimating the number costs is 0.00220. The coefficient moved by three times as much as the decision did.least squares, 2q0.07848± 0.00435whitened at the true Ω0.01551± 0.00140ρ̂ from the residuals0.01771± 0.00156ρ̂ iterated to its fixed point0.01554± 0.00143ρ̂ from the joint likelihood0.01549± 0.00152300 draws, fifteen candidatesthe estimate improves; the choice barely does
Fig. 4 What each rule gives up against the best model available, with the law on the slider.

The answer is: almost nothing.

Under the first-order autoregression, over three hundred draws, least squares gives up 0.07848 of regret and a whitening told the whole covariance gives up 0.01551. Between them the two-step rule sits at 0.01771 and the joint rule at 0.01549 — a paired gain of 0.00222 at 1.8 standard errors, against a total available gain of 0.00220. The joint rule has recovered essentially the whole of what estimating the coefficient costs, and the whole of what estimating the coefficient costs is a fiftieth of what not knowing the law costs.

The dimension the rules choose says the same thing more directly. Averaged over three hundred draws, the two-step rule picks a candidate of 3.73 coefficients and the joint rule one of 3.75, against the oracle’s 3.74. Three rules, three whitenings built at three different correlations, and the models they select are the same models.

The estimate moved by twenty-one standard errors and the decision moved by two. That is not a contradiction and it is not a disappointment; it is a statement about what a criterion is sensitive to. A whitening at ρ = 0.72 and a whitening at ρ = 0.78 produce two different transformed samples, and the fifteen candidates are ranked almost identically on both, because the ranking depends on differences between candidates and the whitening is applied to all fifteen alike.

Under the break the gain is larger and clearer: 0.11483 against 0.10928, a paired difference of 0.00555 at 2.8 standard errors. Under the moving average it is 0.00064 the other way, at 0.8 standard errors, which is nothing. And under long memory it is 0.00130 at 2.1.

The whole of the gain is the first iteration

The three-column comparison has a fourth reading in it that neither of the two obvious ones would find.

Under the first-order autoregression the iterated rule gives up 0.01554 and the joint rule 0.01549: a paired difference of 0.00006 at 0.1 standard errors. Under the break it is 0.00059 at 0.7. Under long memory it is 0.00027 at 3.0, which is real and is a twentieth of the whole gain.

So the picture is: iterating the two-step rule to its fixed point recovers nearly all of what the two-step rule loses, and maximising the likelihood on top of that recovers nearly nothing. The difference between a fixed point and a maximum is real, one-sided and measurable in the coefficient, and it is worth nothing in the decision. That is worth stating plainly because the opposite conclusion is the one a practitioner would draw from the coefficient alone, and the coefficient alone is what a textbook comparison reports.

Where the joint fit does matter

There is a place in this collection where the attenuation is expensive, and it is not the coefficient.

The order an autoregression wants is chosen by a criterion run on a residual series, and that essay reports the gap between the order chosen from residuals and the order chosen from the errors themselves — the second of which nobody has. Under a moving average the two are 7.51 and 9.24. Nearly two orders are lost to the attenuation, and the essay names a joint fit as the repair and does not build one.

Where the order comes from. The autoregressive order a criterion picks under a five-period moving average, from three different series over 150 draws. The residuals of the fullest candidate give 7.51; the errors themselves, which nobody has, give 9.24; and a joint fit — the order chosen by the likelihood of the regression and the dependence together — gives 8.79. The joint route recovers 1.27 of the 1.73 orders the attenuation costs, at 7.2 paired standard errors. It is available where the middle row is not.
Fig. 5 The order three different series ask for, under a moving average.

Built, it recovers 1.27 of the 1.73 orders at 7.2 paired standard errors, landing at 8.79. It is available where the middle row is not: the errors are a simulator’s private quantity and the joint likelihood is a thing anybody can maximise.

The gap is not the same in every world. Under long memory the three read 2.45, 2.77 and 2.53, so the joint fit overshoots the errors slightly — which reproduces from the other side what the order essay established by running the same criterion on the errors: under long memory it is the criterion that is aimed at the wrong part of the sequence, and not the attenuation. Under the first-order autoregression all three sit within a standard error of each other, because there is nothing there to lose.

Three attenuations, and which one the repair reaches

The three quantities that get called the dependence are two differences, and differencing them says which of the two a joint fit is aimed at.

Under the first-order autoregression, the law’s 0.800 falls to 0.7773 at the errors — a loss of 0.0227 to subtracting a sample mean — and then to 0.7338 at the residuals, a further 0.0435 to the fit. The fit costs nearly twice what the mean does. Under long memory the same two steps are 0.2623 and 0.0488: the sample’s own shortfall is more than five times the fit’s.

So the ratio between the two attenuations differs by a factor of ten between the two laws, and it decides what a joint fit can be worth. Under the autoregression the repair is aimed at the larger of the two terms; under long memory it is aimed at the smaller one, and no amount of joint fitting touches the 0.26 that a hundred and twenty rows cost before any regression is run.

There is a third step and it is the smallest. The two-step rule reports 0.7237 where the residuals’ own expectation is 0.7338, so the sample estimate of a residual autocorrelation is a further 0.0101 low — the ordinary finite-sample bias of a ratio of autocovariances, on top of the two structural losses. The full accounting from law to reported number is 0.0227, 0.0435 and 0.0101: 30%, 57% and 13%.

The joint estimate lands on the bound

Set the joint estimate against that accounting and it does something sharper than “recovering two thirds”.

The joint maximum-likelihood estimate is 0.7762. The errors’ own reported dependence — the number a rule handed the true errors would read, and the best any estimator can do on a hundred and twenty rows — is 0.7773. They differ by 0.0011, which is a sixth of one standard error.

The joint fit recovers the whole of the fit’s attenuation and the whole of the ratio’s, and none of the sample’s — which is exactly the right division, since the first two are consequences of the estimator and the third is a consequence of the sample size. Recovering 0.0525 of 0.0763 is 69% of the total; recovering 0.0536 of the 0.0536 that is recoverable is all of it.

The same shape appears one level along in the regret column, where it is easier to miss. The joint rule gives up 0.01549 against the oracle whitening’s 0.01551 — it is not near the bound, it is at it, a fifth of a standard error the better side. The two-step rule’s 0.01771 is 14% worse than the oracle and the whole of that 14% is what estimating one coefficient from residuals costs.

That is why the decision moves by two standard errors while the estimate moves by twenty-one, and it is a better reading than “the criterion is insensitive”: the criterion is not insensitive, it is already at its ceiling under the two-step rule, because 14% of a quantity that is itself a fiftieth of what the law is worth is nothing. The joint fit is a complete repair of a small term.

The two ways the same repair can fail

A joint fit is not free of the thing it repairs. It is a fit, so its own residuals are attenuated too — and the order it selects is selected on a whitening built from an order. What makes the comparison honest is that the three columns are read on the same draws, and that each column is a rule a practitioner could actually run: the residual column is what is run now, the joint column is what could be run instead, and the error column is a bound nobody can reach.

The bound matters because it is what says the repair is nearly complete rather than merely in the right direction. Recovering 1.27 of 1.73 is a statement; recovering 1.27 with no idea what was available is not.

There is a second failure the design of the comparison rules out rather than measures. A joint fit is fitted per candidate, so its nuisance estimate is a candidate’s own — which is exactly the rule the estimated-covariance field refuses, on the grounds that criteria built on different nuisances are not comparable. Every joint estimate quoted here is therefore taken from the fullest candidate and shared across the table, in the same place and in the same way the two-step estimate is, so that the only thing differing between the columns is how the number was arrived at. That is a restriction rather than a discovery, and it is the reason the comparison is a comparison.

The dependence, at four removes. Under long memory at d = 4/9, four different sequences all called the dependence. The top line is the law. The middle line is what a sample of 120 errors reports on average — computable exactly, because the expectation of a sample autocovariance is arithmetic once the covariance is known. The lower line is what a candidate's residuals report, which is what every two-step rule in this collection actually reads: a fit removes variance, and it removes more of the persistent part than of the rest. At the first lag the three are 0.800, 0.5377 and 0.4889. The dots are counted from draws and share no arithmetic with the line they sit on; the worst departure is 0.7 standard errors.
Fig. 6 The same four removes under long memory, where the sample’s own shortfall dwarfs everything a fit adds to it.

What is claimed here, and what is not

This essay takes what a two-step estimate of a dependence gives up, and what a joint estimate gets back. The claims are that a candidate’s residuals report less dependence than the errors do and that the shortfall is exactly computable from the design; that the joint maximum-likelihood estimate of a first-order coefficient is nearer the law than the two-step estimate by 0.0525 at 21.2 paired standard errors; that this buys 0.00222 of regret at 1.8 standard errors, which is the whole of what estimating the coefficient costs and a fiftieth of what not knowing the law costs; and that the same joint fit recovers 1.27 of the 1.73 orders a moving average’s attenuation costs, at 7.2 standard errors.

What stays out, and is named as a decision: the joint fit for the general covariance. The likelihood above is written for one parameter, and the field that estimates a whole covariance has no parameter to put in it — a window and a tapered Ω̂ are not a model with a likelihood, they are an estimate substituted into one. What that means for a joint fit is a different question from this one, and it is the tuning parameter’s question rather than the coefficient’s.

The boundary against the essay that priced the whitening is that it takes what a whitening is worth given an estimate, and this one takes what the estimate is worth given the whitening. Neither is a claim about which of the two is the larger term — except that here, on this design, it is measured and the answer is the first by a factor of fifty.

The checks, and the refusals that make them mean something

Four claims are gated. The exact residual expectation is required to agree with counted residuals, which is a comparison between a matrix identity and four hundred draws — and it is compared as a ratio of averages rather than an average of ratios, because the expectation is linear in γ̂ and not in γ̂(k)/γ̂(0). The first version of this file compared the two the other way and reported a defect of 0.016 that was a comparison between two different quantities. The concentrated profile is required to equal the Gaussian log-likelihood evaluated at its own coefficient vector and variance, with the quadratic form written out, because the concentration is an algebraic step and an algebraic step is the kind that is right in the derivation and wrong in the code. The joint estimate is required to be attenuated too, and to be less attenuated than the two-step one on the same draws. And the joint order is required never to fall below the residual order beyond the noise, which is the direction the whole argument predicts.

The refusal is the one this collection has been walking past. The order read off a candidate’s residuals, quoted as the order of the errors, is refused — the two differ by 1.73 at 8.0 paired standard errors under a moving average, and every two-step rule in the collection has silently treated them as one quantity.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AttenuationAutocorrelationBiasFeasible generalised least squaresGeneralised least squaresHat matrixInformation criterionLong memoryMaximum likelihoodModel selectionNuisance parameterPrais winsten transformRegretResidualWhitening