Differencing — where it appears
Named by 9 essays across 3 fields — each of them below, with the objects they name alongside it.
The observations that repeat each other
Almost every standard error divides by √n, which claims the observations carry independent information. At a lag-one correlation of 0.8 a fifty-point series is worth about six independent observations, and its 95% interval covers 47%.
Three series and a count
A pair of series is either tied together or it is not, so its whole inference is one test with one answer. Three can carry none, one or two relations at once — and the thing being estimated stops being a slope and becomes an integer, read off the gap in a spectrum whose top eigenvalue holds at 0.25 while the rest fall like 1/n.
Counting what is still wandering
The statistic that turns a spectrum into an integer has one name and three distributions. Its 5% point is 8.12, 18.64 or 31.74 depending only on how many series are left wandering under the null being tested — and read against the wrong one of those three, it calls unrelated random walks cointegrated most of the time.
The model that corrects its error
A cointegrated pair can always be written as a mechanism — today's change in y depends on yesterday's disagreement between y and its long-run relation with x. The coefficient of that disagreement is recovered from data that never saw it — and on unrelated series the same fit produces one a t table would call real 41% of the time.
What differencing costs
Differencing takes the false-positive rate between two unrelated walks from 76.7% to 4.9%, and takes a genuine relationship's R² from 0.91 to 0.33. Applied to a series that did not need it, it doubles the variance and installs a correlation of −0.5 that the data never had.
The check before the standard error
One number decides whether every interval in an analysis is trustworthy, and the check for it flags a lag-one correlation of 0.5 nine times in ten — and one of 0.2 only one time in five, where the interval already covers 88.6% instead of 95%.
The cost of differencing a pair
Differencing two cointegrated series makes every standard error honest and throws away the one thing known about where they are going. The error-correction model forecasts better by exactly what a closed form says — and at four hundred observations it is better on four series in five and worse on average.
The repair that keeps the question
A regression between two independent trending series is significant 82.9% of the time on random walks and 100.0% on trend-stationary ones. Subtracting a fitted line leaves 74.2% and 33.5%; differencing leaves 5.0% and 5.2% and throws away the trend the study was about.
Which mistake about the rank costs
On a system with two relations, imposing none costs 29.2% of squared forecast error and imposing three costs 2.5%. The expensive mistake is under-counting, which is the error the procedure's 5% does not bound — so the guarantee protects the cheap side.
Named alongside it
The objects these essays reach for when they reach for this one.
StationarityRandom walkAutocorrelationCointegrationThe error-correction modelCointegrating rankCommon trendCorrelationDependenceOver-differencingReduced-rank regressionSample size