The cost of differencing a pair
Worth reading first: The observations that repeat each other.
What differencing costs measures what happens to a single series when it is differenced unnecessarily: the variance is multiplied by 2(1 − φ), the autocorrelation is driven negative, and an over-differenced series is worse to work with than the one it came from.
For a pair, differencing costs something that has no analogue in the single-series case. It removes the levels, and the relation between the levels was the only thing known about where the two series are going relative to each other.
This essay measures that, and the measurement turns out to have an awkward shape.
The comparison
Two models, both fitted on the first half of each series and scored on the second, so what is measured is a forecast rather than a fit.
The differenced model: Δy regressed on Δx and a constant. Every quantity in it is stationary, every standard error is honest, and it is the safe answer the time-series field arrived at.
The error-correction model: Δy regressed on Δx, a constant, and yesterday’s gap between y and its long-run relation with x. The same model plus one term.
The extra term is the whole difference, and it is worth being precise about what it contributes. The gap is a stationary quantity with a variance of its own, α of it is passed into Δy each step, so it explains α²·Var(gap) of the variance of Δy. That gives a closed form for the ratio of root mean squared errors:
√( η² / (η² + α²·Var(gap)) )
which involves α and the two innovation variances, and nothing about how much data there is. Having a closed form here is what makes the rest of the essay a measurement rather than an impression: without it, a measured ratio of 0.907 is a number, and with it the number has something to be right or wrong about.
Route one: the equilibrium known
At α = −0.2 the closed form gives 0.9045. Measured over four hundred pairs with β supplied rather than estimated: 0.9066 at four hundred observations, and 0.9053 by three thousand two hundred.
The measured curve approaches the closed form from above, and the excess is the cost of fitting three coefficients on the training half. At a hundred observations that half is fifty points, three coefficients from fifty observations is not free, and at the weakest correction tested — α = −0.05 — the cost is enough to put the measured ratio above one: the correction term is worth less than the noise in estimating it.
That was the first thing the assertion for this figure got wrong. Written as the measured ratio equals the closed form at every length, it failed at a hundred observations, and it deserved to: the closed form is what the model is worth once its coefficients are known, and asserting it at every length is asserting that fitting is free. What is checked instead is that the curve settles on the closed form at the longest series and approaches it monotonically, which is the claim that was actually being made.
Route two: the equilibrium estimated
Nobody has β. It has to come from the levels regression on the training half, and this is where the comparison stops being tidy.
| n = 100 | n = 400 | n = 3,200 | |
|---|---|---|---|
| ratio, β known | 0.929 | 0.907 | 0.905 |
| ratio, β estimated | 1.127 | 0.995 | 0.914 |
At four hundred observations, keeping the levels is worth half a per cent. The closed form says it should be worth nine and a half.
The mechanism is worth working through because it is not the obvious one. β̂ is out by some ε. The gap the model uses is then (true gap) − ε·x, and x is a random walk: over four hundred steps it has travelled a distance of order √400 = 20, so a β̂ out by 0.05 contributes a spurious term of order 1 to every gap — comparable to the true gap itself.
So the error-correction term is contaminated by a non-stationary quantity whose size grows like √n while β̂’s own error shrinks like 1/n. The product shrinks like 1/√n, which is why the advantage arrives eventually; and 1/√n is slow, which is why it arrives late.
Better on four series in five, worse on average
The pooled ratio at four hundred is 0.995. The share of individual pairs on which the error-correction model forecasts better is 80.8%, and the median per-pair ratio is 0.929.
Those are not in conflict and the disagreement is the finding.
On four pairs in five, β̂ is good enough and the model does roughly what the closed form promises. On the fifth, β̂ is bad enough that the contaminated gap makes the forecasts much worse — and a squared error that is much worse dominates an average of squared errors. The pooled ratio is a mean over series and the mean is where the tail lives.
This is the site’s own standing rule arriving in a new place. One run is an anecdote, and so is one summary: a procedure that is better four times in five and worse in expectation has two honest descriptions, and which one is quoted is a choice that should be made deliberately.
The three settings together give the rule, and it is a ratio rather than a threshold. What decides whether keeping the levels is worth it is how large α²·Var(gap) is relative to the innovation variance — which is to say, how much of tomorrow’s movement is the correction and how much is new. At α = −0.5 the correction is 67% of the innovation variance and the model is decisively better; at −0.2 it is 22%; at −0.05 it is 5%, and 5% of a forecast variance is not worth the parameter it takes to estimate it.
That quantity is computable in advance from an estimate of α, which the previous essay supplies. So the decision about whether to use an error-correction model for forecasting does not have to be made by trying both: fit the model, read α, and the closed form says what the term is worth before any forecast is made.
One over the root of one plus a share
The three closed forms — 0.775, 0.9045 and 0.975 at α = −0.5, −0.2 and −0.05 — and the three shares of innovation variance quoted beside them, 67%, 22% and 5%, are the same three numbers. Writing s = α²·Var(gap)/η² collapses the expression above to
ratio = 1 / √(1 + s)
and 1/√1.67 = 0.774, 1/√1.22 = 0.905, 1/√1.05 = 0.976. Three settings, one expression, no fitting.
The square root is what makes the trade unattractive at the weak end and only moderately attractive at the strong one. For small s the improvement is about s/2, so a correction carrying five per cent of the innovation variance buys two and a half per cent of forecast error, and one carrying twice as much buys twice as much — while the parameter cost of estimating it is the same in both cases. The returns are linear where the term is cheap relative to nothing and become square-root as soon as it matters: halving the forecast error would need s = 3, a correction carrying three times the innovation variance, which is a far stronger relation than any drawn here.
That is the arithmetic reason a long-run relation is worth less to a forecaster than it is to a modeller. As a description of the mechanism it is everything; as a variance reduction it enters under a square root, and everything under a square root has to be large before it is visible.
The comparison is itself a test, and not a very good one
Two of this essay’s numbers were measured for different purposes and can be read together as one diagnostic. On cointegrated pairs at four hundred observations the error-correction model forecasts better on 80.8% of pairs; on unrelated random walks it forecasts better on 22.5%.
So does keeping the levels improve the one-step forecast is a rule that separates the two kinds of pair with a sensitivity of 81% and a specificity of 78%. It gets four calls in five right on each side, which is a great deal better than nothing and is nowhere near the standard a test is normally held to.
Its virtue is that it needs no critical value. The whole difficulty of the test built for this is that its null distribution is not tabulated and has to be simulated; a forecast comparison has no such problem, because the comparison is against the other model rather than against a distribution. Its vice is that a rule with 22.5% false positives is calling one unrelated pair in five cointegrated, and the whole reason this field exists is that unrelated pairs are the common case.
The two therefore go in a definite order and not the other. Use the test, which is calibrated; use the forecast comparison afterwards to decide whether the relation is worth using, which is the question the test does not answer. A pair that passes the test and fails the comparison is a real relation too weak to forecast with, which is exactly what the α = −0.05 setting looks like — real, detectable, and worth 2.4% of a forecast error.
The pooled figure says where the failures sit. Reading 0.995 as a weighted average of the two groups — approximate, since it is formed from squared errors rather than from ratios — puts the losing fifth of pairs at about 1.27, a quarter worse rather than marginally worse. The pairs the comparison gets wrong are not near the boundary; they are pairs whose β̂ went badly enough to make the extra term actively harmful.
The other thing differencing takes
There is a second cost, older than this field, and the pair case makes it worse rather than better.
Δy and Δx are stationary, so nothing here is over-differenced in the single-series sense — both series genuinely had unit roots. What the differenced model does instead is discard a stationary regressor, which is the opposite mistake and has the opposite signature: the residuals of the differenced model are autocorrelated, because the omitted gap is an autocorrelated quantity that is part of Δy.
So the differenced model of a cointegrated pair is misspecified, and misspecified in a way an ordinary residual diagnostic can see. That is worth knowing because it gives a second, cheaper route to the same conclusion as the test: fit the differenced model, look at the residual correlogram, and if it shows structure the levels are carrying something.
The refusal
A comparison between two models where one is larger needs a case in which the larger one must not win, or it is measuring model size rather than model quality.
On two unrelated random walks the gap carries no information: there is no equilibrium, and y − β̂x is a random walk that β̂ was chosen to make look as stationary as possible. Adding it as a regressor is adding a spurious level to a regression on differences.
The measurement: the median per-pair ratio is 1.032 and the error-correction model is better on only 22.5% of pairs. The extra term costs about three per cent, and it costs rather than pays.
To put the two sides of the comparison beside each other: on a cointegrated pair at four hundred observations the correction term is better on 80.8% of pairs, and on an unrelated pair it is better on 22.5%. The rule separates the two cases well and does not separate them perfectly, which is what two walks and a finding would predict — a spurious level regressed on differences will sometimes appear to help, for the same reason a spurious level regressed on levels usually does.
That is a stronger refusal than “the two come out level”, and it is worth having in that stronger form. The term is not neutral when there is nothing to correct towards; it actively degrades the forecast, because it is a random walk added to a regression of stationary quantities. Which means the practical advice is not include the correction term, it can only help. It is establish that the pair is cointegrated first, and the test for that has its own critical value that no table contains.
What this settles about differencing
The time-series field’s conclusion was: difference trending series before regressing them, because otherwise the regression is spurious three times in four. That conclusion is right and this field has now measured what it costs and when.
It is worth restating the shape once, because a reader who takes only the headline will take the wrong one. The theoretical case for the error-correction model is strong and the empirical case at realistic sample sizes is not, and both are true at once because the theory is about an asymptote that four hundred observations have not reached. That is not an argument against the model — the model is the right description of the data-generating process and the differenced one is not — but it is an argument against expecting a forecasting gain from it on a short series.
When the pair is not cointegrated, differencing costs nothing that matters. There is no long-run relation to throw away. The 22.5% above is the measurement of that: keeping the levels is worse.
When it is cointegrated and long, differencing costs about ten per cent of forecast accuracy at a correction speed of 20% per step, and the amount is given by a closed form in α that can be computed before the data exists.
When it is cointegrated and of moderate length, differencing costs almost nothing — not because the long-run relation is worthless but because it cannot be estimated well enough to use. At four hundred observations the whole advantage has been consumed by the error in β̂.
That third case is the one worth carrying, because it is the case most series are in and it is the opposite of what the theory suggests in isolation. Superconsistency guarantees that β̂ converges at rate 1/n, which sounds like it should make β̂ a non-issue almost immediately. It does not, because what β̂’s error multiplies is a random walk, and √n against 1/n is a race that takes a while to win.
What the field leaves
Three things are settled and one is deliberately not taken.
Settled: that two kinds of trending pair exist and are distinguished by whether a combination of them is stationary; that the fitted relation converges at 1/n in one case and at nothing in the other; that the test separating them needs a critical value of −3.38 where a t table says −1.65; that the relation, once established, is a mechanism with a speed and a half-life; and that keeping it is worth a knowable fraction of a forecast error once there is enough data to estimate it.
Not taken: more than two series. Three or more can carry several independent cointegrating relations at once, counting them is a rank problem rather than a testing one, and the two-step procedure here does not extend to it. That is where the Johansen machinery starts, it is a field of its own, and the honest statement is that this one stops at a pair.
Also not taken: the reverse direction. Everything here forecasts y from x and the gap between them. A symmetric system, in which both series adjust and both have their own speed, is the natural next object — it is what says which of the two is doing the moving — and it needs a vector notation this field does not set up.
Also not taken, and worth naming because the time-series field named it too: forecasting as a craft. Everything here forecasts one step ahead, from a correctly specified model, on stationary innovations. Multi-step forecasts, model selection among lag lengths, and what to do when the specification is wrong are all real and all absent.
What the field does deliver is the thing it was opened to deliver. The time-series field ended with a caveat — unless they are genuinely tied together — that named a case and measured nothing about it. That caveat is now four essays, a library, six figures and a number: 0.905, the fraction of a differenced model’s forecast error that an error-correction model achieves when the pair is genuinely tied and the equilibrium is known, with a closed form beside the count and a case in which the whole thing is refused.
And one number that was not expected and is the more useful of the two. 0.995 — what the same model achieves at four hundred observations once the equilibrium has to be estimated, which is to say almost nothing, while still being the better model on four series in five. The gap between those two summaries is not a defect of the measurement. It is the measurement.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Three series and a count — both name cointegration, differencing, the error-correction model, random walk, stationarity
- Which mistake about the rank costs — both name cointegration, differencing, the error-correction model, forecast error, random walk
- Which series does the moving — both name cointegration, the error-correction model, random walk, spurious regression, stationarity
- How slow a return a sample can see — both name cointegration, the error-correction model, random walk, spurious regression
- Which series goes on the left — both name cointegration, random walk, spurious regression, stationarity
- Counting what is still wandering — both name differencing, random walk, stationarity
Named objects
A flat tag is an object no other essay names yet.
CointegrationDifferencingThe error-correction modelForecast errorOver-differencingRandom walkSpurious regressionStationarity