Series that move together

The cost of differencing a pair

Differencing two cointegrated series makes every standard error honest and throws away the one thing known about where they are going. The error-correction model forecasts better by exactly what a closed form says — and at four hundred observations it is better on four series in five and worse on average.

Worth reading first: The observations that repeat each other.

What differencing costs measures what happens to a single series when it is differenced unnecessarily: the variance is multiplied by 2(1 − φ), the autocorrelation is driven negative, and an over-differenced series is worse to work with than the one it came from.

For a pair, differencing costs something that has no analogue in the single-series case. It removes the levels, and the relation between the levels was the only thing known about where the two series are going relative to each other.

This essay measures that, and the measurement turns out to have an awkward shape.

What the long-run relation is worth, at α = -0.2Root mean squared one-step forecast error of the error-correction model divided by that of the model fitted on differences alone; below one means the levels helped. With the equilibrium known the ratio is 0.929 at 100 observations and settles on 0.905 by 3,200, against a closed form of 0.905 that mentions no sample size at all; the excess at short series is the cost of fitting three coefficients on fifty observations. With the equilibrium estimated as well it is 1.127 at 100 — worse than differencing — and 0.914 at 3,200. The gap between the two curves is the cost of not knowing β.0.90011.1022.302.602.903.203.51length of the series (log scale, ticks at 100 … 3,200)forecast error keeping the levels, ÷ error after differencingdifferencing, for comparisonequilibrium knownequilibrium estimated250 pairs at each length, fitted on the first half and scored on the secondclosed form 0.905 with β known
Fig. 1 One-step forecast error of the error-correction model divided by that of the model fitted on differences alone; below one means keeping the levels helped. With the equilibrium known the ratio settles at 0.905 — matching a closed form that mentions no sample size. With the equilibrium estimated from the data it is 1.127 at a hundred observations, still 0.995 at four hundred, and 0.914 only by three thousand.

The comparison

Two models, both fitted on the first half of each series and scored on the second, so what is measured is a forecast rather than a fit.

The differenced model: Δy regressed on Δx and a constant. Every quantity in it is stationary, every standard error is honest, and it is the safe answer the time-series field arrived at.

The error-correction model: Δy regressed on Δx, a constant, and yesterday’s gap between y and its long-run relation with x. The same model plus one term.

The extra term is the whole difference, and it is worth being precise about what it contributes. The gap is a stationary quantity with a variance of its own, α of it is passed into Δy each step, so it explains α²·Var(gap) of the variance of Δy. That gives a closed form for the ratio of root mean squared errors:

√( η² / (η² + α²·Var(gap)) )

which involves α and the two innovation variances, and nothing about how much data there is. Having a closed form here is what makes the rest of the essay a measurement rather than an impression: without it, a measured ratio of 0.907 is a number, and with it the number has something to be right or wrong about.

Route one: the equilibrium known

At α = −0.2 the closed form gives 0.9045. Measured over four hundred pairs with β supplied rather than estimated: 0.9066 at four hundred observations, and 0.9053 by three thousand two hundred.

The measured curve approaches the closed form from above, and the excess is the cost of fitting three coefficients on the training half. At a hundred observations that half is fifty points, three coefficients from fifty observations is not free, and at the weakest correction tested — α = −0.05 — the cost is enough to put the measured ratio above one: the correction term is worth less than the noise in estimating it.

That was the first thing the assertion for this figure got wrong. Written as the measured ratio equals the closed form at every length, it failed at a hundred observations, and it deserved to: the closed form is what the model is worth once its coefficients are known, and asserting it at every length is asserting that fitting is free. What is checked instead is that the curve settles on the closed form at the longest series and approaches it monotonically, which is the claim that was actually being made.

Route two: the equilibrium estimated

Nobody has β. It has to come from the levels regression on the training half, and this is where the comparison stops being tidy.

n = 100 n = 400 n = 3,200
ratio, β known 0.929 0.907 0.905
ratio, β estimated 1.127 0.995 0.914

At four hundred observations, keeping the levels is worth half a per cent. The closed form says it should be worth nine and a half.

The mechanism is worth working through because it is not the obvious one. β̂ is out by some ε. The gap the model uses is then (true gap) − ε·x, and x is a random walk: over four hundred steps it has travelled a distance of order √400 = 20, so a β̂ out by 0.05 contributes a spurious term of order 1 to every gap — comparable to the true gap itself.

So the error-correction term is contaminated by a non-stationary quantity whose size grows like √n while β̂’s own error shrinks like 1/n. The product shrinks like 1/√n, which is why the advantage arrives eventually; and 1/√n is slow, which is why it arrives late.

A pair pulled back at 20% of the gap per step. Above, the two series. Below, the difference between them. The gap is pulled back towards zero by 20% of itself each step, so it stays inside a band of 14.3 while the series themselves travel much further. Nothing here is stationary except the difference. The faint line below is the gap for two free walks from the same seed, drawn for comparison.
Fig. 2 Why an error in β is so expensive. The lower panel is the gap, which stays inside a band of about 14 over three hundred steps. The upper panel is the levels, which travel several times further — so a β̂ out by a few per cent contributes a term to the gap comparable in size to the gap itself, and that term is a random walk rather than a stationary one.

Better on four series in five, worse on average

The pooled ratio at four hundred is 0.995. The share of individual pairs on which the error-correction model forecasts better is 80.8%, and the median per-pair ratio is 0.929.

Those are not in conflict and the disagreement is the finding.

On four pairs in five, β̂ is good enough and the model does roughly what the closed form promises. On the fifth, β̂ is bad enough that the contaminated gap makes the forecasts much worse — and a squared error that is much worse dominates an average of squared errors. The pooled ratio is a mean over series and the mean is where the tail lives.

This is the site’s own standing rule arriving in a new place. One run is an anecdote, and so is one summary: a procedure that is better four times in five and worse in expectation has two honest descriptions, and which one is quoted is a choice that should be made deliberately.

What the long-run relation is worth, at α = -0.5. Root mean squared one-step forecast error of the error-correction model divided by that of the model fitted on differences alone; below one means the levels helped. With the equilibrium known the ratio is 0.786 at 100 observations and settles on 0.775 by 3,200, against a closed form of 0.775 that mentions no sample size at all; the excess at short series is the cost of fitting three coefficients on fifty observations. With the equilibrium estimated as well it is 0.980 at 100 — worse than differencing — and 0.783 at 3,200. The gap between the two curves is the cost of not knowing β.
Fig. 3 The same comparison at a much faster correction. Both curves sit far lower — the closed form at α = −0.5 is 0.775 — and the estimated-β curve is close to the known-β one at every length, because a strong correction keeps the gap small, which makes β easier to estimate and the contamination smaller relative to what it is contaminating.
What the long-run relation is worth, at α = -0.05. Root mean squared one-step forecast error of the error-correction model divided by that of the model fitted on differences alone; below one means the levels helped. With the equilibrium known the ratio is 1.058 at 100 observations and settles on 0.976 by 3,200, against a closed form of 0.975 that mentions no sample size at all; the excess at short series is the cost of fitting three coefficients on fifty observations. With the equilibrium estimated as well it is 1.264 at 100 — worse than differencing — and 0.985 at 3,200. The gap between the two curves is the cost of not knowing β.
Fig. 4 And at a very weak correction, where the model is barely worth anything even with β known — the closed form is 0.975 — and estimating β costs more than the term is worth at the shorter lengths drawn. A relation this weak is real, is detectable given enough data, and is close to worthless for forecasting.

The three settings together give the rule, and it is a ratio rather than a threshold. What decides whether keeping the levels is worth it is how large α²·Var(gap) is relative to the innovation variance — which is to say, how much of tomorrow’s movement is the correction and how much is new. At α = −0.5 the correction is 67% of the innovation variance and the model is decisively better; at −0.2 it is 22%; at −0.05 it is 5%, and 5% of a forecast variance is not worth the parameter it takes to estimate it.

That quantity is computable in advance from an estimate of α, which the previous essay supplies. So the decision about whether to use an error-correction model for forecasting does not have to be made by trying both: fit the model, read α, and the closed form says what the term is worth before any forecast is made.

One over the root of one plus a share

The three closed forms — 0.775, 0.9045 and 0.975 at α = −0.5, −0.2 and −0.05 — and the three shares of innovation variance quoted beside them, 67%, 22% and 5%, are the same three numbers. Writing s = α²·Var(gap)/η² collapses the expression above to

ratio = 1 / √(1 + s)

and 1/√1.67 = 0.774, 1/√1.22 = 0.905, 1/√1.05 = 0.976. Three settings, one expression, no fitting.

The square root is what makes the trade unattractive at the weak end and only moderately attractive at the strong one. For small s the improvement is about s/2, so a correction carrying five per cent of the innovation variance buys two and a half per cent of forecast error, and one carrying twice as much buys twice as much — while the parameter cost of estimating it is the same in both cases. The returns are linear where the term is cheap relative to nothing and become square-root as soon as it matters: halving the forecast error would need s = 3, a correction carrying three times the innovation variance, which is a far stronger relation than any drawn here.

That is the arithmetic reason a long-run relation is worth less to a forecaster than it is to a modeller. As a description of the mechanism it is everything; as a variance reduction it enters under a square root, and everything under a square root has to be large before it is visible.

The comparison is itself a test, and not a very good one

Two of this essay’s numbers were measured for different purposes and can be read together as one diagnostic. On cointegrated pairs at four hundred observations the error-correction model forecasts better on 80.8% of pairs; on unrelated random walks it forecasts better on 22.5%.

So does keeping the levels improve the one-step forecast is a rule that separates the two kinds of pair with a sensitivity of 81% and a specificity of 78%. It gets four calls in five right on each side, which is a great deal better than nothing and is nowhere near the standard a test is normally held to.

Its virtue is that it needs no critical value. The whole difficulty of the test built for this is that its null distribution is not tabulated and has to be simulated; a forecast comparison has no such problem, because the comparison is against the other model rather than against a distribution. Its vice is that a rule with 22.5% false positives is calling one unrelated pair in five cointegrated, and the whole reason this field exists is that unrelated pairs are the common case.

The two therefore go in a definite order and not the other. Use the test, which is calibrated; use the forecast comparison afterwards to decide whether the relation is worth using, which is the question the test does not answer. A pair that passes the test and fails the comparison is a real relation too weak to forecast with, which is exactly what the α = −0.05 setting looks like — real, detectable, and worth 2.4% of a forecast error.

The pooled figure says where the failures sit. Reading 0.995 as a weighted average of the two groups — approximate, since it is formed from squared errors rather than from ratios — puts the losing fifth of pairs at about 1.27, a quarter worse rather than marginally worse. The pairs the comparison gets wrong are not near the boundary; they are pairs whose β̂ went badly enough to make the extra term actively harmful.

The other thing differencing takes

There is a second cost, older than this field, and the pair case makes it worse rather than better.

The differences of an AR(1) at φ = 0.2. Differencing a series that was already stationary does not clean it: it puts a correlation of -0.38 at lag one where the data had 0.2, and the correlation it installs is negative, which no process here produced.
Fig. 5 Over-differencing, from the time-series field: differencing a series that did not need it drives the lag-one autocorrelation to a large negative value and inflates the variance. The signature is unmistakable once looked for and invisible if not.

Δy and Δx are stationary, so nothing here is over-differenced in the single-series sense — both series genuinely had unit roots. What the differenced model does instead is discard a stationary regressor, which is the opposite mistake and has the opposite signature: the residuals of the differenced model are autocorrelated, because the omitted gap is an autocorrelated quantity that is part of Δy.

So the differenced model of a cointegrated pair is misspecified, and misspecified in a way an ordinary residual diagnostic can see. That is worth knowing because it gives a second, cheaper route to the same conclusion as the test: fit the differenced model, look at the residual correlogram, and if it shows structure the levels are carrying something.

The refusal

A comparison between two models where one is larger needs a case in which the larger one must not win, or it is measuring model size rather than model quality.

On two unrelated random walks the gap carries no information: there is no equilibrium, and y − β̂x is a random walk that β̂ was chosen to make look as stationary as possible. Adding it as a regressor is adding a spurious level to a regression on differences.

The measurement: the median per-pair ratio is 1.032 and the error-correction model is better on only 22.5% of pairs. The extra term costs about three per cent, and it costs rather than pays.

To put the two sides of the comparison beside each other: on a cointegrated pair at four hundred observations the correction term is better on 80.8% of pairs, and on an unrelated pair it is better on 22.5%. The rule separates the two cases well and does not separate them perfectly, which is what two walks and a finding would predict — a spurious level regressed on differences will sometimes appear to help, for the same reason a spurious level regressed on levels usually does.

That is a stronger refusal than “the two come out level”, and it is worth having in that stronger form. The term is not neutral when there is nothing to correct towards; it actively degrades the forecast, because it is a random walk added to a regression of stationary quantities. Which means the practical advice is not include the correction term, it can only help. It is establish that the pair is cointegrated first, and the test for that has its own critical value that no table contains.

How fast a fitted long-run relation settles down. Median absolute error in the fitted slope, over 300 pairs at each length, on log axes so a power law is a straight line. The pulled-back pair has exponent -0.96 — the error falls by a factor of ten when the series get ten times longer — against the −0.5 that every ordinary estimator obeys, drawn as the dashed reference. Two free walks give exponent -0.01: the error at 1,600 observations is 1.01 against 1.05 at 50.
Fig. 6 The rate that decides how long the wait is. β̂’s error falls like 1/n, which is fast — and what it multiplies grows like √n, so the contamination falls only like 1/√n. Superconsistency is real and does not arrive as quickly as the exponent alone suggests.

What this settles about differencing

The time-series field’s conclusion was: difference trending series before regressing them, because otherwise the regression is spurious three times in four. That conclusion is right and this field has now measured what it costs and when.

It is worth restating the shape once, because a reader who takes only the headline will take the wrong one. The theoretical case for the error-correction model is strong and the empirical case at realistic sample sizes is not, and both are true at once because the theory is about an asymptote that four hundred observations have not reached. That is not an argument against the model — the model is the right description of the data-generating process and the differenced one is not — but it is an argument against expecting a forecasting gain from it on a short series.

When the pair is not cointegrated, differencing costs nothing that matters. There is no long-run relation to throw away. The 22.5% above is the measurement of that: keeping the levels is worse.

When it is cointegrated and long, differencing costs about ten per cent of forecast accuracy at a correction speed of 20% per step, and the amount is given by a closed form in α that can be computed before the data exists.

When it is cointegrated and of moderate length, differencing costs almost nothing — not because the long-run relation is worthless but because it cannot be estimated well enough to use. At four hundred observations the whole advantage has been consumed by the error in β̂.

That third case is the one worth carrying, because it is the case most series are in and it is the opposite of what the theory suggests in isolation. Superconsistency guarantees that β̂ converges at rate 1/n, which sounds like it should make β̂ a non-issue almost immediately. It does not, because what β̂’s error multiplies is a random walk, and √n against 1/n is a race that takes a while to win.

What differencing fixes, and what it costs, 100 steps. The first pair is the false-positive rate for two independent random walks: 77% on the levels, 4.9% on the differences. The second pair is how much of a real relationship survives: R² falls from 0.91 to 0.33. The same operation does both.
Fig. 7 The time-series field’s version of the trade, for comparison. It measures what differencing does to the rejection rate and to power, and could not measure what it does to the long-run relation, because there was no model in that field containing one.

What the field leaves

Three things are settled and one is deliberately not taken.

Settled: that two kinds of trending pair exist and are distinguished by whether a combination of them is stationary; that the fitted relation converges at 1/n in one case and at nothing in the other; that the test separating them needs a critical value of −3.38 where a t table says −1.65; that the relation, once established, is a mechanism with a speed and a half-life; and that keeping it is worth a knowable fraction of a forecast error once there is enough data to estimate it.

Not taken: more than two series. Three or more can carry several independent cointegrating relations at once, counting them is a rank problem rather than a testing one, and the two-step procedure here does not extend to it. That is where the Johansen machinery starts, it is a field of its own, and the honest statement is that this one stops at a pair.

Also not taken: the reverse direction. Everything here forecasts y from x and the gap between them. A symmetric system, in which both series adjust and both have their own speed, is the natural next object — it is what says which of the two is doing the moving — and it needs a vector notation this field does not set up.

Also not taken, and worth naming because the time-series field named it too: forecasting as a craft. Everything here forecasts one step ahead, from a correctly specified model, on stationary innovations. Multi-step forecasts, model selection among lag lengths, and what to do when the specification is wrong are all real and all absent.

What the field does deliver is the thing it was opened to deliver. The time-series field ended with a caveat — unless they are genuinely tied together — that named a case and measured nothing about it. That caveat is now four essays, a library, six figures and a number: 0.905, the fraction of a differenced model’s forecast error that an error-correction model achieves when the pair is genuinely tied and the equilibrium is known, with a closed form beside the count and a case in which the whole thing is refused.

And one number that was not expected and is the more useful of the two. 0.995 — what the same model achieves at four hundred observations once the equilibrium has to be estimated, which is to say almost nothing, while still being the better model on four series in five. The gap between those two summaries is not a defect of the measurement. It is the measurement.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CointegrationDifferencingThe error-correction modelForecast errorOver-differencingRandom walkSpurious regressionStationarity