Series that move together

The regression that is not spurious

Two random walks regressed on each other are called significantly related three times in four, so the time-series field ends in a warning. The exception it names and does not measure is here — and when the pair is genuinely tied, the fitted relation converges at rate 1/n rather than the usual 1/√n.

Worth reading first: Two walks and a finding · The observations that repeat each other.

Two walks and a finding ends with a rate: regress one random walk on another and the slope is called significantly different from zero 76.7% of the time, rising to 88.0% at four hundred steps. That is the site’s one measurement that gets worse with more data, and the conclusion drawn from it was the standard one — do not regress trending series on each other in levels.

The essay then adds a sentence that has been an unpaid debt since: unless they are genuinely tied together, in which case the levels regression is the right one. That sentence names a case, says the case exists, and measures nothing about it. This field is what it should have said.

A pair pulled back at 20% of the gap per stepAbove, the two series. Below, the difference between them. The gap is pulled back towards zero by 20% of itself each step, so it stays inside a band of 14.3 while the series themselves travel much further. Nothing here is stationary except the difference. The faint line below is the gap for two free walks from the same seed, drawn for comparison.-40-30-20the two seriesα = -0.202040600100200timey − x, the gapΔy = α(y − x) + η, Δx = ε, over 300 stepsthe gap ranges over 14.3 against 55.2 for free walks
Fig. 1 A pair of series where one is pulled back towards the other, drawn above with the difference between them below. Neither series is stationary — both wander freely over the three hundred steps — but the gap stays inside a band of 14.3, against 55.2 for two unrelated walks from the same seed. Nothing here is stationary except the difference.

Two kinds of pair

The distinction the field turns on is not about the individual series. Both series in both cases are random walks: neither has a mean to return to, both wander without limit, and a unit-root test on either one of them individually finds a unit root. Measured over three hundred pairs, the average Dickey–Fuller statistic is −1.51 for one series and −1.65 for the other, nowhere near any critical value.

The distinction is about what is left when one is subtracted from the other.

For two unrelated walks, every linear combination is another random walk. The difference wanders as freely as the series do, there is no level for it to return to, and the fitted slope is chasing something that does not exist.

For a cointegrated pair, one particular combination is stationary. It has a level, it returns to it, and the fitted slope is estimating something real. The same three hundred pairs give an average statistic of −6.75 on the fitted residual, which is decisive.

That is the whole definition: two series that are individually non-stationary, whose difference — at the right ratio — is stationary.

Neither series knows anything about it

It is worth dwelling on how little the individual series reveal, because the whole difficulty of this subject follows from it.

An AR(1) at φ = 0.99, 300 observations. The bars are the measured correlations; the curve is φᵏ, which is what an AR(1) must have. The band is ±1.96/√n, where an independent series would stay. The first bar is 0.97 against a band of ±0.11.
Fig. 2 The autocorrelation of a single series with a near-unit root. The correlations decay so slowly that the picture is indistinguishable from a genuine random walk’s, which is the reason unit-root testing is hard in the first place: the alternative to a unit root, at any finite sample size, includes processes that look exactly like one.

Take one series out of a cointegrated pair and examine it as carefully as anyone likes. It has a unit root. Its correlogram decays imperceptibly. Its differences are stationary and its levels are not. Every diagnostic says random walk, and every diagnostic is right — it is a random walk.

The relationship lives entirely in the pair. There is no test that can be run on x alone, or on y alone, that says anything about whether they are cointegrated, and this is not a limitation of the tests. It is what the property is: a statement about a linear combination, invisible in any single margin.

The practical consequence is that cointegration cannot be discovered by looking at series one at a time, which is how most series are looked at.

How the pair is built

The generator makes the definition constructive rather than descriptive:

Δx = ε and Δy = α(y₋₁ − βx₋₁) + γΔx + η, with α < 0.

x wanders freely. y changes by an amount that includes a fraction α of yesterday’s disagreement between y and βx. When y has drifted above where x says it should be, the correction term is negative and pulls it back; when below, it pushes up. Neither series has a level; the gap does.

Setting α = 0 gives back the unrelated case exactly. That is why the two are one function rather than two: the null is a point inside the alternative, so nothing about any comparison in this field can turn on the two cases having been coded differently. It is the same device the randomisation null uses, and for the same reason.

Two free walks, and the gap between them. Above, the two series. Below, the difference between them. With no mechanism tying them together the gap is itself a random walk: it ranges over 55.2 in 300 steps and has no level to return to. The faint line below is the gap for two free walks from the same seed, drawn for comparison.
Fig. 3 The same generator with the correction switched off. The two series above look no more or less related than the ones in the first figure — that is the point — and the difference below is now a random walk of its own, ranging over 55.2 with no level to come back to.

The slope converges at a rate nothing else here does

The most surprising fact in this field is about how fast the estimate settles down, and it is best read as an exponent.

How fast a fitted long-run relation settles down. Median absolute error in the fitted slope, over 300 pairs at each length, on log axes so a power law is a straight line. The pulled-back pair has exponent -0.96 — the error falls by a factor of ten when the series get ten times longer — against the −0.5 that every ordinary estimator obeys, drawn as the dashed reference. Two free walks give exponent -0.01: the error at 1,600 observations is 1.01 against 1.05 at 50.
Fig. 4 Median absolute error in the fitted slope against the length of the series, on log axes so a power law is a straight line. The cointegrated pair has exponent −0.96: the error falls by a factor of ten when the series get ten times longer. The dashed reference is slope −½, which is what every other estimator on this site obeys. Two unrelated walks give −0.01 — the error at 1,600 observations is 1.01 against 1.05 at 50.

The three numbers are worth separating.

−0.5 is the ordinary rate. A sample mean, a fitted slope on independent data, a proportion: all have standard errors falling like 1/√n. Four times the data halves the error. Every estimator in the rest of this site does this.

−1 is what a cointegrating slope does. Ten times the data gives a tenth of the error. This is called superconsistency, and the reason for it is that the regressor is a random walk: the sum of squared deviations of x grows like n² rather than like n, because x itself wanders further as the series lengthens. The slope’s precision is that sum, so it grows an order faster than usual.

The middle one deserves a second look because it inverts a habit. Everywhere else on this site, a non-stationary regressor is a problem — it is what makes the spurious regression spurious. Here it is the source of the advantage. The same wandering that destroys the estimate when there is nothing to estimate makes the estimate unusually good when there is: x visiting a wide range of values is exactly what a regressor should do, and a random walk visits an ever-wider range.

That is worth putting beside the slope that borrows, where the same quantity — Σ(x − x̄)², how far the regressor ranges — decides how precisely a group measures its own slope. It is the same statement about the same sum. What is different is that here the regressor supplies its own spread, without anyone designing it, and supplies more of it the longer the study runs.

0 is what two unrelated walks give. The median error is 1.05 at fifty observations and 1.01 at sixteen hundred. Thirty-two times the data has bought nothing at all, because there is nothing to converge to: the fitted slope has a limiting distribution that does not narrow. That is the exact statement behind two walks and a finding’s rejection rate rising with n — the estimate does not improve while its nominal standard error keeps shrinking, so the ratio between them grows.

Where the ratio comes from

One detail has been glossed and it matters for everything after.

The stationary combination is y − βx for one particular β. Any other multiple of x subtracted from y leaves something non-stationary: get β wrong by ε and what remains is (the stationary part) − εx, and εx is a random walk however small ε is. There is exactly one cointegrating ratio, and being close to it is not the same as being at it.

That is what makes the superconsistency above load-bearing rather than decorative. β has to be estimated, the estimate is not the true ratio, and the residual therefore contains a leftover εx that is non-stationary — which is precisely the thing being tested for. The test works only because ε shrinks like 1/n while the leftover it multiplies grows like √n, so the product shrinks like 1/√n and vanishes.

If the slope converged at the ordinary −½ rate, εx would be O(1) and would not vanish, and the whole two-step procedure would be testing a residual permanently contaminated by the error in its own first stage. The rate is not a nice property of this problem. It is the reason the problem is tractable.

Why the exponent is the right thing to assert

The three median errors could each have been asserted separately: 0.354 at fifty, 0.105 at two hundred, 0.025 at eight hundred. Any of those on its own is a number a wrong implementation could match by accident, and all three together are three numbers to keep in step.

Fitting the exponent turns them into one statement that is hard to satisfy without being right. Getting −0.96 requires the error to fall by a factor of ten when n does, across six lengths spanning a factor of thirty-two. An implementation that had, say, computed the slope correctly but generated the data without the correction term would give −0.01; one that had made the series stationary by mistake would give −0.5. Each of those is a specific, recognisable failure and the exponent separates all three.

That is a general habit worth naming. Where a claim is about a rate, assert the rate rather than the values it produces. A rate is one number, it is comparable across settings, and — because it is a property of a whole sequence — it is far harder to hit by coincidence than any single value in it.

The speed the gap closes at

α is not only a switch between two cases. Its size is a quantity with a reading.

A pair pulled back at 80% of the gap per step. Above, the two series. Below, the difference between them. The gap is pulled back towards zero by 80% of itself each step, so it stays inside a band of 8.5 while the series themselves travel much further. Nothing here is stationary except the difference. The faint line below is the gap for two free walks from the same seed, drawn for comparison.
Fig. 5 The same pair with 80% of the gap undone each step instead of 20%. The two series still wander as far as before, and the difference between them is now pinned tightly to zero. The band is a fraction of the first figure’s.

A correction of α means a fraction |α| of any disagreement is removed in one step, so what remains after k steps is (1 + α)ᵏ, and the half-life of a shock is log(0.5)/log(1 + α). At α = −0.2 that is 3.1 steps; at −0.8 it is 0.4; at −0.05 it is 13.5.

That number is what makes cointegration a substantive claim rather than an asymptotic one. Two series tied together with a half-life of thirteen steps are, over any span of twenty observations, effectively untied — the correction is real and too slow to see. A test on a short series will not find it, and a model that assumes it will be relying on something the data cannot support.

How fast a fitted long-run relation settles down. Median absolute error in the fitted slope, over 300 pairs at each length, on log axes so a power law is a straight line. The pulled-back pair has exponent -0.79 — the error falls by a factor of ten when the series get ten times longer — against the −0.5 that every ordinary estimator obeys, drawn as the dashed reference. Two free walks give exponent -0.01: the error at 1,600 observations is 1.01 against 1.05 at 50.
Fig. 6 The convergence exponent at a very weak correction. It is still negative and still far from zero, but the errors at every length are much larger: superconsistency is an asymptotic property and a weak correction pushes the length at which it starts to be visible a long way out.

What superconsistency is worth

The rate is not a curiosity; it changes what a two-stage procedure is allowed to do.

Ordinarily, using an estimate from a first regression as though it were known in a second regression invalidates the second one’s standard errors — the first stage’s error propagates and nothing accounts for it. That is the same defect the hierarchical plug-in has, and it costs sixteen points of coverage there.

Here it does not, and superconsistency is exactly why. The first stage’s error vanishes at rate 1/n while everything in the second stage converges at 1/√n, so by the time the second stage is precise enough to notice the first stage’s error, the first stage’s error has already gone. Asymptotically the second stage behaves as though β were known.

That licence is what makes the two-step procedure in the next essays legitimate, and it is the only place in this site where estimating something and substituting it is allowed without penalty. The condition is narrow and worth remembering: it holds because the regressor is non-stationary, and it does not hold anywhere else.

The finite-sample version is less generous, and the cost of differencing a pair measures it. At four hundred observations the estimated equilibrium is still poor enough to cost most of the model’s advantage. The asymptotic licence is real and it arrives later than the asymptotics suggest.

The exponent arrives, and roughly when

The three median errors quoted for the fitted exponent — 0.354 at fifty observations, 0.105 at two hundred, 0.025 at eight hundred — carry one more piece of information than the single number fitted through them, and it is the piece that says when the asymptotics start.

Taken in pairs, the local exponent between fifty and two hundred is log(0.354/0.105)/log(4) = −0.88, and between two hundred and eight hundred it is log(0.105/0.025)/log(4) = −1.04. The rate is not constant across the range: it is short of −1 on the shorter series and at or just past it on the longer ones, and the −0.96 fitted over the whole sweep is an average of a quantity still settling.

That is the honest reading and it is more useful than the headline. Superconsistency is an asymptotic property, the sweep catches it arriving, and the arrival is somewhere in the low hundreds of observations for a correction of α = −0.2. Below that the estimate is converging faster than an ordinary estimator and slower than the theory promises; above it, the theory is delivering.

It also puts a floor under the two-step licence. The argument that the second stage may treat β̂ as known rests on the first stage converging an order faster, and at fifty observations the first stage is converging at −0.88 against the second stage’s −0.5 — a gap of 0.38 rather than 0.5, so the contamination is dying more slowly than the asymptotic account says. That is one of the reasons the finite-sample cost measured elsewhere in this field is as large as it is at four hundred observations, and it is visible here, in the estimator itself, before any forecast is made.

What the extra half an exponent is worth

The difference between −0.5 and −1 is easy to state and hard to feel, so it is worth converting into the quantity anybody actually spends: data.

To halve the error, an ordinary estimator needs four times the observations and a cointegrating slope needs two. Compounded over the sweep drawn here, that is a large number: going from fifty to sixteen hundred observations is a factor of thirty-two, which buys an ordinary estimator a factor of √32 = 5.7 in accuracy and buys this one the full 32. Two estimators starting level at fifty observations are five and a half times apart by sixteen hundred, with nothing separating them but the exponent.

The unrelated pair is the third term in that comparison and it is not a slower rate, it is no rate. Its median error runs 1.05 at fifty observations and 1.01 at sixteen hundred; an ordinary estimator starting at 1.05 would have reached 0.186 over the same thirty-twofold increase. The gap between what the slope does and what an ordinary estimator would have done is a factor of five in one direction and a factor of five in the other, from the same regression on data that looks the same.

That symmetry is the reason this field is worth having as a positive claim rather than as a footnote to a warning. The levels regression between trending series is not a technique that is usually wrong and occasionally right. It is a technique that is the best estimator on this site when the pair is tied and among the worst when it is not, with no middle ground and nothing in the scatter plot to say which.

Where the two cases actually differ, in one sentence

Every measurement in this essay reduces to one thing and it is worth having it in one place.

For two unrelated walks, the sum Σ(x − x̄)(y − ȳ) is dominated by the shared wandering of two things that have nothing to do with each other, and it does not settle down. For a cointegrated pair, that same sum is dominated by the genuine relation, and the deviation from it is stationary and therefore bounded. Everything else — the exponent, the rejection rates, the residual test, the forecast comparison — follows from which of those two the numerator of the fitted slope is doing.

Two independent random walks, 100 steps. Nothing connects these two series: each is generated from its own independent draws. Regressing one on the other gives a slope with t = -10.9, R² = 0.55 and p = 0.0e+0 — a result that would be reported as a finding by any standard output.
Fig. 7 The failure case as the time-series field drew it: two independent random walks and the regression between them. This is the picture that produced the 76.7%, and the point of this field is that the same picture, with a correction mechanism behind it, is a valid regression.

That also explains why the distinction is invisible to the eye. Both cases produce a scatter plot of y against x that looks like a relationship, because in both cases the two series went up and down together. What differs is whether the deviations from the fitted line accumulate or return, and deviations from a fitted line are exactly what a scatter plot is worst at showing.

What has to come next

The claim so far is that two kinds of pair exist and behave completely differently. What is missing is the ability to tell which of the two a given pair of series belongs to.

The natural test is a unit-root test on the fitted residual: fit the levels regression, take what is left, and ask whether it returns to a level. The statistic is computed as an ordinary t, and the temptation is to read it against an ordinary t table.

That would be a catastrophe, and the size of it is the next essay: at two hundred observations, the t table’s 5% point calls two unrelated random walks cointegrated 70.5% of the time. The reason is not subtle — the regression chose β to make the residual look as stationary as it possibly could, so the residual is not a series anyone handed over, it is the most mean-reverting combination available — but the consequence is that no table anyone owns contains the right critical value, and it has to be simulated.

What each rule rejects, and what it should. The same statistic read against two critical values, on unrelated walks and against a real error-correction mechanism at α = -0.4. Read against a t table the test calls unrelated walks cointegrated 70.5% of the time, against the 5% it claims. Read against the simulated value it fires on 4.8% of unrelated pairs and still finds 100.0% of real ones. The rule that holds its size loses almost nothing.
Fig. 8 What is at stake, before the next essay works out why. The same statistic read against two critical values: one of them holds its size at 5%, the other fires on most of the null, and both find every real mechanism. The choice of table is the whole difference between a test and a formality.

There is a second thing missing and it is more useful than the test. A pair that is cointegrated can always be written as an error-correction model — the generator above is one — and that representation is not a convenience of this site’s simulation. It is Granger’s theorem, and it says that the two descriptions are equivalent: a stationary combination exists if and only if there is a mechanism correcting deviations from it. So a cointegrated pair always has an α, the α is estimable, and its size says how fast the relation reasserts itself. That is what the model that corrects its error takes.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CointegrationConvergence rateLeast squaresRandom walkSpurious regressionStationarityUnit root