Concept

Spurious regression — where it appears

A significant relationship between two independent series that arises because both wander. It appears in most samples rather than occasionally, so a small p-value from such a regression is evidence about the series' persistence rather than about any relationship.

Named by 10 essays across 3 fields — each of them below, with the objects they name alongside it.

A pair pulled back at 20% of the gap per step. Above, the two series. Below, the difference between them. The gap is pulled back towards zero by 20% of itself each step, so it stays inside a band of 14.3 while the series themselves travel much further. Nothing here is stationary except the difference. The faint line below is the gap for two free walks from the same seed, drawn for comparison.

The regression that is not spurious

Two random walks regressed on each other are called significantly related three times in four, so the time-series field ends in a warning. The exception it names and does not measure is here — and when the pair is genuinely tied, the fitted relation converges at rate 1/n rather than the usual 1/√n.

cointegration · Spurious
Where the residual test's statistic actually falls, at n = 200. Four thousand pairs of unrelated random walks, each regressed on the other and each residual tested for a unit root. The statistic is computed as a t and its distribution is not a t: five per cent of it falls below -3.38, where the ordinary one-sided 5% point of a t on 198 degrees of freedom is -1.65. Everything left of -1.65 — 70.2% of the whole distribution — is a pair of unrelated walks that a t table calls cointegrated.

The test with no table

The statistic that separates a real long-run relation from a spurious one is computed as a t and is not a t. At two hundred observations its 5% point is −3.38 where the t table says −1.65, and reading it against the table calls two unrelated random walks cointegrated 70.5% of the time.

cointegration · Spurious
Two independent random walks, 100 steps. Nothing connects these two series: each is generated from its own independent draws. Regressing one on the other gives a slope with t = -10.9, R² = 0.55 and p = 0.0e+0 — a result that would be reported as a finding by any standard output.

Two walks and a finding

Regress one random walk on another, independently generated, and the slope is significant 76.7% of the time with a median R² of 0.17. Nothing connects the two series, nothing in the output says so, and more data makes it worse.

timeseries · Spurious
The same data, one regression per choice of left-hand side. The two-step procedure has to put one series on the left, and with 3 series there are 3 ways to do it. Each returns a relation and a residual test; the 5% point is -3.71, simulated. Here they do not agree: 2 of 3 reject, and the relations they report are written with a 1 in the position of whichever series was on the left, so they can be compared. Nothing in a printed output records which regression was run.

Which series goes on the left

The two-step procedure has to pick a series to regress the others on, and nothing in its output records which. With a pair that choice never changes the verdict. With three series and one relation between them, the three choices disagree about whether the system is cointegrated at all 98.0% of the time.

systems · Rank
The correction, at a generating α of -0.2. Each point is one step: the gap at the end of yesterday against the change in y today. The fitted slope is -0.202 against the -0.2 the data was generated from, which means 20% of any disagreement between y and its long-run relation with x is undone in a single step. A shock therefore has a half-life of 3.1 steps. Neither series is stationary; the relation between them is.

The model that corrects its error

A cointegrated pair can always be written as a mechanism — today's change in y depends on yesterday's disagreement between y and its long-run relation with x. The coefficient of that disagreement is recovered from data that never saw it — and on unrelated series the same fit produces one a t table would call real 41% of the time.

cointegration · Dependence
What the long-run relation is worth, at α = -0.2. Root mean squared one-step forecast error of the error-correction model divided by that of the model fitted on differences alone; below one means the levels helped. With the equilibrium known the ratio is 0.929 at 100 observations and settles on 0.905 by 3,200, against a closed form of 0.905 that mentions no sample size at all; the excess at short series is the cost of fitting three coefficients on fifty observations. With the equilibrium estimated as well it is 1.127 at 100 — worse than differencing — and 0.914 at 3,200. The gap between the two curves is the cost of not knowing β.

The cost of differencing a pair

Differencing two cointegrated series makes every standard error honest and throws away the one thing known about where they are going. The error-correction model forecasts better by exactly what a closed form says — and at four hundred observations it is better on four series in five and worse on average.

cointegration · Dependence
Every equation's adjustment speed, and the one number they make together. Each series gets its own equation, each is regressed on the same lagged disequilibrium, and what comes back is the whole vector α. Averaged over 400 systems at n = 300: α₁ = -0.154 against -0.15 generated, α₂ = 0.104 against 0.1 generated. The gap closes at the combination of them rather than at any one entry — 25% of any disagreement per step, a half-life of 2.41 steps, where the single equation that fits only the first series reports 4.27.

Which series does the moving

“y adjusts towards x” and “x adjusts towards y” are different mechanisms with identical long-run relations, and a single-equation model cannot tell them apart because it only writes one equation. Writing all of them recovers a vector — and a gap that closes at 25% a step where one equation alone reports 15%.

systems · Adjustment
The damage and the warning, against the same dial. Two readings at each persistence. In the darker colour, how often a regression between two independent series of 200 steps is called significant at 5%: 4.9% at φ = 0, 34.2% at 0.8, 52.4% at 0.9, 83.4% at a unit root. In the lighter, how often the standard unit-root test refuses a unit root on one of those series — the chance the analyst is told the series is stationary and may be regressed: 87.2% at φ = 0.9 and 31.9% at 0.95. At φ = 0.9 both are high at once, which is a correct diagnostic licensing a regression that is wrong half the time.

The cliff that is a slope

A regression between two independent series is called significant 4.9% of the time at no persistence, 52.4% at a lag-one correlation of 0.9, and 83.4% at a unit root. The rule the field offers asks whether the last of those holds, and at 0.9 the unit-root test correctly refuses one 87.2% of the time.

timeseries · Spurious
Six cells, and 5% is the right answer in all of them. How often a regression between two independently generated series is called significant at the 5% level, for two worlds and three treatments, at 200 observations. Every pair is independent by construction, so 5% is correct everywhere and every other reading is a failure. Untreated: 82.9% and 100.0%. With a fitted line removed: 74.2% and 33.5%. Differenced: 5.0% and 5.2%. The treatment that controls the rate in both worlds is the one that discards the level and the trend, which is the quantity a study of trending series was about.

The repair that keeps the question

A regression between two independent trending series is significant 82.9% of the time on random walks and 100.0% on trend-stationary ones. Subtracting a fitted line leaves 74.2% and 33.5%; differencing leaves 5.0% and 5.2% and throws away the trend the study was about.

timeseries · Spurious
How fast a gap has to close before a sample can see it close. The power of the test against the half-life of a disagreement, at 100, 200, 400 observations, each read against its own simulated critical value. Every pair in every reading is genuinely tied together, so a non-rejection is a miss. At 200 observations a gap that halves in 3 steps is found 99.9% of the time and one that halves in 12 steps is found 15.3% of the time — and by 35 steps the reading is 6.1%, which is the test's own size. Beyond that the curves are flat because there is nothing left to detect with.

How slow a return a sample can see

At two hundred observations the test finds a gap that halves in five steps four times in five, one that halves in eight 37.3% of the time, and one that halves in fifty 4.95% of the time — which is the rate at which it finds pairs with no mechanism at all. The boundary moves with the sample, not with its square root.

timeseries · Spurious

Named alongside it

The objects these essays reach for when they reach for this one.

Random walkStationarityCointegrationUnit rootThe error-correction modelAutocorrelationCritical valueDifferencingFalse positiveMonte CarloStatistical powerAugmented dickey fuller

All concepts