A pair pulled back at 20% of the gap per step
Above, the two series. Below, the difference between them. The gap is pulled back towards zero by 20% of itself each step, so it stays inside a band of 14.3 while the series themselves travel much further. Nothing here is stationary except the difference. The faint line below is the gap for two free walks from the same seed, drawn for comparison.
Series that move togetherslider: share of the gap undone each step, 6 positionswide7 views
What else it draws
The same object, drawn to answer the other questions the essays put to it.
Median absolute error in the fitted slope, over 300 pairs at each length, on log axes so a power law is a straight line. The pulled-back pair has exponent -0.96 — the error falls by a factor of ten when the series get ten times longer — against the −0.5 that every ordinary estimator obeys, drawn as the dashed reference. Two free walks give exponent -0.01: the error at 1,600 observations is 1.01 against 1.05 at 50.
Four thousand pairs of unrelated random walks, each regressed on the other and each residual tested for a unit root. The statistic is computed as a t and its distribution is not a t: five per cent of it falls below -3.38, where the ordinary one-sided 5% point of a t on 198 degrees of freedom is -1.65. Everything left of -1.65 — 70.2% of the whole distribution — is a pair of unrelated walks that a t table calls cointegrated.
The same statistic read against two critical values, on unrelated walks and against a real error-correction mechanism at α = -0.2. Read against a t table the test calls unrelated walks cointegrated 70.5% of the time, against the 5% it claims. Read against the simulated value it fires on 4.8% of unrelated pairs and still finds 99.7% of real ones. The rule that holds its size loses almost nothing.
Each point is one step: the gap at the end of yesterday against the change in y today. The fitted slope is -0.202 against the -0.2 the data was generated from, which means 20% of any disagreement between y and its long-run relation with x is undone in a single step. A shock therefore has a half-life of 3.1 steps. Neither series is stationary; the relation between them is.
Root mean squared one-step forecast error of the error-correction model divided by that of the model fitted on differences alone; below one means the levels helped. With the equilibrium known the ratio is 0.929 at 100 observations and settles on 0.905 by 3,200, against a closed form of 0.905 that mentions no sample size at all; the excess at short series is the cost of fitting three coefficients on fifty observations. With the equilibrium estimated as well it is 1.127 at 100 — worse than differencing — and 0.914 at 3,200. The gap between the two curves is the cost of not knowing β.
The power of the test against the half-life of a disagreement, at 100, 200, 400 observations, each read against its own simulated critical value. Every pair in every reading is genuinely tied together, so a non-rejection is a miss. At 200 observations a gap that halves in 3 steps is found 99.9% of the time and one that halves in 12 steps is found 15.3% of the time — and by 35 steps the reading is 6.1%, which is the test's own size. Beyond that the curves are flat because there is nothing left to detect with.
Where it is used
5 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 5 different questions.
- The regression that is not spurious Series that move together
- The test with no table Series that move together
- The model that corrects its error Series that move together
- The cost of differencing a pair Series that move together
- How slow a return a sample can see When the observations repeat each other