Series that move together

A relation supplied in advance

When theory supplies the long-run coefficient, the gap between two series can be tested directly and the search the Engle–Granger test pays for disappears: its 5% point moves from −3.38 to −2.89, and at two hundred observations the slowest return found four times in five goes from a half-life of 5.05 steps to 6.88. The price is in what a rejection means. With a gap that halves in three steps, a coefficient a fifth too large is confirmed 99.1% of the time and one half again too large 81.9%, though every such gap contains a random walk — and at a thousand observations the second is confirmed more often, not less.

Worth reading first: Three series and a count.

How slow a return a sample can see measured the boundary of the test that separates two series genuinely tied together from two that merely wander side by side. At two hundred observations the Engle–Granger test finds a gap that halves in five steps four times in five and a gap that halves in fifty at the rate it finds pairs with no mechanism at all. It ended on what a study can do about a relation slower than its sample can see, and the first response it named was to stop estimating the relation.

Where theory supplies the long-run coefficient — purchasing-power parity says an exchange rate should move one for one with relative prices, a no-arbitrage argument fixes a spread at one for one, a mass balance fixes a stoichiometric ratio — the gap between the two series is an observable series in its own right, and whether it returns is a question about one series rather than about the best combination of two. That was expected to recover much of the boundary. It recovers some. And it changes what a rejection is evidence of, in a direction that matters more than the power it buys.

How slow a return each test can see at 200 observations: the relation estimated against the relation suppliedThe estimated relation's test reaches 80% power up to a half-life of 5.05 steps and the supplied relation's up to 6.88. At a half-life of 8 the two find the pair 36.3% and 71.9% of the time; at 12, 16.9% and 39.3%. Their 5% points are −3.38 and −2.89.00.2500.5000.7501half-life of the gap, in steps (log scale)share of tied pairs the test finds at 5%1235812203560100the coefficient suppliedthe coefficient estimated1,500 pairs a point, each test at its own 5% pointdashed: 80% power
Fig. 1 The share of genuinely tied pairs each test finds at 5%, against the half-life of the gap, at two hundred observations: the Engle–Granger test, which estimates the relation, and the Dickey–Fuller test on the gap the true coefficient implies, on the same pairs. The dashed line marks 80%. The slider sets the sample length.

What the search costs

The Engle–Granger test fits the relation first, choosing the coefficient that makes the residual as small — and therefore as stationary-looking — as it can be, and then tests the residual for a unit root. The test with no table measured what that first step does to the second: the statistic looks like a t and is not one, and its 5% point at two hundred observations is −3.38 where a t table says −1.65.

Supplying the coefficient removes the first step. The gap yt−β0xty_t - \beta_0 x_t is computed rather than fitted, and the test on it is the ordinary Dickey–Fuller test with an intercept, whose 5% point is −2.89. That is still far from the t table, because a unit-root test is never read against a t, but it is half a unit to the right of where the search put the Engle–Granger statistic’s.

What choosing the relation costs the test: two null distributions at 200 observations. On unrelated walks, the Engle–Granger statistic's 5% point is −3.38; the Dickey–Fuller statistic on the gap a supplied coefficient implies has its 5% point at −2.89; the t table says −1.65. The distance between the first two is the charge for having chosen the coefficient to make the residual look stationary.
Fig. 2 The two statistics’ distributions on pairs of unrelated walks, twenty thousand each, with their 5% points and the t table’s. The Engle–Granger statistic sits further left because its coefficient was chosen to make the residual look stationary.

The distance between the two histograms is exactly what the search was charged. A statistic built by choosing the most stationary-looking combination of two walks has, even when nothing ties them, a better-looking residual than any combination fixed in advance, and the critical value has to move left until it stops rewarding that. Everything the estimated test loses in power relative to the imposed one is that charge being paid.

What the charge was worth in power

At two hundred observations, on the same pairs, the imposed test finds a gap that halves in eight steps 71.9% of the time where the estimated test finds it 36.3%, and a gap that halves in twelve 39.3% against 16.9%. The half-life at which each reaches 80% power moves from 5.05 steps to 6.88, and the half-life at which each reaches half, from 6.92 to 10.50.

So supplying the coefficient roughly doubles the power in the middle of the range and moves the boundary by about a third. It is a real gain and a modest one. At a hundred observations the 80% half-life goes from 2.34 to 3.40, and at four hundred from 10.11 to 14.24: the imposed test’s boundary scales with the sample the way the estimated test’s does, about proportionally, with a fixed advantage of about 40% in the half-life it can see.

Both tests hold their level. On pairs of unrelated walks the estimated test rejects 4.93% of the time and the imposed test 5.40%, each against its own simulated critical value, so the power curves are comparable.

The boundary the essay on detectable half-lives drew was mostly not the search’s doing. Removing the search entirely moves a detectable half-life of five steps to seven. What remains is the difficulty of telling a series that returns very slowly from one that does not return at all, in a finite sample, which is a property of the evidence and not of the procedure. No choice of test crosses it.

The relations a rejection confirms

The gain has a condition attached, and it is the one that matters. The imposed test assumes the supplied coefficient is right. If it is not, the gap it tests is yt−β0xt=(yt−βxt)−(β0−β)xty_t - \beta_0 x_t = (y_t - \beta x_t) - (\beta_0 - \beta)x_t: the true stationary gap minus a multiple of a random walk. That gap has a unit root. The correct verdict on it is “does not return”, and the test should say so.

It mostly does not.

Which supplied relations the test confirms, at 200 observations, when the true coefficient is one. With a gap that halves in three steps, a supplied coefficient of 1.2 is confirmed 99.1% of the time, 1.5 81.9% and 2 48.5%, though every one of those gaps contains a random walk. With a gap that halves in eight, the true coefficient is confirmed 68.3% of the time and 1.5 45.9%.
Fig. 3 How often the imposed test calls the supplied gap stationary at 5%, against the coefficient supplied, when the true coefficient is one, for a gap that halves in three steps and one that halves in eight, at two hundred observations. Every coefficient except one leaves a random walk in the gap.

With a gap that halves in three steps, a coefficient of 1.2 — a fifth too large — is confirmed 99.1% of the time, the same as the true one. A coefficient of 1.5 is confirmed 81.9% of the time, and a coefficient of two, twice the truth, 48.5%. With a slower gap, halving in eight steps, the true coefficient is confirmed 68.3% of the time and 1.5 is confirmed 45.9%. The curve is flat-topped and wide: the test cannot tell the relation it was handed from any relation near it, and what counts as near is most of the plausible range.

Why the wrong gap looks stationary

The reason is visible in a single pair.

One pair, two supplied coefficients: the gap at the true coefficient and at 1.5. A pair whose gap halves in 3 steps. At the true coefficient the gap's Dickey–Fuller statistic is −5.11; at 1.5 it is −4.86, against a 5% point near −2.9. The second gap is the first minus 0.5 times the wandering series, drawn dashed: a random walk that the fast return around it hides from the test.
Fig. 4 One pair whose gap halves in three steps: the gap at the true coefficient, the gap at a coefficient of 1.5, and the random walk that separates them, dashed. The two gaps’ Dickey–Fuller statistics are both far past the 5% point.

The gap at the wrong coefficient is the true gap plus a slow random walk, and at two hundred steps the random walk is small beside the fast returning part. Step to step, most of the gap’s movement is the returning part being pulled back, so the regression of each change on the previous level finds a strong pull and reports a statistic of −4.86 against the true gap’s −5.11. The unit root is there, in a slow drift under the fast oscillation, and a test that reads the pull between consecutive steps is looking at the wrong timescale to see it.

This is not a flaw of the imposed test in particular. It is the unit-root test’s well-known weakness when a series is a random walk plus a large stationary component, the same arithmetic that makes a series with a large negative moving-average root look stationary; supplying a slightly wrong coefficient manufactures exactly that series. Adding lagged differences to the regression, which is the standard repair, narrows the confirmed set without closing it: with four lags a coefficient of 1.5 is confirmed 49.4% of the time at two hundred observations.

And more data does not narrow it

The intuition that a longer sample would eventually expose the drift is right in the limit and wrong at every sample a study will have. At a thousand observations, with a gap that halves in three steps, a coefficient of 1.2 is confirmed 100.0% of the time, 1.5 91.5% and two 58.3% — each more often than at two hundred. With four lags, 1.5 is confirmed 72.0% of the time at a thousand.

The longer sample gives the test more of the fast returning part to see, and it sees it with more confidence, faster than the slow drift accumulates enough to be noticed. A unit root with a small increment beside a stationary component with a large one is detectable only when the sample is long enough for the walk’s variance to dominate, which for a coefficient 20% wrong on these numbers is far beyond a thousand observations. It is the other half of the difficulty the cliff that is a slope measured: a stationary series close to a unit root is often called non-stationary, and a non-stationary series with a small unit root inside a large stationary part is often called stationary. A unit-root test is a test of the long run run on a sample that is mostly short run, and both of its failure modes come from that.

So the imposed test is two different instruments at once. As a test of “is there a relation of roughly this form”, it is more powerful than the estimated one and holds its level. As a test of “is this coefficient the relation”, it is very nearly blind, and a longer sample makes it more confident rather than more discerning.

What the estimated test was buying

This puts the Engle–Granger test’s search charge in a different light. The estimated coefficient converges at rate 1/n1/n when the pair is tied — the regression that is not spurious measured it — so the estimated test is testing, approximately, the right gap, and the half unit it loses in critical value is the price of not having to be told the coefficient. A study that supplies the coefficient saves that price and accepts a different risk: that a rejection confirms its coefficient when the data would have supported a quite different one.

And the risk is not symmetric in what gets reported. A study whose supplied coefficient passes reports a confirmed theory. The same study, run with the coefficient estimated, would report the estimate — and at two hundred observations with a gap halving in three steps the estimate’s median error is 0.10, enough to tell a coefficient of 1.5 from the truth on almost every sample and not enough to tell 1.1. The imposed test’s extra power is bought partly with information the study discarded.

The half unit of critical value the Engle–Granger test pays is the continuous version of a charge met before, in a different form. A break that was looked for priced a structural break found by scanning every candidate date and reporting the most significant: the statistic is a maximum over the candidates, its critical value has to move to account for the scan, and a break supplied by theory in advance is tested against the ordinary one. Choosing a cointegrating coefficient by least squares is a scan over a continuum of candidate relations, and the residual it reports is the minimum over that continuum. Supplying the coefficient removes the scan in both cases.

The parallel holds on the second half as well. A break date supplied in advance and slightly wrong is still tested at that date, and a real break a few periods away is still detected there, because a break near the tested date produces a shift across it. The test confirms the neighbourhood of the date, not the date. The imposed cointegration test confirms the neighbourhood of the coefficient, not the coefficient, for the same reason: a test built to detect a feature is not built to locate it, and supplying the location only removes the charge for having searched.

Two questions a rejection is asked to answer

A study that tests the gap its theory implies is usually asking two things at once, and the numbers separate them cleanly.

The first is whether the pair returns to a relation at all — whether there is a mechanism that pulls the two series back together, of roughly the form the theory describes. For this question the imposed test is the better instrument. It holds its level, it is more powerful than the estimated test throughout the middle of the range, and the width of its confirmed set does not matter, because every coefficient in the set describes a returning relation of the theory’s form.

The second is whether the theory’s coefficient is the right one — whether purchasing-power parity holds one for one, or whether a spread is the arbitrage-free spread. For this question the imposed test is nearly worthless: a coefficient a fifth wrong passes as often as the right one, and a study that reports “the gap implied by the theory is stationary” as support for the theory’s coefficient has reported a fact about the first question as an answer to the second.

A study that runs the imposed test and the estimated test side by side, and reads agreement between them as reassurance, has learned less than it thinks. On these numbers agreement is what one should expect whether or not the coefficient is right, because the imposed test confirms a wide neighbourhood and the estimated test finds the relation that is there. The informative comparison is between the estimated coefficient and the supplied one, with the estimate’s own uncertainty.

What the walks in the gap look like to someone reading the series

The difficulty is not only the test’s. A reader shown the two gaps in the single-pair figure would not pick the wrong one out either: both oscillate around a level, both return from every excursion within a few steps, and the drift that separates them is slower than anything the eye attends to in a plot of two hundred points. Two walks and a finding showed the opposite illusion — two unrelated walks that look related — and this is its mirror: a relation that is wrong looks right, because what the eye and the test both read is the short-run return, and a wrong coefficient does not disturb that.

That is also why differencing the pair is no remedy here. The differences of the wrong gap and of the right one are almost the same series, since the random walk they differ by contributes only its small increments; everything that distinguishes the two gaps lives in their levels, over long spans, which is exactly the information differencing discards.

What a study that has a theoretical coefficient should do

Estimate the coefficient anyway, and report it beside the imposed test. A supplied coefficient that the data would estimate at 1.5 is not confirmed by a rejection of the unit root in the gap it implies; the estimate is the evidence about the coefficient, and the imposed test is evidence only that some relation of that general form returns. The error-correction model gives the estimate and its speed of adjustment in one regression.

Test the coefficient, not the gap. The hypothesis a theory makes is that β=β0\beta = \beta_0, and with a tied pair that is testable directly, with a standard error that shrinks at rate 1/n1/n. A study that tests the gap for a unit root is testing a consequence of the theory that also follows from many neighbouring theories.

Use the imposed test for what it is good at. When the question is whether a pair returns at all and the coefficient is not in doubt — an accounting identity, a unit conversion, a spread between two quotes of the same instrument — the imposed test’s extra power is honest and the confirmed set does not matter, because nothing near the coefficient is in contention.

Do not read a longer sample as settling the coefficient. On these numbers a thousand observations confirm a coefficient half again too large more often than two hundred do. Precision about the coefficient comes from estimating it.

What is measured here and what is not

Supplying the coefficient moves the 5% point from −3.38 to −2.89 at two hundred observations, and the half-life found with 80% power from 5.05 steps to 6.88, on the same pairs; at a hundred observations from 2.34 to 3.40, at four hundred from 10.11 to 14.24.

With a gap that halves in three steps, the imposed test confirms a coefficient of 1.2 on 99.1% of samples and of 1.5 on 81.9% at two hundred observations, and on 100.0% and 91.5% at a thousand, though each of those gaps contains a random walk.

Every rate is counted over simulated pairs — fifteen hundred a point for the power curves, eight hundred for the confirmed set at two hundred observations and four hundred at a thousand — generated by the error-correction model this field has used throughout, with the true coefficient one and unit-variance shocks. Both tests are read against critical values simulated on twenty thousand pairs of unrelated walks.

Not measured: the imposed test with its lag length chosen from the data, which is what practice does and which would sit between the two versions here. Not claimed either that real series have the separation of timescales these simulations have; a series whose returning part is slow, or whose non-stationary part is large, would expose a wrong coefficient more readily.

Still open: the sample widened rather than lengthened

Supplying the coefficient moved the boundary by a third. The second response the field named was to widen the sample instead of lengthening it: a relation too slow to see in one pair may be visible across many pairs that share it — twenty currencies against the same numeraire, forty regional price indices against a national one, a panel of spreads.

That trades a question about time for a question about homogeneity. A panel test pools the evidence of many pairs, and if they share a half-life, sixteen pairs should see much slower returns than one. What it rejects is a statement about the panel, though, not about any pair in it, and whether a rejection can be read as “they share a relation” when only some of them do is the price that trade has to be charged. Neither the gain nor the price has been measured here.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Augmented dickey fullerCointegrationCritical valueEngle–GrangerHalf-lifeRandom walkStatistical powerUnit root