Series that move together

The test with no table

The statistic that separates a real long-run relation from a spurious one is computed as a t and is not a t. At two hundred observations its 5% point is −3.38 where the t table says −1.65, and reading it against the table calls two unrelated random walks cointegrated 70.5% of the time.

Worth reading first: Two walks and a finding.

Two series are cointegrated when some combination of them is stationary. The obvious way to find out is to construct the combination and test it: fit the levels regression, take the residual, and ask whether it has a unit root.

That is the Engle–Granger procedure and it is exactly right. The statistic is a Dickey–Fuller t on the residual, computed the way any t is computed — a coefficient over its standard error — and there is a number waiting at the end of it that has to be compared with something.

The something is where the whole essay is.

Where the residual test's statistic actually falls, at n = 200Four thousand pairs of unrelated random walks, each regressed on the other and each residual tested for a unit root. The statistic is computed as a t and its distribution is not a t: five per cent of it falls below -3.38, where the ordinary one-sided 5% point of a t on 198 degrees of freedom is -1.65. Everything left of -1.65 — 70.2% of the whole distribution — is a pair of unrelated walks that a t table calls cointegrated.0100200-4-20the residual test's statistic, under two unrelated walkshow many of 4,000simulated 5%: -3.38t table 5%: -1.654,000 unrelated pairs, n = 200, statistic on the fitted residuala t table rejects 70.2% of them at a nominal 5%
Fig. 1 Four thousand pairs of unrelated random walks, each regressed on the other, each residual tested. This is the distribution of the statistic when the null is true. Five per cent of it lies below −3.38. The ordinary one-sided 5% point of a t on 198 degrees of freedom is −1.65, and everything to the left of that — 70.2% of the entire distribution — is a pair of unrelated walks that a t table calls cointegrated.

Why it is not a t

There are two reasons stacked on each other, and separating them is worth doing because the first is well known and the second is what makes this case worse.

A unit-root statistic is never a t. Under the null of a unit root the regressor in the auxiliary regression — the lagged level — is non-stationary, so the usual asymptotics do not apply. The limiting distribution is a functional of Brownian motion, it is shifted well to the left of a t, and this is why the Dickey–Fuller distribution exists as a separate object with its own table.

And this is not the Dickey–Fuller distribution either. The series being tested is not a series anyone handed over; it is a fitted residual, and the fitting chose β to minimise its sum of squares — which is to say, chose the combination that looks as stationary as it possibly can. That extra step biases the statistic further to the left again, by an amount that depends on how many regressors were in the first stage.

The second effect is the larger of the two on this design, and it is the one a reader is most likely to miss, because the first stage looks like preparation rather than like part of the test. It is not preparation. Choosing β to minimise the residual sum of squares is a maximisation over a whole family of candidate combinations, and the statistic that comes out is the best of that family rather than a draw from it.

So the correct critical value is not in a t table, and it is not in the Dickey–Fuller table either. It is in a third distribution that depends on the number of series in the first regression, on whether the regression had a constant, on whether it had a trend, and — very weakly — on the length of the series.

Simulating it

There is a table for this in the literature, and the honest way to get the number on this site is the same way that table was produced: generate pairs under the null, run the procedure, and look at the quantile.

The null is easy to generate correctly because it is a point inside the alternative. The same generator that makes a cointegrated pair, with the correction coefficient set to zero, makes two unrelated random walks. Nothing about the comparison can turn on the two cases having been coded differently.

At two hundred observations, four thousand pairs give a 5% point of −3.38. At eight hundred it is −3.31. The near-constancy across a factor of four in length is the sign that the distribution is converging to something — the asymptotic value for one regressor with a constant is around −3.34 — and it is also a warning: an approximation that is nearly the same at every n is one whose critical value looks like a fixed number and is not one.

Where the residual test's statistic actually falls, at n = 800. Four thousand pairs of unrelated random walks, each regressed on the other and each residual tested for a unit root. The statistic is computed as a t and its distribution is not a t: five per cent of it falls below -3.31, where the ordinary one-sided 5% point of a t on 798 degrees of freedom is -1.65. Everything left of -1.65 — 70.3% of the whole distribution — is a pair of unrelated walks that a t table calls cointegrated.
Fig. 2 The same distribution at eight hundred observations. Its 5% point has moved from −3.38 to −3.31 and its shape is essentially unchanged — while the t table’s critical value, which depends only on degrees of freedom, has not moved at all. The share of unrelated pairs a t table would call cointegrated is 70.3%, against 70.2% at two hundred. The over-rejection does not go away with more data.

That second sentence is the one to keep. Most approximations improve with n. This one does not, because the discrepancy is not a small-sample effect — the t distribution is the wrong limit, not a poor approximation to the right one.

What “no table” means in practice

The phrase in the title is not rhetorical, and it is worth spelling out what a reader in this situation actually faces, because it is unusual.

For a t test, a chi-square test or a normal approximation, the reference distribution is a fixed mathematical object. It does not depend on the design, on the number of variables, or on anything except one or two degrees-of-freedom parameters. A table settles it once.

Here the reference distribution depends on:

  • how many series were in the first regression — each extra regressor gives the fitting step more freedom to make the residual look stationary, and shifts the distribution further left;
  • whether the first regression had a constant, a trend, or neither — three different distributions;
  • how many lagged differences the auxiliary regression includes;
  • and, weakly, on the length of the series.

Published tables cover the common combinations, and any combination outside them has to be simulated. Which is not a hardship — four thousand pairs take under a second — but it does change what the analysis consists of. The critical value becomes part of the work rather than part of the reference material, and it has to be reported alongside the statistic, because a statistic of −3.1 is a rejection under one specification and not under another.

Where the residual test's statistic actually falls, at n = 50. Four thousand pairs of unrelated random walks, each regressed on the other and each residual tested for a unit root. The statistic is computed as a t and its distribution is not a t: five per cent of it falls below -3.49, where the ordinary one-sided 5% point of a t on 48 degrees of freedom is -1.68. Everything left of -1.68 — 68.4% of the whole distribution — is a pair of unrelated walks that a t table calls cointegrated.
Fig. 3 And at fifty observations, where the 5% point is −3.49 rather than −3.38. The drift is the right way round — shorter series need a more extreme critical value — and it is small, about three per cent across a factor of four in length, which is why a single asymptotic value is a usable approximation and a t table is not.

What each rule actually does

A critical value is half of a test. The other half is whether the rule finds anything.

What each rule rejects, and what it should. The same statistic read against two critical values, on unrelated walks and against a real error-correction mechanism at α = -0.2. Read against a t table the test calls unrelated walks cointegrated 70.5% of the time, against the 5% it claims. Read against the simulated value it fires on 4.8% of unrelated pairs and still finds 99.7% of real ones. The rule that holds its size loses almost nothing.
Fig. 4 The same statistic read against two critical values, on unrelated walks and against a real mechanism at α = −0.2. Against a t table: 70.5% of unrelated pairs called cointegrated, at a nominal 5%. Against the simulated value: 4.8% of unrelated pairs, and 99.7% of real ones still found.

The bottom two bars are the reason the correction is not a trade-off. Tightening the critical value from −1.65 to −3.38 removes almost none of the power: the test still finds 99.7% of pairs with a real correction mechanism at α = −0.2. What it removes is the 65 percentage points of false positives.

That happens because the two distributions are far apart. Under the null the statistic averages around −2; against a real mechanism at α = −0.2 with two hundred observations it averages around −7. A critical value at −3.38 sits in the gap between them and separates them almost perfectly.

The uncomfortable corollary is that the t table’s critical value at −1.65 sits inside the null distribution, which is why it rejects most of it. It is not a slightly wrong threshold. It is a threshold in the wrong place entirely.

Worth noting too that the t table finds 100% of the real mechanisms, which is the shape every bad rule has. A test that rejects everything has perfect power, and power quoted without size is not information. That is the same objection what a p-value does not say makes to quoting a p-value without the sample size, and which rate is controlled makes to quoting a correction without saying what it corrects: a rate is only a claim when it is one of a pair.

The check the simulation needed

A simulated critical value has a trap in it that is easy to fall into and produces a perfect-looking result.

If the critical value is computed as the 5% quantile of a set of statistics and then applied to the same set, the rejection rate is 5.0% by construction, exactly, always, whatever the code does. That is not a measurement; it is the definition of a quantile read back.

So the critical value here is computed on one set of seeds and applied on a different set. The rejection rate that comes out is 4.8% rather than exactly 5%, and the small discrepancy is the point: it is Monte Carlo error in a genuine measurement rather than an arithmetic identity dressed as one.

This is the same failure this site has recorded before as a refusal that throws on every path proves nothing, and the same one p-values must be flat is built to catch in another guise. A check that cannot fail is not a check, and a quantile applied to its own sample cannot fail.

The same shape as the site’s other refusals

This is the fourth time this site has met a statistic whose reference distribution is not the one its arithmetic suggests, and the four together make a pattern rather than a list.

Testing at 0.05 every time the data is looked at. The null is true in every one of these trials and the test is correct every time it is run. Looking once rejects 4.9% of the time, as it should; looking ten times rejects 19.2% of the time. Nothing changed except permission to look.
Fig. 5 Sequential testing: five interim looks at the nominal level reject a true null 14.0% of the time. The statistic at each look is a perfectly ordinary z; what is not ordinary is the distribution of the minimum over five looks, which is what the decision rule actually uses.

In when the looking happens the statistic is a z and the reference distribution is the distribution of its running minimum. In twenty analyses of nothing the statistic is a t and the reference distribution is the distribution of the smallest p-value over the analyses that were available. Here the statistic is a t and the reference distribution is what a fitted residual’s unit-root statistic does.

The common structure is that something was chosen, and the reference distribution has to account for the choosing. A stopping rule chooses when; a garden of forking paths chooses which analysis; a levels regression chooses β. In every case the arithmetic of the statistic is untouched and the table it should be read against is not.

That gives a general test to apply to any statistic before looking anything up: was any part of this selected to make the statistic large? If so, the table is wrong, and by how much is a simulation.

The rate this test controls

Everything above is about the size of the test, and the site has a field about what a rate is a rate of.

20,000 p-values from a true null, n = 30. Flat, as it must be: under the null a p-value is uniform on (0,1). The Kolmogorov–Smirnov distance from uniform is 0.0228 (p = 0.01). That flatness is the check that catches an error a single rejection rate would miss.
Fig. 6 The property being claimed, in its familiar form: under the null a p-value is uniform, so a test at 5% rejects 5% of the time. Everything in this essay is that same claim, made for a statistic whose null distribution had to be built rather than looked up.

The 4.8% is a familywise-free single-test rate: one pair, one test, one decision. It is not corrected for anything, and in the setting cointegration testing actually gets used — screening many pairs of series for relationships — every difficulty spending the error rate documents applies on top of it. Testing forty pairs at 4.8% each gives roughly an 86% chance of at least one false positive, and the correction for that is the same correction as anywhere else.

The reason to say so explicitly is that the 70.5% figure makes the uncorrected test look like the whole problem, and it is not. Repair the critical value and there is still a multiplicity problem underneath, and it is the ordinary one.

What the tighter threshold cost, as an exchange rate

The correction is described above as not a trade-off, and the arithmetic behind that word is worth setting out, because “not a trade-off” is a strong claim and this is one of the few places it holds.

Moving the critical value from −1.65 to −3.38 takes the false-positive rate from 70.5% to 4.8% and the power from 100% to 99.7%. So 65.7 points of size were bought with 0.3 points of power: an exchange rate of 219 to 1. Nothing else in this site’s tests comes close — a Bonferroni correction, a wider interval, a boundary spent across looks, each buys size at something between one and five points of power per point.

The reason it is nearly free here is that the two distributions barely overlap. The null averages about −2 and a real mechanism at α = −0.2 on two hundred observations averages about −7, and −3.38 sits in a gap where almost no mass of either lies. A threshold placed in a gap costs nothing to move within the gap — which is also the warning attached to the 99.7%, since a test whose power is that high against the alternative it was measured at has not been asked a difficult question.

The difficult question is a weak correction on a short series, and this field measures where that starts. A correction of speed α has half-life log(0.5)/log(1 + α), and a series needs many complete corrections in it before any statistic can see them. Requiring ten half-lives inside n observations gives log(1 + α) ≤ 10·log(0.5)/n, which at n = 200 means α ≤ −0.034. Two hundred observations can support a relation that undoes about three and a half per cent of a gap per step, and nothing slower.

Screening, and how much simulation it needs

The multiplicity point above — forty pairs at 4.8% each giving about an 86% chance of one false positive — has a consequence for the simulation rather than only for the interpretation, and it is the more awkward half.

Holding a familywise rate of 5% across forty pairs needs a per-pair level of 1 − 0.95^(1/40) = 0.128%. That level is not a correction applied to a p-value; it is a different quantile of the same simulated null distribution, and it sits far out in a tail. Four thousand pairs put only five draws below it, so the critical value estimated there is a statistic based on five numbers, and its own uncertainty is larger than the difference between the critical values at fifty and eight hundred observations that the essay treats as a real drift.

So a screening application needs a simulation roughly forty times the size of a single test’s — not because the arithmetic is harder, but because the quantile being read has moved from the 5% point to the 0.13% point, and the number of draws needed to pin a quantile scales with how far into the tail it is. A hundred and sixty thousand pairs is still under a minute, so this is a note about doing it rather than an obstacle; what it is not is something a published table can supply, since tables stop at the conventional levels.

That is the sharpest form of what “no table” means here. The critical value has to be simulated, and how deep into the tail it has to be simulated is decided by how many pairs are being screened — which is a property of the study rather than of the test.

Reading the picture rather than the number

There is a diagnostic available that costs nothing and is worth more than the test on short series: plot the residual.

A pair pulled back at 10% of the gap per step. Above, the two series. Below, the difference between them. The gap is pulled back towards zero by 10% of itself each step, so it stays inside a band of 19.6 while the series themselves travel much further. Nothing here is stationary except the difference. The faint line below is the gap for two free walks from the same seed, drawn for comparison.
Fig. 7 The lower panel is what the test is testing. A residual that returns to a level, repeatedly, over a span long enough to contain several returns, is evidence a statistic can only summarise. A residual that wanders off and does not come back is the other case, and the eye is good at telling them apart when there are enough complete excursions to see.

What the eye cannot do is calibrate, which is why the test exists. What the test cannot do is notice that the residual returned to its level exactly twice in the whole sample, which is the situation in which a rejection at 5% is technically correct and practically worthless.

The rule of thumb that follows is a statement about the half-life. A correction with half-life h needs a series many multiples of h long before the data contains enough complete corrections to say anything. At α = −0.05 the half-life is 13.5 steps, so two hundred observations contain about fifteen half-lives — enough. At α = −0.01 it is 69 steps, and two hundred observations contain three, which is not.

What it does not settle

Two limits worth stating.

The test is about the pair, not about the ratio. Rejecting the null says a stationary combination exists. It does not say that the β the first regression happened to find is the right one, only that it is close enough that what is left looks stationary — and at short series it can be close enough for the test and not close enough to use, which is what the cost of differencing a pair measures.

Failing to reject says nothing. As everywhere, what a p-value does not say applies: a non-rejection at two hundred observations against a weak correction is exactly what a weak correction produces. The half-life of a shock at α = −0.05 is thirteen steps, and thirteen steps is long enough that two hundred observations contain few complete corrections to learn from.

Both limits point the same way, and it is the direction this whole field points. The interesting quantity is not whether the pair is cointegrated. It is α, how fast the relation reasserts itself — which is a number with units and a half-life, rather than a verdict. The test is a preliminary to estimating it, and the next essay is about what estimating it is worth.

The number to carry

One figure summarises the essay and it is not the critical value.

70.5%. That is the share of unrelated random walk pairs that the ordinary t table’s 5% point calls cointegrated at two hundred observations. It does not fall with more data. It is not a small-sample artefact, a rounding issue or a matter of a slightly conservative threshold. A rule advertised as rejecting one time in twenty rejects fourteen times in twenty.

The reason it is worth carrying rather than the −3.38 is that −3.38 is specific to this design and this specification, and would be different with two regressors or a trend. The 70.5% is about what happens when a statistic is read against a distribution it does not have, and the answer — that the rule can end up rejecting most of its own null — generalises to every case in the list above.

It also settles the question the field opened with. Two walks and a finding measured a 76.7% false-positive rate for the naive levels regression and concluded that levels regressions between trending series should not be trusted. The repair, applied without care, has a 70.5% false-positive rate of its own. A correction that is read against the wrong reference distribution is not much of a correction, and the only thing separating the two numbers is which table was consulted.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CointegrationEngle–GrangerError rateFalse positiveMonte CarloNull hypothesisp-valueSpurious regressionUnit root