Three series, and a count

Three series and a count

A pair of series is either tied together or it is not, so its whole inference is one test with one answer. Three can carry none, one or two relations at once — and the thing being estimated stops being a slope and becomes an integer, read off the gap in a spectrum whose top eigenvalue holds at 0.25 while the rest fall like 1/n.

Worth reading first: The observations that repeat each other · Two walks and a finding.

The cointegration field on this site stops at two series, and stopping there is not a simplification. It is a different question with a smaller answer set.

With two series there is a long-run relation or there is not. The inference is one test, the test has one answer, and the quantity underneath it is a slope — one number, with a standard error, converging at a rate the field spends an essay measuring. Everything an analyst does with a pair is a variation on estimating that number and deciding whether it is there.

With three, the same question has three answers. There may be no relation at all, in which case all three series wander independently. There may be one, tying two of them together and leaving the third free. There may be two, which leaves exactly one thing wandering and everything else pinned to it. The quantity being estimated is a count, and a count is not something a t statistic reports.

Three series and one relation between them. Above, three series generated from Δy = Πy₋₁ + ε with Π of rank 1. Below, the combination y1 −y2. It stays inside a band of 9.5 while the series themselves travel 28.4. The count of combinations that behave this way is the rank of Π, and it is what every method in the field sets out to estimate.
Fig. 1 Three series generated from a mechanism with one long-run relation, and the combination that relation picks out drawn underneath. The series themselves travel a long way; the combination does not leave a narrow band. Nothing in the upper panel says how many such combinations there are.

The object is one matrix

Every claim in this field is a claim about a single matrix, and writing the system down is most of the work.

Δyₜ = Π yₜ₋₁ + εₜ

y is now a vector of three series rather than a number. Δy is the vector of this period’s changes. Π is a 3 × 3 matrix saying how last period’s levels push this period’s changes, and the whole of the field is the question of what Π looks like.

If the three series are unrelated random walks, Π is the zero matrix. Nothing about where a series was yesterday changes what it does today; that is exactly what a random walk is, said in matrix form.

If they are tied together, Π is not zero, and it is also not full rank. A full-rank Π would mean every combination of the levels pushes back on the changes, which would make every series stationary — and these are not stationary, they wander. So Π has a rank strictly between 0 and 3, and it factors:

Π = α β′

β holds the combinations of the levels that are stationary, one per column. α holds how hard each series is pulled back along each of those combinations. If β has one column, there is one long-run relation. If it has two, there are two.

The rank of Π is the number of long-run relations, and the case that matters — rank strictly between 0 and k — is precisely the case a pair does not have. For two series the only options are 0 and 1, and “1” is what the two-step procedure already reports. Three is the smallest number of series for which counting is a question.

What the three cases look like, and how little that helps

The three panels below are the same generator at rank 0, 1 and 2. The upper halves are close to indistinguishable, which is the honest starting point for the field.

Three series with nothing tying them together. Above, three series generated from Δy = Πy₋₁ + ε with Π of rank 0. Below, the combinations y₁ − y₂ and y₂ − y₃. Nothing is stationary: every combination drawn is itself a random walk, ranging over 52.3 in 300 steps with no level to return to.
Fig. 2 Rank zero: three random walks with nothing whatever tying them together. The differences drawn underneath are themselves random walks, because the difference of two independent walks is a walk.
Three series and two relations between them. Above, three series generated from Δy = Πy₋₁ + ε with Π of rank 2. Below, the combinations y1 −y2 and y2 −y3. They stay inside a band of 9.3 while the series themselves travel 17.5. The count of combinations that behave this way is the rank of Π, and it is what every method in the field sets out to estimate.
Fig. 3 Rank two: two independent stationary combinations, which leaves one common trend that all three series ride. Both combinations stay in a band; the series themselves travel with the trend.

Put the two upper panels beside the rank-one panel at the top of this essay and there is nothing to choose between them. Three lines that wander and roughly track each other is what all three systems look like, because in all three cases every individual series is non-stationary. The difference is entirely in what is left when the right combinations are subtracted, and which combinations those are is not known in advance.

That last clause is the difficulty, and it is worth being precise about. With a pair, there is only one combination to consider up to scaling: y₁ − βy₂. With three series there is a two-dimensional space of combinations to search, and the answer is a subspace — the set of all combinations that come out stationary. Its dimension is the count.

Three series with nothing tying them together. Above, three series generated from Δy = Πy₋₁ + ε with Π of rank 0. Below, the combinations y₁ − y₂ and y₂ − y₃. Nothing is stationary: every combination drawn is itself a random walk, ranging over 52.3 in 300 steps with no level to return to.
Fig. 4 The same three unrelated walks with two candidate combinations drawn instead of the system’s own — there are none. Both drift away and neither returns to a level, which is what “rank zero” means when it is drawn rather than asserted.

Counting by looking at a spectrum

The trick that turns this into arithmetic is to stop testing and start decomposing.

Regress the changes Δy on the lagged levels y₋₁, having first taken out a constant from both. That is three regressions, one per series, and what comes out is not directly interesting. What is interesting is the canonical correlation structure between the two blocks: how much of the variation in the changes can be explained by some combination of the levels, then how much of what is left can be explained by another combination, and so on.

That question has an exact answer for each of the three combinations in turn, and the answers are the eigenvalues of a matrix built from three cross-product matrices. Each eigenvalue is a squared canonical correlation, which is what makes it readable: an eigenvalue of 0.25 means there is a combination of the levels that accounts for a quarter of the variance of some combination of the changes. Nothing else in this field has a scale attached like that.

The spectrum of a system with one relation. The 3 eigenvalues of the reduced-rank regression, averaged over 200 systems at n = 300, with each one's trace statistic and the 5% point it is read against. An eigenvalue is a squared canonical correlation between the changes and the levels, so 0.252 means a combination of levels explaining 25.2% of the variance of a combination of changes. The first clears its critical value and the rest do not, and the count of the ones that do is the estimate.
Fig. 5 The three eigenvalues of a rank-one system, averaged over two hundred simulated systems. The first is an order of magnitude larger than the second, and the second is barely larger than the third.

At rank one the average eigenvalues are 0.2519, 0.0276 and 0.0064. At rank zero they are 0.0446, 0.0212 and 0.0050 — small and unremarkable, none of them standing out. At rank two they are 0.3333, 0.1769 and 0.0101, with two large and one small.

The spectrum of a system with no relations. The 3 eigenvalues of the reduced-rank regression, averaged over 200 systems at n = 300, with each one's trace statistic and the 5% point it is read against. An eigenvalue is a squared canonical correlation between the changes and the levels, so 0.045 means a combination of levels explaining 4.5% of the variance of a combination of changes. None of them clears its critical value, which is the right answer for a system with nothing in it.
Fig. 6 The same picture for three unrelated walks. There is no gap: the largest eigenvalue is 0.0446, which is not zero, because a sample correlation between two wandering series is never zero. It is simply not large.
The spectrum of a system with two relations. The 3 eigenvalues of the reduced-rank regression, averaged over 200 systems at n = 300, with each one's trace statistic and the 5% point it is read against. An eigenvalue is a squared canonical correlation between the changes and the levels, so 0.333 means a combination of levels explaining 33.3% of the variance of a combination of changes. The first two clear theirs and the third does not, and the count of the ones that do is the estimate.
Fig. 7 Rank two, where two eigenvalues stand clear of the third. The count is read as “how many are up here”, and the answer is two.

The count is a count of how many eigenvalues are large. That sentence is doing a great deal of work and it needs a reason to be true, because nothing so far explains why the small ones should be small enough to tell apart from the large ones.

Why the gap opens rather than closes

Here is the property that makes rank estimable at all, and it is the sort of claim this site would rather count than cite.

Take the same rank-one system at five lengths and average the eigenvalues over a hundred and twenty systems at each. If every eigenvalue shrank as the series got longer, no threshold could ever separate them, and the count would not be recoverable from data of any length.

The gap the count is read from, as the series get longer. Each eigenvalue of the reduced-rank regression, averaged over 120 systems at five lengths, for a system of rank 1. The top one belongs to real relations and holds station at about 0.244; the rest belong to common trends and fall from 0.1012 to 0.0066. The count is how many stay up.
Fig. 8 Each eigenvalue against the length of the series, on a log scale. The top one is flat. The other two fall in a straight line, which on this scale is what a 1/n rate looks like.

The eigenvalue belonging to the real relation goes 0.3003, 0.2724, 0.2569, 0.2488, 0.2445 as n goes 80, 150, 300, 600, 1200. It is drifting down slightly and settling; it is not going anywhere near zero.

The second, which belongs to a common trend, goes 0.1012, 0.0548, 0.0267, 0.0131, 0.0066. Fifteen times the data, fifteen times smaller. That is the 1/n the theory gives, and it is the whole reason a threshold exists: the ratio λ₁/λ₂ goes from 3.0 at n = 80 to 36.8 at n = 1200.

The gap opens roughly in proportion to the length of the series. So a procedure that puts a cut somewhere in the middle of the spectrum becomes more certain of its answer as more data arrives, which is the behaviour one wants and is not automatic. It is worth saying how unusual it is in this neighbourhood. Two walks and a finding is about a t statistic on a regression between unrelated series that gets larger with more data, so the rejection rate rises towards one rather than settling at 5%; more data makes that analysis more confidently wrong. And more data is not monotone is about an interval whose coverage goes down when n goes up. Neither of those is a pathology of a badly written procedure. They are what happens when a quantity’s sampling behaviour is not the one the analysis assumes.

Here the behaviour is the good one, and the reason is the same reason the pair’s slope converges at rate 1/n rather than 1/√n: a genuine long-run relation is an enormously strong signal, because the series it constrains travel arbitrarily far while the combination does not travel at all. The eigenvalue is picking up that constraint, and it does not weaken. What weakens is everything else.

The rate, checked rather than described

“Fifteen times the data, fifteen times smaller” is the right description and it can be made an exact one, which is worth doing because the whole procedure rests on the rate being 1/n and not something close to it.

Multiply the second eigenvalue by the length it was measured at: 8.10, 8.22, 8.01, 7.86 and 7.92 at n = 80, 150, 300, 600 and 1200. A fifteen-fold change in the sample and the product moves by two per cent either side of 8.0, with no trend in it.

So the nuisance eigenvalue is not approximately 1/n. It is 8.0/n over the whole range measured, and the third eigenvalue is 1.9/n by the same arithmetic at n = 300. Two constants, and the count is read between them and the one that does not move.

The top eigenvalue’s own approach is the same rate seen from the other side. Fitting a constant plus c/n to 0.2488 and 0.2445 at the two longest lengths gives a limit of 0.240 approached from above with c ≈ 5.2, and that fit predicts 0.2574, 0.2746 and 0.3047 at the three shorter lengths against 0.2569, 0.2724 and 0.3003 measured. The real relation’s eigenvalue converges to a population number at rate 1/n; the spurious ones converge to zero at rate 1/n. Same rate, different destinations, and the difference between the destinations is the entire estimand.

The gap in closed form, and where the threshold has to sit

Those two constants give the separation without any further simulation. The ratio is

λ1λ20.2408.0/n=n33,\frac{\lambda_1}{\lambda_2} \approx \frac{0.240}{8.0/n} = \frac{n}{33},

which returns 2.4 at n = 80, 9.1 at 300 and 36 at 1200 against the 3.0, 9.6 and 36.8 measured. The whole of the field’s headline movement is one division.

It also says where a cut can live. A procedure separating the two eigenvalues by a fixed number needs 8.0/n<c<0.2408.0/n < c < 0.240, so the window is empty below n = 33 and is (0.10, 0.24) at eighty observations, widening to (0.007, 0.24) at twelve hundred. The upper end never moves and the lower end falls like 1/n, which is why more data helps here and why it helps asymmetrically: it buys room against under-counting and none at all against over-counting.

That asymmetry is worth carrying into the next section, which is about the two errors not being interchangeable. They are not symmetric in the arithmetic either. The distance from a threshold at 0.15 down to λ₂ grows without limit as the sample grows; the distance up to λ₁ is fixed at 0.09 for ever.

What a count being the estimate actually changes

Three things follow, and each is a reason this field is not the cointegration field with an extra series.

There is no standard error. An estimate of 1.7 relations is not a thing. The output of the procedure is an integer, its error is a probability of returning the wrong integer, and the two kinds of mistake are not symmetric or interchangeable. Under-counting throws away a relation that exists and leaves a model that differences away information it had. Over-counting claims a combination is stationary when it is a random walk, which is the spurious-regression failure re-entering through a door marked “model selection”.

That first mistake is worth dwelling on, because the obvious safe move makes it. A reader who has absorbed the spurious-regression result knows that regressions between wandering series are untrustworthy and reaches for the standard repair: difference everything, and work with the changes. That repair does work on its own terms — the differenced regression holds its size where the levels regression rejects three times in four — and it is exactly wrong here, because differencing removes the levels and the relation is a statement about the levels. An analyst who differences a cointegrated system has not made their analysis safe. They have made it blind, and it will report nothing while looking like it worked.

The count and the common trends are the same fact twice. A system of three with two relations has one thing wandering freely; a system with one relation has two. r relations and k − r common trends are the same number read from opposite ends, and that equivalence is what makes rank the right object rather than a technical device. “How many independent sources of permanent movement are there?” and “how many long-run relations tie the series together?” are one question.

And the procedure has to be sequential. Testing “is the rank at least 1?” and “is the rank at least 2?” are different hypotheses with different null distributions, so the count comes from a chain of tests rather than from one, and each link needs its own critical value. That is the subject of the next essay but one, and the critical values turn out to be in no table this site can point at: 8.12, 18.64 and 31.74 at the 5% level, depending only on how many things are still wandering under the null being tested. That is the same situation the pair’s residual test is in — its 5% point is −3.38 where a t table says −1.65 — one dimension worse, because there is now a different distribution for each link in the chain rather than one distribution for one test.

The relation is recovered, not assumed

One more thing the decomposition gives away, and it is the reason the machinery is worth its complexity.

The eigenvector belonging to the largest eigenvalue, mapped back into the original coordinates, is the estimated cointegrating relation. It is not proposed by the analyst and then tested. It is the combination the data says is most stationary, and on a rank-one system built from y₁ − y₂ it comes back at that, normalised so the first entry is one.

So the procedure answers two questions with one computation: how many relations there are, and what they are. The single-equation approach answers the second by being told the first, and being told the first is precisely what a system of three does not permit — which is why an essay on choosing a left-hand side is the next thing this field has to do.

What the count comes back as, on systems of rank 1500 systems of 3 series at n = 300, each put through the sequential trace procedure at the 5% level with a critical value simulated for every null it tests. The true rank is 1 and it is returned 96.2% of the time. The errors are not symmetric: 0.0% under-count, which throws away a relation that exists, and 3.8% over-count, which claims a stationary combination that is a random walk.rank 00.0%rank 196.2%the true rankrank 23.6%rank 30.2%share of systems500 systems of rank 1 at n = 300, 5% level96.2% correct, 0.0% under, 3.8% over
Fig. 9 Five hundred rank-one systems put through the whole sequential procedure at each of five lengths. Drag it and watch the under-counting vanish while the over-counting stays exactly where it is: the second is the test’s own 5%, and no amount of data removes it.

At n = 300 the procedure returns the true rank 96.2% of the time, and at n = 80 it returns it 77.2% of the time, with 18.4% under-counts — the series are too short for the relation to be visible. By n = 150 the under-counting is 0.2% and it never comes back.

What does not move is the over-counting: 4.4% at n = 80 and 3.8% at n = 300. That is not an error that more data fixes, because it is not an error in the usual sense. It is the 5% level, doing what a 5% level does. A procedure that stops when a test fails to reject will over-shoot about one time in twenty however long the series are, and the only way to change that number is to change the level.

What this field takes on, and what it does not

The claim staked here is systems of three or more series where the number of long-run relations is the object being estimated. That covers the rank, the spectrum it is read from, the critical values that turn it into a decision, and the adjustment vector — which is what makes “which series does the moving” a question with an answer, and is the second half of this field.

What is not claimed, and is named here so a later phase knows it was left deliberately: forecasting as a craft, model selection over lag lengths, deterministic trends inside the cointegrating space, and any system large enough that the arithmetic stops being three-by-three. Each of those is a real subject and none of them is measured here, so none of them is asserted here either.

The check, and the refusal that makes it mean something

The claim this essay rests on is that the spectrum separates: the eigenvalues belonging to relations stay up as n grows, the ones belonging to trends fall, and the gap opens roughly in proportion to n. All three halves are asserted in this field’s library, at two lengths four times apart, with the growth bracketed on both sides rather than merely required to be positive — because “the gap opens” is satisfied by any growth at all, and the claim being made is about the rate.

An assertion that has never rejected anything proves nothing, so the field’s machinery is fed input it must refuse. Three unrelated walks must not produce a relation: the rank procedure returns zero 95.0% of the time on them, and the 5.0% it does not is again the level rather than a defect. And a trace statistic read against the wrong row of the simulated table — the row for one common trend, applied to a three-series system — calls unrelated walks cointegrated most of the time, which is the mistake this field is most likely to produce and the reason its critical values are indexed by the thing they depend on rather than by the statistic’s name.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Cointegrating rankCointegrationCommon trendDifferencingEigenvalueThe error-correction modelRandom walkReduced-rank regressionStationarityUnit root