Field

Three series, and a count

A pair is either tied or it is not, so its whole inference is one test with one answer. Three series can carry none, one or two relations at once, and the quantity being estimated stops being a slope and becomes an integer — read off the gap in a spectrum, against a critical value that depends on how many things are left wandering and on nothing else.
Three series and one relation between them. Above, three series generated from Δy = Πy₋₁ + ε with Π of rank 1. Below, the combination y1 −y2. It stays inside a band of 9.5 while the series themselves travel 28.4. The count of combinations that behave this way is the rank of Π, and it is what every method in the field sets out to estimate.

Three series and a count

A pair of series is either tied together or it is not, so its whole inference is one test with one answer. Three can carry none, one or two relations at once — and the thing being estimated stops being a slope and becomes an integer, read off the gap in a spectrum whose top eigenvalue holds at 0.25 while the rest fall like 1/n.

The same data, one regression per choice of left-hand side. The two-step procedure has to put one series on the left, and with 3 series there are 3 ways to do it. Each returns a relation and a residual test; the 5% point is -3.71, simulated. Here they do not agree: 2 of 3 reject, and the relations they report are written with a 1 in the position of whichever series was on the left, so they can be compared. Nothing in a printed output records which regression was run.

Which series goes on the left

The two-step procedure has to pick a series to regress the others on, and nothing in its output records which. With a pair that choice never changes the verdict. With three series and one relation between them, the three choices disagree about whether the system is cointegrated at all 98.0% of the time.

The trace statistic under the null, and the 5% point it needs. 600 systems of 3 unrelated random walks, each put through the reduced-rank regression, with the statistic for "rank ≤ 0" collected. The 5% point is 31.91. There is no standard table to look that up in: the distribution depends on the number of common trends under the null and is not a chi-square, so the value is simulated on one set of seeds and applied on another — exactly the position the pair's residual test was in one field ago.

Counting what is still wandering

The statistic that turns a spectrum into an integer has one name and three distributions. Its 5% point is 8.12, 18.64 or 31.74 depending only on how many series are left wandering under the null being tested — and read against the wrong one of those three, it calls unrelated random walks cointegrated most of the time.

Every equation's adjustment speed, and the one number they make together. Each series gets its own equation, each is regressed on the same lagged disequilibrium, and what comes back is the whole vector α. Averaged over 400 systems at n = 300: α₁ = -0.154 against -0.15 generated, α₂ = 0.104 against 0.1 generated. The gap closes at the combination of them rather than at any one entry — 25% of any disagreement per step, a half-life of 2.41 steps, where the single equation that fits only the first series reports 4.27.

Which series does the moving

“y adjusts towards x” and “x adjusts towards y” are different mechanisms with identical long-run relations, and a single-equation model cannot tell them apart because it only writes one equation. Writing all of them recovers a vector — and a gap that closes at 25% a step where one equation alone reports 15%.

The bounded error and the unbounded one. How the sequential trace procedure's answer is distributed, against the sample length, for a three-series system with 2 genuine relations. Over-counting — claiming a stationary combination that is a random walk — reads 4.9%, 7.2%, 5.7%, 6.2%, 5.9%, 4.2% across the six lengths, never far from the 5% of a single test. Under-counting reads 69.5%, 40.2%, 14.0%, 0.5%, 0.0%, 0.0%. The procedure is described as a 5% rule and the 5% applies to one of those columns.

The rank is a decision

The sequential procedure's 5% bounds one of its two errors. Over-counting reads between 4.2% and 7.2% at every sample length from fifty observations to three hundred; under-counting reads 69.5% at fifty and 0.0% at three hundred, and nothing in the procedure bounds it.

What each wrong count costs, 4 steps ahead. Squared forecast error 4 steps ahead at each imposed rank, relative to the correctly specified fit, at 200 observations. With 1 genuine relations, imposing 0 costs 13.3% and imposing 2 costs 4.8%. With 2 genuine relations, imposing 1 costs 15.6% and imposing 3 costs 2.5%. Under-counting is the more expensive mistake in both systems, and it is the one the procedure's level does not bound.

Which mistake about the rank costs

On a system with two relations, imposing none costs 29.2% of squared forecast error and imposing three costs 2.5%. The expensive mistake is under-counting, which is the error the procedure's 5% does not bound — so the guarantee protects the cheap side.

One of these converges and the other does not. Two measurements on the same fits, against the sample length, for a system with 2 genuine relations. The distance from the fitted plane to the true plane falls from 0.1438 at 100 observations to 0.0075 at 1600 — halving with each doubling, which is the 1/n rate this field's estimates converge at. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and is flat in between. The plane is an estimate; the relation inside it is not.

A space is not a relation

The fitted plane approaches the true one at rate 1/n — 0.1438 at a hundred observations and 0.0075 at sixteen hundred. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and never moves.

All essays