Three series, and a count

A space is not a relation

The fitted plane approaches the true one at rate 1/n — 0.1438 at a hundred observations and 0.0075 at sixteen hundred. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and never moves.

Worth reading first: Three series and a count.

A system with one cointegrating relation has one long-run equation, and the only thing undetermined about it is its scale: which series goes on the left is a choice of units, and once made, the coefficients are quantities the data has something to say about.

With two relations that stops being true, and the way it stops is not a matter of degree. If y1y2y_1 - y_2 returns to a level and y2y3y_2 - y_3 returns to a level, then so does every combination of them: y1y3y_1 - y_3, and y12y2+y3y_1 - 2y_2 + y_3, and any other. The set of returning combinations is a plane through the origin, and every basis for that plane describes the system exactly as well as every other.

The software prints two vectors anyway. What those two vectors are estimates of is the question, and it has a measurable answer.

One of these converges and the other does not. Two measurements on the same fits, against the sample length, for a system with 2 genuine relations. The distance from the fitted plane to the true plane falls from 0.1438 at 100 observations to 0.0075 at 1600 — halving with each doubling, which is the 1/n rate this field's estimates converge at. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and is flat in between. The plane is an estimate; the relation inside it is not.
Fig. 1 Two measurements on the same fits, against the sample length. The distance from the fitted plane to the true plane falls from 0.1438 at a hundred observations to 0.0075 at sixteen hundred, halving with each doubling. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those lengths, and is flat in between.

One of those lines is an estimate converging on something. The other is a number with no limit to converge to.

The distance that falls, and the rate it falls at

The plane is estimable and it is estimable well. The distance drawn above is the Frobenius norm of the difference between the two projectors — the one onto the fitted plane and the one onto the true plane — which is a quantity defined on subspaces rather than on bases and is therefore indifferent to which vectors were printed.

It halves with each doubling of the sample: 0.1438, 0.0633, 0.0316, 0.0159, 0.0075 at a hundred, two hundred, four hundred, eight hundred and sixteen hundred observations. That is 1/n1/n, not 1/n1/\sqrt{n}, and it is the same superconsistency the two-series case shows in its slope. A cointegrating object is estimated at a rate no ordinary parameter reaches, and the count is only the first thing that benefits from it.

So the field’s central claim survives intact in the two-relation case, and survives in a strong form: the long-run structure is pinned down faster than anything else in the model. It is pinned down as a plane.

The angle that does not move

The second line is the angle between the leading fitted relation — the eigenvector belonging to the largest eigenvalue, which is the first vector any software prints — and the leading vector of the pair the system was generated from.

It reads 29.6° at a hundred observations and 29.0° at sixteen hundred. Sixteen times as much data moves it by half a degree, and what movement there is has no direction: the readings at the intermediate lengths are 28.9°, 29.0° and 29.0°, so the sequence goes down, up, and then nowhere at all.

That is what a quantity with nothing to converge to looks like. The generating pair y1y2y_1 - y_2 and y2y3y_2 - y_3 is one basis for the plane; the eigenvectors are another, chosen by a criterion — the size of a canonical correlation — that has nothing to do with how the system was written down. The angle between them is a property of two arbitrary choices, and no amount of data makes two arbitrary choices agree.

The failure here is not that the estimate is imprecise. The estimate is excellent and it is an estimate of something else. Reading the printed vector as “the long-run relation the system obeys” is a category error that no diagnostic reports, because the vector has a standard error, the standard error shrinks, and everything about the output looks like a parameter behaving well.

The estimated object is the plane, not the arrows in it. The fitted cointegrating plane, drawn in its own coordinates, with both normalisations' basis vectors in it. All four arrows lie in one plane and any two of them that are not parallel describe it completely. The data determines the disc; it does not determine which arrows are drawn on it, and every arrow here corresponds to a relation a reader could be shown as the system's long-run structure.
Fig. 2 The fitted plane drawn in its own coordinates, with two normalisations’ basis vectors in it. Every arrow lies in one plane and any two non-parallel ones describe it completely. The data determines the disc; it does not determine which arrows get drawn on it.

This is a shape of failure that recurs across statistics and that no routine diagnostic looks for: the symptom is absence rather than error. Nothing is reported wrongly. A quantity that should not have been reported at all is reported correctly, with a standard error that means something for a different question, and the only way to notice is to ask what the number would converge to if the sample grew — which is a question no output contains.

Two normalisations, one estimate, two stories

The abstraction becomes concrete the moment the numbers are printed. Take one fitted plane from one system of four hundred observations and write it out twice, solving for a different pair of series each time.

One plane, two sets of relations, both correct. The same fitted cointegrating plane from one system of 400 observations, written under two choices of which series to solve for. The two bases span the same space to 5e-16 — they are the same estimate — and the largest disagreement between corresponding entries is 2.111. Each reads as a pair of relations between named quantities, and which pair a reader is shown is a property of the software's default rather than of the system.
Fig. 3 The same fitted plane written under two choices of which series to solve for. The two bases span the same space to 5e-16 — they are one estimate — and the largest disagreement between corresponding entries is 2.111.

Solved for the first two series, the relations read y1=1.218y3y_1 = 1.218\,y_3 and y2=1.111y3y_2 = 1.111\,y_3: two quantities each tied to the third, which reads as a story in which y3y_3 is the quantity the others are read against and the other two follow it.

Solved for the second and third, they read y2=0.912y1y_2 = 0.912\,y_1 and y3=0.821y1y_3 = 0.821\,y_1: two quantities tied to the first, which reads as a story in which y1y_1 is the quantity the others are read against.

Both are exactly correct. They describe the same plane to fifteen decimal places, which is to say they are the same estimate written in two alphabets. And they support opposite sentences about which quantity the others are read against — a sentence a reader will take from whichever one the software printed.

The plane is genuinely close to the truth: its distance from the true plane is 0.1131 here, against the 2.111 gap between the two printed bases. The disagreement between the two descriptions is twenty times the error in the thing being described.

What a rank-two system looks like when it is drawn

Three series and two relations between them. Above, three series generated from Δy = Πy₋₁ + ε with Π of rank 2. Below, the combinations y₁ − y₂ and y₂ − y₃. They stay inside a band of 9.3 while the series themselves travel 17.5. The count of combinations that behave this way is the rank of Π, and it is what every method in the field sets out to estimate.
Fig. 4 Three series with two relations between them, and beneath them the two combinations that return. Both lower panels are bounded, and so is every combination of them — which is the fact the plane is a description of.

The picture is worth a moment because it makes the ambiguity concrete rather than algebraic. Adding the two lower panels together gives y1y3y_1 - y_3, which is bounded. Subtracting one from the other gives y12y2+y3y_1 - 2y_2 + y_3, also bounded. Nothing distinguishes the two panels that were drawn from any of the others that could have been: they are the two the generator happened to be written with.

A reader shown only the drawn pair would say the system’s long-run structure is “y1y_1 tracks y2y_2, and y2y_2 tracks y3y_3”. A reader shown y1y3y_1 - y_3 and y12y2+y3y_1 - 2y_2 + y_3 would say something else entirely, about the same system, with equal warrant.

Why this does not happen with one relation

It is worth being precise about the boundary, because the one-relation case is the one most readers have in mind and it is genuinely better behaved.

With r = 1 the returning set is a line through the origin. A basis for a line is a single vector, determined up to scale, and fixing the scale — setting one coefficient to one — pins it completely. Two analysts who normalise on different series get vectors that are multiples of each other, and every ratio of coefficients agrees. The choice the two-step procedure has to make is about which equation to estimate, not about which relation exists.

With r = 2 the returning set is a plane, a basis for it is two vectors, and there are infinitely many such bases that are not multiples of one another. Fixing scales does not help: the normalisation solves for two different series in the two cases, and the resulting vectors are not proportional to each other in any pairing.

So the transition from r = 1 to r = 2 is not “more relations, more of the same problem”. It is the point at which the object stops having a canonical description. The same thing happens at r = 2 in any k; three series is simply the smallest system in which it can happen at all, since r must be strictly between 0 and k.

The count is safe, and it is what the earlier essays measured

It is worth saying explicitly what this does not undermine, because “the relations are not identified” sounds like it should reach further than it does.

The count is a property of the plane — its dimension — and is therefore determinate. Every measurement in the two essays before this one is about the count and survives untouched: the procedure’s one-sided error budget, the cost of imposing the wrong rank, the eigenvalue gap the decision is read from. The spectrum itself is determinate too, since eigenvalues do not depend on which basis is printed.

The spectrum of a system with two relations. The 3 eigenvalues of the reduced-rank regression, averaged over 200 systems at n = 300, with each one's trace statistic and the 5% point it is read against. An eigenvalue is a squared canonical correlation between the changes and the levels, so 0.333 means a combination of levels explaining 33.3% of the variance of a combination of changes. The first two clear theirs and the third does not, and the count of the ones that do is the estimate.
Fig. 5 The spectrum of a system with two relations. Two eigenvalues away from zero and one against it — a statement about the plane’s dimension, and one that no choice of basis affects.

So does the forecast. A fitted system at rank two produces identical forecasts under both normalisations, exactly, because the forecast depends on the product of the adjustment matrix and the relations, and re-normalising the relations re-normalises the adjustments inversely. That product is the object that enters the model, and it is determinate.

What is lost is only the interpretation of the printed vectors — which is, unfortunately, the part of the output that gets quoted in a sentence.

What can be said, and what cannot

Three kinds of statement survive the ambiguity and are worth separating, because the difference between them is exactly the difference between a plane and a basis for it.

Statements about the plane. “Some combination of these three series returns to a level” is determinate. So is “the space of returning combinations is two-dimensional”, which is the count the earlier essays were about. So is any statement of the form “this particular combination returns”, for a combination named in advance: whether a given vector lies in the fitted plane is a question with an answer and a test.

Statements about a named relation.y1y_1 and y3y_3 stand in a one-to-one long-run relation” is determinate if it is a claim that a particular vector lies in the plane, and indeterminate if it is a claim that the estimation produced that vector. The two readings are easy to confuse and the output encourages the confusion.

Statements about the basis. “The first relation is between y1y_1 and y3y_3 and the second is between y2y_2 and y3y_3” is not a statement about the system at all. It is a statement about the normalisation, and the same data supports the opposite one.

A fourth kind is worth naming because it is the one most often wanted and is squarely on the wrong side of the line: statements about which series adjusts to which. The adjustment coefficients are attached to the relations, so re-normalising the relations re-normalises the adjustments, and “y2y_2 adjusts to the first relation and y3y_3 does not” is a sentence about the basis. The determinate version of that question exists — whether a particular series adjusts to nothing at all is a restriction on the adjustment matrix that survives any re-normalisation — and it is the question a system’s adjustment vector can genuinely answer.

A fifth kind is worth naming because it is the one a forecast depends on, and it is safe: statements about what the system will do. A fitted model’s predictions are invariant to the normalisation exactly, so every question of the form “where will these series be” has an answer the arithmetic supports — which is the question the previous essay priced. The determinacy runs out precisely where the interpretation begins.

The useful discipline follows directly: bring the relation to the data rather than reading it off. A theory that says a particular combination should return specifies a vector; testing whether that vector lies in the fitted plane is a well-posed question, it has a standard test, and it uses the part of the estimate that actually converges. That is the same move as imposing a long-run coefficient rather than searching for one, and it buys the same thing: a question the data can answer in place of one it cannot.

Why the estimator picks the basis it does

One question the figures raise and do not answer: if the basis is arbitrary, why does the procedure return a particular one, and is it a good one?

The eigenvectors come out ordered by their eigenvalues, which are squared canonical correlations between the changes in the series and their lagged levels. So the first printed relation is the combination of levels that explains the most variance in some combination of changes — the “strongest” relation in a sense that is precise and is about the sample rather than about the system. It is a perfectly reasonable choice of basis and it is a choice.

Two consequences follow. The ordering is itself estimated, so at a short sample the two printed relations can swap places between one sample and the next without anything about the system changing. And the criterion has no interpretation in the subject matter: nothing about an economy or a physical system makes the combination with the largest canonical correlation the one worth naming. The printed basis is the one the arithmetic reached for, which is why it carries no story and is routinely read as though it did.

What the count and the space have in common

The three essays in this part of the field now say the same thing about three different objects, and the pattern is worth naming because it decides which questions a system supports.

What is estimable is what does not depend on a convention. The count is the plane’s dimension, which no basis affects. The plane is the plane. The forecast depends on the product of the adjustments and the relations, which re-normalisation leaves alone. Each of those converges, and the first two converge at 1/n.

What is not estimable is everything a convention had to choose. Which pair of vectors is printed, which variables were solved for, which relation is called first — each is a decision made inside the arithmetic and reported as though it were a reading.

A reader can apply that test without any of the arithmetic: ask whether the quantity would change if the software had been asked a differently-phrased but equivalent question. If it would, the quantity is a property of the phrasing. It is the same test which series goes on the left applies at rank one, where the answer happens to be that only the scale changes — and the reason that essay’s finding looks reassuring is that a line has very few conventions to choose from.

What is claimed here and what is not

Six hundred systems at each length. Each reading on the convergence figure is a mean over six hundred independently generated systems, and the distance and the angle are computed on the same fits, so the two lines are paired rather than independent columns. What separates them is a property of each fit, not a difference between two experiments.

The 29° is not a meaningful number and is not meant to be. It is the angle between the leading eigenvector and one particular generating vector, and both of those are conventions. Its value would change if the system were written with a different pair of generating relations and it would change if the eigenvectors were ordered differently. What is claimed is that it does not move with the sample, and that claim is about the flatness rather than about the level.

The two normalisations are one realisation. The entries quoted — 1.218, 1.111, 0.912, 0.821 — come from a single system of four hundred observations and another draw would give different ones. What does not vary across draws is the structure: the two bases span the same estimated plane exactly, by construction, and they are not proportional to each other, also by construction. The realisation illustrates a fact about linear algebra rather than establishing one by simulation.

Identification is recoverable with restrictions, and none are used here. A structural analysis normally imposes enough restrictions on the relations to pin a basis down — a coefficient set to zero, a combination excluded from an equation — and with enough of them the printed vectors become estimates of something again. Nothing above argues against that; it argues that the restrictions are doing the identifying and that an analysis without them has none. The count of restrictions needed is r2r^2 minus the r scale normalisations, which is two for a rank-two system, and an analysis that supplies none has supplied two too few.

The forecast invariance is exact and is stated rather than measured. It follows from the model’s structure — the fitted system depends on the relations and the adjustments only through their product — and it would be wrong to present it as a simulation result. A reader who wants it checked can note that both normalisations of one fit give the same plane to 5e-16, and a model’s forecast is a continuous function of its coefficients.

And the plane’s convergence is measured against a plane that exists. Every system here is generated at a known rank from known relations, so “the true plane” is a real object to measure a distance to. A study has no such object, and the 1/n rate is therefore a statement about how fast the estimator concentrates rather than a diagnostic anybody can run.

Still open: what a restriction has to be worth

The last paragraph names the repair and does not price it. Two restrictions identify a rank-two system, they are supplied by the analyst rather than by the data, and a restriction that is wrong is a misspecification carried into every coefficient that follows.

That is a trade with two measurable sides and neither has been measured here. On one side, imposing a correct restriction should improve the individual relations from “not estimable” to “estimated at 1/n”, which is the largest possible improvement. On the other, imposing a wrong one takes an estimate of a plane that was converging and turns it into a confident description of a plane the data does not support — and the restriction is testable only if it over-identifies, so an analysis supplying exactly two restrictions has no way to find out.

The question with a shape is whether the usual practice of supplying exactly the number of restrictions needed is the worst available choice. Supplying more allows the surplus to be tested; supplying fewer leaves the basis undetermined and the plane honest. What each costs, on a system where the truth is known, is a measurement this field has not made.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formCointegrating rankCointegrationCommon trendConvergence rateEigenvalueThe error-correction modelEstimation errorMonte CarloNon-identifiabilityNormalisationReduced-rank regressionStructural modelSubset of parameters