Three series, and a count

The rank is a decision

The sequential procedure's 5% bounds one of its two errors. Over-counting reads between 4.2% and 7.2% at every sample length from fifty observations to three hundred; under-counting reads 69.5% at fifty and 0.0% at three hundred, and nothing in the procedure bounds it.

Worth reading first: Three series and a count.

Counting what is still wandering established what the trace statistic is read against: not one distribution but three, keyed on how many series are left wandering under the null being tested. With the right row of that table the test holds its size.

What follows from a table of critical values is a procedure, and the procedure is not a test. It tests whether there are no relations at all; if that rejects, it tests whether there is at most one; if that rejects, at most two; and it reports the first count it cannot reject. Three tests, run in sequence, each at 5%, with the later ones conditional on the earlier having rejected, and a single integer at the end.

Nobody states an error rate for the integer. Counted directly, over a three-series system with two genuine relations, it has one — and it is entirely on one side.

The bounded error and the unbounded one. How the sequential trace procedure's answer is distributed, against the sample length, for a three-series system with 2 genuine relations. Over-counting — claiming a stationary combination that is a random walk — reads 4.9%, 7.2%, 5.7%, 6.2%, 5.9%, 4.2% across the six lengths, never far from the 5% of a single test. Under-counting reads 69.5%, 40.2%, 14.0%, 0.5%, 0.0%, 0.0%. The procedure is described as a 5% rule and the 5% applies to one of those columns.
Fig. 1 How the procedure’s answer is distributed against the sample length, for a system with two genuine relations. Over-counting reads between 4.2% and 7.2% at every length. Under-counting reads 69.5% at fifty observations and 0.0% at three hundred.

The 5% is a bound on over-counting, it holds at every length, and it is a bound on the error that almost never happens. The error that does happen has no bound at all, and at the sample lengths a great many systems are studied at, it is most of what the procedure does.

The system being counted

The object under all of this is three series generated from a stated mechanism, so that the count the procedure is trying to recover is a number that was put in rather than one read off afterwards.

Three series and two relations between them. Above, three series generated from Δy = Πy₋₁ + ε with Π of rank 2. Below, the combinations y1 −y2 and y2 −y3. They stay inside a band of 9.3 while the series themselves travel 17.5. The count of combinations that behave this way is the rank of Π, and it is what every method in the field sets out to estimate.
Fig. 2 Three series generated with two relations between them: y1y2y_1 - y_2 and y2y3y_2 - y_3 both return to a level, and one common trend is left for all three to ride. Every series wanders; the two combinations do not.

Two relations among three series means one thing that genuinely wanders. That is the same statement read from either end, and it is what makes the rank the right object: the count of relations and the count of independent trends add to the number of series, so recovering one is recovering the other.

Why one column is flat and the other is not

The asymmetry is not an accident of these settings and it follows directly from the shape of the procedure.

Each test in the sequence has a null of the form “at most q relations” and rejects in favour of more. Reporting a count larger than the truth therefore requires rejecting a true null, which is what a 5% level is a promise about — and the promise survives the sequence, because the sequence stops at the first non-rejection. To return three where the truth is two, the procedure must reject the true null “at most two”, which happens at 5%. To return four it would have to reject that one too, and it never gets the chance. The level bounds over-counting at the level of a single test, however many tests are in the sequence.

Reporting a count smaller than the truth is a different event: it requires failing to reject a false null, which is a power question, and nothing in the construction of a 5% test says anything about power. So the two errors are not two halves of one budget. One is controlled by design and the other is whatever the sample happens to supply.

This is the mirror image of the structure that a familywise correction is built around: there the concern is that a sequence of tests inflates the chance of a false claim, and here the sequence does not inflate it, because it stops. What is unusual is that the procedure’s reputation rests on the guarantee that is easy to keep.

One consequence worth drawing out: the guarantee is about the reported count and not about the study. Two studies of the same system at the same length will disagree about the integer more often than 5% of the time — at seventy-five observations, two independent samples return different counts about half the time — because disagreement is driven by the unbounded column and not the bounded one. A reader who takes 5% as the rate at which this method produces conflicting structural claims has taken a bound on one error for a bound on both.

What the mistake actually is at a short sample

A percentage correct hides the shape of the error, and here the shape matters, because the procedure can be wrong by more than one.

What the procedure returns when the answer is 2. The whole distribution of the returned count, at two sample lengths, for a system with 2 genuine relations. At 50 observations the right answer comes back 25.6% of the time and the answer "none at all" — wrong by two — comes back 9.9%. At 300 it comes back 95.8% of the time and the error is entirely on the other side. A single share-correct figure would report 25.6% and 95.8% and say nothing about which way.
Fig. 3 The whole distribution of the returned count at two sample lengths, for a system with two genuine relations. At fifty observations the right answer comes back 25.6% of the time, “one relation” 59.6% and “none at all” 9.9%. At three hundred it comes back 95.8% of the time and every error is on the other side.

At fifty observations the modal answer is one, not two. A study reporting it has not made a marginal call: it has claimed a system with a single long-run relation and two independent trends where there are two relations and one trend, and the count of common trends is the same number read from the other end. The structural story that follows — which quantities wander freely and which are tied — is different in kind rather than in degree.

And 9.9% of the time the answer is none at all, which is the verdict “these three series are unrelated random walks” about a system in which every pair of them is tied.

There is a second reason the sequence does not compound, and it is worth separating from the first because it is the one that would fail if the procedure were written differently. The tests are nested: rejecting “at most one” is a precondition for even running “at most two”. A procedure that ran all three tests independently and took, say, the largest count any of them supported would have no such protection, and its over-counting rate would be the union of three 5% events rather than one. The stopping rule is what buys the guarantee, and it buys it by throwing away the later tests in exactly the cases where they could do damage.

That also means the guarantee is fragile in a specific way. Anything that causes the sequence to be restarted, re-specified, or run at several lag lengths with the best-looking result kept, is a search the count has not been charged for, and the one-sided bound is the first thing such a search would break.

The decision is read off a spectrum that is not yet there

The reason is visible one level down, in the object the count is actually read from.

The procedure solves an eigenvalue problem whose k eigenvalues are squared canonical correlations between the changes in the series and their lagged levels. The r belonging to genuine relations stay away from zero as the sample grows; the other k − r go to zero at rate 1/n. The count is the position of the gap.

The gap the count is read from, at two sample lengths. The mean squared canonical correlation at each position of the spectrum, for a system with 2 genuine relations. Each eigenvalue is the share of the variance of a combination of changes that a combination of levels explains. At 300 observations the 2 real ones sit at 0.334 and 0.174 and the remainder at 0.0103, so the gap is plain. At 50 the real ones read 0.415 and 0.228 and the remainder 0.062 — the eigenvalues that converge to zero at rate 1/n have not had the sample to do it, and the decision is being read off a spectrum with no gap in it.
Fig. 4 The mean eigenvalue at each position of the spectrum, at two sample lengths, for a system with two genuine relations. At three hundred observations the two real ones read 0.334 and 0.174 and the third reads 0.0103. At fifty they read 0.415 and 0.228 and the third reads 0.0618 — six times larger, because an eigenvalue converging to zero at rate 1/n has had a sixth of the sample to do it in.

At three hundred observations there is a gap: a factor of seventeen between the smallest real eigenvalue and the largest spurious one. At fifty there is a factor of under four, and a factor of under four between two noisy quantities is not a gap anybody can read an integer off.

That also explains why the first two eigenvalues are larger at the short sample — 0.415 against 0.334 — rather than smaller. Every eigenvalue is inflated by the same finite-sample effect, the real ones included; what changes with the sample is the separation between them, and separation is what a count needs. The same misreading is available here as in a spectrum whose top eigenvalue holds while the rest fall: a large leading eigenvalue is not evidence, because the null produces large leading eigenvalues too.

At a length where the procedure works, the gap is the most legible thing in the field.

The spectrum of a system with two relations. The 3 eigenvalues of the reduced-rank regression, averaged over 200 systems at n = 300, with each one's trace statistic and the 5% point it is read against. An eigenvalue is a squared canonical correlation between the changes and the levels, so 0.333 means a combination of levels explaining 33.3% of the variance of a combination of changes. The first two clear theirs and the third does not, and the count of the ones that do is the estimate.
Fig. 5 The spectrum of one system of three hundred observations with two relations in it. The two eigenvalues belonging to genuine relations sit well away from zero and the third sits against it, and the count is the position of the step between them.

What the short-sample reading says is that the step is a feature of the long sample rather than of the system. The system has two relations at every length; the picture that makes the count readable is something the sample builds, and it builds it at rate 1/n.

Where the crossover sits, and why it is not reassuring

The under-counting column falls fast: 69.5% at fifty observations, 40.2% at seventy-five, 14.0% at a hundred, 0.5% at a hundred and fifty and 0.0% beyond. Above about a hundred and fifty observations the procedure is doing what it is described as doing, and the whole of this essay is about the region below.

Two things stop that being a comfortable reading.

The threshold is a property of these adjustment speeds. The system here is pulled back at thirty per cent of any disagreement per period in one direction and ten in the other — a half-life of about two periods on the fast relation. A system with slower adjustment needs a longer sample to reach the same accuracy, by exactly the arithmetic that governs how slow a return a sample can see in the two-series case. A hundred and fifty is not a general number; it is this system’s number, and it is a number for a fast system.

And the procedure reports nothing about which side it is on. The output is an integer. A study at fifty observations and a study at three hundred print the same kind of object, and nothing in either says that one of them is in the region where the modal answer is wrong.

The counts a short sample returns, drawn

What the count comes back as, on systems of rank 2. 500 systems of 3 series at n = 75, each put through the sequential trace procedure at the 5% level with a critical value simulated for every null it tests. The true rank is 2 and it is returned 52.0% of the time. The errors are not symmetric: 40.0% under-count, which throws away a relation that exists, and 8.0% over-count, which claims a stationary combination that is a random walk.
Fig. 6 The same procedure at seventy-five observations, where it returns the right count 52.6% of the time. The distribution’s mass sits at one and two, which is a system whose structural description is being decided by a coin.

Seventy-five observations is not a pathological length — it is nineteen years of quarterly data — and at it the procedure returns “one relation” 40.2% of the time and “two” 52.6%. A study reporting either has followed the method correctly.

The consequence for what gets written down is worth stating plainly. A count of one and a count of two are not two estimates of a continuous quantity that happen to bracket the truth. They are two different structural claims: one says two of the three series share a wandering component that the third does not, and the other says all three are tied to a single wandering component. The intermediate answer that a confidence interval would provide does not exist, because the object is an integer.

What a defensible use looks like

Report the count with the power that produced it. The useful companion to “two relations” is “at this length and these speeds the procedure returns the right count four times in five” — a quantity computable by simulation from the fitted system, before any interpretation is attached to the integer.

Treat a small count at a short sample as uninformative rather than as a finding. The asymmetry says exactly which direction to be sceptical in. A count of zero from fifty observations is the answer the procedure gives to a tenth of systems that have two relations, so it is not evidence that there are none.

Say which integers the data could not rule out. The procedure returns one count and discards the evidence for the others, which is a strange thing to do with a quantity that takes four possible values. The three trace statistics and their three critical values are all computed on the way to the answer, and printing them gives a reader the set of counts that were not rejected rather than only the smallest — which is closer to what the data supports and costs nothing to report.

And do not read the spectrum as a continuous summary. A count is the position of a gap, and where there is no gap there is no count. Printing the eigenvalues beside the integer is the cheapest available diagnostic, and the ratio between the last retained and the first discarded is the thing to look at — seventeen at three hundred observations here and under four at fifty. That ratio is the same kind of instrument as the width of a band read beside its coverage: a second number that says whether the first one means anything.

What is claimed here and what is not

One system, one set of speeds. Every measurement above is on a three-series system with relations y1y2y_1 - y_2 and y2y3y_2 - y_3, adjustment coefficients of −0.3 and 0.1 on each, and unit innovation variances. The qualitative claim — that over-counting is bounded and under-counting is not — follows from the shape of the procedure and holds for any system. The numbers do not: a system with weaker adjustment is worse at every length and a stronger one is better.

The critical values are simulated at the sample length used. Each null in the sequence gets its own value, generated under the right number of common trends at that n, rather than taken from an asymptotic table. Using asymptotic values at fifty observations would add a second error on top of the one measured here, and separating them is why they are simulated.

The “half the time” figure above is arithmetic, not a separate measurement. Two independent samples at seventy-five observations each return two with probability 0.526 and one with probability 0.402, so they agree with probability about 0.45 if the draws are independent — which they are, being separate simulations. It is quoted as a consequence of the distribution already drawn rather than as a new reading, and nothing rests on it beyond that distribution.

The trace statistic is one of two. The maximum-eigenvalue statistic tests rank = q against rank = q + 1 rather than pooling everything below the cut, and a procedure built on it has the same one-sided structure for the same reason — its nulls are also nulls about “at most” — but different power. Nothing here compares the two, and the comparison is worth making.

Fifty and seventy-five observations are short, and they are not unusual. Annual data over half a century is fifty observations; quarterly data over two decades is eighty. The lengths at which this procedure is most often applied in practice are precisely the lengths at which the measurement above says it under-counts, and that is the reason the region below a hundred and fifty is the subject rather than a footnote about small samples.

And “correct” means the count the system was generated at. That is available here and is not available anywhere else. What the figures measure is a procedure’s ability to recover a number that was put in, which is a lower bound on the difficulty of the real problem, where the number was never put in by anyone.

The same asymmetry, seen from the other side

One more reading makes the structure hard to mistake. The under-counting column is a power curve, and a power curve is always a statement about a particular alternative — here, about this system’s adjustment speeds. The over-counting column is not a power curve and is not about any alternative at all: it is a statement about the null, which is why it does not move when the sample does.

So the two columns are not comparable quantities that happen to differ. They are a fixed promise and a measured consequence, and the reason the procedure is described by the first is that the first is the one that can be promised in advance. The second requires knowing the thing the study is being run to find out, which is the same circularity an allocation rule runs into when it is fed a pilot’s guess: the quantity needed to say how well the method will work is the quantity the method exists to estimate.

Still open: what the mistake costs

An error rate is a count of mistakes and not a measure of them, and the two errors here are not only unequal in frequency. They are unequal in consequence, and the direction of that inequality decides whether the procedure’s one-sided guarantee is protecting the right side.

Imposing one relation too few and imposing one too many are different restrictions on how the system is allowed to move, and each is wrong in its own way: the first differences away a combination that genuinely returns, and the second treats a random walk as though it returned. There is no reason for them to cost the same, and what each of them actually costs is the next thing this field has to measure — because if under-counting is the expensive mistake as well as the unbounded one, the procedure’s guarantee is pointed at the wrong error twice over.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Cointegrating rankCointegrationCommon trendConvergence rateCritical valueEigenvalueError rateMonte CarloMultiple comparisonsRandom walkReduced-rank regressionSequential testingStatistical powerTrace statistic