The rank is a decision
Worth reading first: Three series and a count.
Counting what is still wandering established what the trace statistic is read against: not one distribution but three, keyed on how many series are left wandering under the null being tested. With the right row of that table the test holds its size.
What follows from a table of critical values is a procedure, and the procedure is not a test. It tests whether there are no relations at all; if that rejects, it tests whether there is at most one; if that rejects, at most two; and it reports the first count it cannot reject. Three tests, run in sequence, each at 5%, with the later ones conditional on the earlier having rejected, and a single integer at the end.
Nobody states an error rate for the integer. Counted directly, over a three-series system with two genuine relations, it has one — and it is entirely on one side.
The 5% is a bound on over-counting, it holds at every length, and it is a bound on the error that almost never happens. The error that does happen has no bound at all, and at the sample lengths a great many systems are studied at, it is most of what the procedure does.
The system being counted
The object under all of this is three series generated from a stated mechanism, so that the count the procedure is trying to recover is a number that was put in rather than one read off afterwards.
Two relations among three series means one thing that genuinely wanders. That is the same statement read from either end, and it is what makes the rank the right object: the count of relations and the count of independent trends add to the number of series, so recovering one is recovering the other.
Why one column is flat and the other is not
The asymmetry is not an accident of these settings and it follows directly from the shape of the procedure.
Each test in the sequence has a null of the form “at most q relations” and rejects in favour of more. Reporting a count larger than the truth therefore requires rejecting a true null, which is what a 5% level is a promise about — and the promise survives the sequence, because the sequence stops at the first non-rejection. To return three where the truth is two, the procedure must reject the true null “at most two”, which happens at 5%. To return four it would have to reject that one too, and it never gets the chance. The level bounds over-counting at the level of a single test, however many tests are in the sequence.
Reporting a count smaller than the truth is a different event: it requires failing to reject a false null, which is a power question, and nothing in the construction of a 5% test says anything about power. So the two errors are not two halves of one budget. One is controlled by design and the other is whatever the sample happens to supply.
This is the mirror image of the structure that a familywise correction is built around: there the concern is that a sequence of tests inflates the chance of a false claim, and here the sequence does not inflate it, because it stops. What is unusual is that the procedure’s reputation rests on the guarantee that is easy to keep.
One consequence worth drawing out: the guarantee is about the reported count and not about the study. Two studies of the same system at the same length will disagree about the integer more often than 5% of the time — at seventy-five observations, two independent samples return different counts about half the time — because disagreement is driven by the unbounded column and not the bounded one. A reader who takes 5% as the rate at which this method produces conflicting structural claims has taken a bound on one error for a bound on both.
What the mistake actually is at a short sample
A percentage correct hides the shape of the error, and here the shape matters, because the procedure can be wrong by more than one.
At fifty observations the modal answer is one, not two. A study reporting it has not made a marginal call: it has claimed a system with a single long-run relation and two independent trends where there are two relations and one trend, and the count of common trends is the same number read from the other end. The structural story that follows — which quantities wander freely and which are tied — is different in kind rather than in degree.
And 9.9% of the time the answer is none at all, which is the verdict “these three series are unrelated random walks” about a system in which every pair of them is tied.
There is a second reason the sequence does not compound, and it is worth separating from the first because it is the one that would fail if the procedure were written differently. The tests are nested: rejecting “at most one” is a precondition for even running “at most two”. A procedure that ran all three tests independently and took, say, the largest count any of them supported would have no such protection, and its over-counting rate would be the union of three 5% events rather than one. The stopping rule is what buys the guarantee, and it buys it by throwing away the later tests in exactly the cases where they could do damage.
That also means the guarantee is fragile in a specific way. Anything that causes the sequence to be restarted, re-specified, or run at several lag lengths with the best-looking result kept, is a search the count has not been charged for, and the one-sided bound is the first thing such a search would break.
The decision is read off a spectrum that is not yet there
The reason is visible one level down, in the object the count is actually read from.
The procedure solves an eigenvalue problem whose k eigenvalues are squared canonical correlations between the changes in the series and their lagged levels. The r belonging to genuine relations stay away from zero as the sample grows; the other k − r go to zero at rate 1/n. The count is the position of the gap.
At three hundred observations there is a gap: a factor of seventeen between the smallest real eigenvalue and the largest spurious one. At fifty there is a factor of under four, and a factor of under four between two noisy quantities is not a gap anybody can read an integer off.
That also explains why the first two eigenvalues are larger at the short sample — 0.415 against 0.334 — rather than smaller. Every eigenvalue is inflated by the same finite-sample effect, the real ones included; what changes with the sample is the separation between them, and separation is what a count needs. The same misreading is available here as in a spectrum whose top eigenvalue holds while the rest fall: a large leading eigenvalue is not evidence, because the null produces large leading eigenvalues too.
At a length where the procedure works, the gap is the most legible thing in the field.
What the short-sample reading says is that the step is a feature of the long sample rather than of the system. The system has two relations at every length; the picture that makes the count readable is something the sample builds, and it builds it at rate 1/n.
Where the crossover sits, and why it is not reassuring
The under-counting column falls fast: 69.5% at fifty observations, 40.2% at seventy-five, 14.0% at a hundred, 0.5% at a hundred and fifty and 0.0% beyond. Above about a hundred and fifty observations the procedure is doing what it is described as doing, and the whole of this essay is about the region below.
Two things stop that being a comfortable reading.
The threshold is a property of these adjustment speeds. The system here is pulled back at thirty per cent of any disagreement per period in one direction and ten in the other — a half-life of about two periods on the fast relation. A system with slower adjustment needs a longer sample to reach the same accuracy, by exactly the arithmetic that governs how slow a return a sample can see in the two-series case. A hundred and fifty is not a general number; it is this system’s number, and it is a number for a fast system.
And the procedure reports nothing about which side it is on. The output is an integer. A study at fifty observations and a study at three hundred print the same kind of object, and nothing in either says that one of them is in the region where the modal answer is wrong.
The counts a short sample returns, drawn
Seventy-five observations is not a pathological length — it is nineteen years of quarterly data — and at it the procedure returns “one relation” 40.2% of the time and “two” 52.6%. A study reporting either has followed the method correctly.
The consequence for what gets written down is worth stating plainly. A count of one and a count of two are not two estimates of a continuous quantity that happen to bracket the truth. They are two different structural claims: one says two of the three series share a wandering component that the third does not, and the other says all three are tied to a single wandering component. The intermediate answer that a confidence interval would provide does not exist, because the object is an integer.
What a defensible use looks like
Report the count with the power that produced it. The useful companion to “two relations” is “at this length and these speeds the procedure returns the right count four times in five” — a quantity computable by simulation from the fitted system, before any interpretation is attached to the integer.
Treat a small count at a short sample as uninformative rather than as a finding. The asymmetry says exactly which direction to be sceptical in. A count of zero from fifty observations is the answer the procedure gives to a tenth of systems that have two relations, so it is not evidence that there are none.
Say which integers the data could not rule out. The procedure returns one count and discards the evidence for the others, which is a strange thing to do with a quantity that takes four possible values. The three trace statistics and their three critical values are all computed on the way to the answer, and printing them gives a reader the set of counts that were not rejected rather than only the smallest — which is closer to what the data supports and costs nothing to report.
And do not read the spectrum as a continuous summary. A count is the position of a gap, and where there is no gap there is no count. Printing the eigenvalues beside the integer is the cheapest available diagnostic, and the ratio between the last retained and the first discarded is the thing to look at — seventeen at three hundred observations here and under four at fifty. That ratio is the same kind of instrument as the width of a band read beside its coverage: a second number that says whether the first one means anything.
What is claimed here and what is not
One system, one set of speeds. Every measurement above is on a three-series system with relations and , adjustment coefficients of −0.3 and 0.1 on each, and unit innovation variances. The qualitative claim — that over-counting is bounded and under-counting is not — follows from the shape of the procedure and holds for any system. The numbers do not: a system with weaker adjustment is worse at every length and a stronger one is better.
The critical values are simulated at the sample length used. Each null in the sequence gets its own value, generated under the right number of common trends at that n, rather than taken from an asymptotic table. Using asymptotic values at fifty observations would add a second error on top of the one measured here, and separating them is why they are simulated.
The “half the time” figure above is arithmetic, not a separate measurement. Two independent samples at seventy-five observations each return two with probability 0.526 and one with probability 0.402, so they agree with probability about 0.45 if the draws are independent — which they are, being separate simulations. It is quoted as a consequence of the distribution already drawn rather than as a new reading, and nothing rests on it beyond that distribution.
The trace statistic is one of two. The maximum-eigenvalue statistic tests rank = q against rank = q + 1 rather than pooling everything below the cut, and a procedure built on it has the same one-sided structure for the same reason — its nulls are also nulls about “at most” — but different power. Nothing here compares the two, and the comparison is worth making.
Fifty and seventy-five observations are short, and they are not unusual. Annual data over half a century is fifty observations; quarterly data over two decades is eighty. The lengths at which this procedure is most often applied in practice are precisely the lengths at which the measurement above says it under-counts, and that is the reason the region below a hundred and fifty is the subject rather than a footnote about small samples.
And “correct” means the count the system was generated at. That is available here and is not available anywhere else. What the figures measure is a procedure’s ability to recover a number that was put in, which is a lower bound on the difficulty of the real problem, where the number was never put in by anyone.
The same asymmetry, seen from the other side
One more reading makes the structure hard to mistake. The under-counting column is a power curve, and a power curve is always a statement about a particular alternative — here, about this system’s adjustment speeds. The over-counting column is not a power curve and is not about any alternative at all: it is a statement about the null, which is why it does not move when the sample does.
So the two columns are not comparable quantities that happen to differ. They are a fixed promise and a measured consequence, and the reason the procedure is described by the first is that the first is the one that can be promised in advance. The second requires knowing the thing the study is being run to find out, which is the same circularity an allocation rule runs into when it is fed a pilot’s guess: the quantity needed to say how well the method will work is the quantity the method exists to estimate.
Still open: what the mistake costs
An error rate is a count of mistakes and not a measure of them, and the two errors here are not only unequal in frequency. They are unequal in consequence, and the direction of that inequality decides whether the procedure’s one-sided guarantee is protecting the right side.
Imposing one relation too few and imposing one too many are different restrictions on how the system is allowed to move, and each is wrong in its own way: the first differences away a combination that genuinely returns, and the second treats a random walk as though it returned. There is no reason for them to cost the same, and what each of them actually costs is the next thing this field has to measure — because if under-counting is the expensive mistake as well as the unbounded one, the procedure’s guarantee is pointed at the wrong error twice over.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A space is not a relation — both name cointegrating rank, cointegration, common trend, convergence rate, eigenvalue, monte carlo, reduced-rank regression
- A horizon chosen after looking — both name critical value, monte carlo, multiple comparisons, statistical power
- A simulation that stops when it looks settled — both name error rate, monte carlo, sequential testing, statistical power
- Estimating how many nulls are true — both name error rate, monte carlo, multiple comparisons, statistical power
- False discoveries that arrive together — both name error rate, monte carlo, multiple comparisons, statistical power
- The cliff that is a slope — both name critical value, monte carlo, random walk, statistical power
Named objects
A flat tag is an object no other essay names yet.
Cointegrating rankCointegrationCommon trendConvergence rateCritical valueEigenvalueError rateMonte CarloMultiple comparisonsRandom walkReduced-rank regressionSequential testingStatistical powerTrace statistic