Three series, and a count

A restriction the data can check

A rank-two system of three series is identified by exactly two restrictions on its relations, and any two that describe the estimated plane fit it equally well — their forecasts agree with the unrestricted fit's to the last digit, so the data can neither confirm nor refuse them. Supply one relation in full instead, one restriction more than identification needs, and the surplus becomes a test: a relation a tenth wrong is rejected 37.8% of the time at two hundred observations and 98.5% at eight hundred, the error found half the time halving as the sample doubles. A right relation buys the forecast almost nothing; a fifth wrong costs 13% of it at two hundred observations and 25% at eight hundred.

Worth reading first: Three series and a count.

A space is not a relation estimated a system of three series tied together by two long-run relations and found that what the data determines is the plane those two relations span, not the relations themselves. The fitted plane closed on the true one at rate one over the sample’s length — a distance of 0.1438 at a hundred observations and 0.0075 at sixteen hundred — while the angle between the leading fitted relation and the leading true one sat near thirty degrees at every length. A basis for a plane is a choice, and two choices of basis printed as two different sets of economic relations.

That essay ended on the repair it named and did not price. To say which two relations hold, an analyst supplies restrictions: two of them, for a rank-two system, is the number that pins down a unique basis. Supplied wrongly, a restriction is a misspecification carried into every coefficient after it, and supplied in exactly the number needed it cannot be tested. The question it left was whether supplying exactly the number needed is the worst available choice — whether more restrictions, which can be checked, or fewer, which leave the plane honest, do better. Both halves can be counted.

Three ways to identify one plane

The system is the earlier essay’s. Three series, two relations, y1−y2y_1 - y_2 and y2−y3y_2 - y_3 stationary, the third combination wandering with one common trend that all three share; the first series adjusts to the first gap, the second to both, the third to the second. Johansen’s reduced-rank regression estimates the plane.

An analysis can then say one of three things about the relations. It can decline to identify them, report the plane and the rank, and forecast from the error-correction model the plane implies. It can identify exactly, with the two restrictions a rank-two system needs — the first relation has no weight on y3y_3, the second none on y1y_1, each normalised on its own leading series — and print the two relations the restrictions select. Or it can over-identify, supplying more restrictions than identification needs: here, the first relation supplied in full as y1−c y2y_1 - c\,y_2, with theory’s value c=1c = 1, and only the second estimated.

Exact identification cannot be wrong in the data

The first two choices are the same model. Which series goes on the left met the same fact in the two-step procedure, where the choice of which series to regress on the others decides what relation is printed without changing whether the series are related; here it is the choice of restrictions, and the reduced-rank regression makes the point exact. Any two exactly identifying restrictions pick a basis of the estimated plane, and the error-correction regression that turns the relations into a forecast sees only the plane: the gaps a different basis produces are linear combinations of the same two gaps, so the regression’s fitted values do not change.

That is exact, and it holds on every draw. Fit the system to two hundred observations, normalise the estimated plane on series one and two, on two and three, and on one and three, and forecast eight steps ahead from each: the three forecasts agree with the unrestricted fit’s to the eighth decimal place on every one of twenty systems tried. The likelihood is identical too. So the data cannot distinguish an exactly identifying restriction that matches the true relations from one that does not — “the first relation has no weight on y2y_2” fits as well as “the first relation has no weight on y3y_3” — and an analysis that supplies exactly two restrictions has chosen which two relations to print without any possibility of the data objecting.

Which mistake about the rank costs priced the count’s errors in forecast error. This choice has no forecast price at all, because it changes nothing the forecast uses; its whole price is in the reading. The relations printed are the economics a paper argues from, and exact identification makes them a convention dressed as an estimate.

One relation supplied in full

Supplying the first relation as (1,−c,0)(1, -c, 0) restricts more than the plane’s basis: it says the plane contains that particular vector. For a rank-two plane in three dimensions that is one restriction more than identification needs, and a restriction the data can refuse. The test compares the likelihood with the vector supplied and the second relation estimated against the unrestricted rank-two likelihood; the statistic is the sample length times the difference in log determinants, referred to a chi-square on one degree of freedom.

How often a test of a supplied cointegrating relation rejects it, by how wrong it is and how long the sample is. 800 systems a point. At 100, 200, 400, 800 observations the right relation is rejected 9.4%, 5.4%, 5.4%, 6.4% of the time; one with c = 0.9, 19.9%, 37.8%, 78.5%, 98.5%; c = 1.1, 16.8%, 37.9%, 76.8%, 97.5%.
Fig. 1 How often the chi-square test at 5% rejects a supplied first relation y1−c y2y_1 - c\,y_2, against cc (the truth is 1), at a hundred, two hundred, four hundred and eight hundred observations; eight hundred systems a point.

With the right relation supplied, it is rejected 9.4% of the time at a hundred observations and 5.4%, 5.4% and 6.4% at two hundred, four hundred and eight hundred. Short samples over-reject, as reduced-rank likelihood-ratio tests usually do, and from two hundred observations the level is close to its promise.

With a relation a tenth wrong — c=0.9c = 0.9 — it rejects 19.9% of the time at a hundred observations, 37.8% at two hundred, 78.5% at four hundred and 98.5% at eight hundred. With c=1.1c = 1.1, 16.8%, 37.9%, 76.8% and 97.5%. The power climbs fast with the sample, faster than the square root of the sample an ordinary coefficient test gains with.

The test’s own level

The over-rejection at a hundred observations is worth a figure of its own, because a test that is read against the wrong critical value is a test whose power figures mean less than they say.

The test's own 5% point and mean with the right relation supplied, beside the chi-square values it is read against. 800 systems a length. The statistic's 95th percentile at 100, 200, 400, 800 observations is 5.12, 3.96, 4.02, 4.17 against the chi-square's 3.84, and its mean 1.35, 1.04, 1.03, 1.11 against 1; read at 3.84 it rejects 9.4%, 5.4%, 5.4%, 6.4%.
Fig. 2 With the right relation supplied: the test statistic’s 95th percentile and its mean at each sample length, beside the chi-square’s 3.84 and 1 that it is read against.

With the right relation supplied, the statistic’s 95th percentile is 5.12 at a hundred observations, 3.96 at two hundred, 4.02 at four hundred and 4.17 at eight hundred, against the chi-square’s 3.84; its mean is 1.35, 1.04, 1.03 and 1.11 against 1. At two hundred observations and beyond the test is the chi-square test it claims to be, to within the noise of eight hundred systems. At a hundred it is not: its statistic is a third larger than the chi-square’s on average, and a reader who rejects at 3.84 rejects a right relation nearly one time in ten.

The cause is the one counting what is still wandering met in the trace test. Every reduced-rank statistic is a function of squared canonical correlations estimated from the sample, and with few observations those correlations are biased upwards — noise lines up with the levels more than it would in a long sample. The supplied-relation test inherits that bias through the restricted and the unrestricted likelihoods alike, and the difference between them does not cancel it. A simulated critical value for the system’s own length repairs it, as it did for the trace test, and from two hundred observations it hardly matters.

The error it can see falls like one over the sample

How small an error in a supplied relation the test can find, by the sample's length. The error in c rejected half the time, below and above the truth: 100 observations, beyond 0.2 and beyond 0.2; 200 observations, 0.128 and 0.127; 400 observations, 0.063 and 0.069; 800 observations, 0.033 and 0.033.
Fig. 3 The error in the supplied coefficient that is rejected half the time, below and above the truth, against the number of observations. At a hundred observations it lies beyond the range tried.

The error the test finds half the time is 0.128 below the truth and 0.127 above at two hundred observations, 0.063 and 0.069 at four hundred, and 0.033 and 0.033 at eight hundred. Each doubling of the sample halves it. At a hundred observations no error in the range tried, up to a fifth, is found half the time.

That rate is the plane’s. A wrong supplied vector does not lie in the true plane, so the combination it describes contains a little of the common trend — a fifth of a random walk at c=0.8c = 0.8 — and a random walk’s variance grows with the sample. A relation that is wrong by a fixed amount is wrong by a growing amount in the units the test sees, and the test’s reach scales as one over the sample, the rate at which a space is not a relation found the plane itself converging. A surplus restriction inherits the precision with which the data pins down the plane, which is the highest precision anything in the system has.

What the restriction does to the forecast

What imposing a supplied relation does to an eight-step forecast, by how wrong the relation is. Mean squared error of the eight-step forecast, summed over the three series, relative to the unrestricted rank-two fit. 200 observations: 0.8 → 1.061, 0.9 → 0.999, 1 → 0.990, 1.1 → 1.041, 1.2 → 1.131; 800 observations: 0.8 → 1.195, 0.9 → 1.073, 1 → 0.999, 1.1 → 1.107, 1.2 → 1.252.
Fig. 4 The eight-step forecast error of the fit that imposes a supplied first relation, summed over the three series, relative to the unrestricted rank-two fit’s, against the supplied coefficient, at two hundred and eight hundred observations. The dashed line is the unrestricted fit.

A right relation supplied in full buys the forecast almost nothing: its error is 0.990 of the unrestricted fit’s at two hundred observations and 0.999 at eight hundred. The unrestricted fit’s estimate of the plane was already close, and pinning one direction of it removes very little.

A wrong one costs, and costs more as the sample grows. At two hundred observations a relation a tenth wrong forecasts at 0.999 of the unrestricted error when too small and 1.041 when too large; a fifth wrong, at 1.061 and 1.131. At eight hundred observations the same relations cost 1.073 and 1.107, and 1.195 and 1.252. The wrong restriction stays as wrong as it was while the unrestricted fit gets better, so its relative cost grows with exactly the sample that makes it detectable.

So the trade the earlier essay described has the asymmetric shape it suspected. Imposing a correct restriction gains a fraction of a per cent of forecast error; imposing a wrong one costs up to a quarter. The fitted relation the unrestricted estimate would have printed is itself close to the truth at these lengths, so pinning it to the true value changes the eight-step error by a per cent or less at both lengths. The case for supplying a relation in full is not its forecast gain. It is that the restriction is checked: at the sample lengths where a wrong relation would cost most, the test finds it nearly every time.

What a tenth wrong means

A supplied coefficient a tenth from the truth is not an abstraction. If the first two series are a price and a cost, theory might say the long-run pass-through is complete, c=1c = 1, where the system’s truth is a pass-through of 0.9 — or say 0.9 where the truth is complete. Those are the claims a pricing study is written to make, and the difference between them is the size of error that, at two hundred observations, the test finds in fewer than two systems in five.

At four hundred observations it finds it in more than three in four, and at eight hundred in nearly all. So the honest statement of a supplied relation’s status depends on the length of the record it was tested on. A complete pass-through that passes its test at two hundred observations has survived a test that catches an error of about an eighth half the time; one that passes at eight hundred has survived one that catches an error of about a thirtieth half the time. The same sentence — “the restriction was not rejected” — carries four times the information in the longer record.

The power is the same on both sides of the truth to within the counting noise — 37.8% at c=0.9c = 0.9 and 37.9% at c=1.1c = 1.1 at two hundred observations — because the test sees the wrong relation’s share of the common trend, which depends on the size of the error and not its sign. The forecast cost is not symmetric: an over-stated coefficient costs more than an under-stated one at every length and size of error measured, which is a property of this system’s adjustment speeds that the counts record and do not explain.

Fewer restrictions, more, or exactly enough

Put the three choices beside each other and the ordering is plain.

Declining to identify reports the plane, which the data estimates at rate one over the sample, and forecasts exactly as well as any exactly identified version. It claims no relations, so it cannot be wrong about them.

Exactly identifying reports two relations and forecasts identically to declining. It claims relations the data cannot refuse, so it can be wrong about them without any sign — the earlier essay’s two normalisations of one fitted plane differed by 2.11 in their largest printed entry while fitting identically.

Over-identifying reports relations that include a claim beyond identification, forecasts slightly better when the claim is right and measurably worse when it is wrong, and puts the claim to a test whose power at two hundred observations is about a third for a tenth’s error and whose reach halves with each doubling of the sample.

Exact identification is the worst of the three in exactly one respect, and it is the respect a paper is read for: it is the only one that prints relations as findings without any way of checking them. Declining says less and is never wrong. Over-identifying says more and can be caught.

What a system’s relations should be reported with

Say how many restrictions identify the relations, and how many are surplus. A rank-two system needs two; a third is testable and the first two are not. A paper that prints relations without saying which restrictions chose them has printed a basis, and a relation supplied in advance showed for a single pair what a supplied coefficient can and cannot confirm.

Test a supplied relation, and report the test with its length. At a hundred observations the test over-rejects and finds little; from two hundred it holds its level and finds a tenth’s error about a third of the time; at eight hundred it finds it nearly always. A passed test at a hundred observations is weak evidence for the relation and a passed test at eight hundred is strong.

Do not impose a relation for the forecast’s sake. A correct one gains a per cent and a wrong one costs up to a quarter; the restriction’s value is the check it makes possible, not the precision it adds.

Count the relations before identifying them. Everything here assumes the rank is two. The rank is a decision found the count’s own errors lopsided, under-counting the common one at short lengths, and a restriction supplied to a system whose rank was under-counted is supplied to the wrong plane. The order is count, then identify, then test what was supplied — and three series and a count is why the count comes first: with three series there is something to count, and a pair-wise analysis never asks.

Prefer the plane when nothing is supplied by theory. A forecast, an impulse response built from the plane, a test of the rank — none of these needs the relations identified, and none of them changes when they are.

Exact, and counted

Exact: any exactly identifying restriction is a choice of basis for the estimated plane, and the error-correction model’s fitted values and forecasts do not depend on it — checked to the eighth decimal on twenty systems at three normalisations each.

Counted, over eight hundred systems at each length and supplied coefficient: the test’s level of 9.4%, 5.4%, 5.4% and 6.4%; its power against c=0.9c = 0.9 of 19.9%, 37.8%, 78.5% and 98.5%; the half-power errors of 0.128, 0.063 and 0.033 below the truth at two hundred, four hundred and eight hundred observations; and the eight-step forecast ratios.

Not claimed: anything about restrictions on the adjustment speeds, which are a different set of claims with their own tests, or about systems with short-run dynamics, where the reduced-rank regression’s small-sample distortions are larger than the 9.4% seen here at a hundred observations. The supplied relation is one direction in a two-dimensional plane; supplying both relations in full would over-identify by two and test both.

Still open: a restriction chosen after looking

Every restriction here was supplied in advance, by theory, and tested once. Practice is less tidy: an analyst estimates the plane, looks at the printed relations under some normalisation, notices that a coefficient is close to a round number, and supplies that round number as the restriction. The test is then run on the data that suggested the restriction.

How much a restriction chosen that way over-states its test’s confirmation — how often a relation rounded from its own estimate passes at 5% when it is in fact a fifth wrong — and whether the one-over-the-sample rate that makes the test powerful also makes rounding nearly harmless at long samples, are measurable on the same systems and have not been measured here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CointegrationThe error-correction modelForecast errorIdentificationLikelihood ratioNormalisationReduced-rank regressionStatistical power