Which mistake about the rank costs
Worth reading first: Three series and a count.
The sequential procedure’s error budget is entirely one-sided: over-counting is held near 5% at every sample length and under-counting is bounded by nothing. That is a fact about frequencies, and a frequency says how often a mistake happens rather than what it does.
A rank is not a conclusion. It is a restriction carried into whatever the system is used for, and the restriction says how many combinations of the series are allowed to pull the system back. Imposing too few means treating a combination that genuinely returns as though it wandered; imposing too many means treating a wandering combination as though it returned. There is no reason for those to cost the same.
Measured on four-step-ahead forecasts from systems whose true rank is known, they do not.
Under-counting is the more expensive mistake in both systems, by a factor of about three against the nearest over-count. It is also the unbounded one. The guarantee is pointed at the cheap error.
The system, and what each restriction has to represent
Reading the lower panels is the whole of what a rank restriction is about. A model at rank two represents both of them. A model at rank one represents one and lets the other wander. A model at rank zero says both wander, which is a statement flatly contradicted by the picture and which the forecast will pay for at a rate the rest of this essay measures.
What a rank restriction actually restricts
The system is written as a set of equations in which each series’ change depends on how far the previous period’s levels were from a set of long-run relations. With r relations there are r such gaps and each series is pulled by each of them; with r = 0 there are none, and every series’ change is a constant plus noise.
So the restriction is a statement about the mechanism. Imposing r = 0 on a system that has two relations is fitting three independent random walks with drift to three series that pull each other back, and the fitted model has no way to express the pulling. Imposing r = 3 on the same system is fitting an unrestricted model that has to estimate three pulls where there are two, and the cost is the estimation error in the third.
That difference in kind is why the costs differ in size. Under-counting is a misspecification — the model cannot represent the mechanism, and the error does not go away with more data. Over-counting is an inefficiency — the model can represent the mechanism and is estimating more parameters than it needs, and that error shrinks as the sample grows. The two are the familiar pair, and the familiar asymmetry between them holds here as it does everywhere: a model with too little structure is wrong and a model with too much is noisy, and being wrong does not improve.
There is a third possibility the two labels do not cover, and it is the reason the rank-one system is in this essay at all. Imposing two relations on a system with one is not merely inefficient: the second relation the model imposes is a combination of levels that genuinely wanders, and the model pulls the system back towards it. It is fitting a mechanism that is not there and acting on it. That reads 4.8% — larger than the corresponding over-count on the rank-two system, and still far smaller than either under-count — which says the damage from believing in a relation that does not exist is real and is dominated by the damage from missing one that does.
Two systems is not a survey, and the two are here precisely so that neither result depends on one arrangement. The rank-one system’s ordering (13.3% against 4.8% and 7.0%) and the rank-two system’s (29.2% and 15.6% against 2.5%) agree, and they agree at different true counts and different numbers of common trends.
The horizon decides how much it matters
A forecast error is not one number and the comparison has a shape in the horizon that neither end of the range shows.
The penalty rises to a maximum around four to eight steps and then falls away. Both halves have the same cause and it is worth naming, because the falling half looks like good news and is not.
It rises because a restriction on the mechanism only shows once the mechanism has acted. One step ahead, the difference between a model that knows about the pull and a model that does not is one period of pulling, which is a fraction of a standard deviation. Four steps ahead it is four periods of it, compounded.
It falls because the forecast error itself grows without bound. Every specification here leaves at least one common trend wandering, and at a hundred and twenty-eight steps the uncertainty from that wandering dominates everything else. The correctly specified model’s own error has grown by more than the penalty has, so the ratio falls. The absolute cost of under-counting does not fall; it is simply a smaller share of a much larger number.
That distinction matters for how the figure should be read. A study forecasting far enough ahead is not protected from a wrong count — it is drowning in a different uncertainty, and the correct reading is that the rank stops being the binding constraint rather than that it stops mattering.
Why the cost is smaller than it feels like it should be
Thirteen and twenty-nine per cent of squared error are real and they are not catastrophic, and a reader who expected a wrong structural claim to be disastrous should see why it is not.
The reason is that the specifications are nested and the mistake is partial. Imposing one relation on a system with two still captures one of them: the model gets half the mechanism, and its forecasts are wrong only in the part that the missing relation governs. That is why imposing one costs 15.6% where imposing none costs 29.2% — the penalty is roughly proportional to how much of the mechanism was discarded, rather than being a cliff at the first error.
It also explains why the over-count is so cheap. An unrestricted fit has to estimate a third adjustment where there are two, and the true value of that third one is zero; the estimate is noisy but unbiased, so the damage is the variance of a coefficient rather than the absence of a term. At two hundred observations that variance is small, which is exactly why the inefficiency framing is the right one.
The practical consequence runs against the usual instinct. A study that cannot decide between two relations and three should take three: the penalty measured here for taking one too many is 2.5%, and the penalty for taking one too few is 15.6%, a factor of six. Where the evidence is genuinely ambiguous, the asymmetry of the losses decides, and it decides against the smaller count.
The same restriction, seen as what it throws away
There is one more way to read the under-counting penalty, and it connects this essay to the field it came from.
Imposing rank zero is differencing everything, and differencing everything is the advice the two-series field arrives at when it cannot tell a real relation from a spurious one. This essay prices that advice on a system where the relation is real: 29.2% of squared forecast error, four steps ahead, for the safety of never claiming a relation that is not there.
That is the same trade differencing makes in the two-series case, arriving with a count instead of a verdict. What the count adds is that the trade is no longer all-or-nothing: a system with two relations has an intermediate option, and taking one of the two is worth about half of taking both.
Where the estimation error actually sits
One reading of the four bars deserves to be separated out, because it is the part that decides whether the advice above survives at other sample lengths.
The over-counting penalty is estimation error in a coefficient whose true value is zero. Its size is therefore governed by how precisely that coefficient is estimated, which improves with the sample; at two hundred observations it is 2.5% and at a shorter sample it is larger. The under-counting penalty is the absence of a term that belongs in the model, and its size is governed by how much of the system’s movement that term accounts for — a quantity that does not change with the sample at all.
So the two penalties respond to the sample in opposite directions and the gap between them widens as data accumulates. That is the opposite of what a reader might expect from a model-selection problem, where more data usually makes the choice easier and the consequences of getting it wrong smaller. Here more data makes one consequence smaller and leaves the other where it was, so the case for erring on the side of more relations gets stronger with the sample rather than weaker.
What a defensible procedure looks like
Do not treat the count as settled before the forecast is made. The order of operations in practice is to establish the rank, impose it, and then use the model, which makes the rank look like a fact the later work rests on. It is an estimate with a known error distribution and a known loss attached, and the later work is where both become visible. This is the same complaint a design criterion makes about a quantity chosen before the data: a number fixed early is not thereby more certain.
Choose the count with the loss in mind, not only the level. The sequential procedure is a 5% rule and 5% is a statement about a direction of error that costs 2.5%. Where the counts adjacent to the reported one cannot be rejected, the loss asymmetry says to take the larger, and the evidence for doing so is already computed on the way to the answer.
Say which direction the count is likely to be wrong in. A count from a short sample is a lower bound far more often than an upper one, which is what the procedure’s own error budget says, and a reader given the integer alone has no way to know that. Two sentences — the length, and the direction the method errs in at that length — turn a bare integer into a claim with a shape.
Report the forecast at the counts the data could not rule out. A single number from a single imposed rank hides that the choice was made, and the alternatives are cheap: refitting at rank one, two and three is three fits. The spread between them is the honest uncertainty about the structure, and it is a quantity a reader can use.
Refit rather than reinterpret when the count changes. A rank is not a parameter that can be adjusted after the fact: changing it changes which combinations are estimated, which coefficients exist, and what the forecast propagates. The three fits are genuinely three models, and comparing them is the same discipline two routes to a number demands everywhere else here — compute it the other way rather than argue about which way is right.
And read the horizon before reading the penalty. A comparison made one step ahead understates the cost of a wrong count by about half, and one made a hundred steps ahead understates it by a factor of seven. The horizon at which the penalty is largest — here four to eight steps — is a property of the adjustment speeds, and the same arithmetic that gives a relation its half-life gives the comparison its maximum.
What is claimed here and what is not
The four bars are not four independent experiments. Every rank is imposed on the same simulated systems, so the comparison is paired and the differences are measured on common draws rather than between separately noisy columns. That is what makes a 2.5% difference readable at eight hundred replications; an unpaired comparison at the same size would not resolve it.
The loss is squared forecast error and there are others. A study whose purpose is to estimate a long-run relation rather than to forecast has a different loss, and nothing here prices that one. There is reason to expect the asymmetry to survive — an unrestricted fit still contains the true relation and a restricted one does not — but “reason to expect” is not a measurement, and the direction is stated rather than shown.
Every fit uses the data’s own estimate of the relations. The combinations imposed are the leading eigenvectors of the same fitted spectrum, not the generating ones, so the penalty includes the error in estimating the relations as well as the error in counting them. That is the right comparison for a practitioner and it is not the decomposition of the penalty into its two parts, which would need the true relations supplied and is not done here.
Two hundred observations, and the costs move with it. At a shorter sample the over-counting penalty grows — it is an estimation-error penalty and estimation error is what a short sample has — while the under-counting penalty does not, since misspecification does not care how much data it is fitted on. So the factor of six above is a factor at this length, and it narrows as the sample shortens. The direction of the asymmetry is robust and its size is not.
“Costs 29.2%” means squared error, not error. A 29.2% increase in mean squared error is about a 13.6% increase in root mean squared error, which is the scale a forecast is usually discussed on. Both are stated the same way throughout — every ratio in this essay is a ratio of squared errors — and the square root is the number to carry if the comparison is being made against a published accuracy figure.
And the forecasts are iterated rather than direct. The fitted system is run forward step by step, feeding each period’s fitted change back in, which is what makes a restriction on the mechanism bite. A direct h-step regression would not use the mechanism the same way and would show a smaller difference between the specifications, because it is not asking the model to propagate anything.
One further consequence, and it is the one that makes the asymmetry awkward rather than merely interesting. The sequential procedure’s under-counting is worst at short samples and its over-counting penalty is worst at short samples too. So the length at which the method is most likely to return too small a count is the length at which taking a larger one is most expensive — the advice “when in doubt take the larger count” is cheapest to follow exactly where doubt is least. Nothing above resolves that; it is the honest shape of the trade and the reason the choice cannot be reduced to a rule of thumb.
The count as a model-selection problem
Read as model selection the result has a familiar shape and one unfamiliar feature.
The familiar part: four nested models, a criterion, and an asymmetry between too little structure and too much. Everything a selection criterion does — a penalty for parameters, a preference for parsimony, a trade between fit and complexity — applies here, and what a search costs in parameters is the same arithmetic seen from the penalty’s side.
The unfamiliar part is that the selection is done by a hypothesis test rather than by a criterion, and a test’s level is not a loss. An information criterion trades fit against complexity at a stated exchange rate; the sequential trace procedure rejects at 5% and that 5% corresponds to no exchange rate anybody chose. The measurements above supply the exchange rate the level implies — over-counting costs 2.5% of forecast error and under-counting 15.6% — and it is not the one a 5% level is calibrated for.
Whether a criterion-based selection of the rank does better on forecast error than the sequential test is a natural question and is not measured here. What is measured is that the test’s asymmetry and the loss’s asymmetry point in opposite directions, which is enough to say the two were never aligned.
Still open: what a count of two entitles anyone to say
Both essays so far have treated the rank as the quantity of interest, and for forecasting it is. For interpretation it is not: a study that finds two relations wants to say what they are, and the output obliges by printing two vectors with the series’ names against them.
Those vectors are not the same kind of object as the count. With one relation there is a single combination that returns and the only ambiguity is its scale. With two, any pair of independent combinations drawn from the plane they span returns equally well, and the software prints one particular pair because it was asked to solve for particular variables. Whether the printed vectors are estimates of anything — and what, exactly, converges as the sample grows — is the next thing this field has to settle, because it decides whether the structural story a rank-two system supports is the one that gets written down.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The cost of differencing a pair — both name cointegration, differencing, the error-correction model, forecast error, random walk
- Which series does the moving — both name cointegrating rank, cointegration, common trend, the error-correction model, random walk
- Which series goes on the left — both name cointegrating rank, cointegration, common trend, random walk, reduced-rank regression
- A line that beats two curves — both name estimation error, model selection, monte carlo, nested models
- How slow a return a sample can see — both name cointegration, the error-correction model, monte carlo, random walk
- What the model says next — both name forecast error, forecast horizon, monte carlo, random walk
Named objects
A flat tag is an object no other essay names yet.
Cointegrating rankCointegrationCommon trendDifferencingThe error-correction modelEstimation errorForecast errorForecast horizonModel selectionMonte CarloNested modelsRandom walkReduced-rank regressionTrace statistic