The model that corrects its error
Worth reading first: The observations that repeat each other · Two walks and a finding.
The two essays before this one establish that a cointegrated pair exists, behaves completely differently from an unrelated one, and can be told apart from it by a test whose critical value has to be simulated. None of that says what the relationship is.
The answer is a mechanism, and it is the most useful object in the field because it is the only one with a unit attached.
Δy = c + α·(y₋₁ − β·x₋₁) + γ·Δx + ε
Today’s change in y depends on yesterday’s disagreement between y and where x says it should have been. α is negative, so a positive disagreement produces a negative change: the gap is closed. That is the error-correction model, and α is the speed of adjustment.
What Granger’s theorem says
The representation is not a modelling choice that happens to fit. It is an equivalence.
A pair is cointegrated if and only if it has an error-correction representation. A stationary linear combination exists exactly when there is a mechanism pulling the pair back towards it. The two statements are the same statement, and the theorem that says so is why this model is the canonical way to write a long-run relationship rather than one option among several.
That equivalence is what makes the generator in this field legitimate. The pairs are built from an error-correction mechanism — Δy = α(y − βx) + η with α < 0 — rather than as “y = βx + stationary noise”, and the two constructions are the same construction. Building it the mechanism way means α is a quantity that was put in, so recovering it is a two-route check rather than an estimate with nothing to compare against.
Recovered
The two-step procedure never sees α. It fits the levels regression to get β̂, forms the gap using β̂, and regresses the change in y on that gap and on the change in x. Nothing in that sequence knows what α was.
Over four hundred pairs of four hundred observations at each of three settings:
| generated α | recovered | standard error of the mean |
|---|---|---|
| −0.1 | −0.1078 | 0.0009 |
| −0.2 | −0.2067 | 0.0012 |
| −0.4 | −0.4057 | 0.0015 |
Recovered to within a hundredth at every setting. The excess is systematic rather than noise — 0.008, 0.007 and 0.006 against standard errors of about 0.001 — and it is the known finite-sample bias of an autoregressive coefficient, which pulls the estimated persistence down and therefore the estimated correction speed up. It shrinks with the length of the series and is the same class of effect what differencing costs had to account for when a measured variance ratio came out 12% from its asymptotic value and turned out to be right.
That is worth a sentence on its own, because it is the difference between a check that means something and one that does not. Asserting that the recovered α is close to the generated one would pass with the bias present and would pass with it three times larger. Asserting that it is within a stated number of the simulation’s own standard errors makes the bias visible: at four hundred pairs the standard error is 0.001, the discrepancy is 0.006 to 0.008, and the check has to be written to allow for a real bias of that size rather than to pretend it is zero.
Why the two-step procedure is allowed here
Using β̂ from one regression as though it were known in a second regression is normally a mistake. The first stage’s error propagates, the second stage’s standard errors do not account for it, and the result is the same defect as the hierarchical plug-in — which costs sixteen percentage points of coverage there.
Here it is allowed, and the licence is the superconsistency measured in the regression that is not spurious. β̂ converges at rate 1/n; everything in the second-stage regression converges at 1/√n. By the time the second stage is precise enough to notice the first stage’s error, that error has already vanished. Asymptotically the second stage behaves as though β were known.
This is the only place on this site where estimate-and-substitute is legitimate, and the condition is narrow: it works because the first-stage regressor is non-stationary. Every other plug-in on this site is a defect, and this one is not, for a reason that does not transfer anywhere.
The asymptotic licence also arrives later than the word “asymptotically” suggests, which is the next essay.
The t on α is not a t either
There is a second reference-distribution problem here and it is easy to walk into after the first one has been dealt with.
Having fitted the error-correction model, the obvious summary is the t statistic on α. It is a coefficient over a standard error, the regression is on stationary quantities, and everything looks ordinary.
It is not ordinary, and the measurement says so. Over four hundred pairs of unrelated random walks — where the true α is zero and there is no gap to close — the fitted α’s t statistic falls below −1.96 on 41.5% of them.
The mechanism is the one from the previous essay wearing different clothes. The gap the second stage regresses on is not a series anyone supplied; it is y − β̂x, with β̂ chosen by the first stage to make that combination look as stationary as possible. A combination selected for looking mean-reverting will appear to be mean-reverting, and the coefficient measuring how fast it reverts inherits the selection.
The size of it is worth putting beside the previous essay’s. The residual test read against a t table calls 70.2% of unrelated pairs cointegrated; the α statistic read against a t table calls 41.5% of them real. Both are catastrophic and the second is the smaller number, which is a trap of its own: a reader who has learned to distrust the residual test’s t and reaches for the α statistic instead has moved to a statistic that is wrong by less and is still wrong.
So the same rule applies: the t on the adjustment coefficient in a two-step error-correction model cannot be read against a t table, for exactly the reason the residual test’s statistic cannot. In practice this is why the test is conducted on the residual rather than on α — the residual test’s null distribution is at least tabulated for common cases — and why an α reported with a t of −2.3 and no further comment is not evidence of anything.
Everything in it is stationary
One property of the representation is worth pausing on because it is what makes the whole thing statistically ordinary once β is known.
The levels regression is a regression of a non-stationary series on a non-stationary series, which is the situation the whole time-series field warns about. The error-correction model is not. Δy is stationary, Δx is stationary, and the gap y₋₁ − βx₋₁ is stationary by the definition of cointegration. Every quantity in the second-stage regression is stationary, so ordinary least squares behaves ordinarily: the coefficients are asymptotically normal, the standard errors mean what they say, and the residual diagnostics are the usual ones.
That is a rare thing in this subject and it is worth being explicit that it is not general. The stationarity of the gap is not an assumption the model makes; it is the conclusion of the test in the previous essay. Fit an error-correction model to a pair that is not cointegrated and the gap is a random walk, the second stage is a regression on a non-stationary regressor again, and every guarantee in this paragraph is void — which is exactly the 41.5% measured above.
The error-correction model is well behaved conditional on the pair being cointegrated, and the condition is not free.
What α is worth knowing
The reason to want α rather than a verdict is that α is a rate, and a rate answers questions a verdict cannot.
How long does a disturbance last? log(0.5)/log(1 + α) steps to half-life: 3.1 at α = −0.2, 0.4 at −0.8, 13.5 at −0.05. If the series are monthly, those are three months, a fortnight and just over a year.
Is the relation strong enough to rely on? A correction with a half-life longer than the span over which a decision has to be made is a relation that will not have reasserted itself in time, whatever a test says about its existence.
How much of a shock survives to next year? (1 + α)^k after k steps: at α = −0.2, a quarter of a disturbance is still present after six steps and a twentieth after thirteen. Those are the numbers a decision is actually made against, and none of them is available from a test result.
Which series adjusts? The model above has y correcting towards x. The symmetric version lets both adjust, and the two adjustment coefficients say which series does the moving — which is usually the substantive question, and is invisible in the levels regression, where the fitted slope says only that the two move together.
That last one is worth emphasising because it is the thing the whole field is for. The levels regression between two cointegrated series gives one number, β, and β is symmetric: regressing x on y gives the reciprocal, and neither ordering is privileged. The error-correction representation breaks the symmetry. It says which series returns to the relation and which one wanders freely, and that is a statement about mechanism rather than about association.
The excess has a closed form from another field
The systematic excess in the recovered α — 0.008, 0.007 and 0.006 against standard errors near 0.001 — is described above as the known finite-sample bias of an autoregressive coefficient. That description can be checked rather than accepted, because the site already carries the closed form for it and the three rows are enough to test it against.
The gap closes as gapₜ = (1 + α)gapₜ₋₁ + noise, so the persistence being estimated is φ = 1 + α, and Kendall’s leading term for the downward bias of a least-squares φ̂ is −(1 + 3φ)/n. At n = 400 that predicts 0.00925, 0.00850 and 0.00700 at the three settings, against counted excesses of 0.0078, 0.0067 and 0.0057.
The ratios are 0.84, 0.79 and 0.81. Constant to within four per cent of themselves, across a threefold range of α, from a formula fitted to nothing and derived for a bare autoregression rather than for a two-step procedure with an estimated regressor in it. The shortfall of about a fifth is what the extra terms — the constant, Δx, and the fact that the gap is formed with β̂ rather than β — are worth, and it is a fixed proportion rather than a drift.
That is a second route to the same three numbers, and it does more than confirm them. It says the excess is not an artefact of the two-step construction, since a formula that knows nothing about the two steps predicts four-fifths of it; and it says the excess will scale as 1/n, so a series of sixteen hundred observations should show a quarter of it. Neither of those follows from observing that the three recovered values are close to their targets.
The bias is worse where the half-life matters most
α is reported as a rate and used as a half-life, and the map between them is log(0.5)/log(1 + α), which is violently non-linear near zero. That means a fixed bias in α is not a fixed bias in the quantity a decision is made on.
Take the excess at face value at the two ends. At α = −0.4 the true half-life is 1.36 steps and an estimate biased by 0.0057 gives 1.33 — short by 1.9%, which nothing would notice. At α = −0.05 the true half-life is 13.5 steps and an estimate biased by 0.008 gives 11.6 — short by 14%, or nearly two steps out of thirteen. Same estimator, same sample length, same order of bias in the coefficient, and seven times the error in the number anybody quotes.
The direction is the unhelpful one. Least squares overstates the speed of adjustment, so it understates how long a disturbance persists, and it understates it most where the correction is slowest — which is exactly the case the essay’s own third figure names as the one a short series cannot establish. A weak long-run relation is reported as stronger and faster than it is, in the currency that decides whether it is worth relying on.
The repair is available and is one line, since the bias has the closed form used above: add (1 + 3φ̂)/n to φ̂ before converting to a half-life. What that costs is the subject of the essay that prices it, and the cost is variance rather than bias — which at four hundred observations is a small price for two steps of half-life.
The three coefficients say different things
The model has three fitted quantities and they are answers to three different questions. Keeping them separate is most of what makes the representation useful.
β, the long-run relation. Where y sits relative to x when everything has settled. It comes from the levels regression, it converges superconsistently, and it is the only one of the three that a levels regression on its own supplies.
γ, the short-run response. How much of a change in x is passed through to y immediately, before any correction happens. It is estimated on differences, converges at the ordinary rate, and is the coefficient a differenced regression would report — the only one of the three it can see.
α, the speed of adjustment. How the gap between the two closes. It is the coefficient that exists in neither the levels regression nor the differenced one, and it is the reason for writing the model this way.
That decomposition settles an argument that otherwise runs in circles. A levels regression is accused of being spurious; a differenced regression is accused of throwing information away. Both accusations are correct about what the other model omits, and the error-correction model contains both — β from the levels and γ from the differences — plus α, which neither has.
What it does not settle
The direction is not causation. A series that adjusts towards another is responding to it in a specific and measurable sense, and that is more than an association. It is still consistent with both responding to something unmeasured, and cointegration testing has no more to say about that than any other observational method.
A recovered α is not a discovered α. Everything in the table above was fitted to data that genuinely had a correction mechanism in it. Recovering a number that was put in is a check on the estimator, not evidence that any particular real pair has one.
One pair is one pair. Everything here is two series. Three or more can have several independent cointegrating relations at once, the counting of them is a rank question rather than a testing one, and the two-step procedure does not extend to it — which is where the Johansen machinery starts and where this field stops.
The relation has been assumed constant. α, β and γ are all fixed over the whole sample here, and a long-run relation that changed halfway through would be fitted as a single wrong relation with a residual that fails every diagnostic. Testing for that is a structural-break question, it needs its own non-standard critical values for the same reason everything else here does, and it is not built.
And the model has been correctly specified throughout. The generator has one lag, no drift, and a correction that is constant over time. Real series have all three of those wrong, and the diagnostics for them are the ordinary residual diagnostics that twenty residual plots is about — applied here to the second-stage regression, whose residuals should be stationary and uncorrelated if the model has captured the dynamics.
What remains, and is the last thing this field owes, is the question the caveat in the time-series field raised: differencing both series is the safe repair, it makes every standard error honest, and it throws the long-run relation away. How much is that worth?
The answer is a forecast comparison and it has an awkward shape. With the equilibrium known, the error-correction model’s one-step forecast error is 0.905 of the differenced model’s, and a closed form in α says so before any data is drawn. With the equilibrium estimated at four hundred observations, it is better on four series in five and worse on average, because a handful of pairs whose β̂ came out badly produce very large errors. Which of those two summaries is quoted decides whether the model looks like an improvement, and the difference between them is the subject of the last essay in this field.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The repair that keeps the question — both name autocorrelation, differencing, spurious regression, stationarity, unit root
- Three series and a count — both name cointegration, differencing, the error-correction model, stationarity, unit root
- The cliff that is a slope — both name autocorrelation, spurious regression, stationarity, unit root
- Which series goes on the left — both name cointegration, spurious regression, stationarity, unit root
- How slow a return a sample can see — both name cointegration, the error-correction model, spurious regression
- Which mistake about the rank costs — both name cointegration, differencing, the error-correction model
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationCointegrationDifferencingThe error-correction modelResidual plotSpurious regressionStationarityUnit root