Comparing two forecasters

What the interval is short by

The forecast interval covers 88.42% where it claims 95%. Correcting the persistence recovers 2.66 points, correcting the innovation variance 0.56, propagating the persistence's own standard error 0.40 — and all three together recover 4.20 of the 6.58, leaving a residual none of the standard repairs reaches.

Worth reading first: Correcting the persistence · What the model says next.

Every essay before this one has reported the forecast interval’s coverage in passing and repaired none of it. At φ=0.85\varphi = 0.85 on fifty observations, six steps ahead, a 95% plug-in interval covers 88.42% — a shortfall of 6.58 points — and the account given so far has been that the persistence is estimated too small, so the interval is too narrow.

That account is true and it is a third of the answer.

What the forecast interval is short by, φ = 0.85, 6 steps aheadThe plug-in interval covers 88.42% against a claimed 95%. Correcting the variance recovers 0.56 points, propagating the persistence's own standard error recovers 0.40, correcting the persistence recovers 2.66, and all three together recover 4.20 — leaving 2.38 points unaccounted for.the plug-in interval88.42%+0.00 points, width 6.159+ the variance corrected88.98%+0.56 points, width 6.281+ the parameter's own error88.82%+0.40 points, width 6.306+ the persistence corrected91.08%+2.66 points, width 6.946+ persistence and variance91.82%+3.40 points, width 7.084+ all three92.62%+4.20 points, width 7.351the 95% it claims5,000 series of 50, φ = 0.852.38 points left over
Fig. 1 The interval’s coverage with each repair applied. Correcting the innovation variance recovers 0.56 points, propagating the persistence’s own standard error recovers 0.40, correcting the persistence recovers 2.66, and all three together recover 4.20 — leaving 2.38 points unaccounted for.

The interval, and what is estimated in it

The plug-in interval is

y^n+h±zσ^j<hφ^2j\hat y_{n+h} \pm z\,\hat\sigma\sqrt{\textstyle\sum_{j<h}\hat\varphi^{2j}}

and three things in it are not what the derivation assumed.

φ^\hat\varphi is biased downwards, by (1+3φ^)/n-(1+3\hat\varphi)/n, which is what correcting the persistence is for. A smaller persistence makes every ψ\psi weight smaller and the sum smaller, so the interval is too narrow.

σ^2\hat\sigma^{2} is biased downwards too, for the ordinary reason that residuals from a fitted model are smaller than the errors they estimate — and for an additional reason here, that the regressor is a lagged value of the response.

And both are treated as known. The formula is the conditional variance of the forecast error given the parameters, and the parameters are estimates with their own sampling error. Nothing in the interval accounts for it.

The first two are biases and the third is not. Separating them is the whole of what follows, because the repairs are different and only one of them is the one usually reached for.

Each one, measured

Applying each repair alone and all three together gives four numbers, and the interesting thing is how different they are in size.

coverage width recovers
the plug-in interval 88.42% 6.1586
+ the variance corrected 88.98% 6.2806 0.56
+ the parameter’s own error 88.82% 6.3058 0.40
+ the persistence corrected 91.08% 6.9463 2.66
+ persistence and variance 91.82% 7.0838 3.40
+ all three 92.62% 7.3510 4.20

The persistence is worth six times the other two put together. That is the honest ranking, and it is worth having because the two smaller repairs are the ones a careful analyst is more likely to think of: correcting a residual variance for degrees of freedom is standard practice, and propagating parameter uncertainty is the textbook complaint about plug-in intervals.

The innovation variance turns out barely to be biased at all. Averaged over six thousand series it comes back at 0.9924 against a true 1 — three quarters of a per cent — because least squares already divides by n2n-2 and the extra downward pull from the lagged regressor is O(1/n)O(1/n) and small at fifty observations. The bias everybody corrects for is the one that is not there.

What the widths say that the coverages do not

The width column of the table is worth reading on its own, because it says the three repairs are doing arithmetically similar things and delivering very different amounts.

The plug-in interval is 6.1586 wide. The variance correction takes it to 6.2806 — two per cent wider — and buys 0.56 points. The parameter-uncertainty term takes it to 6.3058, also two and a half per cent wider, and buys 0.40. The persistence correction takes it to 6.9463, thirteen per cent wider, and buys 2.66.

Points of coverage per per cent of width: 0.28, 0.16 and 0.21. The three are within a factor of two of each other, which says something slightly deflating about all of them.

None of the three repairs is doing anything clever. They are all widening the interval, and what each one buys is very nearly proportional to how much it widens it. The persistence correction is the most valuable because it widens the most, not because it widens in a better place.

That is worth knowing because it sets a ceiling on the whole exercise. If coverage is bought by width at a roughly constant rate, then a repair is only worth having if the width it adds is the width the interval was actually missing — and none of these repairs was derived from a statement about how wide the interval should be. Each was derived from a statement about one input being biased.

The honest test for whether a widening is the right widening is whether it lands at the claimed level, and all three together do not.

Why they do not add

The three repairs recover 0.56, 0.40 and 2.66 points, which sum to 3.62. Applied together they recover 4.20.

Super-additive, and the reason is that coverage is not a linear function of width. An interval covers when the forecast error falls inside it, and widening a short interval from 6.16 to 7.35 moves its endpoints through a region of the error distribution where the density is still high — so each increment of width buys more coverage than the last, until the endpoints reach the tail.

That is a general property and it has a practical form worth stating. Partial repairs of an interval under-report what full repair is worth. An analyst who tries the variance correction alone, sees half a point, and concludes the effect is negligible has measured the effect at the wrong width.

It also means the three contributions cannot be attributed. “The persistence accounts for 2.66 of the 6.58 points” is a statement about what happens when only the persistence is corrected, and the decomposition has no unique answer — the same objection that makes variance decomposition in a correlated regression ambiguous, arriving in a place where the components are repairs rather than predictors.

What is left over

Four point two zero of six point five eight is recovered and 2.38 points are not, and the residual is the part this essay cannot close.

The interval's shortfall against the horizon, φ = 0.85. At one step the plug-in interval covers 93.40% and at twelve it covers 86.92%. All three repairs together bring twelve steps to 91.85%, which is still 3.15 points short, and the gap between the repaired curve and the claim widens with the horizon rather than closing.
Fig. 2 The same decomposition across horizons. At one step the plug-in interval covers 93.46% and all three repairs bring it to 94.70%. At twelve steps it covers 86.92% and they bring it to 91.85% — still 3.15 points short, and the residual grows with the horizon rather than closing.

Three things are in the residual and none of them is a bias in a plug-in.

The propagation is linear and the function is not. The extra variance added is (y^/φ)2Var(φ^)(\partial \hat y/\partial\varphi)^{2}\operatorname{Var}(\hat\varphi), which is the delta method, and the forecast is φh\varphi^{h} times a deviation — convex, over a spread that is not small. That is the same failure the previous field measured directly, where the delta method predicted a decay factor of −0.0003 at twelve steps.

The reversion target is estimated too. The forecast decays towards a mean, the mean is the sample mean of the window, and its own uncertainty is not in the interval either.

And the forecast error is not normal. It is a sum of future shocks — normal — plus a product of an estimated power and an observed deviation, which is not. The zz multiplier assumes otherwise.

The first is testable by replacing the delta method with a simulation; the second by adding a term; the third by bootstrapping the whole forecast distribution rather than its variance. None of the three is a plug-in bias, and none is repaired by correcting anything.

At high persistence, where nothing is enough

The setting the field uses is a mild one, and the same decomposition at a hard setting says how far the repairs reach.

What the forecast interval is short by, φ = 0.95, 6 steps ahead. The plug-in interval covers 84.60% against a claimed 95%. Correcting the variance recovers 0.76 points, propagating the persistence's own standard error recovers 1.38, correcting the persistence recovers 5.38, and all three together recover 7.34 — leaving 3.06 points unaccounted for.
Fig. 3 At φ = 0.95 the plug-in interval covers 84.60%. The variance correction recovers 0.76 points, the parameter’s own error 1.38, the persistence 5.38, and all three 7.34 — leaving 3.06 points, which is a larger residual than the whole shortfall at the mild setting.

Two things change and they go in opposite directions. The persistence correction becomes much more valuable — 5.38 points against 2.66 — because the bias it removes is larger. And the residual grows too, from 2.38 to 3.06, because the nonlinearity the delta method is approximating over is more severe.

So the repairs are worth more where they are more needed and they do not keep up. At φ=0.95\varphi = 0.95 an interval with all three repairs applied still covers 91.94% against a claimed 95%, and an analyst who applied all of them would have a better interval and not a correct one.

What the forecast interval is short by, φ = 0.7, 6 steps ahead. The plug-in interval covers 91.34% against a claimed 95%. Correcting the variance recovers 0.40 points, propagating the persistence's own standard error recovers 0.08, correcting the persistence recovers 1.34, and all three together recover 2.16 — leaving 1.50 points unaccounted for.
Fig. 4 And the easy setting, for the other end of the range. At φ = 0.7 the plug-in interval covers 91.34% and all three repairs bring it to 93.50% — a shortfall of 3.66 points before anything is done, and a residual of 1.50 after. Both numbers are roughly half what they are at 0.85, so the whole difficulty scales with the persistence rather than switching on at some point in it.

The residual is the same shape at every setting

Putting the three settings side by side says something the individual decompositions do not, and it is the reason the residual looks like one thing rather than three.

plug-in all three shortfall residual residual as a share
φ = 0.7 91.34% 93.50% 3.66 1.50 41%
φ = 0.85 88.42% 92.62% 6.58 2.38 36%
φ = 0.95 84.60% 91.94% 10.40 3.06 29%

The shortfall nearly triples across the range and the residual roughly doubles, so the share of the shortfall the three repairs fail to reach falls from 41% to 29%. The repairs get relatively more effective as the problem gets worse, and absolutely further behind.

That is the signature of a residual made of something that grows more slowly than the biases do — a nonlinearity that is present at every setting and severe at none of them — rather than of a fourth bias nobody has named. A missing bias correction would scale with the same 1/n1/n the other two do; this does not.

How often the correction leaves the stationary region. The closed-form threshold is φ̂ > (n − 1)/(n + 3), which is 0.8571 at n = 25, 0.9245 at n = 50, 0.9612 at n = 100, 0.9803 at n = 200. Counted over 3,000 series at each cell, the rate reaches 37.9% at φ = 0.99 on 25 observations and falls to 29.5% at 200.
Fig. 5 A reminder of what the persistence correction is doing while it recovers those points: at high persistence and short series it leaves the stationary region on a third of samples, so the repaired interval in the table above is a repaired-and-capped interval. The correction that leaves the region is about what that cap costs, and it is another decision inside the number 91.94%.

What a forecaster should do

The measurements support an ordering, and it is not the ordering a methods checklist would give.

Correct the persistence. It is the largest of the three by a factor of six, it costs one multiplication, and it is the repair that makes the interval’s width approximately right. That it makes the point forecast slightly worse at moderate persistence is a separate accounting and does not change this one.

Do not bother correcting the innovation variance. It is worth half a point at fifty observations and less at a hundred, it is the repair most likely to be applied out of habit, and the bias it removes is three quarters of a per cent.

Scored on the forecast rather than on the parameter. The whole family of corrections on one axis, at φ = 0.85, 50 observations and 6 steps ahead, 4000 series per point. Zero is the plug-in forecast and one is the correction the formula prescribes. The upper curve is the interval's counted coverage, relative to the 95% it claims; the lower is the mean squared forecast error, relative to no correction. They move in opposite directions: coverage rises from 88.2% to 91.3% while squared error rises by 5.1%. The interval improves because a larger φ̂ makes the plug-in band wider, which is not what the correction was for.
Fig. 6 The whole family of persistence corrections on one axis, from none to half again as much, with coverage above and squared forecast error below. The coverage curve rises across the range and the error curve rises with it — which is the trade the first recommendation here accepts and the repair that moves the wrong number scored.

Propagate the parameter uncertainty if the horizon is short. It is worth 0.40 points at six steps and 0.30 at one, and at long horizons its delta-method form is approximating a function it cannot follow. Where the horizon is long, simulate the forecast distribution instead.

And do not expect any of it to reach 95%. With every repair applied, the interval covers 92.62% at the field’s standard setting and 91.85% at twelve steps. A forecaster who needs the stated level needs a different construction — a bootstrap of the whole predictive distribution — rather than a better set of plug-ins.

Why the persistence is worth so much more

The factor of six between the persistence repair and the other two is not an accident of this setting, and the reason is one line of arithmetic.

The interval’s half-width is zσ^j<hφ^2jz\hat\sigma\sqrt{\sum_{j<h}\hat\varphi^{2j}}. The variance enters through σ^\hat\sigma, which is a square root — a 0.76% bias in σ^2\hat\sigma^{2} is a 0.38% bias in the width. The persistence enters through a sum of powers, and at φ=0.85\varphi = 0.85 and six steps the derivative of that sum with respect to φ\varphi is large: the sum is 3.0910 and moving φ\varphi from 0.77 to 0.85 moves it from 2.3497 to 3.0910, which is 32%.

So the same relative bias in the two inputs produces very different relative errors in the width, because one input is under a square root and the other is amplified by the horizon. The persistence’s bias is both larger and more amplified, and the two effects multiply exactly as they did in the gain from pooling one field over.

The corollary is about horizons. The amplification grows with hh — the sum has more terms and each is a higher power — so the persistence repair becomes more valuable as the horizon lengthens while the variance repair stays where it is. At one step the sum is 1 and does not depend on φ\varphi at all, which is why the persistence correction is worth almost nothing there: 93.46% against 93.74%.

The interval nobody quotes, which is the one that works

There is a construction that sidesteps every difficulty in this essay, and this field has already measured it once without saying so.

The plug-in interval is a formula: a normal multiplier times an estimated standard error. Every repair here adjusts an input to that formula, and the residual is where the formula’s own form is wrong. The alternative is to stop using a formula — draw a parameter from its sampling distribution, draw hh shocks, propagate the fitted model forward, repeat, and read the quantiles of the simulated futures.

That construction has no delta method in it, makes no normality assumption about the forecast error, needs no bias correction and has no upper bound for a corrected persistence to cross. It has the property none of the three repairs has: it is a statement about the distribution of the forecast error rather than about its variance.

What it costs is what a bootstrap costs, and what it needs is the one thing this field has been careful about all along — a way of drawing parameters that is not itself the plug-in. Whether it covers is a measurement nobody here has made, and it is the obvious next one.

What is claimed here, and what is not

The claim is a decomposition of a forecast interval’s coverage shortfall: that correcting the persistence recovers 2.66 points of 6.58, correcting the innovation variance 0.56, propagating the persistence’s own standard error 0.40, and all three together 4.20 rather than their sum of 3.62; that the innovation variance is biased downwards by only three quarters of a per cent, so the repair most often applied is the least valuable; and that a residual of 2.38 points at six steps and 3.15 at twelve is not reached by any of the three.

Every number is four to six thousand series with the persistence, the variance and the reversion target all estimated from the same window.

What stays out: the bootstrap predictive interval, which is the construction the essay’s conclusion points at and which would need its own coverage counted before being recommended; the exact finite-sample distribution of the forecast error, which is available for a Gaussian AR(1) in principle and is not a closed form anybody uses; and intervals for the whole forecast path rather than for one horizon, where the multiplicity across horizons is a further problem this field has met in the survival setting and not here.

The variance correction used here scales σ^2\hat\sigma^{2} by 1+2/n1 + 2/n, which is a stated construction rather than a derived finite-sample bias — the exact bias of the conditional least-squares innovation variance for an AR(1) depends on φ\varphi and is not a one-line expression. Since the measured bias is 0.76% and the correction applies 4%, the repair over-corrects, and the 0.56 points it recovers are therefore an upper bound on what an exact variance correction would be worth. That strengthens the essay’s conclusion about it rather than weakening it, which is why the construction is reported rather than tuned.

Still open: the interval that is simulated rather than derived

Every repair here is a correction to a formula, and the formula’s form — a normal multiplier times a plug-in standard error — is the assumption none of the repairs touches. The residual that survives all three is where that assumption lives.

The alternative is to simulate the predictive distribution: draw parameters from their sampling distribution, draw shocks, propagate the series forward, and take the quantiles of what comes out. It needs no delta method, no normality of the forecast error and no correction to anything, and it costs what a bootstrap costs. Whether it covers is a question this field has not asked, and it is the one that would settle whether the residual is the assumption or something else again.

The check, and the refusal

Three claims are gated and they are deliberately stacked. That the plug-in interval covers less than it claims — the premise, and a check that would catch a sign error in the whole apparatus. That every repair moves it in the right direction, which would fail if any of the three were implemented backwards. And that all three together still leave it short, which is the finding.

The refusal is the third of those, read as a refusal: the repairs must not be enough. If all three together reached 95%, the residual this essay is about would not exist and the decomposition would be complete — which would be a better result and a different essay. Requiring it is what stops the finding being an artefact of a repair implemented too weakly, since a too-weak repair is exactly what would leave a residual by accident. The horizon sweep carries the same requirement at every horizon, so a residual that appeared at one and vanished at another would fail rather than be reported.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Bias correctionCoverageDelta methodEstimated varianceForecast horizonForecast intervalMean squared errorParameter uncertaintyPlug in estimateStationarity