Comparing two forecasters

An interval built from simulated futures

A forecast interval read off simulated futures needs no delta method and no normal multiplier, and was expected to close the residual the three repairs to the formula leave. It does not close it on its own: the plain residual bootstrap covers 88.25% six steps ahead at φ = 0.85, the plug-in formula's 88.48% to within the count, because it simulates around the same biased persistence. With that bias removed inside the bootstrap it covers 93.47%, the best of five intervals, and its shortfall stops growing with the horizon — 93.80% at twelve steps, where the repaired formula covers 91.85%.

Worth reading first: What the model says next · Correcting the persistence.

What the interval is short by took a 95% forecast interval for an AR(1) six steps ahead — fifty observations, persistence 0.85 — which covers 88.42%, and repaired it three ways: correcting the persistence’s downward bias, correcting the innovation variance, and propagating the persistence’s own standard error. Together they recovered 4.20 of the 6.58 missing points. The rest it attributed to the formula itself — a normal multiplier times a standard error linearised through the delta method — and it named the construction that has no formula to be wrong: simulate the futures, and read the interval off their quantiles.

That construction is measured here, in three forms, on the same series the formula and its repairs were counted on. The expectation was that simulation would close the residual. What it shows instead is that the residual has two parts, and simulation removes only one of them.

Three ways to simulate a future

Every simulated interval builds many possible values of the series h steps ahead from the fitted model, starting at the last observation, and reports the 2.5th and 97.5th percentiles of those values. None of them assumes the forecast error is normal, and none of them uses a derivative. They differ in how they put the model’s own uncertainty into the futures.

Parameters drawn, then the future. Draw a persistence from the normal centred on the estimate with its standard error, a centre from the sample mean’s own sampling spread, and an innovation variance from the scaled χ2\chi^2 of the residual variance; then run the AR(1) forward six steps with normal shocks. This is the construction the earlier essay described, and the one a Bayesian simulation with flat priors approximates.

The residual bootstrap. Rebuild the whole series from the estimated persistence and resampled residuals, refit the model to the rebuilt series, and run the refitted model forward from the actual last observation with resampled shocks. The spread of the refitted persistence stands in for its sampling distribution, and the residuals stand in for the shocks, so nothing is assumed normal at all.

The bootstrap with the bias removed. The same, except that the series are rebuilt from the corrected persistence, φ^+(1+3φ^)/n\hat\varphi + (1 + 3\hat\varphi)/n, and each refitted persistence is corrected the same way before it is run forward. This is Kilian’s bias-corrected bootstrap, with the analytic correction of correcting the persistence standing in for his inner bootstrap, and capped at 0.995 so that no rebuilt series is explosive.

Each interval uses 499 simulated futures per series, and every one of them is counted on the same 4,000 series as the formula, so differences between them are differences of construction rather than of sample.

Simulation alone reproduces the formula’s failure

Five 95% intervals for an AR(1) forecast 6 steps ahead at φ = 0.85, fifty observationsCoverage on 4,000 series, the same series for every interval, each simulated interval built from 499 simulated futures. The plug-in formula: 88.48%. The formula, all three repairs: 92.80%. Simulated: parameters drawn: 92.55%. Simulated: residual bootstrap: 88.25%. Simulated: bootstrap, bias removed: 93.47%.80%85%90%95%100%the plug-in formula88.48%the formula, all three repairs92.80%simulated: parameters drawn92.55%simulated: residual bootstrap88.25%simulated: bootstrap, bias removed93.47%4,000 series; the axis starts at 80%percentile intervals of simulated futures assume no normality
Fig. 1 Coverage of five nominal 95% intervals for an AR(1) forecast six steps ahead at φ = 0.85 and fifty observations, on the same 4,000 series. The plain bootstrap covers what the plug-in formula covers; the bootstrap with the persistence’s bias removed covers the most.

The plug-in formula covers 88.48% of these series. The residual bootstrap covers 88.25% — the same number, to within a count whose standard error is half a point. Rebuilding the series, refitting them and reading percentiles instead of multiplying a standard error has bought nothing.

The reason is what the bootstrap’s spread is a spread of. Each rebuilt series is generated from the estimated persistence, which is too small; each refit of it is too small again, by the same bias, since fitting a short AR(1) underestimates whatever persistence generated it. So the bootstrap’s distribution of persistences is centred below an estimate that is itself below the truth, and the futures run forward under those persistences decay towards the mean too quickly. The bootstrap faithfully reproduces the sampling distribution of the estimator — and the estimator is biased, so it reproduces the bias too, twice. Its intervals are no wider than the formula’s, 6.05 in median width against the formula’s 6.16 in mean width, and they miss on the low side 6.08% of the time against a target of 2.5%.

The size of the double bias can be read off Kendall’s approximation, which is the correction this field has used throughout. A series with persistence 0.85 and fifty observations is estimated, on average, at 0.85−(1+3×0.85)/50=0.85 - (1 + 3 \times 0.85)/50 = 0.779. A series rebuilt from 0.779 is estimated, on average, at 0.712. Six steps ahead the forecast moves the last deviation by the persistence to the sixth power, and those three persistences give decay factors of 0.377, 0.223 and 0.131. The bootstrap’s typical future keeps about a third of the deviation the real series keeps, which is a statement about where the futures are centred, and centring them wrongly costs coverage however their spread is computed. Correcting the forecast instead aimed its correction at that decay factor directly, because it is the quantity a forecast actually uses; inside a bootstrap the same quantity is biased twice over.

Drawing the parameters from their normal sampling distribution does better, 92.55%, and does it in a way that shows what the formula was missing. Its draws of the persistence are spread symmetrically around the estimate, so half of them land above it, many of them near the persistence the series actually has, and the futures run forward under those draws decay as slowly as the real series does. That adds the width the plug-in formula lacked — a median of 6.94 — and recovers four points of coverage, about as much as the three repairs to the formula together, which cover 92.80%. But it adds the width in the most expensive place. Draws near a unit root produce futures that barely decay, and the mean width of its intervals is 27.73, four times the median: a few series get intervals so wide they say nothing, which is what an interval built on an unbounded draw of the persistence does when the persistence is near the bound.

The waste is not confined to one setting. At one step ahead the parameter draws give intervals with a mean width of 8.33 and a median of 4.10; at twelve steps, 49.33 and 7.56. The gap between the two summaries grows with the horizon because a persistence near one matters more the further it is raised to a power. A reader shown a single interval from this construction usually sees an ordinary one, close to the median; a reader shown many sees that a few of them are unusable, and the average width, which is what a comparison of methods usually reports, is set by those few. That is why the comparisons here quote medians beside means, and why “parameters drawn from their sampling distribution” is a description that needs its distribution stated: the draws here are clipped at 0.995, and a draw kept further from the unit root, or drawn from a posterior rather than a normal, would behave differently and is not measured.

Removing the bias inside the simulation

The third construction does both things the first two did halfway. It uses the bootstrap, so the spread of the futures comes from refitting real-looking series rather than from a normal approximation, and it removes the persistence’s bias both where the series are rebuilt and where they are refitted.

It covers 93.47%, the most of the five, with a median width of 7.46 and misses split 3.40% below and 3.13% above — nearly symmetric, which neither the formula nor the plain bootstrap manages. Against the three repairs to the formula it gains seven tenths of a point and costs a slightly wider interval, 7.54 against 7.36 on average.

So the earlier essay’s residual was not all formula. Part of it was the formula’s form — the normal multiplier and the linear propagation — and that part simulation removes. Part of it was the persistence’s bias being reproduced inside every construction that refits the model, and only correcting the bias inside the simulation removes that. The bias-corrected bootstrap is the first interval in this field that does both, and the shortfall it leaves, 1.52 points, is smaller than any other construction’s here.

The horizon, where the difference is largest

Derived and simulated forecast intervals at φ = 0.85, by the horizon. Coverage at one, six and twelve steps ahead, fifty observations, 4,000 series each. Plug-in formula: 93.40%, 88.48%, 86.92%. Formula, three repairs: 94.67%, 92.80%, 91.85%. Residual bootstrap: 92.90%, 88.25%, 86.88%. Bootstrap, bias removed: 93.38%, 93.47%, 93.80%.
Fig. 2 Coverage at one, six and twelve steps ahead, φ = 0.85, fifty observations: the plug-in formula, the formula with all three repairs, the residual bootstrap, and the bootstrap with the bias removed. The corrected bootstrap’s coverage barely moves with the horizon; every other interval’s falls.

At twelve steps the plug-in formula covers 86.92% and the plain bootstrap 86.88% — again the same interval in different clothes. The formula with all three repairs covers 91.85%, the shortfall the earlier essay found growing with the horizon. The bias-corrected bootstrap covers 93.80%: more than at six steps, not less. Its coverage does not fall with the horizon because it no longer has a linearisation to break. The delta method’s propagation of the persistence’s error was accurate near the estimate and poor far from it, and at twelve steps the forecast is φ12\varphi^{12} times a deviation, a function curved enough that a tangent line misreads it. The bootstrap evaluates the curve at every draw.

At one step the order reverses. The formula with its repairs covers 94.67%, the bias-corrected bootstrap 93.38%, and the plain bootstrap 92.90%. One step ahead the forecast is φ\varphi times a deviation plus one shock, the persistence’s bias hardly matters, and what matters is the size of the shock. The bootstraps draw their shocks from the residuals, which are smaller than the true shocks because the model was fitted to them; the repaired formula scales the innovation variance up and the bootstraps do not. It is the one repair the earlier essay judged least valuable, and at one step it is the one that decides the comparison.

How much the shocks are shrunk can be worked out rather than guessed. Fifty observations give forty-nine residuals from a fit with two parameters, and resampling them with equal probability draws shocks whose variance is 47⁄49 of the estimated innovation variance — which the earlier essay measured as 0.76% below the truth to begin with. Together that is 95.19% of the true variance, a spread 2.44% too small. A normal interval narrowed by that much covers 94.42% rather than 95%. So the shrunken shocks explain about six tenths of a point of the corrected bootstrap’s one-step shortfall of 1.62 points. The other point is not accounted for by anything measured here, and it is the part the last section returns to.

Across the persistence

Five 95% intervals for an AR(1) forecast 6 steps ahead at φ = 0.95, fifty observations. Coverage on 4,000 series, the same series for every interval, each simulated interval built from 499 simulated futures. The plug-in formula: 84.72%. The formula, all three repairs: 92.22%. Simulated: parameters drawn: 90.45%. Simulated: residual bootstrap: 84.28%. Simulated: bootstrap, bias removed: 93.38%.
Fig. 3 The same five intervals at φ = 0.95, six steps, fifty observations. The plug-in formula and the plain bootstrap fall below 85%; the bias-corrected bootstrap holds above 93%.

At φ = 0.95 the persistence’s bias is largest and every construction that reproduces it suffers most: the plug-in formula covers 84.72% and the plain bootstrap 84.28%. The formula with its three repairs covers 92.22%, the parameter draws 90.45% — with a mean width of 569.86 against a median of 8.10, since near a unit root a normal draw of the persistence crosses it often — and the bias-corrected bootstrap 93.38%, barely below its coverage at 0.85. At φ = 0.7 the five are closer together, from 91.07% for the plain bootstrap to 94.13% for the corrected one.

So the corrected bootstrap’s advantage grows with the persistence, from about half a point over the repaired formula at 0.7 to more than a point at 0.95, while its coverage stays between 93.38% and 94.13% across the range. That steadiness is the property the earlier essay found missing from every repair: the repairs got relatively more effective as the problem got worse and absolutely further behind, and this construction does not fall behind.

Where each misses

Where each simulated interval misses, at φ = 0.85 and six steps. The share of series whose future falls below each simulated interval and above it, against 2.5% each side for a correct interval. Simulated: parameters drawn: 3.98% below, 3.48% above, median width 6.94, mean width 27.73. Simulated: residual bootstrap: 6.08% below, 5.68% above, median width 6.05, mean width 6.17. Simulated: bootstrap, bias removed: 3.40% below, 3.13% above, median width 7.46, mean width 7.54.
Fig. 4 The share of series whose future falls below each simulated interval and above it, at φ = 0.85 and six steps, against 2.5% each side for a correct interval, with each interval’s median width.

An interval that misses too often on one side is wrong in a way coverage alone hides, and the three simulated intervals differ there more than in coverage. The plain bootstrap misses 6.08% below and 5.68% above: too narrow on both sides, a little more below, which is the side the too-quick decay towards the mean pushes the futures away from when the last observation sits above it. The parameter draws miss 3.98% below and 3.48% above. The bias-corrected bootstrap misses 3.40% below and 3.13% above — the most symmetric of the three and the closest to 2.5% on each side.

What remains, and where it lives

The corrected bootstrap still covers 93.47% rather than 95%, and the measurements say where the last point and a half is. It is not the horizon, since the shortfall does not grow with it. It is not the persistence, since the coverage hardly moves between 0.85 and 0.95. The one-step result points at the shocks: the residuals the bootstrap resamples are a little smaller than the errors they stand for, because the model was fitted to make them small, and nothing in the construction scales them back up. That costs about the same at every horizon and every persistence, which is the shape of what remains.

The repair is the one the earlier essay found worth half a point on the formula — inflate the residuals by the factor the fit shrank them by before resampling them — and it is cheap. It is also not measured here, and whether it closes the last point and a half, or reveals a third component under it, is the next count.

A second thing remains that no interval here addresses: the cap at 0.995. At φ = 0.95 and fifty observations the corrected persistence leaves the stationary region on 24.6% of series, and the corrected bootstrap’s 93.38% there is a capped-and-corrected number. Every choice inside it is stated and none of it is invisible to the coverage, which is the standard two routes to every number set for this subject; but a different cap would give a different number, and the honest description of the interval is “bias-corrected, capped at 0.995”.

What a forecast interval should say about itself

How the uncertainty in the parameters entered it. A formula with a plug-in persistence, a formula with the persistence’s standard error propagated, a simulation from drawn parameters and a bootstrap are four different statements, and on these series they cover between 88.25% and 93.47%.

Whether the persistence’s bias was corrected, and where. A bootstrap that refits without correcting reproduces the bias, the way the interval that forgets it estimated forgets its parameters; correcting the point estimate alone, without correcting inside the simulation, leaves the refitted persistences biased.

Which side it misses on. A forecast that is too narrow on the low side is a different hazard from one too narrow on the high side, and only the bias-corrected bootstrap here splits its misses evenly.

Counted, on what

Four thousand series at each setting, the same series for all five intervals, generated from an AR(1) with unit innovation variance and a burn-in of two hundred; fifty observations, forecasts one, six and twelve steps ahead; persistences 0.7, 0.85 and 0.95. Each simulated interval uses 499 futures per series, drawn from a random stream separate from the one that generated the series, so the series match the formula’s counts draw for draw. The standard error of a coverage near 93% on four thousand series is 0.4 points, so differences under about a point between two intervals are within the count; the differences the argument rests on — the plain bootstrap against the corrected one, and the corrected one against the repaired formula at twelve steps — are several times that. The constructions are as described above and nothing else: no studentised bootstrap, no bootstrap of the bootstrap, no inflation of the residuals.

Still open: the shocks the bootstrap resamples

The corrected bootstrap leaves about a point and a half at every horizon and every persistence measured, and the one-step comparison suggests it is the size of the resampled shocks. Two repairs are standard and neither is measured here. Residuals can be inflated by (n−1)/(n−3)\sqrt{(n-1)/(n-3)}, the degrees-of-freedom factor the fit removed, or each residual can be divided by 1−ht\sqrt{1 - h_t} with hth_t its leverage in the AR regression, which inflates most the residuals the fit pulled in most.

Whether either closes the last point and a half — and whether, once it does, any shortfall remains that grows with the horizon or the persistence — would settle whether the forecast interval for a short AR(1) can be made to cover what it claims from fifty observations at all. It also bears on where the bootstrap lies: a resample cannot contain an error larger than the largest residual, and an inflated residual is the smallest possible step past that limit.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutoregressionBias correctionBootstrapCoverageDelta methodForecast horizonForecast intervalMonte CarloParameter uncertaintyPercentile intervalPlug in estimateStationarity