Comparing two forecasters

Correcting the forecast instead

The complaint against the usual repair is that a correction aimed at the persistence lands on the wrong quantity. Aiming it at the decay factor the forecast actually uses fixes exactly that — the error stops compounding with the horizon, 69.7% becomes 9.5% at twelve steps — and the forecast still gets worse.

Worth reading first: Correcting the persistence.

The repair that moves the wrong number ends with a named deferral. Correcting the persistence and then raising it to a power is aimed at φ\varphi while the forecast uses φh\varphi^{h}, and the expectation of a power is not the power of an expectation, so the corrected estimate overshoots by 69% at twelve steps. The obvious response is to correct φh\varphi^{h} instead, and that essay set it aside as “a different construction”.

It is one line different, and it is run here.

The average decay factor each route produces, φ = 0.85, 6 steps aheadThe truth is φ^6 = 0.3771. no correction averages 0.2616 with a spread of 0.1646 and a squared forecast error of 3.2516; the formula, on the persistence averages 0.4213 with a spread of 0.2528 and a squared forecast error of 3.4827; the bootstrap, on the persistence averages 0.4355 with a spread of 0.2655 and a squared forecast error of 3.5120; the bootstrap, on the decay factor averages 0.3375 with a spread of 0.2278 and a squared forecast error of 3.4132. 800 series, 100 bootstrap refits each.no correction0.2616bias −0.1155, spread 0.165the formula, on the persistence0.4213bias +0.0442, spread 0.253the bootstrap, on the persistence0.4355bias +0.0583, spread 0.265the bootstrap, on the decay factor0.3375bias −0.0396, spread 0.228the truth, 0.3771800 series of 50, 100 refits eachaimed at the right quantity
Fig. 1 Four decay factors from the same eight hundred series, at φ = 0.85 and six steps ahead. The truth is 0.3771. No correction averages 0.2616; the closed form on the persistence averages 0.4213; the bootstrap on the persistence averages 0.4355; and the bootstrap on the decay factor averages 0.3375.

The one line

The bootstrap bias correction is a general device and its statement does not mention which quantity it is correcting.

Take the fitted model. Simulate from it many times. Refit each simulated series and record the estimate. The average of those estimates is what the estimator does, on average, when the truth is the fitted value — so the difference between that average and the fitted value is an estimate of the bias, and subtracting it gives the correction:

θ^corrected=2θ^θ^\hat\theta_{\text{corrected}} = 2\hat\theta - \overline{\hat\theta^{*}}

Everything about which quantity is being corrected lives in what θ^\hat\theta is. Set it to φ^\hat \varphi and the procedure repairs the persistence, which is what correcting the persistence does by closed form, and what a bootstrap of the same quantity does by simulation. Set it to φ^h\hat\varphi^{h} — record φ^h\hat\varphi^{*h} from each simulated series rather than φ^\hat\varphi^{*} — and it repairs the decay factor.

No new theory, no new derivation, and no formula to get wrong. The correction is whatever the estimator does on data it believes.

It works on the quantity it is aimed at

The first thing to check is whether aiming at the right estimand does what aiming is supposed to do, and it does.

How wrong each decay factor is, against the horizon. Three routes at five horizons, 500 series each. As a share of the truth, the uncorrected decay factor is off by -9.3% at one step and -35.4% at twelve; the persistence-corrected route by 63.5% at twelve and the decay-corrected route by -14.2%.
Fig. 2 The relative error in the decay factor at five horizons. The uncorrected route is 9.3% short at one step and 35.4% short at twelve. The persistence-corrected route is accurate at one step and 63.5% long at twelve. The decay-corrected route is between 0.8% and 14.2% wrong across the whole range.

The shape is the argument. The persistence-corrected route’s error compounds: the correction is additive, an additive offset near the top of the stationary range is a large proportional one, and the horizon raises a proportional error to a power. The decay-corrected route’s error does not compound, because the correction is applied after the power rather than before it.

The crossover is a measurement rather than a principle, and it is worth stating because it is later than the argument suggests. At two steps the persistence-corrected route is still 1.3% short; by four it is 2.9% long. The overshoot needs the convexity of the power to beat the correction’s own accuracy, and at two steps it has not yet.

So that diagnosis is confirmed and the prescription that follows from it works. The estimand was the problem, and correcting the right estimand fixes the estimand.

The average decay factor each route produces, φ = 0.85, 12 steps ahead. The truth is φ^12 = 0.1422. no correction averages 0.0955 with a spread of 0.1089 and a squared forecast error of 4.0109; the formula, on the persistence averages 0.2414 with a spread of 0.2516 and a squared forecast error of 4.4896; the bootstrap, on the persistence averages 0.2600 with a spread of 0.2721 and a squared forecast error of 4.5679; the bootstrap, on the decay factor averages 0.1287 with a spread of 0.1619 and a squared forecast error of 4.1674. 800 series, 100 bootstrap refits each.
Fig. 3 Twelve steps ahead, where the difference between the two aims is largest. The truth is 0.1422; no correction gives 0.0955, the persistence route gives 0.241470% long — and the decay route gives 0.1287. A bootstrap that was never told what it was correcting has stayed within a tenth of the truth while a formula derived for the parameter has missed by most of the quantity.

And the forecast is still worse

Squared forecast error is what decides, and it says something the bias table does not.

Squared forecast error under each route, φ = 0.85, 6 steps ahead. The truth is φ^6 = 0.3771. no correction averages 0.2616 with a spread of 0.1646 and a squared forecast error of 3.2516; the formula, on the persistence averages 0.4213 with a spread of 0.2528 and a squared forecast error of 3.4827; the bootstrap, on the persistence averages 0.4355 with a spread of 0.2655 and a squared forecast error of 3.5120; the bootstrap, on the decay factor averages 0.3375 with a spread of 0.2278 and a squared forecast error of 3.4132. 800 series, 100 bootstrap refits each.
Fig. 4 The same four routes, scored on the thing a forecast is for. No correction costs 3.2516; the closed form on the persistence 3.4827; the bootstrap on the persistence 3.5120; the bootstrap on the decay factor 3.4132. The route with the smallest bias is the best of the three corrections and is still 5% worse than making none.

Every correction loses to doing nothing, and the one aimed at the right quantity loses by least.

The reason is in the third column of the first figure, which is the spread. The uncorrected decay factor has a standard deviation of 0.1646 across series; the decay-corrected one has 0.2278 — 38% larger. Squared error is bias squared plus variance, the correction removes about 0.08 of bias and adds about 0.06 of standard deviation, and at these magnitudes the second wins.

So the estimand was never the whole problem. The finding that the correction is aimed at the wrong quantity — is true, correctable and not the reason the correction loses. What loses is the variance every bias correction adds, and correcting a quantity that is already a power adds the most of it, because a bootstrap sample whose persistence lands near one produces a decay factor that is very large and the average of those is very noisy.

What each route costs the forecast, against the horizon. Three routes at five horizons, 500 series each. As a share of the truth, the uncorrected decay factor is off by -9.3% at one step and -35.4% at twelve; the persistence-corrected route by 63.5% at twelve and the decay-corrected route by -14.2%.
Fig. 5 The same comparison across horizons, on squared forecast error. The three curves stay in the same order at every horizon and the gaps between them are small next to the level, which is the shape of a quantity dominated by shocks that have not happened yet rather than by anything about the estimate.

What the bootstrap is doing that the formula cannot

The twelve-step figure is where the two devices separate most, and the reason is worth one paragraph because it is the whole case for the more expensive one.

Kendall’s formula is an expansion of the bias of φ^\hat\varphi, derived once, valid to order 1/n1/n, and evaluated at the estimate. There is no corresponding expansion for φ^h\hat\varphi^{h} that anybody uses; the device reached for instead is the delta method, and measured elsewhere returning a decay factor of −0.0003 at twelve steps — a linear approximation to a convex function, evaluated over a spread half a standard deviation wide.

The bootstrap needs neither. It does not expand anything, it does not linearise anything, and it does not know what function it is correcting. It simulates from the fitted model, applies the estimator — whatever the estimator is — and measures what comes back. A device that measures rather than derives has no order of approximation to run out of, which is why its relative error at twelve steps is 9.5% where a derived correction’s is 70%.

That is the property worth carrying out of this field, and it is not a property about forecasting. Wherever an estimand is a nonlinear function of a parameter and the bias of the parameter is known, correcting the parameter and transforming is the available move and it is not the right one; the bootstrap on the estimand is one line and is.

Why the forecast is so hard to improve

The last figure says why the whole exercise has such a small ceiling, and it is worth stating plainly because it bounds everything in this field.

An hh-step forecast error is j<hψjεn+hj\sum_{j<h}\psi_{j}\varepsilon_{n+h-j} — the shocks that have not happened yet — plus a term from the estimated parameters. At φ=0.85\varphi = 0.85 and six steps the first part has variance 3.0910, against the best route’s total squared error of 3.2516. Ninety-five per cent of a forecast’s squared error is unavoidable, and every method in this essay is competing over the other five.

That is why the four routes span 3.25 to 3.51 — eight per cent — while their decay factors span 0.26 to 0.44, a factor of 1.7. The quantity being estimated differs a great deal between methods and the quantity being scored barely differs at all. A five per cent budget is not a budget a bias correction can win from, because the correction has to pay for its own noise out of it.

It also says what a bias correction would have to be to win. It would have to remove bias without adding variance, and the bootstrap cannot: it estimates the bias from the same data, so the estimate of the bias is itself noisy, and that noise goes straight into the corrected estimate. A correction whose bias estimate came from somewhere else — a long history, a hierarchy of similar series — would not pay that price, and nothing in this field has one.

Where the correction does win, and it is the other one

The one setting where correcting helps the forecast is the one the repair that moves the wrong number found: high persistence on a short series. Running the four routes there gives the result this essay did not expect.

At φ=0.95\varphi = 0.95 on twenty-five observations the uncorrected decay factor averages 0.3075 against a truth of 0.7351 — not a bias so much as a different forecast — and costs 6.536 in squared error. Correcting the persistence, by either route, brings the decay factor to about 0.59 and the squared error to 6.22, which is a gain of 4.8%. Correcting the decay factor brings it to 0.4489 and the squared error to 6.861, which is a loss of 5.0%.

So at the one setting where bias correction is worth doing, the route aimed at the right estimand is the route that fails. It gets closer on bias than doing nothing and further than the parameter route, and it carries the largest spread of the four — 0.4035 against the parameter route’s 0.3478 — for the reason above: it is averaging powers of bootstrap persistences that sit near one.

The condition for correcting at all is about the ratio of bias to spread, and it is computable from the fitted model before any correction is applied. At φ=0.85\varphi = 0.85, n=50n = 50, h=6h = 6 the uncorrected decay factor’s bias is 0.70 of its own spread, and correcting loses. At φ=0.95\varphi = 0.95, n=25n = 25 it is 1.59, and correcting wins. That rule an analyst can apply.

The condition for which correction is a second question and the answer is the plain one: use the cheapest correction that reduces the bias, because past a point every extra reduction in bias is bought with more variance than it removes. That is the closed form on the persistence, at one multiplication.

Squared forecast error under each route, φ = 0.95, 6 steps ahead. The truth is φ^6 = 0.7351. no correction averages 0.3075 with a spread of 0.2697 and a squared forecast error of 6.5360; the formula, on the persistence averages 0.5808 with a spread of 0.3478 and a squared forecast error of 6.2220; the bootstrap, on the persistence averages 0.5985 with a spread of 0.3507 and a squared forecast error of 6.2189; the bootstrap, on the decay factor averages 0.4489 with a spread of 0.4035 and a squared forecast error of 6.8609. 800 series, 100 bootstrap refits each.
Fig. 6 The setting where correcting is worth doing, scored on the forecast. The two persistence routes are the only ones below the no-correction bar; the route aimed at the decay factor is above it. Aiming better and scoring worse is what a variance-dominated comparison looks like.
The correction does not arrive at the truth, it passes it. The average decay factor a forecast applies to the last observation, at φ = 0.95 and 25 observations, 3000 series per horizon. The middle curve is φʰ, what the model actually does. Below it is the uncorrected forecast, which uses φ̂ʰ and reverts too fast — 47.2% short at h = 4, 56.0% short at h = 6, 61.4% short at h = 8. Above it is the forecast built on the corrected estimate, which overshoots, and the reason is arithmetic rather than a bad correction: raising an unbiased estimate to a power does not give an unbiased estimate of the power, and the higher the power the more the spread of φ̂ is converted into overshoot.
Fig. 7 The same setting drawn in decay factors across the horizon: the uncorrected curve is nowhere near the truth at any horizon past two, and the corrected one — still short — is closer everywhere. This is what “the bias dominates the variance” looks like, and it is the only corner of the field where it does.

Three routes, three different failures

Putting the four routes beside each other at both settings gives a table that is more useful than either setting alone, and it says the three corrections fail in three different ways.

φ=0.85\varphi = 0.85, n=50n = 50 φ=0.95\varphi = 0.95, n=25n = 25
truth 0.3771 0.7351
no correction 0.2616, cost 3.2516 0.3075, cost 6.5360
formula, on the persistence 0.4213, cost 3.4827 0.5808, cost 6.2220
bootstrap, on the persistence 0.4355, cost 3.5120 0.5985, cost 6.2189
bootstrap, on the decay factor 0.3375, cost 3.4132 0.4489, cost 6.8609

No correction is short at both settings and by a great deal at the second — it is the failure the field started from.

The persistence corrections overshoot at the first setting and undershoot at the second, which is the same behaviour seen from two sides: they add a fixed amount to the persistence, and how that lands after the power depends on where the persistence was.

The decay correction is short at both, by less than no correction and more than the persistence routes at the second. It is the only route whose sign of error does not change, which is what aiming at the estimand buys, and it costs the most variance to get it.

There is no row that is best in both columns and no row that is best on both criteria in either column. That is not a failure of the comparison; it is what a bias–variance trade looks like when the trade is close, and it is the reason this field reports settings rather than verdicts.

What the bootstrap costs, and what it buys over the formula

The bootstrap route costs a hundred refits per series where the closed form costs one multiplication, and the comparison between them is worth making because the extra cost buys almost nothing here and buys something real elsewhere.

On the persistence, the closed form and the bootstrap give 0.4213 and 0.4355 at six steps — within one and a half per cent of each other, which is what a correct closed form and a simulation of the same thing should do. The two routes share no arithmetic: one is Kendall’s expansion evaluated at the estimate, the other is two hundred refits of simulated data. Agreeing is evidence that the formula is the bias rather than an approximation to something else, which is the site’s standing method applied to an estimator rather than to a number.

What the bootstrap buys is that it needs no formula at all. There is no closed-form Kendall bias for φ^h\hat\varphi^{h} — the delta-method version of it is the one already measured failing, giving a decay factor of −0.0003 at twelve steps — and the bootstrap produces the correction anyway, because it never needed to know what it was correcting.

That generality is the case for it, and this essay is the case against expecting much from it. A device that can correct any estimand still cannot correct an estimand whose bias is smaller than its own noise, and which estimand is being corrected turns out not to be the binding constraint.

What this says about bias correction generally

One repair has now been taken apart three times over, and the general statement is worth separating from the autoregression it was demonstrated on.

A bias correction is a trade, and the trade has three terms rather than two. It removes a bias, it adds the variance of the estimate of that bias, and — where the quantity of interest is a function of the corrected parameter — it changes what the function does to both. The first term is the one everybody computes, the second is the one that decides, and the third is what the repair that moves the wrong number was about.

The second term is where the useful rule lives, and it is not specific to autoregressions: correcting is worth doing when the bias is large compared with the estimator’s own spread, because that ratio is what decides whether the squared-error trade comes out positive. Under one, correcting loses; well above one, it wins. It is the same arithmetic that decides whether shrinking a group towards a population is worth doing, arrived at in a field with no hierarchy in it, and the ratio plays the role the shrinkage weight plays there.

The third term is why “correct the parameter and then transform” is not a safe default, and why the answer this essay gives is narrower than “aim at the right thing”. Aiming at the right thing is correct, it is cheap with a bootstrap, and it is not what decides.

What is claimed here, and what is not

The claim is what happens when a bias correction is aimed at the quantity a forecast actually uses: that the bootstrap applied to the decay factor keeps its relative error between 0.8% and 14.2% across horizons from one to twelve while the persistence-corrected route reaches 63.5%; that it is nevertheless beaten on squared forecast error by making no correction at all, at 3.4132 against 3.2516; and that the reason is the variance every bootstrap correction adds — 0.2278 of spread against 0.1646.

Every number is eight hundred series with a hundred bootstrap refits each, or five hundred with sixty for the horizon sweep. Those counts are stated because a bootstrap inside a simulation is the most expensive thing in this library and the counts are a compromise rather than a limit.

What stays out: bias correction of the whole forecast path rather than of one horizon’s decay factor, which raises the question of whether the corrected path is still a path any model produces; the bootstrap-after-bootstrap correction, which corrects the bias of the bias estimate and costs B2B^{2} refits; and shrinkage estimators for the persistence, which trade bias for variance deliberately rather than trying to remove the first without paying the second and are the family this essay’s ending points at without measuring.

Still open: the estimate that leaves the region

Every corrected estimate in this essay has been quietly capped. The closed-form correction adds (1+3φ^)/n(1+3\hat\varphi)/n whatever φ^\hat\varphi is, so near the top of the stationary range it produces a persistence above one — which is not a persistence, and for which the forecast diverges and the variance formula returns a negative number.

It is not a rare event. At φ=0.95\varphi = 0.95 on twenty-five observations it happens on 31.1% of series, and the threshold has a closed form that needs no simulation. Every implementation does something about it, none of them documents what, and the five obvious treatments differ by a factor of 2.3 in squared forecast error. That is what happens when the correction leaves the region.

The check, and the refusal

Three claims are gated, and two of them had to be weakened after being measured. That the uncorrected decay factor is short at every horizon, which holds. That some correction beats no correction on the bias, which holds. And that the persistence-corrected route overshoots — required only at four steps and beyond, because at two it is still 1.3% short and demanding it there would have been a claim about one setting stated as a claim about the family.

The refusal is the one that makes the comparison mean anything: the route aimed at the decay factor must stay within a fifth of it at every horizon past four, while the persistence route is required to have the larger worst relative error across the sweep. If both routes behaved the same way, the distinction between correcting a parameter and correcting a function of it would be a distinction without a measurement, and the diagnosis it rests on would be unfalsifiable rather than merely incomplete.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationBias correctionBootstrapDelta methodForecast errorForecast horizonMean squared errorMonte CarloPlug in estimateStationarity