The repair that moves the wrong number
Worth reading first: Correcting the persistence · What the model says next.
The previous essay ends with a correction that works. The bias in the least-squares estimate of a persistence parameter is removed by adding (1 + 3φ̂)/n to it; at φ = 0.8 and fifty observations that takes the bias from −0.0697 to −0.0059 and the root mean squared error from 0.1246 to 0.1096. Every quantity a reader would check has improved.
The forecast field’s complaint was never about φ̂. It was that the forecast reverts to its mean too quickly, and the forecast is φ̂ raised to the power of the horizon. So the obvious next step is to put the corrected estimate into the forecast and count what happens, and the counting produces three answers in three directions.
Raising an unbiased estimate to a power does not give an unbiased power
The first thing to go wrong is arithmetic that has nothing to do with autoregressions.
An h-step forecast multiplies the deviation from the mean by φ̂ʰ. That is a convex function of φ̂ for h ≥ 2, so the average of φ̂ʰ over samples exceeds the h-th power of the average of φ̂. Correcting φ̂ so that it averages the truth therefore leaves φ̂ʰ averaging more than the truth’s h-th power, and the higher the horizon the more of the estimate’s spread is converted into overshoot.
At six steps ahead the three numbers are 0.2624 for the plug-in forecast, 0.3771 for the truth, and 0.4229 for the corrected one. The plug-in forecast is 30% short; the corrected one is 12% long. At twelve steps the plug-in is at 0.0948 against a truth of 0.1422, and the corrected one at 0.2400 — which is now 69% too large.
The correction has not aimed at the target and missed. It has aimed at a different target, hit it, and moved the quantity that was actually being complained about past where it should be.
The delta method gets the direction and not the size
The usual way to carry a bias through a function is the delta method: multiply it by the derivative, here hφ^(h−1). At φ = 0.85, n = 50 and h = 6 that predicts a decay factor of 0.1881 against a true φ⁶ of 0.3771 — a shortfall of half, where the counted shortfall is under a third.
The prediction is wrong in the same way and for the same reason that the correction overshoots: a linear approximation to a convex function evaluated over a spread of half a standard deviation is not accurate, and at φ = 0.85, n = 50, h = 12 the delta method predicts a decay factor of −0.0003, which is not a decay factor at all.
This is worth a sentence rather than a paragraph in most treatments and it earns more here, because the delta method is how the shortfall would normally be quantified and it is the reason a reader might expect the correction to work: if the bias in the forecast were hφ^(h−1) times the bias in φ̂, removing the second would remove the first. It is not, and it does not.
Scored on the forecast, at moderate persistence the repair loses
None of the above is decisive on its own. A forecast whose decay factor is 12% high can still be a better forecast than one whose decay factor is 30% low, because squared error is not a function of the decay factor alone. The comparison has to be run.
At φ = 0.85 and n = 50, six steps ahead, mean squared forecast error rises from 3.4923 with no correction to 3.6700 with all of it. Twelve steps ahead it rises by 9.1%. One step ahead it is flat. The forecast is not improved anywhere in that setting, and it is measurably damaged at the horizons the correction was supposed to be for.
The mechanism is the one above with squared error as the arbiter. The plug-in forecast’s error is biased towards the mean and the corrected forecast’s is biased away from it, and because the h-step forecast error is dominated by the shocks that have not happened yet, the extra variance the correction introduces costs more than the bias it removes buys.
Where it wins, and the boundary is a measurement
At high persistence on a short series the accounting comes out the other way.
The plug-in decay factor at that setting is 0.3214 against a truth of 0.7351, which is not a bias so much as a different forecast: a series that remembers almost everything is being forecast as one that remembers a third. Correcting takes the average decay to 0.6045 — still short — and the squared error falls by 7.8% at six steps and 10.2% at twelve.
So the boundary between the settings where correcting helps a forecast and the settings where it hurts is neither at a value of φ nor at a value of n; it is where the bias in φ̂ʰ is large relative to the shock variance the forecast cannot do anything about, and that is a function of both, plus the horizon. What can be said is what was counted: at φ = 0.85 on fifty observations it costs 5.2%, at φ = 0.95 on twenty-five it buys 7.8%, and at φ = 0.7 on a hundred it does nothing.
One horizon at a time, because the answer moves with it
Reading the sign of the gain off one setting is what makes this look like a paradox, so it is worth laying the horizon out. At φ = 0.85 and fifty observations the correction changes squared forecast error by −0.1% at one step, +5.2% at six and +9.1% at twelve. At φ = 0.95 on twenty-five it changes it by −6.1%, −7.8% and −10.2% at the same three horizons.
Both patterns grow with the horizon and they grow in opposite directions, which is exactly what the decay picture implies: the further ahead the forecast, the more the estimate is amplified, and amplification helps when the plug-in forecast is badly wrong and hurts when it is nearly right.
The interval improves, for the wrong reason
The correction does one thing unambiguously well, and the reason is unedifying.
The plug-in forecast interval undercovers: at φ = 0.85 and n = 50, six steps ahead, it covers 88.2% where it claims 95%. Feed it the corrected estimate and it covers 91.3%. At φ = 0.95 and n = 25 it goes from 76.6% to 84.4%, which is the largest improvement in coverage anywhere in this field.
None of that is because the interval is better centred. It is because the interval is wider: the plug-in standard error is σ̂√(Σψ̂²), the ψ weights are powers of φ̂, and a larger φ̂ makes every one of them larger. The band at φ = 0.85, n = 50, h = 6 goes from 6.185 wide to 6.978 — thirteen percent — and thirteen percent of extra width buys three points of coverage.
That is worth stating plainly because it is a general trap rather than a local one. Any change that widens an interval improves its coverage, and coverage is therefore not evidence that the change was the right one. The interval’s shortfall is a variance problem — the plug-in formula treats estimated parameters as known — and a bias correction is not aimed at it. The correction hits the coverage figure by accident, in the right direction, for a reason that would work just as well if the multiplier had been chosen at random.
What a forecaster should take from this
The correction is not useless and it is not a repair of the thing that was broken, and both halves of that sentence have to survive into practice.
A parameter that is going to be reported should be corrected. It is the quantity the correction was derived for, the improvement is real at every persistence, the cost is six percent of a standard error, and an uncorrected estimate of persistence is a number that is wrong in a known direction by a known amount.
A forecast that is going to be made should be checked rather than corrected on principle. At the settings measured here the correction cost squared error twice and bought it once, and the winning case — high persistence, short series — is the one where a forecaster has the least evidence about anything.
An interval that is going to be quoted should not be repaired by this at all. Its shortfall is the plug-in formula ignoring parameter uncertainty, the correction happens to widen the band, and a repair that works through a side effect will stop working the moment the side effect points elsewhere.
Why the coverage gain is not worth taking anyway
There is one more reason to refuse the widening as a repair, and it is available without any new measurement: the correction cannot be tuned to the interval. The amount it widens the band by is fixed by the bias formula, which was derived for a different quantity, so the coverage it delivers is whatever falls out. At φ = 0.85 and n = 50 that is 91.3% — better than 88.2% and still not 95%. At φ = 0.95 and n = 25 it is 84.4%, which is a large improvement on 76.6% and a long way from what the interval says on the page.
A repair that closes half the gap by accident leaves a reader in the worst position available: the interval is closer to its claim than before, nothing on the output says by how much, and the residual shortfall is now harder to notice. The honest alternative is to widen the interval on purpose by the amount the shortfall actually is, which the forecast field measures directly — at six steps on twenty-five observations the honest interval is 25.7% wider than the plug-in one — and which needs no bias correction at all.
The quantity nobody names, which the correction detonates
There is a second estimate inside every one of these forecasts and it is never discussed. A forecast reverts to a mean, and the mean is estimated. A fitted autoregression with an intercept implies a long-run mean of ĉ/(1 − φ̂) — a ratio, whose denominator is exactly what the correction pushes towards zero.
At φ = 0.95 and n = 50, forecasts built that way have a mean squared error 2,712 times the same forecasts built from the sample mean of the same window, and on one series the implied mean reached 220,890 in units where the series itself has a standard deviation near three. The sample mean of the window estimates the same quantity, is consistent under the same conditions, cannot blow up, and differs from the ratio by nothing anybody would report.
This is the field’s refusal, and it is a refusal about a repair breaking something it was not aimed at: correcting one parameter upwards makes a quantity that depends on it through a denominator enormously worse, and no bias-variance accounting of φ̂ would ever have shown it.
The average decay factor is not what the error is made of
At six steps, φ = 0.85, fifty observations, the corrected forecast’s average decay factor is nearer the truth than the plug-in’s. The truth is 0.3771; the plug-in averages 0.2624, which is 0.1147 below it, and the corrected one averages 0.4229, which is 0.0458 above. By the only summary the decay picture shows, the correction has cut the error to two-fifths of what it was — and squared forecast error at exactly that setting rose by 5.2%.
Both readings are correct, and the contradiction between them is the useful part. An average decay factor is one number summarising a distribution of decay factors, and squared error is a functional of the whole distribution rather than of its centre. Correcting φ̂ adds 3/n times a random quantity to a random quantity and then raises the sum to the sixth power, which amplifies the added spread by more than it moves the centre. The centre moved towards the truth. The spread moved away from it. Squared error counts the second and the picture shows the first.
That is the same trade the correction was scored on in the parameter itself — bias down, variance up — carried through a convex function. What the function does to the two terms is not the same thing: to leading order it multiplies the bias by hφ^(h−1) and the variance by the square of that factor. The horizon therefore enters the two halves at different powers, so a trade that was worth taking on a parameter cannot be assumed to survive to any function of it, and the further from linear the function is, the less of it survives. A bias correction is a claim about one estimand and carries no guarantee about any other, which is the general form of what the comparison of forecasts has to keep separate.
The horizon does not amplify the two errors equally
Reading the same four numbers as proportions rather than as differences separates the two forecasts more sharply than any of the pictures do.
The plug-in forecast’s decay factor is 30.4% short of the truth at six steps and 33.3% short at twelve: near enough flat in the horizon. The corrected forecast’s is 12.1% long at six and 68.8% long at twelve — nearly six times as large. Proportionally, the plug-in forecast makes about the same mistake at every horizon, and the corrected one makes a mistake that compounds.
The asymmetry is in the arithmetic rather than in the estimator. An estimate that sits a constant fraction below the truth stays that same fraction below it when raised to any power, because (aφ)ʰ = aʰφʰ scales the whole curve rather than bending it. The downward bias in φ̂ is close enough to proportional over this range that the plug-in shortfall reads as a near-constant ratio. The correction, on the other hand, is additive: it adds (1 + 3φ̂)/n whatever φ̂ is, and near the top of the stationary range an additive offset is a large proportional one, which the horizon then raises to a power.
So the horizon is the variable that decides, and it decides against the correction rather than for it. The complaint that started this was about long horizons — a forecast reverting to its mean too fast is a statement about h and about nothing else — and the repair aimed at that complaint is the one whose own error grows fastest in h. A reader making a one-step forecast has neither much to gain nor much to lose. A reader at twelve steps has the most of both available, and at moderate persistence has the most of the wrong one.
What is claimed, and what is not
The claim is what a bias-corrected persistence parameter does to the forecast built on it: the decay factor overshooting rather than arriving, the delta method failing to size the effect, squared error rising at moderate persistence and falling at high, the interval improving through its width, and the reversion target’s fragility. Every number is counted on the same series with the same seeds, so no comparison here is between studies.
What stays out: bootstrap bias correction of the forecast rather than of the parameter, which targets the right quantity directly and is a different construction; correcting for the interval’s variance shortfall, which is what the forecast field’s own measurement calls for and which needs a term nobody has computed here; and the whole question at horizons long enough that the forecast is the mean, where none of it matters because φʰ is zero either way.
The checks, and the refusal
Three claims are gated. The uncorrected decay factor must be below the truth and the corrected one above it, at the same setting — which is the overshoot asserted as an inequality rather than read off a picture. Squared forecast error must rise under the correction at moderate persistence and fall at high persistence on a short series, both measured with the same code and the same seeds, because a claim that the sign of a gain depends on where it is measured is only worth making if both signs are demonstrated. And the interval’s coverage must rise while remaining short of what it claims, which is the accidental improvement stated in a form that cannot be mistaken for a repair.
The refusal is the implied reversion target ĉ/(1 − φ̂), against the standard that two consistent estimators of the same mean should not change a forecast’s squared error by an order of magnitude. It fails by a factor of 2,712, and the check requires it to fail.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The interval after the choice — both name coverage, forecast interval, monte carlo, parameter uncertainty
- A charge that reads the draw — both name mean squared error, monte carlo, plug in estimate
- A flat point with more than one direction — both name coverage, delta method, monte carlo
- A line that beats two curves — both name least squares, monte carlo, parameter uncertainty
- Choosing the order — both name forecast error, least squares, monte carlo
- Coverage from exchangeability alone — both name coverage, forecast interval, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Bias correctionCoverageDelta methodForecast errorForecast horizonForecast intervalLeast squaresMean squared errorMonte CarloParameter uncertaintyPlug in estimateStationarity