The observation that has not happened

The interval that forgets it estimated

The forecast band is derived for a model whose parameters are known, and then computed by putting estimates into it. Counted, the 95% interval covers 87.3% six steps ahead on twenty-five observations, and the point forecast inside it returns to the mean a third faster than the series does.

Worth reading first: The observations that repeat each other · What the 95% refers to.

The previous essay derived the forecast interval and checked it. The derivation assumed φ and σ were known. The check ran the simulation with φ and σ known. Both were honest and neither was about anything anybody can do, because a forecaster has a series and not a model.

What happens next is the substitution that every implementation of this makes and that no output mentions. Fit the model, get φ̂ and σ̂, put them where φ and σ were, and quote the interval. The formula is correct. The numbers going into it are not the numbers it was derived for, and the question this essay asks is what that costs, in the only currency the site accepts: how often the interval contains what it said it would.

What a 95% forecast interval covers, counted. 1200 series of 25 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.3% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 92.8% at one step and 87.3% at 6. The interval that would cover what it claims is 6.9% wider at one step.
Fig. 1 Counted coverage of a 95% forecast interval at every horizon, on twenty-five observations. The upper line is the interval a forecaster who knew the parameters would have used. The lower line is the same formula fed estimates. Nothing separates them but the substitution.

What it costs, counted

Twelve hundred series at each horizon, one set of seeds, both intervals computed on the same data. The known-parameter interval sits on 95.0% at six steps and its counterpart on 87.3%.

The shortfall has two properties worth separating, because they have different causes and different remedies.

It grows with the horizon. At one step the plug-in interval covers 92.8% and at six it covers 87.3%. That is not a small drift: the miss rate has gone from one in fourteen to one in eight, on the same data, from the same fit, with nothing changing except how far ahead the statement reaches.

It shrinks with the length of the series. At a hundred observations the one-step interval covers 94.3% and at four hundred 94.4%, both within a standard error of the target. This is estimation error and behaves like estimation error: it is a fact about how much data went into φ̂, and it goes away in the direction more data goes.

What a 95% forecast interval covers, counted. 1200 series of 100 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 94.8% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 94.3% at one step and 93.9% at 6. The interval that would cover what it claims is 0.3% wider at one step.
Fig. 2 The same measurement on a hundred observations. The two lines have nearly closed, which is what “estimation error” means when it is counted rather than described. Drag the length of the series to watch the gap open and shut.
What a 95% forecast interval covers, counted1200 series of 50 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 94.5% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 92.7% at one step and 90.6% at 6. The interval that would cover what it claims is 5.5% wider at one step.0.9000.950123456horizon h, in steps aheadcounted coverage of a 95% interval95%, which all of them claimupper: parameters known · lower: parameters estimated1200 series of 50 at φ = 0.7, counted at each horizon90.6% at h = 6 for an interval claiming 95%
Fig. 3 And at fifty, which is a realistic length for a series somebody actually has. Six steps ahead the interval covers 90.6% where it claims 95%.

Is that a gap or a wobble

Two numbers a few percentage points apart deserve the question this site asks of every counted rate before reading anything into it. Twelve hundred trials at a nominal 5% miss rate gives a standard error of 0.63 percentage points on each line. The gap at six steps is 7.7 points, which is a dozen standard errors, and the gap at one step is 1.9 points, which is three.

The comparison is also within seeds: the same twelve hundred series produce both lines, so the two are not two independent estimates being differenced. Every series contributes a known-parameter verdict and a plug-in verdict, and the pairs disagree on the series where the fit happened to be poor. That is why the upper line can be read as a control rather than as a second measurement — if the machinery had a bug in the forecast recursion, or in the ψ weights, or in the way the future observations are held out, both lines would move and the gap would not be attributable.

The upper line does one more job, and it is the reason it is drawn at all rather than assumed to be 95%: it is a live check on the arithmetic of the whole essay. If it drifted off the target the conclusion below would be unavailable, because the shortfall could be a defect in the simulation rather than in the substitution. It sits within a standard error of 95% at every horizon and every sample size in this essay, which is what makes the lower line’s departure mean something.

How much wider it would have to be

Coverage is the honest measure and it is not the useful one, because a forecaster cannot act on “87%”. What can be acted on is the width, and the width the interval should have is available without any theory at all: simulate, collect the forecast errors, and read off the interval that contains 95% of them.

That number is not a derivation and does not pretend to be. It is an empirical quantile of the error distribution, and comparing it with the mean plug-in width says how much the interval was short by. At twenty-five observations the honest interval is 25.7% wider at six steps and 6.9% wider at one.

A quarter more width is not a rounding error. It is the difference between a band a reader treats as tight and a band a reader treats as barely informative, and the whole of it is missing because a formula derived for two known numbers was handed two estimates.

There is a general habit here that this site has met before under a different name. The plug-in interval for a hierarchical model estimates the population spread and then uses it as though it were known, and covers 78.8% against a claimed 95%. That is the same mistake in a different field, and the two are worth putting side by side because the shape of the repair is the same: carry the uncertainty in the estimated quantity instead of dropping it. There the repair was an integration over the posterior for τ. Here the comparable repair is a bootstrap over the fitted residuals, which this field measures rather than implements — the point of this essay is the size of the hole, and the honest interval above is the measurement of it.

What a 95% forecast interval covers, counted. 1200 series of 400 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 94.9% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 94.4% at one step and 95.8% at 6. The interval that would cover what it claims is 2.3% wider at one step.
Fig. 4 Four hundred observations, where the two lines have closed to within a standard error at every horizon. This is the picture the formula was derived for, and it is the one nobody has.

The other defect, which is in the forecast rather than the interval

The interval is centred on a point forecast, and the point forecast has its own problem. It is a separate defect in the same arithmetic and it is easier to state than to notice.

Least squares estimates the persistence of a series as smaller than it is. The bias is about (1 + 3φ)/n and it is downward at every φ. On forty observations from a series with φ = 0.8, φ̂ averages 0.7074.

By itself that is a bias of a tenth, which is unremarkable. What makes it worth an essay is that the h-step forecast does not use φ̂ — it uses φ̂ raised to the power h, and a proportional error raised to a power is a larger proportional error. The decay factor the forecast applies is

  • 11.6% short of the truth at one step,
  • 24.8% short at three,
  • 30.9% short at five,
  • and 32.8% short at eight.

So the forecast returns to the mean a third faster than the series does, systematically, in every run. This is not visible in any single forecast — a third of a decay factor is well inside the noise of one series — and it is not visible in the coverage either, because the interval is wide enough to absorb it. It is visible only across many series, which is exactly the arrangement the twenty fans show.

Twenty forecasts of twenty series from one model. 20 series of 40 observations from an AR(1) with φ = 0.8, each fitted and each forecast 12 steps. Every path decays towards its own fitted mean at its own φ̂, and the spread of the paths is the quantity the single-forecast picture has no way to show. The fitted φ̂ averages 0.709 against a truth of 0.8 — least squares estimates persistence as smaller than it is, the h-step forecast raises that estimate to the power h, and the paths therefore return to the mean faster than the series they came from.
Fig. 5 Twenty forecasts of twenty series from one model at φ = 0.8, fitted on forty observations each. Each path decays at its own φ̂ and the average of those φ̂ is below the truth, so the swarm as a whole returns to the mean faster than the process does. The bias is in the middle of the picture rather than at its edges.

The practical reading is uncomfortable and worth stating without softening. A forecaster using a short series will, on average, under-predict how long a departure from the mean persists. Not sometimes — on average, in the direction the estimator is biased, at every horizon, and by more the further ahead the question reaches.

The two defects pull in the same direction, which is why neither shows

It is worth putting the two together, because they are not independent and their interaction is what makes the whole thing survivable in practice.

The interval is too narrow: it should be a quarter wider at six steps on a short series. The point forecast is too close to the mean: its decay factor is a third short at five steps. A forecaster therefore publishes a band that is centred too near the average and is too tight around that centre — two errors, both of which make the published statement look more informative than it is.

They do not cancel. A narrow band around a wrong centre misses more often than a narrow band around the right one, and the counted coverage above already contains both, because the same fitted φ̂ that produced the shrunken forecast produced the band. Separating them is possible — put the true parameters in the band and keep the fitted forecast, or the reverse — and the arrangement matters enough that the field’s library computes the fully-known interval rather than the half-known one. The first version of that function used the true variance around the fitted forecast, which is neither of the two things anybody has, and reported the known-parameter line at 93.2% instead of 95.0%. Two per cent of the gap belonged to a hybrid nobody would ever use.

The lesson is one this site keeps meeting from different directions: a control has to be a procedure somebody could run. The known-parameter interval is a control because a forecaster who happened to be told φ and σ could use it. A band computed from the true variance around a fitted forecast is not a procedure at all, and using it as a reference makes the defect look smaller than it is.

One forecast, and the band the arithmetic puts round it. An AR(1) with φ = 0.7, 25 observations, fitted by least squares and forecast 10 steps ahead. The point forecast decays towards the fitted mean at φ̂^h; the band is ±1.96 standard errors from σ̂²Σψ̂², which grows with the horizon and stops at the unconditional spread 1.31. The dashed pair is the same band computed at the true parameters, which nobody has. The marks past zero are what actually arrived: 10 of 10 inside the band this once, which is one draw and settles nothing.
Fig. 6 One forecast from twenty-five observations, which is the setting the numbers above are measured at. The band looks no different from the one drawn on sixty, and it is missing a quarter of its width at the far end.

The shortfall is width and not centre

Two defects are measured — a band that is too narrow and a centre that is too close to the mean — and the coverage figures say how the seven and a half points are split between them.

An interval that should be 6.9% wider covers 2Φ(1.96/1.069) − 1 = 93.3%, and the counted one-step figure is 92.8%. One that should be 25.7% wider covers 88.1%, and the counted six-step figure is 87.3%.

So the width accounts for the coverage shortfall to within about half a point at both horizons, and the biased centre contributes under a point of the seven and a half. That is worth knowing before the two defects are read as equally urgent: the interval misses because it is narrow, and the forecast inside it is wrong for a reason that barely shows up in whether the interval contains the future.

Which is the direction that makes the bias hardest to find, and it is a stronger statement than the essay’s own. The bias is not merely absent from the accuracy measure a forecaster computes; it is nearly absent from the coverage measure as well.

The convexity that softens it is the one that ruins the repair

The decay shortfalls — 11.6%, 24.8%, 30.9% and 32.8% at one, three, five and eight steps — are smaller than the bias in φ̂ alone predicts, and the gap is worth reading.

φ̂ averages 0.7074 against a truth of 0.8, a ratio of 0.884. Raising that ratio to the horizon gives shortfalls of 11.6%, 30.9%, 46.0% and 70.9%. The counted figures match at one step exactly and fall progressively below the prediction after it: the shortfall is 20%, 33% and 54% less than raising the mean would give.

That gap is Jensen’s inequality. E[φ̂ʰ] exceeds (E[φ̂])ʰ because a power is convex, so the spread of φ̂ across samples pushes the average decay factor back up — and it pushes harder the higher the power. It is exactly the mechanism that makes a corrected estimate overshoot when it is raised to a horizon, arriving here to partially rescue an uncorrected one.

So the two halves of that arithmetic are the same fact used twice. The convexity that makes a bias-corrected forecast overshoot is what stops the uncorrected forecast’s shortfall from compounding, and it is why the counted sequence flattens at about a third rather than heading for the 71% the bias alone would imply.

Why the direction is the one that flatters

There is a reason this defect has survived so comfortably, and it is the direction.

A forecast that reverts too quickly is a forecast that is closer to the mean than it should be. A forecast closer to the mean is a less extreme forecast, and less extreme forecasts are, by construction, harder to be badly wrong about. Measured by squared error against the actual future, the shrunken forecast is barely worse than the correct one and is sometimes better — the same arithmetic that makes shrinkage a good idea when it is done on purpose makes it hard to detect when it is done by accident.

So the accuracy measure a forecaster is most likely to compute does not reveal the bias. Nothing reveals it except comparing φ̂ with φ, which requires knowing φ, which is the situation nobody is in. That is the general shape of every finding on this site that has survived a long time: the symptom is absent from the statistic anybody looks at.

One forecast, and the band the arithmetic puts round it. An AR(1) with φ = 0.8, 40 observations, fitted by least squares and forecast 12 steps ahead. The point forecast decays towards the fitted mean at φ̂^h; the band is ±1.96 standard errors from σ̂²Σψ̂², which grows with the horizon and stops at the unconditional spread 1.64. The dashed pair is the same band computed at the true parameters, which nobody has. The marks past zero are what actually arrived: 7 of 12 inside the band this once, which is one draw and settles nothing.
Fig. 7 One of those twenty, drawn on its own with its band. Nothing in this picture is wrong enough to notice. The realised observations are mostly inside the band, the forecast heads sensibly towards the middle, and a reader has no way to tell that it is heading there too fast.

Where in the range it is worst

Everything so far holds the persistence at φ = 0.7. The shortfall is not the same at every φ, and the shape of its dependence is the last thing worth measuring, because it says which series a forecaster should be most careful with.

Two effects run against each other. A more persistent series has a more precisely estimated φ in absolute terms — the regressor has more variance, so the standard error of φ̂ falls — which argues that the interval should be better at high φ. But a more persistent series has a forecast that depends on φ̂ over more horizons before it decays, and the h-step variance is a longer sum of terms each of which involves the estimate, which argues the other way.

One forecast, and the band the arithmetic puts round it. An AR(1) with φ = 0.5, 40 observations, fitted by least squares and forecast 12 steps ahead. The point forecast decays towards the fitted mean at φ̂^h; the band is ±1.96 standard errors from σ̂²Σψ̂², which grows with the horizon and stops at the unconditional spread 1.19. The dashed pair is the same band computed at the true parameters, which nobody has. The marks past zero are what actually arrived: 9 of 12 inside the band this once, which is one draw and settles nothing.
Fig. 8 The same forty-observation fit on a less persistent series. The band reaches its limit within about four steps and the point forecast is at the mean by six, so the horizons over which the estimate can do damage are few. The interval’s problem here is concentrated in the first two or three steps.

The second effect wins, and it wins for a reason that is arithmetic rather than statistical: the error in φ̂ enters the h-step forecast through φ̂ʰ and the h-step variance through a sum of φ̂ to even powers, so the number of places the estimate can be wrong grows with the horizon, and the horizon over which the forecast is doing anything at all grows with φ. A persistent series is the case where a model has the most to say and the case where what it says is least reliable, which is the combination least likely to prompt anybody to check.

The refusal, and it is the mistake a printout invites

A fitted model reports one number for the spread of its errors: the residual standard deviation. It is the only spread on the output, it is correct, and it is correct at one step.

Using it at every horizon is the natural thing to do and it is what an interval computed by hand from a regression printout will be. At six steps on a persistent series that interval covers 75.9% where it claims 95%, against 92.5% for the same forecast given the horizon’s own variance. The refusal in this field’s library requires the first to fail, because the check that the horizon matters is worth nothing unless it has been shown an interval that ignores the horizon and rejected it.

The two numbers are also the cleanest available statement of what the ψ weights are for. They are not a technicality of the derivation; they are the difference between an interval that misses one time in four and one that misses one time in thirteen.

The band on which fitting a model is worth doing, at n = 20. One-step squared error for three forecasts of the same next observation, on the same series and the same seeds, over 1500 series of 20 observations at each φ. The sample mean estimates no dynamics and beats the fitted model below φ = 0.166; the last value carried forward estimates nothing at all and beats it above φ = 0.818. Both crossings are solved from closed forms — σ²(1 + k/n) for the fitted model against 2γ₀(1 − φ) and γ₀(1 + (1+φ)/(n(1−φ))) — and both are functions of the length of the series alone. The band widens at both ends as n grows and never reaches either edge.
Fig. 9 And the benchmark picture at twenty observations, where the band on which fitting is worth doing is at its narrowest. Every defect measured in this essay is a defect of a model that was worth fitting in the first place, which at this length is a narrower claim than it sounds.

What is left, and it is named rather than measured

Three things this essay establishes and does not repair.

The interval is too narrow and the honest width is known. The repair is a bootstrap over the fitted residuals, or an analytic correction with an extra term of order 1/n. Neither is built here; what is built is the measurement that says how much either would have to supply.

The point forecast is biased and the bias is estimable. There are bias-corrected estimators of φ and they are not free — correcting a bias generally costs variance, and whether the corrected forecast is better by squared error is a question this field does not answer.

And both defects are conditional on the order being right. Every number above was computed with the order given: a first-order model fitted to a first-order series. No forecaster is handed the order. Choosing it is the next essay, and choosing it from the same data the interval is then computed on turns out to cost more than the estimation measured here.

That last item is the one worth carrying, because it changes how the numbers above should be read. Everything in this essay is a lower bound on what a real forecaster’s interval is short by. The model was correctly specified, its order was known, the series was genuinely stationary, and the shocks were genuinely independent and identically distributed — four assumptions that were true by construction because the series was simulated from them. A forecaster has none of those guarantees and has, in addition, the one this essay held fixed. The 87.3% is what the interval covers when everything except the parameter values is right.

Finally, a boundary. This essay measures the shortfall and does not price it, because pricing it requires saying what a miss costs, and that is a decision problem rather than a statistical one. The site’s rule holds here as everywhere: nothing is called 95% until it has been counted, and once it has been counted what to do about it is somebody else’s question.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CoverageEstimated varianceForecast errorForecast horizonForecast intervalLeast squaresMonte CarloParameter uncertaintyPlug in estimateStationarity