The observation that has not happened

The interval after the choice

Estimating the coefficients of a known model costs a 95% forecast interval about two points of coverage. Choosing which coefficients to estimate, from the same forty observations, costs another four and a half — so the step nobody records in the output is the more expensive of the two.

Worth reading first: What the model says next · When the looking happens.

Two costs have been measured in this field so far and both were charged to the same interval. The formula is right when the parameters are known. Substituting estimates for them takes a 95% interval down to about 93% on a short series. That was the previous essay, and it held one thing fixed that nobody is ever handed: the order.

The order is chosen. It is chosen by a criterion, and the criterion is computed from the same observations the interval is computed from. That makes the selection part of the procedure, and the site’s rule about procedures applies: a procedure’s advertised property is a claim about what happens when the whole of it is run, and the whole of it includes the step nobody writes down.

What the interval covers once the order is chosen as well. 1200 series of 40 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.5% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 93.6% at one step and 90.8% at 6. The third line chooses the order by AIC from the same data before computing the interval, which costs a further 0.8 points at h = 6.
Fig. 1 Three lines on one axis: the interval at the true parameters, the interval with the parameters estimated at a known order, and the interval with the order chosen by AIC from the same data. Each step down is a thing the 95% on the label does not know about.

Counted, on identical series

Forty observations from a second-order process, twenty-five hundred series, orders up to twelve offered to the criterion. Three intervals per series, all from the same data.

  • 95.0% — the interval at the true parameters and the true order.
  • 93.2% — the same interval with the coefficients estimated, the order given.
  • 90.5% — the order chosen by BIC first.
  • 88.6% — the order chosen by AIC first.

Estimation costs 1.8 points. Selection costs a further 2.7 with BIC and 4.6 with AIC. The step that appears in no output and in no formula is, on this setting, more expensive than the one every textbook derives a correction for.

The comparison is within seeds throughout, which matters more here than usual because three of the four numbers are within a few points of each other. The same series produces all four verdicts: the same shocks, the same history, the same future observation held out. What differs between the rows is only how much of the procedure was allowed to look at the data.

What the interval covers once the order is chosen as well. 1200 series of 40 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.5% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 93.6% at one step and 90.8% at 6. The third line chooses the order by BIC from the same data before computing the interval, which costs a further 1.6 points at h = 6.
Fig. 2 The same picture with BIC doing the choosing. Its third line sits above AIC’s at every horizon, which is the first thing in this field that has separated the two criteria in a way an interval can feel.

Why selection costs anything at all

The mechanism is worth stating because it is not the obvious one.

The obvious story is that a wrongly chosen order gives a wrong variance, and wrong variances sometimes give narrow intervals. That is true and it is the smaller half. The larger half is a selection effect, and it is the same effect this site has already measured on an estimate rather than an interval.

A criterion picks the order whose residual variance, penalised, is smallest. Among the orders that fit about equally well, it picks the one that happened to fit best — and “happened to fit best” means “happened to have the smallest residual variance on this particular realisation”. The selected model’s σ̂² is therefore not an unbiased estimate of σ²: it is the minimum of several, chosen because it was the minimum.

An underestimated σ̂² is a narrow interval. So the interval after selection is narrower than the interval at the true order, on average, for the same reason that the estimate from a chosen arm is inflated: the choice was made on the quantity the estimate is about.

That also explains the ordering. AIC offers a cheaper penalty and therefore selects from a wider effective set — its mean selected order here is 3.00 against BIC’s 1.53 — so it has more chances to pick a low residual variance and the selection effect is larger. The criterion that overfits more is the criterion whose interval is narrower than it should be by more, and the two facts are the same fact.

What each criterion selects, at 40 observations. 700 series from an AR(2) with coefficients 0.6 and -0.3, every order from 0 to 12 fitted to the same 28 responses so the log-likelihoods are comparable. AIC finds the true order 42.6% of the time and lands above it 34.0%; BIC finds it 43.9% and lands above it 7.0%. The closed form for one extra lag is P(χ²₁ > 2) = 15.73% for AIC, which does not depend on n at all, and P(χ²₁ > ln n) = 5.48% for BIC at this size, which falls to zero. Under the true order is the other failure and it is BIC's: 49.1% against 23.4%.
Fig. 3 What the two criteria are selecting on these forty-observation series, with twelve orders on offer. AIC’s mass is spread across the whole range and BIC’s is concentrated below the truth. The spread is what the selection effect feeds on.

The same three steps, as widths

Coverage is the natural unit for counting and it is not the unit anybody reports. Converting each row into how wrong the interval’s width is says the same thing in the quantity that appears in a table.

An interval built at ±1.96 standard errors covers c when the true spread is a factor r larger than the standard error it used, with c=2Φ(1.96/r)1c = 2\Phi(1.96/r) - 1. Inverting:

  • 93.2% with the order known: the interval is 7.4% too narrow.
  • 90.5% after BIC: 17.3% too narrow.
  • 88.6% after AIC: 24.0% too narrow.

In variances that is 1.15, 1.38 and 1.54. So an interval that has had its order chosen by AIC is pricing about two thirds of the uncertainty it has.

What would have to be done to it

The same arithmetic gives the repair, in the crude form that at least states the size of the problem.

To reach 95% coverage the multiplier would have to be 1.96 × 1.240 = 2.43 after AIC selection, 2.30 after BIC, and 2.11 with the order known. Equivalently, the interval labelled 95% after an AIC selection is the interval a 98.5% label would have to be attached to before the multiplier came out right.

That is not a proposal — the factor is a property of this setting and would have to be re-measured at every sample size, horizon and process — and it is the right way to hear the numbers. A 95% label that needs a 98.5% multiplier is not a small calibration error. It is a different interval.

Which of the three steps is the large one

Set the three costs side by side: estimation 1.8 points, selection 2.7 more with BIC and 4.6 more with AIC.

So selection is 1.5 times estimation under BIC and 2.6 times under AIC. Of the total 6.4-point shortfall in the AIC row, 72% is the step that appears in no output.

On twenty-five hundred series a coverage near 90% carries a standard error of 0.6 points unpaired, and every comparison here is within seeds, so the 1.9-point gap between the two criteria is comfortably real and so is each step.

The ordering is what makes the field worth having. The correction textbooks derive is for the smallest of the three effects, and the largest of the three is the one that is not written down anywhere in the procedure — it happens before the interval’s formula is reached, and by the time the formula runs there is nothing left in the data to say it happened.

What the selected order actually is

The mean selected order is the number that makes the mechanism concrete, and it is worth reading carefully because both criteria are wrong about it in ways that look nothing alike.

The truth is two. AIC’s mean is 3.00 and BIC’s is 1.53. So on average AIC is fitting one lag too many and BIC half a lag too few, and neither of those descriptions is what is happening on any individual series — a mean of 3.00 over orders offered up to twelve is a spread, not a habit.

The spread is what the selection effect feeds on. If AIC selected exactly three on every series it would be fitting a slightly over-parameterised model consistently, its σ̂² would be an honest estimate from that model, and its interval would be close to right — a fixed wrong order costs bias in the point forecast and nothing in the coverage. What costs coverage is choosing differently on each series, on the basis of which order happened to fit best, and the wider the range that choice ranges over the more it costs.

That gives a testable consequence and it is the reason the range of orders offered is stated everywhere in this essay. Offer the criterion fewer orders and it has fewer chances to pick a low residual variance; offer it more and it has more. The numbers above use twelve. At eight the selection cost is smaller, and at twenty it would be larger, and a reader who takes 88.6% away from this page without the “twelve orders on offer” attached has taken away a number that is not about anything.

What each criterion selects, at 40 observations. 700 series from an AR(2) with coefficients 0.6 and -0.3, every order from 0 to 8 fitted to the same 32 responses so the log-likelihoods are comparable. AIC finds the true order 48.4% of the time and lands above it 28.3%; BIC finds it 47.6% and lands above it 6.3%. The closed form for one extra lag is P(χ²₁ > 2) = 15.73% for AIC, which does not depend on n at all, and P(χ²₁ > ln n) = 5.48% for BIC at this size, which falls to zero. Under the true order is the other failure and it is BIC's: 46.1% against 23.3%.
Fig. 4 The same forty-observation series with eight orders offered rather than twelve. AIC’s right-hand tail has less room to run into, and its distribution is correspondingly tighter — which is the same picture as before with the selection’s opportunity reduced rather than its behaviour changed.

The criterion that is right more often is also the one that costs less

The previous essay ended with an objection: AIC is not trying to find the true order, it is estimating out-of-sample prediction error, and overfitting is the cheap error for a forecaster because an unnecessary lag costs variance and no bias.

The objection is coherent and it does not survive the count. On these series AIC’s interval covers 88.6% and BIC’s 90.5%, and BIC’s advantage holds at every horizon measured — 89.9% against 87.3% at two steps, 92.4% against 90.1% at four. So at this sample size the criterion that finds the true order more often also produces the better interval, and the argument that overfitting is cheap is an argument about the point forecast rather than about the statement of uncertainty around it.

There is a boundary on that conclusion and it is the one the previous essay drew. At forty observations BIC under-selects heavily, and it still wins here — but at a sample size where BIC’s under-selection is severe enough to leave out a coefficient that matters a great deal, the bias in the point forecast would start to show, and nothing here says where that happens. What is measured is one setting, stated as one setting.

What the interval covers once the order is chosen as well. 1200 series of 25 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.3% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 92.8% at one step and 87.3% at 6. The third line chooses the order by BIC from the same data before computing the interval, which costs a further 5.8 points at h = 6.
Fig. 5 The three lines at twenty-five observations with BIC choosing. The selection line is closer to the fixed-order line than AIC’s was, and both are further from the oracle than at forty — the two costs add, and both grow as the series shortens.

Where the horizon comes into it

The shortfall is not constant across horizons, and its shape is the opposite of what the previous essay found for estimation alone.

Estimation error grew with the horizon: the plug-in interval was worse at six steps than at one, because φ̂ enters the h-step variance through a longer sum. Selection error shrinks with the horizon here — AIC’s shortfall against the fixed-order interval is 4.6 points at one step, 4.8 at two, 3.4 at four and 2.5 at six.

The reason is the ceiling from the field’s first essay. Every forecast interval converges on the same width — the unconditional spread of the series — regardless of which order was fitted, because that width is a property of the process and not of the model. At long horizons every candidate model is quoting nearly the same band, so which one was selected stops mattering. Selection can only do damage in the range where the models still disagree.

What the interval covers once the order is chosen as well. 1200 series of 40 observations from an AR(1) with φ = 0.9, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.7% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 94.3% at one step and 84.8% at 6. The third line chooses the order by AIC from the same data before computing the interval, which costs a further 3.9 points at h = 6.
Fig. 6 The same three lines on a much more persistent series, where the models keep disagreeing for longer because the band takes longer to reach its ceiling. The gap between the second and third lines survives further out than it does at φ = 0.7.

So the two costs are worst in different places, and an interval quoted one step ahead on a short series carries the maximum of both. That is also, by a considerable margin, the interval most often quoted.

What each criterion selects, at 80 observations. 700 series from an AR(2) with coefficients 0.6 and -0.3, every order from 0 to 12 fitted to the same 68 responses so the log-likelihoods are comparable. AIC finds the true order 64.9% of the time and lands above it 27.4%; BIC finds it 68.1% and lands above it 3.3%. The closed form for one extra lag is P(χ²₁ > 2) = 15.73% for AIC, which does not depend on n at all, and P(χ²₁ > ln n) = 3.63% for BIC at this size, which falls to zero. Under the true order is the other failure and it is BIC's: 28.6% against 7.7%.
Fig. 7 Eighty observations with twelve orders on offer. AIC’s tail is shorter than at forty and has not gone, which is the constant rate again — doubling the series halves neither the tail nor what it costs the interval.

What none of this repairs

An interval that covers 88.6% when it claims 95% is a defect with an obvious remedy shape and no obvious remedy. Three things are worth separating.

Widening it by a fixed factor is not a repair. The shortfall depends on the sample size, the horizon, the range of orders offered and the true coefficients, and a constant inflation calibrated at one of those settings is wrong at the others. That is the same objection this site raised against a critical value calibrated at one success rate and used at another, and it is the same objection for the same reason: a correction is only as transferable as the quantity it was computed at.

Not selecting is not available. The alternative to choosing an order is being given one, and nobody is. Fixing the order at some conventional value is itself a selection — made once, by somebody else, with no data at all — and its coverage is whatever it is at the order that got fixed.

Reporting the selection is not the same as accounting for it. A convention exists of printing the criterion’s table alongside the chosen model, so that a reader can see how close the runners-up were. It is a good habit and it does nothing to the coverage: the interval quoted is still the one computed from the winner, and a reader who can see that orders two and three scored within a hundredth of each other has been told the selection was close and given no way to widen anything by the right amount.

And the honest repair is expensive. Conditioning on the selection means the reference distribution has to be the distribution of the whole procedure, selection included, which means simulating the selection. That is available and it is not free, and it is the shape of the repair this site has built twice before — once for an adaptive allocation and once for a sequential design. Both times the cost was real and both times it bought a statement that did not depend on a quantity nobody has.

What the interval covers once the order is chosen as well. 1200 series of 100 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 94.8% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 94.3% at one step and 93.9% at 6. The third line chooses the order by AIC from the same data before computing the interval, which costs a further 0.2 points at h = 6.
Fig. 8 The three lines at a hundred observations. All three have closed towards 95%, which says both costs are estimation costs and both go away in the direction more data goes — slowly, and not at the same rate.

And the model still has to beat two things that estimate nothing

There is a closing question that the whole field has been circling and that a selected model does not escape.

An order was chosen, coefficients were estimated, an interval was computed and it covers 88.6%. Before any of that: is the fitted model better than carrying yesterday’s value forward, or than quoting the long-run average? Neither of those estimates a dynamic parameter and neither can be overfitted, so neither pays any of the costs measured in this field.

The band on which fitting a model is worth doing, at n = 50One-step squared error for three forecasts of the same next observation, on the same series and the same seeds, over 1500 series of 50 observations at each φ. The sample mean estimates no dynamics and beats the fitted model below φ = 0.119; the last value carried forward estimates nothing at all and beats it above φ = 0.923. Both crossings are solved from closed forms — σ²(1 + k/n) for the fitted model against 2γ₀(1 − φ) and γ₀(1 + (1+φ)/(n(1−φ))) — and both are functions of the length of the series alone. The band widens at both ends as n grows and never reaches either edge.0123400.2000.4000.6000.8001persistence φ of the seriesmean squared one-step forecast errorφ = 0.119φ = 0.923flat line: the last value · rising line: the sample mean · lowest inside the band: the fitted model1500 series of 50 at each of 10 values of φthe band runs from 0.119 to 0.923, both solved rather than read off
Fig. 9 The three, over the range of persistence, at fifty observations. The band is where fitting wins, and both of its edges are closed forms in the length of the series alone. Drag it and watch the band open at both ends as n grows.

At fifty observations the band runs from φ = 0.1186 to φ = 0.9231; at twenty it is 0.1656 to 0.8182, and at four hundred 0.0473 to 0.9900. The band never closes and it never reaches either edge of the stationary range, so at every sample size there are series for which the right model is no model.

And one thing that is free. Averaging the fitted forecast with the last value carried forward beats both of them wherever their errors are less than perfectly correlated, which is arithmetic about a quadratic rather than a fact about forecasting: the variance of ½(f₁ + f₂) is ¼(V₁ + V₂ + 2ρ√(V₁V₂)). At φ = 0.9 on fifty observations the two errors correlate at 0.9545, the fitted model gives 1.0544, the last value gives 1.0509, and their average gives 1.0287 — below both, at no cost in data, information or assumption.

That is an odd note to end a field on and it is the honest one. The four essays here have measured a band that is too narrow, a point forecast that is biased, a criterion whose error rate does not fall with n, and a selection step that costs more than the estimation. The one thing that improves a forecast unconditionally is not a better model. It is refusing to pick between two.

It is worth being precise about why the average wins, because it is not because the two forecasts are individually good. At φ = 0.9 both of them are poor in absolute terms — 1.05 against a noise floor of 1.00 — and the average is 1.03. What the average exploits is that their errors are not the same error. The fitted model errs by getting φ̂ slightly wrong and the last-value rule errs by assuming φ is one, and where one is long the other tends to be short. Nothing about that requires either of them to be right, and nothing about it requires knowing which is better, which is exactly why it is free: a forecaster who cannot tell which of two methods to trust is in the best possible position to average them.

The limit is also stated by the same quadratic. As ρ approaches one the improvement goes to zero, because two forecasts making the identical error are one forecast. The measured 0.9545 is high and still leaves room; two methods that share a fitted parameter would leave less.

The checks

Two claims are gated in this field’s library.

Selection costs coverage on top of estimation, asserted as a chain rather than as a single comparison: the oracle interval must be within two points of 95%, the fixed-order interval must be below it by more than three standard errors, and the AIC interval must be below that by more than two points — with BIC required to sit between them. A chain is the right form because each link is what makes the next one attributable.

And the combination beats its parts, asserted against its own closed form rather than only against the two curves: the measured error of the average has to match ¼(V₁ + V₂ + 2ρ√(V₁V₂)) to within two per cent, which is the check that the improvement is the quadratic and not an accident of one setting.

The band’s two edges are gated separately and by three routes, which is more than this field’s other claims get and is warranted by how load-bearing they are. Inside the band the fitted model has to beat both benchmarks at every measured φ; below the lower edge the sample mean has to win; above the upper edge the last value has to win; and the upper edge has to agree with its leading-order form (n − k)/(n + k) to two figures. The lower edge’s leading-order form is deliberately not required to agree, and the check asserts the direction of the disagreement instead — √(k/(n + k)) drops what estimating a dependent series’ mean costs, which makes the mean look better than it is and puts the crossing further right than it belongs. Widening a tolerance until an approximation passed would have buried exactly the term worth knowing about, which is the trap this site has recorded before, when a figure claimed machine precision and delivered four decimal places.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Benchmark forecastCoverageForecast intervalInformation criterionModel selectionMonte CarloOrder-selectionOverfittingParameter uncertaintySelection bias