The interval after the choice
Worth reading first: What the model says next · When the looking happens.
Two costs have been measured in this field so far and both were charged to the same interval. The formula is right when the parameters are known. Substituting estimates for them takes a 95% interval down to about 93% on a short series. That was the previous essay, and it held one thing fixed that nobody is ever handed: the order.
The order is chosen. It is chosen by a criterion, and the criterion is computed from the same observations the interval is computed from. That makes the selection part of the procedure, and the site’s rule about procedures applies: a procedure’s advertised property is a claim about what happens when the whole of it is run, and the whole of it includes the step nobody writes down.
Counted, on identical series
Forty observations from a second-order process, twenty-five hundred series, orders up to twelve offered to the criterion. Three intervals per series, all from the same data.
- 95.0% — the interval at the true parameters and the true order.
- 93.2% — the same interval with the coefficients estimated, the order given.
- 90.5% — the order chosen by BIC first.
- 88.6% — the order chosen by AIC first.
Estimation costs 1.8 points. Selection costs a further 2.7 with BIC and 4.6 with AIC. The step that appears in no output and in no formula is, on this setting, more expensive than the one every textbook derives a correction for.
The comparison is within seeds throughout, which matters more here than usual because three of the four numbers are within a few points of each other. The same series produces all four verdicts: the same shocks, the same history, the same future observation held out. What differs between the rows is only how much of the procedure was allowed to look at the data.
Why selection costs anything at all
The mechanism is worth stating because it is not the obvious one.
The obvious story is that a wrongly chosen order gives a wrong variance, and wrong variances sometimes give narrow intervals. That is true and it is the smaller half. The larger half is a selection effect, and it is the same effect this site has already measured on an estimate rather than an interval.
A criterion picks the order whose residual variance, penalised, is smallest. Among the orders that fit about equally well, it picks the one that happened to fit best — and “happened to fit best” means “happened to have the smallest residual variance on this particular realisation”. The selected model’s σ̂² is therefore not an unbiased estimate of σ²: it is the minimum of several, chosen because it was the minimum.
An underestimated σ̂² is a narrow interval. So the interval after selection is narrower than the interval at the true order, on average, for the same reason that the estimate from a chosen arm is inflated: the choice was made on the quantity the estimate is about.
That also explains the ordering. AIC offers a cheaper penalty and therefore selects from a wider effective set — its mean selected order here is 3.00 against BIC’s 1.53 — so it has more chances to pick a low residual variance and the selection effect is larger. The criterion that overfits more is the criterion whose interval is narrower than it should be by more, and the two facts are the same fact.
The same three steps, as widths
Coverage is the natural unit for counting and it is not the unit anybody reports. Converting each row into how wrong the interval’s width is says the same thing in the quantity that appears in a table.
An interval built at ±1.96 standard errors covers c when the true spread is a factor r larger than the standard error it used, with . Inverting:
- 93.2% with the order known: the interval is 7.4% too narrow.
- 90.5% after BIC: 17.3% too narrow.
- 88.6% after AIC: 24.0% too narrow.
In variances that is 1.15, 1.38 and 1.54. So an interval that has had its order chosen by AIC is pricing about two thirds of the uncertainty it has.
What would have to be done to it
The same arithmetic gives the repair, in the crude form that at least states the size of the problem.
To reach 95% coverage the multiplier would have to be 1.96 × 1.240 = 2.43 after AIC selection, 2.30 after BIC, and 2.11 with the order known. Equivalently, the interval labelled 95% after an AIC selection is the interval a 98.5% label would have to be attached to before the multiplier came out right.
That is not a proposal — the factor is a property of this setting and would have to be re-measured at every sample size, horizon and process — and it is the right way to hear the numbers. A 95% label that needs a 98.5% multiplier is not a small calibration error. It is a different interval.
Which of the three steps is the large one
Set the three costs side by side: estimation 1.8 points, selection 2.7 more with BIC and 4.6 more with AIC.
So selection is 1.5 times estimation under BIC and 2.6 times under AIC. Of the total 6.4-point shortfall in the AIC row, 72% is the step that appears in no output.
On twenty-five hundred series a coverage near 90% carries a standard error of 0.6 points unpaired, and every comparison here is within seeds, so the 1.9-point gap between the two criteria is comfortably real and so is each step.
The ordering is what makes the field worth having. The correction textbooks derive is for the smallest of the three effects, and the largest of the three is the one that is not written down anywhere in the procedure — it happens before the interval’s formula is reached, and by the time the formula runs there is nothing left in the data to say it happened.
What the selected order actually is
The mean selected order is the number that makes the mechanism concrete, and it is worth reading carefully because both criteria are wrong about it in ways that look nothing alike.
The truth is two. AIC’s mean is 3.00 and BIC’s is 1.53. So on average AIC is fitting one lag too many and BIC half a lag too few, and neither of those descriptions is what is happening on any individual series — a mean of 3.00 over orders offered up to twelve is a spread, not a habit.
The spread is what the selection effect feeds on. If AIC selected exactly three on every series it would be fitting a slightly over-parameterised model consistently, its σ̂² would be an honest estimate from that model, and its interval would be close to right — a fixed wrong order costs bias in the point forecast and nothing in the coverage. What costs coverage is choosing differently on each series, on the basis of which order happened to fit best, and the wider the range that choice ranges over the more it costs.
That gives a testable consequence and it is the reason the range of orders offered is stated everywhere in this essay. Offer the criterion fewer orders and it has fewer chances to pick a low residual variance; offer it more and it has more. The numbers above use twelve. At eight the selection cost is smaller, and at twenty it would be larger, and a reader who takes 88.6% away from this page without the “twelve orders on offer” attached has taken away a number that is not about anything.
The criterion that is right more often is also the one that costs less
The previous essay ended with an objection: AIC is not trying to find the true order, it is estimating out-of-sample prediction error, and overfitting is the cheap error for a forecaster because an unnecessary lag costs variance and no bias.
The objection is coherent and it does not survive the count. On these series AIC’s interval covers 88.6% and BIC’s 90.5%, and BIC’s advantage holds at every horizon measured — 89.9% against 87.3% at two steps, 92.4% against 90.1% at four. So at this sample size the criterion that finds the true order more often also produces the better interval, and the argument that overfitting is cheap is an argument about the point forecast rather than about the statement of uncertainty around it.
There is a boundary on that conclusion and it is the one the previous essay drew. At forty observations BIC under-selects heavily, and it still wins here — but at a sample size where BIC’s under-selection is severe enough to leave out a coefficient that matters a great deal, the bias in the point forecast would start to show, and nothing here says where that happens. What is measured is one setting, stated as one setting.
Where the horizon comes into it
The shortfall is not constant across horizons, and its shape is the opposite of what the previous essay found for estimation alone.
Estimation error grew with the horizon: the plug-in interval was worse at six steps than at one, because φ̂ enters the h-step variance through a longer sum. Selection error shrinks with the horizon here — AIC’s shortfall against the fixed-order interval is 4.6 points at one step, 4.8 at two, 3.4 at four and 2.5 at six.
The reason is the ceiling from the field’s first essay. Every forecast interval converges on the same width — the unconditional spread of the series — regardless of which order was fitted, because that width is a property of the process and not of the model. At long horizons every candidate model is quoting nearly the same band, so which one was selected stops mattering. Selection can only do damage in the range where the models still disagree.
So the two costs are worst in different places, and an interval quoted one step ahead on a short series carries the maximum of both. That is also, by a considerable margin, the interval most often quoted.
What none of this repairs
An interval that covers 88.6% when it claims 95% is a defect with an obvious remedy shape and no obvious remedy. Three things are worth separating.
Widening it by a fixed factor is not a repair. The shortfall depends on the sample size, the horizon, the range of orders offered and the true coefficients, and a constant inflation calibrated at one of those settings is wrong at the others. That is the same objection this site raised against a critical value calibrated at one success rate and used at another, and it is the same objection for the same reason: a correction is only as transferable as the quantity it was computed at.
Not selecting is not available. The alternative to choosing an order is being given one, and nobody is. Fixing the order at some conventional value is itself a selection — made once, by somebody else, with no data at all — and its coverage is whatever it is at the order that got fixed.
Reporting the selection is not the same as accounting for it. A convention exists of printing the criterion’s table alongside the chosen model, so that a reader can see how close the runners-up were. It is a good habit and it does nothing to the coverage: the interval quoted is still the one computed from the winner, and a reader who can see that orders two and three scored within a hundredth of each other has been told the selection was close and given no way to widen anything by the right amount.
And the honest repair is expensive. Conditioning on the selection means the reference distribution has to be the distribution of the whole procedure, selection included, which means simulating the selection. That is available and it is not free, and it is the shape of the repair this site has built twice before — once for an adaptive allocation and once for a sequential design. Both times the cost was real and both times it bought a statement that did not depend on a quantity nobody has.
And the model still has to beat two things that estimate nothing
There is a closing question that the whole field has been circling and that a selected model does not escape.
An order was chosen, coefficients were estimated, an interval was computed and it covers 88.6%. Before any of that: is the fitted model better than carrying yesterday’s value forward, or than quoting the long-run average? Neither of those estimates a dynamic parameter and neither can be overfitted, so neither pays any of the costs measured in this field.
At fifty observations the band runs from φ = 0.1186 to φ = 0.9231; at twenty it is 0.1656 to 0.8182, and at four hundred 0.0473 to 0.9900. The band never closes and it never reaches either edge of the stationary range, so at every sample size there are series for which the right model is no model.
And one thing that is free. Averaging the fitted forecast with the last value carried forward beats both of them wherever their errors are less than perfectly correlated, which is arithmetic about a quadratic rather than a fact about forecasting: the variance of ½(f₁ + f₂) is ¼(V₁ + V₂ + 2ρ√(V₁V₂)). At φ = 0.9 on fifty observations the two errors correlate at 0.9545, the fitted model gives 1.0544, the last value gives 1.0509, and their average gives 1.0287 — below both, at no cost in data, information or assumption.
That is an odd note to end a field on and it is the honest one. The four essays here have measured a band that is too narrow, a point forecast that is biased, a criterion whose error rate does not fall with n, and a selection step that costs more than the estimation. The one thing that improves a forecast unconditionally is not a better model. It is refusing to pick between two.
It is worth being precise about why the average wins, because it is not because the two forecasts are individually good. At φ = 0.9 both of them are poor in absolute terms — 1.05 against a noise floor of 1.00 — and the average is 1.03. What the average exploits is that their errors are not the same error. The fitted model errs by getting φ̂ slightly wrong and the last-value rule errs by assuming φ is one, and where one is long the other tends to be short. Nothing about that requires either of them to be right, and nothing about it requires knowing which is better, which is exactly why it is free: a forecaster who cannot tell which of two methods to trust is in the best possible position to average them.
The limit is also stated by the same quadratic. As ρ approaches one the improvement goes to zero, because two forecasts making the identical error are one forecast. The measured 0.9545 is high and still leaves room; two methods that share a fitted parameter would leave less.
The checks
Two claims are gated in this field’s library.
Selection costs coverage on top of estimation, asserted as a chain rather than as a single comparison: the oracle interval must be within two points of 95%, the fixed-order interval must be below it by more than three standard errors, and the AIC interval must be below that by more than two points — with BIC required to sit between them. A chain is the right form because each link is what makes the next one attributable.
And the combination beats its parts, asserted against its own closed form rather than only against the two curves: the measured error of the average has to match ¼(V₁ + V₂ + 2ρ√(V₁V₂)) to within two per cent, which is the check that the improvement is the quadratic and not an accident of one setting.
The band’s two edges are gated separately and by three routes, which is more than this field’s other claims get and is warranted by how load-bearing they are. Inside the band the fitted model has to beat both benchmarks at every measured φ; below the lower edge the sample mean has to win; above the upper edge the last value has to win; and the upper edge has to agree with its leading-order form (n − k)/(n + k) to two figures. The lower edge’s leading-order form is deliberately not required to agree, and the check asserts the direction of the disagreement instead — √(k/(n + k)) drops what estimating a dependent series’ mean costs, which makes the mean look better than it is and puts the crossing further right than it belongs. Widening a tolerance until an approximation passed would have buried exactly the term worth knowing about, which is the trap this site has recorded before, when a figure claimed machine precision and delivered four decimal places.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A table and a list — both name benchmark forecast, information criterion, model selection, monte carlo, order-selection, overfitting
- The eighth that was not a constant — both name benchmark forecast, information criterion, model selection, monte carlo, order-selection, overfitting
- Two factors pointing opposite ways — both name benchmark forecast, information criterion, model selection, monte carlo, order-selection, overfitting
- A line that beats two curves — both name information criterion, model selection, monte carlo, overfitting, parameter uncertainty
- A step that is not a ratio — both name information criterion, model selection, monte carlo, order-selection, overfitting
- The displacement is a parameter count — both name benchmark forecast, information criterion, model selection, monte carlo, overfitting
Named objects
A flat tag is an object no other essay names yet.
Benchmark forecastCoverageForecast intervalInformation criterionModel selectionMonte CarloOrder-selectionOverfittingParameter uncertaintySelection bias