Series

Forecast — the series

9 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. One forecast, and the band the arithmetic puts round it. An AR(1) with φ = 0.75, 60 observations, fitted by least squares and forecast 14 steps ahead. The point forecast decays towards the fitted mean at φ̂^h; the band is ±1.96 standard errors from σ̂²Σψ̂², which grows with the horizon and stops at the unconditional spread 1.72. The dashed pair is the same band computed at the true parameters, which nobody has. The marks past zero are what actually arrived: 12 of 14 inside the band this once, which is one draw and settles nothing.

    What the model says next

    The usual account of a time series stops at estimation. A forecast asks the other question — not what the parameter is but what the next observation will be — and the band round it is a closed form that grows with the horizon and then stops growing, at a value the series was going to reach anyway.

    part 1 · forecast
  2. What a 95% forecast interval covers, counted. 1200 series of 25 observations from an AR(1) with φ = 0.7, at each horizon, on one set of seeds. The upper line is the interval computed at the true parameters — it covers 95.3% on average, which is the check that σ²Σψ² is the right formula rather than a claim about anything a forecaster can do. The lower line is the same formula fed σ̂² and φ̂: 92.8% at one step and 87.3% at 6. The interval that would cover what it claims is 6.9% wider at one step.

    The interval that forgets it estimated

    The forecast band is derived for a model whose parameters are known, and then computed by putting estimates into it. Counted, the 95% interval covers 87.3% six steps ahead on twenty-five observations, and the point forecast inside it returns to the mean a third faster than the series does.

    part 2 · forecast
  3. One comparison, and the two error bars it can be given. 60 rolling origins, a window of 60 observations, forecasts 4 steps ahead, at the persistence φ = 0.8256 where the two benchmarks have exactly equal population mean squared error. Each mark is one origin's difference in squared error; the horizontal line is their mean, 0.6522. The two vertical bars at the right are ±1.96 standard errors round that mean computed two ways — 0.5337 treating the differences as independent, 0.6880 allowing for the overlap between neighbouring forecasts. The null is true here by construction, so an interval that excludes zero is a mistake, and the narrow one does it far more often than the wide one.

    Which forecast is better

    Two forecasters, one series, and a difference in mean squared error. Whether that difference is real is a hypothesis test, its terms are not independent, and the standard error it needs is not the one a t-test computes.

    part 3 · evaluation
  4. A test between nested models, under a null that is true. 1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.1663 against 1.0583 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 67.2% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.9%.

    When one model contains the other

    The comparison a forecaster most often wants is between a model and the same model with one more term. That is exactly the comparison the standard test cannot make — and it fails by declaring the smaller model significantly better, more confidently the more data it is given.

    part 4 · evaluation
  5. Three quantities, and only one of them crosses zero. Two forecasts of an AR(1) — the last value carried forward and the mean of the last 60 observations — at 1 step ahead. The curve through zero is σ₁² − σ₂², the difference in expected squared error that a comparison of accuracy tests; it changes sign at φ = 0.4922. The two curves above it are σ₁² − σ₁₂ and σ₂² − σ₁₂, the quantities the two encompassing tests are about, and neither of them comes near zero anywhere: the smallest value either takes across the range is 0.008 times the variance of the series. All three are closed forms in φ, R and h with no simulation in them. Equal accuracy is one hypothesis about this picture and encompassing is another, and a set of numbers can satisfy either without the other.

    What the other forecast adds

    Two forecasters, one series, and two different questions about them. Which is more accurate has an answer that changes with the persistence of the series; whether either is redundant has an answer that never changes at all.

    part 5 · ranking
  6. What each variant loses before anything has been searched for. The mean loss differential of each of the eight variants against the benchmark, over 600 tables of 60 origins, with every fit given 71 rows. The series is an AR(1) and every variant adds a lag whose coefficient is zero, so in population the two forecasts are the same forecast and the difference drawn here is estimation noise and nothing else. The marked line is σ²(q₁ − q₀)/n = -0.01408, which is an expression in how many coefficients each model has and how many rows it was fitted on — it knows nothing about the series, the persistence or which lag the variant added, and every bar is within a fifth of it. This is the amount a reference distribution recentred at each column's own sample mean believes the candidates are already behind by.

    A table of nested models

    A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.

    part 6 · search
  7. One true null, one table, five readings. every subset of four, fifteen models, at a null where nothing any candidate holds is worth anything, over 500 draws. Each bar is the share of draws on which that reading declares a difference at a nominal 5%. The reading is the whole of the difference between the bars: the data is identical. An open search over all 210 ordered pairs rejects 76.2%; the table's own 5% point is 3.163 against the 1.671 a single comparison uses. Bonferroni takes the open reading to 0.6% — and on the nested ladder the same correction does not reach the nominal level at all, because there the excess is a shift in the mean rather than a maximum over many.

    When the benchmark is a candidate

    A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.

    part 7 · select
  8. The crossing is in the dependence, not in the split. Regret of each rule as the design and the errors are made persistent at the same coefficient, scored on fresh rows because the closed form assumes exactly what is being taken away. An optimism theorem counts rows; when the rows repeat each other there are fewer of them than there are rows, the penalty is too small for the fit it is correcting, and the criterion starts buying coefficients it should not — its average winner grows from 3.31 coefficients to 3.90. The hold-out never used the theorem and overtakes at ρ ≈ 0.81. Schwarz's criterion, worst of the three on independent rows, is best on repeating ones — its heavier penalty is right for the wrong reason.

    Where the two searches cross

    The obvious dial between a criterion and a hold-out is how much of the sample to hold out, and moving it never changes the answer. The dial that does is one nobody chooses — how much each row repeats the one before it — and the two rules change places at about 0.81.

    part 8 · proxy
  9. The row count entered twice, and a penalty is one place. Regret against the best available model as the errors are made persistent. Counting rows more than quadruples; both penalty repairs — the trace, and the scalar effective sample size — are worse than it at every persistence measured; and whitening the sample and keeping the ordinary penalty falls, recovering 86.9% of what counting rows gives up at ρ = 0.85. Doing it at an estimated ρ recovers 80.6%, so having to estimate the dependence from the rows being selected on costs 7.3% of what knowing it is worth. Mallows' forms are drawn beside the logarithmic ones and behave the same, which is what rules the linearisation out.

    The repair that was exact and made it worse

    A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.

    part 9 · effective

All series