One comparison, and the two error bars it can be given
60 rolling origins, a window of 60 observations, forecasts 4 steps ahead, at the persistence φ = 0.8256 where the two benchmarks have exactly equal population mean squared error. Each mark is one origin's difference in squared error; the horizontal line is their mean, 0.6522. The two vertical bars at the right are ±1.96 standard errors round that mean computed two ways — 0.5337 treating the differences as independent, 0.6880 allowing for the overlap between neighbouring forecasts. The null is true here by construction, so an interval that excludes zero is a mistake, and the narrow one does it far more often than the wide one.
Comparing two forecastersslider: steps ahead each forecast is made, 4 positionswide10 views
What else it draws
The same object, drawn to answer the other questions the essays put to it.
1500 comparisons of 60 origins at the persistence where the two forecasts are exactly equally good, so every rejection here is a false one. The same mean difference is divided by five different standard errors. Treating the differences as independent rejects 26.5% of the time where 5% is claimed; truncating the long-run variance at h − 1 lags — the textbook prescription — brings it to 11.6%; the small-sample correction leaves 10.5%; and a test given the true variance of the mean rejects 4.7%. The gap between the last two is the estimator's own bias, and it does not close with more origins.
1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.1663 against 1.0583 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 67.2% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.9%.
3000 series of 50 observations at each persistence. The lower curve is the counted bias of the least-squares estimate of φ, and the open marks on it are −(1 + 3φ)/n, computed rather than fitted. The upper curve is the bias left after adding that quantity back, evaluated at the estimate rather than at the truth nobody has: -0.0020 at φ = 0.3, -0.0023 at φ = 0.5, -0.0039 at φ = 0.7, -0.0059 at φ = 0.8, -0.0108 at φ = 0.9, -0.0165 at φ = 0.95. The formula is a leading-order expression and it understates the bias where the persistence is nearest one — -0.0882 counted against -0.0770 predicted at φ = 0.95, which is the corner of the parameter space every one of these approximations is worst in.
The truth is φ^6 = 0.3771. no correction averages 0.2616 with a spread of 0.1646 and a squared forecast error of 3.2516; the formula, on the persistence averages 0.4213 with a spread of 0.2528 and a squared forecast error of 3.4827; the bootstrap, on the persistence averages 0.4355 with a spread of 0.2655 and a squared forecast error of 3.5120; the bootstrap, on the decay factor averages 0.3375 with a spread of 0.2278 and a squared forecast error of 3.4132. 800 series, 100 bootstrap refits each.
Three routes at five horizons, 500 series each. As a share of the truth, the uncorrected decay factor is off by -9.3% at one step and -35.4% at twelve; the persistence-corrected route by 63.5% at twelve and the decay-corrected route by -14.2%.
The closed-form threshold is φ̂ > (n − 1)/(n + 3), which is 0.8571 at n = 25, 0.9245 at n = 50, 0.9612 at n = 100, 0.9803 at n = 200. Counted over 3,000 series at each cell, the rate reaches 37.9% at φ = 0.99 on 25 observations and falls to 29.5% at 200.
The correction exceeds one on 31.1% of series at this setting. left where it lands: squared forecast error 12.828, average decay factor 0.7974 against a true 0.7351; capped at 0.995: squared forecast error 5.680, average decay factor 0.5950 against a true 0.7351; capped at 1 − 1/n: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction scaled to fit: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction refused where it leaves: squared forecast error 5.868, average decay factor 0.4423 against a true 0.7351.
The plug-in interval covers 88.42% against a claimed 95%. Correcting the variance recovers 0.56 points, propagating the persistence's own standard error recovers 0.40, correcting the persistence recovers 2.66, and all three together recover 4.20 — leaving 2.38 points unaccounted for.
At one step the plug-in interval covers 93.40% and at twelve it covers 86.92%. All three repairs together bring twelve steps to 91.85%, which is still 3.15 points short, and the gap between the repaired curve and the claim widens with the horizon rather than closing.
Where it is used
9 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 9 different questions.
- A table of nested models Searching among fitted models
- Which forecast is better Comparing two forecasters
- A null with a model in it Searching among fitted models
- When one model contains the other Comparing two forecasters
- Correcting the persistence Comparing two forecasters
- The repair that moves the wrong number Comparing two forecasters
- Correcting the forecast instead Comparing two forecasters
- The correction that leaves the region Comparing two forecasters
- What the interval is short by Comparing two forecasters