Comparing two forecasters

When one model contains the other

The comparison a forecaster most often wants is between a model and the same model with one more term. That is exactly the comparison the standard test cannot make — and it fails by declaring the smaller model significantly better, more confidently the more data it is given.

Worth reading first: What the model says next · Choosing the order.

The comparison in the previous essay was between two forecasters with nothing in common: one carries the last value forward, the other averages the window, and neither is a special case of the other. That is not the comparison most forecasters actually want to make, and the machinery that essay built — the loss differential, the long-run variance, the exactly true null — is assumed here throughout.

The question that gets asked is whether adding something helps. Does a second lag improve the forecast? Does the extra variable earn its place? Every one of those is a comparison between a model and the same model with one more term in it, and the smaller model is what the larger one becomes when the extra coefficient is zero. The two forecasters are nested.

Under the null hypothesis that the extra term is worthless, the two forecasts are not merely equally accurate. They are, in population, the same forecast. And a test built to compare two different things has nothing left to divide by.

What is actually being tested, and why it is degenerate

Take an AR(1) truth, and forecast it two ways: with a fitted AR(1), and with a fitted AR(4) whose last three coefficients are estimating zeros. In population the two produce identical forecasts, so the population loss differential is zero — and so is its variance, because the two error sequences are the same sequence.

Nothing in a finite sample is degenerate, of course. What survives is entirely the estimation noise of the larger model: three coefficients that should be zero are estimated as something else, their contribution to the forecast is noise, and noise added to an unbiased forecast increases its squared error. So the sample loss differential is not centred at zero. It is centred below it.

A test between nested models, under a null that is true1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.1663 against 1.0583 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 67.2% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.9%.the smaller model declared better67.2%the larger model declared better0.0%the recentred statistic4.9%5%, which all three claim1000 comparisons, 200 origins, window 40, the extra coefficients zeroMSE 1.0583 against 1.1663
Fig. 1 Two hundred origins, a window of forty, an AR(1) truth. The larger model’s mean squared error is 1.1663 against 1.0583 — worse by ten percent, entirely from estimating coefficients that are not there. The three bars are what a test makes of that.

The mean loss differential is −0.1080, and the ordinary test, applied as it always is, reports what it sees: the smaller model is significantly better, in 67.2% of comparisons, at a nominal 5% level.

That number is not a size distortion in the usual sense. The test is right that the smaller model forecast better in this sample, and right that it will do so again. What it is wrong about is the hypothesis a reader takes it to be testing — whether the extra lags carry information — and there the answer is no, they carry none, which is exactly the null the test claims to be examining.

The size of the shift is not mysterious either. A model estimating k coefficients from a window of R observations carries about k/R of the residual variance as estimation noise in its one-step forecast, and the extra lags contribute their share of that whether or not they are worth anything. Three useless lags on a window of forty is three fortieths of the variance handed to the larger model as a handicap, and the measured difference in mean squared error — 1.1663 against 1.0583, or ten percent — is what that arithmetic predicts once the estimated intercept and first lag are counted on both sides.

One comparison, and the two error bars it can be given. 200 rolling origins, a window of 40 observations, forecasts 1 step ahead, at the persistence φ = 0.4887 where the two benchmarks have exactly equal population mean squared error. Each mark is one origin's difference in squared error; the horizontal line is their mean, 0.1637. The two vertical bars at the right are ±1.96 standard errors round that mean computed two ways — 0.1355 treating the differences as independent, 0.1181 allowing for the overlap between neighbouring forecasts. The null is true here by construction, so an interval that excludes zero is a mistake, and the narrow one does it no more often than the wide one.
Fig. 2 For contrast, the same picture for two forecasters that are not nested, from the previous essay: the differences scatter about a mean that is genuinely zero. In the nested case the whole cloud sits below the axis, and no amount of care with the standard error changes where the cloud is.

Read the way the question is asked, it never rejects

The two-sided reading above is charitable. Anybody comparing a model with its own extension has a direction in mind: does the larger model help? That is a one-sided test, and the statistic has to exceed +1.645 for the larger model to be declared better.

Across two thousand comparisons under a true null, the share exceeding it is 0.0%.

A test between nested models, under a null that is true. 1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR2 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.0912 against 1.0583 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 24.5% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.0%.
Fig. 3 The same picture with one extra lag rather than three. The failure is milder — 24.5% rather than 67.2% — because there is less estimation noise to charge the larger model with, and the one-sided reading is unchanged at 0.0%.

A test that never rejects is not a conservative test. It is a test whose rejection region the statistic cannot reach, and the reason is structural: under the null the statistic’s distribution is shifted left by the estimation noise, and under any alternative small enough to be interesting it is shifted left by almost as much. The test has no power in the region where a forecaster needs it, and it manufactures a finding in the region where nothing is happening.

A test between nested models, under a null that is true. 1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.0736 against 1.0285 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 34.0% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.0%.
Fig. 4 And the same comparison on a window twice as long. Estimation noise falls with the window, so the handicap the larger model carries falls with it too — but the evaluation is unchanged in length, so the smaller standard error finds the smaller difference about as reliably. Nothing about the failure is repaired by fitting on more data.

More data makes it worse, which is the diagnostic

A test that fails from too little data is repaired by more. This one is not.

More data makes the wrong answer more certain. An AR(1) truth, an AR(1) forecast against an AR4 forecast, and the extra coefficients are zero — the two models are the same model in population. The upper curve is how often the ordinary test declares the smaller model significantly better: 18.0% at 50 origins, 35.1% at 100 origins, 67.3% at 200 origins, 94.8% at 400 origins. It rises, because the larger model really is worse out of sample by its own estimation error, and a longer comparison measures that difference more precisely. The lower curve is the recentred statistic, which holds 5% throughout.
Fig. 5 The share of comparisons declaring the smaller model significantly better, against the length of the evaluation: 18.0% at fifty origins, 35.1% at a hundred, 67.3% at two hundred, and 94.8% at four hundred. The lower curve is the recentred statistic, at its nominal level throughout.

At four hundred origins the wrong answer arrives in nineteen comparisons out of twenty. The mean loss differential barely moves across that sweep — −0.1033, −0.1096, −0.1085, −0.1082 — because it is a real feature of the two procedures rather than a sampling accident. What changes is the standard error, which shrinks as the evaluation lengthens, so the same real difference is detected with increasing confidence.

A difference that is real and a hypothesis that is true are not in conflict here. The larger model really does forecast worse out of sample; the extra coefficients really are zero. Both statements are correct, and a test that reports the first as evidence against the second is answering a question nobody asked.

This is the winner’s curse seen from the other side of the same arithmetic. There, a quantity selected for being large is large by more than it should be; here, a model penalised for its own estimation noise is penalised reliably, and repeating the experiment does not average the penalty away because it is not noise in the comparison — it is a property of the pair being compared.

The repair is a subtraction

The fix follows from writing down where the shift comes from. The larger model’s forecast is the smaller model’s forecast plus estimation error; its squared error is therefore the smaller model’s squared error plus the square of that estimation error, minus a cross term with mean zero under the null. So the expected loss differential under the null is not zero but −E[(ŷ₁ − ŷ₂)²] — and that quantity is observable. Both forecasts are in hand at every origin.

Adding it back gives

fₜ = e₁ₜ² − [e₂ₜ² − (ŷ₁ₜ − ŷ₂ₜ)²]

which is centred at zero under the null by construction, and is read one-sided against a normal because the alternative only lives on one side. It is a subtraction rather than a new theory, and everything else about the test — the long-run variance, the bandwidth, the origins — is unchanged.

Its counted size is 4.9% with three extra lags and 4.0% with one, at a nominal 5%.

The repaired test finds what the original cannot. An AR(1) forecast against an AR(2) forecast, with the second coefficient swept from zero — where the two models are the same model — to 0.4, where the larger one is genuinely better. Both curves are one-sided tests of the same hypothesis on the same forecasts. The ordinary statistic finds the larger model better 5.5% of the time at φ₂ = 0.2, which is its own nominal level and therefore no evidence at all, while the recentred one finds it 47.0% of the time. At the left both are honest: 4.1% and 0.0% under a true null.
Fig. 6 The second coefficient swept from zero, where the two models coincide, to 0.4, where the larger one is plainly better. Both curves are one-sided tests of the same hypothesis on the same forecasts. The recentred statistic finds the improvement; the ordinary one does not find it until it is large enough that nobody needed a test.

At φ₂ = 0.2 the larger model is genuinely the better forecaster — mean squared error 1.0980 against 1.1057 — and the ordinary test finds it 5.5% of the time, which is its own nominal level and therefore no evidence whatever. The recentred statistic finds it 47.0% of the time. At φ₂ = 0.3 the two are 35.3% and 87.8%.

The denominator cannot save it

Everything the previous essay was about — the overlap between neighbouring forecasts, the long-run variance, the bandwidth — is orthogonal to this failure and worth saying so explicitly, because the two problems are easy to conflate under the heading forecast comparison is harder than it looks.

The overlap problem is about the denominator: the mean difference is fine and its standard error is too small, so the statistic is inflated in whichever direction the difference happens to lie. It is symmetric, it is repaired by estimating the right variance, and at one step ahead it does not arise at all.

The nesting problem is about the numerator: the mean difference is not centred where the null says, and no denominator can move it. All three of the tests measured here use the same corrected long-run variance; they differ only in what is being averaged. Every comparison in this essay is at one step ahead, where the loss differentials are uncorrelated and the denominator is beyond reproach, and the failure is at its most complete.

Where the statistic actually sits, in standard errors

The four rejection rates along the evaluation sweep — 18.0%, 35.1%, 67.3% and 94.8% at fifty, one hundred, two hundred and four hundred origins — can be turned back into the quantity that produced them, which says more than the rates do.

A two-sided test at 5% rejects when the statistic passes −1.96, and essentially every rejection here is in that tail. Inverting a normal at each rate puts the statistic’s centre at about −1.04, −1.58, −2.41 and −3.59 standard errors. At four hundred origins the ordinary statistic sits three and a half standard errors below zero at a true null, which is a cleaner way of saying what “declares the smaller model significantly better nineteen times in twenty” means.

Those four numbers grow by roughly 1.5 at each doubling, against the √2 = 1.41 that a fixed mean difference divided by a standard error shrinking as 1/√P predicts. Agreeing to within seven per cent over an eightfold range is what says the mechanism is exactly the stated one: the difference is constant — the four mean differentials are −0.1033, −0.1096, −0.1085 and −0.1082 — and only the denominator is moving.

So the test’s failure has a growth rate, and the growth rate is the ordinary one. That is the part worth carrying past this essay. A defect that grows as √P is not distinguishable, from inside one evaluation, from a real effect being detected with increasing confidence; both look like a statistic drifting away from zero at the same rate. The only thing separating them is knowing where the null puts the statistic, and nothing in the output says.

It also prices how long an evaluation would have to be for the ordinary test to stop failing, which is a question with no answer: the shift grows without bound, so the failure gets worse for ever. The only value of P at which the ordinary statistic is honest here is a small one, and it is honest there by having too little power to be wrong.

Completing the k/R arithmetic

The handicap is attributed above to estimation noise at a rate of about k/R of the residual variance, and carrying that count through is worth doing because it says how much of the effect the simple version explains.

A one-step forecast from a model estimating k coefficients on a window of R carries about σ²(1 + k/R). The smaller model estimates an intercept and one lag, so k = 2 and its mean squared error of 1.0583 implies σ² ≈ 1.008. The larger estimates five, so the predicted gap is 3σ²/40 = 0.076, against a measured mean differential of 0.108.

The leading term is therefore about seven-tenths of the handicap, and the remaining three-tenths is the next order in k/R — which at five coefficients on a window of forty is 12.5%, large enough that the first term should not be expected to carry all of it.

That is a useful calibration rather than a discrepancy. It says the mechanism is right, since a formula involving nothing but a count of coefficients and a window length lands within a third of a quantity nobody fitted; and it says the correction cannot be done from the formula, since a third of the shift would remain. The recentring works because it computes the shift from the forecasts themselves at each origin rather than from an asymptotic expression, and the gap between 0.076 and 0.108 is exactly how much that distinction is worth.

What to do with a nested comparison in the wild

Three readings of one comparison, and the arithmetic above says which is wanted when.

Which model should be run tomorrow. Include the estimation cost, because it will be paid. The ordinary comparison answers this and its verdict here — the smaller model — is right. What is not right is describing that verdict as evidence that the extra terms are worthless, which is the sentence such comparisons usually carry.

Whether the extra term carries information. Recentre. The estimation cost is a property of the sample size rather than of the world, and a question about the world should not be answered by a statistic that depends on it.

Whether a difference this size could have arisen by chance. Neither, on its own. Both statistics above are asymptotic, both are read against a normal, and the nested case is the one where that normal is hardest to justify. A reference distribution generated by resampling the comparison itself is the honest answer and it costs a bootstrap per comparison, which is why it is named here and not built.

What the recentring is really doing

The subtraction has an interpretation worth having, because it explains why the repair does not feel like one. Out-of-sample squared error is not an estimate of population predictive accuracy — it is an estimate of accuracy plus the cost of having estimated the model. Those are different quantities, and the larger model pays more of the second.

The recentred statistic removes the estimation cost from the comparison and tests the population quantity. That is a decision about what the comparison is for, and it is defensible in one direction and not in the other. A forecaster choosing which model to run tomorrow should include the estimation cost, because they will pay it; the ordinary comparison is right for them, and its verdict — use the smaller model — is correct. A forecaster asking whether the extra term carries information should exclude it, and then the recentring is what answers the question.

The two questions have different answers here, and that is not a paradox. It is the same distinction the order-selection field makes in a different currency: a criterion that overfits and a criterion that is right about the truth are not the same criterion, and which one is wanted depends on whether the model is going to be used or believed.

What each criterion selects, at 100 observations. 700 series from an AR(2) with coefficients 0.6 and -0.3, every order from 0 to 8 fitted to the same 92 responses so the log-likelihoods are comparable. AIC finds the true order 69.1% of the time and lands above it 26.6%; BIC finds it 81.0% and lands above it 2.3%. The closed form for one extra lag is P(χ²₁ > 2) = 15.73% for AIC, which does not depend on n at all, and P(χ²₁ > ln n) = 3.19% for BIC at this size, which falls to zero. Under the true order is the other failure and it is BIC's: 16.7% against 4.3%.
Fig. 7 The neighbouring machinery, from the field before this one. An information criterion also compares nested models, and it also charges the larger one for its own estimation — that is what its penalty term is. AIC’s willingness to overfit at a rate that does not fall with n is the same trade-off this essay’s two curves are two ends of.
The repaired test finds what the original cannot. An AR(1) forecast against an AR(2) forecast, with the second coefficient swept from zero — where the two models are the same model — to 0.4, where the larger one is genuinely better. Both curves are one-sided tests of the same hypothesis on the same forecasts. The ordinary statistic finds the larger model better 18.6% of the time at φ₂ = 0.2, which is its own nominal level and therefore no evidence at all, while the recentred one finds it 66.9% of the time. At the left both are honest: 3.5% and 0.1% under a true null.
Fig. 8 The power comparison on a longer estimation window. Both curves rise, because the larger model’s handicap shrinks; the gap between them narrows and does not close, and at the point where the two models are equally accurate in population the ordinary test is still finding the wrong one.

The connection to selection, which is not a coincidence

If the comparison between nested models is broken, so is anything built on it. Choosing a model by running the comparison and keeping the winner is model selection by a test whose level is not what it says, and the essays on order selection measure what that costs an interval afterwards: 95.0% at the true order and true parameters, 93.2% with the coefficients estimated, and 88.6% after AIC chose the order from the same data.

The two failures compound in the same direction. A nested comparison that systematically favours the smaller model, used to select, produces a fitted model that is too small; the interval computed from it is then too narrow for the additional reason that it does not know a selection happened.

What is claimed, and what is not

The claim is narrow and it is the one the phase’s brief named: nested forecast comparison as a testing problem, its degeneracy under the null, the direction of the failure, its behaviour as the evaluation lengthens, and the recentring that repairs it. Both statistics are computed from the same forecasts throughout, so nothing here is a comparison between studies.

What stays out: forecast encompassing tests, which ask whether one forecast contains all the information in another and are a different construction on the same data; bootstrap reference distributions for the nested case, which are the alternative to recentring and are more expensive than anything else in this field; and comparisons across a set of many models, where the multiplicity problem sits on top of everything above and the reference distribution has to be over the whole set.

The boundary against the model-selection half of the forecast field is that this essay compares two named models and that field chooses among a ladder of them by a criterion. They meet where a test is used as a criterion, which is the paragraph above, and nowhere else. The two essays that measure that field’s own costs are choosing the order and the interval after the choice.

A test between nested models, under a null that is true. 1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.1679 against 1.0601 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 94.3% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.3%.
Fig. 9 Four hundred origins, where the ordinary test declares the smaller model significantly better in nineteen comparisons out of twenty. Nothing about the models has changed from the first picture in this essay; only the number of origins the same difference is measured over.

The checks, and the refusal

Five claims are gated. Under a nested null the larger model’s mean squared error must exceed the smaller model’s, which is the fact everything else follows from. The ordinary test must declare the smaller model significantly better in more than 40% of comparisons at the settings drawn — a requirement that the failure be gross rather than marginal. It must do so more often as the evaluation lengthens, which is what separates this from a small-sample problem. The one-sided reading must reject essentially never. And the recentred statistic must hold its level and beat the statistic it repairs by more than thirty points of power against a real effect.

The refusal is the ordinary test applied to the nested pair and read against a standard normal, with the standard being that a test’s rejection rate under a true null is its stated level. It fails by a factor of thirteen, and the check requires it to fail, because an assertion that has never rejected anything proves nothing.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Clark–WestDiebold–MarianoError rateForecast errorLeast squaresLong-run varianceLoss differentialMean squared errorModel selectionMonte CarloNested modelsNull hypothesisOverfittingParameter uncertaintyStatistical power