When one model contains the other
Worth reading first: What the model says next · Choosing the order.
The comparison in the previous essay was between two forecasters with nothing in common: one carries the last value forward, the other averages the window, and neither is a special case of the other. That is not the comparison most forecasters actually want to make, and the machinery that essay built — the loss differential, the long-run variance, the exactly true null — is assumed here throughout.
The question that gets asked is whether adding something helps. Does a second lag improve the forecast? Does the extra variable earn its place? Every one of those is a comparison between a model and the same model with one more term in it, and the smaller model is what the larger one becomes when the extra coefficient is zero. The two forecasters are nested.
Under the null hypothesis that the extra term is worthless, the two forecasts are not merely equally accurate. They are, in population, the same forecast. And a test built to compare two different things has nothing left to divide by.
What is actually being tested, and why it is degenerate
Take an AR(1) truth, and forecast it two ways: with a fitted AR(1), and with a fitted AR(4) whose last three coefficients are estimating zeros. In population the two produce identical forecasts, so the population loss differential is zero — and so is its variance, because the two error sequences are the same sequence.
Nothing in a finite sample is degenerate, of course. What survives is entirely the estimation noise of the larger model: three coefficients that should be zero are estimated as something else, their contribution to the forecast is noise, and noise added to an unbiased forecast increases its squared error. So the sample loss differential is not centred at zero. It is centred below it.
The mean loss differential is −0.1080, and the ordinary test, applied as it always is, reports what it sees: the smaller model is significantly better, in 67.2% of comparisons, at a nominal 5% level.
That number is not a size distortion in the usual sense. The test is right that the smaller model forecast better in this sample, and right that it will do so again. What it is wrong about is the hypothesis a reader takes it to be testing — whether the extra lags carry information — and there the answer is no, they carry none, which is exactly the null the test claims to be examining.
The size of the shift is not mysterious either. A model estimating k coefficients from a window of R observations carries about k/R of the residual variance as estimation noise in its one-step forecast, and the extra lags contribute their share of that whether or not they are worth anything. Three useless lags on a window of forty is three fortieths of the variance handed to the larger model as a handicap, and the measured difference in mean squared error — 1.1663 against 1.0583, or ten percent — is what that arithmetic predicts once the estimated intercept and first lag are counted on both sides.
Read the way the question is asked, it never rejects
The two-sided reading above is charitable. Anybody comparing a model with its own extension has a direction in mind: does the larger model help? That is a one-sided test, and the statistic has to exceed +1.645 for the larger model to be declared better.
Across two thousand comparisons under a true null, the share exceeding it is 0.0%.
A test that never rejects is not a conservative test. It is a test whose rejection region the statistic cannot reach, and the reason is structural: under the null the statistic’s distribution is shifted left by the estimation noise, and under any alternative small enough to be interesting it is shifted left by almost as much. The test has no power in the region where a forecaster needs it, and it manufactures a finding in the region where nothing is happening.
More data makes it worse, which is the diagnostic
A test that fails from too little data is repaired by more. This one is not.
At four hundred origins the wrong answer arrives in nineteen comparisons out of twenty. The mean loss differential barely moves across that sweep — −0.1033, −0.1096, −0.1085, −0.1082 — because it is a real feature of the two procedures rather than a sampling accident. What changes is the standard error, which shrinks as the evaluation lengthens, so the same real difference is detected with increasing confidence.
A difference that is real and a hypothesis that is true are not in conflict here. The larger model really does forecast worse out of sample; the extra coefficients really are zero. Both statements are correct, and a test that reports the first as evidence against the second is answering a question nobody asked.
This is the winner’s curse seen from the other side of the same arithmetic. There, a quantity selected for being large is large by more than it should be; here, a model penalised for its own estimation noise is penalised reliably, and repeating the experiment does not average the penalty away because it is not noise in the comparison — it is a property of the pair being compared.
The repair is a subtraction
The fix follows from writing down where the shift comes from. The larger model’s forecast is the smaller model’s forecast plus estimation error; its squared error is therefore the smaller model’s squared error plus the square of that estimation error, minus a cross term with mean zero under the null. So the expected loss differential under the null is not zero but −E[(ŷ₁ − ŷ₂)²] — and that quantity is observable. Both forecasts are in hand at every origin.
Adding it back gives
fₜ = e₁ₜ² − [e₂ₜ² − (ŷ₁ₜ − ŷ₂ₜ)²]
which is centred at zero under the null by construction, and is read one-sided against a normal because the alternative only lives on one side. It is a subtraction rather than a new theory, and everything else about the test — the long-run variance, the bandwidth, the origins — is unchanged.
Its counted size is 4.9% with three extra lags and 4.0% with one, at a nominal 5%.
At φ₂ = 0.2 the larger model is genuinely the better forecaster — mean squared error 1.0980 against 1.1057 — and the ordinary test finds it 5.5% of the time, which is its own nominal level and therefore no evidence whatever. The recentred statistic finds it 47.0% of the time. At φ₂ = 0.3 the two are 35.3% and 87.8%.
The denominator cannot save it
Everything the previous essay was about — the overlap between neighbouring forecasts, the long-run variance, the bandwidth — is orthogonal to this failure and worth saying so explicitly, because the two problems are easy to conflate under the heading forecast comparison is harder than it looks.
The overlap problem is about the denominator: the mean difference is fine and its standard error is too small, so the statistic is inflated in whichever direction the difference happens to lie. It is symmetric, it is repaired by estimating the right variance, and at one step ahead it does not arise at all.
The nesting problem is about the numerator: the mean difference is not centred where the null says, and no denominator can move it. All three of the tests measured here use the same corrected long-run variance; they differ only in what is being averaged. Every comparison in this essay is at one step ahead, where the loss differentials are uncorrelated and the denominator is beyond reproach, and the failure is at its most complete.
Where the statistic actually sits, in standard errors
The four rejection rates along the evaluation sweep — 18.0%, 35.1%, 67.3% and 94.8% at fifty, one hundred, two hundred and four hundred origins — can be turned back into the quantity that produced them, which says more than the rates do.
A two-sided test at 5% rejects when the statistic passes −1.96, and essentially every rejection here is in that tail. Inverting a normal at each rate puts the statistic’s centre at about −1.04, −1.58, −2.41 and −3.59 standard errors. At four hundred origins the ordinary statistic sits three and a half standard errors below zero at a true null, which is a cleaner way of saying what “declares the smaller model significantly better nineteen times in twenty” means.
Those four numbers grow by roughly 1.5 at each doubling, against the √2 = 1.41 that a fixed mean difference divided by a standard error shrinking as 1/√P predicts. Agreeing to within seven per cent over an eightfold range is what says the mechanism is exactly the stated one: the difference is constant — the four mean differentials are −0.1033, −0.1096, −0.1085 and −0.1082 — and only the denominator is moving.
So the test’s failure has a growth rate, and the growth rate is the ordinary one. That is the part worth carrying past this essay. A defect that grows as √P is not distinguishable, from inside one evaluation, from a real effect being detected with increasing confidence; both look like a statistic drifting away from zero at the same rate. The only thing separating them is knowing where the null puts the statistic, and nothing in the output says.
It also prices how long an evaluation would have to be for the ordinary test to stop failing, which is a question with no answer: the shift grows without bound, so the failure gets worse for ever. The only value of P at which the ordinary statistic is honest here is a small one, and it is honest there by having too little power to be wrong.
Completing the k/R arithmetic
The handicap is attributed above to estimation noise at a rate of about k/R of the residual variance, and carrying that count through is worth doing because it says how much of the effect the simple version explains.
A one-step forecast from a model estimating k coefficients on a window of R carries about σ²(1 + k/R). The smaller model estimates an intercept and one lag, so k = 2 and its mean squared error of 1.0583 implies σ² ≈ 1.008. The larger estimates five, so the predicted gap is 3σ²/40 = 0.076, against a measured mean differential of 0.108.
The leading term is therefore about seven-tenths of the handicap, and the remaining three-tenths is the next order in k/R — which at five coefficients on a window of forty is 12.5%, large enough that the first term should not be expected to carry all of it.
That is a useful calibration rather than a discrepancy. It says the mechanism is right, since a formula involving nothing but a count of coefficients and a window length lands within a third of a quantity nobody fitted; and it says the correction cannot be done from the formula, since a third of the shift would remain. The recentring works because it computes the shift from the forecasts themselves at each origin rather than from an asymptotic expression, and the gap between 0.076 and 0.108 is exactly how much that distinction is worth.
What to do with a nested comparison in the wild
Three readings of one comparison, and the arithmetic above says which is wanted when.
Which model should be run tomorrow. Include the estimation cost, because it will be paid. The ordinary comparison answers this and its verdict here — the smaller model — is right. What is not right is describing that verdict as evidence that the extra terms are worthless, which is the sentence such comparisons usually carry.
Whether the extra term carries information. Recentre. The estimation cost is a property of the sample size rather than of the world, and a question about the world should not be answered by a statistic that depends on it.
Whether a difference this size could have arisen by chance. Neither, on its own. Both statistics above are asymptotic, both are read against a normal, and the nested case is the one where that normal is hardest to justify. A reference distribution generated by resampling the comparison itself is the honest answer and it costs a bootstrap per comparison, which is why it is named here and not built.
What the recentring is really doing
The subtraction has an interpretation worth having, because it explains why the repair does not feel like one. Out-of-sample squared error is not an estimate of population predictive accuracy — it is an estimate of accuracy plus the cost of having estimated the model. Those are different quantities, and the larger model pays more of the second.
The recentred statistic removes the estimation cost from the comparison and tests the population quantity. That is a decision about what the comparison is for, and it is defensible in one direction and not in the other. A forecaster choosing which model to run tomorrow should include the estimation cost, because they will pay it; the ordinary comparison is right for them, and its verdict — use the smaller model — is correct. A forecaster asking whether the extra term carries information should exclude it, and then the recentring is what answers the question.
The two questions have different answers here, and that is not a paradox. It is the same distinction the order-selection field makes in a different currency: a criterion that overfits and a criterion that is right about the truth are not the same criterion, and which one is wanted depends on whether the model is going to be used or believed.
The connection to selection, which is not a coincidence
If the comparison between nested models is broken, so is anything built on it. Choosing a model by running the comparison and keeping the winner is model selection by a test whose level is not what it says, and the essays on order selection measure what that costs an interval afterwards: 95.0% at the true order and true parameters, 93.2% with the coefficients estimated, and 88.6% after AIC chose the order from the same data.
The two failures compound in the same direction. A nested comparison that systematically favours the smaller model, used to select, produces a fitted model that is too small; the interval computed from it is then too narrow for the additional reason that it does not know a selection happened.
What is claimed, and what is not
The claim is narrow and it is the one the phase’s brief named: nested forecast comparison as a testing problem, its degeneracy under the null, the direction of the failure, its behaviour as the evaluation lengthens, and the recentring that repairs it. Both statistics are computed from the same forecasts throughout, so nothing here is a comparison between studies.
What stays out: forecast encompassing tests, which ask whether one forecast contains all the information in another and are a different construction on the same data; bootstrap reference distributions for the nested case, which are the alternative to recentring and are more expensive than anything else in this field; and comparisons across a set of many models, where the multiplicity problem sits on top of everything above and the reference distribution has to be over the whole set.
The boundary against the model-selection half of the forecast field is that this essay compares two named models and that field chooses among a ladder of them by a criterion. They meet where a test is used as a criterion, which is the paragraph above, and nowhere else. The two essays that measure that field’s own costs are choosing the order and the interval after the choice.
The checks, and the refusal
Five claims are gated. Under a nested null the larger model’s mean squared error must exceed the smaller model’s, which is the fact everything else follows from. The ordinary test must declare the smaller model significantly better in more than 40% of comparisons at the settings drawn — a requirement that the failure be gross rather than marginal. It must do so more often as the evaluation lengthens, which is what separates this from a small-sample problem. The one-sided reading must reject essentially never. And the recentred statistic must hold its level and beat the statistic it repairs by more than thirty points of power against a real effect.
The refusal is the ordinary test applied to the nested pair and read against a standard normal, with the standard being that a test’s rejection rate under a true null is its stated level. It fails by a factor of thirteen, and the check requires it to fail, because an assertion that has never rejected anything proves nothing.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A line that beats two curves — both name least squares, model selection, monte carlo, nested models, overfitting, parameter uncertainty
- The displacement is a parameter count — both name least squares, mean squared error, model selection, monte carlo, nested models, overfitting
- When the benchmark is a candidate — both name error rate, mean squared error, model selection, monte carlo, nested models, null hypothesis
- A criterion is a prediction of the hold-out — both name mean squared error, model selection, monte carlo, nested models, overfitting
- Where the two searches cross — both name mean squared error, model selection, monte carlo, nested models, overfitting
- A penalty is a trace — both name least squares, mean squared error, model selection, overfitting
Named objects
A flat tag is an object no other essay names yet.
Clark–WestDiebold–MarianoError rateForecast errorLeast squaresLong-run varianceLoss differentialMean squared errorModel selectionMonte CarloNested modelsNull hypothesisOverfittingParameter uncertaintyStatistical power