What the other forecast adds
Worth reading first: What the model says next · The observations that repeat each other.
The evaluation field takes two forecasters, forms the difference in their squared errors, and asks whether its average is far enough from zero to be worth reporting. That is one question about the pair, and the field’s own last paragraph names the other one and leaves it: forecast encompassing, which asks not which forecast is better but whether either of them is redundant.
The two questions have different answers on the same data, and the gap between them is not a technicality about tests. It is arithmetic, and it can be written down before any data exists.
Three numbers, and every question is about them
Two forecasts of the same target produce two error sequences. Everything either question can ask is a statement about three second moments: σ₁², σ₂² and σ₁₂, the two error variances and their covariance. Write the two hypotheses out against those three numbers.
Equal accuracy is σ₁² − σ₂² = 0.
The first forecast encompasses the second is σ₁² − σ₁₂ ≤ 0.
They are different linear combinations of the same three quantities, so knowing one says nothing about the other. Two forecasts can be exactly equally accurate with neither redundant; one can be much worse and still carry something the better one does not.
The pair drawn there is the pair the previous field is built on, for the same reason: their moments are closed forms rather than estimates. A moving average of L observations of an AR(1) has an error covariance that is a sum of powers of φ, so σ₁², σ₂² and σ₁₂ can be written out exactly and every crossing below is solved rather than searched for. The whole family — L = 1 is the last value, L = 60 is the window mean — comes from one covariance matrix, which is worth saying because the field’s later essays put eight members of that family on one axis.
The quantity underneath both questions
There is an object that both hypotheses are statements about, and reaching for it makes the rest obvious. Combine the two forecasts: f = w·f₁ + (1 − w)f₂. The combined error is w·e₁ + (1 − w)e₂, its variance is a quadratic in w, and the minimum is at
w* = (σ₂² − σ₁₂)/(σ₁² + σ₂² − 2σ₁₂)
which is the weight the combination puts on the first forecast. Now read the two hypotheses again. Encompassing is the statement that w* is 0 or 1 — that the combination puts no weight on one of them, which is what “adds nothing” means. Equal accuracy is a statement about a different point of the same axis, and the point is easy to identify: when σ₁² = σ₂² the expression is symmetric in the two forecasts and w* is exactly ½.
So the two hypotheses are not rival answers to one question. They are two points of one interval: 0 and 1 at the ends, ½ in the middle, and everything a comparison can say about a pair of forecasts is a statement about where in that interval the weight sits.
Where the two questions part company, measured
At φ = 0.4922 — the persistence the previous field solves for, reached here by a completely different arrangement of the same equation and agreeing to nine decimals — the two forecasts have identical expected squared error. The accuracy question has the answer “no difference”. The encompassing question has the answer 0.4911 times the variance of the series, in both directions: neither forecast is close to redundant, and the margin is nearly half of everything the series does.
The combination is what makes that concrete. At the crossing, half and half of two forecasts that are individually worth 1.0156 times the series’ variance produces an error variance of 0.7700 — a 24.2% reduction in mean squared error, bought by averaging two things a comparison has just declared indistinguishable.
And the weight never reaches either end. Across persistence from 0.02 to 0.98 it runs from 0.0203 to 0.9841, and the smallest either encompassing quantity gets anywhere in that range is 0.0083 times the variance of the series. In population, on this pair, neither forecast ever encompasses the other. That is not a near miss to be rounded off: it says the combination is always worth forming, at every persistence, which is the forecast field’s finding that an equal-weight average beats both its parts arriving as a theorem rather than as a measurement.
The question a forecaster is actually asked
The two hypotheses are not equally often the one that matters, and it is worth saying which is which before measuring either. An organisation with a forecast in production and an offer of a second one is not choosing between them. It is asking whether the second is worth adding, and that is the encompassing question exactly: if w* is 0 the new forecast can be declined without loss, and if it is anything else the answer is a combination rather than a winner.
Ranking by mean squared error answers a question nobody asked — which single forecast to keep if only one may be kept — and the answer it gives is not even a bound on the other. At the crossing here the ranking is a coin toss and the correct action is to use both, in equal parts, for a quarter less squared error. At φ = 0.9 the ranking is decisive, the last value wins by a wide margin, and the correct action is still to use both, because the weight there is 0.9107 and not 1.
That asymmetry is why the comparison literature has two tests in it rather than one, and why the distinction is worth a section rather than a sentence: they are not competing procedures for one decision, they are procedures for two decisions, and only one of the two is usually being made.
The two questions are linked by a factor of two
At the accuracy crossing the two questions give answers that look unrelated — “no difference” and “0.4911 of the series’ variance” — and they are not unrelated at all. One is exactly twice the other.
Where σ₁² = σ₂² = σ², the optimal weight is a half and the combined error variance is (σ² + σ₁₂)/2, so the reduction from combining is
σ² − (σ² + σ₁₂)/2 = (σ² − σ₁₂)/2
which is half the encompassing quantity. The numbers check to the last digit: 1.0156 − 0.7700 = 0.2456, and 0.4911/2 = 0.2456. The encompassing margin is the gain from combining, doubled.
That turns the essay’s central contrast into an identity rather than a coincidence of one setting. Two forecasts declared indistinguishable by a comparison of accuracy are worth combining by exactly half of what the encompassing test is measuring, so a large encompassing statistic and a worthwhile combination are the same finding reported in two units.
Divided through by σ² it says more. The relative gain at equal accuracy is (1 − ρ)/2 with ρ the correlation between the two error series, so the crossing’s 24.2% implies ρ = 0.516 and nothing else has to be measured to get it. The expression also fixes the ceiling: uncorrelated errors give a gain of exactly a half, which is the familiar statement that averaging two independent estimates halves a variance, arriving here as the ρ = 0 case of a formula whose whole content is the correlation.
So the practical reading of an encompassing statistic is available without any further test. Whatever it reports, in units of the forecasts’ own variance, is twice the fraction of squared error a fifty-fifty combination will remove — and the reason to prefer it to a ranking is that a ranking has no units at all.
The horizons make the same point from the other side. The encompassing quantities are 0.4911, 0.4896, 0.4841 and 0.4659 at one, two, four and eight steps — a fall of 5% — while the persistence at which the two forecasts are equally accurate moves from 0.49 to 0.90. The combination is worth about a quarter of the squared error at every horizon; which forecast to keep, if only one could be kept, is a different answer at each.
How far the weight is from a boundary
The estimated weight never left [0, 1] in fifteen hundred comparisons, and the phrase “never even close” can be given a number.
At a mean of 0.4778 and a standard deviation of 0.1123, the two boundaries sit 4.25 and 4.65 standard deviations away. On a normal that is about two chances in a hundred thousand per side, so the expected number of boundary cases in fifteen hundred comparisons is 0.06 — the observed zero is what the arithmetic predicts rather than a small-sample accident.
That is the encompassing hypothesis measured as a distance rather than as a verdict, and the distance is the useful form. A test reports rejected or not; four and a quarter standard errors says the null is not a near-miss description of this pair at this length, and would still not be at a quarter of the origins, since the distance falls only as √P.
The test, and a null that is true by construction
Measuring what a test does when it should not reject requires a situation where it should not, and the encompassing hypothesis supplies one that needs nothing arranged. The optimally combined forecast encompasses each of its parts by the first-order condition that defines the weight. Differentiating the combined variance and setting it to zero is the statement cov(, e₁ − e₂) = 0, which is the encompassing null written out. So a forecast built at the population weight has an exactly true encompassing null against each of the two forecasts it is made of, and every rejection counted against it is a false one.
The statistic is the one the previous field already has, with a different differential fed into it. Where an accuracy comparison uses dₜ = e₁ₜ² − e₂ₜ², an encompassing comparison uses
cₜ = e₁ₜ(e₁ₜ − e₂ₜ)
whose mean is σ₁² − σ₁₂. Everything the previous field established about the denominator applies unchanged: at horizon h the terms are a moving average of order h − 1 for the same reason, and a standard error that ignores that is too small by the same computable factor — which is the dependence field’s effective sample size arriving in a sequence of products rather than a sequence of observations.
What the two tests say about the same sixty origins
Run both on the same comparisons, at the persistence where the two forecasters are exactly equally accurate.
At the crossing the accuracy test rejects 7.3% of the time at a nominal 5%, which is the over-rejection the previous field attributes to estimating a long-run variance that is not needed one step ahead. The encompassing tests reject 100.0% and 97.5%. A reader given the first result alone would report that the two forecasters cannot be told apart. Both of them are carrying something the other does not, decisively, on the same sixty numbers.
Where the encompassing test’s own failure lives
A test whose null is exactly true is a test whose size can be measured rather than approximated, so the natural next question is whether this one has the level it claims. It does not, and the failure is worth separating into its two possible causes: the statistic, or the denominator it is divided by.
The site’s standing instrument for that separation is the oracle row — the same statistic handed the true variance of its own mean, which is not a test anybody can run because it needs the answer.
Given the true variance, the statistic rejects 3.9% in the upper tail and 5.5% in the lower — 5% to within the noise of three thousand comparisons, in both directions. With the variance estimated from the same data, the same statistic rejects 8.4% upward and 3.6% downward.
The total has hardly moved. What has happened is that the studentised statistic has acquired a skewness of −0.354, and the reason is that the numerator and the denominator are estimated from the same numbers and are correlated: cₜ is a product of two errors, so a sample whose mean is large is a sample whose estimated variance is small, and the ratio is too large more often than it should be on one side and too small on the other.
That is a defect of a shape this site has met before and repaired in only one way: not by a better formula for the denominator, but by reading the statistic against a distribution generated from the null instead of against a normal. The last essay of this field does exactly that for a harder case, and the repair works here for the same reason.
What the horizon does, and what it does not
The accuracy crossing moves a long way with the horizon: 0.4922 at one step, 0.6944 at two, 0.8256 at four and 0.9021 at eight. Those are the four numbers the evaluation field measures its rejection rates at, and they are recovered here from the moment matrix rather than from that field’s fixed-point iteration — two arrangements of one equation, agreeing to nine decimals, which is the check that says the closed forms are the closed forms of the same thing.
At every one of them the optimal weight is exactly ½, which is not a coincidence and is not a property of the horizon: it follows from σ₁² = σ₂², which is what the crossing is defined by. The horizon changes where the two forecasts happen to be equally accurate and changes nothing about what equal accuracy implies.
The encompassing quantities barely move at all — 0.4911, 0.4896, 0.4841 and 0.4659 times the series’ variance at the four horizons. So the answer to “which is better” is a strong function of something the analyst chooses, and the answer to “is either redundant” is nearly invariant to it. Of the two, the second is the more stable description of the pair, and it is the one that never gets reported.
The weight itself, estimated
If the whole comparison is a statement about w*, the honest way to report it is to estimate w* and say what its uncertainty is, rather than to test two of its points and report which test rejected. On sixty origins at the crossing, the estimated weight averages 0.4778 with a standard deviation of 0.1123, and in fifteen hundred comparisons it fell outside the interval [0, 1] 0.0% of the time — so on this pair, at this length, the estimate is never at a boundary and the encompassing hypothesis is never even close to being the best description of the data.
The combination is worth 16.5% of the better part’s mean squared error in a typical sample of sixty origins, against the 24.2% its population weight would deliver — the difference being what estimating one number out of sixty pairs of errors costs.
What is claimed here, and what is not
This field takes the comparison of a set of forecasters, and this essay takes the pair case of it: encompassing, the combination weight, and the fact that they are one object. The evaluation field named encompassing as explicitly not claimed when it was written, and it is claimed now.
What stays out and is named as a decision: forecast combination as a subject in its own right — how weights should be estimated, shrunk or constrained, which is a large literature and a different argument; encompassing tests for more than two forecasters at once, where the same multiplicity problem the next essay is about arrives in a second form; and loss functions other than squared error, where the three moments this essay is built on stop being the sufficient description.
The boundary against the evaluation field is that field’s own statistic: everything about the denominator, the horizon and the dependence between origins is established there and used here without being re-derived.
The checks, and the refusals that make them mean something
Four claims are gated in this field’s library. The closed-form covariance matrix is compared with a rolling comparison that was never told the formula, at four candidates and five hundred origins, and the crossing it implies is required to match the previous field’s fixed point to nine decimals. The optimal weight at the accuracy crossing is required to be exactly ½. The two encompassing quantities are required to stay strictly positive at every persistence drawn, which is the statement that neither forecast is ever redundant. And the encompassing test at the combination’s exactly true null is required to be at its level with the oracle denominator and away from it with the estimated one — because if the first of those failed, the second would be measuring the setup rather than the test.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- When the benchmark is a candidate — both name benchmark forecast, error rate, mean squared error, monte carlo, null hypothesis
- Correcting the forecast instead — both name autocorrelation, forecast horizon, mean squared error, monte carlo
- How long a block a multiplier shares — both name autocorrelation, error rate, monte carlo, null hypothesis
- The corner the test is calibrated at — both name benchmark forecast, error rate, monte carlo, null hypothesis
- The correction that leaves the region — both name autocorrelation, forecast horizon, mean squared error, monte carlo
- The gap a sample shows — both name autocorrelation, long-run variance, mean squared error, monte carlo
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationBenchmark forecastCombination weightDiebold–MarianoError rateForecast combinationForecast encompassingForecast horizonLong-run varianceLoss differentialMean squared errorMonte CarloNull hypothesisSkewness