The best of a set, and what the search costs

What the other forecast adds

Two forecasters, one series, and two different questions about them. Which is more accurate has an answer that changes with the persistence of the series; whether either is redundant has an answer that never changes at all.

Worth reading first: What the model says next · The observations that repeat each other.

The evaluation field takes two forecasters, forms the difference in their squared errors, and asks whether its average is far enough from zero to be worth reporting. That is one question about the pair, and the field’s own last paragraph names the other one and leaves it: forecast encompassing, which asks not which forecast is better but whether either of them is redundant.

The two questions have different answers on the same data, and the gap between them is not a technicality about tests. It is arithmetic, and it can be written down before any data exists.

Three numbers, and every question is about them

Two forecasts of the same target produce two error sequences. Everything either question can ask is a statement about three second moments: σ₁², σ₂² and σ₁₂, the two error variances and their covariance. Write the two hypotheses out against those three numbers.

Equal accuracy is σ₁² − σ₂² = 0.

The first forecast encompasses the second is σ₁² − σ₁₂ ≤ 0.

They are different linear combinations of the same three quantities, so knowing one says nothing about the other. Two forecasts can be exactly equally accurate with neither redundant; one can be much worse and still carry something the better one does not.

Three quantities, and only one of them crosses zero. Two forecasts of an AR(1) — the last value carried forward and the mean of the last 60 observations — at 1 step ahead. The curve through zero is σ₁² − σ₂², the difference in expected squared error that a comparison of accuracy tests; it changes sign at φ = 0.4922. The two curves above it are σ₁² − σ₁₂ and σ₂² − σ₁₂, the quantities the two encompassing tests are about, and neither of them comes near zero anywhere: the smallest value either takes across the range is 0.008 times the variance of the series. All three are closed forms in φ, R and h with no simulation in them. Equal accuracy is one hypothesis about this picture and encompassing is another, and a set of numbers can satisfy either without the other.
Fig. 1 The two benchmark forecasts of a first-order autoregression: the last value carried forward, and the mean of the last sixty observations. Three quantities plotted across the persistence of the series, in units of the series’ own variance. One of them crosses zero.

The pair drawn there is the pair the previous field is built on, for the same reason: their moments are closed forms rather than estimates. A moving average of L observations of an AR(1) has an error covariance that is a sum of powers of φ, so σ₁², σ₂² and σ₁₂ can be written out exactly and every crossing below is solved rather than searched for. The whole family — L = 1 is the last value, L = 60 is the window mean — comes from one covariance matrix, which is worth saying because the field’s later essays put eight members of that family on one axis.

The quantity underneath both questions

There is an object that both hypotheses are statements about, and reaching for it makes the rest obvious. Combine the two forecasts: f = w·f₁ + (1 − w)f₂. The combined error is w·e₁ + (1 − w)e₂, its variance is a quadratic in w, and the minimum is at

w* = (σ₂² − σ₁₂)/(σ₁² + σ₂² − 2σ₁₂)

which is the weight the combination puts on the first forecast. Now read the two hypotheses again. Encompassing is the statement that w* is 0 or 1 — that the combination puts no weight on one of them, which is what “adds nothing” means. Equal accuracy is a statement about a different point of the same axis, and the point is easy to identify: when σ₁² = σ₂² the expression is symmetric in the two forecasts and w* is exactly ½.

The weight that encompassing is a question about. The variance-minimising weight on the last value in a combination of it and the 60-observation window mean, at 1 step ahead, computed from the closed-form error covariance rather than estimated. It runs from 0.020 at φ = 0.02 to 0.984 at 0.98 and is interior at every point between: neither forecast is ever redundant. Where it crosses ½ — φ = 0.4922, the marked line — is exactly where the two have equal expected squared error, because equal variances give equal weights and for no other reason. Forecast encompassing is the hypothesis that this weight is 0 or 1; equal accuracy is the hypothesis that it is ½. They are different points of the same axis.
Fig. 2 The same pair, drawn as the weight instead. It crosses ½ exactly where the previous figure’s accuracy difference crosses zero, and for no other reason than that equal variances give equal weights.

So the two hypotheses are not rival answers to one question. They are two points of one interval: 0 and 1 at the ends, ½ in the middle, and everything a comparison can say about a pair of forecasts is a statement about where in that interval the weight sits.

Where the two questions part company, measured

At φ = 0.4922 — the persistence the previous field solves for, reached here by a completely different arrangement of the same equation and agreeing to nine decimals — the two forecasts have identical expected squared error. The accuracy question has the answer “no difference”. The encompassing question has the answer 0.4911 times the variance of the series, in both directions: neither forecast is close to redundant, and the margin is nearly half of everything the series does.

The combination is what makes that concrete. At the crossing, half and half of two forecasts that are individually worth 1.0156 times the series’ variance produces an error variance of 0.7700 — a 24.2% reduction in mean squared error, bought by averaging two things a comparison has just declared indistinguishable.

Three quantities, and only one of them crosses zero. Two forecasts of an AR(1) — the last value carried forward and the mean of the last 60 observations — at 4 steps ahead. The curve through zero is σ₁² − σ₂², the difference in expected squared error that a comparison of accuracy tests; it changes sign at φ = 0.8256. The two curves above it are σ₁² − σ₁₂ and σ₂² − σ₁₂, the quantities the two encompassing tests are about, and neither of them comes near zero anywhere: the smallest value either takes across the range is 0.000 times the variance of the series. All three are closed forms in φ, R and h with no simulation in them. Equal accuracy is one hypothesis about this picture and encompassing is another, and a set of numbers can satisfy either without the other.
Fig. 3 The same three quantities four steps ahead. The accuracy difference crosses zero at a different place — φ = 0.8256 — and the two encompassing quantities move hardly at all. Which forecast is more accurate depends on the horizon; whether either is redundant does not.

And the weight never reaches either end. Across persistence from 0.02 to 0.98 it runs from 0.0203 to 0.9841, and the smallest either encompassing quantity gets anywhere in that range is 0.0083 times the variance of the series. In population, on this pair, neither forecast ever encompasses the other. That is not a near miss to be rounded off: it says the combination is always worth forming, at every persistence, which is the forecast field’s finding that an equal-weight average beats both its parts arriving as a theorem rather than as a measurement.

The weight that encompassing is a question about. The variance-minimising weight on the last value in a combination of it and the 60-observation window mean, at 4 steps ahead, computed from the closed-form error covariance rather than estimated. It runs from 0.000 at φ = 0.02 to 0.938 at 0.98 and is interior at every point between: neither forecast is ever redundant. Where it crosses ½ — φ = 0.8256, the marked line — is exactly where the two have equal expected squared error, because equal variances give equal weights and for no other reason. Forecast encompassing is the hypothesis that this weight is 0 or 1; equal accuracy is the hypothesis that it is ½. They are different points of the same axis.
Fig. 4 The weight at four steps ahead. It still runs the whole interval and still touches neither end, and it still passes ½ exactly at the accuracy crossing — which is now at a much higher persistence.

The question a forecaster is actually asked

The two hypotheses are not equally often the one that matters, and it is worth saying which is which before measuring either. An organisation with a forecast in production and an offer of a second one is not choosing between them. It is asking whether the second is worth adding, and that is the encompassing question exactly: if w* is 0 the new forecast can be declined without loss, and if it is anything else the answer is a combination rather than a winner.

Ranking by mean squared error answers a question nobody asked — which single forecast to keep if only one may be kept — and the answer it gives is not even a bound on the other. At the crossing here the ranking is a coin toss and the correct action is to use both, in equal parts, for a quarter less squared error. At φ = 0.9 the ranking is decisive, the last value wins by a wide margin, and the correct action is still to use both, because the weight there is 0.9107 and not 1.

That asymmetry is why the comparison literature has two tests in it rather than one, and why the distinction is worth a section rather than a sentence: they are not competing procedures for one decision, they are procedures for two decisions, and only one of the two is usually being made.

The two questions are linked by a factor of two

At the accuracy crossing the two questions give answers that look unrelated — “no difference” and “0.4911 of the series’ variance” — and they are not unrelated at all. One is exactly twice the other.

Where σ₁² = σ₂² = σ², the optimal weight is a half and the combined error variance is (σ² + σ₁₂)/2, so the reduction from combining is

σ² − (σ² + σ₁₂)/2 = (σ² − σ₁₂)/2

which is half the encompassing quantity. The numbers check to the last digit: 1.0156 − 0.7700 = 0.2456, and 0.4911/2 = 0.2456. The encompassing margin is the gain from combining, doubled.

That turns the essay’s central contrast into an identity rather than a coincidence of one setting. Two forecasts declared indistinguishable by a comparison of accuracy are worth combining by exactly half of what the encompassing test is measuring, so a large encompassing statistic and a worthwhile combination are the same finding reported in two units.

Divided through by σ² it says more. The relative gain at equal accuracy is (1 − ρ)/2 with ρ the correlation between the two error series, so the crossing’s 24.2% implies ρ = 0.516 and nothing else has to be measured to get it. The expression also fixes the ceiling: uncorrelated errors give a gain of exactly a half, which is the familiar statement that averaging two independent estimates halves a variance, arriving here as the ρ = 0 case of a formula whose whole content is the correlation.

So the practical reading of an encompassing statistic is available without any further test. Whatever it reports, in units of the forecasts’ own variance, is twice the fraction of squared error a fifty-fifty combination will remove — and the reason to prefer it to a ranking is that a ranking has no units at all.

The horizons make the same point from the other side. The encompassing quantities are 0.4911, 0.4896, 0.4841 and 0.4659 at one, two, four and eight steps — a fall of 5% — while the persistence at which the two forecasts are equally accurate moves from 0.49 to 0.90. The combination is worth about a quarter of the squared error at every horizon; which forecast to keep, if only one could be kept, is a different answer at each.

How far the weight is from a boundary

The estimated weight never left [0, 1] in fifteen hundred comparisons, and the phrase “never even close” can be given a number.

At a mean of 0.4778 and a standard deviation of 0.1123, the two boundaries sit 4.25 and 4.65 standard deviations away. On a normal that is about two chances in a hundred thousand per side, so the expected number of boundary cases in fifteen hundred comparisons is 0.06 — the observed zero is what the arithmetic predicts rather than a small-sample accident.

That is the encompassing hypothesis measured as a distance rather than as a verdict, and the distance is the useful form. A test reports rejected or not; four and a quarter standard errors says the null is not a near-miss description of this pair at this length, and would still not be at a quarter of the origins, since the distance falls only as √P.

The test, and a null that is true by construction

Measuring what a test does when it should not reject requires a situation where it should not, and the encompassing hypothesis supplies one that needs nothing arranged. The optimally combined forecast encompasses each of its parts by the first-order condition that defines the weight. Differentiating the combined variance and setting it to zero is the statement cov(ece_c, e₁ − e₂) = 0, which is the encompassing null written out. So a forecast built at the population weight has an exactly true encompassing null against each of the two forecasts it is made of, and every rejection counted against it is a false one.

The statistic is the one the previous field already has, with a different differential fed into it. Where an accuracy comparison uses dₜ = e₁ₜ² − e₂ₜ², an encompassing comparison uses

cₜ = e₁ₜ(e₁ₜ − e₂ₜ)

whose mean is σ₁² − σ₁₂. Everything the previous field established about the denominator applies unchanged: at horizon h the terms are a moving average of order h − 1 for the same reason, and a standard error that ignores that is too small by the same computable factor — which is the dependence field’s effective sample size arriving in a sequence of products rather than a sequence of observations.

What the two tests say about the same sixty origins

Run both on the same comparisons, at the persistence where the two forecasters are exactly equally accurate.

The two tests, run on the same comparisons. Six persistences, 350 comparisons at each, and two tests on every one of them, at 1 step ahead. The curve that dips is the accuracy comparison: at φ = 0.4922 the two forecasts have identical expected squared error and it is measuring its own size, 7.1%, rising away from there because the difference is real. The other two are the encompassing tests in each direction, and at that same crossing they reject 100.0% and 98.0%. Neither forecast encompasses the other at any persistence — that is a closed form with no simulation in it — so every number on those two curves below 100% is the test running out of power rather than the hypothesis changing, and at long horizons there is a great deal of that: sixty origins at 1 step ahead carry about 60 independent comparisons.
Fig. 5 Six persistences, five hundred comparisons of sixty origins at each, and two tests on every one of them. The lower curve is measuring its own size at the marked crossing; the upper pair are not measuring size anywhere.

At the crossing the accuracy test rejects 7.3% of the time at a nominal 5%, which is the over-rejection the previous field attributes to estimating a long-run variance that is not needed one step ahead. The encompassing tests reject 100.0% and 97.5%. A reader given the first result alone would report that the two forecasters cannot be told apart. Both of them are carrying something the other does not, decisively, on the same sixty numbers.

The two tests, run on the same comparisons. Six persistences, 350 comparisons at each, and two tests on every one of them, at 2 steps ahead. The curve that dips is the accuracy comparison: at φ = 0.6944 the two forecasts have identical expected squared error and it is measuring its own size, 9.1%, rising away from there because the difference is real. The other two are the encompassing tests in each direction, and at that same crossing they reject 99.7% and 89.1%. Neither forecast encompasses the other at any persistence — that is a closed form with no simulation in it — so every number on those two curves below 100% is the test running out of power rather than the hypothesis changing, and at long horizons there is a great deal of that: sixty origins at 2 steps ahead carry about 30 independent comparisons.
Fig. 6 Two steps ahead, where the accuracy crossing has moved to φ = 0.6944 and the encompassing verdict has not moved at all. The two curves are answering different questions and only one of them has a zero to cross.

Where the encompassing test’s own failure lives

A test whose null is exactly true is a test whose size can be measured rather than approximated, so the natural next question is whether this one has the level it claims. It does not, and the failure is worth separating into its two possible causes: the statistic, or the denominator it is divided by.

The site’s standing instrument for that separation is the oracle row — the same statistic handed the true variance of its own mean, which is not a test anybody can run because it needs the answer.

An encompassing test at a null that is exactly true. The forecast being tested is the variance-minimising combination of the last value and the 60-observation window mean, which encompasses each of its parts by the first-order condition that defines the weight — so every rejection counted here is a false one. Four rejection rates at a nominal 5%, over 1,800 comparisons of 120 origins each. Given the true variance of its own mean the statistic is at 3.7% in the upper tail and 5.8% in the lower. With that variance estimated from the same 120 numbers it is at 7.9% and 3.7%: the same total, moved from one side to the other. The standardised statistic's skewness is -0.355, and a denominator correlated with the numerator it divides is what puts it there.
Fig. 7 Four rejection rates at a null that is exactly true. The top pair use the true variance of the statistic’s own mean; the bottom pair estimate it from the same hundred and twenty numbers. The total is nearly the same and it has moved from one side to the other.

Given the true variance, the statistic rejects 3.9% in the upper tail and 5.5% in the lower — 5% to within the noise of three thousand comparisons, in both directions. With the variance estimated from the same data, the same statistic rejects 8.4% upward and 3.6% downward.

The total has hardly moved. What has happened is that the studentised statistic has acquired a skewness of −0.354, and the reason is that the numerator and the denominator are estimated from the same numbers and are correlated: cₜ is a product of two errors, so a sample whose mean is large is a sample whose estimated variance is small, and the ratio is too large more often than it should be on one side and too small on the other.

An encompassing test at a null that is exactly trueThe forecast being tested is the variance-minimising combination of the last value and the 60-observation window mean, which encompasses each of its parts by the first-order condition that defines the weight — so every rejection counted here is a false one. Four rejection rates at a nominal 5%, over 1,800 comparisons of 480 origins each. Given the true variance of its own mean the statistic is at 5.2% in the upper tail and 4.9% in the lower. With that variance estimated from the same 480 numbers it is at 8.5% and 4.4%: the same total, moved from one side to the other. The standardised statistic's skewness is -0.137, and a denominator correlated with the numerator it divides is what puts it there.true variance of the mean, upper tail5.2%true variance of the mean, lower tail4.9%estimated from the sample, upper tail8.5%estimated from the sample, lower tail4.4%nominal 5%the variance of the mean, and where it came from1,800 comparisons of 480 originsφ = 0.4922: every rejection is false
Fig. 8 The same four rates as the comparison lengthens. Drag it: the gap between the two tails narrows and does not close, because the bandwidth the long-run variance is estimated over grows with the sample that estimates it.

That is a defect of a shape this site has met before and repaired in only one way: not by a better formula for the denominator, but by reading the statistic against a distribution generated from the null instead of against a normal. The last essay of this field does exactly that for a harder case, and the repair works here for the same reason.

What the horizon does, and what it does not

The accuracy crossing moves a long way with the horizon: 0.4922 at one step, 0.6944 at two, 0.8256 at four and 0.9021 at eight. Those are the four numbers the evaluation field measures its rejection rates at, and they are recovered here from the moment matrix rather than from that field’s fixed-point iteration — two arrangements of one equation, agreeing to nine decimals, which is the check that says the closed forms are the closed forms of the same thing.

At every one of them the optimal weight is exactly ½, which is not a coincidence and is not a property of the horizon: it follows from σ₁² = σ₂², which is what the crossing is defined by. The horizon changes where the two forecasts happen to be equally accurate and changes nothing about what equal accuracy implies.

The encompassing quantities barely move at all — 0.4911, 0.4896, 0.4841 and 0.4659 times the series’ variance at the four horizons. So the answer to “which is better” is a strong function of something the analyst chooses, and the answer to “is either redundant” is nearly invariant to it. Of the two, the second is the more stable description of the pair, and it is the one that never gets reported.

The weight itself, estimated

If the whole comparison is a statement about w*, the honest way to report it is to estimate w* and say what its uncertainty is, rather than to test two of its points and report which test rejected. On sixty origins at the crossing, the estimated weight averages 0.4778 with a standard deviation of 0.1123, and in fifteen hundred comparisons it fell outside the interval [0, 1] 0.0% of the time — so on this pair, at this length, the estimate is never at a boundary and the encompassing hypothesis is never even close to being the best description of the data.

The combination is worth 16.5% of the better part’s mean squared error in a typical sample of sixty origins, against the 24.2% its population weight would deliver — the difference being what estimating one number out of sixty pairs of errors costs.

The weight that encompassing is a question about. The variance-minimising weight on the last value in a combination of it and the 60-observation window mean, at 8 steps ahead, computed from the closed-form error covariance rather than estimated. It runs from 0.000 at φ = 0.02 to 0.882 at 0.98 and is interior at every point between: neither forecast is ever redundant. Where it crosses ½ — φ = 0.9021, the marked line — is exactly where the two have equal expected squared error, because equal variances give equal weights and for no other reason. Forecast encompassing is the hypothesis that this weight is 0 or 1; equal accuracy is the hypothesis that it is ½. They are different points of the same axis.
Fig. 9 Eight steps ahead. The accuracy crossing has moved to φ = 0.9021 and the weight still runs the full interval — it is only the horizon at which two forecasts happen to be equally good that changes.

What is claimed here, and what is not

This field takes the comparison of a set of forecasters, and this essay takes the pair case of it: encompassing, the combination weight, and the fact that they are one object. The evaluation field named encompassing as explicitly not claimed when it was written, and it is claimed now.

What stays out and is named as a decision: forecast combination as a subject in its own right — how weights should be estimated, shrunk or constrained, which is a large literature and a different argument; encompassing tests for more than two forecasters at once, where the same multiplicity problem the next essay is about arrives in a second form; and loss functions other than squared error, where the three moments this essay is built on stop being the sufficient description.

The boundary against the evaluation field is that field’s own statistic: everything about the denominator, the horizon and the dependence between origins is established there and used here without being re-derived.

The checks, and the refusals that make them mean something

Four claims are gated in this field’s library. The closed-form covariance matrix is compared with a rolling comparison that was never told the formula, at four candidates and five hundred origins, and the crossing it implies is required to match the previous field’s fixed point to nine decimals. The optimal weight at the accuracy crossing is required to be exactly ½. The two encompassing quantities are required to stay strictly positive at every persistence drawn, which is the statement that neither forecast is ever redundant. And the encompassing test at the combination’s exactly true null is required to be at its level with the oracle denominator and away from it with the estimated one — because if the first of those failed, the second would be measuring the setup rather than the test.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationBenchmark forecastCombination weightDiebold–MarianoError rateForecast combinationForecast encompassingForecast horizonLong-run varianceLoss differentialMean squared errorMonte CarloNull hypothesisSkewness