Searching among fitted models

The weight that is a vector

Two forecasts have a best combination and one number describes it. Eight have a best combination too, and the vector describing it puts nothing at all on the forecast with the smallest mean squared error.

Worth reading first: Three series and a count · What the model says next.

Comparing two forecasts turns out to be a statement about a single number. The variance-minimising combination f = w·f₁ + (1 − w)f₂ has a weight w*, equal accuracy is the hypothesis that it is ½ and encompassing is the hypothesis that it is 0 or 1, and everything either test can say is a statement about where in the unit interval that weight sits.

With eight forecasts the weight becomes a vector, and the vector is not a longer version of the same object. It answers a question a ranking cannot be asked, it is often nowhere near the ranking’s answer, and on the set this field has been using it is nearly all zeros.

What the vector is

Combine m forecasts as jwjfj\sum_j w_j f_j with the weights summing to one, so the combination is unbiased for the target whenever the parts are. The combined error is jwjej\sum_j w_j e_j, its variance is wSww'Sw with S the error covariance matrix, and minimising that subject to 1w=1\mathbf{1}'w = 1 gives

w=S111S11,var(ec)=11S11w^* = \frac{S^{-1}\mathbf{1}}{\mathbf{1}'S^{-1}\mathbf{1}}, \qquad \operatorname{var}(e_c) = \frac{1}{\mathbf{1}'S^{-1}\mathbf{1}}

which is one linear solve and no iteration. Two things about that constraint are worth saying before any numbers.

The weights sum to one and are not required to be positive. A negative weight is not a mistake: it means one forecast is being used to correct another, which is what a combination is for. Requiring positive weights is a different estimator with different properties, usually chosen because the unconstrained one is unstable when S is estimated — a fact about estimation rather than about combination, and one this essay ends on.

And S is the matrix of error covariances, not of forecast covariances. Two forecasts can be nearly identical and still have very different errors, because a forecast’s error is the target minus the forecast, and it is the target’s absence from the combination that the matrix has to see.

Eight windows, and six weights of nothing

The set is the one the field has been using: eight moving averages of an autoregression, with window lengths 1, 2, 3, 5, 8, 13, 21 and 34. Their whole error covariance matrix is a closed form in the persistence — sums of powers of φ, no simulation anywhere — so the weight vector is exact.

Take the persistence at which the best of the eight exactly ties the sixty-observation benchmark, φ = 0.4895, which is the configuration the set comparisons are measured at. The eight mean squared errors, in units of the series’ own variance, run 1.0210, 1.0156, 1.0399, 1.0647, 1.0673, 1.0547, 1.0391 and 1.0262. The best of them is the two-observation mean.

The weight vector is

0.5032, 0.0000, −0.0000, 0.0000, −0.0001, 0.0012, −0.0182, 0.5139.

The ranking on the left, the weights on the right. Eight moving-average forecasts of an AR(1) at φ = 0.4895, the persistence at which the best of them exactly ties the 60-observation benchmark. On the left, each candidate's expected squared error in units of the series' own variance: the smallest belongs to L = 2, at 1.0156. On the right, the weight each carries in the variance-minimising combination of all eight — and the best of them carries 0.00000. The two ends of the family carry 1.0172 of the weight between them, and the combination they make is worth 0.7817, which is 23.0% below the best single forecast. Both columns are closed forms in φ. Which forecast to keep and which forecasts to use are different questions, and this is a set where the answers share nothing.
Fig. 1 The ranking on the left and the weights on the right, both computed from the same closed-form covariance matrix. Nothing about the left column predicts the right one.

Six of the eight weights are zero to three decimal places. The two ends of the family carry 1.0172 of the weight between them — slightly more than all of it, the excess being a small negative weight of −0.0182 on the twenty-one-observation mean. And the candidate with the smallest mean squared error carries 0.00000.

That is not a close call to be reported with a caveat. The best of the eight forecasts, the one a comparison table would put at the top and a reader would keep, contributes nothing whatever to the best use of the eight.

Three quantities, and only one of them crosses zero. Two forecasts of an AR(1) — the last value carried forward and the mean of the last 60 observations — at 1 step ahead. The curve through zero is σ₁² − σ₂², the difference in expected squared error that a comparison of accuracy tests; it changes sign at φ = 0.4922. The two curves above it are σ₁² − σ₁₂ and σ₂² − σ₁₂, the quantities the two encompassing tests are about, and neither of them comes near zero anywhere: the smallest value either takes across the range is 0.008 times the variance of the series. All three are closed forms in φ, R and h with no simulation in them. Equal accuracy is one hypothesis about this picture and encompassing is another, and a set of numbers can satisfy either without the other.
Fig. 2 The three second moments a pair of forecasts is described by. The vector generalises this to m(m+1)/2 of them, and the weight vector is a linear solve of the matrix they make.

Why the answer is the two ends

The reason is worth having, because it turns a surprising measurement into an obvious one and because it says what would happen with a different family.

The optimal one-step forecast of this series is φ·yₜ: the last value, shrunk towards the mean. Of the eight candidates, one is the last value and one is nearly the mean — a thirty-four observation average of a series whose persistence is under a half is close to zero. So a combination that puts w on the first and 1 − w on the last produces w·yₜ plus a little noise, and choosing w = φ reconstructs the optimal forecast almost exactly. Nothing in the middle of the family is needed, because a middling window is a worse version of the same shrinkage with no extra information in it.

The measurement confirms the account to two digits: the weight on the last value is 0.5032 at φ = 0.4895, 0.7185 at φ = 0.7 and 0.9180 at φ = 0.9. It exceeds φ by a small amount every time, because the longest window is not exactly the unconditional mean and its slight positive correlation with the target has to be paid for.

Eight weights, and six of them are nothing. The weight each of the eight moving averages carries in the variance-minimising combination, across the persistence of the series, from the closed-form error covariance with no simulation in it. Two curves matter and six lie on the axis: the last value and the 34-observation mean carry between them at least 1.001 of the weight everywhere in the range, and the largest weight any of the other six ever carries is 0.0426 in magnitude. The reason is that the optimal forecast of this series is the last value shrunk towards the mean, and those two are the ingredients that make it — so a set of eight candidates is a set of two, and which two has nothing to do with which of them forecasts best on its own.
Fig. 3 The eight weights across the persistence of the series. Two curves and six flat lines: whichever window forecasts best on its own, the combination is always the same two ingredients in different proportions.

So the sparsity is not a coincidence of one setting. Across the whole range of persistence the two ends carry at least 0.9 of the weight and the largest weight any interior window ever carries is under a tenth. The set of eight is a set of two, and which two has nothing to do with which of them forecasts best.

One consequence is worth stating in the other direction, because it is the practical form of the result. If a set of candidates is a family — windows of one series, lags of one model, smoothings of one signal — then adding members to the middle of it adds nothing at all. Eight candidates here do the work of two, and a ninth window of length 27 would do the work of none. The count of a candidate set is not a measure of how much of the problem it covers, and the quantity that is — how nearly singular the error covariance matrix is — is not something a table of mean squared errors displays.

Eight separate problems, and what a search over them costs. 320 set comparisons of 100 origins, each candidate on its own series at its own exact crossing, so all eight are on the boundary of the null and independent of one another. Five rejection rates at a nominal 5%. Reporting the best candidate's own p-value rejects 29.4% where one stated comparison rejects 4.1%, which makes the set behave like 8.39 independent comparisons rather than eight. A Bonferroni correction charges for eight and lands at 3.1%; the bootstrap over the whole set, which prices the dependence rather than assuming it away, is at 3.4%, and the recentred version at 2.5%.
Fig. 4 The same eight candidates run as eight separate problems, where a ranking and a weight vector would agree — because there is no one series for the two ends of one family to be the ingredients of.

Two questions, two answers, one table

The disagreement between the ranking and the weights is not an ordering dispute. It is a difference in what is being asked.

Which of these eight is worth keeping? is answered by the diagonal of S. Which of these eight should be used? is answered by S⁻¹1. The first question is the one every comparison table answers and it is the one almost nobody has: an organisation with eight forecasts in hand is not obliged to discard seven of them.

At this persistence the two answers share nothing at all. At φ = 0.2 they partly agree — the best single is the thirty-four observation mean and it also carries 0.8024 of the weight — and at φ = 0.9 they agree again, with the last value both best and heaviest. In between they come apart completely, and the place where they come apart is exactly the place a comparison is hardest, which is where the candidates are close.

What the small negative weight is doing

The one interior weight that is not zero is worth a paragraph, because a reader who has been told that six weights are zero will want to know why the seventh is −0.0182 rather than zero too.

It is the twenty-one observation mean, the second-longest window, and it is being subtracted. The thirty-four observation mean is not quite the unconditional mean — it still carries a little of the recent past — and subtracting a shorter long window removes some of what is left. The combination is using the difference between two long windows as a small correction to the second ingredient, which is exactly what a negative weight is for and is invisible to any procedure that ranks candidates.

Whether a reader should be pleased about this depends entirely on whether the weights are known or estimated. In population it is a genuine −1.8% improvement in the second ingredient. In a sample it is the first thing to become nonsense, because it is a difference between two nearly identical quantities, and the estimated version of it is measured below.

The ranking on the left, the weights on the right. Eight moving-average forecasts of an AR(1) at φ = 0.9000, the persistence at which the best of them exactly ties the 60-observation benchmark. On the left, each candidate's expected squared error in units of the series' own variance: the smallest belongs to L = 1, at 0.2000. On the right, the weight each carries in the variance-minimising combination of all eight — and the best of them carries 0.91800. The two ends of the family carry 1.0331 of the weight between them, and the combination they make is worth 0.1938, which is 3.1% below the best single forecast. Both columns are closed forms in φ. Which forecast to keep and which forecasts to use are different questions, and this is a set where the answers share nothing.
Fig. 5 The same two columns at a much more persistent series. The ranking and the weights agree here — the last value is both the best single forecast and the heaviest ingredient — and the interior weights are still nothing.

What the combination is worth

The population combination has an expected squared error of 0.7817 times the series’ variance, against 1.0156 for the best single forecast: 23.0% less. An equal-weight average of all eight — which requires knowing nothing at all — gets 0.8581, or 15.5% less than the best single.

Two thirds of what the optimal weights are worth is available to somebody who refuses to estimate anything.

Estimating eight numbers from sixty pairs of errors

Everything above is population. In practice S is estimated from the same rolling comparison the ranking comes from — thirty-six distinct covariances from P origins — and the weight vector is a linear solve of the estimate.

Scored on the population covariance, so that the number is the expected squared error of the combination that would actually have been used:

origins estimated optimum equal weights best single, picked in sample
30 1.1264 0.8581 1.0347
60 0.9006 0.8581 1.0333
120 0.8336 0.8581 1.0313
240 0.8077 0.8581 1.0294
960 0.7876 0.8581 1.0251

At sixty origins the estimated optimum is worse than the equal-weight average, and beats it on only 39.3% of comparisons. At thirty origins it is worse than every candidate in the set. The crossing is somewhere between 60 and 120 origins, and by 240 the estimated combination beats equal weights on 98.5% of comparisons.

What estimating eight weights costs, against not estimating them. Three combinations of the same eight forecasts, scored on the population covariance so that the number is the expected squared error of the combination that would actually have been used. The flat line is the equal-weight average, which estimates nothing. The curve is the variance-minimising combination estimated on P origins, and it is worse than the flat line until P is about 120 — eight weights estimated from sixty pairs of errors cost more than knowing them is worth. The lower line is the population optimum, 0.7817, which no sample reaches. The upper is what picking the single best forecaster in the same sample delivers, and it is the worst of the four at every length. 400 draws at each of 6 sample sizes.
Fig. 6 Three combinations of the same eight forecasts as the estimation sample lengthens, scored against the population. The flat line estimates nothing and is unbeatable for the first hundred origins.

This is the forecast combination puzzle — that simple averages beat estimated optimal weights — with the population answer known exactly rather than inferred, which makes the accounting unusually clean. The optimum is genuinely better; it is worth 8.9% of the equal-weight average’s mean squared error. What the estimate loses is more than that, because eight weights estimated from sixty pairs of errors is a near-singular problem: the candidates are windows of one series, S is nearly rank-deficient, and the solve responds by producing large offsetting weights. At sixty origins about three of the eight estimated weights are below −0.02, against one in population.

There is a second measurement in the same runs which says how badly the ranking serves as a guide to the weights. Picking the candidate with the smallest sample mean squared error selects the one that carries the heaviest population weight on 20.3% of comparisons at sixty origins — one in five, against one in eight for picking at random from the set. A table sorted by mean squared error is very nearly uninformative about which forecast the best use of the eight would lean on.

And the third column is the one to keep. Picking the single best candidate in the same sample is worse than the equal-weight average at every length, and worse than the population best single by more than the ranking noise would suggest — the winner of eight is selected partly for being lucky, which is the winner’s curse arriving in a comparison of forecasts.

What estimating eight weights costs, against not estimating themThree combinations of the same eight forecasts, scored on the population covariance so that the number is the expected squared error of the combination that would actually have been used. The flat line is the equal-weight average, which estimates nothing. The curve is the variance-minimising combination estimated on P origins, and it is worse than the flat line until P is about 120 — eight weights estimated from sixty pairs of errors cost more than knowing them is worth. The lower line is the population optimum, 0.7817, which no sample reaches. The upper is what picking the single best forecaster in the same sample delivers, and it is the worst of the four at every length. 800 draws at each of 6 sample sizes.0.8000.90011.103.404.094.795.486.176.87origins the weights are estimated onexpected squared error, in units of the series' varianceequal weights, estimating nothingthe population optimum, 0.782800 draws at each of 6 sample sizesscored on the population covariance
Fig. 7 The same comparison with each point averaged over twice as many draws. Drag it: the crossing does not move, because it is a property of how many numbers are being estimated from how many origins and not of how precisely the curve has been drawn.

Eight forecasts, worth one and a third

There is a natural unit for what a combination buys, and it makes the sparsity of the weight vector quantitative rather than merely visible.

If m forecasts were independent and equally accurate, averaging them would divide the error variance by m. So the ratio of a single forecast’s variance to the combination’s is the number of independent forecasts of that quality the set is worth. The best single here carries 1.0156 and the population combination 0.7817, which puts the set at

1.0156 / 0.7817 = 1.30

independent forecasts. Equal weights get 1.0156/0.8581 = 1.18. Eight candidates, optimally combined with a covariance matrix known exactly, are worth one and a third of the best of them.

That is the same finding as the six zero weights, in a form that can be compared across sets. It also sits beside a second effective count of the identical eight windows, measured for an entirely different purpose: the search over them is worth 2.44 independent comparisons rather than eight. Two effective counts of one set, 1.30 and 2.44, both far below the eight the table shows, and both saying the same thing about why — the candidates are views of one series.

They differ because they count different operations. A search is about the maximum of eight statistics, which the extremes of a family dominate and which therefore sees more than two of them; a combination is a sum, which the extremes exhaust. The multiplicity of a set exceeds its combinatorial value, and both are properties of the covariance matrix rather than of the count of rows.

The practical consequence is unwelcome and worth stating. A set assembled by sweeping a parameter across a range pays for eight candidates in the reference distribution of any search over it, and is paid for one and a third of them in any combination of it. Adding candidates to the middle of a family moves the first number and not the second.

How many origins a weight needs

The crossing — where estimating the optimum starts to beat refusing to estimate anything — lies between 60 and 120 origins, and the object being estimated has seven free parameters, since eight weights that sum to one are seven numbers.

That puts the crossing at somewhere between nine and seventeen origins per free weight, which is a usable rule for a decision that is otherwise taken by habit. Below it, an equal-weight average is not a compromise or a robustness measure; it is the better estimator, and the table says by how much — at thirty origins the estimated optimum is at 1.1264 against 0.8581, which is worse than every individual candidate in the set.

The rule also says what to do about a set that is too large for its evaluation, and it is not to run the solve anyway. Seven weights from sixty origins is the near-singular problem measured above, where three of the eight estimated weights come out below −0.02 against one in population. Halving the candidate set halves the free parameters and moves the crossing to about half as many origins, and since six of the eight candidates carry no population weight at all, the set can be halved without losing anything the combination would have used. What makes that hard in practice is that the column deciding which six to drop is the weight vector, and the weight vector is the thing there is not enough data to estimate.

What this leaves for encompassing

Encompassing for a pair is the hypothesis w* = 0. For a set it becomes m hypotheses about one parameter vector — candidate j adds nothing to the others — and the whole apparatus of multiplicity lands on a parameter rather than on a test statistic.

The population weights above say what the answer is here: six of the eight candidates genuinely add nothing, one adds a very little, and the two ends add a great deal. That is an unusually favourable situation for measuring what a set of encompassing tests does, because the truth is known exactly and most of the nulls are exactly true. The next essay makes all eight of them true at once, which the first-order condition defining the weights does for free.

Eight encompassing nulls, all of them true at once. The forecast under test is the variance-minimising combination of the eight candidates, and the first-order condition that defines its weights is cov(e_c, e_j) = var(e_c) for every j — so every one of the eight nulls is exactly true simultaneously and every rejection counted here is false. 800 draws of 60 origins. The largest of the eight statistics rejects 25.0% of the time at a nominal 5% and one comparison stated in advance rejects 3.1%, which makes the set worth about 9.1 independent comparisons — nearly the eight it has. Bonferroni, which is far inside its level on a search over accuracy comparisons of the same eight forecasters, is at 4.8% here.
Fig. 8 Where the next essay goes: the eight encompassing tests against the combination this essay is about, at a configuration where every one of their nulls is exactly true.

What is claimed here, and what is not

This essay takes the combination weight for more than two forecasts: what it is, what it says about this set, and what estimating it costs.

What stays out and is named as a decision: shrinkage and constrained estimators for the weights — positive weights, ridge on S, shrinking towards the equal-weight vector — which are the standard repair for the instability measured above and are a literature rather than a setting; combinations whose weights vary over time; and any loss function other than squared error, under which the covariance matrix stops being a sufficient description of the problem.

The boundary against the set field is that field’s own statistic. Everything about how a comparison of mean squared errors behaves — the denominator, the dependence between origins, the multiplicity of a search — is established there and used here without being re-derived.

The checks, and the refusals that make them mean something

Three claims are gated in this field’s library. The best candidate by mean squared error is required to be an interior window whose weight is under a hundredth in magnitude, and the two ends are required to carry more than 0.98 of the weight — the disagreement between the two questions, stated as something that would fail if the set were ever anything other than a set of two. The combination is required to beat both the best single forecast and the equal-weight average in population. And the estimated optimum is required to be worse than equal weights at sixty origins and better at two hundred and forty, which is the puzzle as a crossing rather than as an anecdote about one sample size.

The refusal beside them is a weight vector scored on the origins that produced it. The estimated weights minimise the sample covariance by construction, so in the sample they were fitted on they beat every other weight vector, including the population optimum — on 100% of comparisons rather than the 39.3% they manage against a covariance they did not choose. A check that scores a combination on its own estimation sample cannot fail, and the body applies the standard and throws.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationBenchmark forecastCombination weightCovariance matrixEstimation errorForecast combinationForecast encompassingMean squared errorModel selectionOut of sampleOverfittingRankingRolling originShrinkage