The best of a set, and what the search costs

Eight forecasters and one benchmark

A set of forecasters is a multiplicity problem on top of a dependence problem, and the two do not separate. Eight windows of one series carry the multiplicity of two and a half independent comparisons; eight separate problems carry eight.

Worth reading first: What the correction corrects · What the model says next.

A forecast comparison in the wild is almost never a pair. It is a table: eight specifications, or forty trading rules, or a hundred and thirty model variants, each with its own average squared error and one of them highlighted. The pair is what the theory is written for, and the table is what gets published.

Turning the theory into something that covers the table is not a matter of applying it eight times. It is two problems arriving together — the search over the set, and the dependence between its members — and the second one changes the answer to the first.

The set, and a null that is true at its boundary

The candidates here are moving averages of the last L observations, for L = 1, 2, 3, 5, 8, 13, 21 and 34; the benchmark is the mean of the last sixty. All nine are members of one family whose error covariance matrix is a closed form, which is what makes the following possible.

The persistence of the series is not chosen. It is solved for: at φ = 0.4895 the best candidate in the set — the average of the last two observations — has exactly the benchmark’s expected squared error, and every other candidate is strictly worse.

Eight candidates, one of them exactly as good as the benchmark. The candidate set: moving averages of the last 1, 2, 3, 5, 8, 13, 21 and 34 observations, each drawn as its expected squared error divided by the benchmark's — the mean of all 60. The persistence is not chosen, it is solved for: at φ = 0.4895 the best candidate in the set, the average of 2, has exactly the benchmark's expected squared error, and every other candidate is worse by between 0.5% and 5.1%. So the null that no candidate beats the benchmark is true, with one candidate on its boundary. Everything a set comparison claims about its own error rate has to be measured here, because anywhere further inside the null every procedure flatters itself.
Fig. 1 The eight candidates and the benchmark at the solved persistence, each drawn as its expected squared error divided by the benchmark’s. One sits exactly on the boundary of the null; the others are inside it by between half a percent and five percent.

That configuration is what a size measurement needs. The null — no candidate beats the benchmark — is true, with one candidate on its boundary, so every rejection counted below is a false one. It is also the configuration a procedure is least flattered by: anywhere further inside the null every correction looks better than it is, and a set in which every candidate is exactly on the boundary has to be built one series at a time, which is what the second half of this essay does.

What a search over the set actually costs

The obvious thing to do with a table is to read off the best row and quote its p-value. Call that the winner’s own p-value: it is what a comparison of one pair would report, applied to a pair chosen by looking at all eight.

Eight windows of one series, and what a search over them costs. 320 set comparisons of 100 origins, all eight candidates on the same series at φ = 0.4895, so they are eight views of the same shocks. Five rejection rates at a nominal 5%. Reporting the best candidate's own p-value rejects 6.9% where one stated comparison rejects 2.8%, which makes the set behave like 2.50 independent comparisons rather than eight. A Bonferroni correction charges for eight and lands at 0.3%; the bootstrap over the whole set, which prices the dependence rather than assuming it away, is at 0.9%, and the recentred version at 1.3%.
Fig. 2 Five hundred set comparisons of a hundred origins each, at the solved null. Five ways of reading the same table, all at a nominal 5%.

The winner’s own p-value rejects 6.9% of the time. A single comparison stated in advance — the candidate on the boundary, named before the data — rejects 2.9%. So searching eight candidates has roughly doubled the error rate rather than multiplying it by eight.

The reason is that the eight candidates are not eight chances. They are eight views of the same shocks: an average of two observations and an average of three observations of one series agree about nearly everything, and the maximum of eight statistics that correlate at 0.9 is barely above the maximum of two. Written as a count,

meff=log(1naive)log(1single)=2.44m_{\text{eff}} = \frac{\log(1 - \text{naive})}{\log(1 - \text{single})} = 2.44

which is the number of independent comparisons that would produce the observed pair of rates. Eight candidates, two and a half comparisons.

A Bonferroni correction charges for eight, and lands at 0.1% against a nominal 5%. That is not a small conservatism to be shrugged at: an error rate eighty times below its claim is bought entirely out of power, and the multiplicity field’s own argument about which rate a procedure controls applies here in a sharper form — the procedure is controlling the right rate for a set that does not exist.

The same eight, on eight separate problems

The measure calibrates itself, which is the thing that makes it worth having. Run the same eight candidates as eight separate forecasting problems: one series each, each drawn at the persistence where that candidate exactly ties its own benchmark — φ = 0.4922, 0.4895, 0.5104, 0.5536, 0.6024, 0.6542, 0.7006 and 0.7391. Now the columns of loss differentials are independent, every candidate is on the boundary of the null rather than one of them, and this is the configuration every correction in the multiplicity field was designed for.

Eight separate problems, and what a search over them costs. 320 set comparisons of 100 origins, each candidate on its own series at its own exact crossing, so all eight are on the boundary of the null and independent of one another. Five rejection rates at a nominal 5%. Reporting the best candidate's own p-value rejects 29.4% where one stated comparison rejects 4.1%, which makes the set behave like 8.39 independent comparisons rather than eight. A Bonferroni correction charges for eight and lands at 3.1%; the bootstrap over the whole set, which prices the dependence rather than assuming it away, is at 3.4%, and the recentred version at 2.5%.
Fig. 3 The same five readings on eight independent problems. The winner’s own p-value has gone from 6.9% to 27.4%, and the two corrections that were far inside their level are now at it.

The winner’s own p-value now rejects 27.4% where one stated comparison rejects 4.0%, and the effective count is 7.84 — the eight they are, which is what says the measure is a measure and not a rescaling. Bonferroni is at 3.4% and the bootstrap over the set at 4.0%, both at their nominal level to within the noise of five hundred comparisons.

So the identical table of eight forecasters costs the multiplicity of two and a half comparisons in one situation and of eight in the other, and nothing in the table says which situation it is. A correction that charges for the number of rows is right in one and wrong by a factor of three in the other; a correction that ignores multiplicity entirely is wrong in both, mildly here and badly there.

Eight separate problems, and what a search over them costs. 320 set comparisons of 200 origins, each candidate on its own series at its own exact crossing, so all eight are on the boundary of the null and independent of one another. Five rejection rates at a nominal 5%. Reporting the best candidate's own p-value rejects 25.0% where one stated comparison rejects 4.1%, which makes the set behave like 6.94 independent comparisons rather than eight. A Bonferroni correction charges for eight and lands at 1.9%; the bootstrap over the whole set, which prices the dependence rather than assuming it away, is at 5.0%, and the recentred version at 2.2%.
Fig. 4 The independent configuration on twice as many origins. Every rate is a property of the set rather than of the length of the evaluation, and the two that were at their level stay there.

The two configurations are the two real situations

Neither configuration is a contrivance, and an evaluation is nearly always recognisably one or the other.

One series, many specifications is the dependent case: a modelling group with a time series and a shelf of methods, or a backtest of forty rules on one market. The candidates share every shock in the data, they agree about most of what happens, and the search over them is worth far less than its count. Correcting for eight there costs power for nothing.

Many series, one method each is the independent case: a forecaster reporting on eight countries, eight products or eight sites, with the best one written up. The comparisons share nothing, the search is worth its full count, and the naive reading is wrong by the factor everybody expects it to be wrong by — 27.4% against a claimed 5%, which is twenty analyses of nothing with forecasts in place of analyses.

The uncomfortable part is that the two look identical on the page. Both are a table with eight rows, a column of mean squared errors and a highlighted winner. The information that decides which correction is right — how much the candidates overlap — is in the loss differentials and not in the summary, so a reader handed the summary cannot tell, and neither can a procedure that is handed only the p-values. That is the argument for a reference distribution built from the differentials themselves rather than from anything computed out of them.

The reference distribution has to be over the set

What a correction has to price is the distribution of the maximum of the eight statistics, and that distribution depends on the whole covariance matrix between them rather than on any summary of it. Nothing writes it down.

It can be generated. The loss differentials form a P × 8 matrix, one row per origin, and resampling blocks of whole rows preserves both dependences at once: blocks, so the serial correlation the previous field measures survives, and whole rows, so the correlation between candidates survives with it. The maximum is then recomputed on each resample, and the observed maximum is read against the resulting distribution.

Three quantities, and only one of them crosses zero. Two forecasts of an AR(1) — the last value carried forward and the mean of the last 60 observations — at 1 step ahead. The curve through zero is σ₁² − σ₂², the difference in expected squared error that a comparison of accuracy tests; it changes sign at φ = 0.4922. The two curves above it are σ₁² − σ₁₂ and σ₂² − σ₁₂, the quantities the two encompassing tests are about, and neither of them comes near zero anywhere: the smallest value either takes across the range is 0.008 times the variance of the series. All three are closed forms in φ, R and h with no simulation in them. Equal accuracy is one hypothesis about this picture and encompassing is another, and a set of numbers can satisfy either without the other.
Fig. 5 Where the pair case ends. Two forecasters, three moments, and closed forms for all of them. Eight forecasters have thirty-six such moments and the question is about the largest of eight statistics built from them, which is what forces a reference distribution rather than a formula.

Both halves of the resampling matter, and each has a refusal in this field’s library. Drawing each candidate’s bootstrap sample from its own indices produces the maximum of eight independent candidates — a set nobody has — and it raises the critical value from 5.71 to 6.85, taking the procedure’s power at φ = 0.65 from 37.2% to 23.2%. A block length of one throws the serial dependence away instead, and at a four-step horizon it takes the rejection rate at a true null from 3.2% to 14.8%.

What the corrections are worth when there is something to find

A rejection rate at the null is size. Above it, the same numbers are power, and the ordering that made Bonferroni look merely cautious becomes the whole argument.

Where the four readings agree, and where they do not. The same four procedures across the persistence of the series, 200 set comparisons of 100 origins at each of five points. The leftmost point is the solved crossing, where the null is true: the winner's own p-value rejects 6.5%, the Bonferroni correction 0.0%, the bootstrap over the set 2.0%. Everywhere to the right of it the best candidate is genuinely better and the same numbers are power: by φ = 0.77 the bootstrap finds it 83.0% of the time against Bonferroni's 59.5%. The gap between those two is the whole value of pricing the dependence rather than assuming it away, and it is largest in the middle of the range where a decision is actually difficult.
Fig. 6 The four readings as the persistence rises past the solved crossing, where the best candidate stops being tied with the benchmark and starts being better than it. Three hundred set comparisons at each of five points.

At φ = 0.6095 the bootstrap over the set finds the better forecaster 24.3% of the time and Bonferroni 8.3%. At φ = 0.6895 it is 58.0% against 32.7%. The winner’s own p-value is above both throughout — 44.0% and 75.3% — but it was above them at the null too, which is what makes it not a procedure.

Where the four readings agree, and where they do not. The same four procedures across the persistence of the series, 200 set comparisons of 200 origins at each of five points. The leftmost point is the solved crossing, where the null is true: the winner's own p-value rejects 7.0%, the Bonferroni correction 0.0%, the bootstrap over the set 1.0%. Everywhere to the right of it the best candidate is genuinely better and the same numbers are power: by φ = 0.77 the bootstrap finds it 98.5% of the time against Bonferroni's 96.0%. The gap between those two is the whole value of pricing the dependence rather than assuming it away, and it is largest in the middle of the range where a decision is actually difficult.
Fig. 7 The same five points on twice the origins. Every curve rises and the gap between the bootstrap and the Bonferroni correction rises with them, because the thing Bonferroni is over-charging for does not get smaller with more data.

The gap is largest in the middle of the range, which is where a decision is actually difficult. At the extremes every procedure agrees: nothing is found when there is nothing, and everything is found when the difference is enormous.

Two changes, measured one at a time

There are two named procedures in circulation for this and they differ in two ways at once, which is exactly the situation where a measurement has to separate them. One studentises the statistics — divides each candidate’s mean differential by its own long-run standard deviation before taking the maximum — and one recentres the reference distribution to drop candidates the data has already ruled out.

Both changes are measured separately here, and a third variant exists in this field’s library for no other reason: studentised and not recentred, which is neither procedure and is what makes the comparison a decomposition rather than a preference. At the null of this essay all three sit together — 1.1%, 1.0% and 1.0% — because with one candidate on the boundary and seven inside it there is nothing for a recentring to remove. The next essay is about what happens when there is.

Studentising is not free, and it is worth saying so here where the recentring cannot be credited with it. At φ = 0.65, with no junk in the set for the recentring to act on, the unstudentised maximum finds the better forecaster 34.3% of the time and the studentised one 26.0% — eight points, spent on a change made for a different reason. On this set the candidates’ long-run variances differ enough that dividing by them promotes a noisy candidate into the maximum more often than it promotes a good one, and that is a fact about this family of moving averages rather than a general result. It is stated as a measurement at a stated setting because that is all it is, and because the alternative — quoting one of the two procedures as better — would be crediting the recentring with the studentisation’s cost and the studentisation with the recentring’s gain.

Which candidate wins, and how little that says

One more number is worth reporting because it is the one a table implicitly claims. Across five hundred comparisons at the null, the maximum landed on the average of 1 in 25% of samples, on 34 in 28%, and on the boundary candidate — the average of 2, the only one that is genuinely tied with the benchmark — in 15%.

Eight candidates, one of them exactly as good as the benchmark. The candidate set: moving averages of the last 1, 2, 3, 5, 8, 13, 21 and 34 observations, each drawn as its expected squared error divided by the benchmark's — the mean of all 60. The persistence is not chosen, it is solved for: at φ = 0.8241 the best candidate in the set, the average of 2, has exactly the benchmark's expected squared error, and every other candidate is worse by between 0.6% and 6.2%. So the null that no candidate beats the benchmark is true, with one candidate on its boundary. Everything a set comparison claims about its own error rate has to be measured here, because anywhere further inside the null every procedure flatters itself.
Fig. 8 The candidate set at a four-step horizon, where the solved crossing moves and a different member of the family sits on the boundary. Which candidate is closest to the benchmark is a property of the horizon, and the winner in a sample is barely a property of anything.

The winner of a set comparison is not an estimate of the best candidate; on this set, at this length, it is nearly a draw from the set. A table that reports the winner and its p-value is reporting two things, and neither of them is what a reader takes from it.

Eight windows of one series, and what a search over them costs320 set comparisons of 50 origins, all eight candidates on the same series at φ = 0.4895, so they are eight views of the same shocks. Five rejection rates at a nominal 5%. Reporting the best candidate's own p-value rejects 5.3% where one stated comparison rejects 1.3%, which makes the set behave like 4.34 independent comparisons rather than eight. A Bonferroni correction charges for eight and lands at 0.3%; the bootstrap over the whole set, which prices the dependence rather than assuming it away, is at 0.0%, and the recentred version at 0.6%.the winner's own p-value5.3%Bonferroni over the eight0.3%bootstrap over the whole set0.0%the same, recentred0.6%one stated comparison1.3%nominal 5%how the set is read320 comparisons of 50 origins, 119 draws eachthe set is worth 4.34 comparisons
Fig. 9 The dependent set at fifty origins. Drag it: the effective count moves with the length of the evaluation as well as with the set, because how alike two candidates look is partly a matter of how long they have been watched.

What the effective count would have charged

The effective count is not only a diagnosis; it is a number a correction could have used, and running it through one says exactly how much Bonferroni over-charges rather than merely that it does.

A Šidák-style level for m independent comparisons is 1 − (1 − α)^(1/m). At the eight the table contains, that is 0.64% per comparison. At the 2.44 the dependent set actually carries, it is 2.08% — three and a third times as much room. That factor is the whole of the conservatism, and it lands where the counted rates land: charging for eight when two and a half were made is what takes a nominal 5% to a realised 0.1%.

The same arithmetic in the independent configuration charges 0.64% against an effective count of 7.84, which asks for 0.65%. The two agree to a hundredth of a percentage point, which is why Bonferroni is at 3.4% there — at its level, within the noise of five hundred comparisons, doing exactly the job it was built for.

So the correction is not conservative and is not right; it is right in one configuration and wrong by a factor of three in the other, and the factor is computable from two rejection rates that any evaluation can produce. That is worth stating as a positive recommendation rather than only as a complaint about Bonferroni: an evaluation that can report both the naive rate and a stated-in-advance rate has measured its own multiplicity, and does not have to guess it from the number of rows.

The power figures price the same gap in the currency a forecaster spends. The bootstrap over the set finds the better forecaster 2.93 times as often as Bonferroni at φ = 0.6095, and 1.77 times as often at φ = 0.6895 — while the absolute gap moves the other way, from 16.0 points to 25.3. Both readings are correct and they answer different questions. A reader deciding whether the machinery is worth building wants the second; a reader with a weak signal and one chance to find it wants the first, and gets the larger number.

The winner is worth two and a half points

The distribution of which candidate wins is the most quotable thing in this essay and it deserves its arithmetic done.

Eight candidates, five hundred comparisons, and a winner picked at random would land on each about 12.5% of the time. The candidate that is genuinely tied with the benchmark — the only member of the set that is not strictly worse — wins 15%. The table’s winner therefore carries about two and a half percentage points more information about which candidate is best than a draw from a hat, and the reader who takes it as an identification is over-reading it by a factor nobody would defend if it were stated that way.

Where the rest of the mass goes is the more instructive half. The two ends of the family, the average of one observation and the average of thirty-four, take 25% and 28% between them — more than half the wins for two of the eight rows. The ends are the members least like the rest of the set: the shortest window shares least of its smoothing with anything, and the longest is the furthest from the short ones. A maximum over correlated statistics is taken by whichever member has the most variation of its own, so the winner of a set comparison is biased towards the extremes of whatever family was enumerated, quite apart from which member is any good.

That has a design consequence for anybody assembling a candidate list. Adding members in the middle of an existing range changes the winner distribution hardly at all and changes the effective count hardly at all; adding one at a new extreme does both, and adds a row that will win disproportionately often for reasons that have nothing to do with forecasting.

What is claimed here, and what is not

This essay claims the multiplicity of a set of forecast comparisons: that its size is not the number of candidates, that it is measurable, and that the reference distribution has to be over the whole set rather than over any one comparison in it.

The boundary against the exact-inference field is worth stating, because that field builds reference distributions too. There the distribution comes from re-running the mechanism that assigned the data, which the experimenter controls and therefore knows exactly. Here nothing was assigned: the mechanism is the series, it is the thing being estimated, and the resampling is an approximation to a distribution rather than an enumeration of one. The two constructions look alike on the page and only one of them is exact.

What stays out: the choice of block length as a subject, which has its own literature and is fixed at a stated value here; sets whose members are chosen sequentially in response to earlier results, where the set itself is data-dependent and none of this applies; and the false discovery rate applied to a forecast table, which is the other promise the multiplicity field distinguishes and which would answer a different question about the same eight rows.

The checks, and the refusal that makes them mean something

Three claims are gated. Eight independent comparisons are required to measure eight effective comparisons, which is what calibrates the measure; the same eight on one series are required to measure fewer than half of that; and the bootstrap over the whole set is required to be at its nominal level in the independent configuration, where every candidate sits on the boundary of the null, while the Bonferroni correction over the same eight is required to be far inside it in the dependent one.

Two refusals sit beside them, and both are ways of resampling that look right. A bootstrap that draws each candidate’s resample from its own indices is refused for pricing eight searches that were never made; a bootstrap with a block length of one is refused for throwing away the dependence that the previous field spent an essay measuring.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Benchmark forecastBlock bootstrapBonferroniData snoopingEffective sample sizeError rateFamilywise error rateLoss differentialMonte CarloMultiple comparisonsNull hypothesisReality checkSelection effectStationary bootstrap