Eight forecasters and one benchmark
Worth reading first: What the correction corrects · What the model says next.
A forecast comparison in the wild is almost never a pair. It is a table: eight specifications, or forty trading rules, or a hundred and thirty model variants, each with its own average squared error and one of them highlighted. The pair is what the theory is written for, and the table is what gets published.
Turning the theory into something that covers the table is not a matter of applying it eight times. It is two problems arriving together — the search over the set, and the dependence between its members — and the second one changes the answer to the first.
The set, and a null that is true at its boundary
The candidates here are moving averages of the last L observations, for L = 1, 2, 3, 5, 8, 13, 21 and 34; the benchmark is the mean of the last sixty. All nine are members of one family whose error covariance matrix is a closed form, which is what makes the following possible.
The persistence of the series is not chosen. It is solved for: at φ = 0.4895 the best candidate in the set — the average of the last two observations — has exactly the benchmark’s expected squared error, and every other candidate is strictly worse.
That configuration is what a size measurement needs. The null — no candidate beats the benchmark — is true, with one candidate on its boundary, so every rejection counted below is a false one. It is also the configuration a procedure is least flattered by: anywhere further inside the null every correction looks better than it is, and a set in which every candidate is exactly on the boundary has to be built one series at a time, which is what the second half of this essay does.
What a search over the set actually costs
The obvious thing to do with a table is to read off the best row and quote its p-value. Call that the winner’s own p-value: it is what a comparison of one pair would report, applied to a pair chosen by looking at all eight.
The winner’s own p-value rejects 6.9% of the time. A single comparison stated in advance — the candidate on the boundary, named before the data — rejects 2.9%. So searching eight candidates has roughly doubled the error rate rather than multiplying it by eight.
The reason is that the eight candidates are not eight chances. They are eight views of the same shocks: an average of two observations and an average of three observations of one series agree about nearly everything, and the maximum of eight statistics that correlate at 0.9 is barely above the maximum of two. Written as a count,
which is the number of independent comparisons that would produce the observed pair of rates. Eight candidates, two and a half comparisons.
A Bonferroni correction charges for eight, and lands at 0.1% against a nominal 5%. That is not a small conservatism to be shrugged at: an error rate eighty times below its claim is bought entirely out of power, and the multiplicity field’s own argument about which rate a procedure controls applies here in a sharper form — the procedure is controlling the right rate for a set that does not exist.
The same eight, on eight separate problems
The measure calibrates itself, which is the thing that makes it worth having. Run the same eight candidates as eight separate forecasting problems: one series each, each drawn at the persistence where that candidate exactly ties its own benchmark — φ = 0.4922, 0.4895, 0.5104, 0.5536, 0.6024, 0.6542, 0.7006 and 0.7391. Now the columns of loss differentials are independent, every candidate is on the boundary of the null rather than one of them, and this is the configuration every correction in the multiplicity field was designed for.
The winner’s own p-value now rejects 27.4% where one stated comparison rejects 4.0%, and the effective count is 7.84 — the eight they are, which is what says the measure is a measure and not a rescaling. Bonferroni is at 3.4% and the bootstrap over the set at 4.0%, both at their nominal level to within the noise of five hundred comparisons.
So the identical table of eight forecasters costs the multiplicity of two and a half comparisons in one situation and of eight in the other, and nothing in the table says which situation it is. A correction that charges for the number of rows is right in one and wrong by a factor of three in the other; a correction that ignores multiplicity entirely is wrong in both, mildly here and badly there.
The two configurations are the two real situations
Neither configuration is a contrivance, and an evaluation is nearly always recognisably one or the other.
One series, many specifications is the dependent case: a modelling group with a time series and a shelf of methods, or a backtest of forty rules on one market. The candidates share every shock in the data, they agree about most of what happens, and the search over them is worth far less than its count. Correcting for eight there costs power for nothing.
Many series, one method each is the independent case: a forecaster reporting on eight countries, eight products or eight sites, with the best one written up. The comparisons share nothing, the search is worth its full count, and the naive reading is wrong by the factor everybody expects it to be wrong by — 27.4% against a claimed 5%, which is twenty analyses of nothing with forecasts in place of analyses.
The uncomfortable part is that the two look identical on the page. Both are a table with eight rows, a column of mean squared errors and a highlighted winner. The information that decides which correction is right — how much the candidates overlap — is in the loss differentials and not in the summary, so a reader handed the summary cannot tell, and neither can a procedure that is handed only the p-values. That is the argument for a reference distribution built from the differentials themselves rather than from anything computed out of them.
The reference distribution has to be over the set
What a correction has to price is the distribution of the maximum of the eight statistics, and that distribution depends on the whole covariance matrix between them rather than on any summary of it. Nothing writes it down.
It can be generated. The loss differentials form a P × 8 matrix, one row per origin, and resampling blocks of whole rows preserves both dependences at once: blocks, so the serial correlation the previous field measures survives, and whole rows, so the correlation between candidates survives with it. The maximum is then recomputed on each resample, and the observed maximum is read against the resulting distribution.
Both halves of the resampling matter, and each has a refusal in this field’s library. Drawing each candidate’s bootstrap sample from its own indices produces the maximum of eight independent candidates — a set nobody has — and it raises the critical value from 5.71 to 6.85, taking the procedure’s power at φ = 0.65 from 37.2% to 23.2%. A block length of one throws the serial dependence away instead, and at a four-step horizon it takes the rejection rate at a true null from 3.2% to 14.8%.
What the corrections are worth when there is something to find
A rejection rate at the null is size. Above it, the same numbers are power, and the ordering that made Bonferroni look merely cautious becomes the whole argument.
At φ = 0.6095 the bootstrap over the set finds the better forecaster 24.3% of the time and Bonferroni 8.3%. At φ = 0.6895 it is 58.0% against 32.7%. The winner’s own p-value is above both throughout — 44.0% and 75.3% — but it was above them at the null too, which is what makes it not a procedure.
The gap is largest in the middle of the range, which is where a decision is actually difficult. At the extremes every procedure agrees: nothing is found when there is nothing, and everything is found when the difference is enormous.
Two changes, measured one at a time
There are two named procedures in circulation for this and they differ in two ways at once, which is exactly the situation where a measurement has to separate them. One studentises the statistics — divides each candidate’s mean differential by its own long-run standard deviation before taking the maximum — and one recentres the reference distribution to drop candidates the data has already ruled out.
Both changes are measured separately here, and a third variant exists in this field’s library for no other reason: studentised and not recentred, which is neither procedure and is what makes the comparison a decomposition rather than a preference. At the null of this essay all three sit together — 1.1%, 1.0% and 1.0% — because with one candidate on the boundary and seven inside it there is nothing for a recentring to remove. The next essay is about what happens when there is.
Studentising is not free, and it is worth saying so here where the recentring cannot be credited with it. At φ = 0.65, with no junk in the set for the recentring to act on, the unstudentised maximum finds the better forecaster 34.3% of the time and the studentised one 26.0% — eight points, spent on a change made for a different reason. On this set the candidates’ long-run variances differ enough that dividing by them promotes a noisy candidate into the maximum more often than it promotes a good one, and that is a fact about this family of moving averages rather than a general result. It is stated as a measurement at a stated setting because that is all it is, and because the alternative — quoting one of the two procedures as better — would be crediting the recentring with the studentisation’s cost and the studentisation with the recentring’s gain.
Which candidate wins, and how little that says
One more number is worth reporting because it is the one a table implicitly claims. Across five hundred comparisons at the null, the maximum landed on the average of 1 in 25% of samples, on 34 in 28%, and on the boundary candidate — the average of 2, the only one that is genuinely tied with the benchmark — in 15%.
The winner of a set comparison is not an estimate of the best candidate; on this set, at this length, it is nearly a draw from the set. A table that reports the winner and its p-value is reporting two things, and neither of them is what a reader takes from it.
What the effective count would have charged
The effective count is not only a diagnosis; it is a number a correction could have used, and running it through one says exactly how much Bonferroni over-charges rather than merely that it does.
A Šidák-style level for m independent comparisons is 1 − (1 − α)^(1/m). At the eight the table contains, that is 0.64% per comparison. At the 2.44 the dependent set actually carries, it is 2.08% — three and a third times as much room. That factor is the whole of the conservatism, and it lands where the counted rates land: charging for eight when two and a half were made is what takes a nominal 5% to a realised 0.1%.
The same arithmetic in the independent configuration charges 0.64% against an effective count of 7.84, which asks for 0.65%. The two agree to a hundredth of a percentage point, which is why Bonferroni is at 3.4% there — at its level, within the noise of five hundred comparisons, doing exactly the job it was built for.
So the correction is not conservative and is not right; it is right in one configuration and wrong by a factor of three in the other, and the factor is computable from two rejection rates that any evaluation can produce. That is worth stating as a positive recommendation rather than only as a complaint about Bonferroni: an evaluation that can report both the naive rate and a stated-in-advance rate has measured its own multiplicity, and does not have to guess it from the number of rows.
The power figures price the same gap in the currency a forecaster spends. The bootstrap over the set finds the better forecaster 2.93 times as often as Bonferroni at φ = 0.6095, and 1.77 times as often at φ = 0.6895 — while the absolute gap moves the other way, from 16.0 points to 25.3. Both readings are correct and they answer different questions. A reader deciding whether the machinery is worth building wants the second; a reader with a weak signal and one chance to find it wants the first, and gets the larger number.
The winner is worth two and a half points
The distribution of which candidate wins is the most quotable thing in this essay and it deserves its arithmetic done.
Eight candidates, five hundred comparisons, and a winner picked at random would land on each about 12.5% of the time. The candidate that is genuinely tied with the benchmark — the only member of the set that is not strictly worse — wins 15%. The table’s winner therefore carries about two and a half percentage points more information about which candidate is best than a draw from a hat, and the reader who takes it as an identification is over-reading it by a factor nobody would defend if it were stated that way.
Where the rest of the mass goes is the more instructive half. The two ends of the family, the average of one observation and the average of thirty-four, take 25% and 28% between them — more than half the wins for two of the eight rows. The ends are the members least like the rest of the set: the shortest window shares least of its smoothing with anything, and the longest is the furthest from the short ones. A maximum over correlated statistics is taken by whichever member has the most variation of its own, so the winner of a set comparison is biased towards the extremes of whatever family was enumerated, quite apart from which member is any good.
That has a design consequence for anybody assembling a candidate list. Adding members in the middle of an existing range changes the winner distribution hardly at all and changes the effective count hardly at all; adding one at a new extreme does both, and adds a row that will win disproportionately often for reasons that have nothing to do with forecasting.
What is claimed here, and what is not
This essay claims the multiplicity of a set of forecast comparisons: that its size is not the number of candidates, that it is measurable, and that the reference distribution has to be over the whole set rather than over any one comparison in it.
The boundary against the exact-inference field is worth stating, because that field builds reference distributions too. There the distribution comes from re-running the mechanism that assigned the data, which the experimenter controls and therefore knows exactly. Here nothing was assigned: the mechanism is the series, it is the thing being estimated, and the resampling is an approximation to a distribution rather than an enumeration of one. The two constructions look alike on the page and only one of them is exact.
What stays out: the choice of block length as a subject, which has its own literature and is fixed at a stated value here; sets whose members are chosen sequentially in response to earlier results, where the set itself is data-dependent and none of this applies; and the false discovery rate applied to a forecast table, which is the other promise the multiplicity field distinguishes and which would answer a different question about the same eight rows.
The checks, and the refusal that makes them mean something
Three claims are gated. Eight independent comparisons are required to measure eight effective comparisons, which is what calibrates the measure; the same eight on one series are required to measure fewer than half of that; and the bootstrap over the whole set is required to be at its nominal level in the independent configuration, where every candidate sits on the boundary of the null, while the Bonferroni correction over the same eight is required to be far inside it in the dependent one.
Two refusals sit beside them, and both are ways of resampling that look right. A bootstrap that draws each candidate’s resample from its own indices is refused for pricing eight searches that were never made; a bootstrap with a block length of one is refused for throwing away the dependence that the previous field spent an essay measuring.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An order that spends the error rate — both name bonferroni, error rate, familywise error rate, monte carlo, multiple comparisons, null hypothesis
- How long a block a multiplier shares — both name block bootstrap, effective sample size, error rate, monte carlo, null hypothesis
- The corner the test is calibrated at — both name benchmark forecast, error rate, monte carlo, null hypothesis, reality check
- Estimating how many nulls are true — both name error rate, monte carlo, multiple comparisons, null hypothesis
- False discoveries that arrive together — both name error rate, familywise error rate, monte carlo, multiple comparisons
- Residuals that keep their own variance — both name block bootstrap, error rate, monte carlo, reality check
Named objects
A flat tag is an object no other essay names yet.
Benchmark forecastBlock bootstrapBonferroniData snoopingEffective sample sizeError rateFamilywise error rateLoss differentialMonte CarloMultiple comparisonsNull hypothesisReality checkSelection effectStationary bootstrap