When every null is true
Worth reading first: What the correction corrects · Three series and a count.
Measuring what a procedure does when it should not reject requires a situation in which it should not reject. For a set of comparisons that is harder than it sounds, and the difficulty is not statistical fastidiousness — it is the reason the standard procedures are built the way they are.
A reality check over m candidates is calibrated at the least favourable configuration: every candidate exactly as good as the benchmark, none better, none worse. That is the boundary of the null at which the maximum is most likely to look large, so a procedure calibrated there is conservative everywhere else. It is also a configuration no real set is in. Solving for a persistence at which the best of eight moving averages exactly ties a benchmark puts one candidate on the boundary and leaves the other seven strictly inside, and that is as close as an accuracy comparison gets.
Ask the other question about the same eight forecasts and the configuration arrives for free.
Eight nulls, one line of algebra
Let be the error covariance matrix of the m candidates and the weights of the variance-minimising combination, which sum to one. The first-order condition that defines those weights is
— every row of is the same number. Read one row of it: for every j. So
and that quantity is exactly what an encompassing test has a sample analogue of. The combination encompasses every one of its own parts, simultaneously, exactly, as an identity rather than as an approximation.
The least favourable configuration is attained. Not approached, not arranged by solving for a parameter, not true of one candidate with the others nearby: all eight nulls are on the boundary at once, and every rejection counted below is a false one.
That the six zero-weight candidates are on the boundary is worth a sentence, because it is the part that looks wrong. A candidate whose optimal weight is zero adds nothing to the combination — that is what “adds nothing” means — and a candidate with a large weight also adds nothing to the combination that already contains it. The null being tested is not “this forecast is useless”; it is “this forecast has nothing the combination has not got”, and against the optimal combination that is true of everything in the set.
There is one more property of this configuration that makes it unusually good to measure in, and it is the reason the arrangement is worth the algebra. Nothing about it has been tuned. The persistence is inherited from the accuracy search, the candidate set is the one that field already uses, and the weights are a closed-form solve of a closed-form covariance matrix — so there is no parameter anywhere that could have been chosen to make a rate come out at 5%. A size measured at a configuration somebody searched for is a measurement of the search as much as of the procedure, and this one was not searched for.
Four readings of the same eight comparisons
Run the eight tests on the same sixty origins, two thousand times, at a nominal 5%.
The largest of the eight statistics, read as though it had been the only one, rejects 26.2% of the time. One comparison stated in advance rejects 3.5%. Bonferroni over the eight rejects 5.2%, and a reference distribution over the whole set — the same block bootstrap of rows the accuracy search uses — rejects 4.5%.
On average 0.55 of the eight individual comparisons are called significant per draw, against the 0.40 that eight independent tests at their level would give.
The critical value behind those rates is worth quoting, because it is the number a reader would otherwise take off a table. The largest of the eight statistics is centred at 1.248 under this null — the maximum of eight statistics is not centred where each of them is — and its 95% point is 2.589, against 1.671 for a single comparison at this many origins. A table of eight encompassing tests with the biggest one starred is reporting a statistic whose distribution is nearly a full unit to the right of the distribution it is being read against.
Why the count comes out above eight
An effective count of 8.63 on a set of eight is worth stopping on, because eight is supposed to be the ceiling: the measure asks how many independent comparisons would produce the observed pair of rates, and eight independent comparisons is the most eight comparisons can be worth.
It exceeds it because the eight statistics are not merely uncorrelated — they are negatively constrained. Multiplying the j-th encompassing quantity by its own weight and summing gives
Σⱼ wⱼ·[var(e_c) − cov(e_c, eⱼ)] = v − v = 0
exactly, by the same first-order condition that put every null on the boundary. So the eight quantities are pinned to a weighted sum of zero: whenever some of them come out high, the rest are obliged to come out low.
A set of statistics pushed apart in that way has a larger maximum than the same number of independent ones, so the naive rate rises relative to the single-comparison rate, so the measure returns a count above the nominal one. The set is worth more than eight independent comparisons because it is worth eight comparisons that cannot all be quiet at once.
That is a structural fact rather than a measurement artefact, and it is the exact opposite of the accuracy search’s situation, where the eight are strongly positively correlated and worth 2.44. The same eight forecasters are sub-additive on one question and super-additive on the other, for the same reason in two directions: what makes them alike as forecasts is what makes their contributions to a combination complementary.
The correction’s error is a factor of fifty for a factor of three in the count
Bonferroni charges for eight in both settings, and the two effective counts are 8.63 and 2.44 — a factor of 3.5.
Its realised sizes are 5.2% here and 0.1% on the accuracy search — a factor of 52.
So a threefold error in the count buys a fiftyfold error in the rate, which is the steepness of a normal tail expressed as a design fact. A correction that is charging for the wrong number of comparisons is wrong in the rate by far more than it is wrong in the count, and the direction is always the same: over-charging costs orders of magnitude and under-charging gains them.
That is the practical reason to measure the overlap rather than assume it. A reference distribution built by resampling the set is within a point of its level on both questions, and it gets there by computing a quantity that the count-based correction has to guess — a quantity that, on one identical set of eight forecasters, takes values on both sides of eight.
A multiplicity three times the one the same eight carry
Read the first two numbers as the site’s usual measure of how many independent comparisons a set is worth — how many separate chances would produce a 26.2% winner’s rate when one comparison gives 3.5% — and the answer is 8.63 of a possible 8.
The identical measurement on the identical eight forecasters, comparing them by mean squared error instead, gives 2.44.
That is the same set of candidates, the same series, the same number of origins and the same nominal level, and the correct charge for having looked at all eight differs by a factor of three and a half depending on which question is being asked of them.
The reason is what each statistic is a statement about. An accuracy comparison asks about , and the eight variances move together almost perfectly: they are windows of one series, a long window contains the short ones, and a sample in which one of them looks good is a sample in which they all do. An encompassing comparison against the combination asks about — what candidate j has that the combination has not — and those eight quantities are nearly orthogonal, because the thing each candidate has that the others do not is, by construction, the thing that distinguishes it.
The practical consequence is a warning about corrections rather than about tests. Bonferroni over these eight is far inside its level on one question and at its level on the other; a reference distribution built by resampling the set is close to right on both, because it measures the overlap rather than assuming a value for it. A correction that charges by the number of models is charging for a quantity that is not the number of models, and how wrong it is depends on which question the models are being asked.
The six exactly-redundant candidates raise the question the accuracy search has its own answer to: what do candidates that are certain not to win do to a procedure? There, adding sixteen stale forecasts nobody would have run took a reality check from 34.3% power to nothing at all, because the recentring is built on the assumption that every candidate might be on the boundary and stale ones are nowhere near it.
Here the situation is inverted and instructive. The six candidates carrying no weight are exactly on the boundary — they add exactly nothing, which is precisely what the null says — so they are not junk in that sense at all. They cost the procedure the multiplicity of six comparisons and contribute no chance of a true rejection, which is the same arithmetic wearing a different face: a set procedure pays for every candidate and collects only from the ones that can move.
What it detects, and what it lets past
A size measurement is half a check. The other half is what the set-level test does when the combination under test is not the optimal one, and this is easy to arrange: test the same eight candidates against a combination formed with the wrong weights.
Three of them, in order of how wrong they are.
The optimal combination is 0.0% above the optimum by definition, and the set test rejects 4.5% of the time. That is the size.
The equal-weight average — which is the best thing available to somebody who refuses to estimate anything, and beats every single forecaster in the set — sits 9.8% above the optimum. The set test catches it 21.5% of the time.
The best single forecast, chosen by mean squared error, sits 29.9% above the optimum. The set test catches it 96.3% of the time, and on average 5.66 of the eight candidates are individually flagged as adding something to it.
The gap between the second and third rows is the useful part. A forecast that is badly wrong — a single candidate where a combination was available — is caught almost always, and the eight tests between them name most of the set as adding something, which is a diagnosis and not merely a rejection. A forecast that is nearly right is let past four times out of five. A set-level encompassing test is an instrument for detecting that a combination is missing something large, and it is not an instrument for tuning one.
Why the equal-weight average is hard to catch
The 21.5% deserves an explanation rather than a shrug, because 9.8% above the optimum is not a small error and sixty origins is not a tiny sample.
The equal-weight average is wrong in a very particular way: it puts a sixth of its weight on the last value where the optimum puts a half, and spreads the rest over six candidates the optimum ignores. What the eight tests are asked is whether any individual candidate adds something to it, and the answer for six of the eight is “hardly” — those six are nearly redundant against each other whatever the weights are. The whole of the deficiency is concentrated in the two ends, so the set test is effectively running two informative comparisons and six uninformative ones, and paying the multiplicity for all eight.
That is visible in the same runs from the other direction: one comparison stated in advance — the last value, which is the candidate the equal-weight average underweights most — rejects 61.0% of the time against the set procedure’s 21.5%. A pre-stated comparison beats a search here by forty points, which is the ordinary price of multiplicity arriving in an unusually sharp form: the search is paying for six comparisons that cannot succeed.
What a rejection would actually mean
It is worth being explicit about what this procedure is for, because “reject the hypothesis that every candidate is redundant” is an awkward sentence to act on.
The set of m encompassing tests against a forecast in production answers one operational question: is there anything in this set of candidates that my current forecast has not got? A rejection says yes and names the candidate that supplied the maximum. That is a shortlist of one, arrived at with the search paid for — which is more than a table of eight statistics with the best one starred provides, and considerably less than a combination.
What it does not answer is how much better the combination would be, which is a question about the weight vector and not about any test. The measurements above give the honest ordering: a forecast 9.8% above the best available use of the set is let past four times in five, and one 29.9% above it is caught almost always. So a non-rejection is evidence that nothing in the set is badly missing, and nothing more than that.
And the asymmetry the two halves of this field share is worth restating. The procedure that holds its level here is the one that resamples the observed comparisons, because these candidates estimate nothing and the null is a statement about the data. Where the candidates are fitted models that same construction recentres in the wrong place and stops rejecting altogether. One set of eight, two questions, and the reference distribution that is right for each of them is the one the other question would be broken by.
What is claimed here, and what is not
This essay takes the size and power of a search over encompassing comparisons, at a configuration where the least favourable case is attained rather than assumed.
What stays out and is named as a decision: joint tests of the whole weight vector — a Wald test that all m − 1 free components equal a stated value, which is one comparison rather than m and a different object; encompassing against a combination whose weights were estimated, where the null is no longer exactly true and the noise in the weights is inside the statistic being tested; and the sequential version, in which candidates are added to a combination one at a time and the multiplicity becomes a stopping problem.
The boundary against the set field is which quantity the maximum is taken over. Everything about the block bootstrap of rows, the studentising, and the recentring is established there and used here unchanged; what changes is that the eight statistics being maximised over are nearly independent rather than nearly identical, and every conclusion that field reached about corrections reverses because of it.
The checks, and the refusals that make them mean something
Three claims are gated in this field’s library. The first-order condition is checked as an identity: the combined error’s covariance with every member is required to equal the combined error’s own variance to within 10⁻⁹, at four persistences, which is what makes every null exactly true rather than approximately so. The winner’s own p-value is required to reject several times as often as one comparison stated in advance, and Bonferroni is required to be within two and a half points of its nominal level — which is the contrast with the accuracy search stated as something that could fail. And the set-level test is required to be at its level against the optimal combination, to catch the best single forecast most of the time, and to catch the equal-weight average much less often, in that order.
The refusals that stand behind them belong to this field’s other half and apply here unchanged: a reference distribution recentred where the null does not put it, and a combination scored on the origins that chose it. Both would produce an apparently well-behaved set procedure at this configuration, and neither would be measuring anything.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Which forecast is better — both name benchmark forecast, effective sample size, error rate, mean squared error, null hypothesis
- Two defects and one resampling — both name block bootstrap, error rate, null hypothesis, reference distribution
- A distribution drawn from the null — both name benchmark forecast, null hypothesis, reference distribution
- A taper and a critical value — both name block bootstrap, error rate, reference distribution
- An order that spends the error rate — both name bonferroni, error rate, null hypothesis
- Errors generated from a fitted model — both name block bootstrap, error rate, reference distribution
Named objects
A flat tag is an object no other essay names yet.
Benchmark forecastBlock bootstrapBonferroniCombination weightEffective sample sizeError rateForecast combinationForecast encompassingLeast favourableMean squared errorMultiplicityNull hypothesisPowerReference distribution