Searching among fitted models

When every null is true

A reality check assumes that every candidate in the set is exactly as good as the benchmark, which is a configuration nobody's data is ever in. Test a combination against its own parts and that configuration is not assumed — it is what the arithmetic makes true.

Worth reading first: What the correction corrects · Three series and a count.

Measuring what a procedure does when it should not reject requires a situation in which it should not reject. For a set of comparisons that is harder than it sounds, and the difficulty is not statistical fastidiousness — it is the reason the standard procedures are built the way they are.

A reality check over m candidates is calibrated at the least favourable configuration: every candidate exactly as good as the benchmark, none better, none worse. That is the boundary of the null at which the maximum is most likely to look large, so a procedure calibrated there is conservative everywhere else. It is also a configuration no real set is in. Solving for a persistence at which the best of eight moving averages exactly ties a benchmark puts one candidate on the boundary and leaves the other seven strictly inside, and that is as close as an accuracy comparison gets.

Ask the other question about the same eight forecasts and the configuration arrives for free.

Eight nulls, one line of algebra

Let SS be the error covariance matrix of the m candidates and ww the weights of the variance-minimising combination, which sum to one. The first-order condition that defines those weights is

Sw=v1,v=var(ec)Sw = v\,\mathbf{1}, \qquad v = \operatorname{var}(e_c)

— every row of SwSw is the same number. Read one row of it: cov(ec,ej)=v\operatorname{cov}(e_c, e_j) = v for every j. So

var(ec)cov(ec,ej)=0for every j,\operatorname{var}(e_c) - \operatorname{cov}(e_c, e_j) = 0 \quad \text{for every } j,

and that quantity is exactly what an encompassing test has a sample analogue of. The combination encompasses every one of its own parts, simultaneously, exactly, as an identity rather than as an approximation.

The least favourable configuration is attained. Not approached, not arranged by solving for a parameter, not true of one candidate with the others nearby: all eight nulls are on the boundary at once, and every rejection counted below is a false one.

The ranking on the left, the weights on the right. Eight moving-average forecasts of an AR(1) at φ = 0.4895, the persistence at which the best of them exactly ties the 60-observation benchmark. On the left, each candidate's expected squared error in units of the series' own variance: the smallest belongs to L = 2, at 1.0156. On the right, the weight each carries in the variance-minimising combination of all eight — and the best of them carries 0.00000. The two ends of the family carry 1.0172 of the weight between them, and the combination they make is worth 0.7817, which is 23.0% below the best single forecast. Both columns are closed forms in φ. Which forecast to keep and which forecasts to use are different questions, and this is a set where the answers share nothing.
Fig. 1 The combination the eight nulls are about. Two of the eight carry all the weight, which does not stop the other six from being exactly on the boundary — a candidate with zero weight adds nothing, which is what the null says.

That the six zero-weight candidates are on the boundary is worth a sentence, because it is the part that looks wrong. A candidate whose optimal weight is zero adds nothing to the combination — that is what “adds nothing” means — and a candidate with a large weight also adds nothing to the combination that already contains it. The null being tested is not “this forecast is useless”; it is “this forecast has nothing the combination has not got”, and against the optimal combination that is true of everything in the set.

There is one more property of this configuration that makes it unusually good to measure in, and it is the reason the arrangement is worth the algebra. Nothing about it has been tuned. The persistence is inherited from the accuracy search, the candidate set is the one that field already uses, and the weights are a closed-form solve of a closed-form covariance matrix — so there is no parameter anywhere that could have been chosen to make a rate come out at 5%. A size measured at a configuration somebody searched for is a measurement of the search as much as of the procedure, and this one was not searched for.

Four readings of the same eight comparisons

Run the eight tests on the same sixty origins, two thousand times, at a nominal 5%.

Eight encompassing nulls, all of them true at once. The forecast under test is the variance-minimising combination of the eight candidates, and the first-order condition that defines its weights is cov(e_c, e_j) = var(e_c) for every j — so every one of the eight nulls is exactly true simultaneously and every rejection counted here is false. 800 draws of 60 origins. The largest of the eight statistics rejects 25.0% of the time at a nominal 5% and one comparison stated in advance rejects 3.1%, which makes the set worth about 9.1 independent comparisons — nearly the eight it has. Bonferroni, which is far inside its level on a search over accuracy comparisons of the same eight forecasters, is at 4.8% here.
Fig. 2 Four readings at a configuration where every rejection is false. The first is what a table of eight statistics with the largest one highlighted reports.

The largest of the eight statistics, read as though it had been the only one, rejects 26.2% of the time. One comparison stated in advance rejects 3.5%. Bonferroni over the eight rejects 5.2%, and a reference distribution over the whole set — the same block bootstrap of rows the accuracy search uses — rejects 4.5%.

On average 0.55 of the eight individual comparisons are called significant per draw, against the 0.40 that eight independent tests at their level would give.

The critical value behind those rates is worth quoting, because it is the number a reader would otherwise take off a table. The largest of the eight statistics is centred at 1.248 under this null — the maximum of eight statistics is not centred where each of them is — and its 95% point is 2.589, against 1.671 for a single comparison at this many origins. A table of eight encompassing tests with the biggest one starred is reporting a statistic whose distribution is nearly a full unit to the right of the distribution it is being read against.

Eight weights, and six of them are nothing. The weight each of the eight moving averages carries in the variance-minimising combination, across the persistence of the series, from the closed-form error covariance with no simulation in it. Two curves matter and six lie on the axis: the last value and the 34-observation mean carry between them at least 1.001 of the weight everywhere in the range, and the largest weight any of the other six ever carries is 0.0426 in magnitude. The reason is that the optimal forecast of this series is the last value shrunk towards the mean, and those two are the ingredients that make it — so a set of eight candidates is a set of two, and which two has nothing to do with which of them forecasts best on its own.
Fig. 3 The weights the eight nulls are about, across the persistence of the series. Six of the eight candidates carry nothing at every persistence, and all eight are exactly on the boundary anyway.

Why the count comes out above eight

An effective count of 8.63 on a set of eight is worth stopping on, because eight is supposed to be the ceiling: the measure asks how many independent comparisons would produce the observed pair of rates, and eight independent comparisons is the most eight comparisons can be worth.

It exceeds it because the eight statistics are not merely uncorrelated — they are negatively constrained. Multiplying the j-th encompassing quantity by its own weight and summing gives

Σⱼ wⱼ·[var(e_c) − cov(e_c, eⱼ)] = v − v = 0

exactly, by the same first-order condition that put every null on the boundary. So the eight quantities are pinned to a weighted sum of zero: whenever some of them come out high, the rest are obliged to come out low.

A set of statistics pushed apart in that way has a larger maximum than the same number of independent ones, so the naive rate rises relative to the single-comparison rate, so the measure returns a count above the nominal one. The set is worth more than eight independent comparisons because it is worth eight comparisons that cannot all be quiet at once.

That is a structural fact rather than a measurement artefact, and it is the exact opposite of the accuracy search’s situation, where the eight are strongly positively correlated and worth 2.44. The same eight forecasters are sub-additive on one question and super-additive on the other, for the same reason in two directions: what makes them alike as forecasts is what makes their contributions to a combination complementary.

The correction’s error is a factor of fifty for a factor of three in the count

Bonferroni charges for eight in both settings, and the two effective counts are 8.63 and 2.44 — a factor of 3.5.

Its realised sizes are 5.2% here and 0.1% on the accuracy search — a factor of 52.

So a threefold error in the count buys a fiftyfold error in the rate, which is the steepness of a normal tail expressed as a design fact. A correction that is charging for the wrong number of comparisons is wrong in the rate by far more than it is wrong in the count, and the direction is always the same: over-charging costs orders of magnitude and under-charging gains them.

That is the practical reason to measure the overlap rather than assume it. A reference distribution built by resampling the set is within a point of its level on both questions, and it gets there by computing a quantity that the count-based correction has to guess — a quantity that, on one identical set of eight forecasters, takes values on both sides of eight.

A multiplicity three times the one the same eight carry

Read the first two numbers as the site’s usual measure of how many independent comparisons a set is worth — how many separate chances would produce a 26.2% winner’s rate when one comparison gives 3.5% — and the answer is 8.63 of a possible 8.

The identical measurement on the identical eight forecasters, comparing them by mean squared error instead, gives 2.44.

That is the same set of candidates, the same series, the same number of origins and the same nominal level, and the correct charge for having looked at all eight differs by a factor of three and a half depending on which question is being asked of them.

The reason is what each statistic is a statement about. An accuracy comparison asks about σj2\sigma_j^2, and the eight variances move together almost perfectly: they are windows of one series, a long window contains the short ones, and a sample in which one of them looks good is a sample in which they all do. An encompassing comparison against the combination asks about cov(ec,ecej)\operatorname{cov}(e_c, e_c - e_j) — what candidate j has that the combination has not — and those eight quantities are nearly orthogonal, because the thing each candidate has that the others do not is, by construction, the thing that distinguishes it.

Eight windows of one series, and what a search over them costs. 320 set comparisons of 100 origins, all eight candidates on the same series at φ = 0.4895, so they are eight views of the same shocks. Five rejection rates at a nominal 5%. Reporting the best candidate's own p-value rejects 6.9% where one stated comparison rejects 2.8%, which makes the set behave like 2.50 independent comparisons rather than eight. A Bonferroni correction charges for eight and lands at 0.3%; the bootstrap over the whole set, which prices the dependence rather than assuming it away, is at 0.9%, and the recentred version at 1.3%.
Fig. 4 The accuracy search on the same eight candidates, where Bonferroni lands at a tenth of its level because it is charging for eight chances that are worth two and a half.

The practical consequence is a warning about corrections rather than about tests. Bonferroni over these eight is far inside its level on one question and at its level on the other; a reference distribution built by resampling the set is close to right on both, because it measures the overlap rather than assuming a value for it. A correction that charges by the number of models is charging for a quantity that is not the number of models, and how wrong it is depends on which question the models are being asked.

An encompassing test at a null that is exactly true. The forecast being tested is the variance-minimising combination of the last value and the 60-observation window mean, which encompasses each of its parts by the first-order condition that defines the weight — so every rejection counted here is a false one. Four rejection rates at a nominal 5%, over 1,800 comparisons of 120 origins each. Given the true variance of its own mean the statistic is at 3.7% in the upper tail and 5.8% in the lower. With that variance estimated from the same 120 numbers it is at 7.9% and 3.7%: the same total, moved from one side to the other. The standardised statistic's skewness is -0.355, and a denominator correlated with the numerator it divides is what puts it there.
Fig. 5 The two-forecast version of the same null, where the same construction makes one null exactly true and the failure that shows up is the denominator’s rather than the multiplicity’s.

The six exactly-redundant candidates raise the question the accuracy search has its own answer to: what do candidates that are certain not to win do to a procedure? There, adding sixteen stale forecasts nobody would have run took a reality check from 34.3% power to nothing at all, because the recentring is built on the assumption that every candidate might be on the boundary and stale ones are nowhere near it.

Here the situation is inverted and instructive. The six candidates carrying no weight are exactly on the boundary — they add exactly nothing, which is precisely what the null says — so they are not junk in that sense at all. They cost the procedure the multiplicity of six comparisons and contribute no chance of a true rejection, which is the same arithmetic wearing a different face: a set procedure pays for every candidate and collects only from the ones that can move.

What estimating eight weights costs, against not estimating them. Three combinations of the same eight forecasts, scored on the population covariance so that the number is the expected squared error of the combination that would actually have been used. The flat line is the equal-weight average, which estimates nothing. The curve is the variance-minimising combination estimated on P origins, and it is worse than the flat line until P is about 120 — eight weights estimated from sixty pairs of errors cost more than knowing them is worth. The lower line is the population optimum, 0.7817, which no sample reaches. The upper is what picking the single best forecaster in the same sample delivers, and it is the worst of the four at every length. 400 draws at each of 6 sample sizes.
Fig. 6 What estimating those weights costs, which is the practical reason the combination under test here is the population one: at sixty origins an estimated combination is worse than an equal-weight average.

What it detects, and what it lets past

A size measurement is half a check. The other half is what the set-level test does when the combination under test is not the optimal one, and this is easy to arrange: test the same eight candidates against a combination formed with the wrong weights.

Three of them, in order of how wrong they are.

The optimal combination is 0.0% above the optimum by definition, and the set test rejects 4.5% of the time. That is the size.

The equal-weight average — which is the best thing available to somebody who refuses to estimate anything, and beats every single forecaster in the set — sits 9.8% above the optimum. The set test catches it 21.5% of the time.

The best single forecast, chosen by mean squared error, sits 29.9% above the optimum. The set test catches it 96.3% of the time, and on average 5.66 of the eight candidates are individually flagged as adding something to it.

What a set-level encompassing test detects, and what it lets past. Three forecasts made of the same eight candidates, each tested against every one of the eight for whether that candidate adds anything to it. The first is the variance-minimising combination, where the answer is no by the first-order conditions and the 4.3% is the test's size. The second is the equal-weight average, which is 9.8% above the optimum and is caught 23.0% of the time. The third is the single forecast with the smallest mean squared error, 29.9% above the optimum, caught 96.3% of the time. 300 draws of 60 origins each. A combination that is badly wrong is detected nearly always and a nearly-right one is not.
Fig. 7 The same set of eight tests against three different forecasts. The first bar is a size and the other two are powers, and the ordering between them is the ordering of how far each is from the best use of the eight.

The gap between the second and third rows is the useful part. A forecast that is badly wrong — a single candidate where a combination was available — is caught almost always, and the eight tests between them name most of the set as adding something, which is a diagnosis and not merely a rejection. A forecast that is nearly right is let past four times out of five. A set-level encompassing test is an instrument for detecting that a combination is missing something large, and it is not an instrument for tuning one.

What a set-level encompassing test detects, and what it lets pastThree forecasts made of the same eight candidates, each tested against every one of the eight for whether that candidate adds anything to it. The first is the variance-minimising combination, where the answer is no by the first-order conditions and the 6.0% is the test's size. The second is the equal-weight average, which is 9.8% above the optimum and is caught 55.0% of the time. The third is the single forecast with the smallest mean squared error, 29.9% above the optimum, caught 100.0% of the time. 300 draws of 120 origins each. A combination that is badly wrong is detected nearly always and a nearly-right one is not.the optimal combination6.0%an equal-weight average55.0%the best single forecast100.0%every null true — this is the size9.8% above the optimum29.9% above the optimumthe forecast being tested300 draws of 120 origins, 8 candidates eachthe reference distribution over the set
Fig. 8 The same three at twice as many origins. Drag it: the size stays where it is and both powers rise, which is what a procedure holding its level while getting more data is supposed to look like.

Why the equal-weight average is hard to catch

The 21.5% deserves an explanation rather than a shrug, because 9.8% above the optimum is not a small error and sixty origins is not a tiny sample.

The equal-weight average is wrong in a very particular way: it puts a sixth of its weight on the last value where the optimum puts a half, and spreads the rest over six candidates the optimum ignores. What the eight tests are asked is whether any individual candidate adds something to it, and the answer for six of the eight is “hardly” — those six are nearly redundant against each other whatever the weights are. The whole of the deficiency is concentrated in the two ends, so the set test is effectively running two informative comparisons and six uninformative ones, and paying the multiplicity for all eight.

That is visible in the same runs from the other direction: one comparison stated in advance — the last value, which is the candidate the equal-weight average underweights most — rejects 61.0% of the time against the set procedure’s 21.5%. A pre-stated comparison beats a search here by forty points, which is the ordinary price of multiplicity arriving in an unusually sharp form: the search is paying for six comparisons that cannot succeed.

What a rejection would actually mean

It is worth being explicit about what this procedure is for, because “reject the hypothesis that every candidate is redundant” is an awkward sentence to act on.

The set of m encompassing tests against a forecast in production answers one operational question: is there anything in this set of candidates that my current forecast has not got? A rejection says yes and names the candidate that supplied the maximum. That is a shortlist of one, arrived at with the search paid for — which is more than a table of eight statistics with the best one starred provides, and considerably less than a combination.

What it does not answer is how much better the combination would be, which is a question about the weight vector and not about any test. The measurements above give the honest ordering: a forecast 9.8% above the best available use of the set is let past four times in five, and one 29.9% above it is caught almost always. So a non-rejection is evidence that nothing in the set is badly missing, and nothing more than that.

And the asymmetry the two halves of this field share is worth restating. The procedure that holds its level here is the one that resamples the observed comparisons, because these candidates estimate nothing and the null is a statement about the data. Where the candidates are fitted models that same construction recentres in the wrong place and stops rejecting altogether. One set of eight, two questions, and the reference distribution that is right for each of them is the one the other question would be broken by.

What is claimed here, and what is not

This essay takes the size and power of a search over encompassing comparisons, at a configuration where the least favourable case is attained rather than assumed.

What stays out and is named as a decision: joint tests of the whole weight vector — a Wald test that all m − 1 free components equal a stated value, which is one comparison rather than m and a different object; encompassing against a combination whose weights were estimated, where the null is no longer exactly true and the noise in the weights is inside the statistic being tested; and the sequential version, in which candidates are added to a combination one at a time and the multiplicity becomes a stopping problem.

The boundary against the set field is which quantity the maximum is taken over. Everything about the block bootstrap of rows, the studentising, and the recentring is established there and used here unchanged; what changes is that the eight statistics being maximised over are nearly independent rather than nearly identical, and every conclusion that field reached about corrections reverses because of it.

The checks, and the refusals that make them mean something

Three claims are gated in this field’s library. The first-order condition is checked as an identity: the combined error’s covariance with every member is required to equal the combined error’s own variance to within 10⁻⁹, at four persistences, which is what makes every null exactly true rather than approximately so. The winner’s own p-value is required to reject several times as often as one comparison stated in advance, and Bonferroni is required to be within two and a half points of its nominal level — which is the contrast with the accuracy search stated as something that could fail. And the set-level test is required to be at its level against the optimal combination, to catch the best single forecast most of the time, and to catch the equal-weight average much less often, in that order.

The refusals that stand behind them belong to this field’s other half and apply here unchanged: a reference distribution recentred where the null does not put it, and a combination scored on the origins that chose it. Both would produce an apparently well-behaved set procedure at this configuration, and neither would be measuring anything.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Benchmark forecastBlock bootstrapBonferroniCombination weightEffective sample sizeError rateForecast combinationForecast encompassingLeast favourableMean squared errorMultiplicityNull hypothesisPowerReference distribution