The best of a set, and what the search costs

The models that were never in the running

A reference distribution for a set has to assume something about every candidate in it. Assuming that all of them are as good as the benchmark is what makes the reality check honest, and it is what sixteen hopeless candidates use to destroy it.

Worth reading first: What the correction corrects · What the model says next.

The null a set comparison tests is not a statement, it is a set of statements. No candidate beats the benchmark is true if every candidate is worse and also true if one of them is exactly tied and the rest are hopeless — and a reference distribution has to be built for a particular one of those configurations before any p-value can come out of it.

The choice made by the standard construction is the least favourable one: assume every candidate in the set is exactly as good as the benchmark. That is the configuration under which the maximum is largest, so a procedure calibrated there cannot reject too often under any other, which is what makes it honest.

It is also an assumption about candidates the data has already disposed of, and once a set contains enough of those, the assumption is the whole answer.

Adding candidates that cannot win

Start from a situation where there is something to find. At φ = 0.65 the best member of the candidate set is genuinely better than the benchmark, and the question is how often each procedure notices.

Now add candidates that are unambiguously worse. The ones used here are stale: the last value carried forward, read two steps late, four steps late, six steps late and so on — a forecaster whose data arrives after everybody else’s, which is a real thing rather than a sabotage. The mildest of them is 43% worse than the benchmark in expected squared error and the worst is 97% worse. None of them is ever the best candidate in a sample.

Sixteen candidates nobody would have run, and what they cost. At φ = 0.65 one candidate in the original set is genuinely better than the benchmark, and the question is how often each procedure finds it. The added candidates are stale copies of the last value — read two, four, six … steps late — every one of them worse than the benchmark by at least 43%, and not one of them is ever the best candidate in a sample. The reality check goes from 35.5% to 0.0% as they are added, because its reference distribution has to assume every candidate is exactly as good as the benchmark and sixteen such assumptions is a critical value nothing reaches. The recentred version, which drops from the recentring the candidates the data has already ruled out — 15.1 of 24 of them — goes from 25.5% to 24.5%. The top line never moves: reporting the winner's own p-value cannot notice a change to a set it never looks at.
Fig. 1 Three hundred set comparisons at each point, a hundred origins each. The horizontal axis is how many stale candidates were added to the eight. Nothing about the data changes along it; only the size of the set does.

The reality check goes from 34.3% to 6.0%, 1.7% and 0.0% as four, eight and sixteen stale candidates are added. With sixteen of them it never rejects at all in three hundred comparisons. The genuinely better forecaster is still there, still better by the same margin, still winning the sample as often as it did; the procedure has simply stopped being able to say so.

The top line of that figure does not move: reporting the winner’s own p-value gives 61.0% regardless, because it looks at one row and cannot notice what happened to the other twenty-three. Bonferroni degrades gently — 17.7%, 14.0%, 10.7%, 8.0% — which is the honest 1/m cost of a correction that charges by the count.

Why the collapse happens, exactly

The reference distribution is the distribution of the largest of m statistics under the assumed configuration. Assume all m are on the boundary, and each added candidate is another draw that could be the maximum. A stale forecaster’s loss differential is enormously variable — it is a bad forecast, so its errors are large and so is the spread of the difference — and a large-variance draw recentred onto the boundary produces large bootstrap maxima routinely.

So the critical value climbs with every candidate added, and it climbs fastest for the candidates that are least plausible. That is the mechanism in one sentence: the assumption that protects the procedure’s error rate is an assumption about the candidates that contribute nothing, and the worse they are the more the assumption costs.

Eight candidates, one of them exactly as good as the benchmark. The candidate set: moving averages of the last 1, 2, 3, 5, 8, 13, 21 and 34 observations, each drawn as its expected squared error divided by the benchmark's — the mean of all 60. The persistence is not chosen, it is solved for: at φ = 0.4895 the best candidate in the set, the average of 2, has exactly the benchmark's expected squared error, and every other candidate is worse by between 0.5% and 5.1%. So the null that no candidate beats the benchmark is true, with one candidate on its boundary. Everything a set comparison claims about its own error rate has to be measured here, because anywhere further inside the null every procedure flatters itself.
Fig. 2 The original set at its own solved null, before anything stale is added. Seven of the eight are inside the null rather than on it — which is already the same effect in miniature, and is why the reality check reads 1.1% at a null where it claims 5%.

The miniature version is visible in the previous essay’s numbers and is worth naming here. At the solved null the reality check rejects 1.1% against a claimed 5%, and the reason is the same: one candidate is on the boundary and seven are between half a percent and five percent inside it, so the least favourable assumption is already wrong about seven eighths of the set. The stale candidates do not introduce a new defect. They take an existing one from a factor of five to a factor of infinity.

Removing what the data has ruled out

The repair is to stop assuming the boundary for candidates the data has placed far from it. A candidate whose standardised mean differential is below −√(2 log log P) keeps its own sample mean in the recentring, which puts it so far below zero that it cannot contribute to the bootstrap maximum; every other candidate is treated as before.

The threshold is a rate rather than a tuning constant — it is the width of the band a sample mean can wander in without the procedure being able to tell it from the boundary — and it is what makes the recentred version’s error rate converge to the right thing rather than to something smaller.

Sixteen candidates nobody would have run, and what they costAt φ = 0.65 one candidate in the original set is genuinely better than the benchmark, and the question is how often each procedure finds it. The added candidates are stale copies of the last value — read two, four, six … steps late — every one of them worse than the benchmark by at least 43%, and not one of them is ever the best candidate in a sample. The reality check goes from 35.5% to 0.0% as they are added, because its reference distribution has to assume every candidate is exactly as good as the benchmark and sixteen such assumptions is a critical value nothing reaches. The recentred version, which drops from the recentring the candidates the data has already ruled out — 15.1 of 24 of them — goes from 25.5% to 24.5%. The top line never moves: reporting the winner's own p-value cannot notice a change to a set it never looks at.00.2000.4000.60004816candidates added to the set that nobody would have runshare of comparisons in which the better forecaster is foundflat: the winner's own p-value · upper curve: recentred · lower curve: the reality check200 comparisons of 100 origins at φ = 0.6535.5% → 0.0% against 25.5% → 24.5%
Fig. 3 The same picture with the recentred procedure on it. Drag the length of the evaluation: the recentring is worth more the longer the comparison runs, because a longer comparison is more certain about which candidates are hopeless.

With sixteen stale candidates the recentring rules out 15.1 of the twenty-four, and the procedure’s power goes from 26.0% with no junk to 24.3% with sixteen — against the reality check’s 34.3% to 0.0%. Almost nothing is lost, and what is lost is the one stale candidate the data has not yet placed with confidence.

And what it does at the null

Power is half the picture. Run the same sweep at the solved null — where no candidate beats the benchmark and every rejection is false — and the procedures do not merely become cautious, they become inert.

Sixteen candidates nobody would have run, and what they cost. At φ = 0.49 one candidate in the original set is genuinely better than the benchmark, and the question is how often each procedure finds it. The added candidates are stale copies of the last value — read two, four, six … steps late — every one of them worse than the benchmark by at least 74%, and not one of them is ever the best candidate in a sample. The reality check goes from 1.0% to 0.0% as they are added, because its reference distribution has to assume every candidate is exactly as good as the benchmark and sixteen such assumptions is a critical value nothing reaches. The recentred version, which drops from the recentring the candidates the data has already ruled out — 16.5 of 24 of them — goes from 1.5% to 0.5%. The top line never moves: reporting the winner's own p-value cannot notice a change to a set it never looks at.
Fig. 4 The junk sweep at the persistence where the best candidate exactly ties the benchmark. These are error rates rather than power, and the two lines that fall are falling below a level nobody asked them to go below.

The reality check goes from 2.0% with no junk to 0.0% with eight and with sixteen, against a claimed 5%. The recentred version holds at 1.0% throughout, and the count of candidates it removes rises exactly as the count added does — 0.7, 8.5 and 16.4 — which is the recentring doing the one thing it was built to do and doing it cleanly.

An error rate of zero is not a success. It is the same fact as the power collapse seen from the other side: a procedure that never rejects has a perfect error rate and no use, and the two numbers have to be read together or the wrong one gets reported. The winner’s own p-value, meanwhile, sits at 9.0% at all three set sizes — nearly twice its claim, and stable, because it has never been a procedure for a set at all.

The two changes, separated

The recentred procedure differs from the reality check in two ways and only one of them is the recentring. It also studentises: each candidate’s mean differential is divided by its own long-run standard deviation before the maximum is taken. Reporting the pair as though it were one improvement would credit each change with the other’s effect, so this field’s library carries a third variant that exists for no other purpose — studentised, not recentred.

Read across the same sweep: with no junk the three sit at 34.3%, 26.0% and 26.0%, so studentising costs eight points and the recentring gains nothing where there is nothing to remove. With sixteen stale candidates they are 0.0%, 17.7% and 24.3%: studentising alone saves most of the collapse and the recentring saves the rest.

That decomposition is the useful result rather than a ranking of the two named procedures. On this set, studentising is a cost when the set is clean and a large protection when it is not, because dividing by each candidate’s own spread is exactly what stops a wildly variable stale forecaster from dominating a maximum.

Eight windows of one series, and what a search over them costs. 320 set comparisons of 100 origins, all eight candidates on the same series at φ = 0.4895, so they are eight views of the same shocks. Five rejection rates at a nominal 5%. Reporting the best candidate's own p-value rejects 6.9% where one stated comparison rejects 2.8%, which makes the set behave like 2.50 independent comparisons rather than eight. A Bonferroni correction charges for eight and lands at 0.3%; the bootstrap over the whole set, which prices the dependence rather than assuming it away, is at 0.9%, and the recentred version at 1.3%.
Fig. 5 The clean set at its null, where the studentisation costs and the recentring is idle. Any comparison of the two procedures that stops here reports the wrong ordering.
Where the four readings agree, and where they do not. The same four procedures across the persistence of the series, 200 set comparisons of 100 origins at each of five points. The leftmost point is the solved crossing, where the null is true: the winner's own p-value rejects 6.5%, the Bonferroni correction 0.0%, the bootstrap over the set 2.0%. Everywhere to the right of it the best candidate is genuinely better and the same numbers are power: by φ = 0.77 the bootstrap finds it 83.0% of the time against Bonferroni's 59.5%. The gap between those two is the whole value of pricing the dependence rather than assuming it away, and it is largest in the middle of the range where a decision is actually difficult.
Fig. 6 And the clean set above its null, where the ordering is the same and is still not the one a set with junk in it produces.

A stale forecaster is a real thing

The candidates added here were chosen to be junk for a stated reason rather than to be junk, and the distinction matters for how far the result travels.

A stale moving average is the last value carried forward, read two or four or six steps late. It is what a forecaster produces when its inputs arrive after everybody else’s — a survey published with a month’s delay, a settlement price available the following week, a data vendor whose revisions land late. Nothing about it is degenerate: it is a competent forecast of the wrong quantity, its errors are the errors of a good rule applied at the wrong moment, and its expected squared error is a closed form in φ and the lag like every other candidate’s here.

That is what makes the measurement about the composition of the set. A candidate built by adding noise to a forecast, or by inverting its sign, would be junk in a way that could be dismissed as contrived; these are the eight members of the family read at a delay, and an evaluation that includes a late-arriving vendor’s forecast alongside seven in-house ones has exactly this set.

The mildest of them is 43% worse than the benchmark, which is not marginal, and none of them wins a sample. The whole effect comes from candidates that any reader of the table would discard at a glance, and the procedure is unable to discard them because it was built not to look.

The composite null is where the difficulty is

Everything above is a consequence of a fact about the hypothesis rather than about forecasting. A null of the form max E[dₖ] ≤ 0 is composite: it is satisfied by an infinite family of configurations, and a test’s size is the supremum of its rejection rate over all of them. Any procedure that attains that supremum is conservative everywhere else, and how conservative depends entirely on where the truth actually sits.

The multiplicity field meets the same structure in a different costume. Bonferroni bounds the chance of any false positive under the configuration where every null is true, and is conservative when some are false; the step-down version recovers part of that by using the data to decide which hypotheses are still live, which is the same move the recentring makes here. What is different is the size of the effect: dropping hypotheses from a Holm procedure buys a few points, and dropping candidates from this reference distribution is the difference between 24% power and none.

Eight separate problems, and what a search over them costs. 320 set comparisons of 100 origins, each candidate on its own series at its own exact crossing, so all eight are on the boundary of the null and independent of one another. Five rejection rates at a nominal 5%. Reporting the best candidate's own p-value rejects 29.4% where one stated comparison rejects 4.1%, which makes the set behave like 8.39 independent comparisons rather than eight. A Bonferroni correction charges for eight and lands at 3.1%; the bootstrap over the whole set, which prices the dependence rather than assuming it away, is at 3.4%, and the recentred version at 2.5%.
Fig. 7 The configuration where the least favourable assumption is exactly right: eight independent problems, every candidate on the boundary. Here the reality check is at its nominal level and has nothing to apologise for, because nothing has been assumed that is not true.

The size of the effect is what is new, not the shape of it

Conservatism under a composite null is not news. Every correction in the multiplicity field is conservative when some of its hypotheses are false, and spending an error rate across looks at a trial is the same arithmetic in a different order. What makes this case worth a field of its own is the size and the direction of the dependence on something the analyst controls.

Three quantities decide how much a set comparison can find, and only one of them is about the data. How many candidates there are: Bonferroni charges for it linearly and the bootstrap charges for what it is worth. How alike they are: the previous essay measures that as an effective count, 2.44 for eight windows of one series. And how far inside the null the hopeless ones sit: which costs nothing under Bonferroni, everything under the reality check, and almost nothing once the recentring is in place.

An analyst can change the third of those without collecting an observation. That is the property that makes it worth measuring rather than deriving, and it is the reason a set comparison should report the composition of its set with the same care a trial reports its stopping rule.

The recentring, read as a classifier

The removal counts at the null — 0.7, 8.5 and 16.4 candidates ruled out when nothing, eight and sixteen stale members are added — are reported as evidence the recentring does what it was built to do, and differencing them says how well.

With nothing added it removes 0.7 of the eight genuine candidates. With eight stale ones added it removes 8.5, so 7.8 of the eight additions; with sixteen, 15.7 of the sixteen. Read as a classification of candidates into hopeless and live, that is a sensitivity of 97.5% and 98.1% against a false-positive rate of 8.8%.

Those are the numbers that make the procedure work, and they are better than the power figures alone suggest. A rule that discarded 90% of the junk would leave one or two wildly variable candidates in the reference distribution, and one is enough to lift the maximum — the reality check’s collapse is what a set with sixteen of them does, and most of that damage is done by the first few. Ruling out better than nineteen in twenty is what turns the repair from partial into near-total.

The false-positive rate is the cost and it is the one visible in the power column: 8.8% of the genuine candidates are dropped from the reference distribution when they should not be, which is most of the 1.7 points the recentring gives up between a clean set and a set with sixteen additions.

Four hopeless candidates reverse the ordering

The three procedures can be ranked two ways and the two rankings are exactly opposite.

On a clean set the reality check finds the better forecaster 34.3% of the time, the recentred procedure 26.0% and Bonferroni 17.7%. On robustness to what is added, the fractions of their own power retained after sixteen stale candidates are 0%, 93.5% and 45.2%. Best to worst on the first reading is worst to best on the second, with no crossings.

So the crossover between the top two is early. The reality check is ahead at nothing added and behind at four — it has fallen to 6.0% by then, and the recentred procedure cannot have fallen below the 24.3% it still holds at sixteen. Four candidates nobody would have run are enough to reverse which procedure is better, on a set of eight, and the reversal is by a factor of four rather than a margin.

That is a small enough number to be an ordinary accident of how an evaluation was assembled. A table of eight in-house specifications plus four vendor feeds that arrive late is not a contrived set; it is what a forecasting group has, and the two procedures it might apply to it have swapped places by the time the fourth vendor is added.

Bonferroni’s position is worth naming separately, because it is the one procedure the comparison retires. It is behind the recentred version at every set size measured — 17.7% against 26.0% clean, 8.0% against 24.3% with sixteen additions — so there is no composition at which charging by the count is the thing to do. It is dominated, and its only remaining virtue is that it needs no bootstrap, which is a computational argument rather than a statistical one.

What this says about assembling a set

The practical reading is uncomfortable and worth stating plainly. Under the reality check, adding a forecaster already believed to be useless makes it harder to detect the one that is not. That inverts the usual instinct about search — that a wider search is a more thorough one — and it is not a small effect: sixteen candidates nobody would run took a procedure from finding a real improvement a third of the time to never. The same instinct has been corrected once before on this site and in the other direction — more data is not monotone — and the shape of the correction is the same: a quantity everybody expects to move one way turns out to depend on something nobody was tracking.

It also means the set has to be declared, because a set is not a neutral container. Two analysts with the same data, the same benchmark and the same favoured model can reach opposite conclusions by including different numbers of hopeless alternatives, and both of them will have applied the same correctly-calibrated procedure.

The two tests, run on the same comparisons. Six persistences, 350 comparisons at each, and two tests on every one of them, at 1 step ahead. The curve that dips is the accuracy comparison: at φ = 0.4922 the two forecasts have identical expected squared error and it is measuring its own size, 7.1%, rising away from there because the difference is real. The other two are the encompassing tests in each direction, and at that same crossing they reject 100.0% and 98.0%. Neither forecast encompasses the other at any persistence — that is a closed form with no simulation in it — so every number on those two curves below 100% is the test running out of power rather than the hypothesis changing, and at long horizons there is a great deal of that: sixty origins at 1 step ahead carry about 60 independent comparisons.
Fig. 8 The pair case, for scale: two forecasters and two questions, where nothing about the composition of a set can change the answer. Everything in this essay appears the moment there is a third forecaster.
The distribution the table does not have. 599 series simulated from the smaller model fitted to one comparison's own data, the whole rolling comparison re-run on each, and the ordinary statistic recorded. Under this null the two forecasts are the same forecast in population, so what is left in a sample is the larger model's estimation error and the statistic is centred at -1.134 rather than at zero. Its 95% point is 0.264; the standard normal drawn behind it puts that point at 1.645. Reading this statistic against that curve is not a poor approximation, it is a different distribution: the share of this one above 1.645 is 0.2%.
Fig. 9 The next essay’s picture: a reference distribution generated from the null rather than assumed. The same instrument, applied where the difficulty is not multiplicity but a statistic whose null distribution is nothing like the table it is read against.

What is claimed here, and what is not

This essay claims that the composition of a candidate set changes what a set comparison can detect, that the size of that effect is measurable, and that recentring on the candidates the data has ruled out recovers most of it while studentisation accounts for the rest.

What stays out: the choice of threshold in the recentring, which is a rate here and is treated as given; sets whose members are added in response to the results, which breaks every calibration in this field; and the question of which candidates should be in a set, which is a question about what is being claimed rather than about how to test it.

The checks, and the refusal that makes them mean something

Two claims are gated in this field’s library. Sixteen stale candidates are required to take the reality check to below a quarter of its own power, and the recentred version is required to keep above four fifths of its own — with the stale candidates required to be worse than the benchmark by a stated margin, so the experiment is about the composition of the set and not about how badly the additions were chosen.

Beside them sits the refusal from the previous essay, which applies here for a different reason: a bootstrap that resamples each candidate on its own indices prices a set whose members are independent. On a set with sixteen stale copies of the same forecaster in it, that assumption is not merely wrong, it is wrong in the direction that hides the effect this essay measures.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Benchmark forecastComposite nullData snoopingError rateFamilywise error rateLeast favourable configurationLoss differentialMonte CarloMultiple comparisonsReality checkRecentringStationary bootstrapStatistical powerSuperior predictive ability