The models that were never in the running
Worth reading first: What the correction corrects · What the model says next.
The null a set comparison tests is not a statement, it is a set of statements. No candidate beats the benchmark is true if every candidate is worse and also true if one of them is exactly tied and the rest are hopeless — and a reference distribution has to be built for a particular one of those configurations before any p-value can come out of it.
The choice made by the standard construction is the least favourable one: assume every candidate in the set is exactly as good as the benchmark. That is the configuration under which the maximum is largest, so a procedure calibrated there cannot reject too often under any other, which is what makes it honest.
It is also an assumption about candidates the data has already disposed of, and once a set contains enough of those, the assumption is the whole answer.
Adding candidates that cannot win
Start from a situation where there is something to find. At φ = 0.65 the best member of the candidate set is genuinely better than the benchmark, and the question is how often each procedure notices.
Now add candidates that are unambiguously worse. The ones used here are stale: the last value carried forward, read two steps late, four steps late, six steps late and so on — a forecaster whose data arrives after everybody else’s, which is a real thing rather than a sabotage. The mildest of them is 43% worse than the benchmark in expected squared error and the worst is 97% worse. None of them is ever the best candidate in a sample.
The reality check goes from 34.3% to 6.0%, 1.7% and 0.0% as four, eight and sixteen stale candidates are added. With sixteen of them it never rejects at all in three hundred comparisons. The genuinely better forecaster is still there, still better by the same margin, still winning the sample as often as it did; the procedure has simply stopped being able to say so.
The top line of that figure does not move: reporting the winner’s own p-value gives 61.0% regardless, because it looks at one row and cannot notice what happened to the other twenty-three. Bonferroni degrades gently — 17.7%, 14.0%, 10.7%, 8.0% — which is the honest 1/m cost of a correction that charges by the count.
Why the collapse happens, exactly
The reference distribution is the distribution of the largest of m statistics under the assumed configuration. Assume all m are on the boundary, and each added candidate is another draw that could be the maximum. A stale forecaster’s loss differential is enormously variable — it is a bad forecast, so its errors are large and so is the spread of the difference — and a large-variance draw recentred onto the boundary produces large bootstrap maxima routinely.
So the critical value climbs with every candidate added, and it climbs fastest for the candidates that are least plausible. That is the mechanism in one sentence: the assumption that protects the procedure’s error rate is an assumption about the candidates that contribute nothing, and the worse they are the more the assumption costs.
The miniature version is visible in the previous essay’s numbers and is worth naming here. At the solved null the reality check rejects 1.1% against a claimed 5%, and the reason is the same: one candidate is on the boundary and seven are between half a percent and five percent inside it, so the least favourable assumption is already wrong about seven eighths of the set. The stale candidates do not introduce a new defect. They take an existing one from a factor of five to a factor of infinity.
Removing what the data has ruled out
The repair is to stop assuming the boundary for candidates the data has placed far from it. A candidate whose standardised mean differential is below −√(2 log log P) keeps its own sample mean in the recentring, which puts it so far below zero that it cannot contribute to the bootstrap maximum; every other candidate is treated as before.
The threshold is a rate rather than a tuning constant — it is the width of the band a sample mean can wander in without the procedure being able to tell it from the boundary — and it is what makes the recentred version’s error rate converge to the right thing rather than to something smaller.
With sixteen stale candidates the recentring rules out 15.1 of the twenty-four, and the procedure’s power goes from 26.0% with no junk to 24.3% with sixteen — against the reality check’s 34.3% to 0.0%. Almost nothing is lost, and what is lost is the one stale candidate the data has not yet placed with confidence.
And what it does at the null
Power is half the picture. Run the same sweep at the solved null — where no candidate beats the benchmark and every rejection is false — and the procedures do not merely become cautious, they become inert.
The reality check goes from 2.0% with no junk to 0.0% with eight and with sixteen, against a claimed 5%. The recentred version holds at 1.0% throughout, and the count of candidates it removes rises exactly as the count added does — 0.7, 8.5 and 16.4 — which is the recentring doing the one thing it was built to do and doing it cleanly.
An error rate of zero is not a success. It is the same fact as the power collapse seen from the other side: a procedure that never rejects has a perfect error rate and no use, and the two numbers have to be read together or the wrong one gets reported. The winner’s own p-value, meanwhile, sits at 9.0% at all three set sizes — nearly twice its claim, and stable, because it has never been a procedure for a set at all.
The two changes, separated
The recentred procedure differs from the reality check in two ways and only one of them is the recentring. It also studentises: each candidate’s mean differential is divided by its own long-run standard deviation before the maximum is taken. Reporting the pair as though it were one improvement would credit each change with the other’s effect, so this field’s library carries a third variant that exists for no other purpose — studentised, not recentred.
Read across the same sweep: with no junk the three sit at 34.3%, 26.0% and 26.0%, so studentising costs eight points and the recentring gains nothing where there is nothing to remove. With sixteen stale candidates they are 0.0%, 17.7% and 24.3%: studentising alone saves most of the collapse and the recentring saves the rest.
That decomposition is the useful result rather than a ranking of the two named procedures. On this set, studentising is a cost when the set is clean and a large protection when it is not, because dividing by each candidate’s own spread is exactly what stops a wildly variable stale forecaster from dominating a maximum.
A stale forecaster is a real thing
The candidates added here were chosen to be junk for a stated reason rather than to be junk, and the distinction matters for how far the result travels.
A stale moving average is the last value carried forward, read two or four or six steps late. It is what a forecaster produces when its inputs arrive after everybody else’s — a survey published with a month’s delay, a settlement price available the following week, a data vendor whose revisions land late. Nothing about it is degenerate: it is a competent forecast of the wrong quantity, its errors are the errors of a good rule applied at the wrong moment, and its expected squared error is a closed form in φ and the lag like every other candidate’s here.
That is what makes the measurement about the composition of the set. A candidate built by adding noise to a forecast, or by inverting its sign, would be junk in a way that could be dismissed as contrived; these are the eight members of the family read at a delay, and an evaluation that includes a late-arriving vendor’s forecast alongside seven in-house ones has exactly this set.
The mildest of them is 43% worse than the benchmark, which is not marginal, and none of them wins a sample. The whole effect comes from candidates that any reader of the table would discard at a glance, and the procedure is unable to discard them because it was built not to look.
The composite null is where the difficulty is
Everything above is a consequence of a fact about the hypothesis rather than about forecasting. A null of the form max E[dₖ] ≤ 0 is composite: it is satisfied by an infinite family of configurations, and a test’s size is the supremum of its rejection rate over all of them. Any procedure that attains that supremum is conservative everywhere else, and how conservative depends entirely on where the truth actually sits.
The multiplicity field meets the same structure in a different costume. Bonferroni bounds the chance of any false positive under the configuration where every null is true, and is conservative when some are false; the step-down version recovers part of that by using the data to decide which hypotheses are still live, which is the same move the recentring makes here. What is different is the size of the effect: dropping hypotheses from a Holm procedure buys a few points, and dropping candidates from this reference distribution is the difference between 24% power and none.
The size of the effect is what is new, not the shape of it
Conservatism under a composite null is not news. Every correction in the multiplicity field is conservative when some of its hypotheses are false, and spending an error rate across looks at a trial is the same arithmetic in a different order. What makes this case worth a field of its own is the size and the direction of the dependence on something the analyst controls.
Three quantities decide how much a set comparison can find, and only one of them is about the data. How many candidates there are: Bonferroni charges for it linearly and the bootstrap charges for what it is worth. How alike they are: the previous essay measures that as an effective count, 2.44 for eight windows of one series. And how far inside the null the hopeless ones sit: which costs nothing under Bonferroni, everything under the reality check, and almost nothing once the recentring is in place.
An analyst can change the third of those without collecting an observation. That is the property that makes it worth measuring rather than deriving, and it is the reason a set comparison should report the composition of its set with the same care a trial reports its stopping rule.
The recentring, read as a classifier
The removal counts at the null — 0.7, 8.5 and 16.4 candidates ruled out when nothing, eight and sixteen stale members are added — are reported as evidence the recentring does what it was built to do, and differencing them says how well.
With nothing added it removes 0.7 of the eight genuine candidates. With eight stale ones added it removes 8.5, so 7.8 of the eight additions; with sixteen, 15.7 of the sixteen. Read as a classification of candidates into hopeless and live, that is a sensitivity of 97.5% and 98.1% against a false-positive rate of 8.8%.
Those are the numbers that make the procedure work, and they are better than the power figures alone suggest. A rule that discarded 90% of the junk would leave one or two wildly variable candidates in the reference distribution, and one is enough to lift the maximum — the reality check’s collapse is what a set with sixteen of them does, and most of that damage is done by the first few. Ruling out better than nineteen in twenty is what turns the repair from partial into near-total.
The false-positive rate is the cost and it is the one visible in the power column: 8.8% of the genuine candidates are dropped from the reference distribution when they should not be, which is most of the 1.7 points the recentring gives up between a clean set and a set with sixteen additions.
Four hopeless candidates reverse the ordering
The three procedures can be ranked two ways and the two rankings are exactly opposite.
On a clean set the reality check finds the better forecaster 34.3% of the time, the recentred procedure 26.0% and Bonferroni 17.7%. On robustness to what is added, the fractions of their own power retained after sixteen stale candidates are 0%, 93.5% and 45.2%. Best to worst on the first reading is worst to best on the second, with no crossings.
So the crossover between the top two is early. The reality check is ahead at nothing added and behind at four — it has fallen to 6.0% by then, and the recentred procedure cannot have fallen below the 24.3% it still holds at sixteen. Four candidates nobody would have run are enough to reverse which procedure is better, on a set of eight, and the reversal is by a factor of four rather than a margin.
That is a small enough number to be an ordinary accident of how an evaluation was assembled. A table of eight in-house specifications plus four vendor feeds that arrive late is not a contrived set; it is what a forecasting group has, and the two procedures it might apply to it have swapped places by the time the fourth vendor is added.
Bonferroni’s position is worth naming separately, because it is the one procedure the comparison retires. It is behind the recentred version at every set size measured — 17.7% against 26.0% clean, 8.0% against 24.3% with sixteen additions — so there is no composition at which charging by the count is the thing to do. It is dominated, and its only remaining virtue is that it needs no bootstrap, which is a computational argument rather than a statistical one.
What this says about assembling a set
The practical reading is uncomfortable and worth stating plainly. Under the reality check, adding a forecaster already believed to be useless makes it harder to detect the one that is not. That inverts the usual instinct about search — that a wider search is a more thorough one — and it is not a small effect: sixteen candidates nobody would run took a procedure from finding a real improvement a third of the time to never. The same instinct has been corrected once before on this site and in the other direction — more data is not monotone — and the shape of the correction is the same: a quantity everybody expects to move one way turns out to depend on something nobody was tracking.
It also means the set has to be declared, because a set is not a neutral container. Two analysts with the same data, the same benchmark and the same favoured model can reach opposite conclusions by including different numbers of hopeless alternatives, and both of them will have applied the same correctly-calibrated procedure.
What is claimed here, and what is not
This essay claims that the composition of a candidate set changes what a set comparison can detect, that the size of that effect is measurable, and that recentring on the candidates the data has ruled out recovers most of it while studentisation accounts for the rest.
What stays out: the choice of threshold in the recentring, which is a rate here and is treated as given; sets whose members are added in response to the results, which breaks every calibration in this field; and the question of which candidates should be in a set, which is a question about what is being claimed rather than about how to test it.
The checks, and the refusal that makes them mean something
Two claims are gated in this field’s library. Sixteen stale candidates are required to take the reality check to below a quarter of its own power, and the recentred version is required to keep above four fifths of its own — with the stale candidates required to be worse than the benchmark by a stated margin, so the experiment is about the composition of the set and not about how badly the additions were chosen.
Beside them sits the refusal from the previous essay, which applies here for a different reason: a bootstrap that resamples each candidate on its own indices prices a set whose members are independent. On a set with sixteen stale copies of the same forecaster in it, that assumption is not merely wrong, it is wrong in the direction that hides the effect this essay measures.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The corner the test is calibrated at — both name benchmark forecast, composite null, error rate, least favourable configuration, monte carlo, reality check, statistical power
- An order that spends the error rate — both name error rate, familywise error rate, monte carlo, multiple comparisons, statistical power
- False discoveries that arrive together — both name error rate, familywise error rate, monte carlo, multiple comparisons, statistical power
- Estimating how many nulls are true — both name error rate, monte carlo, multiple comparisons, statistical power
- The rank is a decision — both name error rate, monte carlo, multiple comparisons, statistical power
- What the other forecast adds — both name benchmark forecast, error rate, loss differential, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Benchmark forecastComposite nullData snoopingError rateFamilywise error rateLeast favourable configurationLoss differentialMonte CarloMultiple comparisonsReality checkRecentringStationary bootstrapStatistical powerSuperior predictive ability