The corner the test is calibrated at
Worth reading first: The experiments that could have happened · What the correction corrects.
Every test in the two fields before this one has a null that is a point. The extra coefficient is zero, the two forecasts are equally accurate, the treatment does nothing. A point null has one distribution and the test is calibrated at it.
A search has a null that is not a point. “No candidate in this table is better than the benchmark” is satisfied by every configuration in which each candidate is as good as the benchmark or worse, which is a region with a face and a corner, and the statistic’s distribution is different at every point of it.
A reality check picks the corner: the configuration in which every candidate is exactly as good as the benchmark, which is the least favourable one because it gives the maximum the most chances to be large. That choice is what makes the procedure valid everywhere in the region, and it is also the entire subject of this essay, because the data is usually not at the corner.
Moving off the corner without leaving the null
Six candidates, each two of four predictors, and a benchmark that is one of them. Give the first predictor — which the benchmark holds — a coefficient. Now the three candidates that also hold it are still exactly as good as the benchmark, and the three that do not are strictly worse.
The null is still true. Not one candidate in the table is better than the benchmark, at any setting. What has changed is only how many of them are on the boundary, and the count falls from 4.46 of five to 2.76 as the coefficient runs from zero to 0.4.
The rejection rate goes with it: 2.7%, 1.7%, 1.7%, 0.0%.
That is a procedure whose size depends on a nuisance configuration nobody chose and nobody can observe. It is conservative, so it is valid — a test that rejects less than its nominal level at every point of a composite null is exactly what “valid” means here — and the conservatism is not free.
What the conservatism costs
Put a real improvement in the table and the cost is visible immediately.
Give the benchmark’s first predictor a coefficient of 1, so that the three candidates missing it are hopeless. Then give a fourth predictor — one the benchmark does not hold — a coefficient of its own, so that some candidates genuinely are better. That is a false null with three hopeless columns in it, which is what any real table looks like: nobody searches over a set of models that are all equally plausible.
The reality check finds the improvement 0.0% of the time at a moderate edge and 2.7% at a large one. The same test with the hopeless columns recentred finds it 31.3% and 51.0%.
The mechanism is entirely mechanical. The reference distribution is the distribution of the maximum over five columns, each recentred at the boundary. Three of those columns are, in the data, hundreds of standard errors below the boundary — but the reference distribution has been told to treat them as though they were on it, so they contribute their full share of upward noise to the maximum. The critical value is set by five columns’ worth of maximum and the observed statistic is produced by two, and the test cannot bridge the gap.
The size falls faster than the count can explain
The two sequences in the section above invite a model, and the model gets the middle of the sweep right and the end of it wrong — which is where the interesting part is.
If the only thing that mattered were how many candidates sit on the boundary, the size would behave like the tail of a maximum over that many statistics. Calibrated at 2.7% with 4.46 of them on the boundary, that model predicts
which at m = 3.89, 3.33 and 2.76 gives 2.4%, 2.0% and 1.7%.
Measured: 1.7%, 1.7% and 0.0%.
So the count explains the first part of the fall and not the last of it. Between the third and fourth settings the boundary count drops by about half a candidate and the size drops from 1.7% to nothing.
Why the candidates that leave do not merely stop counting
The gap says something specific about the mechanism, and it is worth naming because it changes what the conservatism is.
A candidate that leaves the boundary does not become inert. It becomes strictly worse, and a strictly worse candidate still contributes to the bootstrap reference distribution the maximum is compared against — it is recentred, so its own contribution to the reference is a draw from the same null spread, while its contribution to the observed statistic has been pushed down.
So each departing candidate does two things at once: it removes one chance for the observed maximum to be large, and it leaves one full-sized draw in the distribution that maximum is measured against. A maximum-of-m argument counts only the first.
That is what makes the fall to zero possible at all. A pure count argument can never reach zero — with 2.76 candidates still exactly on the boundary there are still 2.76 chances — and the measured rate is a clean zero on the whole run, which on any realistic number of trials bounds the true size well under a per cent.
The conservatism is therefore not a mild loss of power that scales with how far the truth sits from the corner. It is a procedure that can switch off entirely at configurations that are still inside the null, and the switching-off is driven by the candidates the null is not about.
The repair, which is a threshold and an admission
The repair is Hansen’s, and it is one line: a column whose sample differential is far enough below zero is recentred at its own mean rather than at the boundary, so that the simulated version of it sits where the real one does and never wins the maximum.
“Far enough” is a threshold, and a threshold is a decision. The version used here drops a column when its differential is more than one standard error below zero, which keeps 1.80 of five columns at the boundary in the hardest configuration above and 4.46 when the table is at the corner.
Two things about that decision are worth being explicit about.
The procedure is no longer exactly valid. A column just barely below the threshold is treated as hopeless when it might be on the boundary, and at that configuration the test can exceed its level. Measured at the corner, where nothing is dropped, the recentred version rejects 2.7% — identical to the uncorrected one, because the two procedures coincide there. Measured with three columns hopeless it rejects 2.3% at a true null, which is inside the counting error of the level it claims.
And the gain is not marginal. Across the configurations measured here it runs from nothing, when the table is at the corner and there is nothing to drop, to a factor of ten when most of the table is hopeless. There is no setting in which the recentring loses.
The dial between the two, and where it turns
The gap between the two procedures is a function of how far behind the bad columns are, and it opens suddenly rather than gradually.
- Columns mildly behind: the reality check finds a real improvement 11.0% of the time, the recentred version 13.7%.
- Columns moderately behind: 1.7% against 23.3%.
- Columns hopelessly behind: 0.0% against 31.3%.
The reality check gets worse as the table gets easier to read. That is the sentence worth carrying away, and it is not a paradox once the mechanism is in hand: the columns that make a table easy for a human to read — the obviously bad ones — are exactly the columns that inflate the reference distribution and depress the test.
The practical form of this is that padding a table is not free in the direction people assume. The usual worry about adding candidates to a search is a false positive. Here, adding candidates that are obviously bad costs power and nothing else, and an analyst who includes a few hopeless specifications “for completeness” has made their own test unable to see anything.
What a composite null actually asks of a procedure
It is worth slowing down on the logic, because “the test is conservative off the corner” is easy to say and easy to mis-state in either direction.
A test of a composite null has to hold its level at every point of the null region — that is what the guarantee means. A point of the region where it rejects less than its nominal level is not a failure; it is the price of the guarantee at the points where it rejects at the level. So a procedure calibrated at the least favourable configuration is doing the only thing that can be done without further information, and calling it broken is wrong.
What is legitimate to complain about is that the least favourable configuration is being used as a fact rather than as a bound. The reference distribution here does not say “the maximum could be this large if every column were on the boundary”; it says “the maximum is distributed as it would be if every column were on the boundary”, and the data very often contains direct evidence against that. A column sitting eight standard errors below zero is not a column that might be on the boundary.
That is the sense in which the recentring is an improvement and not a loosening. It uses the data to narrow the region the test has to be valid over — from the whole null to the part of it consistent with what has been observed — and it pays for that with a threshold and a small amount of exactness. The trade is the same one every studentised procedure makes, and it is unusually easy to see here because the two versions coincide exactly at the corner and diverge only where the evidence is.
Why this is not the same problem as the reading
Three separate things can miscalibrate a search, and this field has now measured all three. They are independent and it is worth keeping them apart.
The reading decides which maximum is taken — over the rows against a stated benchmark, or over every pair, or against a benchmark that was itself chosen. It moves the same true null between 0.0% and 76.2%.
The displacement is a real difference in expected accuracy that exists before any search, and it comes from candidates of different sizes rather than from any search at all. It is what makes a multiplicity correction the wrong tool on a nested table.
And the configuration is this essay: where in the null region the data sits, which changes the distribution of the same statistic under the same reading with the same candidates.
A procedure can be right about all three or wrong about all three, and the failures do not announce which they are. A search that rejects too often may have the wrong reading or a displacement in it; a search that rejects too rarely may have a searched benchmark, or a table full of hopeless columns, or both.
How many columns are really being maximised over
The count of columns still on the boundary is the quantity that decides everything above, and it is worth looking at directly because it is the one thing in this essay a practitioner can compute from their own table.
At the corner it is 4.46 of five — not five, because the threshold occasionally drops a column whose sample differential wandered a standard error below zero by chance. That is the recentring’s own cost measured at the place it has nothing to gain, and it is why the two procedures agree exactly on rejection rate there rather than the recentred one being slightly better.
As the configuration moves it falls to 4.28, 3.81 and 2.76, and in the power measurements with three hopeless columns it sits at 1.80. So in the configuration where the reality check finds a real improvement 0.0% of the time, its reference distribution is the maximum of five columns while the honest number of live comparisons is fewer than two.
Two live comparisons and a critical value set by five is roughly the whole of the loss, and the size of it can be estimated from that alone: the 95th percentile of the maximum of five approximately independent statistics sits well above that of two, and the observed statistic never reaches it.
There is a tempting shortcut here — count the live columns and apply a correction for that many — and it is not the same procedure. The count is itself random and correlated with the statistic, which is exactly why the recentring is stated as a modification of the reference distribution rather than as a smaller multiplicity. Using the count directly is the mistake the first essay of this field calls reading a search by the result it produced.
Where the least favourable configuration is the truth
There is one setting on this site where the corner is not an assumption, and it is worth pointing at because it is the only measurement of a reality check that is free of everything above.
The variance-minimising combination of a set of forecasts satisfies a first-order condition that makes its error covary with every member’s at exactly its own variance, so every one of the encompassing nulls is on the boundary simultaneously. That is the least favourable configuration arriving as a fact about the arithmetic rather than as a choice about the calibration, and the field that found it measures what a reality check does there.
This field reaches the corner a second and much easier way: candidates of the same size, none containing any other, at a null where nothing is worth anything. Six two-predictor subsets of four predictors are all exactly as accurate as each other in population and all displaced by exactly the same amount, which is none. That is why every measurement in this field is run on that table.
The two routes matter because they say the corner is reachable rather than hypothetical, and a procedure calibrated at a configuration that cannot occur would be a different kind of object from one calibrated at a configuration that merely usually does not.
What is claimed here, and what is not
This essay takes the configuration a set test is calibrated at, and the claims are the size falling as the table leaves the corner, the power at three degrees of hopelessness, and the recentring’s gain and its level.
What stays out and is named as a decision: the choice of threshold, which is set at one standard error here and swept nowhere — the number matters and choosing it well is a literature rather than a measurement; a studentised version of the same test, which changes the shape of the reference distribution and interacts with the recentring in a way nothing here separates; and any statement about tables larger than six, where the number of hopeless columns can be much larger than the number of live ones.
The boundary against the other essays of this field is that the configuration is a property of the world and the reading is a property of the analyst. Both change the same rate by a large factor and only one of them is anybody’s fault.
The checks, and the refusals that make them mean something
One claim is gated in this field’s library, and it is the pair: both procedures are required to hold their level when the extra predictor is worth nothing, and the recentred one is required to find a real improvement many times more often than the uncorrected one when three of five columns are hopeless. The first half is what makes the second half a gain rather than a loosening, and a version of the recentring that simply rejected more often would fail it.
The refusal for this field that bears on this essay is the reference distribution simulated from the selected winner. It is the same error as calibrating at the wrong corner, taken to its limit: instead of assuming the least favourable configuration, it assumes the most favourable one, and the test then finds a genuine improvement almost never while looking like the most careful procedure on the page.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The models that were never in the running — both name benchmark forecast, composite null, error rate, least favourable configuration, monte carlo, reality check, statistical power
- Two defects and one resampling — both name bootstrap, error rate, monte carlo, null hypothesis, out of sample, reference distribution, specification search
- A distribution drawn from the null — both name benchmark forecast, bootstrap, monte carlo, null hypothesis, reference distribution, statistical power
- Eight forecasters and one benchmark — both name benchmark forecast, error rate, monte carlo, null hypothesis, reality check
- How long a block a multiplier shares — both name error rate, monte carlo, null hypothesis, reference distribution, specification search
- The analysis after three arms — both name error rate, monte carlo, null hypothesis, reference distribution, statistical power
Named objects
A flat tag is an object no other essay names yet.
Benchmark forecastBootstrapComposite nullConservative testError rateLeast favourable configurationMonte CarloMultiplicityNull hypothesisOut of sampleReality checkReference distributionSpecification searchStatistical power