A search with no fixed point

The corner the test is calibrated at

"No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.

Worth reading first: The experiments that could have happened · What the correction corrects.

Every test in the two fields before this one has a null that is a point. The extra coefficient is zero, the two forecasts are equally accurate, the treatment does nothing. A point null has one distribution and the test is calibrated at it.

A search has a null that is not a point. “No candidate in this table is better than the benchmark” is satisfied by every configuration in which each candidate is as good as the benchmark or worse, which is a region with a face and a corner, and the statistic’s distribution is different at every point of it.

A reality check picks the corner: the configuration in which every candidate is exactly as good as the benchmark, which is the least favourable one because it gives the maximum the most chances to be large. That choice is what makes the procedure valid everywhere in the region, and it is also the entire subject of this essay, because the data is usually not at the corner.

A null that is a face rather than a point. "No candidate is better than the benchmark" is a region, and a reality check is calibrated at the corner of it where every candidate is exactly as good. As the first predictor becomes worth something, the three candidates that do not hold it become strictly worse without the null ever becoming false — the number of columns still on the boundary falls from 4.46 to 2.76 — and the rejection rate falls with it, from 2.7% to 0.0%. Recentring the columns that are clearly bad, so that they stop contributing to the maximum of the reference distribution, is what puts it back.
Fig. 1 The same true null, with more and more of the table strictly worse than the benchmark. Nothing about the null changes; the rejection rate falls to nothing.

Moving off the corner without leaving the null

Six candidates, each two of four predictors, and a benchmark that is one of them. Give the first predictor — which the benchmark holds — a coefficient. Now the three candidates that also hold it are still exactly as good as the benchmark, and the three that do not are strictly worse.

The null is still true. Not one candidate in the table is better than the benchmark, at any setting. What has changed is only how many of them are on the boundary, and the count falls from 4.46 of five to 2.76 as the coefficient runs from zero to 0.4.

The rejection rate goes with it: 2.7%, 1.7%, 1.7%, 0.0%.

That is a procedure whose size depends on a nuisance configuration nobody chose and nobody can observe. It is conservative, so it is valid — a test that rejects less than its nominal level at every point of a composite null is exactly what “valid” means here — and the conservatism is not free.

What the conservatism costs

Put a real improvement in the table and the cost is visible immediately.

Give the benchmark’s first predictor a coefficient of 1, so that the three candidates missing it are hopeless. Then give a fourth predictor — one the benchmark does not hold — a coefficient of its own, so that some candidates genuinely are better. That is a false null with three hopeless columns in it, which is what any real table looks like: nobody searches over a set of models that are all equally plausible.

What the corner costs when the table is full of hopeless candidates. The benchmark holds two predictors, one of which is worth 1; a third predictor, worth the amount on the horizontal axis, is held only by candidates the benchmark does not contain. At the left the null is true and both procedures hold their level. To the right there is a genuinely better candidate, and the uncorrected reality check finds it 2.7% of the time while the same test with the clearly bad columns recentred finds it 51.0% of the time. The columns doing the damage are the ones nobody would have looked at twice: they are so far behind that they cannot win, and calibrating as though they might is what makes the test blind.
Fig. 2 Three of the five columns hopeless, and a real improvement on the right. The lower line is what the reality check finds.

The reality check finds the improvement 0.0% of the time at a moderate edge and 2.7% at a large one. The same test with the hopeless columns recentred finds it 31.3% and 51.0%.

The mechanism is entirely mechanical. The reference distribution is the distribution of the maximum over five columns, each recentred at the boundary. Three of those columns are, in the data, hundreds of standard errors below the boundary — but the reference distribution has been told to treat them as though they were on it, so they contribute their full share of upward noise to the maximum. The critical value is set by five columns’ worth of maximum and the observed statistic is produced by two, and the test cannot bridge the gap.

The size falls faster than the count can explain

The two sequences in the section above invite a model, and the model gets the middle of the sweep right and the end of it wrong — which is where the interesting part is.

If the only thing that mattered were how many candidates sit on the boundary, the size would behave like the tail of a maximum over that many statistics. Calibrated at 2.7% with 4.46 of them on the boundary, that model predicts

1(10.027)m/4.461 - (1 - 0.027)^{m/4.46}

which at m = 3.89, 3.33 and 2.76 gives 2.4%, 2.0% and 1.7%.

Measured: 1.7%, 1.7% and 0.0%.

So the count explains the first part of the fall and not the last of it. Between the third and fourth settings the boundary count drops by about half a candidate and the size drops from 1.7% to nothing.

Why the candidates that leave do not merely stop counting

The gap says something specific about the mechanism, and it is worth naming because it changes what the conservatism is.

A candidate that leaves the boundary does not become inert. It becomes strictly worse, and a strictly worse candidate still contributes to the bootstrap reference distribution the maximum is compared against — it is recentred, so its own contribution to the reference is a draw from the same null spread, while its contribution to the observed statistic has been pushed down.

So each departing candidate does two things at once: it removes one chance for the observed maximum to be large, and it leaves one full-sized draw in the distribution that maximum is measured against. A maximum-of-m argument counts only the first.

That is what makes the fall to zero possible at all. A pure count argument can never reach zero — with 2.76 candidates still exactly on the boundary there are still 2.76 chances — and the measured rate is a clean zero on the whole run, which on any realistic number of trials bounds the true size well under a per cent.

The conservatism is therefore not a mild loss of power that scales with how far the truth sits from the corner. It is a procedure that can switch off entirely at configurations that are still inside the null, and the switching-off is driven by the candidates the null is not about.

The repair, which is a threshold and an admission

The repair is Hansen’s, and it is one line: a column whose sample differential is far enough below zero is recentred at its own mean rather than at the boundary, so that the simulated version of it sits where the real one does and never wins the maximum.

“Far enough” is a threshold, and a threshold is a decision. The version used here drops a column when its differential is more than one standard error below zero, which keeps 1.80 of five columns at the boundary in the hardest configuration above and 4.46 when the table is at the corner.

Two things about that decision are worth being explicit about.

The procedure is no longer exactly valid. A column just barely below the threshold is treated as hopeless when it might be on the boundary, and at that configuration the test can exceed its level. Measured at the corner, where nothing is dropped, the recentred version rejects 2.7% — identical to the uncorrected one, because the two procedures coincide there. Measured with three columns hopeless it rejects 2.3% at a true null, which is inside the counting error of the level it claims.

And the gain is not marginal. Across the configurations measured here it runs from nothing, when the table is at the corner and there is nothing to drop, to a factor of ten when most of the table is hopeless. There is no setting in which the recentring loses.

What the corner costs when the table is full of hopeless candidates. The benchmark holds two predictors, one of which is worth 0.4; a third predictor, worth the amount on the horizontal axis, is held only by candidates the benchmark does not contain. At the left the null is true and both procedures hold their level. To the right there is a genuinely better candidate, and the uncorrected reality check finds it 34.0% of the time while the same test with the clearly bad columns recentred finds it 34.7% of the time. The columns doing the damage are the ones nobody would have looked at twice: they are so far behind that they cannot win, and calibrating as though they might is what makes the test blind.
Fig. 3 A table where the hopeless columns are only mildly bad. The two lines nearly coincide — 11.0% against 13.7% at the middle point — because there is little to drop.

The dial between the two, and where it turns

The gap between the two procedures is a function of how far behind the bad columns are, and it opens suddenly rather than gradually.

  • Columns mildly behind: the reality check finds a real improvement 11.0% of the time, the recentred version 13.7%.
  • Columns moderately behind: 1.7% against 23.3%.
  • Columns hopelessly behind: 0.0% against 31.3%.

The reality check gets worse as the table gets easier to read. That is the sentence worth carrying away, and it is not a paradox once the mechanism is in hand: the columns that make a table easy for a human to read — the obviously bad ones — are exactly the columns that inflate the reference distribution and depress the test.

What the corner costs when the table is full of hopeless candidatesThe benchmark holds two predictors, one of which is worth 0.7; a third predictor, worth the amount on the horizontal axis, is held only by candidates the benchmark does not contain. At the left the null is true and both procedures hold their level. To the right there is a genuinely better candidate, and the uncorrected reality check finds it 12.0% of the time while the same test with the clearly bad columns recentred finds it 33.3% of the time. The columns doing the damage are the ones nobody would have looked at twice: they are so far behind that they cannot win, and calibrating as though they might is what makes the test blind.00.1000.2000.3000.40000.1000.2000.3000.4000.500how much better the predictor the benchmark does not hold isshare of draws on which the improvement is foundlower line: the reality check · upper: with the bad columns recentred300 draws, 99 resamples, nominal 5%1.9 of five columns on the boundary
Fig. 4 The dial itself. Drag it through how hopeless the hopeless candidates are: the recentred line rises and the uncorrected one falls, and they are furthest apart where a reader would say the table was clearest.

The practical form of this is that padding a table is not free in the direction people assume. The usual worry about adding candidates to a search is a false positive. Here, adding candidates that are obviously bad costs power and nothing else, and an analyst who includes a few hopeless specifications “for completeness” has made their own test unable to see anything.

What a composite null actually asks of a procedure

It is worth slowing down on the logic, because “the test is conservative off the corner” is easy to say and easy to mis-state in either direction.

A test of a composite null has to hold its level at every point of the null region — that is what the guarantee means. A point of the region where it rejects less than its nominal level is not a failure; it is the price of the guarantee at the points where it rejects at the level. So a procedure calibrated at the least favourable configuration is doing the only thing that can be done without further information, and calling it broken is wrong.

What is legitimate to complain about is that the least favourable configuration is being used as a fact rather than as a bound. The reference distribution here does not say “the maximum could be this large if every column were on the boundary”; it says “the maximum is distributed as it would be if every column were on the boundary”, and the data very often contains direct evidence against that. A column sitting eight standard errors below zero is not a column that might be on the boundary.

That is the sense in which the recentring is an improvement and not a loosening. It uses the data to narrow the region the test has to be valid over — from the whole null to the part of it consistent with what has been observed — and it pays for that with a threshold and a small amount of exactness. The trade is the same one every studentised procedure makes, and it is unusually easy to see here because the two versions coincide exactly at the corner and diverge only where the evidence is.

Why this is not the same problem as the reading

Three separate things can miscalibrate a search, and this field has now measured all three. They are independent and it is worth keeping them apart.

The reading decides which maximum is taken — over the rows against a stated benchmark, or over every pair, or against a benchmark that was itself chosen. It moves the same true null between 0.0% and 76.2%.

The displacement is a real difference in expected accuracy that exists before any search, and it comes from candidates of different sizes rather than from any search at all. It is what makes a multiplicity correction the wrong tool on a nested table.

And the configuration is this essay: where in the null region the data sits, which changes the distribution of the same statistic under the same reading with the same candidates.

A procedure can be right about all three or wrong about all three, and the failures do not announce which they are. A search that rejects too often may have the wrong reading or a displacement in it; a search that rejects too rarely may have a searched benchmark, or a table full of hopeless columns, or both.

One true null, one table, five readings. six subsets of two, none nested, at a null where nothing any candidate holds is worth anything, over 500 draws. Each bar is the share of draws on which that reading declares a difference at a nominal 5%. The reading is the whole of the difference between the bars: the data is identical. An open search over all 30 ordered pairs rejects 27.4%; the table's own 5% point is 2.400 against the 1.671 a single comparison uses. Bonferroni takes the open reading to 0.8% — and on the nested ladder the same correction does not reach the nominal level at all, because there the excess is a shift in the mean rather than a maximum over many.
Fig. 5 The reading, on the same six candidates this essay uses. It is a separate lever and it is much larger.
Every candidate is behind by what its parameter count says. Each dot is one of the fifteen subsets of four predictors, fitted on a rolling window of 80 rows and scored against the benchmark out of sample over 60 origins, at a null where every one of them contains the truth. The line is σ²(q₀/(R − q₀ − 1) − q/(R − q − 1)), which is arithmetic on two integers and a window length. Most of these pairs are not nested — a subset of two predictors and a different subset of two share neither model — and the closed form does not care: the displacement is a statement about how many coefficients each side estimates. The candidates of the benchmark's own dimension sit at zero.
Fig. 6 And the displacement, which is zero on this table by construction — every candidate fits the same number of coefficients — which is why the configuration can be measured here in isolation.

How many columns are really being maximised over

The count of columns still on the boundary is the quantity that decides everything above, and it is worth looking at directly because it is the one thing in this essay a practitioner can compute from their own table.

At the corner it is 4.46 of five — not five, because the threshold occasionally drops a column whose sample differential wandered a standard error below zero by chance. That is the recentring’s own cost measured at the place it has nothing to gain, and it is why the two procedures agree exactly on rejection rate there rather than the recentred one being slightly better.

As the configuration moves it falls to 4.28, 3.81 and 2.76, and in the power measurements with three hopeless columns it sits at 1.80. So in the configuration where the reality check finds a real improvement 0.0% of the time, its reference distribution is the maximum of five columns while the honest number of live comparisons is fewer than two.

Two live comparisons and a critical value set by five is roughly the whole of the loss, and the size of it can be estimated from that alone: the 95th percentile of the maximum of five approximately independent statistics sits well above that of two, and the observed statistic never reaches it.

There is a tempting shortcut here — count the live columns and apply a correction for that many — and it is not the same procedure. The count is itself random and correlated with the statistic, which is exactly why the recentring is stated as a modification of the reference distribution rather than as a smaller multiplicity. Using the count directly is the mistake the first essay of this field calls reading a search by the result it produced.

Searching the benchmark makes the table harder to reject with. The benchmark is the best of the first a candidates and the alternatives are the remaining 15 − a, at a true null, over 400 draws. At a = 1 the benchmark is stated in advance and the reading rejects 7.8%. Every candidate added to the family the benchmark is chosen from makes the benchmark a better forecast and the alternatives' job harder, so the rate falls to 0.0% — the opposite direction to the one a search is supposed to move a false-positive rate in, and it happens because there are two searches pushing against each other.
Fig. 7 The other procedure in this field whose rejection rate falls to zero, and for an unrelated reason: a benchmark that was itself searched for. Two ways to build a test that cannot see anything.

Where the least favourable configuration is the truth

There is one setting on this site where the corner is not an assumption, and it is worth pointing at because it is the only measurement of a reality check that is free of everything above.

The variance-minimising combination of a set of forecasts satisfies a first-order condition that makes its error covary with every member’s at exactly its own variance, so every one of the encompassing nulls is on the boundary simultaneously. That is the least favourable configuration arriving as a fact about the arithmetic rather than as a choice about the calibration, and the field that found it measures what a reality check does there.

This field reaches the corner a second and much easier way: candidates of the same size, none containing any other, at a null where nothing is worth anything. Six two-predictor subsets of four predictors are all exactly as accurate as each other in population and all displaced by exactly the same amount, which is none. That is why every measurement in this field is run on that table.

The two routes matter because they say the corner is reachable rather than hypothetical, and a procedure calibrated at a configuration that cannot occur would be a different kind of object from one calibrated at a configuration that merely usually does not.

Eight candidates, one of them exactly as good as the benchmark. The candidate set: moving averages of the last 1, 2, 3, 5, 8, 13, 21 and 34 observations, each drawn as its expected squared error divided by the benchmark's — the mean of all 60. The persistence is not chosen, it is solved for: at φ = 0.4895 the best candidate in the set, the average of 2, has exactly the benchmark's expected squared error, and every other candidate is worse by between 0.5% and 5.1%. So the null that no candidate beats the benchmark is true, with one candidate on its boundary. Everything a set comparison claims about its own error rate has to be measured here, because anywhere further inside the null every procedure flatters itself.
Fig. 8 The reality check in the field that built it, on a set of forecasts nobody estimated.

What is claimed here, and what is not

This essay takes the configuration a set test is calibrated at, and the claims are the size falling as the table leaves the corner, the power at three degrees of hopelessness, and the recentring’s gain and its level.

What stays out and is named as a decision: the choice of threshold, which is set at one standard error here and swept nowhere — the number matters and choosing it well is a literature rather than a measurement; a studentised version of the same test, which changes the shape of the reference distribution and interacts with the recentring in a way nothing here separates; and any statement about tables larger than six, where the number of hopeless columns can be much larger than the number of live ones.

The boundary against the other essays of this field is that the configuration is a property of the world and the reading is a property of the analyst. Both change the same rate by a large factor and only one of them is anybody’s fault.

The checks, and the refusals that make them mean something

One claim is gated in this field’s library, and it is the pair: both procedures are required to hold their level when the extra predictor is worth nothing, and the recentred one is required to find a real improvement many times more often than the uncorrected one when three of five columns are hopeless. The first half is what makes the second half a gain rather than a loosening, and a version of the recentring that simply rejected more often would fail it.

The refusal for this field that bears on this essay is the reference distribution simulated from the selected winner. It is the same error as calibrating at the wrong corner, taken to its limit: instead of assuming the least favourable configuration, it assumes the most favourable one, and the test then finds a genuine improvement almost never while looking like the most careful procedure on the page.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Benchmark forecastBootstrapComposite nullConservative testError rateLeast favourable configurationMonte CarloMultiplicityNull hypothesisOut of sampleReality checkReference distributionSpecification searchStatistical power