Searching among fitted models

A null with a model in it

The distribution to read the winner of a table against cannot be resampled from the data, because the data does not contain the null. It has to be generated from a model — which is the assumption the resampling was chosen to avoid.

Worth reading first: Where the bootstrap lies · What the model says next.

A specification search reports one number: the best row of the table. To say whether that number is surprising, it has to be compared with the distribution of the best row of a table like this one when nothing is there — and there is no table of that distribution, because it depends on how many variants were tried, how much they overlap, and how many rows each fit was given.

This site has two ways of producing such a distribution and they are built on opposite assumptions. One resamples blocks of rows of the observed comparison, which assumes only that the comparison is stationary. The other simulates from a fitted model, which assumes the model and gets a null that is true by construction.

A table of nested variants needs both at once, and neither of them survives the other unchanged.

Why the resampling cannot work here

The row bootstrap does something specific: it draws blocks of rows of the loss-differential matrix — blocks so that the dependence between origins survives, whole rows so that the dependence between candidates survives with it — and it recentres each column at that column’s own sample mean, because the null it is built for is “this candidate is no better than the benchmark” and the least favourable version of that null is “this candidate is exactly as good”.

Recentring at the sample mean is the step that stops being right. In a table of fitted variants the sample mean is not an estimate of the boundary of the null; it is an estimate of the displacement, which at these settings is −0.0152 per origin. So the procedure builds a reference distribution around the belief that every candidate is already behind by that much, and asks whether the observed maximum is surprising relative to candidates that are behind. It almost never is.

Counted at a null where every variant is a variant of the truth, the row bootstrap over the table rejects 1.3% of the time at a nominal 5%.

Six readings of one table, at a null where every variant is a variant of the truth. A benchmark AR(1) and eight variants of it, each adding one further lag, compared over 60 rolling origins with the fits given 71 rows each. The series is an AR(1), so every extra coefficient is zero in population and every rejection counted here is a false one. 120 tables. The winner's own p-value is at 7.5% — near the nominal level, because the multiplicity of eight and the displacement every fitted variant carries push in opposite directions. Recentring each row, which is the repair a nested comparison requires, removes one of the two and takes it to 22.5%. The row bootstrap over the whole table rejects 1.7%: it recentres at each column's own sample mean, which under this null is -0.0152 per origin rather than zero. A distribution simulated from the fitted benchmark is at 4.2%.
Fig. 1 The row bootstrap is the fifth bar. Being at a quarter of the nominal level is not conservatism that has been paid for with something; it is a procedure reading a distribution that describes a situation nobody is in.

Applying the recentring the nested case requires first, and then resampling the recentred rows, repairs the level: 3.3%. That is the honest version of the resampling route, and it is worth noticing what has happened to its selling point. The block bootstrap was chosen because it assumes nothing about the model. The recentring that makes it usable is Clark and West’s, which is derived from the model and is exactly the assumption that was being avoided. The construction that assumes only stationarity is available only after a step that assumes the specification.

The distribution generated instead

The other route builds the null rather than repairing the data. Fit the benchmark to the whole series, resample its residuals, simulate a series in which the extra coefficients are zero by construction, re-run the entire eight-variant search on it, and keep the largest statistic. Do that a couple of hundred times and the result is the distribution of the maximum of a table like this one, under the null, with the overlap between the variants and the estimation noise in every one of them included because they were all actually computed.

The distribution of the largest statistic in the table. Fit the benchmark to the whole series, resample its residuals, simulate 199 series in which the null is true by construction, re-run the entire eight-variant search on each, and keep the largest statistic. That is the distribution drawn here, and it is the distribution of the thing a specification search actually reports. It is centred at 1.045 — the maximum of eight statistics is not centred at zero however well each of them behaves — and its 5% point is 2.536. A table read against 1.671 is reading the distribution of one statistic; a Bonferroni correction reads it against 2.577 and is nearly right here, because eight variants that each add a different lag are nearly eight separate chances.
Fig. 2 The distribution of the largest of eight recentred statistics, from searches run on series in which the null is true. It is not centred at zero, and its upper tail is not a normal’s.

Read against that distribution, the table rejects 4.0% of the time at a nominal 5%, and its p-values are 0.100 from flat by the same uniformity statistic this site applies to every test it builds. The recentred row bootstrap’s are 0.093 from flat, which is the same thing said twice: both resampling readings are honest procedures once the recentring is in place. The winner’s own p-value is 0.529 from flat, which is not a p-value at all.

The critical value it produces is the number worth quoting, because it is the thing a reader would otherwise take from a table. Averaged over comparisons, the 5% point of the maximum is 2.278. A single comparison’s own 5% point at this many origins is 1.671. A Bonferroni correction over the eight rows puts it at 2.577.

One table, and three places to put the line. The eight recentred statistics from a single specification search, at a null where every variant is a variant of the truth, and the three critical values a reader might compare them with. A table's own 5% point is 1.671 and 0 of the eight are past it. Bonferroni's is 2.577. The distribution of the maximum, simulated from the fitted benchmark by re-running the entire search on 199 series in which the null is true by construction, puts it at 2.536. The three answers differ by more than any of the eight statistics differ from each other, and only the third of them is a statement about the maximum of this table.
Fig. 3 One table with all three lines drawn across it. The gap between the first and the third is larger than the gap between the largest statistic in the table and the smallest.
What each variant loses before anything has been searched for. The mean loss differential of each of the eight variants against the benchmark, over 600 tables of 60 origins, with every fit given 71 rows. The series is an AR(1) and every variant adds a lag whose coefficient is zero, so in population the two forecasts are the same forecast and the difference drawn here is estimation noise and nothing else. The marked line is σ²(q₁ − q₀)/n = -0.01408, which is an expression in how many coefficients each model has and how many rows it was fitted on — it knows nothing about the series, the persistence or which lag the variant added, and every bar is within a fifth of it. This is the amount a reference distribution recentred at each column's own sample mean believes the candidates are already behind by.
Fig. 4 The displacement the recentring is about, drawn once more: the amount every column’s sample mean is below zero before anything has been searched.

The maximum of eight is not centred at zero

There is a fact in that picture which is easy to walk past and is most of the reason a table cannot be read against a table. Each of the eight recentred statistics is, individually, roughly standard normal under this null. The largest of eight such statistics is not: it has a distribution of its own, sitting well to the right of zero, and its median is a number no book of tables contains because it depends on how correlated the eight are.

That is why the three critical values above are so far apart, and why the gap between them is not a matter of taste. Reading a table against 1.671 is reading the distribution of one statistic; reading it against the simulated 2.278 is reading the distribution of the largest one, on a table of this shape. A reader who reports the best of eight rows and its p-value has produced two numbers of which the first is the maximum and the second describes something else.

The distribution of the largest statistic in the table. Fit the benchmark to the whole series, resample its residuals, simulate 199 series in which the null is true by construction, re-run the entire eight-variant search on each, and keep the largest statistic. That is the distribution drawn here, and it is the distribution of the thing a specification search actually reports. It is centred at 0.920 — the maximum of eight statistics is not centred at zero however well each of them behaves — and its 5% point is 2.067. A table read against 1.658 is reading the distribution of one statistic; a Bonferroni correction reads it against 2.536 and is nearly right here, because eight variants that each add a different lag are nearly eight separate chances.
Fig. 5 The same distribution when the comparison runs over twice as many origins. The maximum stays to the right of zero and its 5% point stays above a table’s: more origins make each statistic better behaved and do nothing at all about there being eight of them.

The cost of building it is the honest objection to this route and it is worth stating in units. Each reference distribution here is a couple of hundred series, each of which requires the whole eight-variant rolling search to be re-run — sixty origins, nine regressions at each. Everything else in this field is arithmetic on numbers that already exist. This is the one construction that costs more than the comparison it is about, by a factor of the number of simulations, and there is no way around it: the quantity being estimated is a property of the search, so the search is what has to be repeated.

What it is worth when something is there

A procedure that holds its level is not thereby a good procedure — the row bootstrap holds a level far below the one it claims, and a rule that never rejects holds every level there is. The question is what each construction finds when one of the variants is genuinely better.

Give the series a coefficient of 0.3 on the fourth lag, so that the benchmark is now wrong and the variant which adds lag 4 is right.

What each construction finds, as the benchmark becomes wrong. The same table of eight nested variants as the truth acquires a coefficient on the fourth lag that the benchmark does not have. 60 tables at each of four values, 60 origins each. Three readings: the distribution simulated from the fitted benchmark, the row bootstrap over the observed differentials, and a Bonferroni correction over the eight rows. At β = 0 the first is at 0.0% and the second at 1.7% — the row bootstrap has bought its conservatism at a null it is not measuring correctly — and by β = 0.3 the difference between them is 18.3% of the tables. Bonferroni is close to the simulated distribution throughout, which is what a table of eight nearly orthogonal variants does to a correction for eight independent ones.
Fig. 6 Three constructions as the truth moves away from the benchmark. The one that had to be told which model the benchmark was is the one that finds the departure; the one that assumed only stationarity is ten to twenty points behind it the whole way.

The simulated distribution rejects 50.0% of the time. The raw row bootstrap rejects 33.3%, and the recentred one 43.3%. The search picks the variant that is actually right on 72.7% of tables. At a smaller departure — a coefficient of 0.2 — the three are 18.7%, 9.3% and 16.7%, and the right variant is picked 53.3% of the time.

Both of the resampling readings are behind at every one of those settings, and the raw one is behind by more the smaller the departure is — which is the shape a location error makes. A test whose reference distribution sits below where the null puts it needs the truth to move further before the observed maximum clears it, and how much further does not depend on how many origins the comparison has.

So the ordering is the same at every alternative drawn: the construction with a model in it holds its level and finds more; the construction without one either cannot reject or has to import the model through the back door of the recentring, and finds less either way.

The repaired test finds what the original cannot. An AR(1) forecast against an AR(2) forecast, with the second coefficient swept from zero — where the two models are the same model — to 0.4, where the larger one is genuinely better. Both curves are one-sided tests of the same hypothesis on the same forecasts. The ordinary statistic finds the larger model better 5.5% of the time at φ₂ = 0.2, which is its own nominal level and therefore no evidence at all, while the recentred one finds it 47.0% of the time. At the left both are honest: 4.1% and 0.0% under a true null.
Fig. 7 The two-model comparison’s own version of the same trade, where the moment correction and the distributional one can be watched separately as the larger model becomes genuinely necessary.

There is one number in that comparison which is not about tests at all and is worth carrying out of it: at the smallest departure drawn, a coefficient of 0.1, the simulated distribution rejects 6.7% and the search picks the right variant on 21.3% of tables. The benchmark is wrong, the table contains the model that fixes it, and eight rows of comparison at sixty origins mostly cannot tell. The reference distribution is not what is failing there — the departure is genuinely smaller than the estimation noise the table is made of.

And Bonferroni, which is nearly right here

There is a third procedure in these tables and it is the one nobody expects to work: multiply the winner’s recentred p-value by eight. At the null it rejects 2.8%; at the 0.3 alternative it rejects 51.3%, within a point and a half of the simulated distribution.

That is not a general fact about Bonferroni and it is not luck. A table of eight nested variants that each add a different lag is worth 7.12 independent comparisons, so charging for eight is charging for almost exactly what was searched. The same correction applied to eight moving averages of one series charges for eight where two and a half were searched, and lands at a tenth of its nominal level.

The lesson is not that Bonferroni is fine. It is that the correct charge for a search depends on a quantity the table does not report — how much the candidates overlap — and that a reference distribution built by re-running the search is the only one of these procedures that measures that quantity rather than assuming a value for it. Here it happens to agree with the assumption; one field along, on a table that looks identical, it does not.

What the simulation is assuming

The cost of the generated null is stated plainly: it is a parametric reference distribution, and if the benchmark is wrong in a way the search is not looking at, the distribution is drawn from the wrong series.

Three parts of it are worth separating, because they are not equally fragile.

The mean equation — that the null model is an AR(1) — is the assumption being tested against alternatives inside the table, and is the whole point: the reference distribution has to be under the null, and the null is that model.

The innovations are resampled from the benchmark’s own residuals rather than drawn from a normal, so the distribution inherits whatever skewness and heavy tails the series actually has. That is the cheapest part of the construction to get right and it costs nothing.

The dependence the residuals do not contain is the real exposure. If the series has conditional heteroskedasticity, or a break, or anything else the AR(1) does not capture, resampling its residuals one at a time destroys it, and the simulated maximum is drawn from a tamer world than the real one. A wild or block resampling of the residuals is the repair and is not built here; it is named as a decision below.

Most of the difference is calibration, not information

The four readings differ in power and they also differ in the level they actually hold, and the two are entangled in every comparison above. Separating them says something the power column does not.

Take each procedure’s realised level as its true operating point, convert its power at the 0.3 alternative into an implied signal, and read that signal back at a common 5%. On a normal approximation — which is rough, and rough in the same way for all four — the raw row bootstrap runs at 1.3% and 33.3%, which is a signal of 1.79; the recentred one at 3.3% and 43.3%, a signal of 1.67; the simulated distribution at 4.0% and 50.0%, a signal of 1.75; and Bonferroni at 2.8% and 51.3%, a signal of 1.94.

Put all four at 5% and their powers become 55.9%, 51.0%, 54.2% and 61.8%.

The ordering nearly reverses. The raw row bootstrap, which is the worst reading in the essay by a wide margin, is mid-pack once its conservatism is priced out; the recentred one is last; and Bonferroni, which is a correction rather than a reference distribution, is ahead of everything.

That is not an argument that the raw row bootstrap is fine. It cannot be run at 5%, because nobody using it knows what its critical value should be — the whole defect is that its reference distribution is in the wrong place, and an analyst has no way to shift it back. What the calibrated comparison says is where the gain in this field actually comes from: the four statistics extract roughly the same amount of information from the table, within about ten points of each other, and the differences reported are almost entirely differences in whether each one is being read at the level it claims.

The practical form of that is a narrower recommendation than “simulate the null”. The expensive construction — a couple of hundred re-runs of the whole eight-variant search — is buying a correct location and very little else. Anything cheaper that also gets the location right buys nearly the same thing, which is exactly what the Bonferroni row demonstrates on this particular table, and exactly what it fails to do on a table of eight windows of one series where the effective count is 2.44 rather than near eight.

So the question a table poses is not which reference distribution to build. It is whether the cheap correction’s implied count is close to the table’s real one — and that is a single number, measurable by the effective-count arithmetic rather than by simulating the search, which is the one piece of this field that does not cost more than the comparison it is about.

Two constructions, and what each is for

The pair is worth stating as a rule rather than as a result, because the choice recurs everywhere a search is priced.

Resample when the null is a statement about the data. For candidates that estimate nothing, the observed differentials are the object, and the least favourable configuration is a shift of them. Resampling rows carries the dependence for free and needs no model at all.

Simulate when the null is a statement about a model. For candidates that are fitted, the null makes the two forecasts identical in population and everything a sample shows is estimation noise whose distribution is a property of the fitting. There is nothing in the data to shift into the null position, because the null position is not a location — it is a whole distribution the data was never drawn from.

The specification search is the case where both descriptions apply to one table, and this is why the two halves do not compose: the multiplicity is a statement about the data, the nesting is a statement about the model, and each of the two constructions gets one of them right.

What each construction finds, as the benchmark becomes wrongThe same table of eight nested variants as the truth acquires a coefficient on the fourth lag that the benchmark does not have. 60 tables at each of four values, 120 origins each. Three readings: the distribution simulated from the fitted benchmark, the row bootstrap over the observed differentials, and a Bonferroni correction over the eight rows. At β = 0 the first is at 0.0% and the second at 0.0% — the row bootstrap has bought its conservatism at a null it is not measuring correctly — and by β = 0.3 the difference between them is 31.7% of the tables. Bonferroni is close to the simulated distribution throughout, which is what a table of eight nearly orthogonal variants does to a correction for eight independent ones.00.2500.5000.750100.1000.2000.300the coefficient the benchmark is missingshare of tables in which the reading rejects at 5%upper: simulated from the benchmark · middle: Bonferroni · lower: the row bootstrap60 tables at each of 4 alternativesat β = 0 every rejection is false
Fig. 8 The same three constructions as the comparison lengthens. Drag it: the gap between them narrows and does not close, because more origins improve every estimate and change nothing about where the row bootstrap thinks the null is.

What is claimed here, and what is not

This essay takes the reference distribution for a table of fitted variants, and the claim is the comparison: the two constructions, at a null and at three alternatives, on the same tables.

What stays out and is named as a decision: residual resampling that keeps second-moment dependence, which is the repair to the fragility named above and is a different construction rather than a setting; a table whose variants are not all nested in the benchmark, where the displacement differs by row and there is no single closed form for it; and the case where the benchmark is one of the things being searched over, which stops being a test and becomes selection.

The boundary against the set field is where it has been since the fitting arrived: everything about resampling blocks of rows is established there, and every one of its reference distributions is built around a point this table is not at.

The checks, and the refusals that make them mean something

Three claims are gated in this field’s library. The row bootstrap is required to reject far below its nominal level at this null and the simulated distribution to be within three points of it, which is the size comparison stated as something that could fail in either direction. The simulated distribution is required to find a genuinely better variant more often than the row bootstrap at every alternative drawn. And its critical value for the maximum is required to be above a single comparison’s, which is the multiplicity the whole construction exists to price.

The refusal beside them is the most natural mistake available here: a reference distribution simulated from the variant the search selected rather than from the benchmark. It is the fitted model in hand at the end of a search, it produces series in which the null is false, and the maximum it draws is therefore the maximum under the alternative. Handed that construction the check throws, and it throws on a power collapse rather than on a size failure — the test still holds its level, it simply stops being able to find anything, which is the failure mode this field has now met from both directions.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Block bootstrapBonferroniBootstrapClark–WestCritical valueMultiplicityNested modelsNull hypothesisOut of samplePowerReference distributionResidual resamplingSpecification searchStationarity