A null with a model in it
Worth reading first: Where the bootstrap lies · What the model says next.
A specification search reports one number: the best row of the table. To say whether that number is surprising, it has to be compared with the distribution of the best row of a table like this one when nothing is there — and there is no table of that distribution, because it depends on how many variants were tried, how much they overlap, and how many rows each fit was given.
This site has two ways of producing such a distribution and they are built on opposite assumptions. One resamples blocks of rows of the observed comparison, which assumes only that the comparison is stationary. The other simulates from a fitted model, which assumes the model and gets a null that is true by construction.
A table of nested variants needs both at once, and neither of them survives the other unchanged.
Why the resampling cannot work here
The row bootstrap does something specific: it draws blocks of rows of the loss-differential matrix — blocks so that the dependence between origins survives, whole rows so that the dependence between candidates survives with it — and it recentres each column at that column’s own sample mean, because the null it is built for is “this candidate is no better than the benchmark” and the least favourable version of that null is “this candidate is exactly as good”.
Recentring at the sample mean is the step that stops being right. In a table of fitted variants the sample mean is not an estimate of the boundary of the null; it is an estimate of the displacement, which at these settings is −0.0152 per origin. So the procedure builds a reference distribution around the belief that every candidate is already behind by that much, and asks whether the observed maximum is surprising relative to candidates that are behind. It almost never is.
Counted at a null where every variant is a variant of the truth, the row bootstrap over the table rejects 1.3% of the time at a nominal 5%.
Applying the recentring the nested case requires first, and then resampling the recentred rows, repairs the level: 3.3%. That is the honest version of the resampling route, and it is worth noticing what has happened to its selling point. The block bootstrap was chosen because it assumes nothing about the model. The recentring that makes it usable is Clark and West’s, which is derived from the model and is exactly the assumption that was being avoided. The construction that assumes only stationarity is available only after a step that assumes the specification.
The distribution generated instead
The other route builds the null rather than repairing the data. Fit the benchmark to the whole series, resample its residuals, simulate a series in which the extra coefficients are zero by construction, re-run the entire eight-variant search on it, and keep the largest statistic. Do that a couple of hundred times and the result is the distribution of the maximum of a table like this one, under the null, with the overlap between the variants and the estimation noise in every one of them included because they were all actually computed.
Read against that distribution, the table rejects 4.0% of the time at a nominal 5%, and its p-values are 0.100 from flat by the same uniformity statistic this site applies to every test it builds. The recentred row bootstrap’s are 0.093 from flat, which is the same thing said twice: both resampling readings are honest procedures once the recentring is in place. The winner’s own p-value is 0.529 from flat, which is not a p-value at all.
The critical value it produces is the number worth quoting, because it is the thing a reader would otherwise take from a table. Averaged over comparisons, the 5% point of the maximum is 2.278. A single comparison’s own 5% point at this many origins is 1.671. A Bonferroni correction over the eight rows puts it at 2.577.
The maximum of eight is not centred at zero
There is a fact in that picture which is easy to walk past and is most of the reason a table cannot be read against a table. Each of the eight recentred statistics is, individually, roughly standard normal under this null. The largest of eight such statistics is not: it has a distribution of its own, sitting well to the right of zero, and its median is a number no book of tables contains because it depends on how correlated the eight are.
That is why the three critical values above are so far apart, and why the gap between them is not a matter of taste. Reading a table against 1.671 is reading the distribution of one statistic; reading it against the simulated 2.278 is reading the distribution of the largest one, on a table of this shape. A reader who reports the best of eight rows and its p-value has produced two numbers of which the first is the maximum and the second describes something else.
The cost of building it is the honest objection to this route and it is worth stating in units. Each reference distribution here is a couple of hundred series, each of which requires the whole eight-variant rolling search to be re-run — sixty origins, nine regressions at each. Everything else in this field is arithmetic on numbers that already exist. This is the one construction that costs more than the comparison it is about, by a factor of the number of simulations, and there is no way around it: the quantity being estimated is a property of the search, so the search is what has to be repeated.
What it is worth when something is there
A procedure that holds its level is not thereby a good procedure — the row bootstrap holds a level far below the one it claims, and a rule that never rejects holds every level there is. The question is what each construction finds when one of the variants is genuinely better.
Give the series a coefficient of 0.3 on the fourth lag, so that the benchmark is now wrong and the variant which adds lag 4 is right.
The simulated distribution rejects 50.0% of the time. The raw row bootstrap rejects 33.3%, and the recentred one 43.3%. The search picks the variant that is actually right on 72.7% of tables. At a smaller departure — a coefficient of 0.2 — the three are 18.7%, 9.3% and 16.7%, and the right variant is picked 53.3% of the time.
Both of the resampling readings are behind at every one of those settings, and the raw one is behind by more the smaller the departure is — which is the shape a location error makes. A test whose reference distribution sits below where the null puts it needs the truth to move further before the observed maximum clears it, and how much further does not depend on how many origins the comparison has.
So the ordering is the same at every alternative drawn: the construction with a model in it holds its level and finds more; the construction without one either cannot reject or has to import the model through the back door of the recentring, and finds less either way.
There is one number in that comparison which is not about tests at all and is worth carrying out of it: at the smallest departure drawn, a coefficient of 0.1, the simulated distribution rejects 6.7% and the search picks the right variant on 21.3% of tables. The benchmark is wrong, the table contains the model that fixes it, and eight rows of comparison at sixty origins mostly cannot tell. The reference distribution is not what is failing there — the departure is genuinely smaller than the estimation noise the table is made of.
And Bonferroni, which is nearly right here
There is a third procedure in these tables and it is the one nobody expects to work: multiply the winner’s recentred p-value by eight. At the null it rejects 2.8%; at the 0.3 alternative it rejects 51.3%, within a point and a half of the simulated distribution.
That is not a general fact about Bonferroni and it is not luck. A table of eight nested variants that each add a different lag is worth 7.12 independent comparisons, so charging for eight is charging for almost exactly what was searched. The same correction applied to eight moving averages of one series charges for eight where two and a half were searched, and lands at a tenth of its nominal level.
The lesson is not that Bonferroni is fine. It is that the correct charge for a search depends on a quantity the table does not report — how much the candidates overlap — and that a reference distribution built by re-running the search is the only one of these procedures that measures that quantity rather than assuming a value for it. Here it happens to agree with the assumption; one field along, on a table that looks identical, it does not.
What the simulation is assuming
The cost of the generated null is stated plainly: it is a parametric reference distribution, and if the benchmark is wrong in a way the search is not looking at, the distribution is drawn from the wrong series.
Three parts of it are worth separating, because they are not equally fragile.
The mean equation — that the null model is an AR(1) — is the assumption being tested against alternatives inside the table, and is the whole point: the reference distribution has to be under the null, and the null is that model.
The innovations are resampled from the benchmark’s own residuals rather than drawn from a normal, so the distribution inherits whatever skewness and heavy tails the series actually has. That is the cheapest part of the construction to get right and it costs nothing.
The dependence the residuals do not contain is the real exposure. If the series has conditional heteroskedasticity, or a break, or anything else the AR(1) does not capture, resampling its residuals one at a time destroys it, and the simulated maximum is drawn from a tamer world than the real one. A wild or block resampling of the residuals is the repair and is not built here; it is named as a decision below.
Most of the difference is calibration, not information
The four readings differ in power and they also differ in the level they actually hold, and the two are entangled in every comparison above. Separating them says something the power column does not.
Take each procedure’s realised level as its true operating point, convert its power at the 0.3 alternative into an implied signal, and read that signal back at a common 5%. On a normal approximation — which is rough, and rough in the same way for all four — the raw row bootstrap runs at 1.3% and 33.3%, which is a signal of 1.79; the recentred one at 3.3% and 43.3%, a signal of 1.67; the simulated distribution at 4.0% and 50.0%, a signal of 1.75; and Bonferroni at 2.8% and 51.3%, a signal of 1.94.
Put all four at 5% and their powers become 55.9%, 51.0%, 54.2% and 61.8%.
The ordering nearly reverses. The raw row bootstrap, which is the worst reading in the essay by a wide margin, is mid-pack once its conservatism is priced out; the recentred one is last; and Bonferroni, which is a correction rather than a reference distribution, is ahead of everything.
That is not an argument that the raw row bootstrap is fine. It cannot be run at 5%, because nobody using it knows what its critical value should be — the whole defect is that its reference distribution is in the wrong place, and an analyst has no way to shift it back. What the calibrated comparison says is where the gain in this field actually comes from: the four statistics extract roughly the same amount of information from the table, within about ten points of each other, and the differences reported are almost entirely differences in whether each one is being read at the level it claims.
The practical form of that is a narrower recommendation than “simulate the null”. The expensive construction — a couple of hundred re-runs of the whole eight-variant search — is buying a correct location and very little else. Anything cheaper that also gets the location right buys nearly the same thing, which is exactly what the Bonferroni row demonstrates on this particular table, and exactly what it fails to do on a table of eight windows of one series where the effective count is 2.44 rather than near eight.
So the question a table poses is not which reference distribution to build. It is whether the cheap correction’s implied count is close to the table’s real one — and that is a single number, measurable by the effective-count arithmetic rather than by simulating the search, which is the one piece of this field that does not cost more than the comparison it is about.
Two constructions, and what each is for
The pair is worth stating as a rule rather than as a result, because the choice recurs everywhere a search is priced.
Resample when the null is a statement about the data. For candidates that estimate nothing, the observed differentials are the object, and the least favourable configuration is a shift of them. Resampling rows carries the dependence for free and needs no model at all.
Simulate when the null is a statement about a model. For candidates that are fitted, the null makes the two forecasts identical in population and everything a sample shows is estimation noise whose distribution is a property of the fitting. There is nothing in the data to shift into the null position, because the null position is not a location — it is a whole distribution the data was never drawn from.
The specification search is the case where both descriptions apply to one table, and this is why the two halves do not compose: the multiplicity is a statement about the data, the nesting is a statement about the model, and each of the two constructions gets one of them right.
What is claimed here, and what is not
This essay takes the reference distribution for a table of fitted variants, and the claim is the comparison: the two constructions, at a null and at three alternatives, on the same tables.
What stays out and is named as a decision: residual resampling that keeps second-moment dependence, which is the repair to the fragility named above and is a different construction rather than a setting; a table whose variants are not all nested in the benchmark, where the displacement differs by row and there is no single closed form for it; and the case where the benchmark is one of the things being searched over, which stops being a test and becomes selection.
The boundary against the set field is where it has been since the fitting arrived: everything about resampling blocks of rows is established there, and every one of its reference distributions is built around a point this table is not at.
The checks, and the refusals that make them mean something
Three claims are gated in this field’s library. The row bootstrap is required to reject far below its nominal level at this null and the simulated distribution to be within three points of it, which is the size comparison stated as something that could fail in either direction. The simulated distribution is required to find a genuinely better variant more often than the row bootstrap at every alternative drawn. And its critical value for the maximum is required to be above a single comparison’s, which is the multiplicity the whole construction exists to price.
The refusal beside them is the most natural mistake available here: a reference distribution simulated from the variant the search selected rather than from the benchmark. It is the fitted model in hand at the end of a search, it produces series in which the null is false, and the maximum it draws is therefore the maximum under the alternative. Handed that construction the check throws, and it throws on a power collapse rather than on a size failure — the test still holds its level, it simply stops being able to find anything, which is the failure mode this field has now met from both directions.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two defects and one resampling — both name block bootstrap, bootstrap, critical value, null hypothesis, out of sample, reference distribution, specification search
- How long a block a multiplier shares — both name block bootstrap, critical value, null hypothesis, reference distribution, specification search
- The triangle that was not the multiplier's — both name block bootstrap, bootstrap, critical value, reference distribution, specification search
- A criterion is a prediction of the hold-out — both name nested models, out of sample, reference distribution, specification search
- Errors generated from a fitted model — both name block bootstrap, bootstrap, critical value, reference distribution
- Where the two searches cross — both name nested models, out of sample, specification search, stationarity
Named objects
A flat tag is an object no other essay names yet.
Block bootstrapBonferroniBootstrapClark–WestCritical valueMultiplicityNested modelsNull hypothesisOut of samplePowerReference distributionResidual resamplingSpecification searchStationarity