Two defects and one resampling
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
A specification search needs a reference distribution and cannot resample one from its own differentials. The field that establishes why builds it the only way available: fit the benchmark on the whole sample, resample its residuals, build a new series on the same design, and re-search the table. Everything then turns on what the resampling keeps, and that field measures four of them against three worlds and finds each repair matched to its own defect.
It also names the world it did not build. The wild multiplier keeps a residual attached to its row, so a variance that depends on the design survives and any correlation between neighbours does not. The block resampling takes runs of consecutive residuals, so the correlation survives and the tie to the row does not. Against a world with both defects at once, each of them repairs half the problem and each is wrong.
That world is not exotic. A predictive regression on persistent covariates, with errors whose spread depends on the covariates and whose neighbours are correlated, is what most macroeconomic and most epidemiological series look like. This essay builds it.
What the truth is, and why it is available
Every number below is measured against the statistic’s own 95% point in a stated world — the thing every reference distribution is trying to guess. That quantity exists here because these are simulations: draw the world two and a half thousand times, compute the open search’s largest studentised difference each time, take the 95th percentile.
The four worlds are the same table of six equal-dimension candidates, differing only in what the errors carry.
Constant variance, independent rows. Truth 1.2047. Every resampling is close and the blocked multiplier is closest, at 1.2465.
A variance that is a function of the design. Truth 2.0430 — the world is much wilder, and the whole point is whether a resampling notices. The two that detach a residual from its row give 1.5835 and 1.5936, short by about 0.45. The two multipliers that keep it there give 2.0586 and 2.1176.
Rows that repeat each other. Truth 3.0695, wilder again. Now it reverses: the two multipliers drawn once per row give 1.6032 and 1.6040, short by about 1.45, and the block resampling gives 2.4795.
Both at once. Truth 3.8028. The independent draw gives 1.8986, the wild multiplier 2.1078, Mammen’s 2.1417, the block resampling 2.8326. Every one of them is a long way short, and the shortest fall belongs to the two constructions that were each right about one half.
The construction, which is one line
Nothing about the repair is clever, which is the point.
A wild bootstrap multiplies each residual by its own ±1. The residual never leaves its row, so whatever the conditional variance is — a function of the design, a function of the past — it survives untouched. What it destroys is any relationship between neighbours, because two adjacent residuals get independent signs and are as likely to be pushed apart as together.
So draw the sign once per run of ℓ consecutive rows instead of once per row. Inside a run every residual keeps its own size and its own position and they all move the same way, so their correlation survives at strength ℓ; across runs the signs are independent, which is what gives the reference distribution its variation.
That is the whole of it: blockwild in this fleet’s table of resamplings, a fifth entry beside the
four rather than a thing this field invented. At ℓ = 5 it gives 2.9988 in the world with both
defects, against the block resampling’s 2.8326 and the wild multiplier’s 2.1078.
What the rejection rates say, which is what an analyst sees
The critical values are the mechanism. What anybody actually experiences is a rejection rate, and the gap between them matters.
At a nominal 5% in the world with both defects, the independent draw rejects 29.0% of true nulls. The wild multiplier rejects 27.0%, Mammen’s 19.0%, the block resampling 8.0%, and the blocked multiplier 8.0%.
Two things in that list are worth pausing on.
A reference distribution that is 45% too small rejects six times too often. The relationship between a critical value and a rate is steep out in the tail, so an error in a quantile that looks tolerable — 1.90 where 3.80 was wanted — is not tolerable in the thing the quantile is used for.
And 8.0% is still not 5%. The blocked multiplier is the best of the five and it is not correct. It is short of the truth by 0.804, which is a fifth of the truth, and it rejects at more than half again its nominal level. Reporting it as the repair would be reporting a partial one as a complete one, and the size of what is left over is the next essay.
Why each resampling fails where it does
The table has a structure and it is worth stating rather than reading off.
A residual carries three things: its size, its position, and its neighbours. Resampling one at a time keeps the size and throws away the other two. A block resampling keeps the size and the neighbours and throws away the position. A wild multiplier keeps the size and the position and throws away the neighbours. A blocked multiplier keeps all three.
Each defect in the world is a statement about one of those.
A variance that depends on the design is a statement about position: the residual at row t is large because of what is in row t, so a resampling that moves it somewhere else has built a world where the large errors are in the wrong places. Both resamplings that move residuals fail here, and they fail by nearly the same amount — −0.4595 and −0.4494 — which is the signature of a single mechanism rather than two.
Correlation between neighbours is a statement about neighbours, obviously, and both per-row multipliers fail by nearly the same amount — −1.4662 and −1.4655. Again one mechanism.
The pairing is so tidy that it is worth naming as the diagnostic it is: when two resamplings that differ in one respect fail identically, the defect is in the respect they share. That is how the table was read, and it is more useful than any individual number in it.
The truth moves further than any of the guesses
Before reading the guesses it is worth reading the rules they are drawn against, because the largest number in this essay is not any resampling’s error.
The statistic’s own 95% point runs 1.2047, 2.0430, 3.0695, 3.8028 across the four worlds. It triples. Nothing about the table changed — same six candidates, same window, same origins, same true null — and the critical value a correct test would use went up by a factor of three.
That is the quantity a reference distribution exists to track, and it says why none of this can be replaced by a rule of thumb. A practitioner who fixes a critical value from the clean world and applies it in the dependent one is not making a small error of calibration; they are using a threshold a third of the size the situation calls for. Every resampling in the table, however badly it does, is doing better than that: even the independent draw moves from 1.3650 to 1.8986 across the same four worlds, which is a third of the movement it needed but is movement in the right direction.
The reason the truth moves is the reading. An open search takes the largest studentised difference over every ordered pair, and studentising divides by a standard error computed as though the origins were independent. When they are not, that standard error is too small, every one of the thirty statistics is inflated, and the maximum of thirty inflated statistics is inflated more than any of them. Dependence enters twice — once in the numerator’s distribution and once in the denominator’s mistake — which is why the effect is multiplicative rather than additive.
The clean world, where everything errs the same way
The first column is the only one in which no resampling fails, and it is worth reading because everything in it errs in the same direction.
Against a truth of 1.2047 the five give 1.3650, 1.3159, 1.3246, 1.3327 and 1.2465 — all of them above. A reference distribution that is too wide makes a test conservative: it rejects less than its nominal level, which the rates confirm, at 1.0% to 3.0% against a nominal 5%.
The mechanism is the same one that makes an in-sample fit optimistic, running the other way. The reference distribution is generated from residuals of a fitted benchmark, and those residuals are built on the same rows that estimated the coefficients, so the simulated series carries a little more variation about its own fit than the real one carries about the truth. That surplus is a function of how many coefficients were fitted and how many rows there were, which is the same q and the same n as the criterion half of this field — the two halves fail to the same arithmetic from opposite ends.
The practical form: in a world with no defect in it a generated reference distribution is safe, and the safety is not free. It costs power, and how much is not measured here because every world in this essay is a true null.
Mammen’s multiplier, and what a third moment is not for
One row in the table exists to be a negative result.
Mammen’s two-point multiplier has mean zero, variance one and third moment one, where the symmetric ±1 has third moment zero. It is the standard repair for a wild bootstrap’s skewness, and it is here because the previous field measured skewed errors as a third defect and found it earning its keep.
In these four worlds it does almost nothing that the symmetric multiplier does not. Against the design’s heteroskedasticity it gives 2.1176 where the symmetric gives 2.0586, both a little over a truth of 2.0430. Against dependence it gives 1.6040 where the symmetric gives 1.6032 — indistinguishable, and both hopeless. In the world with both it gives 2.1417 against 2.1078.
The reading is not that Mammen’s multiplier is a bad idea. It is that a third moment is a repair for a third-moment defect, and none of the three defects in this essay is one. Adding it to a world it was not built for buys nothing, and the temptation to reach for the more sophisticated multiplier because the simple one is failing is the temptation this row exists to refuse.
What it costs to do nothing
The reason to build the fifth column rather than pick the best of the four is worth pricing, because the honest alternatives are all available.
Do nothing and use the independent draw, which is what almost every implementation does by default. In the world with both defects that rejects 29.0% of true nulls at a nominal 5%. Nearly a third of the searches an analyst runs on a table where nothing is worth anything come back significant.
Use the wild multiplier, on the grounds that heteroskedasticity is the thing everybody worries about: 27.0%. Almost nothing bought, because the defect doing the damage is the other one.
Use the block resampling, on the grounds that the series is persistent: 8.0%. A large improvement, and it arrives by fixing the defect that mattered more rather than by understanding the world.
Use the blocked multiplier: 8.0% on the same draws, on a critical value 0.166 closer to the truth. Better in the mechanism and not distinguishable in the rate at a hundred samples.
The last comparison is the honest one and it does not flatter the construction. Where both defects are present, most of the damage is coming from the dependence, so a resampling that repairs the dependence alone gets most of the way — and the extra that keeping the residual on its row buys is real, visible in the critical value, and small in the rate.
Where it is not small is the middle two columns. A block resampling costs 0.449 of critical value against a design-dependent variance, where the blocked multiplier costs 0.099; a wild multiplier costs 1.466 against dependence, where the blocked multiplier costs 0.716. The case for the fifth column is not that it wins by much in any one world. It is that it is the only one that is never badly wrong, and the analyst choosing between the other four has to know which defect their data has in order to choose — which is exactly the knowledge a resampling was supposed to make unnecessary.
The same defect, twice, in two different machines
The last thing worth saying about this table is that its defect is the same one the other half of this field runs into, and the coincidence is not one.
An optimism theorem counts rows. A bootstrap counts draws. Both are arithmetic over a number of independent things, and persistence is precisely the condition under which there are fewer independent things than there are rows. So the criterion starts buying coefficients it should not at the same point that the reference distribution starts being too tight, and both fail in the direction of finding more than is there.
The two repairs even have the same shape. The repair for the criterion is a penalty computed from an effective sample size rather than from n; the repair for the resampling is a multiplier shared over a run rather than drawn per row. Both replace n independent things with n/ℓ independent things for some ℓ that has to be chosen, and in both cases choosing it is the hard part — which is why one of them is built here and the other is named and left.
That is the field’s whole argument in one sentence. A substitute for data nobody has is a count of independent things, and dependence is what makes the count wrong.
What is claimed here, and what is not
This essay takes which resampling survives two defects at once, and the claims are the four truths, the twenty critical values against them, the pairing of each failure with what its resampling throws away, and the size of the shortfall that is left.
What stays out and is named as a decision: the block length, which is a dial with an interior optimum and is the next essay entire; any question of power, since every world here is a true null; and a resampling of the rows rather than the residuals, which the neighbouring field measures and which is a different construction with a different null.
The boundary against the field that built the four is that this one adds a world rather than a resampling. The three worlds, the four resamplings, the argument that a reference distribution here has to be generated rather than resampled, and the reading being tested are all established there and used here without being re-derived. What is new is the fourth world and the fifth column.
The checks, and the refusals that make them mean something
Two claims are gated in this field’s library. The blocked multiplier is required to be the closest of the five in the world with both defects, and each single-defect world is required to still be repaired by its own resampling rather than by the new one — which fails if the new construction were merely dominating rather than composing. And the repair is required to be partial: the check refuses to pass if the best of the five gets within a tenth of the truth, because a field whose result is this is still wrong by a fifth must not be able to quietly become a field whose result is this works.
The refusal for this essay belongs to the next one and is stated here because it is about this table: a block length chosen because of what it does to the observed p-value. Taking the smallest p-value over eight lengths rejects 20.0% of true nulls against 3.3% for a length chosen in advance. The dial is a search, and a search over reference distributions is still a search.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The triangle that was not the multiplier's — both name autocorrelation, block bootstrap, bootstrap, critical value, error rate, heteroskedasticity, persistence, reference distribution, resampling, residual, specification search, wild bootstrap
- A null with a model in it — both name block bootstrap, bootstrap, critical value, null hypothesis, out of sample, reference distribution, specification search
- The corner the test is calibrated at — both name bootstrap, error rate, monte carlo, null hypothesis, out of sample, reference distribution, specification search
- A block weighted inside itself — both name autocorrelation, block bootstrap, reference distribution, resampling, residual, wild bootstrap
- A length for each instrument — both name block bootstrap, critical value, monte carlo, persistence, reference distribution, resampling
- The instrument and the reading — both name block bootstrap, critical value, monte carlo, persistence, reference distribution, resampling
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationBlock bootstrapBootstrapCritical valueError rateHeteroskedasticityMonte CarloNull hypothesisOut of samplePersistenceReference distributionResamplingResidualSpecification searchWild bootstrap