Scoring a search without spending data

How long a block a multiplier shares

Sharing a sign over more rows keeps more of the dependence and leaves fewer independent signs to build a distribution from. The bias falls from 1.6885 to 0.8479 and the spread rises from 1.3073 to 2.1716, and the rejection rate walks straight through its nominal level on the way from 11.3% to 1.3%.

Worth reading first: The experiments that could have happened · Where the bootstrap lies.

The construction that survives both defects has a number in it, and the previous essay set it to five without argument. A multiplier drawn once per run of ℓ consecutive rows keeps the dependence between neighbours at strength ℓ, and ℓ has to come from somewhere.

It is worth being clear about what kind of number it is. It is not a nuisance parameter that a larger sample makes irrelevant, and it is not something a rule of thumb settles. It is a dial with a failure at each end, and the two failures are of different kinds — one is a bias and one is a variance — so the setting that minimises the first is not the setting that minimises the second, and neither is the setting that gets the rejection rate right.

The two ends

Run the world with both defects — a variance that is a function of the design, rows that repeat each other — and vary nothing but ℓ.

A bias against a variance, with the answer in between. How wrong one sample's reference distribution is, split into the two things it is wrong by. Sharing the multiplier over more rows keeps more of the dependence and closes the bias from 1.688 to 0.835; every row it is shared over also removes an independent sign from the 101 the sample started with, and the spread of the resulting quantile rises from 1.307 to 2.172. The distance a practitioner with one sample is actually exposed to is the two together, and it is smallest at ℓ = 5.
Fig. 1 How wrong one sample’s reference distribution is, split into the two things it is wrong by. The two curves cross and the sum of them has a minimum neither end is near.

At ℓ = 1 the multiplier is drawn per row, which is the ordinary wild bootstrap. It keeps no dependence at all, the simulated world is far tamer than the real one, and the reference distribution’s 95% point comes out at 2.1143 against a truth of 3.8028 — short by 1.6885.

At ℓ = 30 the multiplier is shared over thirty rows. The dependence survives, and the reference distribution is now being built from 4 independent signs. Its 95% point lands at 2.8898, short by 0.9130 — better on average — and it moves between samples with a standard deviation of 2.1716, against 1.3073 at ℓ = 1.

The bias falls and the spread rises, and they are not the same kind of error. The bias is a statement about the method: at ℓ = 1 it is wrong the same way every time. The spread is a statement about the sample: at ℓ = 30 it is right on average and any particular analyst is a long way from the average.

What the two ends say about the errors themselves

The reading at ℓ = 1 is worth one more line than it gets, because it is a measurement of the world rather than of the dial.

At ℓ = 1 the resampled series carries no dependence, so its reference distribution is the one a world of independent rows would produce. Against a truth of 3.8028 it reads 2.1143 — short by 44.4%.

A critical value scales roughly with the square root of the long-run variance, so the ratio 2.1143/3.8028 = 0.556 implies that the errors’ true long-run variance is about 1/0.55621/0.556^2, or 3.2 times what independent rows would give. That is an approximation — the statistic’s shape changes as well as its scale — and it is the right order.

So the rows in this world are worth about a third of the same number of independent rows, and for a first-order process an inflation of 3.2 corresponds to a lag-one correlation near 0.53.

Which is the useful frame for the dial. The sweep is not searching for a block length in the abstract; it is searching for one long enough to reproduce an inflation of 3.2, and the triangle a shared multiplier imposes means it can only ever approach that from below. The lower end of the dial is not merely a bad setting — it is the setting at which the construction is solving a different problem, and everything the sweep does above it is the construction climbing back towards a number it cannot reach.

Where the bias stops falling

The bias curve has a shape worth reading rather than summarising, because it does not fall the way the motivating argument suggests.

It falls steeply — 1.6885 at ℓ = 1, 1.3400 at 2, 1.0531 at 3, 0.8479 at 5 — and then stops. At ℓ = 8 it is 0.8883, at 12 0.8838, at 20 0.8347, at 30 0.9130. Past about five rows, sharing the multiplier over more of them buys essentially nothing.

That is what a correlation length looks like. The errors here are autoregressive at 0.7, so a residual and its neighbour four rows away are correlated at 0.24 and eight rows away at 0.06. A run of five already contains most of the dependence there is to keep, and a run of thirty is keeping the same dependence with a quarter as many signs.

Which means the remaining 0.85 of shortfall — a fifth of the truth — is not dependence the construction is failing to keep. It is something else, and this field does not identify it. Two candidates are worth naming rather than leaving implicit: the residuals of a fitted benchmark are not the errors, and their own dependence is distorted by the fit; and the rolling scheme induces a dependence between forecast errors at neighbouring origins that no resampling of the estimation residuals can reproduce. Neither is measured here, and calling the construction complete would require one of them to be.

The rate walks through the level

The critical values are the mechanism and the rejection rate is what anybody experiences, and the rate does something the critical values do not obviously predict.

One dial, two ways of being wrong. The rejection rate at a nominal 5% as the multiplier is shared over more rows, in a world where the error variance depends on the design and the rows repeat each other. An unshared multiplier throws the dependence away, the reference distribution is too tight and the test rejects 11.3% of true nulls; a multiplier shared over 30 rows leaves 4 independent signs to build a distribution from, the distribution is far too wide, and the test rejects 2.0%. The setting that lands on the level is neither end, and it is not where the smallest bias is either.
Fig. 2 The share of true nulls rejected at a nominal 5% as the multiplier is shared over more rows. The two ends fail in opposite directions and the curve passes through the level it claims.

At ℓ = 1 the test rejects 11.3% of true nulls. At ℓ = 2, 7.3%. At ℓ = 3, 5.3%. At ℓ = 5, 4.0%. At ℓ = 8 and 12, 3.3%. At ℓ = 20, 1.3%.

So the rate is monotone and it crosses 5% between three and five, which is not where the bias is smallest and not where the spread is smallest.

The decisive comparison is ℓ = 5 against ℓ = 20. Their mean critical values are 2.9548 and 2.9680 — the same to within half a per cent, both about 22% short of the truth — and they reject 4.0% and 1.3%. Two settings that are indistinguishable on the quantity the bias curve measures differ by a factor of three on the quantity that is actually a level.

The reason is that a rejection depends on the reference distribution and the observed statistic from the same sample, and those two are correlated: a sample whose errors happen to be tame produces both a small statistic and a tight reference distribution, and the two errors partly cancel. How much they cancel depends on how much the reference distribution moves between samples — which is the spread curve, not the bias curve. At ℓ = 20 the reference distribution’s own quantile has a standard deviation of 2.1080, more than two thirds of what it is estimating, so it is often far too wide on exactly the samples where the statistic is large.

The mean critical value and the realised size are different measurements, and only the second is a level. Anybody tuning this dial by looking at how close the reference distribution’s quantile is to a known truth is tuning the wrong quantity — and in a real problem the truth is not available, which is the whole reason a resampling is being used.

Why the spread rises, and by how much it has to

The variance side of the trade is arithmetic and worth doing, because it says the shape of the curve before any simulation runs.

A reference distribution built from B resamples of a series of n rows, with the multiplier shared over runs of ℓ, has ⌈n/ℓ⌉ independent signs in each resample. Here n is a hundred and one, so that is 101 signs at ℓ = 1, 21 at ℓ = 5, 6 at ℓ = 20 and 4 at ℓ = 30. Every resample is a different assignment of those signs, and the whole variety a reference distribution can show comes from that assignment.

At four signs there are sixteen possible sign patterns. Seventy-nine resamples of a sixteen-point space is not seventy-nine draws of anything; it is the same handful of worlds repeated, and the resulting quantile is a fact about which sixteen worlds this particular residual sequence happens to generate. Its standard deviation across samples is 2.1716, against a truth of 3.8028 — it wanders by more than half of what it is measuring.

So the spread should rise roughly like √ℓ, and it does: 1.3073, 1.4696, 1.7640, 1.7795, 1.7618, 1.8419, 2.1080, 2.1716 against a √ℓ of 1, 1.41, 1.73, 2.24, 2.83, 3.46, 4.47, 5.48. The rise is slower than √ℓ because the residuals inside a run are themselves correlated, so a run of five was never worth five independent pieces of information to begin with. That slowness is the good news in this field: the cost of sharing the multiplier is smaller than the count of signs suggests, for the same reason that sharing it was necessary.

Two failure modes, and only one of them looks like a failure

The asymmetry at the two ends is the practically important part.

At short blocks the test over-rejects, at more than twice its nominal level. That is the failure everybody is watching for, and it announces itself: a specification search that finds something significant on two nulls out of nine will eventually be caught by somebody who tries it on data they know is empty.

At long blocks the test stops rejecting. At ℓ = 20 it rejects 1.3% of true nulls, which means it would also miss most true effects, and nothing about that announces itself at all. An analyst whose searches keep coming back insignificant does not usually conclude that their reference distribution has four independent signs in it. They conclude that there was nothing there.

A conservative test is not a safe test. It is a test with an unmeasured loss of power, and the loss here is large: a reference distribution built from four independent signs has quantiles that wander by more than half of what they are estimating, so the threshold in any given analysis is essentially arbitrary and is arbitrary on the high side often enough to matter.

What the dial is worth in the other three worlds

The whole essay has been run in the world with both defects, and it is worth checking what the setting costs where the defects are absent — because in a real problem nobody knows which world they are in.

In the clean world every one of the five resamplings is already close, and the blocked multiplier at ℓ = 5 is the closest of them at 1.2465 against a truth of 1.2047. Sharing the multiplier there does no harm: there is no dependence to keep, so the shared sign is simply a coarser version of an independent one, and the reference distribution is a little wider than it needs to be.

In the world with only a design-dependent variance the shared multiplier gives 1.9438 against a truth of 2.0430, where the unshared one gives 2.0586. So sharing costs a little there: the sign is doing the work of keeping the conditional variance either way, and the extra coarseness is pure loss.

In the world with only dependence the shared multiplier gives 2.3535 against 3.0695, which is a great deal better than the unshared one’s 1.6032 and slightly worse than the block resampling’s 2.4795.

The pattern is the one that argues for setting ℓ from the residuals rather than from a default: where there is dependence, sharing is most of the answer; where there is none, sharing costs about a twentieth of a critical value. The asymmetry is what makes an ℓ chosen a little too long a safer mistake than one chosen too short — until it becomes long enough that the sign count collapses, at which point it becomes the other failure entirely.

Each repair is for its own defect, and one is for both. The 95% point of the statistic's own distribution in each world, against the mean 95% point of five reference distributions built from one sample. Where the error variance is a function of the design, the two resamplings that detach a residual from its row fall short and the two multipliers that keep it there do not; where the rows repeat each other it is the other way round. With both defects at once the blocked multiplier — drawn once per run of 5 rows, so the residual never moves and its neighbours share a sign — is the closest of the five, at 2.999 against a truth of 3.803. It is still short by 0.804, and that shortfall is the next figure.
Fig. 3 The 95% point of the statistic’s own distribution in each world against the mean 95% point of five reference distributions built from one sample. Where the variance is a function of the design the two resamplings that detach a residual from its row fall short; where the rows repeat each other it is the other way round.

What to do with a dial that has to be set

Three things follow, and only the first is a number.

Set it from the correlation length, not from the sample size. The bias closes at about the point where the autocorrelation has decayed, which here is five rows for an autoregression at 0.7. That is a quantity an analyst can estimate from the residuals directly, before any search runs, and it does not require knowing the answer.

Do not tune it on the statistic. This is the refusal and it is not hypothetical, because the dial is exactly the kind of parameter people try a few settings of. Taking the smallest p-value over eight block lengths rejects 20.0% of true nulls, against 3.3% for a length chosen in advance. That is a search over reference distributions, it is worth a factor of six, and it leaves no trace in the output — the reported p-value looks like any other p-value.

And report the setting. A p-value from this construction is not interpretable without ℓ, in the same way that a p-value is not interpretable without the sampling plan. The number is a function of a choice that is not visible in the number, and the only repair for that is writing the choice down.

The same dial, in the other half of the field

This dial has a twin, and naming it is the point of putting the two halves of this field in one place.

The criterion half ends on a repair it names and does not build: a penalty computed from an effective sample size rather than from n, so that a fit on persistent rows is charged for the independent information it actually used. That is the same substitution as this one. Both replace n independent things with n/ℓ independent things for some ℓ that has to be chosen from the dependence, and both have the same two failure modes — too small an ℓ and the method behaves as though the rows were independent when they are not, too large and it throws away information that was there.

The difference is what the failure costs. Getting the penalty’s effective sample size wrong moves a selection, which costs regret and is bounded by the spread of the candidates’ risks. Getting the block length wrong moves a level, which is unbounded in the sense that matters: an error rate of 11.3% against a claimed 5% is not a slightly worse decision, it is a claim that is false.

That asymmetry is the reason this half of the field builds its dial and the other half names its own and stops. A selection rule that is a bit off produces a slightly worse model. A test whose reference distribution is a bit off produces a number with a percentage sign on it that means something other than what it says.

The crossing is in the dependence, not in the split. Regret of each rule as the design and the errors are made persistent at the same coefficient, scored on fresh rows because the closed form assumes exactly what is being taken away. An optimism theorem counts rows; when the rows repeat each other there are fewer of them than there are rows, the penalty is too small for the fit it is correcting, and the criterion starts buying coefficients it should not — its average winner grows from 3.31 coefficients to 3.90. The hold-out never used the theorem and overtakes at ρ ≈ 0.81. Schwarz's criterion, worst of the three on independent rows, is best on repeating ones — its heavier penalty is right for the wrong reason.
Fig. 4 The same defect one level up: regret of each selection rule as the design and the errors are made persistent together. An optimism theorem counts rows, and when the rows repeat each other the penalty is too small for the fit it is correcting — the criterion’s average winner grows from 3.31 coefficients to 3.90 and the hold-out overtakes it.

What is claimed here, and what is not

This essay takes the block length of a blocked wild bootstrap as a dial, and the claims are the bias falling and stalling, the spread rising monotonically, the rejection rate crossing its nominal level between three and five rows, and the two failure modes being asymmetric in visibility.

What stays out and is named as a decision: what the remaining fifth of shortfall is, which is named above with two candidates and measured for neither; anything about power, since every world here is a true null and the loss at long blocks is inferred from the size rather than counted; and an automatic rule for choosing ℓ, which would need a second reference distribution for a quantity estimated from the same rows and is the same problem one level down.

The boundary against the essay before it is that this one holds the world fixed and moves the setting. The four worlds, the five resamplings and the reason a reference distribution has to be generated at all are established there.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. The bias is required to fall by at least a quarter between one row and five and the spread is required to rise between the shortest run and the longest, which together are the two-sided trade and which either curve alone would not show. And the distance a practitioner is actually exposed to — the two combined — is required to be smallest at neither end of the dial, which fails if the field’s answer were as long as possible or as short as possible.

The refusal is a block length chosen because of what it does to this sample’s p-value. Taking the smallest of eight rejects 20.0% of true nulls against the honest 3.3%, and the check refuses the procedure rather than the number — because the number is a perfectly ordinary p-value and there is nothing about it to refuse.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • The triangle that was not the multiplier's — both name autocorrelation, block bootstrap, critical value, error rate, heteroskedasticity, reference distribution, resampling, residual, specification search, wild bootstrap
  • Errors generated from a fitted model — both name autocorrelation, block bootstrap, critical value, error rate, heteroskedasticity, reference distribution, resampling, residual, wild bootstrap
  • A taper and a critical value — both name block bootstrap, critical value, error rate, heteroskedasticity, reference distribution, resampling, residual, wild bootstrap
  • A block weighted inside itself — both name autocorrelation, bias-variance, block bootstrap, reference distribution, resampling, residual, wild bootstrap
  • A length for each instrument — both name bias-variance, block bootstrap, critical value, monte carlo, reference distribution, resampling
  • The residuals are not the errors — both name autocorrelation, block bootstrap, reference distribution, resampling, residual, wild bootstrap

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationBias-varianceBlock bootstrapCritical valueEffective sample sizeError rateHeteroskedasticityMonte CarloNull hypothesisReference distributionResamplingResidualSpecification searchTuningWild bootstrap