A length for each instrument
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
A rule for choosing a block length is a way of guessing a number nobody has. The field that priced the guessing makes that concrete: the best length on a draw averages 24.57 with a standard deviation of 16.23, a plug-in from the sample’s persistence lands at 14.36, a protocol length is 8 and the rule of thumb is 4.
All four of those are rules for guessing one target. This essay is the observation that there are two.
Two oracles
Everything below is four hundred samples of a hundred and twenty rows from a first-order autoregression at 0.7, with three hundred block resamples formed at each of seven candidate block lengths — 2, 4, 8, 12, 20, 32 and 48 — for each of the two windows. Every reading comes off the same resampled means, which is the design the first essay of this field is about.
The best block length on a draw is the one whose reading is nearest the truth, and each of the three instruments has its own truth. So each has its own oracle, and they are not the same length.
For the rectangular window the implied long-run variance wants 18.92 and the 95% point of the same resampled means wants 16.05. For the tapered window, 21.82 against 17.74.
The quantile wants a shorter block under both windows, by a ratio of 0.848 and 0.813.
That is a difference of three to four block lengths on a hundred and twenty rows, on a grid whose neighbouring points are a factor of one and a half apart. It is not a rounding difference between two ways of saying the same thing.
What an oracle is here
The word is used in the sense the feasible field fixed: the length that minimises the realised error on that draw, chosen with knowledge of the truth and available to nobody. It is a benchmark rather than a rule, and it is the right benchmark for a rule that also sees only that draw.
What is new is that it is computed twice per draw. The same seven candidate lengths, the same resampled means at each, two readings taken off them, and the argmin of each reading’s own error. So the two oracles are two functions of one computation, and the difference between them is not a difference in what was resampled.
Why a quantile wants a shorter block
The mechanism follows from what a longer block does, and it does two things that pull opposite ways.
A longer block keeps more of the dependence. That is what it is for: the resample’s implied autocovariance at lag k is the sample’s times a factor that goes to zero at the block length, so a short block truncates the dependence and biases the variance downwards. The field that separated the bias from the spread shows the bias falling steadily with the length.
And a longer block leaves fewer of them. A resample of n rows built from blocks of ℓ contains ⌈n/ℓ⌉ blocks, so at ℓ = 20 on a hundred and twenty rows the resampled mean is a sum of six nearly independent pieces. Six.
For a variance that hardly matters: the variance of a sum of six independent terms is estimated as well as the variance of a sum of thirty, because the target is a second moment and a second moment is what a variance estimator is good at.
For a quantile it matters a great deal. The resampled mean’s distribution is a sum of six terms rather than thirty, so it is further from normal, and its tail is where the 95% point lives. A short block buys back the central limit theorem, and a quantile reads the part of the distribution that needs it.
So the two instruments are trading the same bias against two different variances, and the one whose variance grows faster with the block length wants a shorter one.
The arithmetic of six blocks
The second half of the mechanism is worth doing with numbers, because fewer blocks sounds like a small effect and it is not.
At a block length of 20 on a hundred and twenty rows, a resample is six blocks. At 8 it is fifteen; at 4 it is thirty. The resampled mean is a sum of that many nearly independent contributions, so its distribution’s departure from normality shrinks like one over the square root of the count: six blocks is about 2.2 times further from normal than thirty.
The 95% point of a distribution is where that departure is largest. A skewness of a tenth moves a normal’s 95% point by about a twentieth of a standard deviation, and a skewness of a quarter by an eighth — while moving its variance by exactly nothing.
So the block length trades bias against normality for a quantile and bias against variance for a variance, and normality is the thing a long block destroys fastest. Six blocks is not many, and the block lengths these rules argue about are all in the range where a resample has between three and thirty of them.
Two dials, read as a two-by-two
The four oracle lengths are a window crossed with an instrument, so they can be read the way a factorial is, and the reading settles which of the two choices moves the target more.
The four are 18.92 and 16.05 for the rectangle, 21.82 and 17.74 for the taper.
- The window’s main effect — the taper’s average against the rectangle’s — is +2.30 lags.
- The instrument’s main effect — the quantile’s average against the variance’s — is −3.48 lags.
- The interaction is 18.92 − 16.05 − 21.82 + 17.74 = −1.21, or −0.60 per cell.
So the instrument moves the target half again as far as the window does, and the interaction is about a quarter of either main effect. The two dials are close to additive and the one nobody turns is the larger.
That is the finding in the form a rule-designer needs it. Choosing between a rectangle and a taper is a decision this collection has spent three fields on; choosing which reading the block length is being tuned for has never been presented as a decision at all, and it is worth more lags.
What a difference of three lags is, on this grid
The candidate lengths are 2, 4, 8, 12, 20, 32 and 48 — neighbouring rungs a factor of one and a half to two apart — so a difference of 2.87 lags is less than one rung in the region where the oracles sit.
That matters for how the number should be read. Neither oracle is at a length the other is not; both are averages over four hundred draws of an argmin that can only take one of seven values, and what differs is how the mass sits on those seven.
Reconstructing the shift makes it concrete. A mean fall of 2.87 lags is what about one draw in seven moving from 32 down to 12 would produce, or about a third of draws moving from 20 down to 12. So the quantile’s oracle is not a systematically different length on every draw; it is the same length on most draws and a shorter one on a substantial minority.
Which is worth knowing before reading the ratios. 0.848 and 0.813 describe a shift in a mean, and a shift in a mean of a seven-valued argmin is a statement about a tail rather than about a typical draw. The difference is real — four hundred paired draws make a 2.87-lag mean shift easy to establish — and its shape is not the uniform three-lag offset the ratio suggests.
What that does to the rules
A rule is a formula that returns a length, and none of the four in this collection knows which instrument it is about to be scored on.
Two of them return the same length whatever: a protocol length is 8, the rule of thumb is 4. So they are three to fourteen short of the quantile’s oracle and seven to eighteen short of the variance’s — badly short of both, and less badly short of the quantile’s, which is the shorter target.
The plug-in reads the sample’s persistence and returns 12.59 for the rectangle and 13.89 for the taper, which is also nearer the quantile’s oracle than the variance’s.
Every feasible rule here is short, and the quantile’s target is the shorter one, so every feasible rule is closer to being right about the quantile than about the variance. That is a fact about where these rules sit rather than about their design — the plug-in’s algebra is derived for a mean squared error in a long-run variance and lands nearer the other target by accident — and it is most of why the ordering between the two windows moves when the instrument does.
Which is the same trade one level up
The shape of this is familiar from three fields back and it is worth naming, because it says the finding is not a curiosity of block bootstraps.
A tuning parameter is chosen to minimise an error, and there is more than one error. The width of a covariance band is chosen to maximise a penalised likelihood and delivers a coefficient error, and the two have different argmaxes. A block length is chosen to minimise an implied variance’s error and delivers a critical value, and the two have different argmins.
In both cases the quantity the rule optimises is an intermediate — a likelihood, a variance — and the quantity a reader acts on is downstream of it. The gap between the two is not an error in the rule; it is the rule answering a question nobody asked, and the size of it is what decides whether that matters.
Here it is fifteen per cent of a block length, and the next essay is what fifteen per cent of a block length does to a recommendation.
One reading that does not move
Not everything differs between the two instruments, and the thing that does not is worth having.
The ordering of the four rules by how good their length is is the same on both. Under either instrument and either window the oracle is best, the plug-in next, then the protocol length, then the rule of thumb — which is the earlier field’s own ordering, unchanged.
So the two instruments agree about which rules are better at guessing and disagree about which window to use once they have guessed. That is a narrower disagreement than it might have been, and it is the reason this field’s conclusion is a correction to one sentence of the earlier one rather than to all of it.
Which target a practitioner should be aiming at
The two targets are not equally worth hitting, and the ordering between them is not the ordering of how fundamental they are.
Nobody reports a long-run variance. It goes into a standard error, and a standard error goes into a t-statistic or an interval — so the variance’s target is one step removed from anything printed, and being right about it is only worth what being right about the thing downstream is worth.
A critical value is printed, and its target is the one a test’s size depends on directly.
So of the two oracles measured here, the shorter one is the one aimed at the quantity a reader acts on, and the longer one is aimed at an intermediate. Every rule in use is derived for the intermediate, which is a reasonable thing to derive for — the algebra is standard and the quantile’s is not — and it means the rules are systematically aimed past the target that matters, in a direction that is now measured.
The oracle’s own spread
One number that does not appear above is worth having, because it decides how much any of this can be pushed.
The oracle is the best length on that draw, so it is a random variable, and the earlier field measures its spread at 16.23 against a mean of 24.57 — it moves more than any rule’s answer does. The averages here, 18.92 and 16.05, are averages of two such variables measured on the same draws.
What makes the comparison between them tighter than either of them alone is that they are paired: the same sample, the same resampled means, the same sorted array, read twice. So the difference of three block lengths is a difference between two readings of one computation rather than between two noisy averages, and its own variation is much smaller than either oracle’s.
What a shorter target does to a comparison
The reason a fifteen per cent difference in targets moves an ordering is that the two windows do not pay the same price for a block that is too long.
A tapered window weights each block down at its ends, so a tapered block of 20 contributes less than a rectangular block of 20 does — its effective length is shorter, which is what its attenuation function says. So at a shared nominal length the taper’s resample has more effective blocks than the rectangle’s, and it is further along the road back to normality.
That predicts what the next essay measures: being scored on a quantile should favour the taper, and being scored at a length that is too long for the quantile should favour it more. The two rules with fixed lengths — a protocol’s 8 and the rule of thumb’s 4 — are the two furthest from either oracle, and they are the two whose ordering reverses.
What the third instrument wants
Coverage has an oracle too, and it is deliberately not in the figure.
A coverage is a rate rather than an error, so its oracle on a single draw is degenerate: the interval either covers or it does not, and the best block length on that draw is any length whose interval happens to cover. The quantity has no per-draw argmin worth taking, which is why the third reading is reported at the length the quantile rule picked rather than at a coverage oracle of its own.
That is a real limitation and it cuts the right way. Reading coverage at somebody else’s chosen length is conservative: it gives coverage no chance to be flattered by a length chosen for it. If the coverage ordering agrees with the quantile’s — and it does, at every rule — the agreement is not an artefact of a shared oracle, because there is only one oracle in it.
What the difference is not
Two readings of the difference are available and one is wrong.
It is not that the quantile is noisier and its argmin is therefore pulled about. A noisier reading would give an oracle with a larger spread and the same centre, not a centre fifteen per cent lower under both windows and on paired draws. Noise moves a variance, not a location.
And it is not the grid. The seven candidate lengths are the same seven for both readings, so an argmin that lands on a lower grid point for one of them is landing there because the reading is lower there. The two oracles average 18.92 and 16.05 on a grid whose points near there are 12, 20 and 32 — so the difference is a shift in how often 12 is picked over 20, which is a shift in the reading rather than in the resolution.
What it would take to fix
The obvious repair is a plug-in rule aimed at the quantile, and it is worth saying why this field does not build one.
The plug-in’s algebra comes from a mean squared error expansion for the long-run variance: the bias goes like 1/ℓ, the variance like ℓ/n, and the optimum is a cube root. The corresponding expansion for a bootstrap quantile’s error needs the third moment of the resampled mean as well as its variance, which is a different and much less standard calculation — and it would produce a rule with the same shape and a different constant.
A different constant is exactly what the measurement says is needed: 16.05 against 18.92 is a ratio of 0.848, and a cube-root rule scaled by 0.848 would land there. So the repair is a coefficient, it is derivable, and deriving it properly is a piece of work this field does not do.
What this field establishes is that it is worth doing, which is the smaller claim and the one the measurement supports: the two instruments’ targets differ by fifteen per cent, and the rules in use are ten to twenty block lengths from either.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An ordering that depends on the rule — both name bias-variance, block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering, tuning parameter
- A block weighted inside itself — both name bias-variance, block bootstrap, dependence, long-run variance, reference distribution, resampling, stationary bootstrap, tapering
- A taper and a critical value — both name block bootstrap, critical value, dependence, long-run variance, reference distribution, resampling, stationary bootstrap, tapering
- Measuring a variance rather than a quantile — both name block bootstrap, block length, critical value, long-run variance, monte carlo, reference distribution, resampling, tapering
- The error no window repairs — both name bias-variance, block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering
- What a multiplier cannot keep — both name block bootstrap, critical value, dependence, estimation error, monte carlo, persistence, reference distribution, resampling
Named objects
A flat tag is an object no other essay names yet.
Bias-varianceBlock bootstrapBlock lengthCritical valueDependenceEstimation errorLong-run varianceMinimisationMonte CarloPersistenceReference distributionResamplingStationary bootstrapTaperingTuning parameter