What choosing the length costs
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
Three essays of one field and three of another have compared two block windows. This one puts the quantity they are comparing beside the quantity a practitioner loses by having to choose a block length at all, and the second is 3.4 times the first.
Three numbers on one scale
At a hundred and twenty rows, in points of the error in a block resample’s implied long-run variance:
What the window choice buys, at the best available length: 2.12 points. That is the whole subject of the comparison this field inherited, and it is a real difference at 22.1 paired standard errors.
What the best feasible rule gives up against that same length: 7.26 points. The plug-in delivers 43.07% for the tapered window where the oracle delivers 35.81%.
What the rule of thumb gives up: 26.01 points. It delivers 61.81%.
So the sentence that follows from the whole area is not use the taper. It is estimate the block length, because that is where the points are, and whether the taper is the right window at all depends on it.
The whole error these sit inside
Even 7.26 points is a part rather than a whole, and the field that measured the whole is the right place to read it against.
The error in a block resample’s implied long-run variance at a hundred and twenty rows is about forty per cent at the best window and the best length, of which 31.0 points are spread, 18.2 the window’s own attenuation and 8.2 the sample’s. The window choice moves four of those points and the block-length choice moves seven.
Set out as a hierarchy: about forty points of error, of which about seven are available to a better choice of block length, about two to a better choice of window, and the remaining thirty-one to nothing either choice can reach — they belong to the sample being a hundred and twenty rows long.
That is the shape of nearly every tuning result in this collection. The window feeding a criterion moves one per cent of the error it feeds; a tuning parameter chosen per candidate moves four thousandths of a regret of two hundredths. The dials are small and the quantity they are set inside is not, and a field that only ever reports the dial’s range has told a reader the wrong thing.
Why the shortfall is as large as it is
Seven points against a two-point window gap is worth a mechanism rather than a shrug.
The oracle’s block length has a standard deviation of 16.23 across draws, against a mean of 24.57. No rule can track a target that moves that much, and the previous field’s essay is about why: the per-draw argmin is the population optimum plus a wobble, and the wobble is much larger than the curvature of the expected error curve near its minimum.
So every feasible rule lands near the population optimum and the oracle lands wherever that draw’s noise put it, and the difference between “the best length for this draw” and “the best length on average” is most of the 7.26 points.
That is not a shortfall anybody can close. It is the price of the benchmark being per-draw, which is the right benchmark for a rule that also sees only that draw — a rule and its benchmark should have the same information, or the comparison measures the information rather than the rule.
What is actually recoverable
Separating the irreducible part from the recoverable part is what makes the number usable, and the comparison between the three feasible rules does it.
The plug-in gives up 7.26 points. The protocol length gives up 11.51. The rule of thumb gives up 26.01. The difference between the best and worst feasible rule is 18.74 points, and it is entirely available: it costs one line of arithmetic to estimate a lag-one correlation and put it into a formula.
So the honest split is that about seven points are the price of not knowing the truth, and up to nineteen more are the price of not looking at the data — and only the second is a choice.
That reframes the practical advice. It is not that a block length is costing seven points with nothing to be done about it. It is a rule that reads the data recovers most of what a rule that does not is losing, and the recovery is much larger than any window comparison.
Why the shortfall is measured against the oracle at all
A shortfall against something unreachable invites the objection that it is not a cost anybody pays, so it is worth defending the benchmark.
The oracle is not there to be reached. It is there to answer how much of this error is about the tuning at all, and that question has no other answer: comparing the three feasible rules against each other gives the 18.74 points that are recoverable and says nothing about whether the remaining error is mostly tuning or mostly the sample.
With the oracle in the table, the split is visible — about seven points that no rule can recover, about nineteen that any data-reading rule does, and about thirty that belong to the sample length. Without it, the same three rules look like three points on an open scale.
That is the same argument the estimated-covariance field makes for its own oracle column: a benchmark that nobody can run is what turns a ranking into a proportion.
The same reading, one field along
This is the second time in two rounds that a comparison turned out to be small beside the thing it sits inside, and the pair is worth putting together.
The whole error of a block resample found the window choice moving four points of forty, and its closing sentence is that a dial with a range of four inside an error of forty is worth setting and is not worth arguing about.
This field finds the same dial’s setting — the block length rather than the window shape — moving seven points, and the rules that read the data recovering nineteen more. So the ordering of the three decisions is: how much data there is, then how the length is chosen, then which window.
Nothing about that ordering was available from inside any of the fields that made the comparisons. Each of them held two of the three fixed and reported the third, which is the only way any of them could be measured, and the proportions only appear when the three are put on one axis.
Where the points would go instead
If seven points are unavailable and nineteen are recoverable by reading the data, it is worth asking what else on the list is worth more than either.
More rows. The whole-error field measures the error at each window’s best setting across four sample sizes — 42.7%, 34.2%, 29.0% and 22.1% for the rectangle across 120, 240, 480 and 960 rows. Each doubling buys about four points, which is more than the window choice and less than the length choice at the smaller sizes.
A model instead of a weighting. A construction that fits an autoregression and generates from it has an autocovariance at every lag rather than at the ones a window keeps, which is why a sieve escapes the ceiling block methods sit under. It buys that by assuming a form, and what a misspecified form costs is a field of its own.
And nothing else on this list. The window shape, the taper’s exact profile, the block-length grid’s resolution: all of them are inside the two points.
The curvature the shortfall implies
The oracle’s length has a mean of 24.57 and a standard deviation of 16.23, a coefficient of variation of 0.66 — the target moves by two thirds of its own size from draw to draw. That number and the 7.26-point shortfall together pin down how steep the error curve is, which is what decides what any displacement in length is worth.
A rule sitting at the population optimum is typically 16.23 lags from the per-draw best. On a locally quadratic curve that costs (c/2) × 16.23², and setting it equal to 7.26 points gives a curvature of about 0.028 points per squared lag.
Run that backwards through the other two rules and it says where they are standing. The protocol length’s 11.51 points imply a displacement of about 12 lags from the optimum; the rule of thumb’s 26.01 points imply about 26.
The second of those is larger than the optimum itself, which cannot be a displacement upwards from 24.57 and would put the rule below zero going downwards. So the quadratic has failed on the short side, and it has failed in the direction the rest of this collection keeps finding: a tuning parameter set too small costs more steeply than the same distance too large, because a block shorter than the dependence discards the dependence rather than merely estimating it noisily. The rule of thumb is short, and the quadratic understates what being short costs it.
What a doubling of the sample is worth, in the same points
The four sample sizes give the exchange rate that makes all of these comparable, and it is worth computing rather than approximating.
The rectangle’s error at its own best setting runs 42.7%, 34.2%, 29.0% and 22.1% across 120, 240, 480 and 960 rows. The three doublings buy 8.5, 5.2 and 6.9 points, so a doubling is worth about 6.9 points on average across the range.
Which prices everything else on this page in rows:
- The window choice, at 2.12 points, is worth about 0.3 of a doubling — a sample a quarter longer again.
- The block length, at 7.26 points, is worth about one doubling — the whole difference between a hundred and twenty rows and two hundred and forty.
- The gap between the best and worst feasible rule, at 18.74 points, is worth about 2.7 doublings — a sample six and a half times as long.
That last line is the one worth carrying out of the field. Replacing a rule of thumb with a plug-in that reads a lag-one correlation buys what six and a half times the data would buy, and it costs one line of arithmetic. Nothing else on the list is within an order of magnitude of that exchange rate.
What a longer sample changes
The proportions move with the sample, and not in the direction that makes the window choice matter more.
The oracle’s block length grows with , so a fixed rule falls further behind: the protocol length’s deficit against the taper goes from 2.04 points at 120 rows to 4.89 at 960. The plug-in keeps pace and its advantage over the rectangle grows from 1.89 to 2.70.
So more data makes the choice of rule matter more rather than less, which is the opposite of the usual reassurance about tuning parameters. The reason is that the target is a function of and two of the three rules are not.
What this does not measure
One thing is missing from every number here and it is worth naming rather than leaving to be inferred.
Everything is scored on the error in the implied long-run variance — the quantity a block resample’s reference distribution has as its second moment, computed from the sample’s own autocovariances without resampling. That instrument is what makes a sweep this wide affordable: it needs no draws at all, so it carries none of the resampler’s noise, and it resolves differences a critical-value table would need twenty thousand draws to see.
What it does not measure is what a practitioner ultimately reads, which is usually a quantile of the reference distribution rather than its variance. The earlier field establishes that the two are not the same reading — the quantile gap between two windows is 0.206 where a pure scale difference predicts 0.143, so the two reference distributions differ in shape as well as in scale.
Whether the ordering reversal in this field survives being read on a quantile is therefore unmeasured, and the honest position is that it is a claim about variances.
The three fields in order
Six essays across two rounds have now been about two block windows, and it is worth setting out what each added, because the sequence is a good example of a comparison being refined until it changes.
The taper field computes each window’s exact attenuation and finds the taper’s advantage has not arrived at any block length a hundred and twenty rows can afford — with the crossing at ℓ = 20 and a difference there of three tenths of a point.
The crossing field finds that reading is on the wrong quantity: a sample reports 4.31 points at that length rather than three tenths, because the autocovariances the window is applied to are themselves attenuated.
The whole-error field finds the four points the window moves sit inside forty, of which thirty-one are a spread no window touches.
And this field finds the two points at the oracle become a reversal under two of three feasible rules, and sit beside a seven-point tuning shortfall and a nineteen-point recoverable gap.
Four refinements, each correct at its own setting, and the conclusion at the end is not a sharpened version of the conclusion at the start. It is a different recommendation about a different decision.
One thing on the list is unusual in being free. Reading the sample’s own persistence and putting it into a formula costs one line and recovers up to nineteen points, which is more than doubling the sample buys and considerably more than any choice of window. That is an unusually good return for an arithmetic operation, and it is the field’s one unambiguous recommendation.
What the field leaves for the next one
Two things this field names and does not measure.
A rule that chooses the window and the length together. Every rule here picks a length for a given window, and a rule choosing both is a search over a product set — which a neighbouring field prices for a different pair and finds sub-additive, so the two charges would not add here either.
And the same table on a law that is not a first-order autoregression. Everything is at a lag-one correlation of 0.7 under a geometric law, which is the world the block-window fields are built in, and a dependence with a horizon or a long tail would move every optimum in the table.
Neither changes the field’s own conclusion, which is about proportions rather than about levels, and both would change the numbers.
What is claimed here, and what is not
This essay takes what estimating a block length costs against the best available one. The claims are that at a hundred and twenty rows the gap between the two windows at the best available length is 2.12 points of a 35.81% error; that the best feasible rule gives up 7.26 points against that same length and the rule of thumb gives up 26.01, so the choice is 3.4 times the size of the comparison it is being made inside; that the difference between the best and worst feasible rule is 18.74 points and costs one line of arithmetic to recover; and that a fixed rule’s deficit grows with the sample, from 2.04 points at 120 rows to 4.89 at 960, because the target grows and the rule does not.
What stays out, and is named as a decision: the same comparison on a quantile. Every number here is an error in an implied variance, which is the instrument that makes the sweep affordable and is not the quantity a practitioner reads. The earlier field establishes that the two windows’ reference distributions differ in shape as well as scale, so a quantile reading could order the rules differently and is not attempted.
Also out: a rule that estimates the whole spectrum. The plug-in here reads one number off the sample. A rule fitting an autoregression of estimated order and deriving a block length from its whole autocovariance sequence would be a better estimate of the population optimum and a second selection problem — it would have to be priced on the same table it is being chosen on, which is what a neighbouring field measures for a different pair and is a field rather than a row.
The boundary against the whole-error field is that it puts the window’s four points inside forty and this one puts the length’s seven beside the window’s two. Both are the same discipline: measure the thing the choice feeds rather than the thing the choice is about.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Bias is not the whole of it — both name bias-variance, block bootstrap, block length, long-run variance, mean squared error, resampling, tapering, tuning parameter
- The reversal that was the instrument's — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering, tuning parameter
- A charge that reads the draw — both name dependence, mean squared error, monte carlo, regret, selection effect, tapering
- A rate times a size — both name closed form, long-run variance, monte carlo, regret, selection effect, tuning parameter
- What studentising costs — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling
- A lag the sample has less of — both name closed form, dependence, long-run variance, monte carlo, tapering
Named objects
A flat tag is an object no other essay names yet.
Bias-varianceBlock bootstrapBlock lengthClosed formDependenceLong-run varianceMean squared errorMonte CarloPlug inRegretResamplingSelection effectTaperingTuning parameter