A block length chosen from the data

The length nobody has

Every comparison of block windows in this collection is made at each window's own best block length. That length has a standard deviation of sixteen across draws and averages twenty-five. No rule is aimed at it.

Worth reading first: Where the bootstrap lies · The observations that repeat each other.

The block-window field compares a rectangular block against a tapered one, and reaches its conclusion by reading each window at its own best block length — 12 against 16 at a hundred and twenty rows, because every window’s error-optimal length differs from every other’s and from its own bias-optimal one.

That is the right comparison and it ends on a sentence about what it cannot do: the best block length is the argmin of a quantity that needs the truth, and nobody has it.

This essay is about what a practitioner has instead.

Four rules, one of which is not a rule

A length written into the protocol. Eight lags, which is this collection’s own default, and the same number on every draw.

The rule of thumb. n1/3n^{1/3}, which is 4 at a hundred and twenty rows, and also the same number on every draw — it does not read the data at all.

A plug-in from the sample’s own persistence. Estimate the lag-one correlation, take the block length the mean-squared-error algebra gives for a first-order autoregression at that value. This is the only one of the three that looks at anything.

The best length on this draw. The argmin of the realised error, available to nobody, and the benchmark the other three are short of.

Three rules and a target none of them is aimed at. Which block length each rule picks, over 400 samples of 120 rows, for the tapered window. Two of the rules are points: a length written into a protocol is 8.00 on every draw and the rule of thumb is 4.00, because n to the one third does not read the data at all. The plug-in reads the sample's own persistence and lands at 14.36 with a standard deviation of 2.93. The length that would actually have been best on that draw averages 24.57 with a standard deviation of 16.23 and runs from 10 to 48 between its tenth and ninetieth percentiles. The target moves five times as much as the best estimate of it does, which is why no rule can be close to it and why the two that do not try are not merely worse — they are somewhere else.
Fig. 1 Which block length each rule picks, over four hundred samples of a hundred and twenty rows, for the tapered window.

What they pick

The two rules that do not read the data are points: 8.00 and 4.00, with standard deviation zero.

The plug-in lands at 14.36 with a standard deviation of 2.93, and its tenth and ninetieth percentiles are 12 and 16 — a tight and sensible-looking distribution.

The best length on the draw averages 24.57 with a standard deviation of 16.23, running from 10 to 48 between its own tenth and ninetieth percentiles.

The target moves five and a half times as much as the best estimate of it does.

Why that is the whole difficulty

A target with a standard deviation of sixteen, on a grid running from two to sixty-four, is not something any rule can track. The plug-in’s spread of 2.93 is not a failure to be more variable; it is an estimate of a population quantity — the block length that minimises the average error at this persistence — and the population quantity is not what the oracle column is.

The oracle is per draw. On a sample that happened to have a long run in it, the best block length is long; on one that did not, it is short; and neither of those is knowable from the sample without knowing the truth it is being scored against.

So the honest description is that the three feasible rules are all estimating one thing and the oracle is a different thing, and the gap between them is irreducible. That is the same shape the specification-search field measures for a different tuning parameter, and it reaches the same conclusion: when the target moves more than any estimate of it can, the differences between feasible rules are second-order and the shortfall against the oracle is not.

Not merely worse — somewhere else

The two rules that do not read the data are not simply short of the plug-in. They are aimed at different quantities.

The rule of thumb’s 4 comes from a rate argument: the block length should grow like n1/3n^{1/3} so that bias and variance vanish at the same order. That is a statement about how the length should scale, and it has a constant in front of it that the argument does not supply. At a hundred and twenty rows it gives 4 where the error is minimised near 24, and the constant is doing all of the work.

The protocol’s 8 comes from nowhere except habit and the fact that a plausible round number was needed before any data existed.

The plug-in’s 14.36 comes from the same algebra with the persistence estimated from the sample, so its constant is at least a function of the problem. It is still short — 14.36 against 24.57 — because the algebra minimises the mean squared error of a long-run variance and what is being scored here is the error of the implied variance at that length, which is a related but different optimum.

Two optima, and only one of them is the answer. At 120 rows, each window's implied long-run variance is measured at every block length and two block lengths are picked out: the one that minimises the bias and the one that minimises the whole error. They are never the same. The rectangle's least biased block is 20 and its least wrong one is 12; the trapezoid's are 24 and 16. The bars are the root mean squared error at each choice, so the difference between the two bars in a pair is what choosing on the bias alone costs. Read at each window's own best block length, the trapezoid is at 40.7% against the rectangle's 42.7% — the reverse of the ordering the same windows have at any shared block length.
Fig. 2 Two optima for one window — the block length that minimises the bias and the one that minimises the error — which is the same distinction one level down.

Three rules, three quantities, none of them the one being scored. That is the ordinary situation for a tuning parameter and it is worth naming, because a comparison that scores every rule on one downstream quantity is the only way it becomes visible.

The curve they are all landing on

The reason any of this is survivable is the shape of the error against the block length, and it is worth reading directly.

The error falls steeply from short blocks, flattens somewhere in the teens, and rises slowly after. A rule landing at 4 is on the steep part; a rule landing at 14 is nearly on the flat; the oracle wanders across the flat part and picks whichever end of it that draw happened to want.

That shape is why the rule of thumb is the worst of the three and why the gap between the plug-in and the oracle is a real number rather than an enormous one. It is also why the ordering between the two windows reverses depending on the rule — a tapered window at a quarter of the right length has thrown away most of what it was weighting, and the next essay is entirely about that.

No rule can explain three per cent of the target

“The target moves five and a half times as much as the best estimate of it” is the field’s headline and it converts into a hard bound on any rule at all.

If a rule’s chosen length is to be an estimate of the oracle’s, the variance it can explain is at most its own variance — a conditional mean cannot be more variable than what it is a mean of. The plug-in’s variance is 2.93² = 8.6; the oracle’s is 16.23² = 263. So the plug-in explains at most 3.3% of the oracle’s variance, and it does so only if it is perfectly aimed.

The two fixed rules explain nothing, exactly, because they have no variance to explain it with.

That is a stronger statement than “the rules are short of the oracle”. It says that a rule five times better than the plug-in — one with a spread of fifteen rather than three, tracking the target draw for draw — is the only kind of rule that could close the gap, and nothing in this collection or its literature proposes one. The gap is not a shortfall in the estimators; it is a variance the estimators do not have and could not acquire from the sample without knowing the truth.

The oracle’s own average is displaced long

The essay’s closing observation — that the arithmetic mean of a noisy argmin is not the argmin of the mean — has a size, and it explains most of the ten lags between the plug-in and the oracle.

The error curve is asymmetric about its minimum. The tapered window’s runs 46.0, 40.8, 40.7, 42.4 and 44.6 across block lengths of 8, 12, 16, 20 and 24, so it falls at 0.66 points a lag approaching the optimum and rises at 0.49 past it — a ratio of 1.35.

A realised curve is that shape plus noise, and the minimum of a noisy asymmetric curve lands preferentially on the flatter side. So E[argmin] exceeds argmin E, and it exceeds it by more the flatter the long side is and the larger the noise.

The 24.57 is therefore not a target. It is the average location of a noisy minimum on a curve that is a third flatter to the right than to the left, and a rule that hit it every time would be sitting past the point where the expected error is smallest. The plug-in’s 14.36 is short of the expected optimum too — the field’s own sweep puts that near 16 — but by four lags rather than by ten, and four lags on a curve this flat is worth under two points of error.

Which reframes the shortfall the field measures. Most of the visible gap is a property of the benchmark rather than of the rules, and the part that belongs to the rules is the smaller half.

Why the flat part is where the rules land

There is a reason the three feasible rules cluster short of the oracle rather than scattering around it, and it is not a coincidence of these three.

Every one of them is derived by minimising an expected error — the mean squared error of the long-run variance, averaged over draws. That average is minimised where the bias curve and the variance curve cross, and both of those are smooth functions of the block length. The per-draw argmin is minimising a realised error, which is the expected one plus a draw-specific wobble, and the wobble is much larger than the curvature of the expected curve near its minimum.

So the oracle’s answer is the population optimum plus noise, and the noise dominates. Any rule aimed at the population optimum lands near the population optimum, and the population optimum is where the expected curve is flattest.

That is why the plug-in’s 14.36 with a spread of 2.93 is not a poor estimate of 24.57 — it is a good estimate of a different and less variable quantity, and the arithmetic mean of a noisy argmin is not the argmin of the mean.

The plug-in’s algebra, checked

The plug-in rule is =(cnα)1/3\ell = (c\,n\,\alpha)^{1/3} with α=(2ρ/(1ρ2))2\alpha = (2\rho/(1-\rho^2))^2 and a constant cc that differs by window. That is a closed form and it has to be checked against something.

At a known persistence the length it returns is within two grid steps of the length the measured error curve is minimised at, on a grid running 2, 3, 4, 6, 8, 10, 12, 16, 20, 24, 32, 40, 48, 64. Two routes, one of which knows nothing about samples.

That check matters because a plug-in with a wrong constant would look exactly like a plug-in with a right one: it would produce a tight distribution of plausible lengths and be short of the oracle by some amount, which is what a correct one does too.

What each instrument costs to read. The number of draws each instrument needs to separate a rectangular block from a trapezoidal one at two standard errors, at a block length of 20 and 120 rows — measured from each instrument's own spread on the same draws. The implied variance needs 7.0 and the 95% point needs 20.2, a factor of 2.90 at this block length. There is a closed form beside it and it does not depend on either the scale or the size of the gap: the standard error of a p-quantile is √(p(1−p))/f(q) over √B where a standard deviation's is σ/√(2B), which at the 95% point of a nearly normal reference distribution is 3.30 times as many draws for the same statement. And the quantile route needs every one of those draws resampled, where the variance route needs none.
Fig. 3 The instrument all of these errors are measured with, which is what makes a sweep this wide affordable at all.

What the oracle column is for

It is not a rule and it is not a recommendation, and it is in every table here for one reason: it says how much of the error is available.

Without it, the three feasible rules deliver 45.28%, 57.28% and 44.96% for the rectangular window and look like three points on a scale with no top. With it, the best available is 37.93% and the three are short by 7.35, 19.35 and 7.03 points — which is a great deal more than the two points the window choice buys, and is the reading this whole field exists to supply.

What a longer block buys and what it costs. A trapezoidal block at 120 rows, with the error split into the two things it is made of. The bias falls with the block length, because a longer block attenuates less, and it flattens at 23.2% because the sample's own autocovariances are short whatever window is applied to them. The spread rises with it, because a longer block means fewer of them. Their sum in quadrature has a minimum at ℓ = 16, which is not where either of the two has one. The faint line is the rectangle's total error, for scale: it is above the trapezoid's from ℓ = 12 onwards.
Fig. 4 What the forty per cent is made of, which is the quantity every shortfall here is a shortfall inside.
Two optima, and only one of them is the answer. At 120 rows, each window's implied long-run variance is measured at every block length and two block lengths are picked out: the one that minimises the bias and the one that minimises the whole error. They are never the same. The rectangle's least biased block is 20 and its least wrong one is 12; the trapezoid's are 24 and 16. The bars are the root mean squared error at each choice, so the difference between the two bars in a pair is what choosing on the bias alone costs. Read at each window's own best block length, the trapezoid is at 40.7% against the rectangle's 42.7% — the reverse of the ordering the same windows have at any shared block length.
Fig. 5 Two optima for one window, which is the distinction between the target and what a rule is aimed at.

What the target actually is

One clarification, because “the best length on this draw” is doing a lot of work.

It is the block length, from the grid, at which the implied long-run variance computed from that sample is closest to the truth. It is not the length that minimises the average error, and it is not the length a clairvoyant practitioner would choose knowing the law — that would be a much less variable quantity and a much less useful benchmark.

The reason to use the per-draw argmin is that it is the right benchmark for a rule that also sees only that draw. A rule and its benchmark should have the same information, or the comparison measures the information rather than the rule. What it costs is that the benchmark is unreachable by construction, which is why every shortfall in this field is positive and none of them is a scandal.

What a sample shows, and what the algebra does. The difference between a rectangular block's implied long-run variance and a trapezoidal one's, as a share of the truth. Above the axis the rectangle is less biased and below it the trapezoid is. The heavy line is exact — computed from the law's own autocovariances — and it crosses at 19.2. The others are what samples of 120, 240, 480, 960 rows report, and every one of them exaggerates whichever window is ahead: at ℓ = 20, where the exact difference is 0.28 points, a sample of 120 rows shows 4.31 points — 15 times larger. That is the number the earlier reading of this comparison was missing: three tenths of a point is what the algebra says and not what a hundred and twenty rows report.
Fig. 6 What a sample shows against what the algebra says, which is the same gap between a draw and a law appearing one field earlier.

What a longer sample changes

The rules and their target both move with the sample size, and not together.

The oracle’s answer grows with nn — that is what the rate argument behind n1/3n^{1/3} says should happen, and it does. The plug-in’s grows too, since nn is in its formula. A block length written into a protocol does not move at all, and the rule of thumb moves by a cube root, which at these sizes is barely.

So a fixed length falls further behind the oracle as the sample grows, and the two rules that read the data keep pace. That is visible in the margin between the two windows at four sample sizes, where the fixed rule’s deficit widens from 2.04 points to 4.89 while the plug-in’s advantage widens from 1.89 to 2.70.

Which is a mild warning about a protocol number: a length chosen once for a study is a length chosen for the sample size the study expected, and studies rarely end at the size they planned for.

Two instruments on the same resample. At a block length of 20 and 120 rows, over 200 draws with 400 resamples each. The implied standard deviation is computed from the sample's own autocovariances through the window's attenuation, with no resampling in it at all: it reads 1.964 for the rectangle and 2.051 for the trapezoid, a paired gap of 0.087 at 10.7 standard errors. The 95% point of the resampled mean reads 2.775 and 2.981, a gap of 0.206 at 6.3. The second instrument is measuring the same difference through a noisier lens.
Fig. 7 Two instruments for reading the same difference, one of which made a sweep over four sample sizes affordable.

Where this sits beside the collection’s other tuning parameters

Four fields now price something chosen from the data, and the four answers are different quantities that are easy to confuse.

Choosing a whitening’s window per candidate rather than once costs four thousandths of regret — the smallest of the four, and the one about when the choice is made rather than whether.

Whether a fit can choose a covariance’s width at all is answered by nested families buying about what a criterion charges, so three defensible rules give widths a factor of three apart and deliver errors within one per cent.

Whether the tuning parameter or the list length is what separates two rules is answered: neither, and the difference had been a missing term in one criterion.

And this one asks what estimating a tuning parameter costs against the best available setting, which is the largest of the four and is the one nobody had put a number on.

The grid, and what it is allowed to decide

Every length here comes from a grid — 2, 3, 4, 6, 8, 10, 12, 16, 20, 24, 32, 40, 48, 64 — and the grid is a choice worth defending, since the oracle’s answer is an argmin over it.

It is geometric-ish rather than uniform, because a block length’s useful range is multiplicative: the difference between 2 and 4 is a different kind of difference from the one between 40 and 42. It runs to 64, which is more than half the sample at a hundred and twenty rows and far past any sensible answer, so the oracle is never at the edge. And it is fine near the bottom, where the error curve is steep and a coarse grid would misprice the rule of thumb.

The one thing it decides is the oracle’s spread, and it decides it in the conservative direction: a finer grid would let the per-draw argmin wander further, not less. So 16.23 is a floor on how much the target moves, and the gap between the plug-in and the oracle would widen rather than narrow under refinement.

Where the taper's case begins, and it is not where the algebra says. The block length at which a trapezoidal block's implied variance stops being more biased than a rectangular one's, against the length of the sample. Computed exactly — from the law's own autocovariances, with no sampling in it — the answer is 19.2 and does not depend on the sample at all. What a sample of 120 rows reports is 13.3, and the reported crossing walks out towards the exact one as the sample grows: 13.3, 15.0, 16.4, 18.0. The mechanism is that the autocovariances the window is applied to are themselves attenuated, worst at the longest lags, and the window that discards those lags loses less of them.
Fig. 8 The same grid used one field earlier to locate a crossing, where a coarse grid would have been a real hazard and a fine one is affordable.

What is claimed here, and what is not

This essay takes what a block length chosen from the data actually is. The claims are that at a hundred and twenty rows a length written into the protocol is 8.00 and the rule of thumb 4.00, both with standard deviation zero; that a plug-in from the sample’s own lag-one correlation lands at 14.36 with a standard deviation of 2.93 and a tenth-to-ninetieth range of 12 to 16; that the length that would have been best on each draw averages 24.57 with a standard deviation of 16.23 and a range of 10 to 48, so the target moves five and a half times as much as the estimate of it; and that the plug-in’s closed form lands within two grid steps of the length the measured error curve is minimised at.

What stays out, and is named as a decision: a cross-validated length. Splitting a dependent series to validate on needs a rule for where to cut and an argument that the two pieces are comparable, and putting an untested fourth rule beside three tested ones would add an answer rather than a measurement.

Also out: a rule that estimates more than the persistence. The plug-in here reads one number off the sample. A rule fitting an autoregression of estimated order and deriving the block length from its whole spectrum is available, and it is a second selection problem inside the first — which this collection prices elsewhere and which would need its own charge here.

The boundary against the field that compared the windows is that it reads each window at its own best length and this one asks what happens when nobody has it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • The error no window repairs — both name bias-variance, block bootstrap, block length, closed form, dependence, long-run variance, mean squared error, monte carlo, resampling, sample autocovariance, tapering
  • The gap a sample shows — both name bias-variance, block bootstrap, block length, closed form, long-run variance, mean squared error, monte carlo, resampling, sample autocovariance, tapering
  • The instrument and the reading — both name block bootstrap, block length, closed form, dependence, long-run variance, monte carlo, resampling, tapering
  • The reversal that was the instrument's — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering, tuning parameter
  • A block weighted inside itself — both name bias-variance, block bootstrap, closed form, dependence, long-run variance, resampling, tapering
  • An interval that carries its scale — both name block bootstrap, block length, closed form, dependence, long-run variance, monte carlo, resampling

Named objects

A flat tag is an object no other essay names yet.

Bias-varianceBlock bootstrapBlock lengthClosed formDependenceLong-run varianceMean squared errorMonte CarloPlug inResamplingSample autocovarianceSelection effectTaperingTuning parameter