Where a taper's case begins

Bias is not the whole of it

A window that reaches zero at its ends attenuates less and uses less of each block. The block length that minimises its bias is not the one that minimises its error, and comparing two windows at one length compares one of them mis-tuned.

Worth reading first: Where the bootstrap lies · The observations that repeat each other.

A block weighted inside itself compares block windows on their bias, which is the right comparison for a statement about orders. Its conclusion — that the taper is worse at every block length a hundred and twenty rows afford — is a statement about a bias at a shared block length, and it is two readings away from the decision a practitioner makes.

The claim is not that the earlier reading was careless. It is that a bias is the quantity an asymptotic argument is about, and a practitioner choosing a window is not making an asymptotic argument — they are choosing between two constructions each of which has a second dial, and the second dial is where the answer turned out to be.

The three quantities

What a longer block buys and what it costsA trapezoidal block at 120 rows, with the error split into the two things it is made of. The bias falls with the block length, because a longer block attenuates less, and it flattens at 23.2% because the sample's own autocovariances are short whatever window is applied to them. The spread rises with it, because a longer block means fewer of them. Their sum in quadrature has a minimum at ℓ = 16, which is not where either of the two has one. The faint line is the rectangle's total error, for scale: it is above the trapezoid's from ℓ = 12 onwards.00.2000.4000.6000.8004681216202432486496block lengthas a share of the long-run varianceℓ = 16the whole errorthe biasthe spread400 draws, n = 120neither part has the minimum
Fig. 1 A trapezoidal block at a hundred and twenty rows, with the error split into what it is made of. The slider is the sample length.

The implied long-run variance of a block resample is Σ_k κ(k)γ̂(k), and over draws it has a mean and a spread.

The bias falls with the block length. A longer block attenuates less, so the shortfall drops — from 59.9% at ℓ = 4 to 23.2% at ℓ = 24 for the trapezoid — and then flattens and turns up, because past a point the sample’s own attenuation dominates and lengthening the block stops helping.

The spread rises with it. A longer block means fewer blocks in a series of fixed length and more weight on the far lags, where γ̂ is noisiest. It runs from 10.9% at ℓ = 4 to 38.1% at ℓ = 24.

The error is their sum in quadrature, and it has a minimum where neither of the other two does: at ℓ = 16 for the trapezoid, where the bias is 26.3% and the spread 31.0%.

The faint line in the picture is the rectangle’s total error, for scale. It is above the trapezoid’s from ℓ = 12 onwards, which is four lags earlier than the bias crosses and eight earlier than the exact bias does. Three different crossings from three different readings of the same pair of windows, and the one a practitioner needs is the earliest of the three.

Two optima, and only one of them is the answer

Two optima, and only one of them is the answer. At 120 rows, each window's implied long-run variance is measured at every block length and two block lengths are picked out: the one that minimises the bias and the one that minimises the whole error. They are never the same. The rectangle's least biased block is 20 and its least wrong one is 12; the trapezoid's are 24 and 16. The bars are the root mean squared error at each choice, so the difference between the two bars in a pair is what choosing on the bias alone costs. Read at each window's own best block length, the trapezoid is at 40.7% against the rectangle's 42.7% — the reverse of the ordering the same windows have at any shared block length.
Fig. 2 Each window’s least biased block length and its least wrong one, with the error at each.

Every window has two block lengths worth naming and they are never the same.

The rectangle’s bias is smallest at ℓ = 20 and its error at ℓ = 12. The trapezoid’s bias is smallest at ℓ = 24 and its error at ℓ = 16. The raised cosine’s are 32 and 16.

Choosing on the bias costs something in every case: the rectangle’s error is 47.16% at its bias optimum and 42.66% at its error optimum, the trapezoid’s 44.62% against 40.67%, and the raised cosine’s 46.93% against 40.35%.

A block length chosen to minimise the bias is chosen on about two thirds of the quantity, and the third it ignores is the third that rises with the dial being set.

The ordering reverses

Here is the reading that changes the earlier conclusion.

At any shared block length the trapezoid is more biased than the rectangle over most of the range a hundred and twenty rows afford — at ℓ = 8 it is 41.8% against 37.5%, and the two cross at ℓ = 16.

At each window’s own best block length, on the error, the trapezoid is at 40.67% and the rectangle at 42.66%. The raised cosine is at 40.35%. The window that loses at a shared length wins at its own.

That is not a paradox and it is not a trick. A window and a block length are two dials on one construction, and a rule is worth what it is worth at its own setting. Comparing two windows at one ℓ is comparing one of them mis-tuned, and which one is mis-tuned depends on the ℓ chosen — which is why the comparison could come out either way and did.

Which quantity a comparison should use

Three quantities have now been used to compare these windows and it is worth saying which is right for which question, because all three are correct about something.

The exact bias, at a shared block length, is right for a statement about orders: a window that reaches zero has a second-order bias and one that does not has a first-order one, and that is a fact about limits which no finite sample can improve on. The taper field establishes it and nothing here disturbs it.

The sampled bias, at a shared block length, is right for asking what a resample of a given series implies about a covariance, which is a question somebody might ask. It differs from the exact one by a factor of fifteen at ℓ = 20 and a hundred and twenty rows.

The sampled error, at each window’s own best block length, is right for choosing a window, because that is what choosing a window means: taking the construction and setting its dials. It is the only one of the three that answers the practitioner’s question, and it is the one nothing had computed.

Three quantities, three answers, and the disagreement between them is entirely about which question was asked. That is a better outcome than one of them being wrong, and it is a worse one for anybody wanting a single sentence about tapered blocks.

The same thing at every sample length

Where the taper's case begins, and it is not where the algebra says. The block length at which a trapezoidal block's implied variance stops being more biased than a rectangular one's, against the length of the sample. Computed exactly — from the law's own autocovariances, with no sampling in it — the answer is 19.2 and does not depend on the sample at all. What a sample of 120 rows reports is 13.3, and the reported crossing walks out towards the exact one as the sample grows: 13.3, 15.0, 16.4, 18.0. The mechanism is that the autocovariances the window is applied to are themselves attenuated, worst at the longest lags, and the window that discards those lags loses less of them.
Fig. 3 Where the two windows change places on the bias, against the sample length.

The reversal is not an artefact of one sample size. At 120, 240, 480 and 960 rows the rectangle’s best error is 42.7%, 34.2%, 29.0% and 22.1%, and the trapezoid’s is 40.7%, 31.5%, 26.4% and 19.6%. The trapezoid is ahead at every one of them, and its lead grows in relative terms — from 5% of the rectangle’s error at a hundred and twenty rows to 11% at nine hundred and sixty.

The block lengths move too. The rectangle’s error optimum runs 12, 12, 16, 20 and the trapezoid’s 16, 16, 20, 24, so both windows want longer blocks as the sample grows and the trapezoid always wants a longer one than the rectangle.

The taper’s case begins at a hundred and twenty rows, on this quantity, at each window’s own setting. Which is a different answer from the one the same comparison gives at a shared length on a bias, and both answers are correct answers to their own questions.

What proportion of the error is bias

The split between the two components is worth reporting because it says the trade is real rather than nominal.

At each window’s own error optimum, the share of the error that is bias rather than spread is 0.72 for the rectangle and 0.65 for the trapezoid at a hundred and twenty rows, and 0.69 and 0.59 at nine hundred and sixty. So the error is bias-dominated throughout, and the taper’s advantage is that it is less bias-dominated — it converts some of a bias into a spread, and quadrature is kind to that trade because the two add in squares.

That is the mechanism in one line. Two errors of 28 and 0 have a total of 28; two of 26 and 31 have a total of 40; two of 23 and 38 have 44. The trade is only worth making while the thing being traded away is the larger of the two, which is why the optimum sits where it does and why it moves as the sample grows.

What a practitioner would actually do

The sweep above sets each window’s block length to whatever turns out to be best, which is a quantity nobody has. It is worth being clear about what that assumption is worth, because it is doing more work than the window choice is.

At a hundred and twenty rows the trapezoid’s error runs 46.0%, 40.8%, 40.7%, 42.4% and 44.6% across block lengths of 8, 12, 16, 20 and 24. The curve is flat near its minimum — anything from 12 to 20 is within two points of the best — so a practitioner who misses the optimum by four lags loses less than the window choice buys.

The rectangle’s is 44.0%, 42.7%, 44.5%, 47.2% and 49.9% over the same range, which is less flat: it rises faster past its optimum, because the rectangle keeps more weight at the long lags where γ̂ is worst.

So the practical statement is not use a trapezoid at sixteen. It is that the trapezoid is more forgiving of a badly chosen block length as well as better at a well chosen one, which is the more useful of the two properties and is invisible in any comparison at a single ℓ.

Where the flattening comes from

One feature of the bias curve deserves a note because it is not the textbook shape.

A block bootstrap’s bias in the long-run variance is usually described as falling monotonically with the block length, towards zero. Here it falls, flattens around ℓ = 20 to 24, and then rises: the trapezoid is 23.2% short at ℓ = 24, 24.3% at 32, 30.1% at 48 and 50.0% at 96.

The rise is the sample rather than the window. At ℓ = 96 on a series of 120 rows there are barely more than one block’s worth of independent placements, the sample autocovariances at those lags are computed from a handful of pairs each, and what the window is weighting is mostly noise that averages towards zero. So the implied variance falls again, and the estimate is short for a completely different reason from the one it was short for at ℓ = 4.

The same estimate is biased downwards at both ends of the dial for opposite reasons, which is the shape the window sweep in the whitening field also has, and it is the reason an interior optimum exists at all.

Four points of bias for one and a half of spread

The bias shares at each window’s own optimum say exactly what the taper is trading, once they are turned back into the two components.

At its error optimum the rectangle’s bias is 0.72 of its error of 42.66 — so 30.7 — and its spread is therefore 29.6. The trapezoid’s bias is 0.65 of 40.67, which is 26.3, and its spread is 31.0. Both recombine in quadrature to the errors quoted.

So the taper buys a bias reduction of 4.4 points for a spread increase of 1.4: a three-to-one exchange. That is the whole of its advantage, and quadrature is what makes the trade profitable — at a point where the bias is the larger component, giving up three times as much of it as one takes on in spread must lower the total, and the optimum is where the exchange rate falls to the ratio of the two components.

The general form is worth having because it says when a taper stops paying. At an interior minimum of b² + s² the bias has to be falling s/b times as fast as the spread is rising — 1.18 for the trapezoid and 0.96 for the rectangle — so a construction whose error is spread-dominated rather than bias-dominated would want the trade run the other way, and a taper would be the wrong choice there for the same arithmetic that makes it the right one here.

How much more forgiving, in lags

“More forgiving of a badly chosen block length” is the essay’s more useful property and the two error curves put a number on it.

The trapezoid’s error runs 46.0, 40.8, 40.7, 42.4 and 44.6 across block lengths of 8, 12, 16, 20 and 24 — so it is within two points of its own best from ℓ = 12 to ℓ = 20, a window eight lags wide. The rectangle’s runs 44.0, 42.7, 44.5, 47.2 and 49.9, and is within two points of its best only from 8 to 12: four lags.

Twice the tolerance, and the asymmetry is on the long side. Past its optimum the rectangle’s error rises at 0.60 points a lag and the trapezoid’s at 0.49, because the rectangle keeps full weight at the long lags where the sample autocovariances are worst. Below their optima the trapezoid falls faster — 0.66 points a lag against 0.33 — so it is also the window that rewards getting the length right.

Put together with the field’s deferral about choosing ℓ from the data, that is the more important of the two findings. A rule that has to estimate the block length will miss it, and the window that loses less when it is missed is the window to use — which is the same window that wins when it is not, and by a larger margin than the tolerance is worth.

The same mistake, three fields along

Comparing two rules with one of them mis-tuned is a shape this collection meets often enough to name.

A criterion minimised over a tuning list is the same error with the sign flipped: there the rule is allowed to tune and the comparison does not charge for it, and here the rule is not allowed to tune and the comparison assumes it was. A basis reported at its best shape is refused for the first reason, and a window compared at a shared block length is refused for the second.

The general rule that covers both is that a comparison has to fix what each rule is allowed to choose and then charge each of them for what it chose. Fixing a dial for one rule and not the other is the first error; letting both choose freely and charging neither is the second.

What makes the second easier to walk into is that it looks fair. Two windows at ℓ = 8 is a controlled comparison in every visible respect — same series, same sample size, same law — and the thing that is not controlled is that eight is a good block length for one of them and a bad one for the other.

What is claimed here, and what is not

This essay takes the whole error of a block resample’s implied long-run variance rather than its bias. The claims are that the bias falls and then rises with the block length while the spread rises throughout, so the error has an interior minimum at neither’s; that every window’s bias-optimal block length differs from its error-optimal one — 20 against 12 for the rectangle, 24 against 16 for the trapezoid, 32 against 16 for the raised cosine at a hundred and twenty rows; that choosing on the bias costs between four and six points of error; that read at each window’s own best block length the trapezoid is at 40.67% against the rectangle’s 42.66%, reversing the ordering the same windows have at a shared length; and that the reversal holds at 240, 480 and 960 rows with the relative advantage growing from 5% to 11%.

What stays out, and is named as a decision: choosing the block length from the data. Everything here is at the block length that happens to be best, which nobody knows. A rule that estimates it is a selection problem of exactly the shape the window a whitening wants prices, and the answer there — that every feasible choice lands short of the best available, sometimes by 89% — is the answer to expect. Running it would turn this essay’s ordering into a comparison of two feasible rules, which is the honest version and is a different measurement.

Also out: the geometric taper. The stationary bootstrap’s attenuation has no block length in the same sense — its runs are geometric rather than exactly ℓ long — so it does not have two optima to compare and it is left out of the sweep rather than shoehorned into it.

The boundary against the gap a sample shows is that it establishes the finite-sample term and this one asks what the term does to a decision. The two together say that the earlier conclusion was right about its own quantity at its own setting and wrong about both.

The checks, and the refusals that make them mean something

Three claims are gated. The error is required to be at least as large as either of the two things it is made of, at every block length and window, because it is their sum in quadrature and a violation would mean the decomposition was wrong. Every window’s bias optimum is required to differ from its error optimum — on the rectangle as well as on the taper, so that the point is about the quantity rather than about the window. And the implied variance is required to agree with realised resamples, which is the identity the whole instrument rests on.

The refusal is the comparison this essay replaces. Two windows compared at one block length, when each has its own best one, is refused — with the rectangle’s 42.66% at ℓ = 12 printed beside the trapezoid’s 40.67% at ℓ = 16, and the observation that the window losing at shared lengths wins at its own.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • The error no window repairs — both name attenuation, bias-variance, block bootstrap, block length, long-run variance, mean squared error, resampling, sample autocovariance, tapering, variance
  • What choosing the length costs — both name bias-variance, block bootstrap, block length, long-run variance, mean squared error, resampling, tapering, tuning parameter
  • A taper and a critical value — both name bandwidth selection, bias, block bootstrap, long-run variance, resampling, tapering
  • The reversal that was the instrument's — both name block bootstrap, block length, long-run variance, resampling, tapering, tuning parameter
  • A charge that reads the draw — both name bandwidth selection, mean squared error, model selection, sample autocovariance, tapering
  • A lag the sample has less of — both name bandwidth selection, long-run variance, model selection, sample autocovariance, tapering

Named objects

A flat tag is an object no other essay names yet.

AttenuationBandwidth selectionBiasBias-varianceBlock bootstrapBlock lengthLong-run varianceMean squared errorModel selectionResamplingSample autocovarianceTaperingTuning parameterVariance