The error no window repairs
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
Three essays of this field compare block windows. This one asks what the comparison is a comparison inside of, and the answer changes how much any of it matters.
The comparison the field inherited was between two block windows, and three essays have now made it three different ways. What none of them asks is how large the thing being improved is, and that turns out to settle the importance of the whole argument in one sweep.
The error, decomposed
At each window’s own best block length on a hundred and twenty rows, the implied long-run variance is wrong by 42.7% for a rectangular block, 40.7% for a trapezoid and 40.4% for a raised cosine. Those are root mean squared errors as a share of the quantity being estimated.
Each splits three ways, and the split is computable because the window’s own contribution is exact.
For the trapezoid at its best block length of sixteen: a bias of 26.3%, of which 18.2 points are the window’s own attenuation — the exact quantity the taper field computes — and 8.2 points are the sample’s. And a spread of 31.0%.
For the rectangle at twelve: a bias of 30.6%, being 22.6 from the window and 8.1 from the sample, and a spread of 29.7%.
The largest single component is the spread, on every window, and no window changes it by much.
Why the decomposition is available at all
The three-way split is not a modelling assumption; it is available because one of the three terms is exact.
The window’s own attenuation is Σ_k κ(k)γ(k) against Σ_k γ(k) with the law’s autocovariances, which the taper field computes in closed form. The sampled bias is the same sum with γ̂ in place of γ, averaged over draws. Subtracting the first from the second leaves what the sample’s own attenuation contributes, with no further assumption — the two quantities differ in exactly one thing, and it is the thing being isolated.
The spread is then whatever is left in the mean squared error once the bias is out, which is a definition rather than a decomposition.
A decomposition is only as good as the exactness of one of its parts, and this one has it. That is why the same split is not offered for a critical value: the window’s exact contribution to a quantile is not a computable quantity, so the three parts would all be estimates and their differences would be noise.
What each part responds to
The three parts have three different remedies and only one of them is a window.
The window’s attenuation is what the window choice moves. From the rectangle’s 22.6 points to the trapezoid’s 18.2 at their own settings — four points, which is what the whole of this field’s comparison is about.
The sample’s attenuation does not respond to the window at all. It is 8.1, 8.2 and 7.0 points for the three windows at a hundred and twenty rows, and it falls with the sample: 4.1, 4.1, 4.6 at two hundred and forty rows, 3.8, 3.6, 3.8 at four hundred and eighty, and 1.6, 1.4, 1.7 at nine hundred and sixty. Roughly halving with each doubling, which is the O(1/n) a sample autocovariance’s bias has.
The spread responds to the sample and to the block length and barely to the window. At each window’s own setting it is 29.7%, 31.0% and 28.5% at a hundred and twenty rows and 16.0%, 15.8% and 17.2% at nine hundred and sixty — falling like one over the square root of the number of blocks, which is what it is.
So of a forty-per-cent error, four points are available to the window choice and the rest belongs to the sample being a hundred and twenty rows long.
The pattern in those three columns is worth naming. One part of the error is a property of the window, one is a property of the sample, and one is a property of both. The first is the only one anybody has been choosing, and it is the smallest of the three at every sample size in the sweep.
At nine hundred and sixty rows the trapezoid’s split is 10.2 points of window, 1.4 of sample and 15.8 of spread. The window’s share has risen relative to the sample’s — the sample term decays faster — and the spread is still the largest. Twenty times the data does not change which component dominates.
The ceiling this sits under
That is not a new kind of result in this collection; it is the ceiling a resampling of residuals runs into, read on a different quantity.
Every one of these constructions resamples the same series. What each of them can carry about the dependence is bounded above by what the series carries, and a series of a hundred and twenty rows carries a sequence of autocovariances that is already short — 0.7773 at the first lag where the law says 0.8000, and shorter still once a fit has been taken out.
A window is a weighting of that sequence. It can decide which of the short numbers to lean on, and it cannot make any of them longer. So the ceiling on what any window achieves is set before any window is chosen, and the whole family of them lives in a four-point band under it.
A dial with a range of four points inside an error of forty is worth setting and is not worth arguing about, and this field spent three essays arguing about it — which was worth doing, because until it was measured nobody could say that either.
The same shape appears one field along and is worth the comparison. Estimating a dependence jointly with the regression rather than after it moves the coefficient by twenty-one standard errors and the decision by two, because the decision is a comparison between candidates all whitened by the same transform. Here the window moves the attenuation by four points and the error by two, because the error is dominated by a spread the window does not touch. Both are the same discipline: measure the thing the choice feeds rather than the thing the choice is about.
Why the spread is where it is
The largest component deserves a mechanism rather than a number, because it is the one nothing in this field has been able to move.
A block resample of n rows with blocks of ℓ lays down about n/ℓ blocks, each drawn independently from about n − ℓ + 1 starting positions. The implied variance is a weighted sum of the sample’s autocovariances, and each of those is itself an average over about n − k products of a persistent series — so it has an effective sample size far below n.
At a hundred and twenty rows and a block of sixteen there are between seven and eight blocks in a resample and the autocovariance at the sixteenth lag averages a hundred and four products of a series whose effective size at a correlation of 0.7 is about a fifth of its length. Everything about the arithmetic points the same way: there is very little independent information in the tail of the sequence being weighted, and the weighting cannot manufacture any.
Lengthening the block makes it worse — the spread runs from 10.9% at ℓ = 4 to 38.1% at ℓ = 24 for the trapezoid — because a longer block puts more weight on the lags with the fewest pairs behind them. Which is the trade the error’s minimum sits at, and is why the minimum exists.
What actually shrinks it
Three things do and none of them is a window.
More rows. The error at each window’s best setting runs 42.7%, 34.2%, 29.0% and 22.1% for the rectangle across 120, 240, 480 and 960 rows. Doubling the sample buys about four points, which is more than the window choice buys and is not usually available.
A longer block, up to a point. The block lengths the windows want lengthen with the sample — the rectangle’s error optimum runs 12, 12, 16, 20 and the trapezoid’s 16, 16, 20, 24 — and a practitioner using a fixed block length across sample sizes is mis-tuned at all but one of them.
A model instead of a weighting. A construction that fits an autoregression and generates from it has an autocovariance at every lag rather than at the ones a window keeps, which is the reason a sieve escapes the ceiling that block methods sit under. It buys that by assuming a form, which is the trade the whole estimated-covariance field prices: the cost of generality where it is not needed and the benefit of it where it is come out the same size.
What a forty per cent error means downstream
A number this large is worth translating, because the long-run variance is forty per cent wrong does not by itself say what a reader should do differently.
A long-run variance is what a standard error is built from, and a standard error enters an interval as a square root. So an implied variance that is forty per cent short gives a standard error about twenty-two per cent short, and an interval about twenty-two per cent too narrow.
What a narrow interval costs in coverage is measured elsewhere in this collection and the answer is that it costs a great deal: at a lag-one correlation of 0.8 and fifty observations, an interval built on an unadjusted variance covers at 47% against a nominal 95%. Twenty-two per cent of width is not the whole of that gap and it is a substantial part of it.
Which puts the field’s own subject in place. A four-point improvement in a forty-point error is about a one-point improvement in an interval’s width, on a shortfall of twenty-two. Choosing the window is a real improvement to a procedure whose main problem is elsewhere, and saying so is a service to anybody deciding where to spend their attention.
What the field ends up saying
Four numbers, in the order of how much they matter.
Forty per cent. The error in a block resample’s implied long-run variance at a hundred and twenty rows, at the best block length and the best window. That is the number a practitioner should carry, and nothing in this field or in the one before it was reporting it.
Thirty-one points of it are spread, which is larger than either bias and responds only to the amount of data.
Four points are the window, which is the entire subject of the comparison the field inherited, and it comes out in favour of the taper once each window is read at its own block length.
Fifteen times. The factor by which a hundred and twenty rows exaggerate the difference between two windows against what the algebra says, which is why the earlier comparison could not resolve what it was looking at.
The first of those is the useful one and it is the one that needed the cheapest instrument to find. Computing the implied variance rather than sampling a quantile made a sweep over three windows, eleven block lengths and four sample sizes affordable, and a sweep is what it takes to see that the thing being argued about is a twentieth of the thing.
What is claimed here, and what is not
This essay takes what the whole error of a block resample’s implied long-run variance is made of. The claims are that at a hundred and twenty rows the best error of the three windows is 42.7%, 40.7% and 40.4%; that the trapezoid’s at its own setting splits into 18.2 points of window attenuation, 8.2 of sample attenuation and 31.0 of spread; that the sample’s contribution is the same for every window and halves with each doubling of the sample, running 8.2, 4.1, 3.6 and 1.4 points; that the spread is the largest single component at every sample size measured; and that the window choice moves four points of the forty.
What stays out, and is named as a decision: a comparison against a sieve. The construction that models rather than weights is named as the thing that escapes the ceiling and is not measured on this quantity. It would be a fair comparison only against a correctly specified model, and what a misspecified one costs is a whole field elsewhere in this collection — so putting one number here would be putting the flattering half of a two-sided answer.
Also out: a recommendation. The field establishes that the trapezoid is better at each window’s own setting at every sample size measured, and that the advantage is small beside the error. Whether that is worth changing a practice for is a judgement about how much four points of a forty-point error matters, and it depends on what the long-run variance is being used for.
The boundary against the taper field is that it asks which window attenuates least and this one asks what the attenuation is a part of. Both are answered and only the second is a scale.
The four essays, and what each one moved
The field started from a named deferral and it is worth recording what each step of it actually changed, because three of the four changed a reading rather than a fact.
The gap a sample shows. The exact difference between two windows at the crossing is 0.28 points and a hundred and twenty rows report 4.31 — a factor of fifteen, from the attenuation in the autocovariances the window is applied to. Nothing about the windows changed; what changed is which number a finite-sample comparison has to resolve.
Bias is not the whole of it. Each window has a bias-optimal block length and an error-optimal one and they are never the same, so a comparison at a shared length compares one rule mis-tuned. Read at each window’s own setting the ordering reverses. Nothing about either window changed; what changed is where each was read.
Measuring a variance rather than a quantile. The implied variance is computable from the sample with no resampling in it and needs a third of the draws for the same statement. Nothing about the constructions changed; what changed is what the comparison was affordable enough to sweep.
And this one. The whole error is forty per cent and the window choice is four points of it. Nothing about anything changed; what changed is the scale everything else is read against.
Four essays, one fact, and four readings of it. That is what a field looks like when its subject turns out to be an instrument rather than a phenomenon.
The checks, and the refusals that make them mean something
Three claims are gated. The error is required to be at least as large as either component at every setting, because it is their sum in quadrature. The exact window contribution used in the decomposition is required to equal the taper field’s own computation to nine decimal places, at every block length and window, because a decomposition that used a different exact bias from the one the other field publishes would be a decomposition of a different quantity. And the implied variance is required to agree with realised resamples, which is the identity every number here rests on.
The refusal is the one this field ends on and it is the one that made the rest possible. An asymptotic gap quoted as what a finite sample shows is refused — 0.28 points against the 4.31 a hundred and twenty rows report at the same block length, with the mechanism attached: the autocovariances the window is applied to are themselves attenuated, and the window that discards the long lags loses less of them.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An ordering that depends on the rule — both name attenuation, bias-variance, block bootstrap, block length, closed form, dependence, long-run variance, mean squared error, monte carlo, resampling, tapering
- The length nobody has — both name bias-variance, block bootstrap, block length, closed form, dependence, long-run variance, mean squared error, monte carlo, resampling, sample autocovariance, tapering
- A length for each instrument — both name bias-variance, block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering
- The instrument and the reading — both name block bootstrap, block length, closed form, dependence, long-run variance, monte carlo, resampling, tapering
- An interval that carries its scale — both name block bootstrap, block length, closed form, dependence, long-run variance, monte carlo, resampling
- The ordering reverses again — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering
Named objects
A flat tag is an object no other essay names yet.
AttenuationBias-varianceBlock bootstrapBlock lengthClosed formDependenceEffective sample sizeLong-run varianceMean squared errorMonte CarloResamplingSample autocovarianceTaperingVariance