A block, weighted inside itself

A taper and a critical value

Two constructions whose tapers visibly differ give the same critical value, and two that share a taper exactly do not. Adding a construction whose taper is a decision rather than an accident says which half of that is true.

Worth reading first: The experiments that could have happened · Where the bootstrap lies.

Five resamplings of the same residuals were measured against the same truth, and the result was flatly negative: an expected autocovariance is not a reference distribution. The fixed block and the stationary bootstrap, whose tapers visibly differ, came out at 10.5% and 10.3% short of the statistic’s own 95% point — indistinguishable. The blocked multiplier, which shares the fixed block’s triangle exactly, came out at 18.1% short, nearly twice as far.

That comparison had a gap in it. All five of those tapers were inherited: nobody chose them, they fell out of what each construction does, and no two of them differ by a window and nothing else. A construction whose taper is a decision is what the comparison was missing, and a block weighted inside itself is one.

The table with the fourth taper in it

How far each reference distribution's 95% point falls shortSeven constructions on rows that repeat each other, at a block length of 5 and 200 draws, against the statistic's own 95% point of 3.0224 computed from three thousand draws of the same world. Reading down: a multiplier on every row keeps no dependence at all and is 44% short; a multiplier shared along a block keeps the triangle; a fixed block keeps the same triangle and is 8% closer, which is the pair that says a taper is not what decides this; the stationary bootstrap; the two tapered blocks, both further short than the untapered one at this block length; and errors generated from a fitted model, which is the only construction here not bounded by what the residuals report.how far the resampled 95% point falls short of the truth — nearer zero is betterrows that repeat each other · ℓ = 5 · 200 drawsa multiplier on every row44.0%a multiplier shared along a block18.1%a fixed block, no taper10.5%the stationary bootstrap10.3%a fixed block, trapezoid12.6%a fixed block, raised cosine16.2%errors from a fitted model3.5%200 draws against a truth from 3000the taper is not what decides it
Fig. 1 Seven constructions on the same residuals, with the world on the slider. The measure is how far each reference distribution’s 95% point falls short of the statistic’s own, computed from three thousand draws of the same world.

On errors that repeat each other, against a truth of 3.0224: a multiplier on every row is 44.0% short, because it keeps no dependence at all. A fixed block is 10.5% short and the stationary bootstrap 10.3%. The same fixed block with a trapezoid inside it is 12.6% short, and with a raised cosine 16.2%. Errors generated from a fitted model overshoot by 3.5%.

So the tapered blocks are further from the truth than the untapered one, and further in the order their windows predict.

The seven, in the units of the critical value itself

Percentages of a truth are easy to compare and hard to feel, so the same seven readings against a truth of 3.0224 are worth writing as the numbers a test would actually use:

1.69 for a multiplier on every row, 2.48 for a blocked multiplier, 2.53 with a raised cosine, 2.64 with a trapezoid, 2.71 for a fixed block, 2.71 for the stationary bootstrap, and 3.13 for errors generated from a fitted model.

Two things stand out that the percentage column does not make obvious.

The sieve is the only one on the far side of the truth, and at 3.5% it is three times closer to it than the best of the others. Overshooting is a different kind of error from falling short — it produces a conservative test rather than an anti-conservative one — so the sign is worth as much as the size.

And the taper is not a small dial here. Every blocked construction is short by at least 0.31 of critical value, which is the floor they share; the four of them then spread over a further 0.24, which is three quarters of that floor. So choosing a taper moves the answer nearly as far as the whole common deficit does, and it moves it in the wrong direction — every taper considered here makes the shortfall worse than leaving the block untapered.

That is the awkward half of the finding. The taper is a decision with a real effect, and the best setting of it available is not to make one.

Inside one family, the taper predicts it exactly

That ordering is not a coincidence and it was available in advance. The implied long-run variance bias at a block length of five, computed in closed form with no sampling in it, is 37.4% low for the geometric taper, 45.7% for the rectangle, 51.9% for the trapezoid and 56.9% for the raised cosine.

The measured shortfalls run in exactly that order: 10.3%, 10.5%, 12.6%, 16.2%.

The taper predicts a critical value inside one family and not across families. Each block construction's implied long-run variance bias at ℓ = 5, computed in closed form, against how far its reference distribution's 95% point falls short of the statistic's own. The four constructions that lay blocks end to end sit in the right order — the stationary bootstrap least biased and least short, then the rectangle, then the trapezoid, then the raised cosine — and the paired differences resolve: a rectangular block against a trapezoid is 0.0626 at 2.0 standard errors, and a trapezoid against a raised cosine 0.1099 at 4.1. The three marks on the axis at the left belong to constructions with no window at all, and one of them shares the rectangle's taper exactly while sitting 8% further from the truth.
Fig. 2 The closed-form bias each window implies against the counted shortfall of its reference distribution. The four constructions that lay blocks end to end sit in order; the three with no window at all are on the axis at the left.

And the differences resolve, because every construction is run on the same draws and the comparisons are paired. A rectangular block against a trapezoid is 0.0626 of critical value at 2.0 paired standard errors, and a trapezoid against a raised cosine 0.1099 at 4.1. The rectangle against the stationary bootstrap is 0.0062 at 0.2, which is nothing — and the closed form says it should be small, since those two are the closest pair in predicted bias.

Inside the family of constructions that lay blocks end to end, the taper predicts the critical value. That is a positive result where the earlier reading was negative, and it does not overturn the earlier reading.

Across families it still says nothing

The three constructions with no window sit off the line entirely, and one of them is the case the earlier finding was built on. A multiplier shared along a block has the fixed block’s triangle exactly — the same attenuation at every lag, proved rather than fitted — and its reference distribution is 18.1% short against the fixed block’s 10.5%, a paired difference of 0.2289 at 6.1 standard errors.

Two constructions with the same taper, six standard errors apart. Two constructions with different tapers, within a fifth of a standard error of each other. Both facts are in the same table, and the resolution is that the taper is a statement about lags and the difference between these families is a statement about rows: a multiplier leaves each residual on the row whose variance produced it, and a block moves it.

Which of the two properties matters depends on the world, and the second world in the table is where it shows. When the error variance depends on the design as well as repeating in time, the blocked multiplier is 11.7% short and every block construction is worse — the fixed block 17.1%, the stationary bootstrap 17.5%, the trapezoid 18.7%, the raised cosine 21.7%. The multiplier moves from worst to best, at 2.5 paired standard errors over the fixed block, while the ordering inside the block family survives unchanged: 17.1, 18.7, 21.7, with the trapezoid against the raised cosine at 2.7 standard errors.

A window orders the constructions that have one. What separates the families is what each does to a row.

The one construction that overshoots

Six of the seven rows are short and one is not. Errors generated from a fitted autoregression give a reference distribution whose 95% point is 3.5% above the truth in the first world, and it is the only construction here that ever lands on the far side.

The reason is the one it was introduced for. Every resampling that reuses the residuals is bounded by what the residuals carry, and the residuals are less persistent than the errors before any resampling begins — so no block length and no window reaches the truth, and the whole table is an argument about how much of a fixed shortfall each construction recovers. A fitted model is not bounded that way: its autocovariance is an infinite decaying sequence rather than a truncated one, and it can carry more dependence at long lags than the sample it was fitted to reports.

It pays for that in the second world, where it is 8.1% short rather than 3.5% over — still the closest of the seven, and now on the same side as everything else. Generating errors from one innovation variance is a misspecification when the variance depends on the design, and a misspecification is a different kind of risk from a ceiling rather than a smaller one.

The two readings together are the useful part. The three block constructions differ from each other by their windows, in an order a closed form predicts. The construction that beats all of them differs by not being a resampling of the residuals at all, and no window would have got there.

What it would take for the taper to win here

The closed forms say the trapezoid overtakes the rectangle at a block length of twenty. It is worth asking what that crossing would look like in this table, because the answer decides whether measuring it is worth anything.

At ℓ = 20 the two implied biases are 13.4% and 13.7% — a gap of 0.3 points, where at ℓ = 5 the gap is 6.2 points. The table above maps a bias gap onto a critical-value gap: 6.2 points of bias moved the 95% point by 0.0626, and the five points between the trapezoid and the raised cosine moved it by 0.1099, so a point of bias is worth somewhere between 0.010 and 0.022 of critical value here.

A 0.3-point gap is therefore worth about 0.005 of critical value, against a paired standard error of 0.031 at two hundred draws. Resolving it at two standard errors would need about twenty thousand draws, which is a hundred times the work of this table for a difference that is one part in six hundred of the quantity being estimated.

So the crossing is real, computable, and not measurable in this design. That is worth recording as a limit of the measurement rather than as a result: the honest statement is that the tapered block is worse here by an amount that is easy to see, and better at longer blocks by an amount that is not.

Where the taper starts paying, and it is not here. The same bias at the block lengths a sample of a hundred and twenty rows can actually support. The ordering is upside down: at ℓ = 8 the rectangle is at -32.3% and the trapezoid at -37.9%, so the window that is better by an order of magnitude asymptotically is worse by six points here. They cross at ℓ = 20. The stationary bootstrap's geometric taper, which has no asymptotic advantage at all, has the smallest bias of the five at every block length up to 32 — because its runs have no hard cut-off, so it keeps something at every lag rather than nothing past ℓ.
Fig. 3 The bias at the block lengths a hundred and twenty rows can support. At ℓ = 8 the rectangle is 32.3% low and the trapezoid 37.9%, so the window that is better by an order of magnitude asymptotically is worse by six points here; they cross at ℓ = 20. The geometric taper, which has no asymptotic advantage at all, is the smallest of the five up to ℓ = 32.

A 95% point that is ten per cent short does not reject ten per cent too often

There is a reading of the table that the table does not support, and it is worth heading off because it is the natural one.

The rejection rates say something different from the shortfalls. Against a nominal 5%, with a standard error of 1.5 points, the fixed block rejects on 3.0% of draws, the stationary bootstrap 4.0%, the trapezoid 3.5% and the raised cosine 4.0% — every one of them at or below nominal, on reference distributions whose 95% points are between 10% and 16% short. Only the multiplier that keeps nothing behaves as the shortfall suggests, rejecting on 25.0% of draws.

The reason is that the shortfall is a marginal statement and a rejection rate is a conditional one. Each draw’s reference distribution is computed from that draw’s own residuals, so a draw whose residuals happen to be large has both a large statistic and a large critical value. Measured across the draws, the correlation between the two is 0.63 in the first world and 0.79 in the second.

A critical value that tracks the statistic that closely is not usefully described by how short it is on average. The average is the right summary for comparing constructions with each other — it is precise, it is paired, and it is what the closed forms predict — and it is the wrong summary for predicting an error rate.

That distinction is the same one this collection keeps meeting from other directions: a procedure’s advertised property has to be counted rather than reasoned about, and two quantities that both look like “how wrong is this” can move in different directions.

What this means for the construction

The tapered block was introduced into this collection because the literature uses it and this collection had not measured it. The measurement says: at the block lengths a sample of this size can support, it is worse than the plain block, in the amount and the order its own window predicts.

That is not a criticism of the construction. Its advantage is an asymptotic one, and the block lengths where it arrives are longer than the ones a hundred and twenty rows leave room for. What the measurement rules out is adopting it here on the strength of that argument, which is what the argument invites.

And there is a second reading, which is more useful. The four block constructions differ by their windows and their critical values differ in the order the windows predict — so a window is a dial that does something predictable, unlike almost everything else in this comparison. A practitioner who wants a slightly shorter reference distribution has a way to get one and a closed form saying how much. That is a better outcome than the tapered block being uniformly better, and it is not the outcome the construction is usually recommended for.

What each block keeps, lag by lag. The attenuation each construction applies at each lag, in a block of 20. A rectangular block gives exactly the triangle 1 − k/ℓ, which is not a fact about blocks: it is the self-convolution of a rectangle, and every other window has one of its own. The trapezoid keeps 0.987 at the first lag against the rectangle's 0.950 and 0.263 at the half-block against 0.500, so it holds the short lags almost intact and gives up the long ones faster. The dashed line is the stationary bootstrap's geometric taper, which comes from randomising the block length rather than from weighting inside it, and is the only one of the five that is never exactly zero.
Fig. 4 What each construction keeps at each lag in a block of twenty. A rectangular block gives exactly the triangle 1 − k/ℓ; the trapezoid keeps 0.987 at the first lag against 0.950 and 0.263 at the half-block against 0.500. Two constructions with visibly different attenuations are the ones the table finds indistinguishable.

What a practitioner would actually change

Read as advice rather than as a measurement, the table says three things and they are not the ones the constructions are usually chosen for.

The window is the smallest dial here. Between the best and worst window at ℓ = 5 there are six points of critical value; between a construction that moves rows and one that does not there are eight in one world and five the other way in the other; and between any of them and a multiplier that keeps no dependence at all there are thirty-three. Anybody choosing a resampling should settle what it does to a row before thinking about what it does to a lag.

The choice depends on which defect is present, and both defects are usually present. The blocked multiplier wins where the variance depends on the design and loses where the dependence is the only problem; the block constructions do the reverse. There is no row of this table that is best in both worlds except the one that is not a resampling of the residuals — which is an argument for fitting a model of the errors rather than for any window.

And the reference distribution should be judged by what it does, not by what it keeps. Every number in the first half of this essay is an expectation and every number in the second half is a quantile or a rate, and the three readings disagree with each other in ways that are not small: a construction 10.5% short in the mean rejects below nominal, and a construction whose taper is identical to another’s is six standard errors away from it. The habit this collection runs on — count the advertised property rather than the property that implies it — is the whole of what separates those readings.

What is claimed here, and what is not

This essay takes whether a resampling’s taper predicts its reference distribution. The claims are that with a fourth taper in the table the four block constructions order exactly as their closed-form biases do — 10.3%, 10.5%, 12.6%, 16.2% against predicted biases of 37.4%, 45.7%, 51.9%, 56.9% — with the adjacent pairs resolved at 2.0 and 4.1 paired standard errors; that two constructions sharing a taper exactly still differ by 6.1 standard errors, so nothing crosses between families; that when the variance depends on the design the row-preserving multiplier moves from worst to best while the ordering inside the block family is unchanged; and that a reference distribution 10% short in the mean rejects at or below nominal, because its critical value tracks the statistic at a correlation of 0.63.

What stays out and is named as a decision: one block length. Everything above is at ℓ = 5, which is where the constructions in this collection have always been compared, and the whole of the tapered block’s case is that it wins at larger ℓ. Measuring the table again at ℓ = 20 would be a fair test of that case and it is not made here, because the sample has a hundred and twenty rows and twenty-row blocks leave six of them.

Also out: one statistic. The reading is a studentised maximum over an open search across a table of candidates, which is the hardest thing in this collection for a resampling to get right and is not representative of a resampling’s job in general.

The boundary against the essay that derives the windows is that it computes what each construction keeps and this one counts what each construction’s reference distribution does. The first is an expectation and the second is a quantile, and the point of running both is that a quantile is not a function of an expectation.

The checks, and the refusal

Two claims are gated. Two constructions with visibly different tapers are required to agree within two and a half standard errors and two constructions sharing a taper exactly to differ by more than two, which is the earlier finding stated as a test rather than as a remark. And inside the block family, more attenuation is required to produce a shorter critical value, in the order the closed form puts them.

The refusal is the construction this essay was written to add. A tapered block offered as an improvement to a reference distribution is rejected at these block lengths: it is 2.1 points further short than the untapered block in the first world and 1.6 in the second, in the direction its own window predicts, and the argument for adopting it is about a regime this sample size does not reach.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A length for each instrument — both name block bootstrap, critical value, dependence, long-run variance, reference distribution, resampling, stationary bootstrap, tapering
  • How long a block a multiplier shares — both name block bootstrap, critical value, error rate, heteroskedasticity, reference distribution, resampling, residual, wild bootstrap
  • The reversal that was the instrument's — both name block bootstrap, critical value, dependence, long-run variance, reference distribution, resampling, stationary bootstrap, tapering
  • What the interval covers — both name block bootstrap, critical value, dependence, long-run variance, reference distribution, resampling, stationary bootstrap, tapering
  • Residuals that keep their own variance — both name block bootstrap, critical value, error rate, heteroskedasticity, reference distribution, residual, wild bootstrap
  • The ordering reverses again — both name block bootstrap, critical value, dependence, long-run variance, reference distribution, resampling, tapering

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionBiasBlock bootstrapCritical valueDependenceError rateHeteroskedasticityLong-run varianceReference distributionResamplingResidualStationary bootstrapStatistical powerTaperingWild bootstrap