The block length read on a quantile

The reversal that was the instrument's

On an implied variance the rectangle wins at a protocol length and at the rule of thumb. On the 95% point a test reads, and on the coverage an interval delivers, the taper wins at all four rules.

Worth reading first: Where the bootstrap lies · The observations that repeat each other.

The field this one re-reads has a headline that is a reversal: the tapered block window beats the rectangular one at the best available block length and at one estimated from the sample’s persistence, and loses at a length written into a protocol and at the rule of thumb, at every sample size measured. The sentence it arrives at is not use the taper; it is estimate the block length, and then use the taper.

Every number in that is an error in an implied long-run variance.

The table

Everything below is four hundred samples of a hundred and twenty rows from a first-order autoregression at 0.7, with three hundred block resamples formed at each of seven candidate block lengths, and all three readings taken from the same sorted array of resampled means — which is the design the first essay of this field is about and is what makes a disagreement between the columns a disagreement about the reading.

Four rules, three readings of one set of resampled means, four hundred draws, signed so that a positive margin is the taper winning.

The reversal is a property of the instrumentThe margin between the two block windows under each of four rules, on three readings of the same resampled means, signed so that a positive bar is the tapered window winning. On the implied long-run variance the taper wins at the best available block length and at one estimated from the sample and loses at a length written into a protocol and at the rule of thumb — which is the reversal the earlier field's whole argument turns on, at 1.48 and 4.52 points. On the 95% point a test actually reads, the taper wins at all four, by 6.13 to 7.08 points. On the coverage the interval actually delivers, the taper wins at all four again, by 2.50 to 5.75 percentage points. Two of the four rules change sign between the first reading and the other two, and the two that change are exactly the two the earlier field's recommendation is about.the two windows tiethe implied variance · a protocol length-1.48the implied variance · the rule of thumb-4.52the implied variance · the plug-in1.11the implied variance · the best length2.12the 95% point · a protocol length6.13the 95% point · the rule of thumb3.16the 95% point · the plug-in7.08the 95% point · the best length6.09the coverage · a protocol length5.00the coverage · the rule of thumb2.50the coverage · the plug-in5.75the coverage · the best length4.00400 draws, 300 resamples each, pairedpositive is the taper winning
Fig. 1 The margin between the two windows under each rule, on all three readings. The slider changes how many draws the sweep takes.

On the implied variance the taper wins at the plug-in by 1.108 points and at the oracle by 2.123, and loses at a protocol length by 1.476 and at the rule of thumb by 4.525. That is the earlier field’s reversal, reproduced here with a resampled variance rather than a closed-form one, sign for sign.

On the 95% point the taper wins at all four: 6.130, 3.156, 7.077 and 6.091 points, at eight to eleven paired standard errors.

On the coverage the taper wins at all four again: 5.00, 2.50, 5.75 and 4.00 percentage points, at three to five standard errors.

Two of the four rules change sign, and they are the protocol length and the rule of thumb — the two the earlier field’s recommendation is about.

Reading the table’s three columns

The three readings are not three attempts at one number and the table should not be read as a majority verdict.

The implied variance is the estimand of the whole block-bootstrap literature and the quantity every rule in this collection is derived to estimate. It is right about its own target.

The 95% point is what a one-sided test compares its statistic to. Its target is the finite-sample point of the standardised mean, 3.8891 here, found by drawing forty thousand series from the law rather than by an asymptotic formula.

The coverage is what a two-sided interval delivers, against the 95% it promises.

So the two that agree are the two a reader acts on, and the one that disagrees is the intermediate. That ordering matters: the disagreement is not two against one, it is downstream against upstream.

What a flip proves, which is more than that the table moves

A margin changing sign between two readings of the same sorted array is a stronger statement than it looks, and it is worth saying exactly what it rules out.

Suppose the tapered window’s resample distribution were simply the rectangular one scaled by a constant c. Then the variance ratio would be c2c^2, the 95%-point ratio would be c, and the interval-width ratio would be c. All three are on the same side of one, so every reading would order the two windows identically and no flip could occur at any rule or any sample size.

So each flip is a proof that the two distributions differ in shape and not only in scale. It is the only such proof available here without measuring a shape statistic directly, and it comes for free out of a comparison that was set up to do something else.

The direction says which way. Where a margin is positive on the variance and negative on the 95% point, the tapered distribution is relatively short-tailed against the rectangular one — its spread is the larger of the two while its tail is not — so a test reading the tail gets a different ordering from an interval reading the spread.

And the two rules that do not flip are consistent with a pure scale difference. That is a weaker conclusion than it sounds and it is the honest one: not that their two distributions have the same shape, but that nothing in these three readings distinguishes them from a pair that does.

What that does to the recommendation

The earlier field’s sentence has two halves and they come apart.

Estimate the block length survives everything. It survives on the implied variance, where the plug-in beats the two fixed rules; it survives on the quantile, where the plug-in’s margin is the largest of the four; and it survives as the larger finding, since what a rule gives up against the oracle is several times what the window choice buys, on every reading here as on the one it was established on.

And then use the taper was a conditional, and the condition is gone. The taper wins at every rule on both instruments a practitioner reads. A practitioner who writes a block length into a protocol and uses the taper is doing the right thing on a critical value and on an interval, and the wrong thing only on a quantity nobody reports.

So the correction is narrow and it is to the half of the sentence that was doing the qualifying.

What “the taper” and “the rectangle” are

Two windows, and the difference between them is one line.

The rectangle is the ordinary moving block bootstrap: fixed-length blocks of the residual series laid end to end, each starting at a uniform position, every observation inside a block counted at full weight.

The taper is the same construction with each block weighted down at its ends before it is laid — the trapezoid that rises over the first 43% of the block, holds, and falls again. The field that measured what that buys shows the weighting removing the join discontinuity, which is what buys an order of convergence, and costing attenuation at short lags, which is what it pays with.

Everything in the table below is those two, at the same nominal block length, on the same residual series, with the same seed.

Why the two fixed rules are the ones that move

The previous essay has the mechanism and this is what it predicts.

The two instruments want different block lengths: the implied variance’s oracle is at 18.92 and 21.82 for the two windows, and the quantile’s is at 16.05 and 17.74. The quantile wants a shorter block, by about fifteen per cent.

The two fixed rules pick 8 and 4 — far short of either target, and less far short of the quantile’s. So they are the two rules whose length is most wrong for the variance and least wrong for the quantile, and moving between the two instruments moves them furthest.

And the windows do not pay the same price for being short. A tapered block weights its ends down, so at a nominal length of 8 its effective length is shorter than a rectangle’s 8 — which is exactly what the earlier field says costs it there: a tapered window at a quarter of the right length has thrown away most of what it was weighting. That is true and it is a statement about the variance. On the quantile, having fewer effective observations per block means having more effective blocks, and more blocks is what a tail needs.

The same property makes the taper worse at short lengths on one reading and better on the other. That is the whole reversal, and it is not a coincidence of these numbers: it follows from what the window does.

The two rules that do not move

The plug-in and the oracle keep their sign, and that is worth as much as the two that do not.

Both pick a length near either oracle — the plug-in returns 12.59 and 13.89 for the two windows, and the oracles sit at 16 to 22 — so both are in the range where the two instruments’ targets nearly coincide. Their margins grow rather than reverse: the plug-in’s goes from 1.108 points on the variance to 7.077 on the quantile, and the oracle’s from 2.123 to 6.091.

So the taper’s advantage is larger on the quantile everywhere, and at the two short rules it is large enough to cross zero. The instrument does not flip the ordering; it shifts every margin in the taper’s favour, and two of the four were close enough to zero to change sign. That is a milder description of the same table and it is the accurate one.

It also says which way a different setting would go. A rule that picked a length nearer the oracle would move less; a rule shorter than four would move more.

What the two fixed rules are doing wrong

A protocol length of eight and a rule of thumb of four are not arbitrary and it is worth saying what each is for, because both are defensible and both are badly short here.

Eight is a convention. It is a round number in the range applied work uses, and it is what this collection’s own figures have used as a default since the block bootstrap first appeared in it. Its virtue is that it is declared in advance, which is a real virtue — a block length chosen after seeing the answer is a searched parameter and this collection has priced what a searched parameter costs.

Four is n1/3n^{1/3}, which is the rate the mean-squared-error algebra gives for a block length and is right about the rate. What it omits is the constant, which depends on the persistence — and at a first-order autoregression of 0.7 the constant is large.

So both rules are short for a reason rather than by carelessness, and both are short by a factor of two to five against either oracle. What the instrument decides is not whether they are short but what being short costs, and the answer is different for the two windows on the two readings.

How large the flips are

The margins that change sign are not small on either side.

At the rule of thumb the taper loses by 4.525 points on the implied variance at 70.8 paired standard errors, and wins by 3.156 on the quantile at 8.1. At a protocol length it loses by 1.476 at 8.5 and wins by 6.130 at 11.3.

Neither reading is marginal. Both orderings are established at many standard errors and they point opposite ways, which is the strongest form this kind of finding takes: it is not that one instrument is undecided and the other has an opinion.

The levels the margins sit inside are large — the errors run from 36.93% to 68.17% across the eight cells — so a margin of a few points is a few points of a fifty per cent error. That is the same proportion the earlier field reports and the same caveat applies: the comparison is between two ways of being badly wrong.

A quantile is the dearer reading, everywhere. The error each rule and window delivers on the two error readings, over 400 draws. The lower pair of lines is the implied long-run variance — the instrument the earlier field uses — and the upper pair is the 95% point of the standardised resampled mean, read against the finite-sample truth of 3.889 found by simulating the law directly. The quantile costs more at every one of the eight cells: at the plug-in rule it is 59.1% against 45.3% for the taper. That is not a defect in the bootstrap; a quantile is a statement about the shape of a distribution as well as its scale, and a fixed number of resamples estimates a tail worse than a variance. What matters for the comparison is that the two orderings between the windows are not the same, which the margins figure is about.
Fig. 2 The error each rule and window delivers on the two readings, over four hundred draws: the implied long-run variance below and the 95% point above, the latter read against the finite-sample truth of 3.889. The quantile costs more in all eight cells — 59.1% against the taper’s 45.3% at the plug-in rule — because a fixed number of resamples estimates a tail worse than a variance.

What a practitioner should take from the pair of fields

Three sentences, in the order they matter.

Estimate the block length rather than writing one into a protocol. That is the larger finding by a factor of three and it is unchanged: what the best feasible rule gives up against the best available length is several times what the window choice buys at that length, on every instrument here.

Use the tapered window, whichever length the rule returned. That is the change. The earlier field’s condition — only once the length has been estimated — is an artefact of the reading it was established on, and on both readings a practitioner actually uses the taper wins at every rule including the two the condition was about.

And a bootstrap’s ordering should not be read off a variance when a quantile is what will be quoted. That is the transportable half. The two instruments are two readings of one computation, they cost nothing extra to take together, and where they disagree they disagree at many standard errors on both sides.

Two instruments, two block lengths. The block length that would actually have been best on each draw, for each of the two error readings, averaged over 400 samples of 120 rows. For the rectangular window the implied long-run variance wants 18.92 and the 95% point wants 16.05; for the tapered window, 21.82 against 17.74. The quantile wants a shorter block under both windows — a ratio of 0.848 and 0.813. That is the mechanism the whole field turns on: a rule for choosing a block length is a way of guessing a target, and the two instruments do not have the same target. A rule tuned to one is systematically long for the other, and the two windows do not pay the same price for being long.
Fig. 3 The block length that would actually have been best on each draw, by instrument. The rectangle’s implied variance wants 18.92 and its 95% point 16.05; the taper’s, 21.82 against 17.74. A rule is a way of guessing a target, and the two instruments do not have the same one.

The shape of the correction, and where it belongs

It is worth being precise about what kind of error this is, because it is a common one and it is not a mistake in any of the arithmetic.

Nothing the earlier field computes is wrong. Its implied variances are right, its rules are right, its reversal is real, and this field reproduces it. What is wrong is one step of inference: from an ordering on a variance to a recommendation about a practice. A practice reads a quantile.

That step is invisible in the write-up because the instrument is never named as an instrument — it is the error, and an error is what a comparison is about. The moment it is named as one reading among several, the question of whether the ordering transports becomes askable, and it turns out not to.

The same shape has come up twice more in this collection. A charge derived for a covariance’s dimension is levied on a likelihood and read for a coefficient error, which are two quantities with different argmaxes. A tuning parameter’s regret is a rate times a size reported as one number. In each case the reported quantity is an intermediate and the reader’s is downstream, and in each case the two do not move together.

The tell is a comparison whose units are not the units anybody quotes, and it is worth checking for by default rather than by accident.

What a reader of the earlier field should now do

The earlier field is not withdrawn and its table is not wrong, so the practical question is which of its numbers to keep.

Keep the lengths. 24.57 for the oracle, 14.36 for the plug-in, 8 and 4 for the fixed rules, and the spreads beside them. Those are facts about the rules and do not depend on the reading.

Keep the cost of choosing. 7.26 points of shortfall for the best feasible rule against 2.12 for the window choice at the oracle, a factor of 3.4. That comparison is between two quantities on the same instrument, so changing the instrument moves both and the ratio survives — as it does here.

And replace the conditional. The taper wins once the length has been estimated becomes the taper wins, on both readings a practitioner uses, at every rule.

One reading the table forbids

A reader who has got this far may be tempted by a tidier story than the table supports: that the implied variance is simply the wrong instrument and should be dropped.

The table forbids it. On the two rules whose ordering does not flip — the plug-in and the oracle — all three readings agree, so the variance is right about them; and the variance is what the whole comparison was affordable through, since separating two windows on a quantile takes about two orders of magnitude more draws. The field that measured that exchange rate chose it for a good reason and the reason still holds.

What the table forbids is quoting a variance’s ordering for a quantile’s decision, which is a much narrower prohibition and costs nothing: taking both readings off one bootstrap is one extra sort.

What is not established

That the taper is better on a quantile in general. Everything here is one persistence, one sample size and one law. The earlier field sweeps four sample sizes on its own instrument and finds no crossing; nothing here does that on the quantile, and a sweep is the obvious next thing.

That the implied variance is the wrong instrument to sweep with. It is a hundred times cheaper to separate two windows with, which is what made four fields of sweeping possible, and it reproduces here to the sign. What this field establishes is that an ordering read off it should be checked on the reading it is about to be quoted for — not that it should not be read.

And that coverage and the quantile will always agree. They agree here, at every rule, and they are sensitive to different features of the same distribution: the fourth essay of this field is where that is measured, and where the levels turn out to matter more than the ordering.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A taper and a critical value — both name block bootstrap, critical value, dependence, long-run variance, reference distribution, resampling, stationary bootstrap, tapering
  • An interval that carries its scale — both name block bootstrap, block length, coverage, dependence, long-run variance, monte carlo, reference distribution, resampling
  • The length nobody has — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering, tuning parameter
  • What a multiplier cannot keep — both name block bootstrap, critical value, dependence, estimation error, monte carlo, persistence, reference distribution, resampling
  • What choosing the length costs — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering, tuning parameter
  • Errors generated from a fitted model — both name block bootstrap, critical value, dependence, estimation error, persistence, reference distribution, resampling

Named objects

A flat tag is an object no other essay names yet.

Block bootstrapBlock lengthCoverageCritical valueDependenceEstimation errorLong-run varianceMonte CarloPersistenceReference distributionResamplingRobustnessStationary bootstrapTaperingTuning parameter