The ordering reverses again
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
The field that read one block bootstrap three ways reports a reversal: two of its four rules rank the two block windows one way on an implied long-run variance and the other way on the 95% point a test reads. Two of four, between two readings of one resampling.
Change instead what is done with the resamples — build a studentised interval rather than a percentile one — and the reversal is four of four.
Why this is not the same reversal
Before the numbers, one distinction, because the word has been used twice in this collection already and means something different here.
An earlier field found the ordering reversing between two instruments: an implied long-run variance and a 95% point are different quantities, measured against different truths, and it is not paradoxical that they rank two windows differently. Its argument was that one of the two is what a practitioner reads.
This field changes neither the quantity nor the truth. Both intervals are two-sided ninety-five per cent intervals for the same mean, measured against the same promise, built from the same resamples of the same draws. The only difference is whether each resample is divided by its own scale before its quantile is taken.
Two constructions of one interval, and they order the windows oppositely.
The margins
The margin between the two windows in points of coverage, signed so that positive is the taper winning, at each of four rules and on each of three intervals built from the same resamples:
| rule | percentile | studentised | normal |
|---|---|---|---|
| a protocol length | +5.33 | −4.33 | 0.00 |
| the rule of thumb | +1.67 | −7.00 | 0.00 |
| the plug-in | +5.00 | −4.33 | −0.33 |
| the best length | +4.00 | −2.67 | +0.33 |
On the percentile interval the taper wins at all four. On the studentised interval the rectangle wins at all four.
What the third column says
The normal interval is what makes this readable rather than merely surprising.
It uses no resampling at all — x̄ ± 1.96 σ̂/√n on the block-means variance — so the only way a window can affect it is through the block length the rule picked. Its margins are 0.00, 0.00, −0.33 and +0.33 points: the two windows are indistinguishable.
So the difference between the windows is not in the data, not in the block lengths the rules choose, and not in the scale. It is manufactured entirely by what is done with the resamples, and two different things done with the same resamples produce opposite orderings.
What is held fixed
Everything about the resampling except what is done with the resamples afterwards, which is what makes the comparison attributable.
The draws are the same: three hundred series of a hundred and twenty rows from a first-order autoregression at a persistence of 0.7, the world the whole line of fields works in.
The resamples are the same: two hundred and fifty per draw per block length per window, formed once and read three ways. A difference between the three intervals cannot be a difference in the seed, the replication count or the resampling scheme.
The rules are the same four: a protocol length of eight, the rule of thumb’s n^(1/3), a plug-in from the sample’s own lag-one correlation, and an oracle defined on the quantile exactly as the earlier field defines it.
And the grid is the same seven lengths. What changes is one thing: whether each resample is standardised by its own scale before its quantile is taken.
Why studentising helps one window and hurts the other
The mechanism is available and it follows from what the two windows do.
A rectangular moving-block bootstrap lays fixed-length blocks end to end. A tapered one weights each block down at its ends before laying it, which removes the join discontinuity and buys a better order of convergence — the taper field’s own measurement.
The weighting attenuates the variance each block contributes. So at the same nominal block length a tapered resample’s block means are more variable relative to their own mean than a rectangular one’s, and the block-means variance that studentises them is noisier.
A noisier denominator makes the resampled t heavier-tailed, which widens the interval; and it makes the t less pivotal, which is exactly the property studentising was supposed to buy. The rectangle, whose blocks are unweighted and whose block means are the cleanest thing available, gets more of what studentising promises.
The widths agree with that reading: at the two long-block rules the taper’s studentised interval is ×5.042 as wide as its percentile one and the rectangle’s is ×4.369, and at every rule the taper’s ratio is the larger of the two — ×1.204 against ×1.171 at the protocol length, ×1.140 against ×1.070 at the rule of thumb, ×1.428 against ×1.325 at the plug-in.
Four rules, four times the taper’s ratio larger. A wider interval covers better, other things equal, so the taper’s extra width should have pushed it towards better coverage rather than away — and it covers worse at every rule anyway. That is what says the noisier denominator is not merely widening the interval but displacing it, which is the failure a non-pivotal quantity produces.
The unresampled interval sits at the midpoint
The third column is used to show that the block lengths are not doing the work, and averaging the three columns says something more.
The percentile margins average +4.00 points and the studentised ones −4.58. Their midpoint is −0.29, and the normal interval’s four margins average 0.00.
So the two resampled constructions displace the window ordering by about four and a half points in opposite directions from a baseline that uses no resamples at all, and the baseline is where the midpoint of the two lands. Neither construction is the neutral one and neither is the distorted one — they are symmetric departures, and the interval that never resamples is the fixed point between them.
That is a stronger reading than “the difference is manufactured by what is done with the resamples”. It says the two things done with them are equal and opposite, which no argument about either construction alone would predict, and it puts a number on how much each one moves: about four and a half points of window margin, from a resampling that both share.
Per rule the mirroring is close but not exact. The two margins sum to +1.00, −5.33, +0.67 and +1.33 at the four rules — three of the four are inside a point and a third of cancelling exactly, and the rule of thumb is the one that does not.
The extra width does not explain the loss
The essay argues that the taper’s studentised interval covers worse despite being wider, and ranking the four cells makes that quantitative.
The taper’s studentised width, as a multiple of the rectangle’s, is 1.028, 1.065, 1.078 and 1.154 at the four rules. Its coverage margins are −4.33, −7.00, −4.33 and −2.67.
Rank the four by extra width and by size of coverage loss and the two orderings run against each other — a rank correlation of about −0.55 on four points. The cell where the taper’s interval is fifteen per cent wider than the rectangle’s is the cell where it loses least, and the cell where it is only three per cent wider is where it loses most.
Four points cannot establish a relationship and they can rule one out. If the coverage loss were the width being spent badly, the two would rank together, and they rank oppositely — so whatever the extra width is doing, it is not what costs the taper its coverage. The displacement of the resampled t, which the essay attributes the failure to, is a separate quantity that the width does not track.
The cells behind the margins
A margin is a difference of two coverages, and the eight cells underneath are worth reading because they say which window moved.
On the percentile interval: the rectangle covers 83.7%, 80.0%, 84.7% and 81.3% at the four rules; the taper covers 89.0%, 81.7%, 89.7% and 85.3%. The taper is ahead everywhere and by most at the protocol length and the plug-in.
On the studentised interval: the rectangle covers 88.7%, 82.7%, 92.3% and 85.0%; the taper covers 84.3%, 75.7%, 88.0% and 82.3%.
So both windows moved, in opposite directions. The rectangle gained 5.00, 2.67, 7.67 and 3.67 points; the taper lost 4.67, 6.00, 1.67 and 3.00. Neither stayed put and neither’s movement alone accounts for the reversal.
That is worth knowing because a reversal produced by one window moving a long way would be a statement about that window. This one is two windows moving a few points each in opposite directions across a gap that was four or five points wide, and the gap is narrow enough for that to be enough.
Which reading is right
None of them, and that is the finding rather than an evasion.
Not one of the twenty-four cells reaches ninety-five per cent. The studentised column runs from 75.7% to 92.3%, the percentile column from 80.0% to 89.7%, and the normal column from 81.7% to 89.0%. So the ordering between the two windows — whichever way it goes, on whichever interval — is a choice inside a shortfall that no cell escapes.
The margins are two to seven points. The shortfalls are three to nineteen.
That is the same shape the earlier field arrives at, and it is the reading that should survive this field: the choice between two block windows is a small choice inside a large failure, and which window wins depends on which of several equally defensible intervals is built from the resamples.
How much of this is measurement
Three hundred draws, and coverage is a proportion, so the standard error on any single cell is about 2.2 points. A margin is a difference of two coverages measured on the same draws, so it is paired and its standard error is smaller than that — but not by much, because the two windows’ intervals do not cover the same draws nearly as often as two intervals from one window do.
The paired t-statistics for the studentised-against-percentile comparison, cell by cell, are 2.82, 1.89, 4.57, 2.13, −2.89, −3.93, −1.09 and −1.62. Five of the eight separate at two standard errors; three do not.
So no single cell in this table is decisive, and the claim does not rest on one. What it rests on is that all four rules point the same way on each interval, and the two intervals point opposite ways — eight readings, four positive under one construction and four negative under the other, with the two windows’ individual movements all in the direction their construction predicts.
Eight coin flips landing that way is one in a hundred and twenty-eight. The mechanism in the section above is what makes it more than a count.
What that does to a recommendation
The earlier fields end with a recommendation, and it is worth tracing what happens to it.
The field that re-ran the window comparison on feasible rules recommends estimate the block length, and then use the taper — on the grounds that the taper wins at the plug-in and at the oracle and loses at the two fixed rules, on an implied long-run variance.
The field that read the same resampling three ways finds the taper winning at all four rules on the 95% point and on the coverage, so the recommendation survives and gets stronger: whatever a practitioner reads, use the taper.
This field finds the taper losing at all four rules once the interval carries its own scale.
So the recommendation now depends on which interval a practitioner builds. Percentile: the taper. Studentised: the rectangle. Normal: it does not matter. Three defensible constructions, three answers, one set of resamples.
What stays out of this field
Four things, and each of them would change what the reversal means.
A second sample size. Everything is a hundred and twenty rows. The width penalty this field measures is the block count, and the block count is n/ℓ — so a longer sample at the same block lengths would have more blocks, a better scale, and plausibly a studentised interval that behaves the way the textbook says. Whether the reversal survives that is a guess, and it is the guess most worth testing.
A second persistence. A first-order autoregression at 0.7 throughout, inherited from the fields this one re-reads. The block lengths the rules pick are functions of the persistence, so a weaker dependence would push every rule towards shorter blocks and more of them.
A bias-corrected interval. The standard family has a third member — a percentile interval corrected for bias and acceleration — which addresses the skewness a percentile interval inherits without dividing by an estimated scale. That is the construction this field’s own findings most obviously point at, and it is not run.
And a one-sided reading. Everything here is a two-sided interval. A one-sided test reads one tail, which is a different functional of the same resampled distribution, and the field that separated a critical value from an interval is explicit that the two are sensitive to different features of it.
A stronger reversal than the one it follows
It is worth being precise about how this differs from the reversal it is modelled on, because they are two different kinds of thing.
The earlier field’s reversal is between two quantities. An implied long-run variance and a 95% point are different numbers with different truths, and the field’s whole argument is that the second is the one anybody reads. That two different quantities order two windows differently is surprising but not paradoxical.
This one is between two constructions of the same quantity. Both intervals are two-sided ninety-five per cent intervals for the same mean, built from the same resamples of the same draws, differing only in whether each resample carries its own scale. There is no sense in which one of them is measuring something else.
So the earlier field could say read the quantity a practitioner reads. This one cannot say the equivalent, because both of these are intervals a practitioner reads, and the standard advice — studentise — points at the one that reverses the ordering and covers worse under the window the earlier fields recommend.
What a practitioner is supposed to do
The honest answer is short and it is not comfortable.
Do not choose a window on the strength of any of these orderings. They are two to seven points wide, they reverse between two constructions of the same interval, and the shortfall they sit inside is three to nineteen points on every cell.
Choose the block length carefully instead. Across every interval on this table the spread between the four rules is larger than the spread between the two windows, and the rule of thumb is the worst rule on all three. The plug-in — a length estimated from the sample’s own lag-one correlation — is the best feasible rule on the percentile and studentised columns under both windows.
And do not expect ninety-five per cent. On a hundred and twenty rows at a persistence of 0.7, a block bootstrap interval for a mean covers between three quarters and nine tenths of the time, whichever window, whichever rule and whichever of three constructions is used. A reader told an interval is ninety-five per cent is being told something that is not true by a margin larger than any of the choices being argued over.
That last is the only one of the three that would change what somebody does, and it is the one none of the fields in this line set out to find.
What is left standing
Three things survive the whole line of fields and they are worth separating from what does not.
Nothing reaches its promise. Twenty-four cells, three intervals, four rules, two windows, and the best of them is 92.3%. That reading is stable across every construction tried and it is the one a practitioner should act on.
The choice of block length dominates the choice of window. The rule of thumb’s cells are the worst on every interval — 80.0%, 81.7% and 75.7% under the two windows on the percentile and studentised columns — and the spread across rules is larger than the spread across windows on all three intervals.
And the ordering between the windows is not a property of the windows. It is a property of the pair (window, construction), it has now been shown to reverse on the instrument and on the interval, and a recommendation that names a window without naming what will be read off it is a recommendation with a free variable in it.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The instrument and the reading — both name block bootstrap, block length, confidence interval, coverage, critical value, dependence, long-run variance, monte carlo, reference distribution, resampling, tapering
- Measuring a variance rather than a quantile — both name block bootstrap, block length, critical value, long-run variance, monte carlo, reference distribution, resampling, tapering
- A taper and a critical value — both name block bootstrap, critical value, dependence, long-run variance, reference distribution, resampling, tapering
- An ordering that depends on the rule — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering
- The error no window repairs — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering
- The length nobody has — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering
Named objects
A flat tag is an object no other essay names yet.
Block bootstrapBlock lengthConfidence intervalCoverageCritical valueDependenceEstimated varianceLong-run varianceMonte CarloPercentile intervalReference distributionResamplingSkewnessStudentised bootstrapTapering