The instrument and the reading
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
Three fields in a row have compared two block windows, and every number in all three is the same quantity: the error in an implied long-run variance, Σκ(k)γ̂(k), computed in closed form from the residuals’ own autocovariances with no resampling in it at all.
The field that introduced it says why: a quantity computed rather than sampled has no resampling noise, so what a sweep measures is the noise in γ̂ rather than the noise in the bootstrap. It also says, in its own title, what the trade is. A variance is not a quantile, and it is the quantile a test reads.
The field that re-ran the comparison on feasible rules closes on the same admission and sharpens it: two windows’ reference distributions differ in shape as well as in scale, so a quantile reading could order the rules differently.
This field takes the reading. Three of them, from one computation.
Three readings of one bootstrap
The point of taking all three from one computation is that it removes every explanation for a disagreement except the one the field is testing. Three separate measurements could differ because of three seeds, three replication counts, three resampling schemes; three readings of one sorted array cannot.
The design is the part that matters, because three instruments measured three ways would be three measurements and could differ for three reasons.
For each draw and each block length, the resampled means are formed once — a few hundred block resamples of the residual series, each averaged and standardised by √n — and sorted. Everything else is read off that one sorted array:
the variance of it, which is the implied long-run variance the earlier fields use, and which is checked to be that quantity rather than assumed;
its 95% point, which is the critical value a one-sided test compares its statistic to;
and its 2.5% and 97.5% points, which are the ends of the interval a reader is handed.
So a difference between the three readings cannot be a difference in the resampling, in the seed, or in the number of replications. It can only be a difference in what is being read.
One thing the design gives up
Reading all three from one bootstrap buys the comparison and costs something, and the cost is worth naming.
The implied variance the earlier fields use has no resampling noise in it at all — it is a closed-form sum over the sample’s own autocovariances, so its error is entirely the error in γ̂. The version read here is the variance of three hundred resampled means, which has the same target and a little extra noise on top.
That makes the comparison conservative in the right direction. If the resampled variance ordered the two windows differently from the closed-form one, the difference would be resampling noise rather than a finding — and it does not: the ordering it gives is the earlier field’s, sign for sign, at every rule. The instrument that this field says is the wrong one is reproduced here before it is argued with, which is what makes the argument about the reading rather than about the implementation.
Each against its own truth
An error needs a truth, and the three do not share one.
The variance is measured against the law’s long-run variance summed exactly — 5.5811 at a persistence of 0.7 on a hundred and twenty rows, which is a closed form.
The coverage is measured against 95%, which needs no argument.
The quantile is the one that needs care, and getting it wrong would have decided the field.
The asymptotic answer is 1.645 times the square root of the long-run variance, which is 3.9155. The truth at a hundred and twenty rows, found by drawing forty thousand series from the law and taking the quantile directly, is 3.8891 — a ratio of 1.0068.
The gap is small and reading a bootstrap against the wrong one would have charged it for a finite-sample correction it did not cause. The whole subject of these fields is that a finite sample is not its own limit, so the truth here is the finite-sample one, computed by simulating the law rather than by a formula about it.
A quantile is the dearer reading
At every one of the eight cells the quantile costs more. At the plug-in rule the tapered window’s implied variance carries a root mean squared relative error of 45.34% and its 95% point carries 59.15%. At the rectangle’s protocol length, 45.83% against 65.51%.
That is not a defect in the bootstrap and it is not news. A quantile is a statement about the shape of a distribution as well as its scale, and a fixed number of resamples estimates a tail worse than it estimates a second moment. The field that first priced the two instruments establishes exactly that, and uses it to justify the cheaper one.
What that justification does not cover is the ordering. Being dearer to measure says nothing about which window comes out ahead, and the earlier field’s own finding — that the two reference distributions differ in shape — is precisely a reason to expect the ordering to move.
How much of the quantile’s extra error is the resampling
The explanation offered for the quantile costing more — that a fixed number of resamples estimates a tail worse than a second moment — is correct and it is worth measuring, because on these numbers it accounts for almost none of the gap.
With a few hundred resamples the Monte Carlo contribution can be computed rather than guessed. For a roughly normal resample distribution the relative standard error of a 95% point is , which at B = 300 is 7.4%; the relative standard error of a standard deviation is , which is 4.1%.
Subtract each in quadrature from the measured errors. The tapered window at the plug-in rule carries 59.15% on the quantile and 45.34% on the variance; removing the resampling noise leaves 58.68% and 45.16%.
So of the 13.8-point gap between the two readings, the extra Monte Carlo noise is 0.5 points. The other 13.3 are the quantile being a harder thing to get right from a hundred and twenty rows, whatever the resampling budget.
The second cell says the same. The rectangle at its protocol length reads 65.51% and 45.83%, a gap of 19.7 points, of which the resampling accounts for 0.24.
Which means more resamples would not help
That decomposition settles a question a reader would reasonably ask, and settles it against the obvious remedy.
Running the bootstrap with an unlimited number of resamples — B = ∞, no resampling noise at all — would move the quantile’s error from 59.15% to 58.68% and the variance’s from 45.34% to 45.16%. Under one point in the first case and under a fifth of a point in the second.
The two instruments are not separated by how well they are being sampled; they are separated by how well the bootstrap’s reference distribution matches the truth in the tail against how well it matches it in the middle. A resample distribution can have very nearly the right spread and the wrong shape, and the second reading is the one that notices.
Which is exactly why the ordering was worth re-taking rather than deduced. If the quantile’s extra cost were resampling noise it would be noise around the same answer, and the ranking of the windows could not move. It is a shape error instead, and a shape error is a systematic thing that can favour one window over the other.
What the three readings are for
It is worth being concrete about who reads which, because the case for the quantile is not that it is more fundamental.
A long-run variance is read by nobody and is used by everybody. It is an intermediate: a standard error is its square root over √n, and a standard error is what goes into a t-statistic. So the implied variance is one step from a reading rather than no steps, and a rule that gets it right gets the standard error right.
A critical value is read by a test. A block bootstrap’s 95% point is what a one-sided test compares its statistic to, and it is the object the taper field’s own critical-value essay is about. It uses the tail of the reference distribution and nothing else.
And an interval is read by a reader. It uses both tails, so it is sensitive to the reference distribution’s skewness in a way neither of the others is — and a bootstrap distribution of a mean under dependence is not symmetric.
Three quantities, increasingly far from the sample and increasingly close to what somebody acts on. A field that measures only the first is measuring the thing furthest from the decision, which is fine when the three agree.
The grid, and what it gives up
One economy is worth stating because it is the only place this field is coarser than the one it corrects.
The field this one re-reads sweeps fourteen block lengths; this field sweeps seven — 2, 4, 8, 12, 20, 32 and 48 — over the same range. Every length here costs three hundred resamples where there it costs one closed-form sum, and seven lengths at four hundred draws and two windows is already 1.7 million resampled series.
What that coarsens is the oracle, which picks the best length from the grid it is given: a seven-point grid’s best is a little worse than a fourteen-point grid’s. It does not coarsen the three feasible rules, which pick a length by a formula and are then snapped to the nearest grid point — and the protocol length of eight and the rule of thumb’s four are both on the grid exactly, so the two rules whose ordering reverses are measured at the lengths they actually choose.
Where a shape difference comes from
The earlier field’s shape as well as scale finding is the premise of everything here, so it is worth restating what produces it rather than citing it.
A moving block bootstrap lays fixed-length blocks end to end. A tapered one weights each block down at its ends before laying it, which is what removes the join discontinuity and buys the order of convergence the taper field measures. The weighting changes two things at once: it attenuates the covariance the resample keeps, which is a scale effect and is what the implied variance sees, and it changes how many effective observations each block contributes, which alters the resampled mean’s distribution and not merely its variance.
A rectangle’s resampled means are a sum of nearly independent block sums; a taper’s are a weighted sum of the same blocks with the weights concentrated in the middle. Same variance target, fewer effective terms, a heavier tail. That is a difference no variance can report and every quantile can.
What the field does not change
Two things stay exactly as the earlier fields left them, and saying so is what keeps this a re-reading rather than a replacement.
The rules. A length written into a protocol, the rule of thumb , a plug-in from the sample’s own lag-one correlation, and the best length on the draw. Same four, same definitions, same plug-in algebra.
And the finding that the choice of length dominates the choice of window. The best feasible rule’s shortfall against the oracle is larger than the gap between the windows at the oracle, on every instrument here as on the one the earlier field used. Nothing in this field makes the window choice matter more; it makes it point the other way.
Which it does
The three readings do not agree about which window is better.
On the implied variance the taper wins at the best available block length and at one estimated from the sample, and loses at a length written into a protocol and at the rule of thumb. That reversal is the whole of the earlier field’s recommendation — estimate the block length, and then use the taper.
On the 95% point the taper wins at all four rules. On the coverage the interval delivers, the taper wins at all four again.
Two of the four rules change sign between the first reading and the other two, and they are exactly the two the recommendation is about. The third essay of this field is the table, and the second is why: the two instruments want different block lengths, so a rule that is systematically long for one of them is being scored on a target it was not aimed at.
The persistence everything is measured at
One setting is fixed throughout and inherited rather than chosen: a first-order autoregression at 0.7, on a hundred and twenty rows, which is the world the crossing field and the feasible field both work in.
Keeping it is the point. This field’s job is to re-read those fields’ comparison on a different instrument, so changing the world at the same time would make the two differences inseparable. What it means is that every number here is one persistence and one sample size, and the earlier field’s own sweep over four sample sizes has no counterpart here.
What the numbers are, before the comparison
The eight cells are worth reading once as levels rather than as differences, because the differences the rest of this field is about are small inside them.
On the implied variance the errors run from 36.93% at the tapered window’s own best block length to 61.47% at the rule of thumb — a root mean squared relative error, so a typical draw’s implied long-run variance is out by a third to two thirds of the truth. On a hundred and twenty rows at a persistence of 0.7, estimating a long-run variance is hard.
On the quantile they run from 47.55% to 68.17%.
Every margin this field measures is a point or two inside errors of forty to seventy per cent. That is the same shape the earlier field reports — its window margins are two points of a thirty-six point error — and it is worth carrying: the comparison between two windows is a comparison between two ways of being badly wrong, and what changes when the instrument changes is which of them is slightly less so.
Why the coverage is a third reading rather than a restatement
A reader may reasonably ask whether coverage adds anything the quantile does not, since an interval is built out of quantiles. It does, and the reason is a general one about intervals.
A critical value’s error is a one-number miss: the 95% point is too high or too low, and the test’s size is wrong in one direction. An interval uses both ends, so its coverage depends on the two errors together — and two errors of the same size in the same direction shift the interval without changing its width, which loses nothing at all. Two of opposite sign narrow or widen it, which loses a great deal.
So a bootstrap distribution that is accurate on average and skewed the wrong way reads well on the quantile and covers badly, and one that is systematically shifted reads badly and covers fine. The two readings are sensitive to different features of the same reference distribution, which is exactly the situation the earlier field’s shape as well as scale finding says to expect.
Whether it happens here is the fourth essay’s subject. The short answer is that the coverage ordering agrees with the quantile’s and the levels do not: the best cell covers at 91.0% and the worst at 80.8%, on a promise of 95%.
What the cheap instrument bought and what it cost
It is worth being fair to the choice, because it was a good one and it was not free.
The implied variance made four fields possible. A sweep over fourteen block lengths, four rules, three windows and four hundred draws is fifty thousand evaluations; done by resampling at three hundred replications apiece it is fifteen million, and the fields that used it swept more than that. The field that measured the exchange rate puts it at about two orders of magnitude in the draws needed to separate two windows at two standard errors.
What it cost is one assumption that was never stated as one: that an ordering established on a variance is an ordering on a quantile. That is true when two distributions differ only in scale and false when they differ in shape, and the field that introduced the instrument is the field that established they differ in shape.
So the defect is not that the wrong instrument was used. It is that a known difference between the two instruments was measured, reported, and then not carried through into the comparison the instrument was used for.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The ordering reverses again — both name block bootstrap, block length, confidence interval, coverage, critical value, dependence, long-run variance, monte carlo, reference distribution, resampling, tapering
- What a multiplier cannot keep — both name block bootstrap, closed form, critical value, dependence, estimation error, monte carlo, persistence, reference distribution, resampling
- An ordering that depends on the rule — both name block bootstrap, block length, closed form, dependence, long-run variance, monte carlo, resampling, tapering
- The error no window repairs — both name block bootstrap, block length, closed form, dependence, long-run variance, monte carlo, resampling, tapering
- The interval with no resampling in it — both name block bootstrap, block length, closed form, confidence interval, coverage, dependence, long-run variance, resampling
- The length nobody has — both name block bootstrap, block length, closed form, dependence, long-run variance, monte carlo, resampling, tapering
Named objects
A flat tag is an object no other essay names yet.
Block bootstrapBlock lengthClosed formConfidence intervalCoverageCritical valueDependenceEstimation errorLong-run varianceMonte CarloPersistenceReference distributionResamplingStationary bootstrapTapering