The interval with no resampling in it
Worth reading first: Where the bootstrap lies.
The essay that separated the block count from the block length ended on three things it could have run and did not, and it ranked them. The cheapest, and the one it judged most likely to change a recommendation, was an interval that took the block count’s degrees of freedom at face value: the normal interval on the block-means variance, with its 1.96 replaced by Student’s on one fewer degrees of freedom than there are whole blocks. It is one line. It needs no resample.
Run on the same grid and the same draws, that line covers at least as often as the studentised bootstrap interval at all twenty-four cells, by between 0.42 and 10.42 points. Where seven or fewer whole blocks are left it is also narrower, down to 0.614 times the studentised interval’s width at three blocks. And at fifteen blocks of 32 rows it covers 95.0% of its draws, which no resampled interval reaches at any cell of the grid.
Studentising was supposed to be the repair for an interval that did not know how noisy its own scale was. On this grid the repair is available without the resampling, and the resampling, where it differs from the repair, makes the interval wider and covers less.
One line, and no resampling
Every cell of the grid draws a first-order autoregression at a coefficient of 0.7, of rows, and asks for a 95% interval for its mean. Cut the series into whole blocks of length , and let be times the sample variance of the block means — the scale every interval on the grid already uses, on the series and on every resample alike. The normal interval is . The interval here is
which is the batch-means interval of the simulation literature under another name, and it has one justification that is not an approximation. If the block means were independent and normal, would be a scaled chi-square on degrees of freedom and the ratio would be exactly Student’s . The block means of a dependent series are neither, and the grid is where that matters.
Two things are true of it by construction, and both are checked rather than assumed. Its width on every draw is the normal interval’s width times , the multiplier the earlier essay called , so the ratio of the mean widths across a cell is that multiplier — to within 8.9e-16 at every cell. And it consumes no random numbers, so under the rectangular window and the taper it is the same interval, and adding it to the sweep changes none of the numbers the grid already reported. That is two routes to one number in the most literal form: the resampled intervals and this one share every draw, every block mean and every scale, and differ only in where the multiplier comes from.
The grid’s harshest cell is 120 rows in blocks of 32, which leaves three whole blocks and a variance estimated from three numbers. The normal interval covers 82.1% there at an average width of 0.708, and the percentile interval covers the same 82.1% at 0.632. The studentised interval covers 90.4%, and pays for it with a width of 2.495, three and a half times the normal interval’s.
The interval covers 94.2% at a width of 1.555. It is wider than the normal interval by , exactly, and narrower than the studentised one by more than a third. Three blocks is where studentising was supposed to earn its cost, because it is where a normal multiplier is most wrong, and three blocks is where the closed multiplier beats it by the widest margin on both columns at once.
At every cell of the grid
One cell could be a coincidence of three blocks. The grid has twenty-four.
At every one of the twenty-four cells the interval covers more often than the studentised bootstrap interval. The margin runs from 0.42 points, at 480 rows in blocks of 4 under the rectangle, to 10.42, at 240 rows in blocks of 8 under the taper, and averages 4.11. No cell is a tie.
The largest margins are all under the taper, and that is not a property of the interval, which has no window. It is the reversal the field found earlier: the tapered window helps the percentile interval and hurts the studentised one, at every cell. The interval sits at one height for both windows, so the taper’s damage to the studentised interval shows up as extra distance above the diagonal. Under the rectangle alone the margin runs from 0.42 to 3.75 points; under the taper, from 4.58 to 10.42.
What the picture does not show is a region where resampling wins. A reader looking for the cells where the bootstrap earns its keep — few blocks, where a normal approximation is worst, or many, where resampling has the most to work with — finds the interval ahead at both ends and in between.
Where the width goes
Coverage bought with width is not a repair, and the earlier essays in this field were careful to read the two columns together. So the width is set beside it, cell by cell.
At three blocks the interval is 0.614 to 0.623 times the studentised interval’s width, and at seven it is 0.925 to 0.971. From fifteen blocks up the two are within 2.7% of each other, from 0.993 to 1.027 — so at the cells where the interval is slightly wider it is wider by a couple of per cent and covers more by a point or more, and at the cells where it is much narrower it also covers more.
The earlier essay measured the studentised interval’s width against and found it above that floor at every cell, by a factor of 1.80 at three blocks falling to 1.005 at a hundred and twenty. The interval sits exactly on the floor. The reading that fits both is that the studentised interval pays for a noisy variance twice: once in the degrees of freedom, which any honest interval must pay, and again in the resampled distribution of a ratio whose denominator is resampled from the same few blocks — a resample of three blocks drawn with replacement from three blocks repeats a block more often than not, and its variance is noisier than a chi-square on two degrees of freedom. That second payment is what the excess over measures, and nothing in the coverage column says it bought anything.
It is worth being plain that this is a reading and not a derivation. What is measured is that the interval’s width is times the normal interval’s on every draw, that the studentised interval’s is more, and that the extra width comes with less coverage rather than more.
What the multiplier alone is worth
The interval and the normal interval share everything except the multiplier, so the difference between their coverages is what the degrees of freedom are worth at each cell, and nothing else.
At three blocks that difference is 12.08 points: the normal interval covers 82.1% and the interval 94.2%, so the multiplier alone closes 94% of the normal interval’s shortfall. The studentised interval, starting from the same scale on the same draws, closes 65% of it under the rectangle and 52% under the taper. At seven blocks the multiplier is worth 5.00 points at both cells that leave seven, and at a hundred and twenty blocks it is worth 1.25, closing 9% of a shortfall of 13.75 points.
That is the shape a correction for estimating a variance from numbers ought to have. Where there are few blocks it is nearly the whole of the repair. Where there are many, the normal interval is short for a reason that has nothing to do with how many numbers the variance came from, and the multiplier correctly declines to pretend otherwise. The studentised interval does not have that shape. At three blocks it closes less of the shortfall than the multiplier does and adds more than twice as much width to get there, and at a hundred and twenty blocks under the taper it covers 77.5%, 3.75 points below the normal interval it was built to improve on.
Fifteen blocks of thirty-two
At 480 rows cut into fifteen blocks of 32, the normal interval covers 91.7% at a width of 0.392, the percentile interval 87.5% at 0.374, and the studentised interval 92.5% at 0.423. The interval covers 95.0% — 228 of 240 draws — at 0.429, 1.01 times the studentised interval’s width.
That is the only cell on the grid where any of the four intervals covers what it promises, and it is the same cell under both windows because the interval does not see the window. The field that first measured these intervals ended on the reading that none of them reached 95% anywhere, and the essay that tripled its sample sizes found the same at twenty-four cells. The reading survives for every resampled interval. It does not survive for the one that resamples nothing.
A standard error keeps that honest. Over 240 draws, a coverage of 95% has a standard error of about 1.4 points, so the cell is consistent with the promise rather than proof of it, and the neighbouring cells at 94.2% and 94.6% are consistent with it too. What the grid does establish is the ordering around it. The best resampled cell anywhere is the tapered percentile interval at 93.8%, inside a standard error of the promise at a cell where the interval covers 92.9%; the best studentised cells are at 92.5%, 1.78 standard errors short, and the interval covers 94.6% and 95.0% at those two cells.
The count and the length, again
The essay that separated the two dials found that the studentised interval’s width followed the block count and its coverage followed the block length. The interval makes the first half of that true by construction, since its width multiplier is a function of the count alone. The second half is a question it has to answer by measurement.
Holding the length at eight while the count runs 15, 30, 60, the interval covers 90.8%, 90.8% and 89.6%. Holding the count at fifteen while the length runs 8, 16, 32, it covers 90.8%, 92.9% and 95.0%. The first path moves it by −1.25 points and the second by 4.17.
So it follows the same dial as the studentised interval, for the reason the earlier essay gave. Its width is fixed from the count, and the degrees of freedom are priced exactly. What it covers depends on whether the scale is the right scale, and that depends on how much of the series’ dependence a block of rows carries. A longer block leaves less correlation between neighbouring block means for the variance to miss. The multiplier has nothing to say about that, and neither does resampling.
Where it is still short
The interval is not a repair for everything, and the cells where it fails say which problem it does not touch. Its shortfall is largest at blocks of four at every sample size: 10.42 points at 120 rows, 12.50 at 240 and 12.50 at 480. Those cells have 30, 60 and 120 blocks, so is between 1.04 and 1.01 and the degrees of freedom are not the problem. The scale is.
The correlation between neighbouring block means of a series at 0.7 has a closed form, and it is 0.41 for blocks of four, 0.23 for eight, 0.10 for sixteen and 0.05 for thirty-two. A block-means variance treats those means as uncorrelated, so at blocks of four it leaves out a large positive covariance and reports a scale that is too small. At blocks of 32 the covariance it leaves out is small, which is why the interval’s own justification — independent block means — nearly holds there, and why its coverage reaches the promise there and nowhere else. A scale that is too small is not rescued by any multiplier chosen for how many blocks there are, which is the error no window repairs, one interval further along.
Beside it, the studentised interval’s shortfall at the same cells is 11.67 and 17.50 points at 120 rows, and 13.75 and 21.25 at 240. At no cell of the twelve is the studentised interval closer to its promise than the interval.
Where resampling still earns something
The claim is against the studentised interval, and it would be wrong to widen it to every resampled interval.
The percentile interval under the taper covers more than the interval at two cells, both at 480 rows: 90.4% against 89.6% in blocks of 8, and 93.8% against 92.9% in blocks of 16. It ties it at two more. Those margins are under a point — two draws in 240 — and the percentile interval is narrower at both, 0.324 against 0.354 and 0.366 against 0.400. That is the one corner of the grid where a resampled interval is ahead on both columns at once, and it is the tapered window at the larger sample sizes, which is where the ordering that reversed put the taper’s advantage.
Everywhere else the percentile interval is short of the interval, by as much as 12.08 points at three blocks under the rectangle. So the reading is not that resampling is useless for a mean; it is that the resampled interval built to fix the scale’s noise does worse than the closed interval that prices the noise directly, and that the one resampled interval that holds its own does so at a handful of cells and by a margin the draws barely resolve.
What a reader should take from it
For a mean, choose the block length for the dependence and let the count set the multiplier. The coverage follows the length, so the length is the choice that matters; the width follows the count, and Student’s on degrees of freedom prices it in closed form before any data are resampled. The rules that choose a length choose it on other grounds, and none of them states the count it leaves.
Studentising by resampling is dominated here. On every cell of this grid the studentised bootstrap interval covers less often than the interval on the same draws, and where it is much wider it is also further short. That is a statement about the mean of one autoregression on one grid, and it is a strong one on that grid.
And the scale is still the problem the multiplier is not. Where blocks are short the interval is as far short as twelve and a half points, and the only thing that moved it was a longer block.
Still open, and what comes next
A statistic that is not a mean. The interval’s justification is that a mean of block means is close to normal; the studentised bootstrap’s is that it does not need that. A median, a quantile or a ratio of two means is where the second justification should be worth something, and where a closed multiplier has less claim to be right. Rerunning the grid for one of those is the next measurement, and it is the one that could reverse this essay.
A rule that prices both terms. The width term now has a closed form with no resampling in it, , and the coverage term follows the length. A block length chosen to minimise the expected width subject to a coverage floor is buildable from those two pieces, and it needs its own oracle and its own sweep to be scored on data it was not fitted to.
Overlapping blocks and weaker dependence. Every block here is disjoint and every series is at 0.7. Overlapping batch means raise the effective number of blocks by about half, which would narrow the interval at the harsh cells; and at weaker dependence shorter blocks suffice, the counts rise, and the two intervals should converge. Neither is measured, and the number of observations that repeat each other sets how fast that convergence would be.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An interval that carries its scale — both name block bootstrap, block length, closed form, confidence interval, coverage, dependence, estimated variance, long-run variance, percentile interval, resampling, studentised bootstrap
- The instrument and the reading — both name block bootstrap, block length, closed form, confidence interval, coverage, dependence, long-run variance, resampling
- An ordering that depends on the rule — both name block bootstrap, block length, closed form, dependence, long-run variance, resampling
- The length nobody has — both name block bootstrap, block length, closed form, dependence, long-run variance, resampling
- The reversal that was the instrument's — both name block bootstrap, block length, coverage, dependence, long-run variance, resampling
- A block weighted inside itself — both name block bootstrap, closed form, dependence, long-run variance, resampling
Named objects
A flat tag is an object no other essay names yet.
Block bootstrapBlock countBlock lengthClosed formConfidence intervalCoverageDegrees of freedomDependenceEstimated varianceInterval widthLong-run variancePercentile intervalResamplingStudentised bootstrap