A block at every starting row
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
The interval with no resampling in it replaced the 1.96 in a normal interval for the mean of a dependent series with Student’s on one fewer degrees of freedom than there are whole blocks, and found that it covered at least as often as the studentised bootstrap at every cell of a grid of series lengths and block lengths, and was narrower wherever few blocks were left. Its scale is the block-means variance: cut the series into disjoint blocks of rows, take each block’s mean, and multiply the sample variance of those numbers by . At 120 rows and blocks of thirty-two, that is a variance of three numbers, and the multiplier on two degrees of freedom is 4.30.
A series of 120 rows does not contain three blocks of thirty-two. It contains eighty-nine of them: one starting at the first row, one at the second, and so on to the block that ends on the last row. The disjoint scale reads three of those and discards the rest. The question here is what the other eighty-six are worth — whether reading every starting row changes what the scale finds, or only how reliably it finds it — and whether a taper applied inside each block, which the essay that derived the windows showed changes the attenuation itself, changes the answer.
Three readings of one series
The three scales are built the same way and differ in which blocks they read. Each takes a set of blocks, sums the deviations of the observations from the sample mean inside each block, squares the sums, and averages them, divided by the block length so that the result estimates the long-run variance that governs the variance of the mean. The disjoint scale reads the blocks that tile the series. The overlapping scale reads all blocks, one at every starting row; it is the estimator simulation practice calls overlapping batch means. The tapered scale reads the same overlapping blocks with each observation inside a block weighted by the trapezoid that rises over the first 43% of the block and falls over the last 43%, normalised to unit mean square.
Centring at the sample mean subtracts a little from every block’s sum, and the overlapping scales are divided by to put it back, where is one for the rectangle and 0.76 for the trapezoid, whose small end weights lose less of each block’s sum to the mean. For the rectangle this is the familiar of overlapping batch means. It is not a small adjustment when blocks are long and the series short: at blocks of thirty-two in 120 rows it multiplies the rectangle’s scale by 1.36 and the trapezoid’s by 1.26. It is exact for independent observations and the leading term under dependence, and every expectation quoted below is the expectation of the corrected scale, so nothing about it is assumed away.
Each scale is then read against Student’s on its own degrees of freedom. For the disjoint scale they are , as before. For the overlapping scales the standard rule divides by the integrated square of the lag window the blocks imply, . For the rectangle is the triangle and approaches two thirds, which is the textbook statement that overlapping batch means are worth half as many degrees of freedom again as disjoint ones. For the trapezoid is smaller still, and the rule promises more.
The hero figure sets the three side by side at the cell where the disjoint interval is weakest on width. Read from three disjoint blocks, the interval covers 94.4% at a width of 1.560. Read from every starting row it covers 94.8% at 1.054, with a multiplier of 2.74 on 4.1 degrees of freedom. Tapered, it covers 95.0% at 1.020. The interval has lost a third of its width and none of its coverage, and nothing was resampled to get there.
The same triangle, read more often
The first thing to settle is whether reading more blocks changes what the scale is measuring. A block weighted inside itself showed that what a block-based variance keeps of each autocovariance is the block window’s self-convolution: for a rectangular block of rows, a triangle that keeps all of the variance, of the first autocovariance, and nothing beyond lag . That is a statement about a single block. A disjoint scale and an overlapping one average the same statement over different sets of blocks, so they should expect the same thing.
They do. Every scale here is a sum of squared linear combinations of the observations, so its expectation under a known autocovariance is a finite sum, computed exactly rather than estimated. At 120 rows the disjoint and overlapping rectangles expect 0.461 and 0.461 of the truth at blocks of four, 0.655 and 0.653 at eight, 0.805 and 0.801 at sixteen, and 0.886 and 0.877 at thirty-two. Across all twelve cells of the grid — 120, 240 and 480 rows against the four block lengths — the two never differ by more than 0.0086 of the truth, and the difference that remains is the centring correction acting on blocks that overlap the ends of the series differently. The drawn averages over eight thousand series agree with the exact ones to the third decimal.
So reading every starting row does not touch the attenuation. It is the same triangle, read eighty-nine times instead of three, and whatever the interval loses to the triangle it loses either way.
The taper is a different window and a different self-convolution, and its expectation does move. At blocks of thirty-two it keeps 0.914 of the truth where the rectangle keeps 0.877: a window that reaches zero at its ends has a self-convolution flat at the origin, so it keeps more of the short-lag autocovariances, which at a correlation of 0.7 are most of what there is. At blocks of four it keeps 0.404 where the rectangle keeps 0.461, because a four-row trapezoid weights its middle two rows and its effective length is nearer two than four. That reversal at short blocks is the one the windows essay found for resampled means, where the taper’s advantage arrived only at block lengths longer than a short series can afford. It holds here for the scale too.
What the other eighty-six blocks are worth
If the expectation does not change, what can? Only the spread. Each disjoint block is close to independent of the next when blocks are long, so a disjoint scale on blocks behaves like a variance with degrees of freedom. Overlapping blocks share most of their rows with their neighbours and are far from independent, but averaging over all of them still reduces the scale’s variance, and the three-halves rule says by how much.
The realised degrees of freedom are the measurement that matters, because the multiplier is a claim about them: reading the interval on degrees of freedom claims that the scale varies like a . Twice the squared mean of the scale over its variance is what would have to be for that to hold, and it is directly measurable on the draws.
At long blocks the rule is close to honest. At three blocks of thirty-two the disjoint scale realises 2.0 degrees of freedom against the 2 it is read on, and the overlapping one realises 3.8 against the rule’s 4.1 — 91.9% of what is promised, and nearly twice the disjoint scale. The taper realises 4.8 against 5.0. Where the disjoint scale is a variance of three numbers, the overlapping one is roughly a variance of five, and the multiplier falls from 4.30 to 2.74.
At short blocks the rule fails, and fails badly. At blocks of four in 120 rows the overlapping scale is read on 42.2 degrees of freedom and realises 21.4, which is 50.8% of the promise and hardly more than the disjoint scale’s 21.2. The disjoint scale itself realises only 73.1% of its own there. Blocks of four rows in a series whose lag-one correlation is 0.7 are not close to independent of their neighbours, so neither reading of them has the degrees of freedom a count of blocks suggests. The three-halves rule is derived for blocks long enough that the dependence between them is negligible, and below that it overstates what the extra blocks are worth by a factor of two.
That overstatement costs little in practice, for a reason the next figure shows: at forty degrees of freedom or at twenty, a multiplier is within a few hundredths of 1.96. The rule is wrong exactly where being wrong is cheap, and nearly right where it is expensive.
Where the extra degrees of freedom narrow the interval
A multiplier on degrees of freedom is wide only when is small, so the extra degrees of freedom should narrow the interval only where the disjoint blocks were few.
At three disjoint blocks, reading every starting row gives 67.6% of the disjoint interval’s width, and the taper 65.4%. At seven blocks — 240 rows in blocks of thirty-two, or 120 in blocks of sixteen — the overlapping interval is 0.634 wide against 0.686, and 0.843 against 0.913. From fifteen blocks up the overlapping rectangle is within 3.0% of the disjoint one at every cell, because a multiplier on fourteen degrees of freedom is 2.14 and on twenty-one it is 2.08.
What makes the narrowing worth having is that coverage does not move with it. At 120 rows in blocks of thirty-two, the cell where the width falls by a third, the coverage goes from 94.4% to 94.8%, a difference smaller than its own standard error of about a quarter of a point. At seven blocks of thirty-two in 240 rows the disjoint interval covers 94.19% and the overlapping one 94.09%; at fifteen blocks of thirty-two in 480 rows, 93.95% and 93.71%. The overlapping interval is narrower by exactly the amount its scale is more reliable, and no narrower.
This is the comparison the studentised bootstrap lost. At three blocks of thirty-two in 120 rows it covered 90.4% at a width of 2.495 on the resampled grid, because it pays for a scale resting on three numbers twice: once in the degrees of freedom, and again in resampled quantiles of a ratio with that noisy scale underneath. The disjoint interval pays the first and not the second, at 1.560. Reading every starting row reduces the first payment itself, to 1.054, and the interval that results is less than half as wide as the studentised one at a coverage four points higher.
Coverage follows what the scale expects to find
The widths say that reading more blocks buys precision in the scale. The coverages say what it cannot buy.
Plotted against the share of the truth each scale expects to find, the thirty-six coverages fall on one band, whichever scale they belong to. At blocks of four every scale expects between 40.4% and 47.4% of the true long-run variance and covers between 78.8% and 82.8%. At blocks of thirty-two they expect 87.7% to 93.2% and cover 93.7% to 95.0%. Where the scale is attenuated, the interval is short, and how the blocks were read only moves a point along the band rather than off it.
That is the triangle again, and the reason the error no window repairs found most of a block estimator’s error at short series lengths outside anything a window choice could reach. Reading every starting row reduces the scale’s variance; it leaves its bias exactly where it was, and the bias is what keeps short-block intervals near 80%. The only dial that moves a point up the band is the block length, which is a length nobody has when it has to be chosen from the same 120 rows.
The taper sits on the band too, and that is its whole story in one picture. At thirty-two rows it expects more of the truth than the rectangle and so sits higher, at 95.0% where the overlapping rectangle covers 94.8%. At four rows it expects less and sits lower: 78.8% against 81.7%. A taper is a decision about the attenuation, and the attenuation decides the coverage, so a taper helps exactly where it keeps more and hurts exactly where it keeps less.
A number the earlier grid could not resolve
Eight thousand draws a cell were affordable here because nothing is resampled; the grid that introduced the interval drew 240 series a cell, with two hundred resamples each, and its coverages carried a standard error near one and a half points. That grid reported the disjoint interval at fifteen blocks of thirty-two in 480 rows covering 95.0%, the one cell where any interval on it reached the promise.
On eight thousand draws the same interval at the same cell covers 93.95%, with a standard error of 0.27 points. The difference is within what 240 draws could produce, and it means the earlier reading was a favourable draw rather than a cell that delivers 95%. Nothing in the earlier essay’s argument depends on it — the interval still covers at least as often as the studentised one wherever both were measured — but the claim that it reaches the promise at that cell does not survive being counted more carefully. At this grid’s resolution only the tapered overlapping scale at three blocks of thirty-two reaches 95.0%, and it does so within a standard error.
What reading every starting row is for
It is for short series with long blocks. Where the block length leaves few disjoint blocks — three to seven, which is where a series of a hundred or two hundred rows at strong dependence ends up — reading a block at every starting row cuts the interval’s width by between about 8% and a third at no measurable cost in coverage, and the realised degrees of freedom are close to the three-halves rule’s.
It is not a repair for attenuation. Where the interval is short of 95%, it is short because the scale expects too little of the long-run variance, and every way of reading rectangular blocks expects the same. The coverage moves with the block length and not with how the blocks are read.
A taper on top of the overlap is a block-length decision. At thirty-two rows the trapezoid keeps 0.914 of the truth against 0.877, realises 4.8 degrees of freedom against 3.8, and gives the narrowest interval on the grid at the highest coverage. At four rows it keeps less and covers three points worse. The ordering that depends on the rule between windows has the same root: whether a taper helps is a question about the block length it is used at, not about the taper.
The three-halves rule should not be trusted at short blocks, where it promises twice what the overlapping scale realises. It does no harm there only because the multiplier is already close to 1.96.
How the numbers were made, and where they stop
Every expectation is exact: each scale is written as a sum of squared linear combinations of the observations, and its expectation under a first-order autoregression at 0.7 with unit marginal variance is computed term by term from the autocovariances. Every coverage, width and realised degree of freedom is over eight thousand draws per cell, with the three scales read off the same series at every draw, so a difference between scales is a difference in how the series was read. The disjoint and overlapping expectations are required to agree within a hundredth at every cell, and the drawn and exact expectations within sampling error. The three-halves rule is refused as a statement of what overlapping blocks of four are worth: it promises 42.2 degrees of freedom and the scale realises about half that.
Not measured: dependence other than a first-order autoregression, and other strengths of it. At a correlation of 0.3 the triangle’s bias at blocks of four would be small, the band would sit near 95%, and the question would become almost entirely one of width. A statistic other than the mean is not measured either; a median or a ratio has no block-means scale of this form, and the overlapping trick has no direct analogue there without resampling.
Still open: a block length chosen for the overlapping scale
Every block length here is stated rather than chosen. The rules that choose one — a plug-in for the bias-variance trade-off, a rule of thumb written into a protocol — were all derived for a scale whose variance is that of degrees of freedom. An overlapping scale has a smaller variance at every length, so the trade-off it faces is different: the same bias costs the same, and the variance it is traded against is smaller, which moves the best length longer.
How much longer, and whether a rule built for the overlapping scale — or one built for the tapered overlap, whose bias falls faster and whose variance is smaller still — closes the gap to 95% at 120 rows that no stated length on this grid closes, is computable with the same exact expectations and realised degrees of freedom used here. It would say whether the eighty-six discarded blocks, once read, change which block length a practitioner should ask for.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The instrument and the reading — both name block bootstrap, block length, closed form, confidence interval, coverage, dependence, long-run variance, tapering
- An interval that carries its scale — both name block bootstrap, block length, closed form, confidence interval, coverage, dependence, long-run variance
- The ordering reverses again — both name block bootstrap, block length, confidence interval, coverage, dependence, long-run variance, tapering
- What choosing the length costs — both name bias-variance, block bootstrap, block length, closed form, dependence, long-run variance, tapering
- What the interval covers — both name block bootstrap, block length, confidence interval, coverage, dependence, long-run variance, tapering
- A length for each instrument — both name bias-variance, block bootstrap, block length, dependence, long-run variance, tapering
Named objects
A flat tag is an object no other essay names yet.
Bias-varianceBlock bootstrapBlock countBlock lengthClosed formConfidence intervalCoverageDegrees of freedomDependenceInterval widthLong-run varianceTapering