A block, weighted inside itself

A length chosen for the interval

The squared-error rule for a block length weighs the scale's bias against its variance. Rebuilt for a block at every starting row, whose variance is two thirds as large, it asks for blocks fourteen per cent longer and buys four tenths of a point of coverage. A t interval has already paid for the scale's variance in its multiplier, so its own coverage is predicted by the bias alone — within seven tenths of a point up to blocks of twenty-eight rows — and a rule that chooses the shortest length that prediction says will cover asks for blocks two and a half to three times longer and lands within about a point of 95% at 120, 240 and 480 rows. The price is width, and it is paid where the series is short.

Worth reading first: Where the bootstrap lies · The observations that repeat each other.

A block at every starting row found that reading the block-means scale from every block a series contains, rather than from the disjoint ones, leaves its attenuation exactly where it was and roughly halves its variance. The t interval built on it is narrower wherever few disjoint blocks are left, at no measurable cost in coverage. What it could not do was move the coverage, because coverage follows what the scale expects to find, and every way of reading rectangular blocks expects the same.

It ended on the one dial that does move coverage: the block length. Every length in that essay was stated. The rules that choose a length were all derived for a disjoint scale, whose variance is that of k−1k - 1 degrees of freedom; an overlapping scale has a smaller variance at every length, so a rule that trades bias against variance should ask for longer blocks. The questions it left are how much longer, and whether a rule built for the overlapping scale closes the gap to 95% at 120 rows that no stated length on the old grid closed.

The answers are: fourteen per cent longer, and no. The rule that does close the gap is built on a different criterion, and the reason it is different is the most useful thing in this essay.

Coverage against length, on a fine grid

The series is the one this field has used throughout: a first-order autoregression at 0.7 with unit variance, the interval a t interval for its mean with the overlapping scale and its nominal degrees of freedom. The old grid had four lengths. The first step is a fine one — every length from two rows to a third of the series — so that an optimum can be placed to within a few rows.

Coverage against block length, 120 rowsThe coverage of the t interval for the mean of a first-order autoregression at 0.7 on 120 rows, at every block length on a fine grid, with its scale read from the disjoint blocks and from a block at every starting row, over 8,000 draws. The squared-error length for disjoint blocks at the true persistence is 13.9 rows, and the same algebra for the overlapping scale gives 16.0, marked. The overlapping interval first reaches 95% — to within a tenth of a point — at blocks of 32 rows, where it covers 94.9% at a width of 1.056; at the overlapping squared-error length it covers about 92.5%.0.7000.8000.900block length, rows (log scale)coverage of a nominal 95% interval2481632a block at every starting rowdisjoint blockssquared-error lengths, disjoint then overlapping120 rows, AR(1) at 0.7the dashed line is 95%
Fig. 1 The t interval’s coverage against block length at 120 rows, with its scale read from the disjoint blocks and from a block at every starting row, and the two squared-error lengths marked. The slider changes the length of the series.

At 120 rows both curves rise with the block length, as the triangle’s bias recedes. The overlapping interval first reaches 95%, to within a tenth of a point, at blocks of 32 rows, where it covers 94.9% at a width of 1.056. At 240 rows it first reaches it at 56 rows, covering 95.1% at 0.732; at 480 rows at 80, covering 95.3% at 0.481.

So the gap the earlier essay could not close at 120 rows is closable by a stated length, and the length is about a quarter of the series. That is a length the disjoint scale cannot afford: at thirty-two rows the disjoint interval rests on three blocks and is half as wide again. The overlapping scale’s whole contribution, in this light, is to make long blocks cheap enough to use.

What the squared-error rule asks for

The standard rule chooses the length that minimises the scale’s mean squared error as an estimate of the long-run variance. Its bias falls like 1/ℓ1/\ell and its variance rises like ℓ/n\ell/n; the optimum is the cube root of c n αc\,n\,\alpha, with α\alpha a function of the persistence — (2ρ/(1−ρ2))2(2\rho/(1-\rho^2))^2 for the rectangle — and cc a constant set by the scale’s variance. For disjoint blocks c=3c = 3. An overlapping scale’s variance is two thirds of the disjoint one’s, which multiplies cc by three halves and the length by their cube root, 1.1451.145.

At the true persistence of 0.7, the disjoint rule’s length is 13.9 rows at 120, 17.6 at 240 and 22.1 at 480. The overlapping rule’s is 16.0, 20.1 and 25.3. Both are marked on the figure, and both sit well to the left of where coverage reaches 95%: at the overlapping squared-error length the interval covers about 92.5% at 120 rows, 92.9% at 240 and 93.8% at 480.

The squared-error rule is doing its job. It finds the length at which the scale is the best estimate of the long-run variance, in the sense of squared error. That is simply not the length at which an interval built on the scale covers.

The interval has already paid for the variance

The reason is in what the t multiplier is for. A t interval on ν\nu degrees of freedom is wider than a normal one by exactly the amount the scale’s randomness requires, provided the scale varies like a chi-square on ν\nu. The essay on every starting row measured how nearly that holds and found the realised degrees of freedom close to the nominal ones at long blocks. So within the interval, the scale’s variance is not an error: it has been priced. What has not been priced is the bias. If the scale expects a share ss of the long-run variance, the interval is too narrow by s\sqrt s, and its coverage is

P(∣Tν∣≤tνs)P\big(|T_\nu| \le t_\nu \sqrt{s}\big)

— a number that needs only the scale’s exact expectation and its degrees of freedom, both of which are closed forms here.

The coverage the bias predicts, beside the coverage counted. At 120 rows, the coverage of the t interval on the overlapping scale at each block length, counted over 8,000 draws, beside the coverage predicted by the exact share of the long-run variance the scale expects and nothing else: P(|T| ≤ t√s) on the scale's nominal degrees of freedom. Up to blocks of twenty-eight rows the two agree within 0.7%; at blocks of 40 the prediction is 94.2% and the count 96.0%. The scale's randomness is already in the multiplier, so what is left to predict the coverage is the bias.
Fig. 2 At 120 rows, the coverage the scale’s bias alone predicts at each block length, beside the coverage counted over eight thousand draws.

Up to blocks of twenty-eight rows the prediction and the count agree within 0.7%. At the longest blocks the prediction is cautious: at forty rows it says 94.2% and the count is 96.0%, so on a scale resting on three blocks’ worth of degrees of freedom something the formula leaves out works in the interval’s favour. Over the range where a practitioner would choose, the bias predicts the coverage, and nothing else needs to.

That is the whole difference between the two criteria. Squared error charges the scale for its variance; the interval has already paid that charge in its multiplier, and charging it again is what pulls the squared-error length short. For an interval the right objective is the bias, subject to the width the multiplier then demands — and the shortest length whose bias the multiplier can absorb.

Three rules on the same draws

A criterion is not a rule until it can be computed from the sample. All three rules here read the sample’s own lag-one autocorrelation, ρ^\hat\rho, and nothing else:

  • the squared-error rule for disjoint blocks, (3nα^)1/3(3n\hat\alpha)^{1/3};
  • the squared-error rule for the overlapping scale, (4.5nα^)1/3(4.5n\hat\alpha)^{1/3};
  • the coverage rule: the shortest length at which the predicted coverage, computed from the overlapping scale’s exact share at ρ^\hat\rho and its nominal degrees of freedom, reaches 94.5%, capped at a third of the series.

Each is applied to every draw, the overlapping scale is read at the length it chooses, and the interval is scored against the truth.

What three block-length rules deliver in coverage. The coverage of the t interval with its scale read from a block at every starting row, when the block length is chosen on each draw by each of three rules, over 8,000 draws of a first-order autoregression at 0.7 at each size. The squared-error rule derived for disjoint blocks covers 91.3%, 92.7%, 93.6% at 120, 240, 480 rows; the same rule rebuilt for the overlapping scale covers 91.8%, 93.0%, 93.9%; the rule that chooses the shortest length predicted to cover covers 96.1%, 95.3%, 94.6%.
Fig. 3 The coverage of the t interval on the overlapping scale when each of three rules chooses the block length on each draw, at 120, 240 and 480 rows.

The squared-error rule for disjoint blocks covers 91.3%, 92.7% and 93.6% at 120, 240 and 480 rows. Rebuilt for the overlapping scale it covers 91.8%, 93.0% and 93.9% — four tenths of a point better at the shortest series and three tenths at the longest, which is what fourteen per cent more length is worth on a curve this shape. The coverage rule covers 96.1%, 95.3% and 94.6%.

The rebuilt squared-error rule is the answer to the question the earlier essay asked, and it is a small one. Taking the overlapping scale’s variance into account moves the length in the right direction and by the right amount for the criterion it serves, and it buys almost nothing for the interval. The coverage rule closes the gap at every size — over-covering by a point at 120 rows, where the cap binds, and short by four tenths at 480.

Where the rules land, and what they cost

The block lengths three rules choose. The median block length each rule chooses across draws, with the tenth and ninetieth percentiles. The squared-error rule for disjoint blocks chooses a median of 13, 17, 22 rows at 120, 240, 480; rebuilt for the overlapping scale, 15, 19, 25; aimed at coverage, 40, 59, 63, between 38 and 40, 47 and 73, 53 and 74.
Fig. 4 The median block length each rule chooses, with the tenth and ninetieth percentiles across draws, at each series length.

The squared-error rules choose median lengths of 13, 17 and 22 rows, and 15, 19 and 25, a little short of their true-persistence values because the sample’s lag-one autocorrelation is biased towards zero: it averages 0.668 at 120 rows where the truth is 0.7. The coverage rule chooses 40, 59 and 63. At 120 rows its median is the cap of a third of the series, and its tenth percentile is 38: the prediction never quite reaches 94.5% on a series that short, so the rule asks for the longest blocks it is allowed.

The ratio between the two families is the measurement worth carrying. The length an interval needs is between two and a half and three times the length a variance estimate needs, at every size here, and the gap does not close as the series grows. The length nobody has found the best length for a variance estimate on each draw to wander by sixteen rows. The coverage rule’s lengths spread too — from 47 to 73 rows between the tenth and ninetieth percentiles at 240 rows — but the spread is the plug-in’s noise passed through a flat prediction, not a target that moves from draw to draw.

What three block-length rules cost in width. The average width of the same intervals. The squared-error rule for disjoint blocks gives 0.815, 0.574, 0.411 at 120, 240, 480 rows; rebuilt for the overlapping scale, 0.841, 0.585, 0.416; the rule aimed at coverage, 1.208, 0.756, 0.463. The coverage it buys is paid for in width, by 43.7%, 29.3%, 11.2% over the overlapping squared-error rule.
Fig. 5 The average width of the same intervals under each rule, at each series length.

The coverage is paid for in width. The overlapping squared-error rule gives intervals 0.841, 0.585 and 0.416 wide at 120, 240 and 480 rows; the coverage rule gives 1.208, 0.756 and 0.463. That is 43.7% wider at 120 rows, 29.3% at 240 and 11.2% at 480. The squared-error rule’s narrower intervals are narrower because they are biased: the 3.2 points of coverage they give up at 120 rows are the price of the 0.37 of width they save.

Whether that trade is worth making is not a statistical question. A protocol that promises 95% and delivers 91.8% has made a claim it cannot keep; one that delivers 96.1% at a width forty per cent larger has kept it and spent more than it needed to. At 480 rows the choice nearly disappears, since the coverage rule costs a ninth more width for seven tenths of a point of coverage. At 120 rows it is a real choice, and the squared-error rule makes it silently.

With the persistence known

Two things push the squared-error rules short of 95%, and they can be separated. One is the criterion. The other is the plug-in: the sample’s lag-one autocorrelation is biased towards zero, so every rule that reads it asks for shorter blocks than it would at the truth. Handing each rule the true persistence instead of the sample’s measures the criterion alone.

With the persistence known, the overlapping squared-error rule covers 92.1%, 93.3% and 94.0% at 120, 240 and 480 rows — four tenths of a point better than with the plug-in at 120 rows, a third at 240 and almost nothing at 480 — and still short of 95% at every size. The coverage rule covers 96.3%, 95.7% and 94.8%. So the plug-in’s bias costs each rule at most a few tenths of a point, and the criterion costs the squared-error rule nearly three points at 120 rows. The count or the length found that an interval’s coverage tracks the block length and its width tracks the count; a criterion that weighs the length against the scale’s variance is weighing coverage against something the interval does not spend coverage on.

What the rule is aimed at

The coverage rule has a target, and the target is a choice. It was set at 94.5% rather than 95% for a reason the predicted coverage figure shows: at 120 rows the prediction never reaches 95% at any length a third of the series allows — its highest is 94.2%, at forty rows — so a rule aimed at 95% simply asks for the longest blocks it is permitted on every draw. Aiming lower makes the rule’s choice informative, and the cost of aiming lower can be read directly.

At a target of 93% the rule covers 92.2%, 92.8% and 93.2% at the three sizes, with widths of 0.877, 0.581 and 0.405 — a squared-error rule in all but name, choosing median lengths of 17, 18 and 19 rows. At 94% it covers 93.9%, 93.8% and 94.1% at widths of 1.055, 0.637 and 0.430. At 95% it covers 96.3%, 97.0% and 97.5%: the prediction never reaches its target at any of the three sizes, the cap binds on every draw, and the interval is wider than it needs to be. Between 94% and 94.5% the rule moves from just short of the promise to just past it, and the width moves with it.

So the target is a dial, and it sets a trade between coverage and width that the squared-error rule also makes — only the squared-error rule makes it at a point nobody chose, while this one makes it at a point the protocol writes down. Bias is not the whole of it made the same point about windows from the other side: a rule tuned to one quantity reads as a verdict about another, and the honest version names what it was tuned to.

Why the disjoint scale could not do this

It is worth seeing why the coverage rule needs the overlapping scale. At 120 rows it asks for blocks of about forty rows. On disjoint blocks that leaves three, a t multiplier on two degrees of freedom of 4.30, and on this essay’s draws the disjoint interval at three blocks of thirty-two is 1.571 wide — the cell the interval with no resampling in it first measured. The coverage rule’s overlapping interval at forty rows is 1.208 wide on average, and at thirty-two rows 1.056. Long blocks on a disjoint scale are paid for twice — once in the bias they remove, which is the point, and once in degrees of freedom the disjoint count throws away. The overlapping scale removes the second payment, and that is what makes a length chosen for coverage affordable at all on a short series.

The earlier essays in this field kept finding that the decision that mattered was the block length and that nobody could make it well from 120 rows. Both halves survive. The block length is still the decision, and the plug-in still reads a persistence that is biased low. What changes is the target: aimed at the interval’s coverage rather than the scale’s error, even a biased plug-in lands within about a point of the promise, because the prediction it feeds is flat near its target and the rule only has to clear it.

How the numbers were made, and where they stop

Every coverage and width is over eight thousand draws per cell, with the disjoint and overlapping scales read off the same series. Every share of the long-run variance is exact, from the autocovariances, with nothing drawn. The coverage rule’s predicted coverage uses the exact share at the sample’s ρ^\hat\rho, tabulated on a grid of persistence values and interpolated, and the scale’s nominal degrees of freedom — the three-halves rule, which the essay on every starting row found close to honest at long blocks and too generous at short ones, where this rule never goes.

Not measured: any law other than a first-order autoregression, whose exact share the coverage rule leans on, and any persistence other than 0.7. Under weaker dependence every rule’s length shrinks and the triangle’s bias with it, and the two families should converge; under long memory the observations that repeat each other do so for longer than any block of a short series can hold, and the coverage rule’s cap would bind everywhere.

Still open: the share the rule assumes

The coverage rule reads its prediction from the exact share a first-order autoregression implies at the sample’s ρ^\hat\rho. A series that is not a first-order autoregression has a different triangle, and the share the rule computes is then the wrong one — too generous where the true dependence decays more slowly than geometrically, too cautious where it stops sooner. The next measurement is a rule that reads the share from the sample instead: the attenuation is ∑k(1−∣k∣/ℓ)γ(k)/∑kγ(k)\sum_k (1 - |k|/\ell)\gamma(k) / \sum_k \gamma(k), and both sums can be estimated from sample autocovariances out to some lag, at the price of a second choice that resembles the first. Whether that rule keeps the coverage rule’s accuracy on a moving average and under a break in the persistence — the two laws on which a geometric share is most wrong in opposite directions — and what its own lag choice costs, is the measurement that would say whether a length chosen for the interval survives a series nobody has modelled.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Bias-varianceBlock countBlock lengthClosed formConfidence intervalCoverageDegrees of freedomDependenceInterval widthLong-run varianceMean squared errorPlug in estimate