The count or the length
Worth reading first: Where the bootstrap lies.
The essay that priced studentising explains its own headline in one sentence: a resample of rows at block length holds whole blocks, so the variance that studentises it is a sample variance of that many numbers — sixty at the shortest length on its grid and two at the longest — and a whose denominator is that noisy has tails no interval built from it can be narrow.
Read that sentence again with the sample size in mind. At one sample size the block length and the block count are one number read two ways. Everything that essay attributes to the length is equally an attribution to the count, and nothing it measures can tell them apart.
What one sample size cannot separate
The two quantities are related by , so fixing makes a function of alone. A finding of the form long blocks make the interval wide and a finding of the form few blocks make the interval wide are then the same finding, and the two recommend opposite things: the first says use shorter blocks and the second says get more rows.
They also predict different things about a series with more of them. If the length is what matters, a study of 480 rows at blocks of 32 is in the same trouble as one of 120 rows at blocks of 32. If the count is what matters, the first has fifteen blocks and the second has three, and they are nothing like each other.
Separating them costs a second sample size and a third. Doubling and holding doubles ; doubling and doubling holds fixed. So a grid whose sample sizes double and whose block lengths double contains both paths at once, and they leave from the same cell.
The grid here is against , under both block windows: twenty-four cells, each of 240 draws with 200 resamples apiece. Nothing chooses anything. The four rules the field this extends reads its table through are absent, because a rule that picks a longer block at a larger sample size moves both dials at once, which is the confusion rather than a way out of it.
Two paths from the same cell
Fifteen blocks of eight rows at is the shared starting point. From it, one path holds the length at eight and lets the sample size quadruple, so the count runs 15, 30, 60; the other holds the count at fifteen and lets the length run 8, 16, 32 over the same three sizes.
Along the fixed length, the studentised interval’s width against the percentile one falls: , , under the rectangular window and , , under the taper. Along the fixed count it barely moves: , , and , , .
Averaged over the two windows the first path travels −0.1280 and the second −0.0626. The penalty is a penalty of having few blocks; the block length is only how a sample size is turned into a count, and quadrupling the sample size while holding the count buys almost nothing.
The whole grid says it more plainly than the two paths do, because the two paths are four cells each and the grid is twenty-four. Every cell at fifteen blocks reads between and , across three sample sizes and three block lengths. Every cell at blocks of 32 reads between and , because those cells hold three, seven and fifteen blocks. One of those two numbers is a property and the other is a coincidence of the grid.
What is the same in every cell
A grid that moves two things has to hold everything else, and it is worth listing what is held, because a sweep across three sample sizes has more ways of quietly changing than a sweep across four block lengths does.
One resampling per draw, three intervals off it. At each cell every draw produces one set of two hundred resamples, and the normal, percentile and studentised intervals are all read from it. Three intervals from three resamplings could differ because of three sets of resamples; three from one cannot, and that is the discipline the field that read one resampling three ways established and this grid inherits.
One variance functional, on the series and on every resample. The scale is times the sample variance of the non-overlapping block means, applied identically to the original series and to each resample. Studentising with two different estimators is the commonest way to make a bootstrap- worse than the percentile interval it replaces, and it would show up here as an artefact that moved with — which is one of the two dials the whole grid is trying to read.
And the degenerate resamples stay rare. A resample whose blocks happen to agree has a near-zero scale and an enormous , and enough of those would make the studentised quantiles unusable at exactly the cells with fewest blocks, which is where the finding lives. Across all twenty-four cells the share is 0.0004% — four resamples in a million — so the width at three blocks is the degrees of freedom rather than a handful of divisions by nothing.
What is not held is the dependence. Every cell draws from the same first-order autoregression at a coefficient of 0.7, so a longer block at a larger sample size is a longer block on the same series rather than on a more strongly dependent one. That is what makes the number of observations that repeat each other a constant of the grid and the block length a free choice against it.
And there is a closed form waiting for it
If the penalty is about how many numbers a variance is estimated from, it should be predictable from that number and from nothing else — and there is an obvious candidate with no simulation in it. A variance built from block means carries degrees of freedom, so the multiplier a two-sided 95% interval needs is the Student- 97.5% point on rather than 1.96, and the ratio of the two is
That is 2.1953 at three blocks, 1.2484 at seven, 1.0943 at fifteen, 1.0435 at thirty, 1.0209 at sixty and 1.0103 at a hundred and twenty. It contains no sample size, no block length, no window and no resampling.
Across the twenty-four cells, accounts for 0.9911 of the variation in the log of the measured width penalty. The block length on its own accounts for 0.2927, and the sample size for 0.2493.
Two things follow, and the second is the more useful.
It is a floor rather than a fit. Every one of the twenty-four cells is wider than the multiplier alone would make it. The excess runs from a factor of 1.7984 at three blocks down to 1.0049 at a hundred and twenty — monotone in the count, which is what a resampling distribution converging on its own limit looks like. The multiplier is what is left of the penalty when the resampling has stopped contributing.
And it is available before any resampling is done. A practitioner choosing a block length knows before drawing a single resample, and of it is the smallest width penalty studentising can cost. At blocks of 32 in 120 rows that is already in closed form, and the measured cost is . Nothing about that needs a simulation to notice.
The coverage tracks the other dial
The width is half the trade. The other half is whether the interval covers, and the expectation going in was that it would follow the same dial: an interval that is wide because it has few blocks should be short of its promise because it has few blocks.
It does not. Along the fixed length, with the count running to sixty, the studentised interval gains 0.00 points of coverage averaged over the two windows — 88.8%, 89.2%, 87.9% under the rectangle and 84.2%, 80.4%, 85.0% under the taper, which is a wobble rather than a trend. Along the fixed count, with the length running to 32, it gains 4.79 points: 88.8%, 91.7%, 92.5% and 84.2%, 86.7%, 90.0%.
The percentile interval does the same thing more strongly still, gaining 4.17 and 6.25 points along the fixed-count path against 2.50 and 3.75 along the fixed-length one. Over the whole grid the block length accounts for 0.6677 of the studentised interval’s coverage and 0.3840 of the percentile interval’s; the block count accounts for 0.4070 and 0.0874.
The sample size itself explains almost none of either column: 0.2493 of the width penalty and 0.0004 of the studentised coverage. That is the reading that makes the two dials worth separating rather than merely worth naming. A larger sample is not a better sample here in any direct sense — it is an opportunity to have more blocks or longer ones, and which of the two it is spent on decides which column improves. Spending it on more blocks narrows the interval and leaves the coverage where it was; spending it on longer blocks improves the coverage and leaves the width penalty where it was. The two are not alternatives a practitioner is usually offered, because a rule that chooses the length makes the choice on other grounds entirely and never states which of the two it has bought.
So the two columns of this table are about different things, and the sentence that explained the width does not explain the coverage. It should not have been expected to. A longer block carries more of the series’ dependence into each resample, and a resample of fifteen long blocks is a better model of the series than a resample of fifteen short ones; the variance that studentises it is a sample variance of fifteen numbers either way and does not care which numbers they are.
That is a separation the earlier essays could not have made and would not have thought to look for, because at one sample size the two dials are one and the single explanation covered both columns without strain.
The reversal is not about the sample size either
The finding this field’s third essay is about is a change of sign: the tapered window covers better on the percentile interval and worse on the studentised one, at all four block-length rules. It was measured at one sample size and at lengths the rules chose.
At every one of the twelve cells of this grid, the taper wins on the percentile interval — by between 0.83 and 5.83 points — and loses on the studentised one, by between 1.67 and 8.75. Twenty-four sign agreements out of twenty-four, across a quadrupling of the sample size and an eightfold range of block lengths.
And the control behaves as it must. The normal interval, which resamples nothing and uses only the block-means variance of the original series, puts the two windows in exactly the same place at every cell — to machine precision, because the window enters only through the resampling and the normal interval does none. So the disagreement between the two windows is manufactured entirely by what is done with the resamples, which is what that essay argues and what this grid now says at twelve settings rather than at one.
Nothing reaches the promise at any size
The reading the field that measured these intervals first ends on is that not one of its cells reaches the 95% it offers. That was eight cells at one sample size, and the obvious escape route is that a hundred and twenty rows is simply a small sample.
It is not the explanation. Of the twenty-four cells here, not one reaches 95% on the normal interval, the percentile interval or the studentised one. The best cell in the whole grid is the percentile interval at with blocks of 16 under the taper, at 93.8%. The studentised column runs from 73.8% to 92.5%.
Quadrupling the sample size moves the shortfall and does not close it, which is what says the shortfall belongs to the resampling scheme and to the dependence it is trying to reproduce rather than to a small sample that a larger one would fix. A series at an autoregressive coefficient of 0.7 has a long-run variance several times its marginal one, and the error that no window repairs is the same error here at four times the rows.
What this says about the number the field reported
The width penalty this field reports is 2.0938, averaged over the eight cells its four rules choose. The grid here averages 1.3795 over twenty-four cells, and the difference is not a disagreement.
The oracle rule picks a block length of 18.1 rows under the rectangle and 20.7 under the taper, which at 120 rows leaves six and five whole blocks. At those cells the measured penalty is and . The rule of thumb picks 4.0 rows — thirty blocks — and the penalty there is and .
So the 2.09 is not a fact about studentising. It is a fact about the block counts that field’s rules happened to leave, and every one of those rules is choosing a length for a reason that has nothing to do with how many blocks it leaves: the oracle picks the length whose resampled quantile is nearest the truth, and the plug-in picks from the estimated autocorrelation. Neither has a term in it for the degrees of freedom the choice costs a studentised interval.
That is the practical form of this essay. A rule choosing a block length is also choosing a , silently, and the choice is worth a factor of four between the shortest length on that field’s grid and the longest. A rule that priced the second alongside the first would be a different rule, and nothing in that field or this one has built it.
What a reader should carry away
Three sentences, and the first two are the ones a practitioner can act on before running anything.
The width penalty is or worse, and it is known in advance. A block length chosen for any reason at all fixes a number of whole blocks, and the Student- multiplier on one fewer degrees of freedom is the least that studentising can cost against a percentile interval on the same resamples. At three blocks that floor is already and the measured cost is .
The coverage is a different question with a different answer. It improves with longer blocks at a fixed count and not with more blocks at a fixed length, which is the opposite arrangement, and it means the two halves of the trade cannot be improved by the same change.
And neither of them reaches the promise. Every cell of this grid is short of 95%, at four times the rows the earlier essays used, which is the reading the field that first measured these intervals ends on and which four times the data does not repair.
What this opens and does not measure
Three things, each named where it could have been run.
A rule that prices the count is not built here. The obvious one is a length chosen to minimise something like the interval’s expected width subject to its coverage, which would trade against how much of the dependence a block of carries. The two terms are both measured above and the rule that combines them is not, because a rule needs its own sweep and its own oracle to be scored against, and scoring it on this grid would be scoring it on the grid it was fitted to.
The degrees-of-freedom correction is not applied. If is the floor, an interval that used the quantile on degrees of freedom rather than the resampled one would be the natural repair, and it is one line. It is not run here because it is a fourth interval and this sweep is about separating two dials on the three it already has — but it is the cheapest thing on this list and the one most likely to change a recommendation.
And the sample sizes double rather than range. Three sizes a factor of two apart is enough to separate a count from a length and not enough to say anything about the rate at which either effect decays. The measured excess over the closed form falls from 1.7984 to 1.0049 across the counts on this grid, which looks like convergence and is four points; whether it is or something slower is a question a grid built for it would answer and this one cannot.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The reversal that was the instrument's — both name block bootstrap, block length, coverage, dependence, long-run variance, resampling, tapering
- A block weighted inside itself — both name block bootstrap, closed form, dependence, long-run variance, resampling, tapering
- Measuring a variance rather than a quantile — both name block bootstrap, block length, closed form, long-run variance, resampling, tapering
- The gap a sample shows — both name block bootstrap, block length, closed form, long-run variance, resampling, tapering
- A lag the sample has less of — both name closed form, degrees of freedom, dependence, long-run variance, tapering
- A taper and a critical value — both name block bootstrap, dependence, long-run variance, resampling, tapering
Named objects
A flat tag is an object no other essay names yet.
Block bootstrapBlock countBlock lengthClosed formConfidence intervalCoverageDegrees of freedomDependenceEstimated varianceInterval widthLong-run variancePercentile intervalResamplingStudentised bootstrapTapering