What studentising costs
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
A repair is worth its cost or it is not, and both halves have to be measured. The construction is one extra variance per resample. What it buys and what it costs are two columns of the same table.
It costs a factor of 2.09 in width. It buys 0.46 points of coverage.
The two columns
Over three hundred draws, cell by cell, the studentised interval’s width as a multiple of the percentile interval’s and the coverage it gains:
| window and rule | width | coverage gained |
|---|---|---|
| the rectangle, a protocol length | ×1.171 | +5.00 |
| the rectangle, the rule of thumb | ×1.070 | +2.67 |
| the rectangle, the plug-in | ×1.325 | +7.67 |
| the rectangle, the best length | ×4.369 | +3.67 |
| the taper, a protocol length | ×1.204 | −4.67 |
| the taper, the rule of thumb | ×1.140 | −6.00 |
| the taper, the plug-in | ×1.428 | −1.67 |
| the taper, the best length | ×5.042 | −3.00 |
Averaged, ×2.0938 for +0.46 points.
What is being compared
A word on the two intervals being set against each other, because they are not two estimates of one thing.
The percentile interval is the 2.5% and 97.5% points of the resampled means, read directly. Its width is the spread of the resampled distribution and nothing else, so it is as wide as the bootstrap says the estimate is uncertain.
The studentised interval is x̄ minus the resampled t-quantiles times σ̂/√n. Its width is the spread of the resampled t distribution scaled by the sample’s own standard error, so it is as wide as the bootstrap says a standardised statistic is uncertain, converted back.
Those coincide when the scale is estimated perfectly — the t-quantiles are then the mean-quantiles divided by a constant, and multiplying back gives the same interval. They part when the scale is noisy, and they part by however noisy it is.
So the width ratio in the table is a direct reading of how badly the scale is estimated, and it is the same reading as the block count, arrived at from the other side. A ratio of 1.07 says the scale is estimated well enough not to matter; a ratio of 5.04 says it is not.
That is why this field reports the ratio rather than the two widths. The earlier field’s own intervals have widths that barely move across the table, and putting a number beside them that moves by a factor of five is what makes the mechanism visible.
Where the width goes
The averages hide the shape, and the shape is the finding.
At the two rules that choose short blocks — a protocol length of eight and the rule of thumb’s four — the two intervals are within a fifth of each other: ×1.070 to ×1.204.
At the oracle’s lengths, which average 18.11 and 20.69 rows, the studentised interval is four to five times as wide.
The plug-in, which picks 12.49 and 13.73, sits between at ×1.325 and ×1.428.
So the width is a function of the block length, and it is not a gentle one.
The pairing that makes these readable
Every comparison here is between two intervals computed from the same resamples on the same draw, and that is what makes a gain of 2.67 points reportable at all.
Coverage is a proportion over three hundred draws, so its own standard error is about 2.2 points at the levels in this table. An unpaired comparison of 80.0% against 82.7% would be well inside one standard error and would say nothing.
Paired, the comparison is: on this draw, did the percentile interval cover and did the studentised one? Most draws agree — both cover or neither does — and what is left is the draws where they differ. The paired t-statistics for the eight cells are 2.82, 1.89, 4.57, 2.13, −2.89, −3.93, −1.09 and −1.62, so five of the eight separate at two standard errors and three do not.
That is the right resolution for the claim being made. The individual cells are mostly separated; the pattern — four positive under one window and four negative under the other — is what the field rests on, and eight cells all pointing the way their window predicts is a stronger statement than any one of them.
The arithmetic behind it
A resample of n rows at block length ℓ holds ⌊n/ℓ⌋ whole blocks. That is how many numbers the variance that studentises it is a sample variance of.
On a hundred and twenty rows, across the grid:
60 blocks at ℓ = 2, 30 at 4, 15 at 8, 10 at 12, 6 at 20, 3 at 32, and 2 at 48.
A tail of fewer than ℓ rows is dropped rather than counted as a short block, because a block mean over three rows and one over twenty have different variances and averaging them would bias the scale by an amount that grows with how badly ℓ divides n. At ℓ = 48 that discards twenty-four of the hundred and twenty rows, which is the price of an unbiased scale on two blocks.
At the top of the grid the scale is a sample variance of two numbers. A t-statistic whose denominator is estimated from two or three observations has tails no interval built from it can be narrow — that is the ordinary behaviour of a t on very few degrees of freedom, and it is not a defect in the resampling.
The rules that choose long blocks are exactly the rules that leave the scale resting on two or three numbers, and the oracle chooses long blocks.
Why the oracle chooses long blocks
The two cells with the four- and five-fold width are the oracle’s, and it is worth saying why the oracle is where it is, because it is not an arbitrary choice.
The oracle here is the earlier field’s: the block length whose resampled 95% point is nearest the finite-sample truth on that draw. It averages 18.11 rows under the rectangle and 20.69 under the taper.
Long blocks are what a quantile wants under dependence. A short block breaks the series into too many nearly independent pieces and loses the dependence the bootstrap is supposed to preserve, which shrinks the resampled distribution and puts its 95% point too low; a long block keeps the dependence and pays for it in variance. The field that separated bias from spread is the measurement of that trade.
So the length that is best for the quantity being read is a length at which a resample holds six blocks, or three, or two. The rule that picks the best block length is the rule that leaves the fewest numbers to estimate a scale from, and that is not a coincidence — it is the same trade-off seen from the other end.
A practitioner is not choosing the oracle, of course. But the plug-in picks 12.49 and 13.73, which is ten blocks and eight, and its width penalty is already 33% and 43%.
The width ratio is a t multiplier, until it is not
If the widening is the ordinary behaviour of a t on few degrees of freedom, the ratio should be with b the block count. That prediction can be checked against the table and it holds at one end and fails at the other.
At the rule of thumb’s four rows a resample holds 30 blocks, so t₂₉/1.96 = 1.043 against a counted 1.070. At the protocol length of eight it holds 15, so t₁₄/1.96 = 1.094 against 1.171. Both within three points.
At the oracle’s eighteen rows a resample holds six, so t₅/1.96 = 1.31 — against a counted 4.37, more than three times as large.
So the blow-up at long blocks is not the ordinary t effect, and the essay’s explanation covers the first two cells and not the last two. What separates them is that the counted ratio is a mean over draws, and at six blocks the width’s distribution is heavily skewed: the 2.5% and 97.5% points of two hundred and fifty resampled t values, on six blocks, are themselves very variable, and a mean over draws of a quantity with a long right tail sits far above its typical value.
The t model predicts the typical width and the table reports the average one, and the two agree until the block count falls far enough for the difference between a mean and a median to matter. That is the same reason this collection reports medians for widths everywhere else, arriving here as a three-fold discrepancy in an explanation.
What each cell pays per point of coverage
Dividing the coverage gained by the extra width bought says which cells the repair is efficient in, and the spread is a factor of thirty.
Under the rectangular window the four cells buy 29.2, 38.1, 23.6 and 1.1 points of coverage per unit of relative width. The first three are the short and middling block lengths; the last is the oracle’s.
The repair is efficient exactly where the block length is chosen worst and ruinous where it is chosen best. A rule of thumb picking four rows gets its coverage for a seven per cent widening; the oracle’s eighteen rows pay a three-hundred-and-forty per cent widening for less coverage than either.
That inverts what a practitioner would assume about the two decisions. Choosing the block length well and studentising are not two independent improvements to be applied together: the better the length, the fewer blocks a resample holds, the noisier the scale, and the more the repair costs. The two repairs are substitutes at best and antagonists at worst, and the field’s own oracle is the cell where they conflict most.
Which is not the degenerate tail
There is a different explanation available and it is worth ruling out.
A bootstrap-t is known to fail when a resample’s scale comes out near zero, because then its t is enormous and one such resample can set a quantile. If that were what is happening, the width would be a property of a handful of pathological resamples rather than of the estimator.
It is measured. Across the whole sweep — three hundred draws, seven lengths, two windows, two hundred and fifty resamples each — the share of resamples with a scale too small to studentise is 0.1996%, and those are dropped and counted rather than clamped.
One in five hundred is not what produces a factor of five. The width is the block count, and the block count is arithmetic.
What a width is, on this scale
The widths quoted are in the units of the data, and it is worth having a sense of them because a factor of five is easier to read against an absolute number.
At the protocol length of eight, the three intervals are 0.668 wide for the normal one, 0.626 for the percentile one and 0.733 for the studentised one, under the rectangular window. At the oracle’s length the percentile interval is 0.627 and the studentised one is 2.738.
So the percentile interval barely changes width with the block length — 0.478 to 0.665 across the whole table — and the studentised one runs from 0.580 to 3.233.
That asymmetry is itself informative. A percentile interval is the spread of the resampled means, and the resampled means’ spread does not blow up at long block lengths; it grows a little, because a longer block preserves more dependence and therefore more variance in the mean. A studentised interval is that spread divided by an estimated scale, and dividing by something estimated from two numbers is what produces a three-unit interval on a series whose mean has a standard error of about a sixth of a unit.
The numerator is well behaved and the denominator is not.
What the coverage does
The other column is the one the repair was for, and it splits by window rather than by rule.
Under the rectangular window studentising helps at every rule: 83.7% to 88.7%, 80.0% to 82.7%, 84.7% to 92.3%, 81.3% to 85.0%. Four gains, from 2.7 to 7.7 points, and the largest cell reaches 92.3%.
Under the tapered window it hurts at every rule: 89.0% to 84.3%, 81.7% to 75.7%, 89.7% to 88.0%, 85.3% to 82.3%. Four losses, from 1.7 to 6.0 points, and the worst cell falls to 75.7% — below anything the earlier field reported on any interval.
Averaged over the eight, the two directions nearly cancel and the mean gain is +0.46 points.
That the repair helps one window and hurts the other is the field’s third finding, and it is a stronger result than the width. Here it is enough to say that the mean gain is small because it is a mean over two opposite effects rather than because the repair is uniformly weak.
What a wider interval would have bought anyway
The comparison a reader should want is not “is the studentised interval better?” but “is it better than widening the percentile interval by the same amount?”, and that comparison is available.
A percentile interval widened by a factor of 2.09 would cover a great deal better than 0.46 points more. The eight cells run from 80.0% to 89.7%, so their shortfalls against 95% are five to fifteen points; a normal interval with twice the width would have coverage in the high nineties at every one of them.
It is worth checking that against the table rather than leaving it as an argument. The cell with the largest width penalty — the taper at the oracle’s length, ×5.042 — loses three points of coverage. The cell with the largest coverage gain — the rectangle at the plug-in, +7.67 points — has a width penalty of ×1.325, the third smallest in the table. The two columns are not merely weakly related; across the eight cells they are close to unrelated.
So the studentised interval is not buying its coverage with its width. If it were, the trade would be a bad one but a comprehensible one. What it is doing is moving the coverage around by a few points in different directions while being twice as wide, which is a worse outcome than a simple widening would have produced.
That is not an argument for widening intervals arbitrarily — a width chosen to fix coverage is a width chosen after seeing the coverage, which no practitioner has. It is an argument that the studentised interval’s extra width is not where its coverage comes from, and therefore that the width is a cost with no matching benefit.
What the normal interval says about the trade
The third interval on the table is the one nobody computes, and it prices the whole exercise.
A normal interval on the same block-means variance covers 87.3%, 81.7%, 89.0% and 82.0% under the rectangle and 87.3%, 81.7%, 88.7% and 82.3% under the taper. Its widths are 0.668, 0.565, 0.717 and 0.695 — between the percentile interval’s and the studentised one’s at the short lengths, and far below the studentised one’s at the long ones.
Against it:
Under the rectangle the studentised interval is nearer 95% at every rule — 88.7% against 87.3%, 82.7% against 81.7%, 92.3% against 89.0%, 85.0% against 82.0%. So there the bootstrap is doing something a normal interval is not.
Under the taper it is not, at any rule. 84.3% against 87.3%, 75.7% against 81.7%, 88.0% against 88.7%, 82.3% against 82.3%.
Half the table is a bootstrap that beats the interval anybody could have written down without one, and half is a bootstrap that does not. That is a poor return for two hundred and fifty resamples per cell, and it is a comparison the earlier fields never made because they never put a non-resampled interval on the table.
Where a bootstrap-t would work
It is worth saying what this measurement is not, because studentising is a good idea in most of the places it is used.
The scale has to be estimable. In the independent case a bootstrap resample of n rows has n observations to estimate a variance from, and studentising is famously worth doing. Under dependence, with a block bootstrap, the effective count is the number of blocks, and a resample designed to preserve dependence is a resample that has few independent pieces by construction.
And the block length has to be short relative to the sample. At ℓ = 8 on a hundred and twenty rows there are fifteen blocks and the width penalty is 17% and 20%; at ℓ = 48 there are two and it is four- and five-fold. A sample ten times longer at the same block lengths would have ten times the blocks and the penalty would nearly vanish.
So the finding here is specific: on a hundred and twenty rows, at the block lengths four sensible rules choose, the scale a bootstrap-t divides by is estimated from too few blocks to be worth dividing by. A longer sample is a different measurement, and this field does not make it.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The interval with no resampling in it — both name block bootstrap, block length, confidence interval, coverage, dependence, estimated variance, interval width, long-run variance, percentile interval, resampling, studentised bootstrap
- The error no window repairs — both name block bootstrap, block length, dependence, effective sample size, long-run variance, monte carlo, resampling
- The reversal that was the instrument's — both name block bootstrap, block length, coverage, dependence, long-run variance, monte carlo, resampling
- An ordering that depends on the rule — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling
- Measuring a variance rather than a quantile — both name block bootstrap, block length, long-run variance, monte carlo, resampling, standard error
- The length nobody has — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling
Named objects
A flat tag is an object no other essay names yet.
Block bootstrapBlock lengthConfidence intervalCoverageDependenceEffective sample sizeEstimated varianceInterval widthLong-run varianceMonte CarloPercentile intervalPivotal quantityResamplingStandard errorStudentised bootstrap