The interval, studentised

The interval with no resampling in it

Replace 1.96 in a normal interval on the block-means variance with Student's t on one fewer degrees of freedom than there are whole blocks, and resample nothing. Across twenty-four cells it covers at least as often as the studentised bootstrap interval at every one, by 0.42 to 10.42 points; it is narrower wherever seven blocks or fewer are left; and at fifteen blocks of 32 it covers 95.0%, which no resampled interval on the grid reaches.

Worth reading first: Where the bootstrap lies.

The essay that separated the block count from the block length ended on three things it could have run and did not, and it ranked them. The cheapest, and the one it judged most likely to change a recommendation, was an interval that took the block count’s degrees of freedom at face value: the normal interval on the block-means variance, with its 1.96 replaced by Student’s tt on one fewer degrees of freedom than there are whole blocks. It is one line. It needs no resample.

Run on the same grid and the same draws, that line covers at least as often as the studentised bootstrap interval at all twenty-four cells, by between 0.42 and 10.42 points. Where seven or fewer whole blocks are left it is also narrower, down to 0.614 times the studentised interval’s width at three blocks. And at fifteen blocks of 32 rows it covers 95.0% of its draws, which no resampled interval reaches at any cell of the grid.

Studentising was supposed to be the repair for an interval that did not know how noisy its own scale was. On this grid the repair is available without the resampling, and the resampling, where it differs from the repair, makes the interval wider and covers less.

One line, and no resampling

Every cell of the grid draws a first-order autoregression at a coefficient of 0.7, of nn rows, and asks for a 95% interval for its mean. Cut the series into k=n/k = \lfloor n/\ell \rfloor whole blocks of length \ell, and let sB2s^2_B be \ell times the sample variance of the block means — the scale every interval on the grid already uses, on the series and on every resample alike. The normal interval is xˉ±1.96sB/n\bar x \pm 1.96\, s_B/\sqrt n. The interval here is

xˉ  ±  t0.975,  k1  sB/n,\bar x \;\pm\; t_{0.975,\;k-1}\; s_B/\sqrt n,

which is the batch-means interval of the simulation literature under another name, and it has one justification that is not an approximation. If the block means were independent and normal, sB2s^2_B would be a scaled chi-square on k1k - 1 degrees of freedom and the ratio would be exactly Student’s tt. The block means of a dependent series are neither, and the grid is where that matters.

Two things are true of it by construction, and both are checked rather than assumed. Its width on every draw is the normal interval’s width times t0.975,k1/1.96t_{0.975,k-1}/1.96, the multiplier the earlier essay called τ(k)\tau(k), so the ratio of the mean widths across a cell is that multiplier — to within 8.9e-16 at every cell. And it consumes no random numbers, so under the rectangular window and the taper it is the same interval, and adding it to the sweep changes none of the numbers the grid already reported. That is two routes to one number in the most literal form: the resampled intervals and this one share every draw, every block mean and every scale, and differ only in where the multiplier comes from.

Four intervals at 3 blocks of 32 rows. What four 95% intervals for the mean of a first-order autoregression at 0.7 cover, and how wide they are on average, at 120 rows cut into 3 whole blocks of 32, over 240 draws with 200 resamples each under the rectangle. The normal interval, the block-means variance with 1.96, covers 82.1% at a width of 0.708. The percentile interval covers 82.1% at 0.632 and the studentised one 90.4% at 2.495. The fourth resamples nothing: it is the normal interval with 1.96 replaced by Student's t on 2 degrees of freedom, and it covers 94.2% at 1.555, 0.62 times the studentised interval's width.
Fig. 1 Four intervals at the grid’s harshest cell, three whole blocks of 32 rows at 120 rows, with each one’s average width beside its coverage. The interval that resamples nothing covers most, at a width between the percentile interval’s and the studentised one’s.

The grid’s harshest cell is 120 rows in blocks of 32, which leaves three whole blocks and a variance estimated from three numbers. The normal interval covers 82.1% there at an average width of 0.708, and the percentile interval covers the same 82.1% at 0.632. The studentised interval covers 90.4%, and pays for it with a width of 2.495, three and a half times the normal interval’s.

The tt interval covers 94.2% at a width of 1.555. It is wider than the normal interval by τ(3)=2.195\tau(3) = 2.195, exactly, and narrower than the studentised one by more than a third. Three blocks is where studentising was supposed to earn its cost, because it is where a normal multiplier is most wrong, and three blocks is where the closed multiplier beats it by the widest margin on both columns at once.

At every cell of the grid

One cell could be a coincidence of three blocks. The grid has twenty-four.

Every cell on or above the line. The coverage of a t interval on the block count against the studentised bootstrap interval's, at all 24 cells of the grid — three sample sizes, four block lengths, two windows — on the same 240 draws a cell. No point is below the diagonal: the t interval covers more often at every one of them, by 0.42 to 10.42 points and 4.11 on average. It reaches the 95% line at 2 cells, 15 blocks of 32 at 480 rows under both windows, where no resampled interval on the grid reaches it at all. The t interval resamples nothing, so under the two windows it is the same interval and each of its values appears twice, against two different studentised ones.
Fig. 2 Every cell of the grid, with the studentised bootstrap interval’s coverage across and the t interval’s up, on the same draws. No cell falls below the diagonal, and only the t interval reaches the dashed 95% line.

At every one of the twenty-four cells the tt interval covers more often than the studentised bootstrap interval. The margin runs from 0.42 points, at 480 rows in blocks of 4 under the rectangle, to 10.42, at 240 rows in blocks of 8 under the taper, and averages 4.11. No cell is a tie.

The largest margins are all under the taper, and that is not a property of the tt interval, which has no window. It is the reversal the field found earlier: the tapered window helps the percentile interval and hurts the studentised one, at every cell. The tt interval sits at one height for both windows, so the taper’s damage to the studentised interval shows up as extra distance above the diagonal. Under the rectangle alone the margin runs from 0.42 to 3.75 points; under the taper, from 4.58 to 10.42.

What the picture does not show is a region where resampling wins. A reader looking for the cells where the bootstrap earns its keep — few blocks, where a normal approximation is worst, or many, where resampling has the most to work with — finds the tt interval ahead at both ends and in between.

Where the width goes

Coverage bought with width is not a repair, and the earlier essays in this field were careful to read the two columns together. So the width is set beside it, cell by cell.

Narrower exactly where studentising is dear. How wide the t interval is against the studentised bootstrap interval, at every cell of the grid, against the number of whole blocks the cell leaves. At three blocks it is 0.614 to 0.623 times as wide; at seven, 0.925 to 0.971; from fifteen blocks up the two are within 2.7% of each other, from 0.993 to 1.027. The studentised interval pays for a noisy variance twice, once in the degrees of freedom and again in resampled quantiles of a ratio with that noisy variance underneath; the t interval pays the first and not the second.
Fig. 3 The t interval’s average width against the studentised interval’s, at every cell, by the number of whole blocks the cell leaves. Below the dashes it is narrower: at every cell of seven blocks or fewer, and within a few per cent of level from fifteen up.

At three blocks the tt interval is 0.614 to 0.623 times the studentised interval’s width, and at seven it is 0.925 to 0.971. From fifteen blocks up the two are within 2.7% of each other, from 0.993 to 1.027 — so at the cells where the tt interval is slightly wider it is wider by a couple of per cent and covers more by a point or more, and at the cells where it is much narrower it also covers more.

The earlier essay measured the studentised interval’s width against τ(k)\tau(k) and found it above that floor at every cell, by a factor of 1.80 at three blocks falling to 1.005 at a hundred and twenty. The tt interval sits exactly on the floor. The reading that fits both is that the studentised interval pays for a noisy variance twice: once in the degrees of freedom, which any honest interval must pay, and again in the resampled distribution of a ratio whose denominator is resampled from the same few blocks — a resample of three blocks drawn with replacement from three blocks repeats a block more often than not, and its variance is noisier than a chi-square on two degrees of freedom. That second payment is what the excess over τ\tau measures, and nothing in the coverage column says it bought anything.

It is worth being plain that this is a reading and not a derivation. What is measured is that the tt interval’s width is τ(k)\tau(k) times the normal interval’s on every draw, that the studentised interval’s is more, and that the extra width comes with less coverage rather than more.

What the multiplier alone is worth

The tt interval and the normal interval share everything except the multiplier, so the difference between their coverages is what the degrees of freedom are worth at each cell, and nothing else.

At three blocks that difference is 12.08 points: the normal interval covers 82.1% and the tt interval 94.2%, so the multiplier alone closes 94% of the normal interval’s shortfall. The studentised interval, starting from the same scale on the same draws, closes 65% of it under the rectangle and 52% under the taper. At seven blocks the multiplier is worth 5.00 points at both cells that leave seven, and at a hundred and twenty blocks it is worth 1.25, closing 9% of a shortfall of 13.75 points.

That is the shape a correction for estimating a variance from kk numbers ought to have. Where there are few blocks it is nearly the whole of the repair. Where there are many, the normal interval is short for a reason that has nothing to do with how many numbers the variance came from, and the multiplier correctly declines to pretend otherwise. The studentised interval does not have that shape. At three blocks it closes less of the shortfall than the multiplier does and adds more than twice as much width to get there, and at a hundred and twenty blocks under the taper it covers 77.5%, 3.75 points below the normal interval it was built to improve on.

Fifteen blocks of thirty-two

Four intervals at 15 blocks of 32 rows. What four 95% intervals for the mean of a first-order autoregression at 0.7 cover, and how wide they are on average, at 480 rows cut into 15 whole blocks of 32, over 240 draws with 200 resamples each under the rectangle. The normal interval, the block-means variance with 1.96, covers 91.7% at a width of 0.392. The percentile interval covers 87.5% at 0.374 and the studentised one 92.5% at 0.423. The fourth resamples nothing: it is the normal interval with 1.96 replaced by Student's t on 14 degrees of freedom, and it covers 95.0% at 0.429, 1.01 times the studentised interval's width.
Fig. 4 The same four intervals at 480 rows in blocks of 32, fifteen whole blocks. The t interval covers 95.0% at nearly the studentised interval’s width; every resampled interval on the grid falls short of that.

At 480 rows cut into fifteen blocks of 32, the normal interval covers 91.7% at a width of 0.392, the percentile interval 87.5% at 0.374, and the studentised interval 92.5% at 0.423. The tt interval covers 95.0% — 228 of 240 draws — at 0.429, 1.01 times the studentised interval’s width.

That is the only cell on the grid where any of the four intervals covers what it promises, and it is the same cell under both windows because the tt interval does not see the window. The field that first measured these intervals ended on the reading that none of them reached 95% anywhere, and the essay that tripled its sample sizes found the same at twenty-four cells. The reading survives for every resampled interval. It does not survive for the one that resamples nothing.

A standard error keeps that honest. Over 240 draws, a coverage of 95% has a standard error of about 1.4 points, so the cell is consistent with the promise rather than proof of it, and the neighbouring cells at 94.2% and 94.6% are consistent with it too. What the grid does establish is the ordering around it. The best resampled cell anywhere is the tapered percentile interval at 93.8%, inside a standard error of the promise at a cell where the tt interval covers 92.9%; the best studentised cells are at 92.5%, 1.78 standard errors short, and the tt interval covers 94.6% and 95.0% at those two cells.

The count and the length, again

The essay that separated the two dials found that the studentised interval’s width followed the block count and its coverage followed the block length. The tt interval makes the first half of that true by construction, since its width multiplier is a function of the count alone. The second half is a question it has to answer by measurement.

The t interval follows the length as well. What the t interval covers along the two paths out of the shared cell, 15 blocks of 8 rows at n = 120. Holding the length at 8 while the count runs 15, 30, 60, it covers 90.8%, 90.8%, 89.6%, a change of −1.25 points. Holding the count at 15 while the length runs 8, 16, 32, it covers 90.8%, 92.9%, 95.0%, a change of 4.17. The degrees of freedom fix its width from the count; what it covers depends on how much of the series' dependence a block of that length carries, which is the same dial the studentised interval's coverage follows. The two windows are not shown separately because this interval is the same under both.
Fig. 5 The t interval along the two paths out of fifteen blocks of eight rows at 120 rows. Letting the count run leaves its coverage where it was; letting the length run carries it to 95%.

Holding the length at eight while the count runs 15, 30, 60, the tt interval covers 90.8%, 90.8% and 89.6%. Holding the count at fifteen while the length runs 8, 16, 32, it covers 90.8%, 92.9% and 95.0%. The first path moves it by −1.25 points and the second by 4.17.

So it follows the same dial as the studentised interval, for the reason the earlier essay gave. Its width is fixed from the count, and the degrees of freedom are priced exactly. What it covers depends on whether the scale sB2s^2_B is the right scale, and that depends on how much of the series’ dependence a block of \ell rows carries. A longer block leaves less correlation between neighbouring block means for the variance to miss. The tt multiplier has nothing to say about that, and neither does resampling.

Where it is still short

Short of the promise by less, and at one cell not at all. How many points short of 95% the t interval falls at each of the 12 cells of the grid, beside how far short the studentised bootstrap interval falls under the rectangle and under the taper. The t interval's shortfall runs from 12.50 points at n = 240, 60 blocks of 4 down to nothing at n = 480, 15 blocks of 32. The studentised interval is never closer to its promise than the t interval at the same cell. At every sample size the t interval falls furthest short at blocks of 4, where a block-means variance leaves out most of the correlation between neighbouring blocks — and a multiplier chosen for the count of blocks cannot repair a scale that is too small to begin with.
Fig. 6 How many points short of 95% the t interval falls at each cell, beside the studentised interval’s shortfall under each window. It is never further short than either, and it is furthest short at blocks of four at every sample size.

The tt interval is not a repair for everything, and the cells where it fails say which problem it does not touch. Its shortfall is largest at blocks of four at every sample size: 10.42 points at 120 rows, 12.50 at 240 and 12.50 at 480. Those cells have 30, 60 and 120 blocks, so τ(k)\tau(k) is between 1.04 and 1.01 and the degrees of freedom are not the problem. The scale is.

The correlation between neighbouring block means of a series at 0.7 has a closed form, and it is 0.41 for blocks of four, 0.23 for eight, 0.10 for sixteen and 0.05 for thirty-two. A block-means variance treats those means as uncorrelated, so at blocks of four it leaves out a large positive covariance and reports a scale that is too small. At blocks of 32 the covariance it leaves out is small, which is why the tt interval’s own justification — independent block means — nearly holds there, and why its coverage reaches the promise there and nowhere else. A scale that is too small is not rescued by any multiplier chosen for how many blocks there are, which is the error no window repairs, one interval further along.

Beside it, the studentised interval’s shortfall at the same cells is 11.67 and 17.50 points at 120 rows, and 13.75 and 21.25 at 240. At no cell of the twelve is the studentised interval closer to its promise than the tt interval.

Where resampling still earns something

The claim is against the studentised interval, and it would be wrong to widen it to every resampled interval.

The percentile interval under the taper covers more than the tt interval at two cells, both at 480 rows: 90.4% against 89.6% in blocks of 8, and 93.8% against 92.9% in blocks of 16. It ties it at two more. Those margins are under a point — two draws in 240 — and the percentile interval is narrower at both, 0.324 against 0.354 and 0.366 against 0.400. That is the one corner of the grid where a resampled interval is ahead on both columns at once, and it is the tapered window at the larger sample sizes, which is where the ordering that reversed put the taper’s advantage.

Everywhere else the percentile interval is short of the tt interval, by as much as 12.08 points at three blocks under the rectangle. So the reading is not that resampling is useless for a mean; it is that the resampled interval built to fix the scale’s noise does worse than the closed interval that prices the noise directly, and that the one resampled interval that holds its own does so at a handful of cells and by a margin the draws barely resolve.

What a reader should take from it

For a mean, choose the block length for the dependence and let the count set the multiplier. The coverage follows the length, so the length is the choice that matters; the width follows the count, and Student’s tt on k1k - 1 degrees of freedom prices it in closed form before any data are resampled. The rules that choose a length choose it on other grounds, and none of them states the count it leaves.

Studentising by resampling is dominated here. On every cell of this grid the studentised bootstrap interval covers less often than the tt interval on the same draws, and where it is much wider it is also further short. That is a statement about the mean of one autoregression on one grid, and it is a strong one on that grid.

And the scale is still the problem the multiplier is not. Where blocks are short the tt interval is as far short as twelve and a half points, and the only thing that moved it was a longer block.

Still open, and what comes next

A statistic that is not a mean. The tt interval’s justification is that a mean of block means is close to normal; the studentised bootstrap’s is that it does not need that. A median, a quantile or a ratio of two means is where the second justification should be worth something, and where a closed multiplier has less claim to be right. Rerunning the grid for one of those is the next measurement, and it is the one that could reverse this essay.

A rule that prices both terms. The width term now has a closed form with no resampling in it, τ(n/)\tau(\lfloor n/\ell \rfloor), and the coverage term follows the length. A block length chosen to minimise the expected width subject to a coverage floor is buildable from those two pieces, and it needs its own oracle and its own sweep to be scored on data it was not fitted to.

Overlapping blocks and weaker dependence. Every block here is disjoint and every series is at 0.7. Overlapping batch means raise the effective number of blocks by about half, which would narrow the tt interval at the harsh cells; and at weaker dependence shorter blocks suffice, the counts rise, and the two intervals should converge. Neither is measured, and the number of observations that repeat each other sets how fast that convergence would be.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Block bootstrapBlock countBlock lengthClosed formConfidence intervalCoverageDegrees of freedomDependenceEstimated varianceInterval widthLong-run variancePercentile intervalResamplingStudentised bootstrap