What the interval covers
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
The premise of this site is that nothing is called 95% until it has been counted. Three fields have now compared two block windows and none of them has counted one.
The count
Every earlier comparison has been a comparison of errors — how far an implied variance is from the law’s, how far a 95% point is from the finite-sample truth. An error is a distance and a coverage is a rate, and a rate is the only one of the three that can be held against a promise.
The interval is the ordinary percentile one: the 2.5% and 97.5% points of the resampled means, taken off the same sorted array the other two readings come from, at whatever block length the rule chose.
The eight cells run from 80.8% to 91.0%.
Not one reaches 95%. The best is the tapered window at a plug-in block length, at 91.0% ± 1.4. The worst is the rectangle at the rule of thumb, at 80.8% ± 2.0. A rectangle at a protocol length of eight gives 84.8%; the taper at the same length gives 89.8%.
The shortfall runs from four points to fourteen, on a promise of five per cent error, and every comparison the three fields before this one have made is a comparison of a point or two inside it.
What the interval is, exactly
Three choices go into building it and all three are the conventional ones.
The percentile form. The interval is the 2.5% and 97.5% points of the resampled means themselves, rather than a studentised or bias-corrected version. It is the simplest bootstrap interval and the one a reader is most likely to be handed.
The block length the rule chose. Coverage is read at whatever length the rule picked on that draw, because a practitioner has one block length and reads one interval from it — not at a length chosen to make the coverage good.
And the truth is zero. The errors have mean zero by construction, so an interval covers when it contains zero. There is no estimation of the target and no second sample.
So the count is what it says: over four hundred samples, how often the interval a practitioner would have written down contains the number it is about.
Which does not make the comparison pointless
Two readings of that are available and only one is right.
The wrong one is that a two-point difference inside a fourteen-point shortfall is not worth measuring. It is — a block bootstrap under dependence is a hard problem and nobody expects nominal coverage at a hundred and twenty rows; the field that first measured what dependence does to an interval finds an ordinary interval covering far worse than any of these. Two points of coverage recovered for free is two points.
The right one is that the ordering and the level are different findings and the level is the larger one. A practitioner choosing between two windows is choosing between 84.8% and 89.8% at a protocol length. A practitioner deciding whether to believe a 95% interval at all is looking at a number between 80.8 and 91.0, and that decision is not helped by knowing which of the two is better.
Converting the earlier fields’ units into this one
The three quantities are not incomparable, and the conversion is worth having because it says what the earlier comparisons were implying about coverage without measuring it.
A two-sided interval built at a 95% point that is x per cent short covers . So:
- a 95% point 10% short delivers 92.2%
- 20% short delivers 88.3%
- 40% short delivers 76.0%
Which puts the block-resampling shortfalls this collection has been reporting on the scale a promise is made in. A construction ten per cent short of the truth is a 92% interval — a shortfall a reader might tolerate. One forty per cent short, which is what a multiplier on every row delivers, is a 76% interval, and no reader would.
The relationship is not proportional, which is the reason the conversion is worth doing rather than assuming. Four times the error in the critical value produces three times the shortfall in coverage, because the normal density is falling away from 1.96 — so the errors compress as they grow, and a comparison made in error units flatters the worst constructions.
What the promise was
One clarification, since 95% is doing the work of a claim here.
The interval is not promising that 95% of intervals contain the estimate; it is promising that 95% of them contain the true mean of the process the errors came from, which is zero by construction. That is the frequentist coverage the site’s name is about, and it is countable exactly because the truth is known.
A practitioner has no such luxury and that is the point of counting it here. A coverage is a property of a procedure rather than of a dataset, so it can only ever be counted in simulation — and a procedure whose coverage has never been counted is a procedure making an unchecked promise. The oldest figure on this site is twenty intervals at a known truth, and this is the same count four fields later on a much more elaborate procedure.
The ordering, which agrees
The tapered window covers better under every one of the four rules: by 5.00, 2.50, 5.75 and 4.00 percentage points at the protocol length, the rule of thumb, the plug-in and the oracle, at three to five paired standard errors.
That is the same ordering the 95% point gives and the opposite of the implied variance’s at two of the four rules, which is the previous essay’s subject.
The agreement is worth more than it looks, because the two readings are sensitive to different things. A critical value uses one tail; an interval uses both, so it is sensitive to the resampled distribution’s skewness in a way the critical value is not. A bootstrap that is accurate on average and skewed the wrong way would read well on the quantile and cover badly.
It does not happen here. The two readings order the four rules and two windows identically, which says the taper’s advantage is in the whole reference distribution rather than in one tail of it.
What a count of four hundred can say
The standard errors on the eight cells run from 1.4 to 2.0 percentage points, which is what four hundred Bernoulli draws give at these rates.
That is enough for the level and not much more than enough for the margins. A shortfall of fourteen points is seven standard errors and is not in doubt; a margin of two and a half points between two windows is a paired comparison at 3.2 standard errors and is established; a margin of one point would not have been.
The paired comparison is what makes the margins readable at all. The two windows are run on the same four hundred samples with the same seeds and the same resamples, so the difference in coverage is a difference within each draw rather than between two independent rates — which turns a comparison of two numbers each carrying 1.5 points of error into one carrying about a third of that.
Where the shortfall comes from
Two sources, and separating them says what a better rule could and could not fix.
The block length is too short. Every feasible rule picks between 4 and 13.89 where the oracles sit at 16 to 22, and a short block truncates the dependence — so the resampled means are less spread than the true sampling distribution and the interval is too narrow. That is the larger part and it is fixable in principle: the plug-in, which picks the longest of the three feasible lengths, gives the best coverage of the three under both windows.
And the reference distribution is not the sampling distribution. Even at the oracle length the coverage is 81.8% and 85.8% — worse than the plug-in’s, which is the one place the ordering by rule inverts. The oracle here is chosen to make the 95% point accurate, not to make the interval cover, and those are different objectives; a length that centres one tail need not centre both.
That inversion is the sharpest thing in the table. The rule with the most accurate critical value does not have the best coverage, on either window, and the gap is three and a half points for the rectangle and five for the taper.
The one place the ordering by rule inverts
The oracle rows deserve more than a clause, because they are the only place in four fields where a benchmark does worse than a rule that can be run.
On the implied variance and on the quantile, the oracle is best by construction: it is the length that minimises that draw’s own error, so nothing can beat it. On coverage it is 81.8% for the rectangle and 85.8% for the taper, against the plug-in’s 85.3% and 91.0%.
The oracle is not an oracle for coverage. It is the argmin of the 95% point’s error, and it is being read on a quantity it was not chosen for — which is precisely the mistake this whole field is about, committed here by this field’s own benchmark and reported rather than repaired.
Repairing it would mean defining a coverage oracle, and a coverage on a single draw is a coin: the interval covers or it does not, and there is no per-draw argmin worth taking. The second essay of this field says the same thing about why coverage has no oracle column. So the benchmark row of the coverage table is an approximation and is labelled as one.
What a fourteen-point shortfall looks like from inside
The number worth carrying is not the average but what it means for one analysis.
An interval covering 80.8% is an interval whose stated 5% error rate is really 19.2% — nearly four times what it says. A test built on the same reference distribution rejects a true null about four times as often as it claims, which is the difference between a finding at the five per cent level and a finding at the twenty.
At 91.0% it is 9.0% against a claimed 5%, which is a factor of 1.8.
So the choice between the worst cell and the best is the difference between a procedure that is twice its nominal size and one that is four times it. Neither is the procedure a reader thinks they are being handed, and the gap between them is most of what a block bootstrap under this much dependence has to offer.
Which is a refusal rather than a curiosity
It is worth stating as the negative result it is, because it closes the obvious shortcut.
If coverage always followed the critical value, this essay would be unnecessary: measure the quantile, infer the interval. The oracle rows say it does not. An interval reads two tails and a critical value reads one, the resampled distribution of a mean under dependence is skewed, and a block length that puts one tail in the right place puts the other somewhere else.
So the three readings really are three, and the cheapest of them — the implied variance three fields have used — is the furthest from any of them.
The rule that is short and the rule that is right
One pattern in the eight cells is worth stating on its own, because it is the practical half.
The three feasible rules order themselves by their block length in every column: the rule of thumb picks 4 and covers worst, the protocol length picks 8 and covers next, the plug-in picks 12.59 and 13.89 and covers best. Under both windows, without exception.
That is not automatic. On the implied variance the plug-in is best of the three too, but the ordering between the two fixed rules depends on the window; on coverage there is no such wobble. Longer is better throughout the feasible range, which is what a shortfall driven by truncated dependence predicts and is the cleanest signal in the table for what to do.
Where it stops being true is past the feasible rules, at the oracle, which picks 16 to 22 and covers worse than the plug-in — the inversion the next section is about.
What would close the gap
Three things would and it is worth ranking them by how much.
A longer block. The largest part of the shortfall is that every feasible rule picks a length well short of either oracle, so the resample truncates the dependence and the interval comes out too narrow. The plug-in, which picks the longest of the three, covers best of the three under both windows — 85.3% and 91.0% against the rule of thumb’s 80.8% and 83.3%.
A studentised interval. The percentile interval used here inherits the resampled distribution’s shape directly. A studentised version divides each resampled mean by its own standard error, which removes most of the skewness and is the standard repair for exactly this. It is not measured here and it is the obvious next thing.
And a longer sample. Everything is a hundred and twenty rows at a persistence of 0.7, which is a hard case: the effective sample size there is about twenty-one rows’ worth of information. Coverage at a thousand rows would be a different measurement and is not made.
What would not close it is choosing the window better. The two windows differ by two to six points of coverage; the promise is short by four to fourteen. The window is the smallest of the four dials, which is the same ordering the earlier fields report on their own instrument and is the reason the recommendation ends where it does.
Why nobody had counted
It is worth asking, because the count is four hundred draws and three hundred resamples and takes six seconds.
The answer is the one this field has been about throughout: the instrument that made the sweeps possible does not have a coverage in it. An implied long-run variance is a closed-form sum over autocovariances — there are no resampled means to take quantiles of, so there is no interval to check. Getting a coverage means resampling, and resampling is what the instrument was chosen to avoid.
So the absence is structural rather than an oversight. Three fields swept a comparison on a quantity that is cheap precisely because it discards the object a coverage is a property of, and the coverage could not have been counted without giving up the sweep that made the fields affordable.
What makes it affordable now is that the sweep is already done. This field does not re-derive the rules, the lengths or the ordering; it takes the four rules as given and reads three numbers off one bootstrap at seven block lengths rather than fourteen. That is a hundredth of the earlier fields’ work, and it is only a hundredth because they did theirs first.
What the four fields come to
The block-length line of fields now has four instruments and one recommendation, and it is worth setting them out once.
A block length matters more than a window. What the best feasible rule gives up against the best available length is several times what the window choice buys there, on every instrument measured. That is the earlier field’s finding and nothing here touches it.
Estimate the length rather than fixing it. The plug-in gives the best coverage of the three feasible rules under both windows, the largest margin on the quantile, and the only positive margin on the variance.
Use the taper. On both readings a practitioner acts on, it wins at every rule — including the two where the implied variance says it loses.
And do not believe the interval. 80.8% to 91.0% on a promise of 95%, at a hundred and twenty rows and a persistence of 0.7. Every recommendation above is a recommendation about where inside that range to sit.
The last of those is the one the whole line of fields could not have reached on the instrument it was using, and it is the one this site is named for.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A taper and a critical value — both name block bootstrap, critical value, dependence, long-run variance, reference distribution, resampling, stationary bootstrap, tapering
- Measuring a variance rather than a quantile — both name block bootstrap, block length, critical value, long-run variance, monte carlo, reference distribution, resampling, tapering
- What a multiplier cannot keep — both name block bootstrap, critical value, dependence, estimation error, monte carlo, persistence, reference distribution, resampling
- An ordering that depends on the rule — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering
- Errors generated from a fitted model — both name block bootstrap, critical value, dependence, estimation error, persistence, reference distribution, resampling
- The error no window repairs — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering
Named objects
A flat tag is an object no other essay names yet.
Block bootstrapBlock lengthConfidence intervalConservative intervalCoverageCritical valueDependenceEstimation errorLong-run varianceMonte CarloPersistenceReference distributionResamplingStationary bootstrapTapering