An ordering that depends on the rule
Worth reading first: Where the bootstrap lies · The observations that repeat each other.
Read at each window’s own best block length, a tapered block beats a rectangular one — 40.67% against 42.66% at a hundred and twenty rows, and at every sample size in that field’s sweep. That is the comparison the field ends on, and it is the comparison a practitioner cannot make, because the best block length is a quantity nobody has.
Re-run on the rules a practitioner does have, the answer is not the same answer with a smaller margin. It changes sign.
Four rules, two each way
At the best available length on each draw, the taper is ahead by 2.12 points of a 35.81% error, at 22.1 paired standard errors.
At a length estimated from the sample’s own persistence, it is ahead by 1.89 points, at 8.0.
At a length written into the protocol — eight lags, the field’s own default — the rectangle is ahead by 2.04 points, at 13.8.
At the rule of thumb, , the rectangle is ahead by 4.54 points, at 71.7.
Every rule sees the same four hundred samples, and the comparisons are paired on the draw, so these are not four independent measurements that happen to disagree. They are one measurement read at four settings, and it changes sign between them.
Why a short length reverses it
The mechanism is in the lengths beside the margins.
The best available length averages 24.57 for the taper. The plug-in picks 14.36. The protocol picks 8 and the rule of thumb picks 4.
A tapered window weights the lags inside its block by a triangle that reaches zero at both ends. At the right length that is exactly what buys the taper its advantage: it attenuates less at short lags and throws away more of the series, so its implied variance has a smaller bias and a larger spread, and at the length where those balance it wins.
At a quarter of the right length the triangle has almost nothing left to weight. A block of four lags tapered to zero at both ends keeps very little of what is inside it, where a rectangle of four lags keeps all of it — so the rectangle, which throws away nothing inside its window, is much less damaged by being short.
The taper’s advantage is an advantage at the right block length. Every rule that is systematically short reverses it, and two of the three feasible rules are systematically short by a factor of three or six.
Not a sample-size effect
A reversal that appears at one sample size invites the question of whether it is about that size.
At 120, 240, 480 and 960 rows the four rules give margins of −2.04, −4.54, 1.89 and 2.12; −3.40, −4.98, 2.31 and 1.98; −4.32, −4.32, 2.59 and 2.00; and −4.89, −3.29, 2.70 and 2.16. Every rule is on the same side of zero at every size.
Twenty times the data does not change which window wins under which rule. What it changes is the size of the fixed rules’ deficits, which grow — from 2.04 to 4.89 points for the protocol length — because the oracle’s own answer grows with the sample and a fixed length falls further behind it.
So the reversal is not an artefact of a short series. It is a property of a length that does not track the target, and the more data there is, the worse it gets.
A data-driven length makes the comparison noisier
The four margins come with four t statistics, and dividing one by the other gives the paired standard error of each comparison. They are not the same size, and the pattern is worth naming.
- best available length: 2.12 at 22.1, so 0.096
- the plug-in: 1.89 at 8.0, so 0.236
- the protocol’s eight lags: 2.04 at 13.8, so 0.148
- the rule of thumb’s four: 4.54 at 71.7, so 0.063
A factor of 3.7 between the noisiest comparison and the quietest, on the same four hundred draws, comparing the same two windows.
The ordering explains itself once the lengths are read beside it. The rule of thumb and the protocol pick a fixed length on every draw, so the only thing varying between draws is the sample, and the pairing removes most of that. The plug-in picks a different length on every draw, and the margin between two windows depends on the length, so its comparison inherits the length’s variation on top of the sample’s.
So a rule that reads the data costs precision in the comparison as well as delivering a different answer. The plug-in’s margin is the second largest of the four and the least well measured of the four, and nothing about the margin itself says so.
That has a practical edge. The plug-in is the rule this field recommends, and its 1.89-point advantage for the taper rests on a t of 8.0 — comfortable, but an order of magnitude less comfortable than the 71.7 attached to the rule of thumb’s reversal. A reader taking the four rows as four equally solid findings would have the strength of the evidence almost exactly backwards relative to how much anybody should care about each row.
The oracle row is the interesting exception. Its length varies from draw to draw as much as the plug-in’s does — a standard deviation of 16.23 against a mean of 24.57 — and its comparison is the second quietest at 0.096. The reason is that each window is read at its own best length rather than at a shared one, so whatever moved the length on a given draw moved both windows’ lengths together and the pairing removes it. A shared data-driven length is the one arrangement that lets the length’s noise into the difference.
The errors behind the margins
The margins are differences and they sit inside errors that are worth having in view, because they are large.
Under the rectangle: 45.28% at the protocol length, 57.28% at the rule of thumb, 44.96% at the plug-in and 37.93% at the oracle. Under the taper: 47.32%, 61.81%, 43.07% and 35.81%.
So the largest difference in the table is not between the two windows at all. It is between the rules: the rule of thumb delivers 57.28% where the oracle delivers 37.93%, a gap of 19.35 points, against window margins of two to four and a half.
That is the reading the next essay is entirely about, and it is the reason this one’s recommendation is about the length rather than about the window. A practitioner arguing between windows is arguing about two points inside a decision worth nineteen.
Which comparison a practitioner should read
Not the oracle’s, and this is the essay’s whole recommendation.
A comparison at the best available setting answers which window would be better if the tuning problem were solved. It is the right question for somebody choosing what to study and the wrong one for somebody choosing what to run, and the two answers here are opposite.
The comparison to read is the one at the rule the practitioner will actually use. If that is a plug-in from the sample, the taper wins by 1.89 points. If it is a number in a protocol, the rectangle wins by 2.04. If it is , the rectangle wins by 4.54.
Which turns the recommendation inside out. The sentence that follows from this field is not use the taper; it is estimate the block length, and then use the taper — and if the block length is not going to be estimated, the taper is the wrong choice.
What the taper is actually buying
It is worth being clear about what the 2.12 points at the oracle are, because they are a real thing and the reversal does not make them unreal.
A tapered window’s implied variance has less bias than a rectangular one’s at the same block length, because the attenuation it applies is its window’s self-convolution and a window reaching zero at both ends attenuates less at the short lags that carry most of the dependence. That is exact, it is computed in closed form, and it does not depend on any sample.
What it costs is spread: a window that discards the ends of its block has fewer effective pairs behind its estimate. So the taper’s error curve is lower at its own optimum and steeper on the short side, which is exactly the property that makes it lose under a rule that lands short.
The two facts together are the field’s whole content: the taper is better where it is right and worse where it is wrong, and the rules a practitioner has land where it is wrong.
The plug-in is the interesting row
Three of the four rows are what they look like. The plug-in’s is worth reading twice.
It picks 14.36 against an oracle of 24.57 — well short — and it still reproduces the oracle’s ordering, at 8.0 paired standard errors. So the reversal is not simply a matter of being short of the target; it is a matter of being short enough.
Somewhere between 14 and 8 lags the two windows change places, and the plug-in is on the right side of it while the protocol’s eight is not. That is a narrow margin for a recommendation to rest on, and it is the reason this essay’s advice is estimate the length rather than estimate it well: the difference between a rule that reads the data and one that does not is the difference between the two orderings, and refining the estimate further buys much less than that.
Where the crossing is
The plug-in at 14.36 lands on the taper’s side and the protocol’s 8 on the rectangle’s, so the two windows change places somewhere between them, and it is worth saying what that crossing is.
At a shared block length the rectangle wins at short lengths and the taper at long ones — which is the crossing the earlier field locates exactly, at ℓ = 19.15 in the algebra and between 13.28 and 17.99 in a finite sample depending on its size.
The crossing here is a different object: each rule gives the two windows different lengths, because the plug-in’s constant differs by window and the oracle’s argmin does too. So what is being asked is not at which shared length do they cross but which rules put each window near enough to its own optimum for the taper’s advantage to survive.
The two questions have similar answers at these sizes, which is a coincidence of the numbers rather than a fact: the plug-in gives the rectangle 13.2 and the taper 14.4, and the shared-length crossing is around 13 to 18. A rule whose constant differed more between windows would separate the two questions.
What the reversal does not say
It does not say the earlier field was wrong. That comparison is correct, it is made at a stated setting, and it says in its own closing section that the setting is not available. This field supplies the missing half rather than contradicting it.
Nor does it say a tapered window is a poor construction. At its own optimum it delivers 35.81% against the rectangle’s 37.93%, and that is the best either of them ever does. The failure is a tuning failure and it is the same one every tuning parameter in this collection has: the rule that would make the better construction better is a rule nobody runs.
The paired comparison, and why it is necessary
Every margin here is a paired difference: the same series goes through both windows on every draw, and the standard error is the standard error of the per-draw difference.
That is not a refinement. The errors themselves have standard deviations comparable to the margins being measured — a 2.12-point difference between two quantities each around 36% — so an unpaired comparison at four hundred draws would resolve almost none of the four rows. Paired, the same four hundred draws give 22.1, 8.0, 13.8 and 71.7 standard errors.
The reason pairing works so well here is that most of the variation in the realised error is a property of the draw rather than of the window: a sample with a long run in it gives both windows a bad estimate, and a quiet one gives both a good estimate. Differencing removes it.
This collection reaches for the same construction wherever two rules are read on one sample — the estimated-covariance field says so explicitly, and the reason is always this one.
One practical consequence
The advice that follows is unusually concrete, so it is worth stating in the form a protocol would take.
If the analysis plan fixes a block length, use a rectangular block. The taper’s advantage does not survive a fixed length at any sample size measured, and at the rule of thumb’s four it loses by four and a half points.
If the analysis plan estimates the block length from the data, use a tapered block. It wins by 1.89 points at the plug-in, and it wins by more the closer the estimate gets to the target.
And in either case the length matters more than the window, by a factor of three or more. A protocol that specifies a window and leaves the length to habit has specified the smaller half of the decision.
One caution goes with all three. Everything here is measured on a first-order autoregression at a lag-one correlation of 0.7, which is the world the block-window fields are built in, and a dependence with a different shape would move the optima and could move the crossing. What should transport is the mechanism — a taper landing short loses more than a rectangle landing short — because that is a statement about the windows rather than about the law.
It is worth adding what none of this touches. The two windows are compared here on one construction — a block resample’s implied long-run variance, at one dependence, on a stationary series — and the reversal is a statement about that. What would move it is anything that changes where the error curve sits relative to the lengths the rules pick: a shorter series, a weaker dependence, or a rule with a different constant. Two of those are common. So the safe form of the recommendation is not that the rectangle is better at eight lags; it is that a window comparison made at a setting the practitioner will not use answers a question the practitioner does not have.
What is claimed here, and what is not
This essay takes whether the ordering between two block windows survives a feasible choice of block length. The claims are that at a hundred and twenty rows the tapered block beats the rectangular one by 2.12 points at the best available length on each draw, at 22.1 paired standard errors, and by 1.89 points at a length estimated from the sample’s persistence, at 8.0; that the rectangle beats the taper by 2.04 points at a protocol length of eight, at 13.8, and by 4.54 at the rule of thumb’s four, at 71.7; that every rule is on the same side of zero at 120, 240, 480 and 960 rows, with the fixed rules’ deficits growing from 2.04 to 4.89 points as the sample grows; and that the mechanism is the lengths — 24.57, 14.36, 8 and 4 — with a tapered window at a quarter of the right length having thrown away most of what it was weighting.
What stays out, and is named as a decision: the raised-cosine window. The earlier field measures three windows and this one measures two, because the comparison being re-run is the two-window one the earlier field’s conclusion is about. A third window would add a row and would not change the question, which is whether an ordering survives the tuning problem.
Also out: a rule that chooses the window and the length together. Every rule here picks a length for a given window. A rule choosing both jointly is a search over a product set, which a neighbouring field prices and which would need its own charge before its ordering meant anything.
The boundary against the field that compared the windows is that it reads each window at its own optimum and this one at a rule. Both are right and they disagree, which is what makes the setting part of the claim.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The error no window repairs — both name attenuation, bias-variance, block bootstrap, block length, closed form, dependence, long-run variance, mean squared error, monte carlo, resampling, tapering
- A length for each instrument — both name bias-variance, block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering, tuning parameter
- The instrument and the reading — both name block bootstrap, block length, closed form, dependence, long-run variance, monte carlo, resampling, tapering
- An interval that carries its scale — both name block bootstrap, block length, closed form, dependence, long-run variance, monte carlo, resampling
- Measuring a variance rather than a quantile — both name block bootstrap, block length, closed form, long-run variance, monte carlo, resampling, tapering
- The ordering reverses again — both name block bootstrap, block length, dependence, long-run variance, monte carlo, resampling, tapering
Named objects
A flat tag is an object no other essay names yet.
AttenuationBias-varianceBlock bootstrapBlock lengthClosed formDependenceLong-run varianceMean squared errorMonte CarloPaired comparisonPlug inResamplingTaperingTuning parameter