Where a taper's case begins

Measuring a variance rather than a quantile

A resample's implied long-run variance can be computed from the sample with no resampling in it at all. A critical value cannot, and the difference is a factor of three in the draws before any of the resampling is counted.

Worth reading first: Where the bootstrap lies · The observations that repeat each other.

The taper field ends where its instrument does. Its comparison of block windows runs through a table of critical values — the 95% point of each construction’s reference distribution, over draws — and it reports that resolving the difference it cares about would need about twenty thousand draws.

The measurement it could not afford is affordable through a different reading of the same resample, and the saving is not only in time.

The saving is worth stating in advance because it is unusual. The two instruments are not a fast approximate one and a slow exact one; they are two exact readings of the same object, and one of them has no simulation in it. What separates them is not accuracy but what they are quantities of, and that turns out to decide more than the cost does.

Two instruments on one resample

Two instruments on the same resample. At a block length of 20 and 120 rows, over 200 draws with 400 resamples each. The implied standard deviation is computed from the sample's own autocovariances through the window's attenuation, with no resampling in it at all: it reads 1.964 for the rectangle and 2.051 for the trapezoid, a paired gap of 0.087 at 10.7 standard errors. The 95% point of the resampled mean reads 2.775 and 2.981, a gap of 0.206 at 6.3. The second instrument is measuring the same difference through a noisier lens.
Fig. 1 The two readings of the same construction, on the same draws.

The implied long-run variance is Σ_k κ(k)γ̂(k) — the window’s attenuation times the sample’s own autocovariances. The identity behind it is the one the taper field establishes: a block resample’s autocovariance at lag k is κ(k) times the residuals’. So the second moment of the whole reference distribution is a sum over lags, computable from the series with no resampling at all.

The 95% point of the resampled mean is a quantile of a distribution that has to be generated. Four hundred resamples per draw, in the measurements here.

At a block length of twenty on a hundred and twenty rows, the implied standard deviations read 1.964 for a rectangular block and 2.051 for a trapezoid, a paired gap of 0.087 at 10.7 standard errors over two hundred draws. The 95% points read 2.775 and 2.981, a gap of 0.206 at 6.3.

Same construction, same draws, same difference — and one instrument sees it at ten and a half standard errors and the other at six.

The ratio has a closed form

What each instrument costs to read. The number of draws each instrument needs to separate a rectangular block from a trapezoidal one at two standard errors, at a block length of 20 and 120 rows — measured from each instrument's own spread on the same draws. The implied variance needs 7.0 and the 95% point needs 20.2, a factor of 2.90 at this block length. There is a closed form beside it and it does not depend on either the scale or the size of the gap: the standard error of a p-quantile is √(p(1−p))/f(q) over √B where a standard deviation's is σ/√(2B), which at the 95% point of a nearly normal reference distribution is 3.30 times as many draws for the same statement. And the quantile route needs every one of those draws resampled, where the variance route needs none.
Fig. 2 The draws each instrument needs to separate the two windows at two standard errors.

For two nearly normal reference distributions differing by a scale factor δ:

A standard deviation has standard error σ/√(2B), so B = 2σ²/δ² separates them at two standard errors.

A p-quantile has standard error σ√(p(1−p))/(φ(z_p)√B), and the same scale difference moves the quantile by z_p δ, so B = 4σ²p(1−p)/(φ(z_p)²z_p²δ²).

The ratio is 2p(1−p)/(φ(z_p)z_p)², which depends on neither σ nor δ. At the 95% point it is 3.301: three and a third times the draws for the same statement, whatever the scale and whatever the gap.

Measured on the same draws rather than assumed, the implied variance needs 7.0 draws to separate the two windows and the 95% point needs 20.2 — a ratio of 2.90 against the closed form’s 3.30. The two agree to within what a two-hundred-draw estimate of a spread supports, and they are not the same calculation: the closed form assumes a normal reference distribution and a pure scale difference, and the measurement assumes neither.

Why a quantile is dearer

The closed form is two lines and it is worth reading rather than quoting, because the reason a quantile is dear is the reason it is a quantile.

A variance is an average. Every draw contributes to it, so its standard error falls like one over the square root of the number of draws with the full sample behind it.

A quantile is a local statistic. What determines the 95% point is how many draws land near it, and near it the density is low: at the 95% point of a normal distribution the density is 0.1031 against 0.3989 at the centre. So the draws that inform the estimate are the ones in a thin band, and the estimate’s standard error carries a 1/f(q) in it — which at the 95% point is a factor of about ten.

The p(1−p) in the numerator pulls the other way, and the two combine into √(0.0475)/0.1031 = 2.114 against a standard deviation’s 1/√2 = 0.707. Three times the standard error, nine times the draws for the same precision on the same statistic — and a factor of three rather than nine in the ratio quoted, because a scale difference δ moves the quantile by 1.645δ and moves the standard deviation by δ.

A quantile is expensive because it is a statement about one place in a distribution, and everything a reference distribution is used for is a statement about one place.

Where the measured ratio inverts

The closed-form ratio is a property of the two readings and the measured one is not, and the difference shows up sharply at one block length.

At ℓ = 8 the two windows’ implied standard deviations differ by 0.0625 at 15.6 standard errors and their 95% points by 0.0417 at 1.4 — the variance route needs 3.3 draws and the quantile route 406. At ℓ = 32 the gaps are 0.1439 and 0.3831 and the two routes need 4.9 and 8.0.

At ℓ = 12 the variance gap is 0.0038 at 0.7 standard errors and the quantile gap is 0.0977 at 3.3, and the ratio inverts: the variance route needs fifteen hundred draws and the quantile route seventy-three.

That is not the instrument failing. Twelve is where the two windows’ implied variances cross at this sample size, so there is no gap for the variance route to resolve — and what the quantile route is resolving there is the shape difference below rather than a scale difference at all.

A ratio of draws needed is a ratio of two noise-to-signal readings, and it is undefined where the signal is zero. What is scale-free, and what holds at every block length here, is that the variance route carries less of its own noise as a share of what it reads.

What a longer block buys and what it costs. A trapezoidal block at 120 rows, with the error split into the two things it is made of. The bias falls with the block length, because a longer block attenuates less, and it flattens at 23.2% because the sample's own autocovariances are short whatever window is applied to them. The spread rises with it, because a longer block means fewer of them. Their sum in quadrature has a minimum at ℓ = 16, which is not where either of the two has one. The faint line is the rectangle's total error, for scale: it is above the trapezoid's from ℓ = 12 onwards.
Fig. 3 A trapezoidal block at 120 rows with the error split into its two parts. The bias falls with the block length and flattens at 23.2%; the spread rises with it; their sum in quadrature has a minimum at ℓ = 16, which is where neither of the two has one.

Where the closed form is wrong, and by how much

The disagreement between 2.90 and 3.30 is worth a paragraph rather than a shrug, because it points at something.

A pure scale difference would move the 95% point by exactly z × the standard-deviation gap: 1.645 × 0.087 = 0.143. The measured quantile gap is 0.206 — 44% larger. So the two windows’ reference distributions do not differ only in scale; they differ in shape, and the tapered one has a heavier tail relative to its own spread.

That is a real difference and it is not the one anybody was measuring. A window comparison read off a quantile is reading a scale difference and a shape difference added together, and the sum is what appears in a critical-value table. The variance route separates them by measuring only the first.

Two instruments that disagree by more than their noise are measuring two things, and here the two things are separable and only one of them is what a long-run variance is.

What the cheap instrument makes possible

The saving in draws is the smaller half of what the change buys.

No resampling noise. The implied variance is computed rather than sampled, so the only noise in it is the noise in γ̂ — which is the thing being studied. The quantile route adds the resampler’s own variability on top, and four hundred resamples per draw is not enough to make that negligible: it is the reason the quantile route’s spread is 0.462 against the variance route’s 0.114.

Sweeps become affordable. The gap between the exact and sampled bias is measured at six block lengths on four sample sizes with four hundred draws each — 9,600 evaluations, each of which is a sum over lags. Through the quantile route the same sweep is 9,600 × 400 resamples, which is what made it not exist.

The quantity is interpretable. A 95% point of a resampled mean is a number about a test. An implied long-run variance is the quantity every one of these constructions is an estimator of, so its bias and its spread are statements about the estimator rather than about a decision downstream of it. That is what lets the error be split into its parts, which a quantile cannot be.

The instrument the field already had

There is a third instrument in this collection and it is worth placing beside the two here, because it makes the same trade in the opposite direction.

The cost of a usable draw prices a chain-based sampler by how many evaluations it spends per independent draw, and the quantity it ends up caring about is the effective sample size — draws divided by an autocorrelation time. That is an instrument about the sampler rather than about what the sampler estimates, and it is expensive for the opposite reason: it needs the whole sequence rather than a summary of it.

Three instruments, three questions. How well does the estimator estimate the thing (a variance). How well does the reference distribution place its tail (a quantile). How efficiently does the sampler produce draws (an effective size). A comparison that picks the wrong one of the three does not get a wrong answer; it gets a right answer to a question nobody asked, which is harder to notice.

Where the taper's case begins, and it is not where the algebra says. The block length at which a trapezoidal block's implied variance stops being more biased than a rectangular one's, against the length of the sample. Computed exactly — from the law's own autocovariances, with no sampling in it — the answer is 19.2 and does not depend on the sample at all. What a sample of 120 rows reports is 13.3, and the reported crossing walks out towards the exact one as the sample grows: 13.3, 15.0, 16.4, 18.0. The mechanism is that the autocovariances the window is applied to are themselves attenuated, worst at the longest lags, and the window that discards those lags loses less of them.
Fig. 4 The block length at which the trapezoid stops being the more biased window, against the length of the sample. Computed exactly from the law’s own autocovariances it is 19.2 and does not depend on the sample at all; a hundred and twenty rows report 13.3, and the reported crossing walks out to 15.0, 16.4 and 18.0 as the sample grows.

What the expensive instrument is for

None of this says the critical-value table was the wrong thing to build. It answers a question the variance does not.

A resampling is used to get a reference distribution, and what a practitioner does with it is compare a statistic against a quantile of it. So the quantity that decides whether one construction is better than another for that purpose is the quantile, and the taper field’s table measures exactly that: four constructions that lay blocks end to end come out 10.3%, 10.5%, 12.6% and 16.2% short of the statistic’s own 95% point, ordering exactly as their closed-form biases do.

That table also produces a result the variance route cannot: two constructions sharing a taper exactly still differ by 6.1 standard errors, so a window predicts a critical value inside one family of constructions and not across families. The shape difference above is presumably part of why.

The variance is the right instrument for comparing estimators of a variance and the quantile is the right instrument for comparing reference distributions, and the mistake is not using one of them — it is using the second where the first was wanted, because the second is what the field happened to have built.

What it costs to be wrong about the instrument

The concrete cost in this field is one deferral and one wrong conclusion, and both are worth counting.

The deferral is the taper field’s own: it names the crossing at ℓ = 20, states that resolving the difference there would need about twenty thousand draws, and stops. That is a correct assessment of what the quantile route needs for the exact difference of three tenths of a point. What the variance route needs for the difference a hundred and twenty rows actually show is seven draws, and the difference is fifteen times larger than the one that was being priced — so two separate factors, each of an order of magnitude, were compounding.

The wrong conclusion is that the taper’s advantage has not arrived. Read at each window’s own block length on the error, it has, at the sample size in question.

Neither of those is a mistake in a calculation. They are the consequence of an instrument that could only see one quantity at one setting, and a field that then said what that instrument saw. The check that would have caught it is asking what the instrument is measuring and whether it is the quantity in the question — which is not a check anything in this collection currently runs, and is the habit this essay is an argument for.

What is claimed here, and what is not

This essay takes what a comparison’s instrument decides about what it can see. The claims are that a block resample’s implied long-run variance is computable from the sample’s autocovariances with no resampling, through the attenuation identity; that at a block length of twenty on a hundred and twenty rows the two windows’ implied standard deviations differ at 10.7 paired standard errors and their 95% points at 6.3, on the same draws; that a quantile needs 3.301 times the draws of a variance for the same scale statement, by a closed form that depends on neither the scale nor the gap, and 2.90 times when measured; and that the excess — a quantile gap of 0.206 against the 0.143 a pure scale difference predicts — is a shape difference between the two reference distributions rather than noise.

What stays out, and is named as a decision: the shape difference itself. It is identified as the reason the two routes disagree and it is not characterised. What would characterise it is a comparison of the two constructions’ whole reference distributions rather than of two summaries, and that is a different essay from this one.

Also out: whether the variance route’s ordering survives to the critical values. The variance route says the trapezoid is better at each window’s own block length and the critical-value table says the constructions order by their biases at a shared one. The two are not in conflict — different quantities at different settings — and whether the tuned ordering also holds on quantiles is not measured here, because measuring it is what the draw count forbids.

The boundary against the taper field is that it names the instrument as its limit and this one builds the alternative. Its own conclusion about critical values stands; what changes is that the question it could not answer now has an answer.

Where the identity comes from and where it stops

The instrument rests on one identity and it is worth being precise about its scope, because outside that scope the cheap route is not available.

A moving block resample of a residual series has autocovariance κ(k)γ̂(k), where κ is the block window’s normalised self-convolution. That is established three ways in the taper field — the closed form, the expectation on a particular residual series, and the average of realised resamples — and it is exact for any construction that lays blocks end to end from a fixed series.

It stops at three places. A stationary bootstrap randomises its block lengths, so its attenuation is geometric rather than a self-convolution and the identity has to be replaced by the geometric one, which is also closed. A sieve generates a new series from a fitted model rather than laying pieces of the old one down, so there is no κ at all and its implied variance is the fitted model’s, which is a different closed form. And a wild or blocked multiplier resampling keeps each residual in its own row, so its attenuation is a fact about the multiplier rather than about a block.

All four have a computable implied variance and none of them needs a resample to get it. What none of them has is a computable quantile, which is why the expensive instrument is the only one that covers every construction on one axis — and is the reason a table of critical values existed in the first place.

The checks, and the refusals that make them mean something

Three claims are gated. The implied variance is required to agree with the variance of realised resamples on three windows at two block lengths, which is a check on the identity the whole instrument rests on and is a comparison between a sum over lags and a sampler that has never heard of one. The closed-form ratio is required to exceed three and to be independent of the gap and the scale, which is checked by computing it at two different gaps and requiring the same number. And the two instruments are required to order the same way on the same draws, because an instrument that were cheaper and disagreed would be a different measurement rather than a better one.

The refusal is what the cheaper instrument makes visible. Two windows compared at one block length, when each has its own best one, is refused — a comparison the expensive instrument could not have made either way, because the sweep it needs is the one the draw count forbids.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A length for each instrument — both name block bootstrap, block length, critical value, long-run variance, monte carlo, reference distribution, resampling, tapering
  • The ordering reverses again — both name block bootstrap, block length, critical value, long-run variance, monte carlo, reference distribution, resampling, tapering
  • What the interval covers — both name block bootstrap, block length, critical value, long-run variance, monte carlo, reference distribution, resampling, tapering
  • An ordering that depends on the rule — both name block bootstrap, block length, closed form, long-run variance, monte carlo, resampling, tapering
  • The length nobody has — both name block bootstrap, block length, closed form, long-run variance, monte carlo, resampling, tapering
  • The count or the length — both name block bootstrap, block length, closed form, long-run variance, resampling, tapering

Named objects

A flat tag is an object no other essay names yet.

Block bootstrapBlock lengthClosed formCritical valueLong-run varianceMonte CarloPrecisionReference distributionResamplingSampling variationStandard errorTail probabilityTaperingVariance