Counting what is independent

What a multiplier cannot keep

Two reasons were named for the quarter a blocked resampling falls short, and taking either away makes the gap larger. What is left is a bound — a multiplier can only take dependence out, and the residuals' own is already below the errors'.

Worth reading first: The experiments that could have happened · Where the bootstrap lies.

The construction that survives a world with two defects at once — errors whose variance depends on the design and rows that repeat each other — is a multiplier drawn once per run of consecutive rows. It beats the four alternatives and it is short of the truth by about a quarter, and the essay that measured it named two candidates for the residue and measured neither.

The residuals are not the errors, and their dependence has been distorted by the fit that produced them. The rolling scheme couples neighbouring origins: two forecasts made one step apart come from windows sharing all but one row, so their errors are correlated by the estimation and not by the errors, and no resampling of estimation residuals can reproduce that.

Both are true statements. Neither is the explanation, and the way to find that out is to take each away and look.

The two-by-two

Taking the residuals away is possible here for a reason worth stating: the null being resampled under is that every coefficient is zero, so the series is the error. Handing the same construction the true errors rather than a fitted benchmark’s residuals is not a procedure anybody can run; it is the upper bound on what the procedure could be worth if the residual problem were solved perfectly.

Taking the coupling away means scoring the same table at a fixed origin — every forecast made from one fit, estimated once on the first sixty rows — so that two neighbouring origins share no estimation rows at all. That changes the statistic, so it changes the truth as well, and each cell is measured against its own.

Neither of the two named candidates is the shortfall. The blocked multiplier's critical value against the truth, as a share of that truth, under two schemes and two sources. Handing the resampling the true errors rather than the fitted benchmark's residuals — possible here because the null is β = 0 — makes it worse. Scoring at a fixed origin, where neighbouring forecasts share no estimation rows, leaves a shortfall of the same share against a truth half as large again (6.4303 rather than 4.0405). Each cell is measured against its own truth, because changing the scheme changes the statistic.
Fig. 1 The critical value the resampling produces against the truth it is trying to reproduce, as a share of that truth, under two schemes and two sources.

The rolling scheme with residuals — the construction as it stands — reaches 2.9849 against a truth of 4.0405, which is 26.1% short.

Hand it the true errors and it reaches 2.4716 against the same truth: 38.8% short. The gap gets bigger by half.

Score at a fixed origin with residuals and it reaches 4.6259 against a truth of 6.4303: 28.1% short. The truth is half as large again and the shortfall is the same share of it.

Both changes at once: 4.4639 against 6.4303, 30.6% short.

Neither candidate is the explanation. The fixed scheme falls short by essentially the same fraction as the rolling one, so the coupling between neighbouring origins is not what is missing; and handing the construction the errors themselves makes it worse rather than better, so the residuals are not what is missing either.

Both cells are measured on a hundred and fifty samples with ninety-nine replicates each, against truths computed on two thousand — so the truths are far more precisely known than the guesses, which is the right way round when the truth is the thing every row is measured against. The standard errors on the four critical values run from 0.11 to 0.22, against differences between cells of half a unit and more, so none of the four comparisons is a close call.

Why the true errors make it worse

The second result reads as a contradiction of the essay before this, and it is not. The residuals really are smoother than the errors. What that essay does not claim is that smoother is worse for this statistic.

The construction builds each replicate as the benchmark’s fitted values plus a resampled residual. With the true errors the fitted part is nothing, so every replicate is a resampled error path and the whole statistic comes from it; with residuals, part of each replicate is a fixed component that does not change between replicates. The reference distribution is a distribution of a maximum over pairs of candidates, and a component held fixed across replicates does not merely shift it — it changes how far the maximum can travel.

So the two sources are not the same construction with more or less dependence in it. They are two different constructions, one of which happens to be closer to the truth for a reason that has nothing to do with fidelity. That is worth saying plainly: the ablation refutes the candidate, and the direction it moves in is not itself evidence for anything.

A fit takes the low frequencies out of what it leaves behind. The autocorrelation of the errors, of the residuals of a fitted benchmark, and of those residuals rescaled by their own leverage. (I − H) removes the component of the errors lying in a column space that is itself slow-moving, so the residuals are less persistent at every lag — by 5.9% at the first and 26.6% by the fourth. The leverage correction is the standard repair for what a fit does to a residual's size; drawn here against what it does to a residual's dependence, it does nothing.
Fig. 2 What the fit takes out before any resampling happens: the autocorrelation of the errors, of the residuals of a fitted benchmark, and of those residuals rescaled by their own leverage. The residuals are less persistent at every lag — by 5.9% at the first and 26.6% by the fourth — and the leverage correction does nothing to it.

What the fixed scheme does say

The coupling between neighbouring origins is real even though it is not the shortfall, and the measurement puts a size on it that is worth keeping.

The statistic under the rolling scheme has a 95% point of 4.0405; under the fixed scheme it is 6.4303, half as large again. That is the coupling, seen from the other side: a rolling window re-estimates the coefficients at every origin, and re-estimation on nearly the same rows produces forecast errors that move together and therefore an average that varies less between samples. Taking the re-estimation away lets each origin’s error be what the errors make it, and the maximum over pairs of candidates travels much further.

So the scheme changes the object by fifty per cent, and changes the resampling’s ability to reproduce the object by two points. A large effect on the target and almost none on the shortfall is exactly the signature of something that is not the explanation, and it is the reason the ablation is worth running rather than reasoning about: a mechanism that plainly matters this much is the obvious suspect, and it is not guilty.

The two-by-two read as two main effects

Three of the four cells are on the table and they separate the two candidates cleanly, because each change is made with the other held fixed.

Changing the scheme — rolling to fixed, with residuals throughout — moves the shortfall from 26.1% to 28.1%. Two points.

Changing the source — residuals to true errors, with the rolling scheme throughout — moves it from 26.1% to 38.8%. Twelve and a half points.

So the source is worth six times the scheme, and both make the shortfall worse. The coupling between neighbouring origins, which was the more sophisticated of the two candidates, is worth two points of a twenty-six point gap.

What the true errors being worse actually measures

The direction of the second effect is the surprising one and it is not a paradox once the triangle is in view.

A blocked resampling keeps (1k/)+(1-k/\ell)^+ of whatever dependence is in the series it is handed. Hand it a more persistent series and it discards more, in absolute terms, against a truth that has not moved. So a construction fed the true errors falls further behind than one fed residuals whose dependence has already been shrunk by the fit.

That inverts the usual reading of “the residuals are not the errors”. The residuals’ shortfall is not adding to the resampling’s shortfall; it is partly cancelling it, because both losses are losses of the same dependence and the resampling can only lose what is still there.

Taken at face value the two shortfalls put a number on the cancellation. If the loss were proportional to the dependence present, the ratio 38.8 / 26.1 = 1.49 would say the residuals carry about two thirds of the errors’ dependence. That is a crude reading — the relation between dependence and critical value is not linear — and it is the right order for a fit that has taken out a mean and a handful of coefficients from a hundred and twenty rows.

The shortfall is a share and not a level

One more thing falls out of the two residual cells, and it is what licenses quoting any of this as a percentage.

The two schemes are not measured against the same truth: 4.0405 rolling, 6.4303 fixed, which is 59% larger. The absolute shortfalls are correspondingly different — 1.0556 and 1.8044, a factor of 1.71.

And the shares are 26.1% and 28.1%, two points apart.

So the construction loses a roughly fixed fraction of whatever critical value it is aiming at, across a scheme change that moves that value by three fifths. That is what a multiplicative attenuation does and it is not what an additive error would do, which is the strongest evidence in the field that the residue is the triangle rather than any of the things a level shift would come from.

The bound

What is left is not a residue to be closed but a ceiling, and it is one line.

A wild-type resampling forms e*ₜ = eₜ·wₜ with a multiplier drawn independently of the residual, mean zero and unit variance. Then

E[etet+k]=etet+k  γw(k),γw(k)1,\operatorname{E}[e^{*}_t\, e^{*}_{t+k}] = e_t\, e_{t+k}\;\gamma_w(k), \qquad |\gamma_w(k)| \le 1,

so, averaged over the sample, the resample’s autocovariance is the residuals’ own, multiplied by the multiplier’s. Two things follow and neither is a matter of tuning.

A multiplier can only take dependence out. |γ_w(k)| ≤ 1, so the reference distribution’s dependence is bounded above by the residuals’ at every lag — and the residuals’ is already below the errors’ at every lag. The two shortfalls compose, and the ceiling is a product of them rather than either one.

The bound is attained only where the reference distribution is empty. γ_w(k) = 1 at every lag means one multiplier for the whole sample: a reference distribution built from a single sign, which has two points in it. So the block length is a dial along the bound rather than a route past it.

The ceiling a multiplier cannot reach pastA wild-type resampling forms e*_t = e_t·w_t with the multiplier independent of the residual, so what comes out has autocovariance γ_resid(k)·γ_w(k) — the residuals' own, multiplied by the multiplier's. Since |γ_w| ≤ 1 the reference distribution's dependence is bounded above by the residuals', and the residuals' is already below the errors'. The two shortfalls compose. For a block of ℓ the multiplier's autocorrelation is exactly the triangle (1 − k/ℓ)⁺, drawn here as the dashed prediction against the realised resamples at ℓ = 5; the bound is attained only at ℓ = n, where the reference distribution is built from one sign.00.2000.4000.6001234568lagautocorrelation of what the resampling producesthe errorsthe residuals — the ceilingwhat ℓ = 5 keeps80 samples, 40 resamples eachℓ = 5 keeps 73.6% of the first lag
Fig. 3 The errors’ autocorrelation, the residuals’ — which is the ceiling — and what a blocked multiplier actually produces, against the closed prediction. The block length is on a slider.

For a blocked multiplier the multiplier’s own autocorrelation is exactly the triangle (1 − k/ℓ)⁺: the share of pairs k apart that fall inside one block. So the whole prediction is closed, and the check is that the realised resamples reproduce it — which they do, at five block lengths and seven lags, with a worst departure of a hundredth and a half.

The triangle is worth a sentence on its own, because it is the part that makes the whole thing arithmetic. A blocked multiplier draws one sign per run of ℓ consecutive rows, so two rows k apart share a sign exactly when they fall in the same run — which, for k < ℓ, happens for ℓ − k of every ℓ starting positions. Their contribution to the resample’s autocovariance is preserved when they share a sign and cancels in expectation when they do not. Hence (1 − k/ℓ)⁺, with nothing estimated and nothing fitted.

That closed form is what turns the bound from a plausible inequality into something that can be checked. The realised resamples are measured against it at five block lengths and seven lags and match within two standard errors on all but one cell, where the last partial block at ℓ = 3 leaves a hundredth of a discrepancy — a boundary effect, and the only one.

Reading the dial as what it is

The numbers the bound puts on the sweep are worth having, because they turn a tuning parameter into an accounting.

The errors’ first-lag autocorrelation is 0.6900. The residuals’ is 0.6302, which is 91.3% of it and is the ceiling. What a blocked multiplier actually reaches at the first lag is 0.0013 at ℓ = 1 — nothing, correctly, since a fresh sign per row destroys everything — then 0.4114 at ℓ = 3, 0.5077 at ℓ = 5, 0.5692 at ℓ = 10 and 0.6013 at ℓ = 20. As shares of the errors’ own: 0.2%, 59.6%, 73.6%, 82.5%, 87.1%.

So at ℓ = 5, which is where the essay that measured the trade found the bias stopping and the spread still rising, the reference distribution carries 73.6% of the errors’ first-lag dependence. Push ℓ to twenty and it carries 87.1% — and the reference distribution is then built from five independent signs, which is what makes its own quantiles useless.

The dial does not run from bad to good. It runs from one failure to another along a ceiling that is below the truth at every point on it. That reframes the interior optimum the earlier essay found: it is not a compromise between a repairable bias and an unavoidable variance, it is the least bad point on a curve neither of whose ends reaches the target.

Two shortfalls, and only one of them is a dial

The accounting divides cleanly and it is worth doing once.

At the first lag the errors carry 0.6900 and the best block length measured reaches 0.6013, so 12.9% of the dependence is missing. Of that, 8.7 points are the residuals — what the projection removed before any resampling started — and 4.2 points are the taper the block length imposes. The smaller share is the one with the tuning parameter attached to it.

That ratio moves with the benchmark rather than with anything a practitioner adjusts afterwards. A benchmark with more persistent columns takes more of the low-frequency part of the errors out, which lowers the ceiling; a benchmark with white columns takes almost none. So the largest single decision about how well a blocked multiplier will do is which candidate is the null, and that decision is made for reasons that have nothing to do with reference distributions.

What a construction would have to do instead

The bound is about multipliers, and naming what it does not cover is the honest way to end.

It applies to any resampling of the form keep the residual where it is and multiply it — the wild bootstrap, Mammen’s two-point version, the blocked multiplier. It does not apply to a resampling that moves residuals, because then the pairing etet+ke_t e_{t+k} is not preserved and the autocovariance of the result is not the residuals’ times anything. A plain block bootstrap moves them, which is why it survives dependence and destroys the tie between a residual and its row — the trade the field’s table of resamplings is organised around.

Nor does it apply to a construction that generates errors from a fitted model of the dependence rather than resampling them: fit an autoregression to the residuals, draw fresh innovations, and produce a series with whatever autocorrelation the fit says. That has a different failure — it is a statement about the model of the dependence rather than about the data — and it is not bounded by the residuals’ own autocovariance, because it is not multiplying them.

Which of those is worth having in a world with two defects at once is a question this field does not settle. What it settles is that no choice of block length settles it, and that the quarter the blocked multiplier is short of the truth is not a tuning failure to be closed.

There is one more thing the bound is good for, and it is diagnostic rather than constructive. Because the prediction is closed — γ_resid(k) times a triangle — the shortfall of any particular run can be computed in advance, from the residuals alone, before a single replicate is drawn. A practitioner with a sample can ask what share of the residuals’ dependence a block of five will carry, and get an answer, and compare it with what a block of twenty would carry and how many independent signs each leaves. None of that requires knowing the truth. It requires knowing that the object being tuned has a ceiling, and where the ceiling is.

A reference distribution whose shortfall is computable is in a different position from one whose shortfall is a mystery, even when neither can be closed. The first can be reported.

What is claimed here, and what is not

This essay takes why a blocked multiplier is short of the truth, and the claims are that neither of the two named candidates explains it — the fixed scheme is 28.1% short against 26.1% for the rolling one, and the true errors make it 38.8% short — that the resample’s autocovariance is the residuals’ times the multiplier’s, that a blocked multiplier’s multiplier autocorrelation is exactly the triangle (1 − k/ℓ)⁺, and that no block length reaches more than 87.1% of the errors’ first-lag dependence against a ceiling of 91.3%.

What stays out and is named as a decision: a resampling that moves residuals rather than multiplying them, and a construction that generates errors from a fitted model of the dependence. Both escape the bound and both have failures of their own; neither is measured here, and pricing them against the blocked multiplier in a two-defect world is the obvious next thing.

The boundary against the essay that measured the dial is that it prices the trade and this one says what the trade is a trade along.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. The realised resamples are required to reproduce γ_resid(k)·(1 − k/ℓ)⁺ at five block lengths and seven lags, which is the closed prediction and is the only reason the bound can be read as arithmetic rather than as an observation. And no block length is allowed to exceed the residuals’ own autocorrelation, which is the bound itself and would be violated by any construction that put dependence in rather than failing to take it out.

The refusal for this essay is a block length defended by what it keeps. The share kept rises with ℓ and the number of independent signs the reference distribution is built from falls with it — a hundred and one at ℓ = 1 and four at ℓ = 30 — and only one of those two is usually in the sentence.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationBlock bootstrapClosed formCritical valueDependenceEstimation errorForecast errorHeteroskedasticityMonte CarloPersistenceReference distributionResamplingResidualRolling originWild bootstrap