Errors generated from a fitted model
Worth reading first: The experiments that could have happened · Where the bootstrap lies.
Two constructions were named as escaping the multiplier’s ceiling. The first is not an escape: a fixed-length block bootstrap has the same triangle, because the attenuation is the block boundary rather than the multiplier — and the construction this collection actually runs under that name has a geometric taper instead, which turns out not to move its critical value either. The second is, and the reason is a difference in kind rather than in degree.
Every resampling considered so far builds its new error series out of the residuals. A multiplier reuses each one in place; a moving block reuses runs of them somewhere else. Both are limited by what the residuals contain, which is a truncated, noisy version of what the errors contained, and the residuals are systematically smoother than the errors besides.
A sieve does not resample the residuals at all. It fits an autoregression to them and generates a fresh series from the fit. What the resample carries is then the model’s autocovariance — an infinite, smoothly decaying sequence — rather than the sample’s. That is the whole difference, and it has two consequences that pull opposite ways.
The sequence a model has, against the sequence a sample has
At the first lag the sieve carries 0.6198 where the residuals have 0.6409 — slightly below, which is what it must be: a Yule–Walker fit reproduces the sample autocovariances exactly out to the order it was fitted at, and a finite series generated from it recovers them with a little attenuation of its own.
At the sixth lag the sieve carries 0.0638 where the residuals have 0.0300 and the blocked multiplier has −0.0011. It has more dependence at that lag than the series it was fitted to.
That is not an error; it is the point. A fourth-order fit is told about four lags and then implies every lag after them, by its own recursion. The implied sixth-lag autocorrelation is a smooth continuation of the first four rather than a noisy sample quantity that happened to come out near zero — and on a design whose true errors are autoregressive at 0.7, the truth at the sixth lag is 0.1176. The residuals report 0.0300 and the sample errors themselves report 0.0389, both of them badly short because a sample autocorrelation at a long lag on a hundred rows is mostly noise. The sieve’s 0.0638 is closer to the truth than either.
A model extrapolates, and a truncated sample sequence cannot. Everything in this essay follows from that sentence, in both directions.
Where the sieve crosses the residuals, which is almost immediately
The two readings quoted are at the first lag and the sixth, and between them they pin down the whole shape — including the lag at which the sieve stops being below the series it was fitted to and starts being above it.
At the first lag the sieve carries 0.6198 against the residuals’ 0.6409, a shortfall of 3.3%. That is the generation attenuation: a finite series drawn from a fitted model recovers the model’s own autocovariances a little short, for the same reason any finite series does.
At the sixth the sieve carries 0.0638 against the residuals’ 0.0300 — 2.1 times as much.
Now read each as a decay rate. From the first lag to the sixth the residuals fall by a factor of 0.0468 over five lags, a rate of 0.542 a lag. The sieve’s own extrapolation is geometric at the rate its fit implies: , and after the 3.3% handicap that is 0.0670 — close to the 0.0638 measured, so the fit is behaving like a low-order model decaying at about 0.64 a lag.
Setting the two curves equal puts the crossing at lag 1.2. So the sieve is below the residuals at the first lag only, and above them from the second onwards.
The escape is not something that happens out in the tail. It happens immediately, and by the sixth lag it is a factor of two.
Three constructions at one lag
The sixth lag is the cleanest place to see all three side by side, because it is past the block length.
- The blocked multiplier: −0.0011. Dead, as the triangle requires — a construction with a block of five keeps , which is zero from the fifth lag on, and what is left is noise.
- The residuals: 0.0300. What the series being resampled actually contains.
- The sieve: 0.0638. Twice the residuals, from a model that was fitted to them.
The middle row is the ceiling every resampling of the residuals is under, and the first row is a construction sitting a long way below it. The third row is above the ceiling, which is the only thing in this collection that is, and it is above it because the quantity it carries was never a function of what the residuals reported at that lag — it is a continuation, and a continuation can exceed its own anchor wherever the anchor has fallen faster than the model says it should.
Which closes a shortfall no block length reaches
The essay that found the ceiling leaves a fifth of the truth unaccounted for and rules out two explanations for it by ablation. What is left is the bound: a multiplier’s autocovariance is the residuals’ times something no larger than one, so it cannot reach a truth the residuals themselves fall short of.
Take the bound away and the shortfall goes with it.
Three per cent over, against a standard error of about three per cent on the critical value: the sieve is exact in that world to the accuracy the measurement has. Its test rejects a true null 0.5% of the time at a nominal 5%, which is conservative in the same direction as everything else in this family and is much closer to the level than any resampling of residuals gets.
The gap that four rounds of block lengths could not close is closed by not resampling the residuals.
And pays for it where there are two defects
The second world is the one this collection keeps returning to: the error variance depends on the design and the rows repeat each other. There the sieve falls to 0.922 of the truth — short by 7.8% — and the blocked multiplier improves to 0.883.
The sieve is still the best of the four, and it has lost most of its advantage. The reason is the same one that undoes both block bootstraps: a sieve generates its errors from a model with a single innovation variance, so every row in the resampled series carries the same variability. On a design where the real error variance is a function of the covariates, the reference distribution is built from a homoskedastic world.
So the inventory comes out like this, and it is worth stating as an inventory because the three constructions are usually discussed as though one of them were simply better:
- a blocked multiplier keeps the row and truncates the dependence;
- a fixed-length block truncates the dependence in exactly the same way and loses the row;
- a stationary bootstrap truncates it geometrically instead and loses the row;
- a sieve keeps the dependence and loses the row.
Nothing here keeps both. That is not a gap in the search; it is close to a statement about what a residual is. To keep the row is to reuse the residual where it was, and to reuse it where it was is to be limited by what it contains.
The failure that is a misspecification rather than a ceiling
There is a second difference between the sieve and the others and it is the one that should make a practitioner careful.
A multiplier’s shortfall is a ceiling: it is bounded by a quantity that can be computed from the sample in hand, it is in a known direction, and no choice of dial escapes it. That is an unpleasant property and a predictable one.
A sieve’s error is a misspecification. The resample’s autocovariance is whatever the fitted model says, which is right to the extent the model is right and unbounded to the extent it is not. On this design the errors really are autoregressive of low order, so a fourth-order sieve is nearly correctly specified and its extrapolation is nearly right. On errors with a moving-average component, or long memory, or a break, the same extrapolation would be confidently wrong at exactly the long lags where it is doing the work — and there would be nothing in the output to say so, because a sieve resample looks like a series either way.
A construction whose error is bounded and a construction whose error is a model’s error are not comparable by a single number, and the numbers above are from the world where the model is right. Naming that is more useful than the numbers.
There is a version of this trade in every field of this collection and it is usually the other way round. A parametric method beats a non-parametric one when its model is right and loses when it is not; the non-parametric one is slower and safer. What is unusual here is that the safe construction is not merely slower — it is bounded away from the truth by an amount that does not shrink with any dial, and the bound is a property of the residuals rather than of the sample size. So the usual advice, that a practitioner unsure of the error structure should take the non-parametric route, buys a guarantee of being about a fifth short rather than a guarantee of being right eventually.
That is not an argument for the sieve. It is an argument for saying which of the two failures a particular analysis can live with, and the two are different enough that the question has an answer: a test whose critical value is short by a known fraction in a known direction is over-rejecting in a way that can be reported, and a test whose critical value depends on whether a fitted order captured the error process is not.
What the two worlds are, and why there are only two
Every number in this essay comes from one of two worlds, and it is worth saying what they are because the choice of two rather than four is deliberate.
The first has persistent errors and persistent predictors and nothing else: the variance is constant across rows, so a construction that moves a residual off its row loses nothing. It exists to isolate the dependence, which is the quantity the whole field is about.
The second adds a variance that is a function of the first predictor, scaled so the average variance is unchanged. It exists because that is the combination the table of nine has no column for, and because it is the combination a real regression on a time series usually has.
The two worlds a fuller table would add — heteroskedasticity alone, and neither defect — are already measured in that field, and every resampling here behaves there as it does there. Repeating them would add rows and no argument.
The check the escape needed
The claim that a sieve reproduces its sample out to its own order and extrapolates past it is an identity rather than an approximation, and it is checked as one. A fit of order p to a series, then the fitted process’s own autocovariance computed from its coefficients by the Yule–Walker recursion, is required to reproduce the sample’s autocovariances at lags 0 through p to within a hundred-millionth.
That check is what separates the sieve carries more dependence from the sieve carries different numbers. If the model’s sequence did not agree with the sample’s where the sample told it anything, the excess at lag six would be an artefact of the fit rather than an extrapolation from it.
The recursion is Levinson–Durbin rather than a least-squares fit, and the reason is worth one sentence: conditional least squares can return coefficients whose implied process is not stationary, and a resample generated from a non-stationary fit has no autocovariance to compare against at all. The Yule–Walker solution cannot, because every partial autocorrelation it produces is inside (−1, 1) whenever the sequence it is given is positive definite — which is the property the sample autocovariance sequence has and its truncation does not.
A ceiling and a misspecification are not two sizes of the same risk
The sieve escapes the bound, and it is worth being precise about what it has traded for that, because the two risks are usually compared as though they were commensurable and they are not.
A ceiling is bounded, known in advance, and computable from the sample. A construction that reuses residuals one at a time cannot carry more dependence than the residuals carry, and the residuals’ sixth-lag autocovariance is 0.0300 against a truth of 0.1176. That shortfall does not depend on anything unobserved. It can be written down before the resampling runs, it is the same for every dataset of that shape, and a reader who knows the construction knows the number.
A misspecification is none of those things. The sieve generates from an autoregression’s own autocovariance, which is an infinite decaying sequence rather than a truncated one — 0.0638 at the sixth lag, where the multiplier has −0.0011 — and that is exactly why it is not bounded by the sample. But the same property is the risk: what it carries is the model’s dependence, and if the model is wrong the error is a function of how wrong, which is unobservable, unbounded, and does not announce itself in any diagnostic the resampling produces. A sieve fitted to a process with a moving-average component, or a break, or long memory, will produce a confident reference distribution for a dependence the data does not have.
The measurements here show both faces in one table. In a world whose only defect is dependence the sieve is exact to the measurement’s own accuracy, at 1.029 of the truth against a blocked multiplier’s 0.819. Add a variance that depends on the design — a defect its single innovation variance cannot represent — and it gives most of that back, at 0.922 against 0.883, landing barely ahead of the construction it comprehensively beat one column earlier. Nothing about the sieve changed. The world did, in a direction its parameterisation has no room for.
So the choice is not between a worse method and a better one. It is between an error that can be computed and an error that cannot, and which of those is preferable depends on how much is known about the dependence before the resampling starts — which is the same question the whole field keeps returning to, one level further down.
What is claimed here, and what is not
This essay takes the one construction that escapes the multiplier’s ceiling. The claims are that a sieve resample carries the fitted model’s autocovariance rather than the sample’s, so it exceeds the residuals at long lags — 0.0638 against 0.0300 at the sixth, where the truth is 0.1176; that it is exact to the measurement’s own accuracy in a world whose only defect is dependence, at 1.029 of the truth against a blocked multiplier’s 0.819; and that it gives most of that back where the error variance also depends on the design, at 0.922 against 0.883, because it generates from one innovation variance and so builds a homoskedastic reference for a heteroskedastic world.
What stays out and is named as a decision: the order. Every measurement here uses a fourth-order sieve on a design whose errors are first-order autoregressive, so the model is correctly specified and generously so. Choosing the order from the residuals is a selection problem with the same shape as choosing a window, it would have to be priced on the same table, and the sieve’s whole advantage lives in the extrapolation that the order controls. What a misspecified sieve costs is therefore the obvious next measurement and is not made here.
The boundary against the essay before this one is that it is about a construction that was expected to escape and does not, and this one is about the construction that does.
The checks, and the refusals that make them mean something
Two claims are gated. A fitted autoregression’s own autocovariance is required to reproduce the sample’s exactly out to the order it was fitted at, to a hundred-millionth, which is what makes the excess at longer lags an extrapolation rather than an artefact. And the sieve is required to carry more dependence at some lag than a blocked multiplier does, because that is the escape, and a sieve that failed to would mean the ceiling was not about the construction after all.
The refusals in this field are the two that bracket the sieve on either side: a truncated covariance estimate used without asking whether it is a covariance matrix, and a fixed-length block resample offered as an escape from a bound it sits exactly on.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- How long a block a multiplier shares — both name autocorrelation, block bootstrap, critical value, error rate, heteroskedasticity, reference distribution, resampling, residual, wild bootstrap
- Residuals that keep their own variance — both name block bootstrap, bootstrap, critical value, error rate, heteroskedasticity, reference distribution, residual, wild bootstrap
- A block weighted inside itself — both name autocorrelation, block bootstrap, dependence, reference distribution, resampling, residual, wild bootstrap
- A length for each instrument — both name block bootstrap, critical value, dependence, estimation error, persistence, reference distribution, resampling
- The instrument and the reading — both name block bootstrap, critical value, dependence, estimation error, persistence, reference distribution, resampling
- The reversal that was the instrument's — both name block bootstrap, critical value, dependence, estimation error, persistence, reference distribution, resampling
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationBlock bootstrapBootstrapCritical valueDependenceError rateEstimation errorHeteroskedasticityIndependenceModel misspecificationPersistenceReference distributionResamplingResidualWild bootstrap