The shape a dependence has

Where the generality runs out

A covariance that changes half way through a sample is not one a window can estimate. One number, a window and an order are worth the same as each other on it — and letting the model change once, at a point nobody can locate, is worth as much again as all three.

Worth reading first: The observations that repeat each other · A design is a number.

Every estimated covariance in this collection is a Toeplitz matrix. The estimate reads the sample autocovariance at each lag and puts it on the corresponding diagonal, so the covariance of two rows is a function of the gap between them and of nothing else. That is not a detail of the implementation; it is the assumption that makes an estimate possible at all, because a general n × n covariance has n(n + 1)/2 numbers in it and a sample has n rows.

So estimating the dependence rather than naming it buys generality in one direction: any decay in the gap, geometric or triangular or hyperbolic. This essay is about a departure in the other direction.

A covariance that changes

The errors here follow a first-order autoregression at 0.95 for the first sixty rows of a hundred and twenty and at 0.65 for the rest. Every marginal variance is one, so nothing has been rescaled; what changes is how much a row repeats the row before it.

Averaged over the pairs at each gap it reports 0.7987 at the first lag, which is the same first lag as the three stationary laws it is measured beside — so a rule told the errors are a first-order autoregression finds the same number here as it finds in a world with no break in it. There is nothing in a lag-one estimate that could report otherwise.

The three general rules are one rule

Generality in the wrong direction buys nothing. Regret on a sample whose persistence changes from 0.95 to 0.65 at row 60, over 200 draws. The three stationary rules — told one number, told a window, told an order — are within 0.4 standard errors of each other, and all three stop in the same place: they are general in the lag direction, and the departure is in the other one. Letting the model change once, at a point estimated from the same residuals, is worth 0.05021 more at 4.5 paired standard errors — about as much again as the whole of the first repair. Being told where the break is adds 0.01926, and being told the entire covariance adds 0.02465.
Fig. 1 Regret on a covariance that changes at row sixty. The three stationary rules are indistinguishable from each other, and letting the model change once is worth as much again as any of them.

Least squares with the ordinary penalty gives up 0.16925. The rule told the form recovers 35.1% of that, at 0.10991. A Bartlett-tapered covariance estimated from twenty lags reads 0.11058 and a whitening built from a fitted autoregression reads 0.10941: the paired difference between the first and the second is 0.00066 at 0.4 standard errors, and between the first and the third 0.00050 at 0.4.

One number, a window and an order. Three constructions that differ enormously when the dependence has a shape — where the window earns 0.00831 over the parameter and the fitted model earns 0.01173 — and here they are worth the same as each other to within a rounding. Whatever is stopping them is not something more lags would fix, and it is not something a better estimator of the same object would fix either, because all three are estimating different objects and arriving at the same wall.

The wall, with no sampling in it

The wall is computable without drawing anything. Build Ω_L from the sequence a sample of this length reports on average — the taper’s weight at each lag times the (1 − k/n) that comes from having n − k pairs to average — and apply its factor to the true covariance from both sides. If the window were right the result would be the identity, and how far it is from the identity is the bias half of the window’s trade, alone, with the variance taken out of the question entirely.

Where a wider window stops helping, and where it never does. The window's bias, alone, with no sampling anywhere in it: Ω_L is built from the sequence a sample of 120 rows reports on average — the taper's weight times the (1 − k/n) that comes from having n − k pairs at lag k — and its factor is applied to the true covariance from both sides. If the window were right the result would be the identity, and what is plotted is how far it is from one. Under all three stationary laws it falls the whole way, so a wider window would keep improving the whitening if variance were free, and the interior optimum in the sweep beside this one is therefore a fact about variance rather than about fit. Under the break it stops at L = 20 and turns back up: the object being approximated is not a function of the gap between two rows, and no matrix that is can get closer.
Fig. 2 The bias half of the trade under each law. Three of the four curves fall the whole way; the fourth has a floor and turns back up.

Under the three stationary laws that quantity falls all the way to the end of the range: at sixty lags it is 0.0900 under the geometric law, 0.3452 under the moving average and 0.2443 under long memory, and still falling. A wider window would keep improving those whitenings if variance were free, which is what makes the interior optimum in the window sweep a fact about variance rather than about fit.

Under the break it bottoms at 0.6570 at twenty lags and then rises, reaching 0.6806 at sixty. There is a closest Toeplitz matrix to this covariance and it is not close. No estimator of a Toeplitz matrix, however good, can do better than the best Toeplitz matrix — and the best one leaves seven times as much behind as the geometric law’s does.

Where 0.7987 comes from

The lag-one reading is not approximately the average of the two regimes; it is exactly the average a pair count produces, and working it out says precisely how much information the Toeplitz summary discards.

A hundred and twenty rows give 119 pairs at the first lag. Fifty-nine of them lie inside the first sixty rows and carry 0.95, fifty-nine lie inside the second and carry 0.65, and one straddles row sixty and carries whatever the correlation across the break is — 0.65, since the row after the break is generated at the new persistence.

59×0.95+59×0.65+0.65119=0.7987.\frac{59 \times 0.95 + 59 \times 0.65 + 0.65}{119} = 0.7987 .

The single crossing pair moves the answer by a thousandth. So the estimate is a pair-weighted mean of the two regimes and nothing else, and it would report the same 0.7987 if the two halves were interleaved row by row instead of split at row sixty. The estimate cannot see where the break is because a Toeplitz average has no place to put that information, which is the assumption the field is about, arriving at the first lag rather than at the twentieth.

The three rules recover the same third

Read as shares of what least squares gives up, the three general rules land on top of each other.

Against the ordinary penalty’s 0.16925, the rule told the form recovers 35.1%, the tapered covariance at twenty lags 34.7%, and the fitted autoregression 35.4%. Three constructions with nothing in common but their inputs, spread over seven tenths of a percentage point.

Which leaves 64.6% of the regret standing whatever is estimated and however well. That is the wall as a share rather than as a level, and it is the number worth carrying, because it does not depend on the scale the regret happens to be measured on.

The comparison with a world that has no break makes the collapse legible. There the window earns 0.00831 over the rule told the form and the fitted model earns 0.01173. Here the window earns −0.00066 — it is behind — and the fitted model earns 0.00050, which is four per cent of what it earns when the covariance holds still.

So generality is not merely failing to help. Two of the three rules are estimating strictly more than the parameter rule is, from the same sample, and converting that extra estimation into nothing: the window’s twenty lags and the sieve’s fitted order are both being spent on a sequence that is a pair-weighted average of two regimes, and no amount of accuracy about that average recovers a quantity it does not contain.

Room for the covariance to change

The repair is not more lags. It is one restart: whiten the first segment as one autoregression and the second as another, which makes the implied covariance block diagonal and lets the two blocks disagree.

Told where the break is, and left to estimate a lag-one autocorrelation on each side of it, the rule reads 0.04043 — a paired gain of 0.06948 over the rule told the form, and 76.1% of what least squares gives up recovered against 35.1%.

That is more than the entire first repair was worth. Whitening at all is worth 0.05933 against counting rows; letting the whitening change once is worth 0.06948 against whitening at one persistence. The second half of the sample is not a small correction to the first.

And the break point cannot be found

The obvious objection is that nobody is told where the break is. So estimate it: for every admissible position, fit the two segments, read off the concentrated Gaussian likelihood, and take the maximiser.

The estimated break leans towards the side with more memory. Where the profile likelihood puts the break, over 200 draws, on a sample that changes at row 60. When the persistent half comes first the estimate averages 74.6 — 14.6 rows late, at 17 standard errors. Turning the law over, so that the persistent half comes second, moves the average to 46.1, which is 13.9 rows early. A bias with the same sign both ways would be the trim or the estimator's arithmetic; one that reverses is the likelihood being steeper where the dependence is stronger, so the search hands that side rows it does not own.
Fig. 3 Where the profile likelihood puts the break, over two hundred draws, on a sample that changes at row sixty.

It is a poor estimate. The average is 74.6 where the truth is 60, its standard deviation across draws is 12.3, and it lands within twelve rows of the truth on 48.5% of them. On the evidence of a hundred and twenty rows the break is barely there: the two segments differ by a lag-one autocorrelation of 0.3, each is estimated from about sixty rows, and the profile is nearly flat across the middle half of the sample.

And the rule built on it works anyway. At the estimated break point the whitening reads 0.05970, recovering 64.7% of what least squares gives up, against the 35.1% every stationary rule reaches. Locating the break costs 0.01926 at 3.1 paired standard errors — real, and a third of what having the room is worth.

A break in the wrong place is still most of the repair. Regret when a two-regime whitening is told to change at each of seven positions, on a sample whose persistence really changes at row 60. The upper line estimates a lag-one autocorrelation on each side; the lower one is told both of them, and is minimised exactly where the break is, which is the check that the sweep is measuring what it claims. What matters is how shallow the curve is: at row 48 — twelve rows early — the fitted rule gives 0.06369 against 0.04043 at the truth, where the best of the stationary rules sits at 0.10941. Room for the covariance to change is worth several times more than knowing where it changes.
Fig. 4 Regret when the whitening is told to change at each of seven positions. The lower line is told both persistences; it is minimised where the break is, which is the check that the sweep measures what it claims.

The sweep over positions says why. Placed twelve rows early the rule reads 0.06369 and twelve rows late 0.04860, against 0.04043 at the truth — while every stationary rule sits above 0.109. The curve through the middle of the sample is shallow and the floor around it is not, so an estimate scattered over twenty rows spends most of its time somewhere that is nearly as good as being right.

Room for the covariance to change is worth several times more than knowing where it changes, and the practical form of that sentence is that a badly estimated break point is not an argument against allowing one.

Why this world is harder than its first lag says

One column of the table is worth explaining before anything is concluded from it. Least squares gives up 0.16925 here and about 0.077 in each of the other three worlds — more than twice as much — even though all four have the same lag-one autocorrelation. The break is not merely a world where the repairs work badly; it is a world where there is more to repair.

The reason is convexity and it is arithmetic. What dependence costs a fit is governed by the variance inflation (1 + ρ)/(1 − ρ), which is 9.0 at ρ = 0.8, 39.0 at 0.95 and 4.71 at 0.65. Half a sample at each of the last two averages 21.9, which is 2.43 times the inflation of a sample at the average persistence. The measured ratio of what least squares gives up is 2.21.

Those two numbers agree more closely than the argument deserves — the inflation formula is about the variance of a mean and the regret is about a table of fitted models — but the direction is not in doubt and the size is right. A persistence that varies is worse than the same persistence held at its own average, because the expensive end of the range is very expensive indeed, and averaging a convex function is not the function of the average.

It is the same shape as the arithmetic that makes a blocked design’s weights matter: a quantity that enters through a convex function cannot be replaced by its mean without changing the answer, and the error is always in the same direction. Here it means a reader who summarises a sample by one autocorrelation has understated how much dependence is costing, before any question about which rule to use has been asked.

Which way the estimate leans

The break estimate is not merely noisy, it is biased, and the direction is a mechanism rather than an accident. The profile likelihood is steeper on the persistent side — a segment at 0.95 gains far more from being modelled correctly than a segment at 0.65 — so the search hands the persistent side rows it does not own. With the persistent half first the estimate averages 74.6, fourteen and a half rows late, at seventeen standard errors.

That is a claim about a direction, and a claim about a direction can be refuted by turning the law over. Run the same estimator on a sample whose persistence goes from 0.65 to 0.95 at the same row and the average is 46.1, thirteen and nine tenths rows early, at seventeen standard errors the other way. A bias with the same sign both ways would have been the trim, or the estimator’s own arithmetic; one that reverses when the law does is the likelihood being steeper where the memory is longer.

What is left after the break is allowed for

Being told the whole covariance reads 0.01579, so the two-regime rule at a known break still gives up 0.02465 at 5.5 paired standard errors against it. That residue is worth naming because it is not the break.

Two things are in it. The first is that each segment’s persistence is estimated from residuals rather than from errors, and a fit removes dependence along with signal: the estimate is attenuated, and attenuation at 0.95 costs far more than attenuation at 0.65, because the variance inflation (1 + ρ)/(1 − ρ) is 39 at one and 4.7 at the other. The second is that the block-diagonal transform drops exactly one link — the pair straddling the break — out of a hundred and nineteen, which is the only approximation in the construction and is much too small to be the residue.

So what is left is the price of estimating two numbers rather than being told them, concentrated almost entirely in the number that matters most. The same asymmetry that biases the break estimate is what makes the residue large.

It also says where effort would go. Sharpening the break point is worth 0.01926 and cannot be done — the profile is flat and the data are what they are. Sharpening the persistence estimate in the segment that has most of it is worth more than that, and is an ordinary estimation problem with an ordinary answer: fit the autoregression to the segment jointly with the regression rather than to its residuals afterwards, so that the attenuation a fit introduces is not there to be inherited.

What is claimed here, and what is not

This essay takes the direction in which an estimated covariance is not general. The claims are that one number, a window and an order are within half a standard error of each other on a covariance that changes at row sixty, all recovering about 35% of what least squares gives up; that the closest Toeplitz matrix to this covariance leaves a floor of 0.6570 which no window reaches past, computed with no sampling in it; that a two-regime whitening at an estimated break recovers 64.7% and at a known one 76.1%; that the break point itself is estimated with a standard deviation of 12.3 rows and a bias of fourteen and a half towards the persistent side; and that the bias reverses when the law is turned over.

What stays out and is named as a decision: one break, in the persistence, at a point in the middle. A covariance that drifts smoothly rather than jumping, a break in the error variance rather than in the persistence, and two breaks rather than one are all outside this measurement, and the last of them is where the construction stops being obvious — a second break is a second search on the same rows, and the profile likelihood that already scatters over twenty rows would then be maximised over pairs of them.

The other thing that stays out is a penalty for having searched. The break point is chosen on the residuals the criterion then reads, and nothing here charges for that. It is the same shape as a criterion that reads a window chosen from its own sample and the same shape as an order chosen the same way, and the honest statement is that this field measures the first two and not the third.

The boundary against the essay that priced the four laws is that it asks what shape the dependence has and this one asks whether it has one shape at all.

The checks, and the refusal that made the field turn over

The refusal is the essay’s own title. A stationary estimate offered as the repair for a break is rejected: at 0.4 paired standard errors against the rule told one number, the generality that was bought is in the wrong direction, and it is refused in favour of the construction that has one change point in it and estimates that change point badly.

Two claims are gated. The bias-only mismatch is required to fall monotonically under every stationary law and to have an interior minimum under the break, which is the assertion that the wall is a fact about the object and not about the estimator. And the sweep over break positions, when the two persistences are given rather than estimated, is required to be minimised exactly where the break is.

That second gate is not decoration. In the first version of this field the break was defined as the half-way point of whatever length the series was drawn to, and the analysis reads the first hundred and twenty rows of a series drawn to a hundred and sixty — so the break sat at row eighty while every rule assumed row sixty. What the sweep reported was that the best place to break a sample is a fifth of the way past the truth, and the estimated break point agreed with it. Two instruments agreeing on an impossible number is what caught it; either one alone would have looked like a finding.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationBiasCovariance matrixEstimated varianceGeneralised least squaresInformation criterionModel misspecificationModel selectionNuisance parameterProfile likelihoodRegretStationarityStructural breakWhitening