A charge for a covariance's own dimension

What a window leaves free

A Bartlett window's weights sum to exactly half its width, which is a candidate for what the band costs. Varying the weights without varying anything else says the weights are the mechanism; varying the shape at the same weight says they are not the arithmetic.

Worth reading first: A design is a number · The observations that repeat each other.

The measurement one essay back says that a band of thirty lags costs 10.855 log-likelihood units where a criterion charges 29. It also offers an explanation — the band’s entries are shrunk by the window’s weights, and a shrunk parameter is worth the shrinkage factor rather than one — and that explanation predicts 14.500. Being wrong by a quarter is a better outcome than being wrong by a factor of three, and it is still being wrong.

So the explanation has to be tested rather than accepted, and testing it means finding a way to change the weights while changing nothing else. This essay does that, and the answer separates into two halves that point in different directions. The weights are the mechanism. Scale them and the charge scales with them, exactly. The weights are not the arithmetic. The level is three quarters of the sum, a window of a different shape sits somewhere else, and the control that would pin down why is unavailable in principle.

The claim, stated so it can fail

A plug-in tapered autocovariance is γ̂(k) w(k, L): the free estimate multiplied by a fixed number below one. That is a linear shrinkage towards zero, and the effective dimension of a linear shrinkage is the sum of its factors, because the derivative of the estimate with respect to the thing being estimated is the factor itself.

For the Bartlett window the sum has a closed form and it is exact:

k=1L(1kL+1)=L2\sum_{k=1}^{L}\left(1-\frac{k}{L+1}\right) = \frac{L}{2}

Half the width, at every width, to machine precision — 15 at a width of thirty, 4 at a width of eight, and no approximation in the argument. For the Parzen window the sum is 0.375(L + 1) − ½, which is 11.125 at a width of thirty and agrees with the summed weights to four decimals from about eight lags up.

The area under the window is what the band actually costs. The three windows' weight sequences at a width of 30 lags, drawn against the lag as a share of the window. A truncated window applies a weight of one to every lag inside it and zero outside, which is why its sum is the width and why every conventional charge is right for it — and it is a covariance matrix on almost no sample, so it cannot be used. The Bartlett window falls linearly to zero and its weights sum to exactly 15.000000000000004, which is half the width, at every width: Σ(1 − k/(L+1)) over k = 1 … L is L − L/2. The Parzen window sums to 11.13 here, three eighths of the width, and it gets there by holding a weight near one over the first few lags and then falling faster. A plug-in estimate multiplied by a weight below one is a shrunk estimate, and a shrunk estimate is worth less than a free one — which is the whole of why a charge levied per lag is a charge for parameters the window has already spent.
Fig. 1 The weight sequences. The area under each curve is the quantity the claim is about, and the two shapes reach comparable areas by very different routes.

The claim, then, is that the measured optimism equals the summed weights. Measured, it is 10.855 ± 0.281 against 14.500 — thirteen standard errors low. The claim as stated is false. What is left is the question of whether the weights are doing anything, and there is a clean way to ask it.

Why an effective dimension is the right kind of object at all

It is worth being explicit about what is being claimed, because effective degrees of freedom is a phrase that gets used for several different quantities and only one of them is on offer here.

The quantity a criterion needs is the optimism: how much the objective read on the sample it was fitted to overstates the same objective read on a sample it was not. For an ordinary fit that quantity is a trace, and the essay that derives it shows the trace collapsing to the parameter count exactly when the fit has that many free directions and the rows are independent. The collapse is what makes a count usable; the trace is what is true.

A linear shrinkage has a trace too, and it is the sum of the shrinkage factors. Multiply an estimate by 0.6 and it moves six tenths as far in response to the same noise, so it can chase six tenths as much of it. Nothing in that argument requires the fit to be linear in the data or the objective to be a residual sum — it requires only that the estimate be a fixed multiple of the free one, which γ̂(k) w(k, L) is by construction.

So the object is right. What is in doubt is whether the rest of the setting — an objective with a determinant in it, a parameter that enters a matrix inverse, a sample whose rows are correlated by the very thing being estimated — leaves the argument intact. That is not a question a derivation settles, and it is why the family below exists.

The ratio is three quarters, to within a twentieth of a standard error

“Three quarters of the sum” is a description in the essay’s opening and it is worth testing as a number, because a clean fraction is either a clue or a coincidence and only measurement can say which.

The measured optimism is 10.855 ± 0.281 against a summed weight of 14.500, a ratio of 0.7486. The standard error carries across as 0.281 / 14.5 = 0.0194.

So the ratio is 0.749 ± 0.019, and exactly three quarters sits 0.07 standard errors away. On this measurement alone, a claim that the optimism is three quarters of the summed weights cannot be distinguished from the data at all.

That is a much sharper statement of the failure than “thirteen standard errors low”. The summed-weight claim is not merely wrong; it is wrong by a factor that is itself a simple fraction, which is what a missing factor in a derivation looks like rather than what an approximation breaking down looks like.

What the fraction has to explain, and what rules it out

Two constraints on any explanation follow from what is already measured, and together they are quite restrictive.

It must survive scaling. Scaling the weights scales the charge exactly, which the field establishes directly, so whatever produces the three quarters is homogeneous of degree one in the weight sequence. A fraction cannot come from a threshold, a clamp or anything else that would break under a rescaling.

It must depend on shape as well as sum. If the ratio were a universal three quarters, the Parzen window’s summed weight of 11.125 at the same width would predict an optimism of 8.34, and the field reports that a window of a different shape sits somewhere else. So the constant is not a constant across sequences with the same sum.

Those two together say the quantity is a homogeneous functional of the weights that is not the sum — a weighted sum with a second sequence in it, or a ratio of two sums of the same degree. That is a narrow enough family to be worth naming even without identifying which member it is.

One last figure for scale. A band of thirty lags has 29 free entries and its measured optimism is 10.855, so the window is removing 62.6% of the parameters the band would otherwise cost. The summed-weight rule says it removes fifty per cent. The rule is not a small correction to the naive count; it is most of the way from 29 to 10.9, and the argument here is about the last quarter of the distance rather than about the first three.

Three windows that differ only in how much they shrink

Take the Bartlett weight and raise it to a power. The square and the cube of a sequence falling linearly to zero fall to zero the same way and more steeply, and their sums are smaller in a known ratio: Σ(1 − k/(L+1))^p is about (L + 1)/(p + 1), so the three of them stand at roughly 6 : 4 : 3.

Nothing about that is safe by default — an arbitrary sequence is not a covariance sequence, and a window whose matrix has a negative eigenvalue produces no objective value at all rather than a bad one. These three are safe, and for a reason worth naming. A Bartlett-weighted Toeplitz matrix is positive semidefinite, and the elementwise product of two positive semidefinite matrices is positive semidefinite; that is Schur’s theorem, and it makes every integer power of the Bartlett window a legitimate window. Their availability rate is a hundred per cent over every draw in the sweep, which matters more than it sounds — a window that survives on some draws and not others is measured on a selected subsample, and the last section of this essay is what that looks like.

At thirty lags the three sums are 14.500, 9.589 and 7.133 above the narrowest band, and the three measured optimisms are 10.855 ± 0.281, 7.390 ± 0.323 and 5.574 ± 0.329.

Four windows, one line, and one that is off it. The optimism measured for each window at a band of 30 lags, against what that window's weights sum to, on 2000 pairs of independent samples of 120 rows. The diagonal is where a window that spent exactly its summed weights would sit. Three of the points are one shape at three levels — the Bartlett window, its square and its cube, whose sums stand in the ratio 6 : 4 : 3 — and they lie on a line through the origin at 0.767 of the diagonal, with 0.033 between the highest and the lowest. Scaling the weights scales the charge by the factor the weights predict, which is what makes the weights the mechanism. The Parzen window has a comparable sum and a different shape, and it sits at 0.871: its weights stay near one over the first few lags, and the first few lags are where the information is. A weight sum treats every lag as equally informative and no sample does.
Fig. 2 Each window’s measured charge against its own summed weights. The dashed diagonal is where a window spending exactly its weights would sit; the solid line is fitted to the three points of one shape.

Divided through, the three shares are 0.7486, 0.7707 and 0.7815 — an average of 0.7669 with 0.0329 between the highest and the lowest, against standard errors of 0.019, 0.034 and 0.046. The weight sums differ by a factor of two and the shares do not differ at all. Three points on a line through the origin, with one constant fitted and nothing else.

That is what it means for the weights to be the mechanism. A window that shrinks harder costs less, by exactly the factor it shrinks harder by, across the whole range the family spans.

And one that differs in where the weight sits

The Parzen window sums to 10.875 above the narrowest band, which sits between the Bartlett square and the Bartlett itself. If the sum were the whole story, its measured charge would sit between theirs on the same line.

It does not. It measures 9.471 ± 0.363, a share of 0.8709 ± 0.0333 against the family’s 0.7669 — higher by 0.1039, at 3.12 standard errors.

How much of its own weight sum each window spends. The measured optimism as a share of the summed weights, for four windows at a band of 30 lags, over 2000 pairs of independent samples of 120 rows. One would mean the weights are the arithmetic; the three Bartlett powers read 0.749, 0.771, 0.781, which averages 0.767 and spans 0.033. Their weight sums differ by a factor of two and their shares do not differ at all, so the weights carry the whole of the difference between those three windows and about three quarters of the level. Parzen reads 0.871 ± 0.033, which is 0.104 above the family at 3.1 standard errors: the same amount of weight, put at the lags where the sample knows most, costs more. So a derived charge needs the window's shape and not only its area, and no charge that reads L alone can have either.
Fig. 3 The same four windows as shares of their own summed weights. The three of one shape agree to 0.0329; the fourth is a tenth above them.

The mechanism is legible in the weight figure above. The Parzen window holds a weight near one over the first few lags and then falls away fast; the Bartlett window starts falling immediately. Two windows can reach the same total while putting it in different places, and the places are not interchangeable, because the first few lags are where the sample knows most. A sample of a hundred and twenty rows estimates γ(1) from a hundred and nineteen products and γ(29) from ninety-one, and the estimates at short lags are the ones a persistent process makes precise. Weight concentrated there is weight on well-determined directions, which is exactly the weight a criterion should charge for.

So the sum is not enough. A derived charge needs the window’s shape as well as its area — and no charge written as a function of L alone can carry either.

The control that would settle it, and cannot be run

There is one window where the two candidate answers coincide by construction. Truncate rather than taper: every weight inside the band is one, the sum is the width, and a measurement of the optimism has to come back at the width or the measurement is measuring something other than optimism. That is the control this field wanted.

The one window a convention is right for cannot be measured. How often each window's estimate is a covariance matrix at all, over 1200 draws of 120 rows out to 3 lags. A truncated band applies a weight of one to every lag it keeps, so the summed weights are the width and the derived charge and the convention coincide by construction — which makes it the control this field wanted. It survives on 0.5% of draws. What is left is the subsample where a truncated sample autocovariance sequence happened to be positive definite, and it reports 7.76 log-likelihood units of optimism for 2 lags — several times more than a band of that width could carry, which is what a selected subsample looks like from the inside. The tapered band survives on 100.0%, because the Bartlett window is positive semidefinite for every sequence there is.
Fig. 4 How often each estimate is a covariance matrix at all, out to three lags. The control is the row that is not usable.

Over twelve hundred draws out to three lags, a truncated sample autocovariance sequence is a covariance matrix on 0.5% of them. Six draws survive. Those six report 7.76 log-likelihood units of optimism for two lags — several times more than two lags could carry under any theory of what a lag costs — because they are not a sample of draws, they are the sample of draws on which a sequence that is usually not positive definite happened to be. The tapered band survives on 100.0% of the same draws.

That is the finding rather than a gap in it. The field that put the window there established that truncation destroys the one property that makes a band usable, and the consequence here is a second one: the only window a conventional charge is exactly right for is a window nobody can use, and the charge is therefore untestable at the place it is correct. A convention that is right where it cannot be checked and wrong everywhere it can be is not a convention with an exception in it; it is a convention derived for a different problem.

What the family cannot separate

Two things move together in the power family and one comparison cannot pull them apart, which is worth stating before the shares are read as more than they are.

Raising the Bartlett weight to a power shrinks every lag, and it shrinks the far lags proportionally more than the near ones — that is what a power does to a number below one. So the square is not only a smaller-sum window, it is also a window whose weight has moved towards the origin relative to its own sum. If short lags cost more per unit of weight, as the Parzen comparison says they do, then the square should spend a slightly larger share of its sum than the Bartlett does.

It does: 0.7707 against 0.7486, and the cube is higher again at 0.7815. The three shares are ordered, and they are ordered in the direction the shape argument predicts. The spread is 0.0329 against standard errors of 0.019 to 0.046, so the ordering is well inside noise and cannot be called a finding — but it is not evidence against the shape argument either, and reporting the three as identical would be reporting the convenient half of it.

The honest statement is that the family holds the shape nearly fixed and not exactly fixed, which is the weaker version of the design and the one it actually implements. A family that varied the sum with the shape held exactly constant would need a window scaled by a constant rather than raised to a power, and a Bartlett window multiplied by 0.6 is not a covariance sequence with a unit diagonal — the diagonal is the one entry the normalisation fixes at one, so scaling the off-diagonals is scaling the correlation, which is a different object. That is a real obstruction rather than an oversight, and it is why the design is powers.

What a wide band and a narrow band do not share

One more reading falls out of the sweep and complicates every number above.

The share is not constant in the width. Across the Bartlett sweep it runs from 0.9528 at two lags to 0.7486 at thirty, falling monotonically: 0.8292 at four, 0.7921 at eight, 0.7732 at twenty. At the narrow end the weights are very nearly the arithmetic and at the wide end they overstate by a quarter.

That has a consequence for anything built on this. The right charge is not linear in L, so a rule of the form charge c times the summed weights is a straight line fitted to a curve, and where it is fitted decides what it says. Fitted at the wide end it charges too little near the origin; fitted at the narrow end it charges too much far out. Every charge considered in this field, including the two derived ones, is a straight line through the origin, and the width table is read at widths where the discrepancy is a few per cent of the charge rather than a few per cent of the width.

Why the share should fall is the same question as where the missing quarter goes, one level down. A band’s far lags are estimated from fewer products and are correlated with each other more strongly than its near ones are, so a wide band’s parameters are less nearly independent than a count of them assumes — and a trace over correlated directions is smaller than a count of them. That is a mechanism rather than a measurement, and it is in the same place as the level: named, plausible, and not established here.

What survives

Three statements, in decreasing order of how well they are established.

Scaling a window’s weights scales its charge by the same factor. Established at 0.0329 of spread across a factor of two in the weights, on one shape, with independent samples on both sides of the comparison.

A window’s shape moves the charge at a fixed sum, in the direction that says short lags are the informative ones. Established at 3.12 standard errors on one comparison, which is one comparison.

And the level is about three quarters of the sum and is not derived. The candidates for the missing quarter are all shrinkages the measurement inherits rather than imposes: the sequence is normalised by γ̂(0), the sample autocovariance divides by n rather than by n − k, and the whole thing is computed from residuals rather than from errors — each of which shrinks the estimate on its own. Three compounding shrinkages of a few per cent apiece would not reach a quarter; the largest of them is the residual one, and separating it needs a design that varies one at a time.

What a band of lags actually costs, against what it is chargedThree quantities against the width of a tapered band, on 2000 pairs of independent samples of 120 rows under a five-period moving average. The upper line is what a criterion charges — one log-likelihood unit a lag, which is Akaike's penalty applied to the band as though every lag were a free coefficient. The middle line is what the window's own weights predict, Σ(1 − k/(L+1)), which is exactly half the width. The lower line with its error bars is the measurement: the objective at the fitted covariance, minus the same objective on a second sample drawn independently from the same law, differenced from the narrowest band so that the coefficients' own optimism drops out. At 30 lags it is 10.812 ± 0.235 against a charge of 29 — 2.68 times too large — and against 14.50 predicted by the weights. The measurement is nearer the weights than the convention and is under both, at every width in the sweep.01020301481216202430lags in the bandlog-likelihood units of optimismwhat a criterion chargeswhat the window's weights predictwhat it measures2000 independent pairs, 120 rows each2.68× at the widest band
Fig. 5 The sweep the shares are read off, at the moving average rather than the autoregression. The slider changes the law; the share is not the same under all four.

The last of those is the one that should be stated loudest, because the per-law figure in the essay that measured the charge shows the share itself moving: 0.738 under a first-order autoregression, 0.746 under a moving average, 0.834 under long memory and 0.944 under a break. The share is not a constant of the window either. It is a constant of the window and the law, and under the law a band represents worst it is very nearly one.

That last number is the one that would repay another round. A break in the persistence is the law under which a band of lags is most misspecified and the law under which its lags cost most nearly what a criterion charges for them — which is either a coincidence or the clearest statement available of what the missing quarter is.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BandwidthClosed formCovariance matrixDegrees of freedomEigenvalueInformation criterionInformation matrixLong-run varianceModel misspecificationOptimismPersistencePlug in estimateSample autocovarianceShrinkageTapering