A charge for a covariance's own dimension

The charge nobody derived

A band of lags is charged one log-likelihood unit apiece, because that is what a regression coefficient costs. A band's numbers are not regression coefficients, and measuring what they actually cost puts the convention out by a factor of nearly three.

Worth reading first: A design is a number · The observations that repeat each other.

A tapered band is a parameterised family, so the width of the band is a selection problem, and the essay that established that closes by choosing the width with the two penalties every selection problem in this collection has used. Akaike’s charges one log-likelihood unit a lag. Schwarz’s charges half a log n, which at a hundred and twenty rows is 2.394. Four defensible readings of a flat objective, three widths a factor of three apart, and a shrug.

The shrug is the wrong response and the reason is one sentence long. Neither charge was derived for this parameter. Both are derived for a coefficient in a regression: a number the fit may move anywhere it likes, whose contribution to the difference between an in-sample fit and an out-of-sample one is exactly one unit because that is what a free direction is worth. A band’s numbers are not free, and they do not enter the objective the same way. Nobody has ever checked what they cost.

This essay checks. The answer is 0.374 of a log-likelihood unit a lag where the convention charges one — a factor of 2.672 — and it is not a constant either, because it moves with the law the errors follow.

What the band’s parameters actually are

Ω̂ is banded Toeplitz with bandwidth L, so it is described by γ(1) … γ(L) with the diagonal fixed at one. That is L numbers, which is where the charge of L comes from and is the only part of the reasoning that survives inspection.

The numbers are not estimated freely. What goes into the matrix is the tapered sample autocovariance γ̂(k) w(k, L), where w is a fixed sequence falling from one at the origin to zero at the edge of the window. The field that made the band a covariance matrix at all put that sequence there for a reason unrelated to charges — a truncated sequence is not a covariance matrix on most draws, and a Bartlett-weighted one is a covariance matrix on every draw there is — but the side effect is that every entry of Ω̂ is a shrunk estimate rather than an estimate.

The area under the window is what the band actually costs. The three windows' weight sequences at a width of 30 lags, drawn against the lag as a share of the window. A truncated window applies a weight of one to every lag inside it and zero outside, which is why its sum is the width and why every conventional charge is right for it — and it is a covariance matrix on almost no sample, so it cannot be used. The Bartlett window falls linearly to zero and its weights sum to exactly 15.000000000000004, which is half the width, at every width: Σ(1 − k/(L+1)) over k = 1 … L is L − L/2. The Parzen window sums to 11.13 here, three eighths of the width, and it gets there by holding a weight near one over the first few lags and then falling faster. A plug-in estimate multiplied by a weight below one is a shrunk estimate, and a shrunk estimate is worth less than a free one — which is the whole of why a charge levied per lag is a charge for parameters the window has already spent.
Fig. 1 The three windows’ weights across the band. A truncated window applies one to every lag it keeps, and its sum is the width; a Bartlett window’s sum is exactly half the width, at every width.

A shrunk parameter is worth less than a free one. That is the whole content of an effective degrees of freedom, and for a linear shrinkage the amount it is worth is the sum of the shrinkage factors. For the Bartlett window that sum is arithmetic:

k=1L(1kL+1)=L1L+1L(L+1)2=L2\sum_{k=1}^{L}\left(1-\frac{k}{L+1}\right) = L - \frac{1}{L+1}\cdot\frac{L(L+1)}{2} = \frac{L}{2}

Exactly half the width, at every width, with no approximation anywhere in it. For the Parzen window the sum is 0.375(L + 1) − ½, which is 11.125 at a width of thirty. For a truncated band every weight is one and the sum is the width — so the conventional charge is exactly right for the one band that is not a covariance matrix, and twice what it should be for the one that is.

That is a derivation, and this site does not ship derivations. It ships measurements beside them.

The determinant, which is the difference that is not the difference

Before measuring, there is a second thing that makes a covariance parameter unlike a coefficient, and it is the one that looks decisive.

A coefficient enters the concentrated likelihood through the residual sum and nowhere else. A covariance enters twice: through the quadratic form ê′Ω⁻¹ê and through −½log|Ω|. The field that had to put the determinant back found the second term worth 23.875 across a table of candidates and exactly zero across a table that shares it — a term that cancels out of every difference until it does not. It would be reasonable to expect it to matter here too.

It does not, and the reason is that the determinant is a deterministic function of γ. It contains no data. The optimism of an objective is the expected gap between the same objective read on the sample it was fitted to and on a sample it was not; a term that is identical on both samples contributes nothing to that gap. What the determinant does is bend the objective, which moves where the maximiser sits and not what the maximiser costs. The measurement below separates the two by holding the band’s dimension fixed and varying the window, which changes the shrinkage and leaves the determinant’s role untouched.

Measuring it

The charge a criterion should levy is the optimism, and an optimism is a difference between two samples, so it needs two samples.

Two worlds are drawn per trial, from two separate streams. Ω̂ is estimated on the first at each width in turn. The concentrated likelihood is then evaluated at that Ω̂ on both worlds, with the coefficients and the scale re-profiled on each — so what is left across widths is the optimism attributable to estimating γ, and the narrowest band in the sweep is the constant the rest is differenced against. That constant is not zero and is not expected to be: it is the coefficients’ own optimism, which every width carries equally.

Two separate streams rather than the two halves of one long series, which is the obvious economy and is wrong here. One of the four laws has its break at an absolute row and another is still at 0.57 correlation at the twentieth lag; a gap wide enough to make the second half independent under one of them is not wide enough under the other, and a control that holds under three laws of four is not a control.

What a band of lags actually costs, against what it is chargedThree quantities against the width of a tapered band, on 2000 pairs of independent samples of 120 rows under AR(1) at 0.8. The upper line is what a criterion charges — one log-likelihood unit a lag, which is Akaike's penalty applied to the band as though every lag were a free coefficient. The middle line is what the window's own weights predict, Σ(1 − k/(L+1)), which is exactly half the width. The lower line with its error bars is the measurement: the objective at the fitted covariance, minus the same objective on a second sample drawn independently from the same law, differenced from the narrowest band so that the coefficients' own optimism drops out. At 30 lags it is 10.855 ± 0.281 against a charge of 29 — 2.67 times too large — and against 14.50 predicted by the weights. The measurement is nearer the weights than the convention and is under both, at every width in the sweep.01020301481216202430lags in the bandlog-likelihood units of optimismwhat a criterion chargeswhat the window's weights predictwhat it measures2000 independent pairs, 120 rows each2.67× at the widest band
Fig. 2 What a band of lags costs, measured, against what it is charged. The slider changes the law the errors follow; the two straight lines do not move, because neither convention knows what it is charging for.

At thirty lags the measured optimism is 10.855 ± 0.281 log-likelihood units, on two thousand independent pairs of samples of a hundred and twenty rows under a first-order autoregression. The charge for those twenty-nine additional lags is 29. The summed weights predict 14.50. The measurement is under both, and it is under the convention by a factor of 2.672.

The per-lag rate is 0.374, and the shape of the sweep is worth reading rather than summarising. At two lags the measured rise is 0.476 ± 0.090 against 0.500 summed — the weights are the arithmetic there, to within a fifth of a standard error. At eight lags it is 2.772 against 3.500, at twenty 7.346 against 9.500, and at thirty 10.855 against 14.500. The share falls monotonically from 0.95 to 0.75 across the sweep. A narrow band costs what its weights say; a wide one costs three quarters of that, and the gap opens smoothly.

Which is a fact about the covariance, not about the criterion

The next figure is the one that decides what any of this can be turned into.

A charge for a covariance depends on the covariance. The measured optimism of a Bartlett band, per lag of width, under each of the four laws, on 1200 pairs of independent samples of 120 rows each. A charge for a regression coefficient does not depend on what is being regressed; this one depends on what is being estimated, because the information a sample carries about the autocovariance at lag k is itself a functional of the autocovariance. It is 0.369 of a log-likelihood unit under a first-order autoregression, 0.373 under a five-period moving average, 0.417 under long memory and 0.472 under a break in the persistence — a spread of 1.28 times, and every one of them well under the 1 unit a criterion charges. The two laws a band cannot represent cost the most, which is the ordering a misspecified family predicts and is not the ordering any convention has.
Fig. 3 The measured charge per lag under each of the four laws. A charge for a regression coefficient does not depend on what is being regressed; this one depends on what is being estimated.

Per lag of Bartlett band, the measured charge is 0.369 under a first-order autoregression, 0.373 under a five-period moving average, 0.417 under long memory and 0.472 under a break in the persistence. The spread is a factor of 1.279 and the ordering is not arbitrary: the two laws a band can very nearly represent cost the least, and the two it cannot cost the most.

The mechanism is not mysterious once stated. The information a sample of n rows carries about γ(k) is itself a functional of γ — the variance of a sample autocovariance at lag k involves the whole autocovariance sequence — so a persistent process gives up its autocovariances less precisely than a short-memory one does, and an imprecise estimate has more room to overfit. That is a fact about the object being estimated, and no charge written as a function of L and n alone can carry it.

So there is no constant to look up, and that is the finding rather than a limitation of the measurement. What every reading here agrees on is the direction and the order of magnitude: the convention is between two and three times too large, under every law measured and every window measured, and it is never too small.

The share has a limit, and the limit is a third under the weights

The share of the summed weights the measurement recovers falls from 0.95 to 0.75 across the sweep, and the four readings say where it is going.

Fitting the simplest thing with a floor in it — a + b/L — to the four shares gives 0.735 + 0.435/L, which reproduces 0.952 and 0.792 and 0.749 to within three thousandths and the twenty-lag reading to 0.017. So the share does not keep falling: it approaches 0.735 and the sweep has very nearly reached it by thirty lags.

Read as a rule that is what a practitioner needs. The summed weights are exactly right for a very narrow band and about 27% too high for a wide one, with the whole of the transition happening inside the first eight lags. A charge of 0.735 × Σ w(k) is therefore right at the wide end by construction and over-charges the narrow end by a few per cent, which is the opposite error from the one the convention makes and a hundredth of its size.

That constant has an independent check. Fitting a straight line in the summed weights to the same profile over the plateau — a different field, a different weighting, a different set of widths — gives a rate of 0.7557 per unit of summed weight. Against the 0.735 the asymptote implies, that is agreement to 3%, from two routes that share the measurement and nothing else about how it is read.

Two ordinary laws and two awkward ones

The per-lag charges of 0.369, 0.373, 0.417 and 0.472 are read as an ordering by how well a band represents each law, and the four numbers say the split is binary rather than graded.

The first-order autoregression and the moving average sit at 0.369 and 0.373 — within one per cent of each other, and the moving average is if anything the dearer of the two despite being the one law a band contains exactly. So representability does not order the cheap pair at all. Long memory at 0.417 and the break at 0.472 are 13% and 28% above them.

What separates the two groups is not how much of each law a band omits but whether the law has an ordinary summable dependence at all: long memory’s autocorrelations do not sum, and the break’s covariance is a function of position rather than of gap. Two laws with a finite dependence length cost 0.37 a lag and two without one cost 0.42 to 0.47, and inside each pair the ordering is noise.

The practical consequence is a bound rather than a formula. A rule calibrated on the ordinary case and applied to the awkward one under-charges by up to 21% — 0.374 levied where 0.472 is owed — and an under-charged penalty picks a wider band than it should. That is a real error and it is a quarter of the size of the one the convention makes in the other direction, so a practitioner choosing between charge one a lag and charge 0.37 a lag whatever the law should take the second and know that it is the second-best available answer rather than the right one.

What the number is not

Three readings are available from the same measurement and two of them are wrong, so it is worth closing them before going on.

It is not a claim that a criterion is over-penalising by a factor of three in general. The quantity measured here is the optimism of the covariance’s dimension only. The coefficients in the same fit are ordinary free parameters and cost an ordinary unit apiece; nothing in this sweep touches them, which is exactly what differencing from the narrowest band achieves. A criterion selecting among regressors with a fixed band is correct as written, and one selecting the band is not.

It is not the optimism of the maximiser. Everything here is measured at the plug-in — the tapered sample autocovariance substituted into the objective — because that is the estimator every rule in this collection actually uses. The essay that maximised the same family reports the plug-in sitting several units below the maximum, and a maximiser has more freedom than a plug-in and would cost more. So this number is a charge for the rule as run, not for the family as such, and a practitioner who switched to maximising would need it measured again.

And it is not an approximation whose error is unknown. The standard errors are on the figure and they are small relative to the gap: at thirty lags the measurement is 10.855 with a standard error of 0.281, which is nearly seventy standard errors below the charge of 29 and thirteen below the summed weights. Whatever is wrong with the derivation, it is not that the measurement cannot see the difference.

What a charge of the right size does

The obvious next question is what levying it buys, and the honest answer arrives in the essay that levies it. It is worth stating the shape here because it changes how this measurement should be read.

Cut the charge and the width follows it. The band width each charge picks, averaged over 400 draws of 120 rows under AR(1) at 0.8, with the standard deviation across draws beside it. Schwarz's charge — half a log n a lag, which is 2.39 here — picks 3.67. Akaike's picks 6.02. Charging the numbers the window actually leaves free, which is half the width, picks 10.12; charging what the optimism measures, 0.767 of that, picks 14.15. A charge and the width it buys are very nearly reciprocal, which is what a likelihood rising at a fixed rate a lag implies and is why the four answers span a factor of 3.86. The width that was actually best on the draw averages 13.90 and moves by 10.30 from draw to draw — three times as much as any rule's answer does.
Fig. 4 The width each charge picks. A charge and the width it buys are very nearly reciprocal, which is what a likelihood rising at a fixed rate a lag implies.

Cutting the charge to a third roughly triples the width. Schwarz’s picks 3.67 lags, Akaike’s 6.02, the summed weights 10.12, and the measured charge 14.15 — a factor of 3.86 across the table. The error those widths deliver runs from 1.11290 to 1.10728, which is half a per cent, and the reason is the one the width field already measured: the likelihood rises at about a unit a lag whatever is in the data, so a penalised objective is nearly flat in L and its argmax is nearly arbitrary.

Which makes the right reading of this essay’s number an explanation rather than a correction. The width table was noisy because the charge was wrong by a factor of three, and it stays noisy after the charge is fixed, because the thing the charge is levied on has almost no curvature to find an argmax in.

Where the convention came from, and why it survived

It is worth being precise about what the two conventions are, because neither is careless.

Akaike’s penalty is the optimism of a maximised log-likelihood in a correctly specified model with q free parameters, and the essay that derives it as a trace shows the trace collapsing to q exactly when the fit has q free directions and the rows are independent. Both conditions fail here. The band’s directions are not free — they are shrunk by a fixed sequence — and the rows are the very thing whose dependence is being estimated.

Schwarz’s is not an optimism at all; it is an approximation to a marginal likelihood, and reading it beside Akaike’s as though the two were the same kind of quantity is a confusion this collection has already priced. It is here because it is what a practitioner reaches for second.

What kept both in place is that nothing in the pipeline could detect the error. A criterion returns a width, the width returns a fit, the fit returns an error, and the error is half a per cent away from the error any other width would have returned. A penalty that is wrong by a factor of three produces no symptom at all when the objective it penalises is flat, and the flatness is exactly the condition under which somebody would want a principled penalty.

What is left open

Three things, and naming them is more useful than the number.

The level has no derivation. The weights predict the ordering and the direction and three quarters of the size; where the remaining quarter goes is not established here. The candidates are the normalisation by γ̂(0), the divisor in the sample autocovariance, and the fact that the sequence is computed from residuals rather than errors — each of which is a shrinkage in its own right, and the three of them compound. Separating them needs a design that varies one at a time and this field does not build one.

The measurement is per law and a practitioner has one sample. Everything here is an expectation over draws from a known law. The charge a practitioner needs is the charge under the law their sample came from, which is the thing being estimated, so a genuinely usable derived charge would have to be a plug-in — computed from Ω̂ itself, which is the object whose dimension is being charged for. That circularity is not obviously fatal and is not addressed.

And a wider band is charged less per lag than a narrow one. The share falls from 0.95 to 0.75 across the sweep, so the right charge is not linear in L at all, and every rule considered here — including the derived one — is a straight line through the origin. Fitting the curve would move the widths again, by less than the factor of three already on the table.

None of the three changes the sentence this essay exists for, which survives all of them: a band of lags is charged as though its lags were free, and they are not free, and nobody had counted.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BandwidthCovariance matrixDegrees of freedomInformation criterionLong-run varianceModel selectionNested modelsNuisance parameterOptimismOut of samplePersistencePlug in estimateSample autocovarianceShrinkageTapering