A charge that is not a straight line

The width a band is measured in

A tapered covariance band spends 84% of its own weights at two lags and 74% at thirty. Every charge in the collection is a straight line through the origin in those weights, so it is too dear at one end and too cheap at the other.

Worth reading first: A design is a number · The observations that repeat each other.

The field that derived a charge for a covariance band ends by naming the thing it did not do. Its charge is a share of the band’s summed weights: measure the optimism a band of thirty lags actually costs, divide by Σ w(k), and levy that fraction at every width. The number it lands on is 0.7486 at two thousand draws, against a convention that charges one unit a lag.

The trouble is in the word share. A share is a constant, and this one is not.

At two lags the band spends 0.8447 of its own summed weights. At thirty it spends 0.7399. In between it falls without interruption — 0.8338, 0.7902, 0.7768, 0.7745, 0.7730, 0.7683, 0.7621, 0.7546 — so a charge levied as a fixed fraction of Σ w(k) under-charges a narrow band and over-charges a wide one, by about a tenth at the narrow end.

That is not a small defect in an otherwise good rule. It is the rule. Every charge the earlier field considered, including the two it derived, is a straight line through the origin in the summed weights, and the measurement says the optimism is not one.

The curvature is in the denominatorThe optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.0.8000.9001102030the band's width, in lagsthe measured charge, per unit of the band's widththe plateau startsin the pairs the band usesin the weights it spends2000 draws of 120 rows, Bartlettflat means a straight line is right
Fig. 1 The optimism a band of each width costs, divided by that width, on two ways of measuring the width. The upper line is the one this field arrives at; the lower one is what the earlier field levies. The slider changes how many draws the sweep takes.

What is being divided by what

The numerator is not in dispute and is worth restating, because the whole of this field is about the denominator.

For each draw, a world is generated and a second, independent world is generated beside it. A tapered plug-in covariance is fitted to the first at width L; the concentrated likelihood is evaluated at that Ω̂ on the first world and again on the second, with the coefficients and the error variance re-profiled on the replicate. The gap between the two is the optimism: how much better the objective looks on the data that chose Ω̂ than on data that did not. It is measured as a rise from the narrowest width in the sweep, so the coefficients’ own optimism — a constant every width pays — is differenced away.

At eight thousand draws the rises are 0.4224, 0.8338, 1.1852, 1.9420, 2.7106, 4.2516, 5.7623, 7.2400, 8.6776 and 10.7279, with standard errors from 0.0440 at the narrowest to 0.1388 at the widest.

Those numbers are not the subject of this field. The field that measured them is, and it also establishes that the ratio between them and the summed weights is a property of the window’s shape rather than of any one window’s scale: three Bartlett windows whose weight sums stand as 6 : 4 : 3 give charges within 0.033 of each other as fractions of their own sums.

What this field is about is the denominator those rises are divided by, and the finding is that the division is being done in the wrong variable.

The two sequences are one sequence

The ten shares and the ten optimism readings are printed in different sections and they are the same measurement, which is worth checking because the check identifies every width in the sweep.

Divide each optimism by its share and the summed weights fall out: 0.5, 1.0, 1.5, 2.5, 3.5, 5.5, 7.5, 9.5, 11.5 and 14.5, each to four decimal places. Since a Bartlett window’s summed weight is exactly half its width, the sweep runs at widths of 1, 2, 3, 5, 7, 11, 15, 19, 23 and 29 free entries — the last being what “a band of thirty lags” holds once its own variance is counted separately.

Nothing was fitted to produce that. It is a consistency check that would have failed on any mis-indexing between the two tables, and it passes at every one of the ten widths.

The optimism is a quadratic in the summed weights

With the widths recovered, the shares can be read as a curve rather than as a list, and the curve has a simple shape over the range that matters.

The first two points fall steeply — 0.8447 at half a unit of summed weight, 0.8338 at one — and from 2.5 upwards the share falls almost exactly linearly:

share0.78450.00308Σw.\text{share} \approx 0.7845 - 0.00308\,\Sigma w .

Multiplying through, the optimism itself is

0.7845Σw0.00308(Σw)2,0.7845\,\Sigma w - 0.00308\,(\Sigma w)^2 ,

which at the widest point gives 11.375 − 0.648 = 10.727 against 10.7279 measured, and at Σw = 5.5 gives 4.222 against 4.2516.

So the earlier field’s straight line through the origin is not merely the wrong slope; it is missing a curvature term whose size is 0.0031 per squared unit of summed weight. Small, and it accumulates: by the widest band it has taken 0.65 log-likelihood units off a linear charge of 11.4, which is six per cent of the total.

Where the constant share is exactly right

A fixed fraction has to be correct somewhere, and locating it says what the earlier field measured.

Interpolating the shares, 0.7486 sits at a summed weight of about 12.7 — a band of roughly twenty-five lags, near the wide end of this sweep and close to the thirty the earlier field measured at. So its number is not wrong; it is the local value at the width it was calibrated on, presented as a constant.

The cost of treating it as one runs in the direction the shares give. At the narrowest band the true share is 0.8447 and the fixed rule charges 0.7486 — eleven per cent too little. At the widest the true share is 0.7399 and the rule charges 0.7486 — one per cent too much. The rule is nearly right where it was measured and it is the narrow bands it misprices, which are exactly the bands a criterion comparing widths has to get right if it is ever to choose one.

A share that is not a share

The plainest way to see it is to plot the ratio rather than the two columns, which is what the figure above does.

A charge that is a straight line through the origin is exactly a charge whose ratio is flat. So a flat profile says the line is right and a sloped profile says it is not, and the profile slopes: it starts at 0.8447, falls fast for the first two widths, and then keeps falling for the rest of the sweep. Between four lags and thirty — which is where the first steep drop has already happened — it still falls from 0.7902 to 0.7399, at every one of the seven steps, and to 0.9363 of its starting value.

The right way to price that is a fit rather than a difference between two readings, because the eight come off the same draws and are correlated. Weighting each width by the reciprocal of its own squared standard error, the best constant share misses the plateau by 0.6323 per width in units of those errors — which is what a profile with a slope in it looks like, and is the number every comparison below is against.

What a band of lags actually costs, against what it is charged. Three quantities against the width of a tapered band, on 2000 pairs of independent samples of 120 rows under AR(1) at 0.8. The upper line is what a criterion charges — one log-likelihood unit a lag, which is Akaike's penalty applied to the band as though every lag were a free coefficient. The middle line is what the window's own weights predict, Σ(1 − k/(L+1)), which is exactly half the width. The lower line with its error bars is the measurement: the objective at the fitted covariance, minus the same objective on a second sample drawn independently from the same law, differenced from the narrowest band so that the coefficients' own optimism drops out. At 30 lags it is 10.855 ± 0.281 against a charge of 29 — 2.67 times too large — and against 14.50 predicted by the weights. The measurement is nearer the weights than the convention and is under both, at every width in the sweep.
Fig. 2 The optimism itself, width by width, in the field that measured it — the numerator this one divides.

What a standard error looks like here

Eight thousand draws is a large sweep for this collection, and the number is chosen rather than inherited.

The rises are measured as paired differences from the narrowest width — the same draw’s optimism at L and at the base, subtracted — so the draw-to-draw variation in how hard a particular world is to fit cancels before anything is averaged. What is left is the standard error of a difference, and it grows with the width because a wider band has more room to overfit and therefore more spread: 0.0440 at two lags, 0.1124 at eight, 0.1388 at thirty.

Those are absolute errors on rises that grow from 0.4224 to 10.7279, so as relative errors they shrink steeply — from about a tenth of the rise at two lags to about a seventy-seventh at thirty. That is the right shape for what follows, because the fits in the third essay weight each width by the reciprocal of its own squared standard error, and a width the sweep resolves badly should not be allowed to choose the shape.

And two thousand would not have done. The field this one follows reads its own sweeps at two thousand draws, which is right for the quantity it reports: the optimism at the widest band, whose standard error there is three per cent of itself. It is not right for a ratio read at eleven widths and then compared across them. At a thousand draws the comparison below comes out the other way round; at two thousand and four thousand it comes out this way at different strengths; and the level of the share settles only from about eight thousand. A ratio needs more draws than either of the things it is a ratio of, and a statement about the shape of a profile needs more again.

The two candidates a reader will reach for

Faced with a curve, there are two obvious things to do, and this field’s next three essays are largely about why neither is the answer.

Fit the curve. Take the eleven measured rises, fit a two-constant shape to them — a power of the summed weights, or a rate that decays with the width — and levy that. It is one extra fitted constant and it will describe the profile better than a line, because it has a constant more.

Or ignore it. The charge already beats a convention by a factor of two, the widths it picks are already near the best available, and a curvature of a fifth across a fifteenfold range of width might be a refinement nobody can act on.

Both readings assume the curve is a property of the charge. The alternative — that it is a property of the scale the charge is read on — is the one nothing in the earlier field considers, and it is the next essay’s subject.

The width that is not the width

Here is the shape of the answer, stated now so that the rest of this essay can be read against it.

A tapered plug-in autocovariance at lag k is γ̂(k) w(k). The window’s weight w(k) says how much of γ̂(k) the band keeps, and Σ w(k) adds those up and calls the total the band’s width in free parameters — which is the right count for a linear shrinkage towards zero, and is derived carefully in the earlier field.

What it says nothing about is how much of γ̂(k) there was to keep. The sample autocovariance at lag k is an average over n − k products, not over n. A hundred and twenty rows give a hundred and nineteen products at the first lag and ninety at the thirtieth, so the thirtieth lag’s estimate is built from three quarters of the evidence the first lag’s is.

Counting a long lag at the same weight as a short one therefore over-states how much a wide band is spending. Counting each lag at the share of the pairs it actually has —

P(L)=k=1Lw(k)(1k/n)P(L) = \sum_{k=1}^{L} w(k)\,(1 - k/n)

— is one line of arithmetic, has no fitted parameter in it at all, and is what the upper line in this essay’s opening figure is drawn against.

Read on that width, the same eleven measurements give 0.8566, 0.8479, 0.8058, 0.7967, 0.7989, 0.8066, 0.8111, 0.8141, 0.8158 and 0.8145 — no longer falling from four lags up, and a constant share misses them by 0.0632 per width against 0.6323 on the summed weights.

Why the curvature and the correction are the same size

It is worth checking that this is not a coincidence of arithmetic, because a correction that happens to be the right size at one sample length and the wrong size at another would be a fit with the fitting hidden.

The correction is largest where the lags are longest relative to the sample. At a hundred and twenty rows a Bartlett band of thirty lags has a pairs width of 13.6667 against a summed-weight width of 15, a ratio of 0.9111; the same band on sixty rows gives 0.8222 and on four hundred and eighty rows 0.9778. So the correction shrinks towards nothing as the sample lengthens, which is what it must do if it is about the sample’s own arithmetic and not about the window.

It also does nothing at all at the narrow end. At two lags the correction moves the width from 0.5 to 0.4931, a change of a hundred and forty parts in ten thousand, and the share moves from 0.8447 to 0.8566 — up rather than down, and by almost nothing. Which is why the two narrowest widths are still off the plateau on both scales, and why this field draws its plateau from four lags rather than from one.

That limit is stated rather than smoothed over. A two-lag band costs more per unit of width than a thirty-lag band does, on either scale, and nothing here explains why.

A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it.
Fig. 3 How far each candidate shape sits from the measurement across the plateau, in units of each width’s own standard error — the third essay’s table.
The repair is one window's. How far the measured charge per unit of width sits from a constant across the plateau, in units of each width's own standard error, on each of two scales, for each of four windows, over 2000 draws. A reading near zero means a straight line through the origin is the right shape in that scale. Counting the band's width in pairs rather than in summed weights improves the Bartlett window by a factor of 28.16 — from 0.2359 to 0.0084 — and does far less for the other three: 1.80, 1.61 and 1.31, on profiles that are ten to seventy times further from flat to begin with. So the correction, which has no fitted parameter and is stated for windows in general, repairs the one window the earlier field measured and leaves the rest wanting a curve.
Fig. 4 And how much the correction flattens each of the four windows, which is where the repair stops.

What a fifth of a charge is worth

Before any of this is worth doing, it is fair to ask what the whole quantity decides.

The field that measured the charge is unusually direct about this: getting the charge right moves every width and almost no error. Its four charges pick 3.67, 6.02, 10.12 and 14.15 lags — a factor of nearly four — and deliver root mean squared errors of 1.11290, 1.10896, 1.10728 and 1.10739, which is half a per cent.

So the honest expectation going in is that a fifth of a charge moves the width by something and the error by nothing. The fourth essay of this field measures exactly that and finds exactly that, and it is not a reason to stop: a criterion that picks a width three times too narrow is wrong in a way somebody will eventually notice, and a criterion that is right for the wrong reason is wrong in a way nobody will.

The one thing the profile cannot be

There is a reading of a falling share that would make it uninteresting, and it is worth closing it off.

A share that fell because the optimism saturates — because a band past some width stops buying any extra fit at all — would be a statement about the likelihood rather than about the charge, and the right response would be to cap the width rather than to re-scale the charge. That is not what is happening. The rises are very nearly proportional to the width throughout: dividing consecutive differences by the width they span gives 0.823, 0.703, 0.757, 0.769, 0.770, 0.755, 0.739, 0.719 and 0.683 per unit of summed weight, which drifts by a sixth and does not fall off a cliff.

The optimism keeps growing at very nearly a constant rate. What is not constant is the rate at which the summed weights grow relative to what the sample has, and that is a fact about the denominator.

Why the denominator was never looked at

There is a reason the earlier field divided by Σ w(k) and did not ask whether it should, and it is worth naming because it is the kind of reason that recurs.

Σ w(k) is derived and it is derived correctly. A tapered plug-in autocovariance is a linear shrinkage of γ̂(k) towards zero with a fixed factor; the effective dimension of a linear shrinkage is the sum of the derivatives of the fitted values with respect to the things they estimate; those derivatives are the weights; so the band is worth Σ w(k) parameters rather than L of them. That argument is right, it is checked against a closed form, and it is the argument that makes the whole earlier field possible.

What it computes is the dimension of the shrinkage given the γ̂ it shrinks. It takes the sample autocovariances as the objects being shrunk and asks how much of them survives. It does not ask what those objects cost to have, and there is no reason from inside that derivation to ask — the count is exact for what it counts.

So the defect is not an error in the derivation. It is that a derivation which is right about one quantity was used as the denominator for a different one, and the only way to notice is to read the ratio across widths rather than at the widest. The next essay is what happens when the ratio is read that way.

What this essay claims, and what it does not

It claims one thing: the measured charge for a tapered covariance band is not a fixed share of the band’s summed weights, and the departure is large enough to see and systematic enough to have a mechanism.

It does not yet claim that the pairs count is the mechanism — that is the next essay, which has to show that the correction is arithmetic rather than fitted, that it survives changing the law, and that it fails where it should.

It does not claim that a curve is unnecessary — that is the third, which puts a fitted power law and a fitted decaying rate beside the two straight lines and reads off which fits the plateau at how many constants.

And it does not claim that any of this repairs an error somebody is making. That is the fourth, and the answer there is mostly no.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A charge that reads the draw — both name bandwidth selection, dependence, estimation error, information criterion, model selection, monte carlo, optimism, plug in estimate, sample autocovariance, tapering
  • A penalty is a trace — both name closed form, degrees of freedom, dependence, information criterion, model selection, optimism, overfitting
  • A rate times a size — both name bandwidth selection, closed form, estimation error, information criterion, long-run variance, model selection, monte carlo
  • A window for every candidate — both name bandwidth selection, information criterion, long-run variance, model selection, optimism, overfitting, tapering
  • The plug-in and the maximum — both name degrees of freedom, dependence, monte carlo, overfitting, sample autocovariance, shrinkage, tapering
  • The window that has to be chosen, and the term that was dropped — both name degrees of freedom, dependence, estimation error, information criterion, long-run variance, model selection, tapering

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionClosed formDegrees of freedomDependenceEstimation errorInformation criterionLong-run varianceModel selectionMonte CarloOptimismOverfittingPlug in estimateSample autocovarianceShrinkageTapering