A charge that is not a straight line

A lag the sample has less of

A sample autocovariance at lag k is an average over n − k products, not n. Count a band's width in the pairs it actually has and the curvature in its charge goes away, on a correction with nothing fitted in it.

Worth reading first: A design is a number · The observations that repeat each other.

The first essay of this field leaves a curve on the table: the charge a tapered covariance band actually costs, read per unit of the band’s summed weights, falls from 0.8447 at two lags to 0.7399 at thirty. This essay proposes that the curve is in the wrong variable, and then tries to break the proposal.

The proposal is one line long.

The line

The proposal is not that the earlier field’s denominator is wrong about what it counts. It is that a covariance band has two widths, that they are different numbers, and that a charge is levied in the second.

A sample autocovariance at lag k is

γ^(k)=1nt=k+1n(xtxˉ)(xtkxˉ)\hat\gamma(k) = \tfrac{1}{n}\sum_{t=k+1}^{n} (x_t - \bar x)(x_{t-k} - \bar x)

and the sum has n − k terms, not n. On a hundred and twenty rows the first lag is built from a hundred and nineteen products and the thirtieth from ninety. A tapered band keeps w(k) of each of those, so a band of thirty lags is spending its widest weights on the estimates the sample knows least about.

Σ w(k) counts every lag at the weight the window gives it. It says nothing about how much of γ̂(k) there was to keep. So count each lag at the share of the pairs it has:

P(L)=k=1Lw(k)(1k/n)P(L) = \sum_{k=1}^{L} w(k)\,(1 - k/n)

That is the whole of it. There is no parameter in it, nothing is estimated, and it is computable from the window, the width and the sample size before any data exist.

The curvature is in the denominatorThe optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.0.8000.9001102030the band's width, in lagsthe measured charge, per unit of the band's widththe plateau startsin the pairs the band usesin the weights it spends2000 draws of 120 rows, Bartlettflat means a straight line is right
Fig. 1 The measured charge per unit of width, on the two denominators. The upper line is the pairs count and the lower is the summed weights. The slider changes how many draws the sweep takes.

What it does

Read on the pairs, the eleven measurements from the earlier essay become 0.8566, 0.8479, 0.8058, 0.7967, 0.7989, 0.8066, 0.8111, 0.8141, 0.8158 and 0.8145.

From four lags up they stop falling. Weighting each width by the reciprocal of its own squared standard error, a constant share misses the plateau by 0.0632 per width against 0.6323 on the summed weights — a factor of 10.008 — and the reading at thirty lags is 1.0108 of the reading at four against 0.9363.

The mean over the plateau is 0.8079, which is the constant a straight line through the origin would be levied at. On the summed weights the same mean is 0.7674 and is not a constant anything should be levied at.

A correction with nothing fitted in it takes an order of magnitude off the misfit. That is the finding, and the rest of this essay is the four ways it could be an accident.

Which of the two widths a reader has been handed

Both widths are a count of parameters and they mean different things, so it is worth setting them side by side once.

Σ w(k) is what the band spends. It is the effective dimension of the shrinkage: how many free numbers survive after the window has multiplied the sample autocovariances down. For the Bartlett window it is exactly L/2 at every width, which is a closed form and a satisfying one, and it is the count the earlier field derives and checks.

Σ w(k)(1 − k/n) is what the band gets. It discounts each lag by the share of the sample’s products that went into it. On a hundred and twenty rows the discount is under one per cent at the first lag and a quarter at the thirtieth, so the two widths agree at the narrow end and part at the wide one — 0.5 against 0.4931 at two lags, 15 against 13.6667 at thirty.

The gap between them is therefore a width-dependent discount, which is exactly the shape a curvature in their ratio would need. That is not proof; a discount of the right shape and the wrong size would leave a residual slope, and the measurement is what says the size is right too.

It is arithmetic, and it is checked as arithmetic

The first thing that could be wrong is the sum itself, so it is computed twice.

P(L) = Σ w(k) − (1/n) Σ k w(k), and for the Bartlett window both sums are closed:

k=1Lw(k)=L/2k=1Lkw(k)=L(L+1)2L(2L+1)6\sum_{k=1}^{L} w(k) = L/2 \qquad \sum_{k=1}^{L} k\,w(k) = \tfrac{L(L+1)}{2} - \tfrac{L(2L+1)}{6}

so P(L) = L/2 − (1/n)(L(L + 1)/2 − L(2L + 1)/6). The loop and the closed form agree to a part in 10¹² at every width from one to sixty-four and at sample sizes of sixty, a hundred and twenty and four hundred and eighty.

That is a small check and it is not a formality. The one way this correction could be silently wrong is an off-by-one in the lag index, which would move every number here by about 1/n and would leave the whole picture looking exactly as it does.

It reads no data

The second thing that could be wrong is subtler: a “correction with no fitted parameter” that quietly reads the sample is a fit with the fitting hidden, and the difference is invisible from the output.

So the claim is asserted directly. The pairs width is computed twice on the same arguments and required to return identical bits, and its ratio to the summed weights is required to move with the sample size and with nothing else. At thirty lags that ratio is 0.8222 on sixty rows, 0.9111 on a hundred and twenty and 0.9778 on four hundred and eighty.

The direction is the one the mechanism predicts and would be hard to arrange by accident: a longer sample knows more about a long lag, so the correction has less to correct, and the two denominators converge. A correction that grew with the sample size, or that did not move with it at all, would not be about the pairs.

The discount is (L + 2) over three n

The two widths’ ratio is reported at three sample sizes and it has a closed form, which is worth having because it turns the correction into one expression rather than a table.

For the Bartlett window, Σw(k) = L/2 and Σk·w(k) = L(L+1)/2 − L(2L+1)/6, so

P(L)/Σw(L) = 1 − (L + 2)/(3n)

At thirty lags and a hundred and twenty rows that is 1 − 32/360 = 0.9111; at sixty rows, 0.8222; at four hundred and eighty, 0.9778 — the three ratios the essay reports, from an expression with no sum in it.

So the whole correction is a multiplication by 1 − (L + 2)/(3n), which is 1.1% at two lags, 1.7% at four and 8.9% at thirty on this sample. That is the width-dependent discount the essay describes, and now it has a size and a shape: linear in the width, inverse in the sample.

The conversion checks against both profiles at once. Dividing the summed-weight shares by the discount should give the pairs shares, and at two lags 0.8447/0.9889 = 0.8542 against a reported 0.8566, and at thirty 0.7399/0.9111 = 0.8121 against 0.8145. Both ends to within three thousandths, from one expression applied to the earlier field’s own numbers.

The residual curvature is one law’s

Four laws, four reductions, and the four residuals are not alike.

After the correction the misfits are 0.0679, 0.1265, 3.0165 and 0.4716 under the autoregression, the moving average, long memory and the break. Long memory’s residual is forty-four times the autoregression’s, and its reduction factor of 2.48 is the smallest of the four.

So the law the correction helps least is also the law it leaves worst, and it is the one whose autocorrelations do not sum. That is not a coincidence: the pairs discount is a statement about how many products went into each γ̂(k), and it says nothing about whether the γ(k) being estimated themselves die away. A band on a summable dependence is missing a finite tail; a band on long memory is missing an infinite one, and no denominator counts that.

The practical residue is a boundary rather than a caveat. The pairs width is the right count wherever the dependence is summable, which is three of the four laws here and every law a band family contains; on long memory it is still an improvement by a factor of two and a half and still leaves a profile no straight line through the origin describes.

It survives the law

The third thing that could be wrong is that a first-order autoregression at 0.7 is a convenient world.

The profile is re-measured under all four laws this line of fields works in — the autoregression, a moving average, a long-memory process and one with a break in it — at four thousand draws apiece, half the standard because it is four profiles rather than one. How far a constant share sits from the plateau falls under every one:

0.2663 to 0.0679 under the autoregression, 0.7189 to 0.1265 under the moving average, 7.4759 to 3.0165 under long memory, and 2.1376 to 0.4716 under the break — factors of 3.92, 5.69, 2.48 and 4.53.

Four for four is what a piece of arithmetic should give. It is not four flat profiles, and that is the honest half: under long memory the corrected profile still misses a constant by 3.0165 per width, and under a break its level is 1.0899 — above one, meaning the band costs more optimism than its pairs width says it should.

So the pairs count improves every law by a factor of two to six and finishes the job in one of them, which is the case the earlier field measured in.

The correction flattens under every law. How far the measured charge per unit of width sits from a constant across the plateau, in units of each width's own standard error, on two scales, under each of the four laws the collection works in, over 4000 draws apiece. Counting the band's width in pairs brings the profile nearer a constant under all four — by factors of 3.92, 5.68, 2.48 and 4.53 — which is what makes the correction arithmetic rather than a fit to one world. It does not leave all four flat: under a first-order autoregression the corrected profile sits at 0.0679 and under long memory at 3.0165, so the correction improves every law and finishes the job in one.
Fig. 2 The spread of the charge across the plateau, on both denominators, under each of the four laws.
Four dependences a single parameter cannot tell apart. Every law here is standardised to a lag-one autocorrelation of 0.8, so a rule told the errors are a first-order autoregression finds the same number in all four and has no way of seeing what separates them. The geometric decay is the world in which estimating a covariance rather than naming it was priced, and found to cost. The five-period moving average has 0.200 at the fourth lag and exactly nothing past it, where the geometric law says 0.328 at the fifth. Long memory at d = 4/9 is still at 0.576 by the twentieth lag, where the geometric law has reached 0.012. The break has no autocorrelation function at all: what is drawn for it is the average over the pairs at each gap, which is what a stationary estimate converges to.
Fig. 3 The four laws themselves, in the field that assembled them.

What the correction is worth against what it replaces

It is fair to ask how much of a change this is in the units anybody levies a charge in, because “a spread of 0.081 becomes a spread of 0.029” is a statement about a ratio and not about a price.

At thirty lags, on a hundred and twenty rows, the summed-weight width of a Bartlett band is 15 and its pairs width is 13.6667. So the correction takes about one and a third parameters off a band that a convention would charge thirty for and the earlier field’s rule would charge about eleven for. In log-likelihood units, at the plateau’s own rate, that is a little over a unit.

A unit is not nothing at the scale a criterion works at — the difference between two candidate widths on this grid is routinely under two — but it is not a repair either. What the correction buys is not a better price at any one width. It is that the same price works at every width, which is a different kind of gain: a rule that is right on average and wrong at both ends will pick badly at both ends, and a reader has no way to know which end they are at.

That is why the comparison this field makes is between shapes rather than between levels, and why the number reported throughout is a spread rather than a mean.

The obvious second correction, and why it is not here

A reader who accepts the pairs argument will immediately propose a second one, and it is worth answering because the answer says what kind of quantity is being corrected.

If a long lag is estimated from fewer products, then it is also estimated worse — its sampling variance is larger, roughly in proportion to 1/(n − k) rather than 1/n — so perhaps the right weight is not 1 − k/n but its square, or some other power. That would be a one-parameter family, and fitting the parameter would certainly fit the profile better than the pairs count does.

It would also stop being arithmetic. The whole claim being made here is that a correction with nothing fitted in it takes an order of magnitude off the misfit, and a correction with an exponent chosen to remove the misfit removes it by construction. A fitted power of the pairs count would be a curve, and the curves are already on the table in the next essay, which prices them against exactly this line.

The relevant question is therefore not whether some transform of the pairs count fits better. It is whether the untransformed one fits well enough that the transform is not worth a constant, and that comparison is made there rather than asserted here.

And it fails where it should not be expected to work

The fourth thing that could be wrong is the most serious, and it is what makes this a limited result rather than a general one.

The correction is stated for a window and measured on four. It repairs one.

Its flattening factor is 10.008 for the Bartlett window, 2.111 for that window squared, 1.884 for it cubed and 1.348 for Parzen. The three it barely touches are not marginal cases: Parzen’s corrected profile still misses a constant by 4.3491 per width against the Bartlett window’s 0.0632, which is a profile with a great deal of shape left in it.

Three of the four windows still want a curve, and this file does not have one that is not fitted.

Why is not established here. The obvious candidate is that a window whose weight is concentrated at short lags — which is what squaring or cubing a Bartlett weight does, and what Parzen does by construction — spends most of its budget where the pairs correction is smallest, so there is less for the correction to do and whatever else is going on is left visible. That is a story, not a measurement, and it is left as one.

The repair is one window's. How far the measured charge per unit of width sits from a constant across the plateau, in units of each width's own standard error, on each of two scales, for each of four windows, over 2000 draws. A reading near zero means a straight line through the origin is the right shape in that scale. Counting the band's width in pairs rather than in summed weights improves the Bartlett window by a factor of 28.16 — from 0.2359 to 0.0084 — and does far less for the other three: 1.80, 1.61 and 1.31, on profiles that are ten to seventy times further from flat to begin with. So the correction, which has no fitted parameter and is stated for windows in general, repairs the one window the earlier field measured and leaves the rest wanting a curve.
Fig. 4 The spread of the charge across the plateau on both denominators, for each of the four windows — where the repair stops.

The narrow end, which the correction cannot be about

One more limit, and this one follows from the arithmetic rather than surprising it.

At two lags the correction moves the width from 0.5 to 0.4931 and the share from 0.8447 to 0.8566 — up rather than down, and by a hundredth. At thirty lags it moves the width from 15 to 13.6667 and the share from 0.7399 to 0.8145, by seven hundredths. The correction is nearly nothing where k/n is nearly nothing, which is exactly what it should be.

But the share at two lags is above the plateau on both scales, by more than a tenth. So there is a second thing going on at the narrow end, and the pairs count is not it.

Stating that plainly is the point. A field that reported a flattening factor of 2.8 and did not say that its two narrowest widths were excluded from the range the factor is computed over would be reporting a number a reader could not check. The plateau starts at four lags, it starts there because the two below it are off it on both scales, and the two below it are shown in every figure.

One thing this does not touch

The correction changes the denominator and leaves the numerator exactly where it was, and it is worth being explicit that this is not a repair to the measurement of the optimism.

The field that measured it draws two independent worlds per trial from separate streams rather than two halves of one long one, re-profiles the coefficients and the error variance on the replicate, and differences everything from the narrowest width so that the coefficients’ own optimism drops out. Every one of those choices is upstream of anything here, and every number in this essay inherits them.

What that means practically is that the pairs count could be right and the profile could still be wrong, if the optimism were being mis-measured in a way that happened to look like a slope. The strongest evidence against that is the sample-size check above: a mis-measurement of the optimism has no reason to move with n in the direction the pairs ratio does, and the two move together.

What a denominator is for

It is worth stepping back one line, because the shape of this result recurs.

The earlier field derived Σ w(k) as the effective dimension of a linear shrinkage, and that derivation is correct. It is checked against a closed form — the Bartlett window leaves exactly half its width free, at every width, exactly — and the check passes.

What it computes is the dimension of the shrinkage given the objects being shrunk. A charge, though, is not a dimension; it is a price, and a price has to be in the currency the buyer pays in. The buyer here is a sample of a hundred and twenty rows, and what it pays for a long lag is not what it pays for a short one.

So the two quantities were never the same quantity. They agree in the limit — the ratio between them goes to one as the sample lengthens — which is exactly why a derivation about one can be used for the other for a long time without anybody noticing.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • What a better charge buys — both name bandwidth selection, degrees of freedom, dependence, estimation error, information criterion, model selection, monte carlo, optimism, plug in estimate, tapering
  • A rate times a size — both name bandwidth selection, closed form, estimation error, information criterion, long-run variance, model selection, monte carlo
  • The window that has to be chosen, and the term that was dropped — both name degrees of freedom, dependence, estimation error, information criterion, long-run variance, model selection, tapering
  • A covariance with no parameter in it — both name dependence, estimation error, information criterion, long-run variance, model selection, tapering
  • A penalty is a trace — both name closed form, degrees of freedom, dependence, information criterion, model selection, optimism
  • A window for every candidate — both name bandwidth selection, information criterion, long-run variance, model selection, optimism, tapering

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionClosed formDegrees of freedomDependenceEstimation errorInformation criterionLong-run varianceModel selectionMonte CarloNormalisationOptimismPlug in estimateSample autocovarianceShrinkageTapering