A lag the sample has less of
Worth reading first: A design is a number · The observations that repeat each other.
The first essay of this field leaves a curve on the table: the charge a tapered covariance band actually costs, read per unit of the band’s summed weights, falls from 0.8447 at two lags to 0.7399 at thirty. This essay proposes that the curve is in the wrong variable, and then tries to break the proposal.
The proposal is one line long.
The line
The proposal is not that the earlier field’s denominator is wrong about what it counts. It is that a covariance band has two widths, that they are different numbers, and that a charge is levied in the second.
A sample autocovariance at lag k is
and the sum has n − k terms, not n. On a hundred and twenty rows the first lag is built from a hundred and nineteen products and the thirtieth from ninety. A tapered band keeps w(k) of each of those, so a band of thirty lags is spending its widest weights on the estimates the sample knows least about.
Σ w(k) counts every lag at the weight the window gives it. It says nothing about how much of γ̂(k) there was to keep. So count each lag at the share of the pairs it has:
That is the whole of it. There is no parameter in it, nothing is estimated, and it is computable from the window, the width and the sample size before any data exist.
What it does
Read on the pairs, the eleven measurements from the earlier essay become 0.8566, 0.8479, 0.8058, 0.7967, 0.7989, 0.8066, 0.8111, 0.8141, 0.8158 and 0.8145.
From four lags up they stop falling. Weighting each width by the reciprocal of its own squared standard error, a constant share misses the plateau by 0.0632 per width against 0.6323 on the summed weights — a factor of 10.008 — and the reading at thirty lags is 1.0108 of the reading at four against 0.9363.
The mean over the plateau is 0.8079, which is the constant a straight line through the origin would be levied at. On the summed weights the same mean is 0.7674 and is not a constant anything should be levied at.
A correction with nothing fitted in it takes an order of magnitude off the misfit. That is the finding, and the rest of this essay is the four ways it could be an accident.
Which of the two widths a reader has been handed
Both widths are a count of parameters and they mean different things, so it is worth setting them side by side once.
Σ w(k) is what the band spends. It is the effective dimension of the shrinkage: how many free numbers survive after the window has multiplied the sample autocovariances down. For the Bartlett window it is exactly L/2 at every width, which is a closed form and a satisfying one, and it is the count the earlier field derives and checks.
Σ w(k)(1 − k/n) is what the band gets. It discounts each lag by the share of the sample’s products that went into it. On a hundred and twenty rows the discount is under one per cent at the first lag and a quarter at the thirtieth, so the two widths agree at the narrow end and part at the wide one — 0.5 against 0.4931 at two lags, 15 against 13.6667 at thirty.
The gap between them is therefore a width-dependent discount, which is exactly the shape a curvature in their ratio would need. That is not proof; a discount of the right shape and the wrong size would leave a residual slope, and the measurement is what says the size is right too.
It is arithmetic, and it is checked as arithmetic
The first thing that could be wrong is the sum itself, so it is computed twice.
P(L) = Σ w(k) − (1/n) Σ k w(k), and for the Bartlett window both sums are closed:
so P(L) = L/2 − (1/n)(L(L + 1)/2 − L(2L + 1)/6). The loop and the closed form agree to a part in 10¹² at every width from one to sixty-four and at sample sizes of sixty, a hundred and twenty and four hundred and eighty.
That is a small check and it is not a formality. The one way this correction could be silently wrong is an off-by-one in the lag index, which would move every number here by about 1/n and would leave the whole picture looking exactly as it does.
It reads no data
The second thing that could be wrong is subtler: a “correction with no fitted parameter” that quietly reads the sample is a fit with the fitting hidden, and the difference is invisible from the output.
So the claim is asserted directly. The pairs width is computed twice on the same arguments and required to return identical bits, and its ratio to the summed weights is required to move with the sample size and with nothing else. At thirty lags that ratio is 0.8222 on sixty rows, 0.9111 on a hundred and twenty and 0.9778 on four hundred and eighty.
The direction is the one the mechanism predicts and would be hard to arrange by accident: a longer sample knows more about a long lag, so the correction has less to correct, and the two denominators converge. A correction that grew with the sample size, or that did not move with it at all, would not be about the pairs.
The discount is (L + 2) over three n
The two widths’ ratio is reported at three sample sizes and it has a closed form, which is worth having because it turns the correction into one expression rather than a table.
For the Bartlett window, Σw(k) = L/2 and Σk·w(k) = L(L+1)/2 − L(2L+1)/6, so
P(L)/Σw(L) = 1 − (L + 2)/(3n)
At thirty lags and a hundred and twenty rows that is 1 − 32/360 = 0.9111; at sixty rows, 0.8222; at four hundred and eighty, 0.9778 — the three ratios the essay reports, from an expression with no sum in it.
So the whole correction is a multiplication by 1 − (L + 2)/(3n), which is 1.1% at two lags, 1.7% at four and 8.9% at thirty on this sample. That is the width-dependent discount the essay describes, and now it has a size and a shape: linear in the width, inverse in the sample.
The conversion checks against both profiles at once. Dividing the summed-weight shares by the discount should give the pairs shares, and at two lags 0.8447/0.9889 = 0.8542 against a reported 0.8566, and at thirty 0.7399/0.9111 = 0.8121 against 0.8145. Both ends to within three thousandths, from one expression applied to the earlier field’s own numbers.
The residual curvature is one law’s
Four laws, four reductions, and the four residuals are not alike.
After the correction the misfits are 0.0679, 0.1265, 3.0165 and 0.4716 under the autoregression, the moving average, long memory and the break. Long memory’s residual is forty-four times the autoregression’s, and its reduction factor of 2.48 is the smallest of the four.
So the law the correction helps least is also the law it leaves worst, and it is the one whose autocorrelations do not sum. That is not a coincidence: the pairs discount is a statement about how many products went into each γ̂(k), and it says nothing about whether the γ(k) being estimated themselves die away. A band on a summable dependence is missing a finite tail; a band on long memory is missing an infinite one, and no denominator counts that.
The practical residue is a boundary rather than a caveat. The pairs width is the right count wherever the dependence is summable, which is three of the four laws here and every law a band family contains; on long memory it is still an improvement by a factor of two and a half and still leaves a profile no straight line through the origin describes.
It survives the law
The third thing that could be wrong is that a first-order autoregression at 0.7 is a convenient world.
The profile is re-measured under all four laws this line of fields works in — the autoregression, a moving average, a long-memory process and one with a break in it — at four thousand draws apiece, half the standard because it is four profiles rather than one. How far a constant share sits from the plateau falls under every one:
0.2663 to 0.0679 under the autoregression, 0.7189 to 0.1265 under the moving average, 7.4759 to 3.0165 under long memory, and 2.1376 to 0.4716 under the break — factors of 3.92, 5.69, 2.48 and 4.53.
Four for four is what a piece of arithmetic should give. It is not four flat profiles, and that is the honest half: under long memory the corrected profile still misses a constant by 3.0165 per width, and under a break its level is 1.0899 — above one, meaning the band costs more optimism than its pairs width says it should.
So the pairs count improves every law by a factor of two to six and finishes the job in one of them, which is the case the earlier field measured in.
What the correction is worth against what it replaces
It is fair to ask how much of a change this is in the units anybody levies a charge in, because “a spread of 0.081 becomes a spread of 0.029” is a statement about a ratio and not about a price.
At thirty lags, on a hundred and twenty rows, the summed-weight width of a Bartlett band is 15 and its pairs width is 13.6667. So the correction takes about one and a third parameters off a band that a convention would charge thirty for and the earlier field’s rule would charge about eleven for. In log-likelihood units, at the plateau’s own rate, that is a little over a unit.
A unit is not nothing at the scale a criterion works at — the difference between two candidate widths on this grid is routinely under two — but it is not a repair either. What the correction buys is not a better price at any one width. It is that the same price works at every width, which is a different kind of gain: a rule that is right on average and wrong at both ends will pick badly at both ends, and a reader has no way to know which end they are at.
That is why the comparison this field makes is between shapes rather than between levels, and why the number reported throughout is a spread rather than a mean.
The obvious second correction, and why it is not here
A reader who accepts the pairs argument will immediately propose a second one, and it is worth answering because the answer says what kind of quantity is being corrected.
If a long lag is estimated from fewer products, then it is also estimated worse — its sampling variance is larger, roughly in proportion to 1/(n − k) rather than 1/n — so perhaps the right weight is not 1 − k/n but its square, or some other power. That would be a one-parameter family, and fitting the parameter would certainly fit the profile better than the pairs count does.
It would also stop being arithmetic. The whole claim being made here is that a correction with nothing fitted in it takes an order of magnitude off the misfit, and a correction with an exponent chosen to remove the misfit removes it by construction. A fitted power of the pairs count would be a curve, and the curves are already on the table in the next essay, which prices them against exactly this line.
The relevant question is therefore not whether some transform of the pairs count fits better. It is whether the untransformed one fits well enough that the transform is not worth a constant, and that comparison is made there rather than asserted here.
And it fails where it should not be expected to work
The fourth thing that could be wrong is the most serious, and it is what makes this a limited result rather than a general one.
The correction is stated for a window and measured on four. It repairs one.
Its flattening factor is 10.008 for the Bartlett window, 2.111 for that window squared, 1.884 for it cubed and 1.348 for Parzen. The three it barely touches are not marginal cases: Parzen’s corrected profile still misses a constant by 4.3491 per width against the Bartlett window’s 0.0632, which is a profile with a great deal of shape left in it.
Three of the four windows still want a curve, and this file does not have one that is not fitted.
Why is not established here. The obvious candidate is that a window whose weight is concentrated at short lags — which is what squaring or cubing a Bartlett weight does, and what Parzen does by construction — spends most of its budget where the pairs correction is smallest, so there is less for the correction to do and whatever else is going on is left visible. That is a story, not a measurement, and it is left as one.
The narrow end, which the correction cannot be about
One more limit, and this one follows from the arithmetic rather than surprising it.
At two lags the correction moves the width from 0.5 to 0.4931 and the share from 0.8447 to 0.8566 — up rather than down, and by a hundredth. At thirty lags it moves the width from 15 to 13.6667 and the share from 0.7399 to 0.8145, by seven hundredths. The correction is nearly nothing where k/n is nearly nothing, which is exactly what it should be.
But the share at two lags is above the plateau on both scales, by more than a tenth. So there is a second thing going on at the narrow end, and the pairs count is not it.
Stating that plainly is the point. A field that reported a flattening factor of 2.8 and did not say that its two narrowest widths were excluded from the range the factor is computed over would be reporting a number a reader could not check. The plateau starts at four lags, it starts there because the two below it are off it on both scales, and the two below it are shown in every figure.
One thing this does not touch
The correction changes the denominator and leaves the numerator exactly where it was, and it is worth being explicit that this is not a repair to the measurement of the optimism.
The field that measured it draws two independent worlds per trial from separate streams rather than two halves of one long one, re-profiles the coefficients and the error variance on the replicate, and differences everything from the narrowest width so that the coefficients’ own optimism drops out. Every one of those choices is upstream of anything here, and every number in this essay inherits them.
What that means practically is that the pairs count could be right and the profile could still be wrong, if the optimism were being mis-measured in a way that happened to look like a slope. The strongest evidence against that is the sample-size check above: a mis-measurement of the optimism has no reason to move with n in the direction the pairs ratio does, and the two move together.
What a denominator is for
It is worth stepping back one line, because the shape of this result recurs.
The earlier field derived Σ w(k) as the effective dimension of a linear shrinkage, and that derivation is correct. It is checked against a closed form — the Bartlett window leaves exactly half its width free, at every width, exactly — and the check passes.
What it computes is the dimension of the shrinkage given the objects being shrunk. A charge, though, is not a dimension; it is a price, and a price has to be in the currency the buyer pays in. The buyer here is a sample of a hundred and twenty rows, and what it pays for a long lag is not what it pays for a short one.
So the two quantities were never the same quantity. They agree in the limit — the ratio between them goes to one as the sample lengthens — which is exactly why a derivation about one can be used for the other for a long time without anybody noticing.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a better charge buys — both name bandwidth selection, degrees of freedom, dependence, estimation error, information criterion, model selection, monte carlo, optimism, plug in estimate, tapering
- A rate times a size — both name bandwidth selection, closed form, estimation error, information criterion, long-run variance, model selection, monte carlo
- The window that has to be chosen, and the term that was dropped — both name degrees of freedom, dependence, estimation error, information criterion, long-run variance, model selection, tapering
- A covariance with no parameter in it — both name dependence, estimation error, information criterion, long-run variance, model selection, tapering
- A penalty is a trace — both name closed form, degrees of freedom, dependence, information criterion, model selection, optimism
- A window for every candidate — both name bandwidth selection, information criterion, long-run variance, model selection, optimism, tapering
Named objects
A flat tag is an object no other essay names yet.
Bandwidth selectionClosed formDegrees of freedomDependenceEstimation errorInformation criterionLong-run varianceModel selectionMonte CarloNormalisationOptimismPlug in estimateSample autocovarianceShrinkageTapering