The plug-in and the maximum
Worth reading first: The observations that repeat each other · A design is a number.
Once the band is a family with a parameter, the estimate every field in this collection uses becomes a point in it rather than a matrix — and a point can be asked whether it is the best point.
It is not. Under a first-order autoregression at a hundred and twenty rows, maximising the likelihood over the same eight-lag band raises it by 5.72 from where the tapered plug-in sits. On the scale a criterion charges parameters in, that is eleven parameters’ worth.
The number is real and reporting it that way would be a mistake, and finding out why is most of what this essay is.
The rise a maximiser always finds
Maximising free numbers raises a log-likelihood whatever it is started from. Under a correctly specified model the rise is about in expectation — the same arithmetic that makes a searched parameter cost more than it looks — so an eight-lag band ought to gain about four units from any starting point, including the right one.
Against that, 5.72 is an excess of 1.7 rather than a shortfall of 5.7, and the honest question is what the baseline actually is here. The asymptotic figure will not do: the plug-in is not a random start, the family is constrained, and the optimiser is a local one.
So the baseline is measured rather than asserted, by the only control available. Take the plug-in’s own covariance , generate a fresh error series from it on the same design, and run the whole procedure again. On that sample the band family is correctly specified by construction, the plug-in is as right as an estimate of a member of the family can be, and whatever the optimiser gains is what maximising eight numbers manufactures rather than what the plug-in was missing.
Under the autoregression the control produces 4.79 of the 5.72. What is left is 0.93, at 1.8 standard errors — which is not a finding. Under long memory the raw rise is 4.04 and the control is 3.89, leaving 0.14 at 0.2 standard errors, which is nothing at all.
The law’s own band cannot play the part of this control, and the reason is the previous essay’s: truncated at eight lags it is not a covariance matrix under three of these four laws, so there is no point in the family to start an optimiser from.
The correction is the same size as the finding
The asymptotic baseline and the measured one differ by enough to matter, and the amount is worth stating because it is what the control was built to supply.
Half the free numbers is 4.00. The control returns 4.79 under the autoregression — the optimiser manufactures twenty per cent more than the asymptotic figure — so the correction is 0.79.
The finding, after the correction, is 0.93.
Those are the same size. Had the asymptotic baseline been used, the reported excess would have been 1.72 rather than 0.93 — nearly double — and it would have carried the same standard error, so it would have read at more than three rather than at 1.8. The control does not refine the answer here; it halves it and takes it below the threshold at which anybody would report it.
And the baseline is a property of the law, not of the optimiser
The second law settles that the correction could not have been supplied by any general argument.
Under long memory the control returns 3.89, which is below the asymptotic 4.00 rather than above it. So the two laws’ baselines are 4.79 and 3.89 — a spread of 0.90, which is larger than either law’s measured excess of 0.93 and 0.14.
A single number for what maximising eight parameters manufactures would therefore have been wrong by more than the quantity being measured, in both directions, depending on which law it had been calibrated in. That is the strongest possible case for measuring a baseline rather than citing one, and it is not an argument that could have been made without running the control twice.
What 0.93 is, in the units the field charges in
A criterion charges half a unit a parameter, so the three numbers convert directly.
The raw rise of 5.72 is 11.4 parameters’ worth. The control’s 4.79 is 9.6. The excess of 0.93 is 1.9 parameters — and at 1.8 standard errors, with a standard error of about 0.52, the two-sigma range runs from roughly zero to 3.9 parameters.
So the honest sentence is that the plug-in may be short of the maximum by about two parameters’ worth and may be short by none, and the eleven-parameter reading the raw rise supports is off by a factor of six before any question of significance arises.
Closing that would be cheap. The standard error is 0.52 on sixty draws, so reaching three standard errors on the same 0.93 takes about 170 draws — under three times the count, on a control that has to be run anyway.
Where the plug-in really is short
One of the four laws answers differently, and by a margin nothing else in this field comes near.
Under a five-period moving average the raw rise is 17.72, the control is 5.85, and the excess is 11.87 at 19.4 standard errors — on every draw, not on a majority of them.
The mechanism is visible in one sentence once the two sequences are drawn beside each other. A tapered plug-in multiplies the sample’s -th autocovariance by : a shrinkage, chosen because it makes the result positive definite, and one that is nearly harmless when the true sequence is already decaying smoothly. The moving average’s sequence does not decay smoothly. It runs 0.8, 0.6, 0.4, 0.2 and then stops dead, and the taper is shrinking real structure at exactly the lags where there is most of it to shrink.
So the shortfall is a fact about the taper rather than about the data, and the direction it points is the same everywhere: the maximiser’s sequence is larger than the plug-in’s, by 1.26 times under the autoregression and 1.60 under long memory, summed over the lags that carry weight. The shrinkage is being undone.
Under the autoregression the first lag moves from 0.636 to 0.717. The law’s own value is 0.800 and the sample’s untapered estimate is 0.690, so the likelihood walks about two thirds of the way back from the taper’s answer to the untapered one and stops well short of the truth — which is right, because a fit takes memory out of a residual series before any taper touches it, and no amount of maximising can put back what the design removed.
Two shrinkages, one of them known
There are two reasons the plug-in’s first lag is 0.636 where the law says 0.800, and they are different sizes and have different remedies.
The design takes 0.11 of it. A candidate’s residuals are annihilated by that candidate’s own projection, and the expectation of their autocovariance is computable before any sample is drawn. That is the exact residual expectation, it is 0.7773 at a hundred and twenty errors and 0.7338 for the fullest candidate’s residuals, and it is not something a likelihood can recover: it is a fact about what the residuals are.
The taper takes the rest. That part is a choice, it is undone by maximising, and — this is the finding — undoing it is worth nothing on three of the four laws.
The pair is worth holding together because it is the ordinary shape of an estimation problem in this collection: one component is a bias with a closed form and one is a choice with a cost, and only the second is negotiable. The whole error of a block resample decomposes the same way one field along, and the ratio there is similar — the part anybody has been choosing is the small one.
The four laws, in order
Set out plainly, the excesses are 0.93, 11.87, 0.14 and 1.72 under the autoregression, the moving average, long memory and the break, at 1.8, 19.4, 0.2 and 3.1 standard errors. Two of the four are nothing, one is small and one is enormous, and the ordering is not the ordering of anything else in these fields — it is not by persistence, not by how far the law is from the family, and not by how much the fit removes.
It is by how sharp the true sequence is. The moving average has an edge at the fourth lag; the break has a mild one, because a series whose persistence changes half way through has an autocovariance sequence that is a mixture of two geometrics and bends; and the two smooth laws have none. The taper’s shrinkage is a low-pass filter on the sequence, and a low-pass filter costs something exactly where there is a corner.
That has a practical reading. A dependence that is genuinely banded — an outcome averaged over a fixed window, a design with a fixed carry-over, anything with a horizon rather than a decay — is where estimating the covariance jointly with the regression is worth doing, and a dependence that decays is where it is not. Neither of those is a statement about how strong the dependence is, which is the quantity anybody would have reached for.
What the excess predicts
A measurement that names one law is worth more if something else confirms it, and something does.
The likelihood gap is an internal quantity: it says the plug-in is not where the objective is maximised. It says nothing directly about whether a practitioner is worse off. The comparison that does is what the joint fit delivers in coefficients, and the two readings agree completely — the joint fit beats the two-step under the moving average and is a tie under the other three, in the same order the excess above puts them in.
Two instruments that share no arithmetic and rank four laws identically is the strongest statement this field makes about anything, and it is worth naming the shape: the likelihood knew, before any coefficient was compared, which law the plug-in was wrong about.
What the raw number would have said
It is worth being blunt about the reading that was avoided, because it is the natural one.
Reported without a control, the four raw rises are 5.72, 17.72, 4.04 and 6.61 — every one of them several parameters’ worth, every one of them positive on nearly every draw, and a perfectly coherent story available: estimated covariances are systematically short of the likelihood’s answer, by between four and eighteen units, on every law. Nothing in the numbers contradicts it. The assertion in the library that the machinery has to refuse is exactly that sentence, and it refuses it by asking for the control.
This is the same trap a single sweep’s growth factor sprang one field along, and the same one an average of ratios sprang one field earlier. The common shape: a quantity that is manufactured by the procedure is reported as though it were produced by the data, and the only defence is a run of the procedure on data where there is nothing to find.
Not a sample-size effect
A reader who has met the whole error of a block resample will ask whether this is the same thing at a different address: an estimate short of a truth because a hundred and twenty rows is not many. It is not, and the control is what separates them.
A sample-size effect would show up in the control as well — the generated series is the same length, the design is the same, the plug-in on it is estimated from the same number of residuals. What the control removes is everything that is a property of estimating eight numbers from a hundred and twenty rows, and what it leaves is the mismatch between the shape the taper imposes and the shape the law has.
There is a second reason to believe that. If the shortfall were about the number of rows it would shrink as the rows grew, and the excess under the moving average would be the smallest of the four rather than the largest — the moving average is the law whose sequence a short sample estimates best, because it has only four lags to get right.
The optimiser, and what it is allowed to prove
Coordinate ascent from the plug-in is a local method, and a local method’s answer is a lower bound on the maximum. That cuts one way for each of the two claims here.
For the plug-in is not the maximiser, a lower bound is enough: a point strictly above the plug-in proves the plug-in is not the top, whatever else is up there. That claim is asserted on every draw rather than on an average, and it holds on every draw.
For the shortfall is only this large, a lower bound is not enough — a better optimiser might find more. What protects the comparison is that the control runs the same optimiser, so a shortfall in the search cancels between the two readings as long as it is the same size on both samples. That is an assumption and it is stated as one. It is also checkable in one direction: the control’s rise is 4.79 against the the asymptotic argument gives, which is close enough to say the optimiser is finding most of what is available and far enough to say it is not finding a spurious excess.
What a practitioner should take from it
Three sentences, and the third is the one worth carrying.
A tapered covariance estimate is not the maximum likelihood estimate of the covariance, and nothing in the way it is usually described suggests otherwise — it is presented as an estimator, with a bias and a positive-definiteness guarantee, and neither of those is a claim about a likelihood.
How far short it falls is decided by the shape of the true dependence, not by its strength, and the shape is the thing a practitioner is least likely to know and most likely to be able to reason about from the subject: a horizon is a band, an exponential decay is not.
And a rise found by an optimiser is not evidence of anything until it has been run where there is nothing to find. That is the reusable half. It costs one extra run of the whole procedure on data generated from the answer it produced, and on this measurement it turned an apparent five-and-a-half unit shortfall into nine tenths of a unit at under two standard errors, on the law the field is mostly about.
What is claimed here, and what is not
This essay takes how far an estimated covariance is from maximising the likelihood it is substituted into. The claims are that the raw rise from the tapered plug-in to the maximum over the same eight-lag band is 5.72 log-likelihood units under a first-order autoregression; that 4.79 of it is what the same optimiser produces on a sample generated from the plug-in’s own covariance, where nothing is missing; that the excess is therefore 0.93 at 1.8 standard errors under that law and 0.14 at 0.2 under long memory, and 11.87 at 19.4 standard errors under a five-period moving average, on every draw; that the maximiser’s sequence is 1.26 times the plug-in’s under the autoregression, which is the taper’s shrinkage being partly undone; and that the first lag moves from 0.636 to 0.717 against a law of 0.800 and an untapered sample estimate of 0.690.
What stays out, and is named as a decision: a global optimiser. Coordinate ascent gives a lower bound on the maximum, which is all the first claim needs and is not all the second would like. Replacing it would change what the control is controlling for as well as what the measurement reports, and the two would have to move together — so the honest form of the second claim is a comparison of two runs of one method rather than a statement about the maximum.
Also out: a taper chosen to be right rather than to be positive definite. The whole shortfall here is the Bartlett triangle’s shrinkage, and a window with a faster decay would shrink differently. What each window’s attenuation actually is is computed in the taper field; putting a second window here would be a comparison of two shrinkages when the argument is about there being one at all.
The boundary against the previous essay is that it establishes what the family is and this one asks where in the family the estimate sits. The boundary against what the fit buys is that this one is measured on the objective and that one on the coefficients, and they agree — which is why both are worth having.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The width a band is measured in — both name degrees of freedom, dependence, monte carlo, overfitting, sample autocovariance, shrinkage, tapering
- A lag the sample has less of — both name degrees of freedom, dependence, monte carlo, sample autocovariance, shrinkage, tapering
- The comparison that was not made — both name covariance matrix, degrees of freedom, dependence, monte carlo, tapering, whitening
- What the correction assumes — both name degrees of freedom, dependence, monte carlo, sample autocovariance, shrinkage, tapering
- A dependence with a shape — both name bias, covariance matrix, long memory, sample autocovariance, whitening
- The gap a sample shows — both name attenuation, bias, monte carlo, sample autocovariance, tapering
Named objects
A flat tag is an object no other essay names yet.
AttenuationBiasCovariance matrixDegrees of freedomDependenceLong memoryMonte CarloOverfittingParametric bootstrapProfile likelihoodSample autocovarianceShrinkageTaperingWhitening