A covariance with no parameter

The plug-in and the maximum

A tapered covariance estimate sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. Four fifths of that is what the optimiser would have found if nothing were missing.

Worth reading first: The observations that repeat each other · A design is a number.

Once the band is a family with a parameter, the estimate every field in this collection uses becomes a point in it rather than a matrix — and a point can be asked whether it is the best point.

It is not. Under a first-order autoregression at a hundred and twenty rows, maximising the likelihood over the same eight-lag band raises it by 5.72 from where the tapered plug-in sits. On the scale a criterion charges parameters in, that is eleven parameters’ worth.

The number is real and reporting it that way would be a mistake, and finding out why is most of what this essay is.

The rise a maximiser always finds

Maximising LL free numbers raises a log-likelihood whatever it is started from. Under a correctly specified model the rise is about L/2L/2 in expectation — the same arithmetic that makes a searched parameter cost more than it looks — so an eight-lag band ought to gain about four units from any starting point, including the right one.

Against that, 5.72 is an excess of 1.7 rather than a shortfall of 5.7, and the honest question is what the baseline actually is here. The asymptotic figure will not do: the plug-in is not a random start, the family is constrained, and the optimiser is a local one.

So the baseline is measured rather than asserted, by the only control available. Take the plug-in’s own covariance Ω^\hat\Omega, generate a fresh error series from it on the same design, and run the whole procedure again. On that sample the band family is correctly specified by construction, the plug-in is as right as an estimate of a member of the family can be, and whatever the optimiser gains is what maximising eight numbers manufactures rather than what the plug-in was missing.

Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.
Fig. 1 The rise from the plug-in to the maximum, beside the rise the same optimiser produces when nothing is missing.

Under the autoregression the control produces 4.79 of the 5.72. What is left is 0.93, at 1.8 standard errors — which is not a finding. Under long memory the raw rise is 4.04 and the control is 3.89, leaving 0.14 at 0.2 standard errors, which is nothing at all.

The law’s own band cannot play the part of this control, and the reason is the previous essay’s: truncated at eight lags it is not a covariance matrix under three of these four laws, so there is no point in the family to start an optimiser from.

The correction is the same size as the finding

The asymptotic baseline and the measured one differ by enough to matter, and the amount is worth stating because it is what the control was built to supply.

Half the free numbers is 4.00. The control returns 4.79 under the autoregression — the optimiser manufactures twenty per cent more than the asymptotic figure — so the correction is 0.79.

The finding, after the correction, is 0.93.

Those are the same size. Had the asymptotic baseline been used, the reported excess would have been 1.72 rather than 0.93 — nearly double — and it would have carried the same standard error, so it would have read at more than three rather than at 1.8. The control does not refine the answer here; it halves it and takes it below the threshold at which anybody would report it.

And the baseline is a property of the law, not of the optimiser

The second law settles that the correction could not have been supplied by any general argument.

Under long memory the control returns 3.89, which is below the asymptotic 4.00 rather than above it. So the two laws’ baselines are 4.79 and 3.89 — a spread of 0.90, which is larger than either law’s measured excess of 0.93 and 0.14.

A single number for what maximising eight parameters manufactures would therefore have been wrong by more than the quantity being measured, in both directions, depending on which law it had been calibrated in. That is the strongest possible case for measuring a baseline rather than citing one, and it is not an argument that could have been made without running the control twice.

What 0.93 is, in the units the field charges in

A criterion charges half a unit a parameter, so the three numbers convert directly.

The raw rise of 5.72 is 11.4 parameters’ worth. The control’s 4.79 is 9.6. The excess of 0.93 is 1.9 parameters — and at 1.8 standard errors, with a standard error of about 0.52, the two-sigma range runs from roughly zero to 3.9 parameters.

So the honest sentence is that the plug-in may be short of the maximum by about two parameters’ worth and may be short by none, and the eleven-parameter reading the raw rise supports is off by a factor of six before any question of significance arises.

Closing that would be cheap. The standard error is 0.52 on sixty draws, so reaching three standard errors on the same 0.93 takes about 170 draws — under three times the count, on a control that has to be run anyway.

Where the plug-in really is short

One of the four laws answers differently, and by a margin nothing else in this field comes near.

Under a five-period moving average the raw rise is 17.72, the control is 5.85, and the excess is 11.87 at 19.4 standard errors — on every draw, not on a majority of them.

The mechanism is visible in one sentence once the two sequences are drawn beside each other. A tapered plug-in multiplies the sample’s kk-th autocovariance by 1k/(L+1)1 - k/(L+1): a shrinkage, chosen because it makes the result positive definite, and one that is nearly harmless when the true sequence is already decaying smoothly. The moving average’s sequence does not decay smoothly. It runs 0.8, 0.6, 0.4, 0.2 and then stops dead, and the taper is shrinking real structure at exactly the lags where there is most of it to shrink.

Three sequences, and one of them is not in the familyOne sample of 120 rows under a five-period moving average. The law's own autocorrelations are the top line: 0.800, 0.600, 0.400, 0.200 at the first four lags. The residuals' untapered sample sequence is already far below it, because a fit takes memory out and a hundred and twenty rows take more; the tapered plug-in every whitened rule in this collection uses is lower again, at 0.649 for a first lag of 0.800. The sequence that maximises the likelihood over the same band lies between the two — 0.709 — which is the taper's shrinkage being partly undone. The law's own line is drawn rather than offered as a candidate because, cut off at 8 lags, it is still a covariance matrix and is in the family.-0.50000.50012345678lagautocorrelation of the errorsthe lawthe sample, untaperedthe tapered plug-inthe maximiserone sample, seed 88150the taper shrinks and the likelihood pulls back
Fig. 2 The three sequences under each law: what the law has, what a taper reports, and what the likelihood pulls back to. The slider changes the law.

So the shortfall is a fact about the taper rather than about the data, and the direction it points is the same everywhere: the maximiser’s sequence is larger than the plug-in’s, by 1.26 times under the autoregression and 1.60 under long memory, summed over the lags that carry weight. The shrinkage is being undone.

Under the autoregression the first lag moves from 0.636 to 0.717. The law’s own value is 0.800 and the sample’s untapered estimate is 0.690, so the likelihood walks about two thirds of the way back from the taper’s answer to the untapered one and stops well short of the truth — which is right, because a fit takes memory out of a residual series before any taper touches it, and no amount of maximising can put back what the design removed.

What the band family does not contain. A band at eight lags holds every autocorrelation past the eighth at exactly zero. Three of these four laws are still correlated there — 0.168 for AR(1) at 0.8, 0.637 for long memory at d = 4/9, 0.663 for a break in the persistence — and cutting a sequence off does not merely approximate it: the truncated sequence stops being a covariance matrix, so there is no point in the family to call the truth. Only the moving average, whose autocorrelations are zero past the fourth lag by construction, is inside. That is what makes it the law the joint fit is worth anything under.
Fig. 3 Which laws the band family holds at all, which decides whether there is a shortfall to find.

Two shrinkages, one of them known

There are two reasons the plug-in’s first lag is 0.636 where the law says 0.800, and they are different sizes and have different remedies.

The design takes 0.11 of it. A candidate’s residuals are annihilated by that candidate’s own projection, and the expectation of their autocovariance is computable before any sample is drawn. That is the exact residual expectation, it is 0.7773 at a hundred and twenty errors and 0.7338 for the fullest candidate’s residuals, and it is not something a likelihood can recover: it is a fact about what the residuals are.

The taper takes the rest. That part is a choice, it is undone by maximising, and — this is the finding — undoing it is worth nothing on three of the four laws.

The dependence, at four removes. Under AR(1) at 0.8, four different sequences all called the dependence. The top line is the law. The middle line is what a sample of 120 errors reports on average — computable exactly, because the expectation of a sample autocovariance is arithmetic once the covariance is known. The lower line is what a candidate's residuals report, which is what every two-step rule in this collection actually reads: a fit removes variance, and it removes more of the persistent part than of the rest. At the first lag the three are 0.800, 0.7773 and 0.7338. The dots are counted from draws and share no arithmetic with the line they sit on; the worst departure is 1.2 standard errors.
Fig. 4 What a fit removes from a dependence, exactly, which is the first of the two shrinkages and the one no maximiser can reverse.

The pair is worth holding together because it is the ordinary shape of an estimation problem in this collection: one component is a bias with a closed form and one is a choice with a cost, and only the second is negotiable. The whole error of a block resample decomposes the same way one field along, and the ratio there is similar — the part anybody has been choosing is the small one.

The four laws, in order

Set out plainly, the excesses are 0.93, 11.87, 0.14 and 1.72 under the autoregression, the moving average, long memory and the break, at 1.8, 19.4, 0.2 and 3.1 standard errors. Two of the four are nothing, one is small and one is enormous, and the ordering is not the ordering of anything else in these fields — it is not by persistence, not by how far the law is from the family, and not by how much the fit removes.

It is by how sharp the true sequence is. The moving average has an edge at the fourth lag; the break has a mild one, because a series whose persistence changes half way through has an autocovariance sequence that is a mixture of two geometrics and bends; and the two smooth laws have none. The taper’s shrinkage is a low-pass filter on the sequence, and a low-pass filter costs something exactly where there is a corner.

The sequence a whitening is built from is already short. The solid lines are the law; the dashed ones are what a sample of 120 reports on average, computed exactly rather than simulated — γ̂(k) subtracts a sample mean, and subtracting a mean from a persistent series removes a large part of the series. Under long memory the first lag arrives as 0.538 against a truth of 0.800, and by the twelfth it is 0.089 against 0.609. Under the geometric law the shortfall is smaller and in the same direction: 0.777 against 0.800. Every estimated whitening in this field is built from the dashed line, and the counted values agree with it to 2.3 standard errors.
Fig. 5 What a sample reports of each law’s sequence, which is the object the taper then shrinks.

That has a practical reading. A dependence that is genuinely banded — an outcome averaged over a fixed window, a design with a fixed carry-over, anything with a horizon rather than a decay — is where estimating the covariance jointly with the regression is worth doing, and a dependence that decays is where it is not. Neither of those is a statement about how strong the dependence is, which is the quantity anybody would have reached for.

What the excess predicts

A measurement that names one law is worth more if something else confirms it, and something does.

The likelihood gap is an internal quantity: it says the plug-in is not where the objective is maximised. It says nothing directly about whether a practitioner is worse off. The comparison that does is what the joint fit delivers in coefficients, and the two readings agree completely — the joint fit beats the two-step under the moving average and is a tie under the other three, in the same order the excess above puts them in.

Fitting them together is worth something under one law. Four fits of the same regression under four dependences: least squares, the two-step plug-in every whitened rule in this collection runs, the coefficients and the band maximised together, and a whitening at the law's own covariance that nobody has. Under the moving average — the one law the band family contains — the joint fit beats the two-step by 0.0077 at 3.3 paired standard errors. Under the autoregression, long memory and the break it is a tie: 0.4, 1.0, 0.3 standard errors. That is the same ordering the likelihood gap gave, arrived at through the coefficients rather than through the objective.
Fig. 6 The same ordering reached through the coefficients rather than through the objective.

Two instruments that share no arithmetic and rank four laws identically is the strongest statement this field makes about anything, and it is worth naming the shape: the likelihood knew, before any coefficient was compared, which law the plug-in was wrong about.

A wider band is always a better fit. The likelihood maximised over the band, at five widths, averaged over 30 samples. A band at L lags is a band at L + 1 with the last entry held at zero, so the families are nested and the maximised likelihood cannot fall — it does not, on any draw. What it does is rise at 1.579 of log-likelihood a lag. A parameter that is doing nothing buys half a unit in expectation and Akaike's criterion charges one, so this is a criterion very nearly indifferent between every width on offer. The dashed line is what a charge of one unit a lag would exactly cancel. Nothing in the fit chooses a width, and what does choose one is a charge somebody has to pick.
Fig. 7 The other place in this field where a rise has to be read against what a nested family buys anyway.

What the raw number would have said

It is worth being blunt about the reading that was avoided, because it is the natural one.

Reported without a control, the four raw rises are 5.72, 17.72, 4.04 and 6.61 — every one of them several parameters’ worth, every one of them positive on nearly every draw, and a perfectly coherent story available: estimated covariances are systematically short of the likelihood’s answer, by between four and eighteen units, on every law. Nothing in the numbers contradicts it. The assertion in the library that the machinery has to refuse is exactly that sentence, and it refuses it by asking for the control.

This is the same trap a single sweep’s growth factor sprang one field along, and the same one an average of ratios sprang one field earlier. The common shape: a quantity that is manufactured by the procedure is reported as though it were produced by the data, and the only defence is a run of the procedure on data where there is nothing to find.

Where the general fit becomes the parametric one. The band family's objective at the autoregression's own geometric sequence, cut off at each width, on one sample of 60 rows. The horizontal line is the profile likelihood the parametric fit maximises, written independently through a different whitening. At the full width the two are the same number to 3e-14, which is what says the general construction contains the parametric one rather than resembling it. Below 10 lags there is no line at all: the geometric sequence cut off short is not a covariance matrix, so the objective has nothing to evaluate. Between the two the truncation is briefly above the parametric likelihood — a wrong covariance can fit one sample better than the right one, which is the whole reason a width has to be charged for rather than chosen.
Fig. 8 The reduction check, which is the one place in this field two routes to one number are available.

Not a sample-size effect

A reader who has met the whole error of a block resample will ask whether this is the same thing at a different address: an estimate short of a truth because a hundred and twenty rows is not many. It is not, and the control is what separates them.

A sample-size effect would show up in the control as well — the generated series is the same length, the design is the same, the plug-in on it is estimated from the same number of residuals. What the control removes is everything that is a property of estimating eight numbers from a hundred and twenty rows, and what it leaves is the mismatch between the shape the taper imposes and the shape the law has.

There is a second reason to believe that. If the shortfall were about the number of rows it would shrink as the rows grew, and the excess under the moving average would be the smallest of the four rather than the largest — the moving average is the law whose sequence a short sample estimates best, because it has only four lags to get right.

The optimiser, and what it is allowed to prove

Coordinate ascent from the plug-in is a local method, and a local method’s answer is a lower bound on the maximum. That cuts one way for each of the two claims here.

For the plug-in is not the maximiser, a lower bound is enough: a point strictly above the plug-in proves the plug-in is not the top, whatever else is up there. That claim is asserted on every draw rather than on an average, and it holds on every draw.

For the shortfall is only this large, a lower bound is not enough — a better optimiser might find more. What protects the comparison is that the control runs the same optimiser, so a shortfall in the search cancels between the two readings as long as it is the same size on both samples. That is an assumption and it is stated as one. It is also checkable in one direction: the control’s rise is 4.79 against the L/2=4L/2 = 4 the asymptotic argument gives, which is close enough to say the optimiser is finding most of what is available and far enough to say it is not finding a spurious excess.

Both halves turn over before the edge. The plug-in sequence multiplied by a constant, from 0.40 up to the last value at which it is still a covariance matrix. The natural guess is that the residual sum alone runs away — a covariance free to become singular ought to explain any set of residuals — and it does not: it peaks at 0.58 and falls. The unit diagonal is the reason. It fixes the sum of the eigenvalues at 120, so mass moved into one direction has to come out of the others, and the quadratic form pays for it there. The likelihood peaks further out, at 1.15, because the determinant term rewards exactly the concentration the quadratic form is being charged for. What the determinant does here is move the answer, not bound it.
Fig. 9 One dimension of the same surface, on the law where the excess is smallest, with both halves of the objective drawn.

What a practitioner should take from it

Three sentences, and the third is the one worth carrying.

A tapered covariance estimate is not the maximum likelihood estimate of the covariance, and nothing in the way it is usually described suggests otherwise — it is presented as an estimator, with a bias and a positive-definiteness guarantee, and neither of those is a claim about a likelihood.

How far short it falls is decided by the shape of the true dependence, not by its strength, and the shape is the thing a practitioner is least likely to know and most likely to be able to reason about from the subject: a horizon is a band, an exponential decay is not.

And a rise found by an optimiser is not evidence of anything until it has been run where there is nothing to find. That is the reusable half. It costs one extra run of the whole procedure on data generated from the answer it produced, and on this measurement it turned an apparent five-and-a-half unit shortfall into nine tenths of a unit at under two standard errors, on the law the field is mostly about.

What is claimed here, and what is not

This essay takes how far an estimated covariance is from maximising the likelihood it is substituted into. The claims are that the raw rise from the tapered plug-in to the maximum over the same eight-lag band is 5.72 log-likelihood units under a first-order autoregression; that 4.79 of it is what the same optimiser produces on a sample generated from the plug-in’s own covariance, where nothing is missing; that the excess is therefore 0.93 at 1.8 standard errors under that law and 0.14 at 0.2 under long memory, and 11.87 at 19.4 standard errors under a five-period moving average, on every draw; that the maximiser’s sequence is 1.26 times the plug-in’s under the autoregression, which is the taper’s shrinkage being partly undone; and that the first lag moves from 0.636 to 0.717 against a law of 0.800 and an untapered sample estimate of 0.690.

What stays out, and is named as a decision: a global optimiser. Coordinate ascent gives a lower bound on the maximum, which is all the first claim needs and is not all the second would like. Replacing it would change what the control is controlling for as well as what the measurement reports, and the two would have to move together — so the honest form of the second claim is a comparison of two runs of one method rather than a statement about the maximum.

Also out: a taper chosen to be right rather than to be positive definite. The whole shortfall here is the Bartlett triangle’s shrinkage, and a window with a faster decay would shrink differently. What each window’s attenuation actually is is computed in the taper field; putting a second window here would be a comparison of two shrinkages when the argument is about there being one at all.

The boundary against the previous essay is that it establishes what the family is and this one asks where in the family the estimate sits. The boundary against what the fit buys is that this one is measured on the objective and that one on the coefficients, and they agree — which is why both are worth having.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AttenuationBiasCovariance matrixDegrees of freedomDependenceLong memoryMonte CarloOverfittingParametric bootstrapProfile likelihoodSample autocovarianceShrinkageTaperingWhitening