A covariance with no parameter

A family before a fit

A regression's coefficients and one correlation can be maximised together. Replace the correlation with an estimated covariance and there is nothing left for "jointly" to mean — until a set of covariances is named, and the set turns out not to contain the truth.

Worth reading first: The observations that repeat each other · A design is a number.

Two fields of this collection fit a regression whose errors are dependent, and they answer the same question in incompatible ways.

Fitting the dependence with the line maximises the Gaussian likelihood over the coefficients and the correlation together, and reports what that is worth: 0.0525 of correlation at 21.2 standard errors, buying 0.00222 of selection regret. The whole construction rests on the dependence being one number. A covariance with no parameter removes that number. It estimates the whole error covariance from a set of residuals — a tapered band, as many numbers as it has lags — and substitutes the result into a formula.

A substituted estimate is a plug-in, and a plug-in has nothing to be maximised over. So the sentence the second field leaves standing is that a joint fit is unavailable once the dependence stops being a parameter.

That sentence is false, and saying why precisely turns out to be most of a field.

What “jointly” needs

A joint fit needs a family: a set of covariances indexed by something a likelihood can be maximised over. The band is one. Its parameter is the sequence γ(1),,γ(L)\gamma(1), \ldots, \gamma(L) — the autocorrelations it keeps, with everything past the LL-th held at zero — and the concentrated Gaussian log-likelihood

(β,γ)=n2log ⁣(e^Ω(γ)1e^/n)12logΩ(γ)\ell(\beta, \gamma) = -\tfrac{n}{2}\log\!\left(\hat e' \Omega(\gamma)^{-1} \hat e \,/\, n\right) - \tfrac{1}{2}\log\left|\Omega(\gamma)\right|

is an ordinary function of it. The coefficients are profiled out by generalised least squares on the whitened sample and the scale by its own maximum, which leaves a function of the covariance alone.

So the object exists. What does not exist is any reason to expect the plug-in to sit at its maximum, and the first thing worth measuring is how far from it the plug-in is — which is the next essay’s subject and turns out to depend entirely on which law the errors follow.

Three sequences, and one of them is not in the familyOne sample of 120 rows under AR(1) at 0.8. The law's own autocorrelations are the top line: 0.800, 0.640, 0.512, 0.410 at the first four lags. The residuals' untapered sample sequence is already far below it, because a fit takes memory out and a hundred and twenty rows take more; the tapered plug-in every whitened rule in this collection uses is lower again, at 0.613 for a first lag of 0.800. The sequence that maximises the likelihood over the same band lies between the two — 0.731 — which is the taper's shrinkage being partly undone. The law's own line is drawn rather than offered as a candidate because, cut off at 8 lags, it is no longer a covariance matrix at all, so the family does not contain it.00.50012345678lagautocorrelation of the errorsthe lawthe sample, untaperedthe tapered plug-inthe maximiserone sample, seed 88150the taper shrinks and the likelihood pulls back
Fig. 1 The three sequences on one sample: what the law has, what the residuals report, and what the likelihood wants. The slider changes the law.

The family’s dimension is LL, the width of the band. That is a second thing the parameterised version did not have — one number has no width — and it is the parameter nothing in the fit can choose, which is a whole essay further along.

Three defensible rules and three different widths. The error of the fitted line against the width of the band, under a five-period moving average, with the widths four rules choose marked on it. The criterion charging a unit a lag picks 11.7; a heavier charge picks 5.8; the rule of thumb every applied long-run variance uses picks 4; the width that actually minimises the error is 19.1, and it is available to nobody — chosen per draw it has a standard deviation of 10.3, which is most of the range on offer. What the three rules deliver differs by less than a hundredth of the error they are all paying: 1.0665, 1.0706 and 1.0736 against the best available 1.0619. The curve is nearly flat between eight and twenty lags, which is why three rules can disagree by a factor of three and cost almost nothing.
Fig. 2 The error against the width under the one law the family contains, which is where the family’s own dimension first has an answer worth having.

The determinant, which is not optional here

The second term is what makes two covariances comparable at all. Estimating a covariance per candidate rather than once needed logΩ^\log|\hat\Omega| for the same reason: a fit made under one error model and a fit made under another are not comparable through their residual sums, because the two models assign different volumes to the same region of sample space. There the term was needed to compare candidates. Here it is needed to compare covariances, which is what a fit over a family does at every step.

A reader who has met the term in one place will expect it in the other, and it is worth saying that the expectation is right and the reasoning behind it is usually wrong. The natural guess is that without the determinant the fit runs away: a covariance free to become singular can explain any residual vector, so the quadratic form ought to be driveable to zero and only the determinant stops it.

It does not run away, and the reason is one line. The band is normalised to a unit diagonal, so the eigenvalues of Ω\Omega sum to nn whatever else happens. Mass moved into one eigendirection is taken out of the others, and a residual vector with any component in those others pays for it there. Both halves of the objective turn over well inside the region where the sequence is still a covariance matrix.

Both halves turn over before the edge. The plug-in sequence multiplied by a constant, from 0.40 up to the last value at which it is still a covariance matrix. The natural guess is that the residual sum alone runs away — a covariance free to become singular ought to explain any set of residuals — and it does not: it peaks at 0.65 and falls. The unit diagonal is the reason. It fixes the sum of the eigenvalues at 120, so mass moved into one direction has to come out of the others, and the quadratic form pays for it there. The likelihood peaks further out, at 1.12, because the determinant term rewards exactly the concentration the quadratic form is being charged for. What the determinant does here is move the answer, not bound it.
Fig. 3 The plug-in sequence scaled outwards until it stops being a covariance. Neither half of the objective runs away, and the two peak in different places.

What the determinant does is move the answer outwards rather than bound it. On the path drawn above the residual sum alone peaks at 0.65 of the plug-in sequence and the likelihood at 1.12. That is a real difference and it is not the one the argument for the term is usually made from.

One more thing follows from the trace being fixed, and it is worth a sentence because it is the kind of fact that decides an argument without appearing in it. A family whose members all have the same trace cannot buy fit by inflating the scale, so every improvement it finds is an improvement in shape. That is what separates this construction from the ordinary variance-component problem, and it is why the scale never appears below: the residual sum carries it, and a criterion reads it there.

A window with two wrong ends. The regret of a rule whitened by a Bartlett-tapered Ω̂, as the window widens, against three rules that need no window at all. At L = 0 the estimate is the identity and the rule is exactly least squares — 0.08556, the same number to five places. It falls to 0.01754 at L = 20 and rises again by L = 30, because a quarter of the sample's lags are then being estimated from it. The automatic bandwidth a practitioner would reach for, 4(n/100) to the power 2/9, which at this sample size is 4, gives 0.02963 — 69% above the best available window. The rule told the dependence is an AR(1) sits at 0.01440 throughout, which is the price of not knowing the form.
Fig. 4 What an estimated covariance costs a criterion as the window widens, which is the same dial read for a different purpose.

Where the general fit becomes the parametric one

A general construction that claims to contain a special one has to be checked against it, and this one can be checked exactly rather than approximately.

Restrict the band’s sequence to a geometric one, γ(k)=ρk\gamma(k) = \rho^k, and take the width to the whole sample. At L=n1L = n - 1 the banded Toeplitz matrix is the first-order autoregression’s covariance matrix, so the objective above must equal the profile likelihood the parametric field maximises — a different whitening, a different derivation, written months apart.

It does. At sixty rows and a correlation of 0.8 the two agree to 2.8 × 10⁻¹⁴, which is the arithmetic of the two routes rather than a tolerance anybody chose. The hero of this essay is that comparison drawn across every width in between.

Where the general fit becomes the parametric one. The band family's objective at the autoregression's own geometric sequence, cut off at each width, on one sample of 60 rows. The horizontal line is the profile likelihood the parametric fit maximises, written independently through a different whitening. At the full width the two are the same number to 3e-14, which is what says the general construction contains the parametric one rather than resembling it. Below 10 lags there is no line at all: the geometric sequence cut off short is not a covariance matrix, so the objective has nothing to evaluate. Between the two the truncation is briefly above the parametric likelihood — a wrong covariance can fit one sample better than the right one, which is the whole reason a width has to be charged for rather than chosen.
Fig. 5 The band family’s objective at the autoregression’s own sequence, against the width of the band, with the parametric likelihood as a horizontal line.

Two things in that picture are worth more than the agreement at the right-hand end.

Below ten lags there is no curve at all. The geometric sequence cut off at nine lags is not a covariance matrix: the truncated Toeplitz form has a negative eigenvalue, the factorisation fails, and the objective has nothing to evaluate. So the family does not merely approximate the parameterised model at small widths — it does not contain it. Truncation is not a covariance is a result the estimated field already has about sample autocovariances; the same statement about the law’s own sequence is stronger and is the one that matters here.

Between the two ends the truncation is above the parametric likelihood. At twelve lags the objective reads 42.958 against the parametric 42.137. A wrong covariance can fit one sample better than the right one, because it has more numbers in it. That is the ordinary fact behind every selection problem in this collection and it is worth seeing it happen to a covariance, because a covariance is the one thing in a regression nobody usually thinks of as having a dimension.

The family does not contain the truth

The four laws these fields are measured on are a first-order autoregression, a five-period moving average, long memory at d=4/9d = 4/9, and a break in the persistence. All four are standardised to the same first-lag correlation of 0.8, which is what makes them comparable at all.

Cut each law’s own autocorrelation sequence off at eight lags and ask whether what is left is a covariance matrix. For the moving average it is, and trivially: the law is zero past its fourth lag by construction, so the cut removes nothing. For the other three it is not.

What the band family does not contain. A band at eight lags holds every autocorrelation past the eighth at exactly zero. Three of these four laws are still correlated there — 0.168 for AR(1) at 0.8, 0.637 for long memory at d = 4/9, 0.663 for a break in the persistence — and cutting a sequence off does not merely approximate it: the truncated sequence stops being a covariance matrix, so there is no point in the family to call the truth. Only the moving average, whose autocorrelations are zero past the fourth lag by construction, is inside. That is what makes it the law the joint fit is worth anything under.
Fig. 6 What each law still has at the eighth lag, and whether the band family contains it.

The autoregression is still at 0.168 there and long memory at 0.637. Zeroing those is not a small approximation, and the object it produces is not a member of the set the estimate is chosen from. So under three of the four laws the phrase the true covariance names a point outside the family, and the estimate is an estimate of something the family cannot represent.

This is not a defect of the band. It is what makes the band a general construction: a family that contained every law would be a family with no dimension to trade. But it decides which comparisons in this field mean anything, and the pattern runs through everything that follows — what the joint fit is worth is essentially nothing under three of these laws and real under the fourth, and the fourth is the one the family contains.

Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.
Fig. 7 What the plug-in is short of, by law, which is the first thing a family makes it possible to ask.

What the determinant is worth, in the units of the sequence

The two peaks along the scaling path — the residual sum alone at 0.65 of the plug-in sequence and the whole objective at 1.12 — say how much the determinant moves the answer, and the gap is larger than anything else in this field.

Between them is a factor of 1.72 in the amplitude of the sequence. Relative to the plug-in, the likelihood’s answer sits 12% above and the residual sum’s 35% below, so dropping the determinant would move the estimate more than three times as far as the whole distance from the plug-in to the maximum.

That is the useful form of the term’s importance, and it is a different argument from the one the essay corrects. The determinant is not a guard against a runaway; the trace is fixed and nothing runs away. It is the larger of the two forces deciding where the answer sits. A fit that omitted it would not fail, it would land somewhere plausible and wrong, at two thirds of the amplitude the data actually supports — and there is nothing in the output of such a fit to say so, since a sequence shrunk by a third is still a perfectly respectable covariance.

Outside the family, twice over

Three of the four laws lie outside the band family, and the two that are quantified lie outside it in qualitatively different ways.

Summing what a band at eight lags discards makes the difference explicit. The autoregression’s autocorrelations are 0.8^k, so the tail past lag eight sums to 0.8⁹/(1 − 0.8) = 0.671 against a total of 4.0 — the band omits 17% of the summed dependence, which is a real approximation and a finite one.

Long memory at d = 4/9 has autocorrelations falling like k^(2d−1) = k^(−1/9), and that sum does not converge. There is no fraction of the dependence a band at eight lags omits, because the quantity it would be a fraction of is infinite. Every band, at every width, discards infinitely much.

So “the family does not contain the truth” is two statements. For the autoregression it means the family contains something 17% away, and widening the band closes the gap geometrically — at sixteen lags it is 2.8%, at thirty-two 0.08%. For long memory it means the family contains nothing that approaches the truth in that norm at all, and widening the band improves the approximation without ever making the omission small.

That is why the pattern the field goes on to report is the one it is. A joint fit over a family can only find what the family holds, so under the autoregression it is chasing a target the band gets arbitrarily close to and under long memory it is chasing one the band cannot reach from any width. The law the joint fit helps on is the one the family contains exactly, and the ordering of the other three is an ordering of how badly they are missed rather than of anything about the fit.

Why the band and not something else

Three families are available and this field uses one of them, so the choice is worth stating.

A truncated sequence — the sample autocovariances cut off at LL with no weighting — is the obvious construction and it is the one that fails, on a measurable fraction of draws, by not being a covariance matrix at all. The rate at which it fails is measured in the estimated field and it is not small.

A tapered sequence multiplies each lag by a weight that reaches zero, which restores positive definiteness by construction. That is the family here. The weight is the Bartlett triangle, and the reason it works is the same self-convolution identity the taper field establishes about block windows: a sequence that is the autocorrelation of something is a covariance sequence, and a triangle is the self-convolution of a rectangle.

A sieve — an autoregression of order pp fitted to the residuals — is the third, and it is a family too, with its own dimension. It is not used here for a reason that is a finding rather than a preference: its whitening’s own volume has never been carried in this collection’s criteria, and what that omission was doing is a field of its own two along from this one.

What a fit over a family looks like

Nothing about the arithmetic is unusual once the family is named. The optimiser is coordinate ascent from the plug-in, and it is coordinate ascent rather than anything cleverer for one reason: the feasible set is the positive definite cone intersected with the banded Toeplitz matrices, its boundary has no closed form, and a step that leaves it produces no objective value at all rather than a bad one. A coordinate step that fails the factorisation is refused and the step size halved, which needs nothing to be known about the shape of the region.

That is a small point of technique with a general moral. Every optimiser in this collection that searches a constrained set has the same structure — the walk over the admissible assignments of a trial refuses a step outside the set rather than penalising it — and the reason is the same: a constraint that can be evaluated and not described is a constraint a rejection rule can handle and a gradient cannot.

A wider band is always a better fit. The likelihood maximised over the band, at five widths, averaged over 30 samples. A band at L lags is a band at L + 1 with the last entry held at zero, so the families are nested and the maximised likelihood cannot fall — it does not, on any draw. What it does is rise at 0.984 of log-likelihood a lag. A parameter that is doing nothing buys half a unit in expectation and Akaike's criterion charges one, so this is a criterion very nearly indifferent between every width on offer. The dashed line is what a charge of one unit a lag would exactly cancel. Nothing in the fit chooses a width, and what does choose one is a charge somebody has to pick.
Fig. 8 The likelihood maximised over the band at five widths. Nested families cannot fall, and what they rise by is the next essay but one.

What is claimed here, and what is not

This essay takes what a joint fit is over when the dependence has no parameter in it. The claims are that a tapered band at LL lags is a family whose parameter is its own sequence, so the Gaussian likelihood over the coefficients and the covariance together is well defined; that the objective carries a determinant for the same reason a per-candidate criterion does, and that the usual argument for it — that the fit would otherwise run away — is wrong, because a unit diagonal fixes the trace; that at the geometric band taken to the full width the objective is the parametric profile likelihood to 2.8 × 10⁻¹⁴, by two whitenings written independently; that the same sequence cut off below ten lags is not a covariance matrix at sixty rows; and that under three of the four laws here the law’s own sequence, cut off at eight lags, is outside the family altogether.

What stays out, and is named as a decision: a family with a smoothness penalty rather than a cut. A band is the crudest possible restriction — every lag inside is free and every lag outside is zero — and a penalised spectral fit would be a family with a continuous dimension rather than an integer one. It is a different object and it would need its own calibration, and the honest reason it is not here is that the integer dimension is exactly what makes the width’s own selection problem measurable in the units the rest of this collection charges parameters in.

Also out: the sieve as the family. It is a family, its dimension is an order rather than a width, and its criterion in this collection has been missing a term the whole time — so putting it here would be comparing two fits under two criteria and calling the difference a difference between families. That comparison is made properly where the missing term is priced.

The boundary against the estimated-covariance field is that it asks what an estimated Ω̂ costs a criterion and this one asks what set the estimate was chosen from. Both are about the same matrix and only the second gives it a dimension.

What generality costs, and where it pays. The error of the general joint fit minus the error of the parametric one — a first-order autoregression at ρ̂, which is the rule three fields of this collection run — under each law, paired on the draw. Positive is the general fit losing. Under the law the parameterisation is true in it loses 0.0205 at 3.15 paired standard errors, which is the cost of a generality the world does not reward. Under the moving average it wins 0.0097 at 3.09. Under long memory neither, at 0.63: the band cannot represent a sequence that is still 0.637 at the eighth lag any better than one number can. The two costs are the same size, which is the answer the estimated-covariance field reached from the other direction.
Fig. 9 What generality costs where the form is known and buys where it is not, which is the field’s last measurement.

What the field is going to say

Four essays, and the shortest statement of each.

A family supplies the parameter, and it does not supply the truth: three of these four laws are outside the set their estimate is chosen from.

The plug-in is not the maximiser, and how far short it falls is a fact about the taper rather than about the data — nothing under two of the laws and 11.9 units of log-likelihood under the one whose autocorrelations stop dead.

Nothing in the fit picks the width. Nested families buy about a unit of log-likelihood a lag, which is the order of what a criterion charges, so three defensible rules give three widths a factor of three apart and deliver errors within one per cent of each other.

And fitting them together buys something under exactly one law — the one the family contains — where it beats the two-step by 3.3 paired standard errors, and is a tie under the other three.

The thread running through it is the one the estimated-covariance field states about criteria and this field states about the covariance itself: the cost of generality where it is not needed and the benefit of it where it is come out the same size.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationCholeskyClosed formCovariance matrixDependenceGeneralised least squaresIdentificationLong memoryModel selectionNuisance parameterPositive definiteProfile likelihoodTaperingWhitening