A covariance with no parameter

What fitting them together buys

Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.

Worth reading first: The observations that repeat each other · A design is a number.

Three essays of this field have been about an objective. This one is about a coefficient, which is what a practitioner actually receives, and the two answers agree in a way that is worth more than either.

Four fits of the same regression under four dependences: least squares with no whitening at all; the two-step plug-in every whitened rule in this collection runs — residuals, a tapered band from them, one generalised least squares fit through it; the coefficients and the band maximised together over the family; and a whitening at the law’s own covariance, which nobody has.

Fitting them together is worth something under one law. Four fits of the same regression under four dependences: least squares, the two-step plug-in every whitened rule in this collection runs, the coefficients and the band maximised together, and a whitening at the law's own covariance that nobody has. Under the moving average — the one law the band family contains — the joint fit beats the two-step by 0.0077 at 3.3 paired standard errors. Under the autoregression, long memory and the break it is a tie: 0.4, 1.0, 0.3 standard errors. That is the same ordering the likelihood gap gave, arrived at through the coefficients rather than through the objective.
Fig. 1 What each fit delivers, by law, paired on the draw.

The answer, and it is one law

Under a five-period moving average the joint fit beats the two-step by 0.00767 at 3.3 paired standard errors. Under a first-order autoregression, long memory and a break in the persistence it does not: the paired differences are 0.4, 1.0 and 0.4 standard errors, in two cases the wrong way.

The moving average is the law the band family contains — its autocorrelations are zero past the fourth lag by construction, so a band at eight lags can represent it exactly, and the other three laws are outside the set their estimate is chosen from.

That is the whole result, and its interest is that it was predictable from a quantity computed without looking at a single coefficient. The likelihood gap — the rise from the plug-in to the maximum, less what the same optimiser manufactures on a sample where nothing is missing — ranks the four laws 11.87, 1.72, 0.93 and 0.14, and this ranking is the same ranking. Two instruments sharing no arithmetic, one internal to the objective and one external to it, agreeing on which law the plug-in was wrong about.

Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.
Fig. 2 The objective’s version of the same ranking, computed before any coefficient was compared.

The numbers, set out

Under the autoregression the four fits deliver 1.1621, 1.1105, 1.1119 and 1.1058. Least squares is worse than either whitening by about five hundredths, which is the whole value of whitening at all and is what the estimated-covariance field is about. The joint fit and the two-step are apart by 0.0014, which is a tenth of that and is not resolvable.

Under the moving average: 1.1184, 1.0681, 1.0604 and 1.0533. The joint fit takes about a third of the remaining gap to the oracle.

Under long memory: 1.6892, 1.6745, 1.6764 and 1.6619. Every fit is worse in absolute terms because the law is harder — a sequence still at 0.637 at the eighth lag is a sequence no eight-lag band reaches — and the whitening buys much less of what is there to buy.

Under the break: 1.2744, 1.2067, 1.2054 and 1.0815. The gap to the oracle is 0.125, an order of magnitude larger than anything else in the table, and neither estimated rule closes any of it. That is not a failure of the joint fit. A break in the persistence is not a stationary law at all, so its covariance is not Toeplitz, and no banded Toeplitz family — however wide, however well fitted — contains it. The right response to a break is to model it, not to widen a band.

What each rule is worth, under one lawRegret against the best model available, on the fifteen-candidate table at n = 120 over 200 draws, when the errors follow a break in the persistence — 0.95 for the first half of the sample and 0.65 for the second. Reading down: least squares with the ordinary penalty; the whitening told the whole covariance; the first-order autoregression at ρ̂, which is the rule the previous round measured; a Bartlett-tapered Ω̂ at L = 20; an autoregression of an order chosen by the residuals; and the truncated Ω̂, which exists on 24% of draws and is averaged over those. The paired difference between the parameterised rule and the tapered one is -0.00066 at -0.4 standard errors — neither is worth anything against the other here.regret against the best model available — smaller is bettera break in the persistence · n = 120 · 200 draws · window L = 20least squares, 2q0.16925whitened at the true Ω0.01579AR(1) at ρ̂0.10991tapered Ω̂ at L0.11058AR(p̂) whitening0.10941truncated Ω̂ at L — on 24% of draws0.08976200 draws, regret against the best available fitnot stationary
Fig. 3 Each rule against each law, which is where the break’s size in the table above comes from.

Reading the table the other way

There is a second reading of the same four rows, and it is the one a reader who has followed the estimated-covariance field will want.

Take the gap between least squares and the oracle as what is on offer under each law: 0.056, 0.065, 0.027 and 0.193. Then take what the two-step captures of it: 0.052, 0.050, 0.015 and 0.068 — that is 92%, 77%, 55% and 35%. The joint fit adds nothing to the first, third and fourth, and takes the second from 77% to 89%.

So the picture is not a general fit helps under one law. It is an estimated covariance already captures most of what is available under the two smooth stationary laws, and captures a minority of it under the two that are not — and the joint fit is a repair to the middle case only.

Three routes to one number. The lag-one coefficient of the errors, estimated three ways on the same 200 draws under a five-period moving average. The law is 0.800. Reading it off the least-squares residuals gives 0.7418; iterating that rule to its fixed point gives 0.7802; maximising the Gaussian likelihood over the coefficients and the correlation together gives 0.7880. All three are attenuated, because a hundred and twenty rows are a hundred and twenty rows, and the paired gap between the first and the last is 0.0462 at 20.9 standard errors. The spread hardly moves — 0.056 against 0.050 — so what separates them is bias and not noise.
Fig. 4 How much of the dependence each construction recovers, which is the numerator of those shares.

The two laws where a large gap is left open are the two where something other than a better fit of a band is the remedy: long memory needs a sequence with a tail, which a band by definition does not have, and a break needs a model with two regimes in it. Both of those are named in the field that priced the laws, and neither is a tuning problem.

The likelihood gap predicts the size, not only the order

The two instruments are reported as agreeing on a ranking, and the agreement goes further than that: the likelihood gap predicts how large each coefficient gain should be, and it predicts that three of the four are unmeasurable.

The moving average’s gap of 11.87 buys 0.00767 of coefficient error, which is 6.5 × 10⁻⁴ per unit of log-likelihood. Applying that rate to the other three gaps — 1.72, 0.93 and 0.14 — predicts gains of 0.0011, 0.0006 and 0.0001.

Every one of those is inside the standard error of the paired comparison it belongs to, which is why the essay’s three null results are at 0.4, 1.0 and 0.4 standard errors and two of them point the wrong way. The three laws where the joint fit does nothing are three laws where it should do something too small to see, and the exchange rate says how much: about a thousandth of coefficient error per unit of likelihood.

That is a stronger use of the objective than a ranking. It converts an internal quantity, available from the fit itself, into a prediction about the external one a practitioner receives — so a practitioner who has computed the likelihood gap already knows whether the joint fit is worth running, without running it.

The law it helps least is the law it helps most

The capture shares — 92%, 77%, 54% and 35% — read as an ordering of how well the whitening does, and the absolute numbers behind them say the opposite.

What the two-step actually recovers is 0.0516, 0.0503, 0.0147 and 0.0677 of regret under the autoregression, the moving average, long memory and the break. The break is the law where the whitening recovers the most, by a third more than anywhere else, and it is the law whose capture share is lowest.

Both readings are correct and they answer different questions. The share says how much of what was available the rule got; the absolute figure says how much better the practitioner’s estimate is. On a law with an enormous gap, a rule can take a minority of it and still deliver more than a rule taking almost all of a small one.

So “the break is where the whitening does least” is exactly wrong as a description of what an experimenter receives. The break is where whitening is worth most and where the most is still left, and the right reading of its 35% is not that the rule failed but that a band cannot represent a covariance that changes with position — which is a reason to fit the break rather than a reason to stop whitening.

What generality costs where it is not needed

The comparison above is between two general rules. The other comparison a practitioner faces is between a general rule and a parametric one, and it runs the other way.

Against a first-order autoregression at ρ̂ — the rule three fields of this collection run — the joint band fit loses 0.0205 at 3.2 paired standard errors under the law the parameterisation is true in. It wins 0.0097 at 3.1 under the moving average. Under long memory neither, at 0.6 standard errors; under the break the parametric rule wins 0.0096 at 2.5.

What generality costs, and where it pays. The error of the general joint fit minus the error of the parametric one — a first-order autoregression at ρ̂, which is the rule three fields of this collection run — under each law, paired on the draw. Positive is the general fit losing. Under the law the parameterisation is true in it loses 0.0205 at 3.15 paired standard errors, which is the cost of a generality the world does not reward. Under the moving average it wins 0.0097 at 3.09. Under long memory neither, at 0.63: the band cannot represent a sequence that is still 0.637 at the eighth lag any better than one number can. The two costs are the same size, which is the answer the estimated-covariance field reached from the other direction.
Fig. 5 The general fit against the parametric one, by law, paired on the draw.

Two costs the same size, pointing opposite ways, decided by whether the form assumed is the form present. That is the identical shape the estimated-covariance field reaches from the other direction, and finding it again through a different construction is worth as much as finding it the first time: it says the trade is a property of the problem rather than of the rule that measured it.

The practical form of it is a question rather than a recommendation. Is the dependence’s form known? If it is — a carry-over of a known length, an exponential decay from a known mechanism — the parametric fit is better and estimating a general covariance is paying two hundredths for a generality nothing rewards. If it is not, the general fit is the one that cannot be badly wrong.

Three sequences, and one of them is not in the family. One sample of 120 rows under a five-period moving average. The law's own autocorrelations are the top line: 0.800, 0.600, 0.400, 0.200 at the first four lags. The residuals' untapered sample sequence is already far below it, because a fit takes memory out and a hundred and twenty rows take more; the tapered plug-in every whitened rule in this collection uses is lower again, at 0.649 for a first lag of 0.800. The sequence that maximises the likelihood over the same band lies between the two — 0.709 — which is the taper's shrinkage being partly undone. The law's own line is drawn rather than offered as a candidate because, cut off at 8 lags, it is still a covariance matrix and is in the family.
Fig. 6 How little the maximiser moves the sequence under the law it helps on, which is why the coefficients move less.

Why the joint fit wins so little where it wins

Even under the moving average the joint fit takes about a third of the two-step’s remaining gap to the oracle, and not more. The reason is worth a paragraph because it explains why the field’s answer is small everywhere.

A regression coefficient under generalised least squares is a weighted average, and the weights come from Ω1\Omega^{-1}. Two covariance estimates that differ by a shrinkage of the off-diagonals produce weights that differ by less, because inverting compresses differences in the middle of the spectrum, and the coefficient is a further average on top of that. So a correction that is worth twelve log-likelihood units on the objective is worth well under one per cent on the fit.

That is the same insensitivity the joint fit of a single correlation reports — a coefficient moving twenty-one standard errors and a decision moving two — and it is the reason both fields end in the same place. Measure the thing the choice feeds, not the thing the choice is about.

A wider band is always a better fit. The likelihood maximised over the band, at five widths, averaged over 30 samples. A band at L lags is a band at L + 1 with the last entry held at zero, so the families are nested and the maximised likelihood cannot fall — it does not, on any draw. What it does is rise at 0.984 of log-likelihood a lag. A parameter that is doing nothing buys half a unit in expectation and Akaike's criterion charges one, so this is a criterion very nearly indifferent between every width on offer. The dashed line is what a charge of one unit a lag would exactly cancel. Nothing in the fit chooses a width, and what does choose one is a charge somebody has to pick.
Fig. 7 The sweep the joint fit has to be run across when the width is not fixed, which is where its cost comes from.

What it costs to run

The joint fit is a coordinate ascent over eight numbers with a banded Cholesky inside it, restarted at every width when the width is being swept. The two-step is one pass. On the sweeps here the joint fit is between thirty and eighty times the arithmetic, and it is the reason every measurement in this field runs at sixty draws where the rest of the collection runs at two or four hundred.

That is not a small consideration for a practitioner and it is worth stating alongside the delivered gains. A rule that is thirty times dearer and better under one dependence in four, by under one per cent of the error, is not a rule to reach for by default.

Where it is worth reaching for is exactly where the objective says the plug-in is short, and that is a diagnostic anybody can run: fit the plug-in, maximise over the same band, and compare the rise against the rise on a series generated from the plug-in’s own covariance. It is two extra fits rather than a decision, it costs what one joint fit costs, and it answers the question the coefficients take a Monte Carlo study to answer.

One more consideration decides it in practice, and it is not a statistical one. A two-step rule produces the same answer twice; a joint fit produces the answer its optimiser reached, from the start it was given, in the time it was allowed. Every measurement in this field warm-starts the optimiser from the plug-in for exactly that reason, and the sweep over widths has to warm-start each width from the previous one to make a monotonicity that is arithmetic actually hold. A rule whose output depends on how it was started is a rule that needs its starting point written down beside its answer.

What the oracle column is for

Three of the four columns are rules and the fourth is not, and it is in the table for a specific reason rather than for completeness.

The oracle whitening — at the law’s own covariance, through a dense factor because three of the four laws are not banded — is what says how much of the error is available to any covariance rule at all. Under the autoregression it is 1.1058 against least squares at 1.1621, so five hundredths are on offer and the two-step already takes nine tenths of them. Under the break 1.0815 against 1.2744, so nineteen hundredths are on offer and neither estimated rule takes any.

Without that column the two-step and the joint fit look like near-neighbours under every law. With it, the four laws separate into ones where the estimated rules have nearly finished and ones where they have barely started, and the joint fit’s failure to help under the break stops being a disappointment and becomes a statement about what the family can represent.

One thing the comparison cannot settle

Every number here is a risk: the expected squared error of the fitted line’s predictions on fresh rows, averaged over draws. That is the right quantity for a practitioner choosing a procedure, and it is not the quantity a practitioner reporting a result cares about, which is whether the interval around the coefficient is honest.

The two come apart in this collection more often than they agree. An interval that forgets it estimated something is a whole essay about a risk that is fine and a coverage that is not, and the same gap could easily be open here: a covariance maximised on the same sample the coefficient is fitted from is exactly the construction that makes a plug-in standard error too small.

Nothing in this field measures that, and the honest position is that the risk comparison above is silent about it. It is named in the closing section as a deliberate omission rather than left to be inferred.

The two-step is not a bad rule

It is worth saying plainly, because a field named after a joint fit can read as a case against the alternative and this one is not.

The two-step plug-in captures nine tenths of what is available under the law this collection’s regressions are mostly written under, costs one pass, cannot fail to converge, and has no optimiser in it to be tuned or to stop early. Against it the joint fit offers under one per cent of the error on one dependence in four and thirty times the arithmetic.

What the joint fit is for is not routine use. It is for the case where something about the subject says the dependence has a horizon rather than a decay, and for the diagnostic above, which costs one run and says whether the case applies.

What is claimed here, and what is not

This essay takes what a joint fit over a general covariance delivers against the plug-in. The claims are that the joint fit beats the two-step by 0.00767 at 3.3 paired standard errors under a five-period moving average and ties under the other three laws, at 0.4, 1.0 and 0.4 standard errors; that the law it wins under is the one the band family contains; that the ranking of the four laws by this comparison is the ranking the likelihood gap gives, through arithmetic that shares nothing with it; that against a parametric first-order fit the general one loses 0.0205 at 3.2 paired standard errors where the parameterisation is true and wins 0.0097 at 3.1 where it is not; and that under a break the gap between any estimated rule and the law’s own covariance is 0.125, which no banded family closes.

What stays out, and is named as a decision: a family that contains the break. A law whose persistence changes half way through has a covariance that is not Toeplitz, so the whole construction of this field — a sequence, a band, a banded Cholesky — does not apply to it. A two-regime version is buildable and it is a different field’s object, where the break point has to be searched for and charged for; joining the two would mean pricing a joint fit and a searched break at once, which is a field of its own and reaches a different conclusion.

Also out: a standard error for the joint estimate. Everything reported here is a risk averaged over draws, which is what a practitioner experiences. What a joint fit does to the reported uncertainty of a coefficient — whether the usual sandwich is still right when the covariance was maximised rather than plugged in — is a separate question with a separate answer, and putting an unchecked interval beside a measured risk would be putting two different kinds of claim in one table.

The boundary against the parametric joint fit is that it maximises over one number and this one over eight, and the two agree about the shape of the answer: the estimate moves a great deal and what it buys moves very little.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Covariance matrixDependenceEfficiencyGeneralised least squaresLong memoryMisspecificationModel selectionMonte CarloNuisance parameterPaired comparisonProfile likelihoodStructural breakTaperingWhitening