What fitting them together buys
Worth reading first: The observations that repeat each other · A design is a number.
Three essays of this field have been about an objective. This one is about a coefficient, which is what a practitioner actually receives, and the two answers agree in a way that is worth more than either.
Four fits of the same regression under four dependences: least squares with no whitening at all; the two-step plug-in every whitened rule in this collection runs — residuals, a tapered band from them, one generalised least squares fit through it; the coefficients and the band maximised together over the family; and a whitening at the law’s own covariance, which nobody has.
The answer, and it is one law
Under a five-period moving average the joint fit beats the two-step by 0.00767 at 3.3 paired standard errors. Under a first-order autoregression, long memory and a break in the persistence it does not: the paired differences are 0.4, 1.0 and 0.4 standard errors, in two cases the wrong way.
The moving average is the law the band family contains — its autocorrelations are zero past the fourth lag by construction, so a band at eight lags can represent it exactly, and the other three laws are outside the set their estimate is chosen from.
That is the whole result, and its interest is that it was predictable from a quantity computed without looking at a single coefficient. The likelihood gap — the rise from the plug-in to the maximum, less what the same optimiser manufactures on a sample where nothing is missing — ranks the four laws 11.87, 1.72, 0.93 and 0.14, and this ranking is the same ranking. Two instruments sharing no arithmetic, one internal to the objective and one external to it, agreeing on which law the plug-in was wrong about.
The numbers, set out
Under the autoregression the four fits deliver 1.1621, 1.1105, 1.1119 and 1.1058. Least squares is worse than either whitening by about five hundredths, which is the whole value of whitening at all and is what the estimated-covariance field is about. The joint fit and the two-step are apart by 0.0014, which is a tenth of that and is not resolvable.
Under the moving average: 1.1184, 1.0681, 1.0604 and 1.0533. The joint fit takes about a third of the remaining gap to the oracle.
Under long memory: 1.6892, 1.6745, 1.6764 and 1.6619. Every fit is worse in absolute terms because the law is harder — a sequence still at 0.637 at the eighth lag is a sequence no eight-lag band reaches — and the whitening buys much less of what is there to buy.
Under the break: 1.2744, 1.2067, 1.2054 and 1.0815. The gap to the oracle is 0.125, an order of magnitude larger than anything else in the table, and neither estimated rule closes any of it. That is not a failure of the joint fit. A break in the persistence is not a stationary law at all, so its covariance is not Toeplitz, and no banded Toeplitz family — however wide, however well fitted — contains it. The right response to a break is to model it, not to widen a band.
Reading the table the other way
There is a second reading of the same four rows, and it is the one a reader who has followed the estimated-covariance field will want.
Take the gap between least squares and the oracle as what is on offer under each law: 0.056, 0.065, 0.027 and 0.193. Then take what the two-step captures of it: 0.052, 0.050, 0.015 and 0.068 — that is 92%, 77%, 55% and 35%. The joint fit adds nothing to the first, third and fourth, and takes the second from 77% to 89%.
So the picture is not a general fit helps under one law. It is an estimated covariance already captures most of what is available under the two smooth stationary laws, and captures a minority of it under the two that are not — and the joint fit is a repair to the middle case only.
The two laws where a large gap is left open are the two where something other than a better fit of a band is the remedy: long memory needs a sequence with a tail, which a band by definition does not have, and a break needs a model with two regimes in it. Both of those are named in the field that priced the laws, and neither is a tuning problem.
The likelihood gap predicts the size, not only the order
The two instruments are reported as agreeing on a ranking, and the agreement goes further than that: the likelihood gap predicts how large each coefficient gain should be, and it predicts that three of the four are unmeasurable.
The moving average’s gap of 11.87 buys 0.00767 of coefficient error, which is 6.5 × 10⁻⁴ per unit of log-likelihood. Applying that rate to the other three gaps — 1.72, 0.93 and 0.14 — predicts gains of 0.0011, 0.0006 and 0.0001.
Every one of those is inside the standard error of the paired comparison it belongs to, which is why the essay’s three null results are at 0.4, 1.0 and 0.4 standard errors and two of them point the wrong way. The three laws where the joint fit does nothing are three laws where it should do something too small to see, and the exchange rate says how much: about a thousandth of coefficient error per unit of likelihood.
That is a stronger use of the objective than a ranking. It converts an internal quantity, available from the fit itself, into a prediction about the external one a practitioner receives — so a practitioner who has computed the likelihood gap already knows whether the joint fit is worth running, without running it.
The law it helps least is the law it helps most
The capture shares — 92%, 77%, 54% and 35% — read as an ordering of how well the whitening does, and the absolute numbers behind them say the opposite.
What the two-step actually recovers is 0.0516, 0.0503, 0.0147 and 0.0677 of regret under the autoregression, the moving average, long memory and the break. The break is the law where the whitening recovers the most, by a third more than anywhere else, and it is the law whose capture share is lowest.
Both readings are correct and they answer different questions. The share says how much of what was available the rule got; the absolute figure says how much better the practitioner’s estimate is. On a law with an enormous gap, a rule can take a minority of it and still deliver more than a rule taking almost all of a small one.
So “the break is where the whitening does least” is exactly wrong as a description of what an experimenter receives. The break is where whitening is worth most and where the most is still left, and the right reading of its 35% is not that the rule failed but that a band cannot represent a covariance that changes with position — which is a reason to fit the break rather than a reason to stop whitening.
What generality costs where it is not needed
The comparison above is between two general rules. The other comparison a practitioner faces is between a general rule and a parametric one, and it runs the other way.
Against a first-order autoregression at ρ̂ — the rule three fields of this collection run — the joint band fit loses 0.0205 at 3.2 paired standard errors under the law the parameterisation is true in. It wins 0.0097 at 3.1 under the moving average. Under long memory neither, at 0.6 standard errors; under the break the parametric rule wins 0.0096 at 2.5.
Two costs the same size, pointing opposite ways, decided by whether the form assumed is the form present. That is the identical shape the estimated-covariance field reaches from the other direction, and finding it again through a different construction is worth as much as finding it the first time: it says the trade is a property of the problem rather than of the rule that measured it.
The practical form of it is a question rather than a recommendation. Is the dependence’s form known? If it is — a carry-over of a known length, an exponential decay from a known mechanism — the parametric fit is better and estimating a general covariance is paying two hundredths for a generality nothing rewards. If it is not, the general fit is the one that cannot be badly wrong.
Why the joint fit wins so little where it wins
Even under the moving average the joint fit takes about a third of the two-step’s remaining gap to the oracle, and not more. The reason is worth a paragraph because it explains why the field’s answer is small everywhere.
A regression coefficient under generalised least squares is a weighted average, and the weights come from . Two covariance estimates that differ by a shrinkage of the off-diagonals produce weights that differ by less, because inverting compresses differences in the middle of the spectrum, and the coefficient is a further average on top of that. So a correction that is worth twelve log-likelihood units on the objective is worth well under one per cent on the fit.
That is the same insensitivity the joint fit of a single correlation reports — a coefficient moving twenty-one standard errors and a decision moving two — and it is the reason both fields end in the same place. Measure the thing the choice feeds, not the thing the choice is about.
What it costs to run
The joint fit is a coordinate ascent over eight numbers with a banded Cholesky inside it, restarted at every width when the width is being swept. The two-step is one pass. On the sweeps here the joint fit is between thirty and eighty times the arithmetic, and it is the reason every measurement in this field runs at sixty draws where the rest of the collection runs at two or four hundred.
That is not a small consideration for a practitioner and it is worth stating alongside the delivered gains. A rule that is thirty times dearer and better under one dependence in four, by under one per cent of the error, is not a rule to reach for by default.
Where it is worth reaching for is exactly where the objective says the plug-in is short, and that is a diagnostic anybody can run: fit the plug-in, maximise over the same band, and compare the rise against the rise on a series generated from the plug-in’s own covariance. It is two extra fits rather than a decision, it costs what one joint fit costs, and it answers the question the coefficients take a Monte Carlo study to answer.
One more consideration decides it in practice, and it is not a statistical one. A two-step rule produces the same answer twice; a joint fit produces the answer its optimiser reached, from the start it was given, in the time it was allowed. Every measurement in this field warm-starts the optimiser from the plug-in for exactly that reason, and the sweep over widths has to warm-start each width from the previous one to make a monotonicity that is arithmetic actually hold. A rule whose output depends on how it was started is a rule that needs its starting point written down beside its answer.
What the oracle column is for
Three of the four columns are rules and the fourth is not, and it is in the table for a specific reason rather than for completeness.
The oracle whitening — at the law’s own covariance, through a dense factor because three of the four laws are not banded — is what says how much of the error is available to any covariance rule at all. Under the autoregression it is 1.1058 against least squares at 1.1621, so five hundredths are on offer and the two-step already takes nine tenths of them. Under the break 1.0815 against 1.2744, so nineteen hundredths are on offer and neither estimated rule takes any.
Without that column the two-step and the joint fit look like near-neighbours under every law. With it, the four laws separate into ones where the estimated rules have nearly finished and ones where they have barely started, and the joint fit’s failure to help under the break stops being a disappointment and becomes a statement about what the family can represent.
One thing the comparison cannot settle
Every number here is a risk: the expected squared error of the fitted line’s predictions on fresh rows, averaged over draws. That is the right quantity for a practitioner choosing a procedure, and it is not the quantity a practitioner reporting a result cares about, which is whether the interval around the coefficient is honest.
The two come apart in this collection more often than they agree. An interval that forgets it estimated something is a whole essay about a risk that is fine and a coverage that is not, and the same gap could easily be open here: a covariance maximised on the same sample the coefficient is fitted from is exactly the construction that makes a plug-in standard error too small.
Nothing in this field measures that, and the honest position is that the risk comparison above is silent about it. It is named in the closing section as a deliberate omission rather than left to be inferred.
The two-step is not a bad rule
It is worth saying plainly, because a field named after a joint fit can read as a case against the alternative and this one is not.
The two-step plug-in captures nine tenths of what is available under the law this collection’s regressions are mostly written under, costs one pass, cannot fail to converge, and has no optimiser in it to be tuned or to stop early. Against it the joint fit offers under one per cent of the error on one dependence in four and thirty times the arithmetic.
What the joint fit is for is not routine use. It is for the case where something about the subject says the dependence has a horizon rather than a decay, and for the diagnostic above, which costs one run and says whether the case applies.
What is claimed here, and what is not
This essay takes what a joint fit over a general covariance delivers against the plug-in. The claims are that the joint fit beats the two-step by 0.00767 at 3.3 paired standard errors under a five-period moving average and ties under the other three laws, at 0.4, 1.0 and 0.4 standard errors; that the law it wins under is the one the band family contains; that the ranking of the four laws by this comparison is the ranking the likelihood gap gives, through arithmetic that shares nothing with it; that against a parametric first-order fit the general one loses 0.0205 at 3.2 paired standard errors where the parameterisation is true and wins 0.0097 at 3.1 where it is not; and that under a break the gap between any estimated rule and the law’s own covariance is 0.125, which no banded family closes.
What stays out, and is named as a decision: a family that contains the break. A law whose persistence changes half way through has a covariance that is not Toeplitz, so the whole construction of this field — a sequence, a band, a banded Cholesky — does not apply to it. A two-regime version is buildable and it is a different field’s object, where the break point has to be searched for and charged for; joining the two would mean pricing a joint fit and a searched break at once, which is a field of its own and reaches a different conclusion.
Also out: a standard error for the joint estimate. Everything reported here is a risk averaged over draws, which is what a practitioner experiences. What a joint fit does to the reported uncertainty of a coefficient — whether the usual sandwich is still right when the covariance was maximised rather than plugged in — is a separate question with a separate answer, and putting an unchecked interval beside a measured risk would be putting two different kinds of claim in one table.
The boundary against the parametric joint fit is that it maximises over one number and this one over eight, and the two agree about the shape of the answer: the estimate moves a great deal and what it buys moves very little.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The comparison that was not made — both name covariance matrix, dependence, model selection, monte carlo, nuisance parameter, paired comparison, tapering, whitening
- A charge that depends on the rule — both name dependence, model selection, monte carlo, nuisance parameter, profile likelihood, structural break, whitening
- The fit that takes the memory out — both name covariance matrix, dependence, long memory, model selection, nuisance parameter, structural break, whitening
- A list is not a rule — both name dependence, model selection, monte carlo, nuisance parameter, paired comparison, whitening
- The window a whitening wants — both name covariance matrix, generalised least squares, long memory, model selection, nuisance parameter, whitening
- The window that has to be chosen, and the term that was dropped — both name covariance matrix, dependence, generalised least squares, model selection, nuisance parameter, tapering
Named objects
A flat tag is an object no other essay names yet.
Covariance matrixDependenceEfficiencyGeneralised least squaresLong memoryMisspecificationModel selectionMonte CarloNuisance parameterPaired comparisonProfile likelihoodStructural breakTaperingWhitening