Nothing in the fit picks the width
Worth reading first: A design is a number · The observations that repeat each other.
A band at eight lags is a band at nine with the last entry held at zero. That one sentence settles what a likelihood can say about the width of a band, and the answer is nothing.
Nested families cannot lose. The maximised likelihood at nine lags is at least the maximised likelihood at eight, because the eight-lag answer is a feasible point of the nine-lag problem, and the family the previous essays set up is nested in exactly that way. So a rule that chose the width by the fit it buys would take the widest band on offer, on every draw, for ever.
That much is arithmetic. What has to be measured is the rate.
About one unit a lag
Maximised over the band at five widths, the likelihood reads 57.77, 65.38, 69.14, 76.58 and 79.43 at two, four, eight, sixteen and twenty-four lags. Over the whole sweep that is 0.984 of log-likelihood per lag, and it never falls on any draw.
The number matters because of what it sits between. A parameter that is doing nothing raises a maximised log-likelihood by about a half in expectation — the chi-square-on-one over two that every selection argument in this collection runs on. Akaike’s criterion charges one. So a rate of 0.984 a lag is a criterion very nearly indifferent between every width on offer: each extra lag buys almost exactly what it is charged.
That is a much more awkward position than either of the two it sits between. A rate well below the charge would mean the criterion picks a narrow band decisively; a rate well above it would mean it picks a wide one decisively. A rate at the charge means the choice is made by noise, and two runs of the same rule on two samples from the same law will land somewhere else.
It also means the answer is sensitive to what is charged, which is the ordinary situation whenever the criterion and the decision want different charges — and here there is no decision to calibrate against until the width feeds something.
Three rules and three answers
Three ways to pick the width are defensible and in use, and a fourth is the benchmark none of them can reach.
A criterion charging a unit a lag picks a mean width of 6.20, with a standard deviation of 1.41 across draws.
A heavier charge — half the log of the sample size a parameter, which is the other standard convention — picks 3.81.
The rule of thumb every applied long-run variance uses, , picks 4 at a hundred and twenty rows and picks it on every draw, because it does not read the data at all.
The width that actually minimises the error of the fitted line picks 13.29, and it is available to nobody: chosen per draw it has a standard deviation of 10.17, which is most of the range on offer.
Two of those rules are three times narrower than a third, and the third is not a rule. That is the state of a question that looks, from inside any one of the three, entirely settled.
What the three rules are actually disagreeing about
It is worth separating two things the three answers could be disagreeing about, because only one of them is a disagreement.
They could be estimating different quantities. The criterion is estimating the width that maximises a penalised likelihood; the rule of thumb is estimating the width that minimises the mean squared error of a long-run variance, which is a different object with a different optimum; and the error-minimising width is a third. Three rules aimed at three targets are not in conflict and their disagreement says nothing.
Or they could be estimating the same quantity badly. That is the case that would matter, and it is the one the sweep is set up to test: every rule is scored on the same downstream quantity, the error of the coefficients the whitened fit produces. Whatever each rule was derived to optimise, what is compared is what it delivers.
Read that way the three are genuinely in conflict, and the conflict is small. Which is the answer, and the reason it takes a paired comparison to see at all.
The rate is not one a lag, and that is why the rules are decisive
“About one unit a lag” is an average, and differencing the five likelihoods says it is an average over a rate that falls by a factor of ten.
From two lags to four the likelihood rises 3.81 units a lag. From four to eight, 0.94. From eight to sixteen, 0.93. From sixteen to twenty-four, 0.36.
So the criterion is not indifferent between widths at all. It is decisive at the narrow end, where each lag buys nearly four times what it costs, and decisive the other way past sixteen, where each lag buys a third of its charge. The indifference is a property of the middle of the range, which is where the average of 0.984 comes from.
Read that way the two criteria’s choices stop being conventions and become first-order conditions. A charge of one should stop where the rate crosses one — somewhere between four and eight lags — and the criterion picks 6.20. A charge of half the log of a hundred and twenty is 2.394, which the rate crosses between two and four lags, and the heavier criterion picks 3.81.
Both rules land exactly where their own charge meets the likelihood’s local rate, which is what a penalised maximum is and is a check on the whole construction rather than a description of it. It also says how to move either answer: the width is a function of the charge alone, through a rate that is measured and published, so a practitioner who wants a wider band has to defend a smaller charge rather than a different width.
What the whole decision is worth
The error curve’s ends put the three rules’ disagreement in proportion.
Across every width measured the error runs from 1.1341 at one lag to 1.1083 at its minimum, so the entire width decision is worth 2.3% of the quantity it is choosing over. The gap between the best and worst of the three rules is 0.00285, which is 0.26%.
So the three rules disagree about a tenth of what the decision is worth, and the decision is worth a fortieth of what whitening at all is worth. A factor of three in the width is a quarter of a per cent, and the reason a paired comparison is needed to see it is that it is a quarter of a per cent rather than that the sweep is short.
The oracle’s 13.29 sits where the rate is still 0.93 — just below the charge of one — which is why the criterion’s 6.20 is short rather than wrong: the two are on the same flat stretch, separated by seven lags that the likelihood values at almost exactly what the criterion charges for them.
And it hardly matters
The error curve is what rescues the situation, and it does so in a way worth being explicit about because it is the opposite of a reassurance.
Read across the widths, the error of the fitted line runs 1.1341, 1.1227, 1.1166, 1.1133, 1.1103, 1.1090, 1.1083, 1.1083, 1.1088, 1.1096 and 1.1109 at one through thirty lags. It falls steeply to about six, is flat between eight and twenty, and rises slowly after. The minimum of the averaged curve is 1.1083 at twelve and sixteen lags together.
So the three rules deliver 1.1104, 1.1141 and 1.1133 against the best available 1.1031. The criterion is short by 0.00734 at 7.9 paired standard errors, the rule of thumb by 0.01019 at 8.5, and the heavier charge is worse than the criterion by 0.00370 at 4.4.
Every one of those gaps is real — the paired standard errors say so, and pairing on the draw is what makes them resolvable at all — and every one of them is under one per cent of the error being paid. A factor of three in the width is worth a hundredth of the quantity it is chosen to minimise.
Why the oracle is unreachable rather than merely unknown
The width that minimises the error on a given draw has a standard deviation of 10.17 across draws, against a mean of 13.29. That is not an estimation problem anybody can win.
It is the shape the specification-search field measures for a different tuning parameter and reaches the same conclusion about: when the target moves more than the estimate of it can, every feasible rule is short by roughly the same amount, and the differences between feasible rules are second-order.
The mechanism is the flat curve again, read the other way. A curve that is flat between eight and twenty lags is a curve on which the argmin is decided by noise, so the argmin is noisy — and the noise costs almost nothing, because the curve is flat. The unreachability of the oracle and the harmlessness of missing it are the same fact.
What the criterion is charging for
There is a term worth naming here, because it is the one that makes the charge meaningful at all.
A criterion charging a unit a lag is charging for the covariance’s own dimension, and this collection went eleven rounds without one. Every whitened criterion before this field treats the window as given: it is chosen from the residuals by a separate rule, and the candidate’s score carries no charge for it. That is defensible when the same window is used for every candidate, because a constant subtracted from every score changes nothing — a shared nuisance cancels out of every difference.
It stops being defensible the moment the width is being chosen, and it is the same omission that costs a per-candidate sieve its whole comparison two fields along. The determinant and the dimension are the two terms an estimated covariance owes a criterion, and this collection had neither until it started varying one.
The width and what it is a width of
One more thing separates this dial from the ones the rest of the collection turns, and it is easy to miss because the arithmetic looks identical.
A regression’s dimension is a claim about the world: a candidate with four predictors says four things matter. The covariance’s dimension is a claim about an approximation — the true sequence is not zero past any lag under three of the four laws here, so no width is right and every width is an admission. That changes what a charge is for. Charging a regression’s parameters trades bias for variance in a quantity that has a true value; charging a covariance’s trades the same two things in a quantity that has none.
The practical consequence is the flat curve. A dial with no correct setting, feeding a quantity that is an average over the whole sequence, is a dial whose exact position matters less than a dial with a correct setting would — and that is what the sweep reports rather than what anybody assumed.
What a practitioner does
The reading is not use this width, and pretending otherwise would be the mistake this whole field is about.
If the width feeds a coefficient, anything between eight and twenty is within a hundredth of the best available and the rule of thumb’s four is not: it is short by 0.0102 where the criterion’s six is short by 0.0073. Widening the rule of thumb costs nothing and buys a third of the available gap.
If the width feeds a selection, the answer is elsewhere. A criterion comparing candidates through an estimated covariance is comparing them through the same covariance, and what the width does to that comparison is a different sweep with a different shape — where the cost is in the middle rather than at the ends.
And if the width is being chosen per candidate rather than once, this essay is not the relevant measurement at all. That is a selection problem inside a selection problem and it is priced in its own field, where the answer is that the two ends of the argument — a window and an order — cost the same once they are scored by the same criterion.
The width and the law
One column of the sweep is worth reporting separately because it is the only place the four laws disagree about anything here.
Under the moving average the criterion’s width lands close to the law’s own band, which is what a correctly specified family should do. Under long memory nothing lands anywhere useful, because the law’s sequence is still 0.637 at the eighth lag and no width inside the sweep captures it — the family does not contain that law and widening the band walks towards it without arriving.
That is the honest limit of a criterion here. A charge per lag prices the cost of a wider band correctly and prices the benefit only within the family; against a law outside it, the benefit keeps arriving and the charge keeps being levied, and the two nearly cancel — which is exactly the 0.984 a lag the sweep reports.
The one thing the likelihood does settle
None of this says the likelihood is useless about the covariance. It settles the sequence at a given width completely — that is what the maximiser the plug-in is not is about, and the answer there is sharp and law-dependent — and it settles nothing at all about how many lags there should be.
The division is exactly the one every nested-model problem in this collection has. A table of nested models has a likelihood that ranks the fits within a dimension and cannot rank across dimensions; the order a sieve is fitted at has the same shape; and the covariance turns out to be one more instance of it rather than a new kind of difficulty. What is new is only that nobody was treating the covariance as having a dimension, so the charge was missing rather than disputed.
There is one asymmetry worth naming. In a table of regressions the dimension is what a practitioner is trying to learn; here it is a nuisance, and nobody wants to know how many lags the errors have. That makes the flatness of the delivered error a genuinely good outcome rather than a disappointment: a nuisance whose misestimation costs one per cent is a nuisance that has been handled.
What is claimed here, and what is not
This essay takes whether anything in a fit chooses the width of an estimated covariance. The claims are that the band families are nested, so the maximised likelihood is non-decreasing in the width and is so on every draw; that it rises by 0.984 of log-likelihood a lag over widths from two to twenty-four, which is the order of what a criterion charges; that a criterion charging a unit a lag picks a mean width of 6.20, a heavier charge 3.81, the rule of thumb 4 and the error-minimising width 13.29 with a standard deviation of 10.17; and that the three feasible rules deliver 1.1104, 1.1141 and 1.1133 against a best available 1.1031, so a factor of three in the width is worth under one per cent of the error.
What stays out, and is named as a decision: a charge derived rather than assumed. Both charges used here are the standard conventions, and neither is derived for this problem — the covariance’s parameters are not the regression’s, they enter the objective through a determinant as well as a quadratic form, and the effective number of them a taper leaves free is smaller than . Deriving the right charge is a real piece of work and it would change the widths in the table; what it would not change is the rate, which is measured and is what makes the choice noisy in the first place.
Also out: a comparison against a cross-validated width. Splitting a dependent series to validate on is a construction with its own assumptions about where to cut, and putting an untested one beside three tested rules would be putting a fourth answer in a table that already has three too many.
The boundary against the window a whitening wants is that it chooses the width from the residuals’ own criterion and this one asks what the fit can say about it. They are different criteria and they do not have to agree, which is the whole subject of the field after next.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A line that beats two curves — both name closed form, degrees of freedom, information criterion, model selection, nested models, overfitting, tapering
- A list is not a rule — both name bandwidth, closed form, information criterion, model selection, nuisance parameter, regret, whitening
- A window for every candidate — both name information criterion, model selection, nuisance parameter, overfitting, regret, tapering, whitening
- Where the generality runs out — both name covariance matrix, information criterion, model selection, nuisance parameter, profile likelihood, regret, whitening
- A dependence with a shape — both name covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
- The displacement is a parameter count — both name closed form, degrees of freedom, information criterion, model selection, nested models, overfitting
Named objects
A flat tag is an object no other essay names yet.
BandwidthBias-varianceClosed formCovariance matrixDegrees of freedomInformation criterionModel selectionNested modelsNuisance parameterOverfittingProfile likelihoodRegretTaperingWhitening