A covariance with no parameter

Nothing in the fit picks the width

A wider band is always a better fit, and it is better by about one unit of log-likelihood a lag — which is the order of what a criterion charges for a parameter. Three defensible rules choose widths a factor of three apart.

Worth reading first: A design is a number · The observations that repeat each other.

A band at eight lags is a band at nine with the last entry held at zero. That one sentence settles what a likelihood can say about the width of a band, and the answer is nothing.

Nested families cannot lose. The maximised likelihood at nine lags is at least the maximised likelihood at eight, because the eight-lag answer is a feasible point of the nine-lag problem, and the family the previous essays set up is nested in exactly that way. So a rule that chose the width by the fit it buys would take the widest band on offer, on every draw, for ever.

That much is arithmetic. What has to be measured is the rate.

About one unit a lag

Maximised over the band at five widths, the likelihood reads 57.77, 65.38, 69.14, 76.58 and 79.43 at two, four, eight, sixteen and twenty-four lags. Over the whole sweep that is 0.984 of log-likelihood per lag, and it never falls on any draw.

A wider band is always a better fit. The likelihood maximised over the band, at five widths, averaged over 30 samples. A band at L lags is a band at L + 1 with the last entry held at zero, so the families are nested and the maximised likelihood cannot fall — it does not, on any draw. What it does is rise at 0.984 of log-likelihood a lag. A parameter that is doing nothing buys half a unit in expectation and Akaike's criterion charges one, so this is a criterion very nearly indifferent between every width on offer. The dashed line is what a charge of one unit a lag would exactly cancel. Nothing in the fit chooses a width, and what does choose one is a charge somebody has to pick.
Fig. 1 The likelihood maximised over the band at five widths, with a line at one unit a lag.

The number matters because of what it sits between. A parameter that is doing nothing raises a maximised log-likelihood by about a half in expectation — the chi-square-on-one over two that every selection argument in this collection runs on. Akaike’s criterion charges one. So a rate of 0.984 a lag is a criterion very nearly indifferent between every width on offer: each extra lag buys almost exactly what it is charged.

That is a much more awkward position than either of the two it sits between. A rate well below the charge would mean the criterion picks a narrow band decisively; a rate well above it would mean it picks a wide one decisively. A rate at the charge means the choice is made by noise, and two runs of the same rule on two samples from the same law will land somewhere else.

It also means the answer is sensitive to what is charged, which is the ordinary situation whenever the criterion and the decision want different charges — and here there is no decision to calibrate against until the width feeds something.

Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.
Fig. 2 Where the sequence sits at a given width, which is the question this one is asked on top of.

Three rules and three answers

Three ways to pick the width are defensible and in use, and a fourth is the benchmark none of them can reach.

A criterion charging a unit a lag picks a mean width of 6.20, with a standard deviation of 1.41 across draws.

A heavier charge — half the log of the sample size a parameter, which is the other standard convention — picks 3.81.

The rule of thumb every applied long-run variance uses, 4(n/100)2/94(n/100)^{2/9}, picks 4 at a hundred and twenty rows and picks it on every draw, because it does not read the data at all.

The width that actually minimises the error of the fitted line picks 13.29, and it is available to nobody: chosen per draw it has a standard deviation of 10.17, which is most of the range on offer.

Three defensible rules and three different widthsThe error of the fitted line against the width of the band, under AR(1) at 0.8, with the widths four rules choose marked on it. The criterion charging a unit a lag picks 6.2; a heavier charge picks 3.8; the rule of thumb every applied long-run variance uses picks 4; the width that actually minimises the error is 13.3, and it is available to nobody — chosen per draw it has a standard deviation of 10.2, which is most of the range on offer. What the three rules deliver differs by less than a hundredth of the error they are all paying: 1.1104, 1.1141 and 1.1133 against the best available 1.1031. The curve is nearly flat between eight and twenty lags, which is why three rules can disagree by a factor of three and cost almost nothing.1.111.121.131481216202430lags the band keepserror of the fitted linea heavier chargethe rule of thumbthe criterionthe best width, available to nobody80 draws, AR(1) at 0.8the curve is flat where they disagree
Fig. 3 The error against the width, with the widths four rules choose marked on it. The slider changes the law.

Two of those rules are three times narrower than a third, and the third is not a rule. That is the state of a question that looks, from inside any one of the three, entirely settled.

What the three rules are actually disagreeing about

It is worth separating two things the three answers could be disagreeing about, because only one of them is a disagreement.

They could be estimating different quantities. The criterion is estimating the width that maximises a penalised likelihood; the rule of thumb is estimating the width that minimises the mean squared error of a long-run variance, which is a different object with a different optimum; and the error-minimising width is a third. Three rules aimed at three targets are not in conflict and their disagreement says nothing.

Or they could be estimating the same quantity badly. That is the case that would matter, and it is the one the sweep is set up to test: every rule is scored on the same downstream quantity, the error of the coefficients the whitened fit produces. Whatever each rule was derived to optimise, what is compared is what it delivers.

Repairing the penalty, and repairing the fit. What each rule gives up against the best model available from any of them, at ρ = 0.85. The first three fit by least squares and differ only in the penalty; the next two are the same substitutions on Mallows' scale, where there is no linearisation to blame; the last two whiten the sample and keep the ordinary penalty 2q. A penalty computed from the trace is exact about the optimism and selects worse than the parameter count it corrects. Whitening is better than every penalty, and doing it at an estimated ρ rather than a known one costs 0.00544.
Fig. 4 Two places a row count enters a criterion, which is the same distinction one level down: what a rule is derived from and what it is scored on.

Read that way the three are genuinely in conflict, and the conflict is small. Which is the answer, and the reason it takes a paired comparison to see at all.

The rate is not one a lag, and that is why the rules are decisive

“About one unit a lag” is an average, and differencing the five likelihoods says it is an average over a rate that falls by a factor of ten.

From two lags to four the likelihood rises 3.81 units a lag. From four to eight, 0.94. From eight to sixteen, 0.93. From sixteen to twenty-four, 0.36.

So the criterion is not indifferent between widths at all. It is decisive at the narrow end, where each lag buys nearly four times what it costs, and decisive the other way past sixteen, where each lag buys a third of its charge. The indifference is a property of the middle of the range, which is where the average of 0.984 comes from.

Read that way the two criteria’s choices stop being conventions and become first-order conditions. A charge of one should stop where the rate crosses one — somewhere between four and eight lags — and the criterion picks 6.20. A charge of half the log of a hundred and twenty is 2.394, which the rate crosses between two and four lags, and the heavier criterion picks 3.81.

Both rules land exactly where their own charge meets the likelihood’s local rate, which is what a penalised maximum is and is a check on the whole construction rather than a description of it. It also says how to move either answer: the width is a function of the charge alone, through a rate that is measured and published, so a practitioner who wants a wider band has to defend a smaller charge rather than a different width.

What the whole decision is worth

The error curve’s ends put the three rules’ disagreement in proportion.

Across every width measured the error runs from 1.1341 at one lag to 1.1083 at its minimum, so the entire width decision is worth 2.3% of the quantity it is choosing over. The gap between the best and worst of the three rules is 0.00285, which is 0.26%.

So the three rules disagree about a tenth of what the decision is worth, and the decision is worth a fortieth of what whitening at all is worth. A factor of three in the width is a quarter of a per cent, and the reason a paired comparison is needed to see it is that it is a quarter of a per cent rather than that the sweep is short.

The oracle’s 13.29 sits where the rate is still 0.93 — just below the charge of one — which is why the criterion’s 6.20 is short rather than wrong: the two are on the same flat stretch, separated by seven lags that the likelihood values at almost exactly what the criterion charges for them.

And it hardly matters

The error curve is what rescues the situation, and it does so in a way worth being explicit about because it is the opposite of a reassurance.

Read across the widths, the error of the fitted line runs 1.1341, 1.1227, 1.1166, 1.1133, 1.1103, 1.1090, 1.1083, 1.1083, 1.1088, 1.1096 and 1.1109 at one through thirty lags. It falls steeply to about six, is flat between eight and twenty, and rises slowly after. The minimum of the averaged curve is 1.1083 at twelve and sixteen lags together.

So the three rules deliver 1.1104, 1.1141 and 1.1133 against the best available 1.1031. The criterion is short by 0.00734 at 7.9 paired standard errors, the rule of thumb by 0.01019 at 8.5, and the heavier charge is worse than the criterion by 0.00370 at 4.4.

Every one of those gaps is real — the paired standard errors say so, and pairing on the draw is what makes them resolvable at all — and every one of them is under one per cent of the error being paid. A factor of three in the width is worth a hundredth of the quantity it is chosen to minimise.

The whole error, which is what the margins are inside of. The error in a block resample's implied long-run variance under each rule, for both windows, on 400 samples of 120 rows. The best available block length gives 35.8% and no rule reaches it: the plug-in gives 43.1%, a length in the protocol 47.3%, the rule of thumb 61.8%. The gap between the two windows at any one rule is at most 4.54 points and the gap between the rules is 26.01. Which is the reading a practitioner should carry away: the block length is the decision, the window is a detail of it, and the earlier field's ordering is a statement about the detail at a setting the decision has to supply.
Fig. 5 The same shape one field along, on a block length rather than a band width: the rules differ more than the thing they are choosing between does.

Why the oracle is unreachable rather than merely unknown

The width that minimises the error on a given draw has a standard deviation of 10.17 across draws, against a mean of 13.29. That is not an estimation problem anybody can win.

It is the shape the specification-search field measures for a different tuning parameter and reaches the same conclusion about: when the target moves more than the estimate of it can, every feasible rule is short by roughly the same amount, and the differences between feasible rules are second-order.

The mechanism is the flat curve again, read the other way. A curve that is flat between eight and twenty lags is a curve on which the argmin is decided by noise, so the argmin is noisy — and the noise costs almost nothing, because the curve is flat. The unreachability of the oracle and the harmlessness of missing it are the same fact.

What the criterion is charging for

There is a term worth naming here, because it is the one that makes the charge meaningful at all.

A criterion charging a unit a lag is charging for the covariance’s own dimension, and this collection went eleven rounds without one. Every whitened criterion before this field treats the window as given: it is chosen from the residuals by a separate rule, and the candidate’s score carries no charge for it. That is defensible when the same window is used for every candidate, because a constant subtracted from every score changes nothing — a shared nuisance cancels out of every difference.

It stops being defensible the moment the width is being chosen, and it is the same omission that costs a per-candidate sieve its whole comparison two fields along. The determinant and the dimension are the two terms an estimated covariance owes a criterion, and this collection had neither until it started varying one.

The width and what it is a width of

One more thing separates this dial from the ones the rest of the collection turns, and it is easy to miss because the arithmetic looks identical.

A regression’s dimension is a claim about the world: a candidate with four predictors says four things matter. The covariance’s dimension is a claim about an approximation — the true sequence is not zero past any lag under three of the four laws here, so no width is right and every width is an admission. That changes what a charge is for. Charging a regression’s parameters trades bias for variance in a quantity that has a true value; charging a covariance’s trades the same two things in a quantity that has none.

The practical consequence is the flat curve. A dial with no correct setting, feeding a quantity that is an average over the whole sequence, is a dial whose exact position matters less than a dial with a correct setting would — and that is what the sweep reports rather than what anybody assumed.

What a practitioner does

The reading is not use this width, and pretending otherwise would be the mistake this whole field is about.

If the width feeds a coefficient, anything between eight and twenty is within a hundredth of the best available and the rule of thumb’s four is not: it is short by 0.0102 where the criterion’s six is short by 0.0073. Widening the rule of thumb costs nothing and buys a third of the available gap.

If the width feeds a selection, the answer is elsewhere. A criterion comparing candidates through an estimated covariance is comparing them through the same covariance, and what the width does to that comparison is a different sweep with a different shape — where the cost is in the middle rather than at the ends.

And if the width is being chosen per candidate rather than once, this essay is not the relevant measurement at all. That is a selection problem inside a selection problem and it is priced in its own field, where the answer is that the two ends of the argument — a window and an order — cost the same once they are scored by the same criterion.

Fitting them together is worth something under one law. Four fits of the same regression under four dependences: least squares, the two-step plug-in every whitened rule in this collection runs, the coefficients and the band maximised together, and a whitening at the law's own covariance that nobody has. Under the moving average — the one law the band family contains — the joint fit beats the two-step by 0.0077 at 3.3 paired standard errors. Under the autoregression, long memory and the break it is a tie: 0.4, 1.0, 0.3 standard errors. That is the same ordering the likelihood gap gave, arrived at through the coefficients rather than through the objective.
Fig. 6 What each fit delivers by law, which is the quantity the width is chosen to protect.

The width and the law

One column of the sweep is worth reporting separately because it is the only place the four laws disagree about anything here.

Under the moving average the criterion’s width lands close to the law’s own band, which is what a correctly specified family should do. Under long memory nothing lands anywhere useful, because the law’s sequence is still 0.637 at the eighth lag and no width inside the sweep captures it — the family does not contain that law and widening the band walks towards it without arriving.

What the band family does not contain. A band at eight lags holds every autocorrelation past the eighth at exactly zero. Three of these four laws are still correlated there — 0.168 for AR(1) at 0.8, 0.637 for long memory at d = 4/9, 0.663 for a break in the persistence — and cutting a sequence off does not merely approximate it: the truncated sequence stops being a covariance matrix, so there is no point in the family to call the truth. Only the moving average, whose autocorrelations are zero past the fourth lag by construction, is inside. That is what makes it the law the joint fit is worth anything under.
Fig. 7 What each law still has at the eighth lag, which is what a width is trying and failing to reach.

That is the honest limit of a criterion here. A charge per lag prices the cost of a wider band correctly and prices the benefit only within the family; against a law outside it, the benefit keeps arriving and the charge keeps being levied, and the two nearly cancel — which is exactly the 0.984 a lag the sweep reports.

Three sequences, and one of them is not in the family. One sample of 120 rows under long memory at d = 4/9. The law's own autocorrelations are the top line: 0.800, 0.743, 0.711, 0.688 at the first four lags. The residuals' untapered sample sequence is already far below it, because a fit takes memory out and a hundred and twenty rows take more; the tapered plug-in every whitened rule in this collection uses is lower again, at 0.437 for a first lag of 0.800. The sequence that maximises the likelihood over the same band lies between the two — 0.522 — which is the taper's shrinkage being partly undone. The law's own line is drawn rather than offered as a candidate because, cut off at 8 lags, it is no longer a covariance matrix at all, so the family does not contain it.
Fig. 8 The sequence at one width under long memory, where the likelihood has a sharp answer and the width does not.

The one thing the likelihood does settle

None of this says the likelihood is useless about the covariance. It settles the sequence at a given width completely — that is what the maximiser the plug-in is not is about, and the answer there is sharp and law-dependent — and it settles nothing at all about how many lags there should be.

The division is exactly the one every nested-model problem in this collection has. A table of nested models has a likelihood that ranks the fits within a dimension and cannot rank across dimensions; the order a sieve is fitted at has the same shape; and the covariance turns out to be one more instance of it rather than a new kind of difficulty. What is new is only that nobody was treating the covariance as having a dimension, so the charge was missing rather than disputed.

There is one asymmetry worth naming. In a table of regressions the dimension is what a practitioner is trying to learn; here it is a nuisance, and nobody wants to know how many lags the errors have. That makes the flatness of the delivered error a genuinely good outcome rather than a disappointment: a nuisance whose misestimation costs one per cent is a nuisance that has been handled.

What is claimed here, and what is not

This essay takes whether anything in a fit chooses the width of an estimated covariance. The claims are that the band families are nested, so the maximised likelihood is non-decreasing in the width and is so on every draw; that it rises by 0.984 of log-likelihood a lag over widths from two to twenty-four, which is the order of what a criterion charges; that a criterion charging a unit a lag picks a mean width of 6.20, a heavier charge 3.81, the rule of thumb 4 and the error-minimising width 13.29 with a standard deviation of 10.17; and that the three feasible rules deliver 1.1104, 1.1141 and 1.1133 against a best available 1.1031, so a factor of three in the width is worth under one per cent of the error.

What stays out, and is named as a decision: a charge derived rather than assumed. Both charges used here are the standard conventions, and neither is derived for this problem — the covariance’s parameters are not the regression’s, they enter the objective through a determinant as well as a quadratic form, and the effective number of them a taper leaves free is smaller than LL. Deriving the right charge is a real piece of work and it would change the widths in the table; what it would not change is the rate, which is measured and is what makes the choice noisy in the first place.

Also out: a comparison against a cross-validated width. Splitting a dependent series to validate on is a construction with its own assumptions about where to cut, and putting an untested one beside three tested rules would be putting a fourth answer in a table that already has three too many.

The boundary against the window a whitening wants is that it chooses the width from the residuals’ own criterion and this one asks what the fit can say about it. They are different criteria and they do not have to agree, which is the whole subject of the field after next.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A line that beats two curves — both name closed form, degrees of freedom, information criterion, model selection, nested models, overfitting, tapering
  • A list is not a rule — both name bandwidth, closed form, information criterion, model selection, nuisance parameter, regret, whitening
  • A window for every candidate — both name information criterion, model selection, nuisance parameter, overfitting, regret, tapering, whitening
  • Where the generality runs out — both name covariance matrix, information criterion, model selection, nuisance parameter, profile likelihood, regret, whitening
  • A dependence with a shape — both name covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
  • The displacement is a parameter count — both name closed form, degrees of freedom, information criterion, model selection, nested models, overfitting

Named objects

A flat tag is an object no other essay names yet.

BandwidthBias-varianceClosed formCovariance matrixDegrees of freedomInformation criterionModel selectionNested modelsNuisance parameterOverfittingProfile likelihoodRegretTaperingWhitening