A charge for a covariance's own dimension

A width that moves and an error that does not

Four charges give four widths a factor of four apart and four errors half a per cent apart. The derived charge wins, significantly, by a quarter of what was on offer — and none of the four is an estimate of anything.

Worth reading first: A design is a number · The observations that repeat each other.

Two essays have now established that a criterion charging one log-likelihood unit a lag for a covariance band is charging between two and three times too much. The obvious next sentence writes itself: correct the charge and the width improves and the error falls. This essay measures how much, and the answer is 0.15 per cent.

That is not a disappointing result to be buried at the end. It is the reason the width was ever an open question, and it arrives with a second measurement that is sharper than the first: no rule’s width is an estimate of the width that was actually best. Not approximately, not with a bias — the correlation between the width a charge picks on a draw and the width that would have been best on that same draw is 0.06, and for one of the four charges it is 0.006.

Four charges, four widths

Every rule reads the same objective on the same draws, differing only in what it subtracts.

Schwarz’s charges half a log n a lag, which at a hundred and twenty rows is 2.394 apiece. Akaike’s charges one. The weights charge what the window leaves free, which for a Bartlett window is exactly half the width. The measured charge is 0.767 of that, which is the share the optimism actually comes to. And a fifth answer is in the table for scale: the rule of thumb every applied long-run variance uses, 4(n/100)^(2/9), which at this sample size is four and does not read the data at all.

Cut the charge and the width follows it. The band width each charge picks, averaged over 400 draws of 120 rows under AR(1) at 0.8, with the standard deviation across draws beside it. Schwarz's charge — half a log n a lag, which is 2.39 here — picks 3.67. Akaike's picks 6.02. Charging the numbers the window actually leaves free, which is half the width, picks 10.12; charging what the optimism measures, 0.767 of that, picks 14.15. A charge and the width it buys are very nearly reciprocal, which is what a likelihood rising at a fixed rate a lag implies and is why the four answers span a factor of 3.86. The width that was actually best on the draw averages 13.90 and moves by 10.30 from draw to draw — three times as much as any rule's answer does.
Fig. 1 The band width each charge picks, with the spread across draws beside it. The infeasible row is the width that would actually have been best on the draw.

Over four hundred draws of a hundred and twenty rows under a first-order autoregression, the widths are 3.67, 6.02, 10.12 and 14.15. A factor of 3.86 from end to end, produced by nothing but the coefficient in front of L.

The near-reciprocality is not a coincidence and it is the mechanism. The field that measured the rise found the maximised likelihood climbing at about a unit a lag whatever is in the data, because the families are nested and a wider band cannot do worse. A penalised objective is therefore a rise of about one a lag minus a charge of c a lag, and its argmax sits wherever the small curvature left over happens to turn — which moves a long way for a small change in c. Halve the charge and the width roughly doubles.

And four errors that are the same error

The widths are only interesting through what they deliver, and what they deliver is the coefficient error against a fresh replicate — the quantity the whole line of fields is ultimately about, rather than the likelihood the charge is levied on.

Every width in the table moves and almost no error does. The error each charge's chosen width delivers, on 400 draws of 120 rows under AR(1) at 0.8, with the charge that chose it named under each point. The widths run from 3.67 to 14.15 — a factor of 3.86 — and the errors from 1.1129 to 1.1073, which is 0.51 per cent of the error. The derived charge does beat the convention, by 0.0017 at 6.1 paired standard errors, and that margin is 24% of the 0.0070 the convention gives up against the width that was actually best. So the correction is real, significant and negligible — which is the answer to what deriving the charge properly buys, and it is not the answer the derivation invites.
Fig. 2 What each charge’s width costs, with the width it picked printed beneath. The vertical scale spans half a per cent of the error.

The errors are 1.11290, 1.10896, 1.10728 and 1.10739, against 1.11177 for the rule of thumb and 1.10192 for the best width available on the draw. From worst feasible to best feasible is 0.00562, which is 0.51 per cent of the error. The factor of 3.86 in the width buys half a per cent in the thing the width is for.

Inside that half a per cent, the ordering is real and the margins are paired.

The derived charge beats the convention. The weights charge lowers the error by 0.00169 against Akaike’s, at 6.13 paired standard errors on four hundred draws. The measured charge lowers it by 0.00158 at 4.38. Neither is noise, and neither is a rounding error in the third decimal that a different seed would reverse.

And it closes a quarter of what was there to close. Akaike’s charge gives up 0.00704 against the best width available, at 17.76 standard errors. The derived charge recovers 0.00169 of that — 24% — and leaves 0.00536, still at 15.88 standard errors from the best available. So the correction is real, significant and small, and the honest summary of it is that most of what a width rule gives up is not recoverable by getting the charge right.

Schwarz’s is the outlier in the other direction, at 0.00393 worse than Akaike’s at 8.88 standard errors. It is not a charge derived for this problem either — it is an approximation to a marginal likelihood rather than an optimism — and at this sample size it charges 2.394 where the truth is 0.374, which is a factor of six.

Halving the charge does not double the width, and the elasticity is not constant

“Halve the charge and the width roughly doubles” is the right mechanism and the arithmetic is worth being exact about, because the departure from it says where the objective’s curvature is.

Going from Akaike’s charge of 1 to the weights’ charge of 0.5 takes the width from 6.02 to 10.12 — a factor of 1.68, not 2. Reading each adjacent pair as a power law gives elasticities of

  • −0.57 between Schwarz’s 2.394 and Akaike’s 1,
  • −0.75 between Akaike’s 1 and the weights’ 0.5,
  • −1.26 between the weights’ 0.5 and the measured 0.3835.

So the width is accelerating in the charge. Near the heavy end a halving buys about half again; near the light end it buys more than a doubling.

That is the shape a flattening objective has to produce. The penalised objective is a rise of about a unit a lag minus a charge, and the residual curvature that decides where it turns is smallest where the band is widest — so the same proportional change in the charge moves the argmax further the further out it already is.

At the operating point the elasticity is −1.26, which prices the whole question this field exists to answer: a ten per cent error in the charge is a thirteen per cent error in the width. That is what makes the charge look like something worth getting right, and it is why the next section’s finding — that thirteen per cent of the width is worth 0.15 per cent of the error — is the surprising half.

Two numbers that ought to be far apart and are not

One row of the table is not a charge at all, and it lands closer to a charge than the charges land to each other.

The rule of thumb picks 4 at this sample size, reading nothing. Schwarz’s charge picks 3.67, reading the whole sample. They are 9% apart.

Against that, the four data-reading rules span 3.67 to 14.15 — a factor of 3.86. Any two of them picked at random are further apart than the heaviest of them is from a rule that never opened the file.

And the correlations say why that is not the embarrassment it sounds like. A charge’s width correlates with the width that would actually have been best at 0.06, and for one of the four at 0.006 — which is 0.36% and 0.004% of the best width’s variance explained. A rule explaining a third of one per cent of the variation in its own target is not, in any useful sense, aiming at it, and a rule explaining none of it is not doing worse.

Why almost nothing moves

The error curve against the width says the rest.

Across widths of 1, 2, 3, 4, 6, 8, 12, 16, 20, 24 and 30 the error runs 1.1323, 1.1205, 1.1149, 1.1118, 1.1088, 1.1077, 1.1073, 1.1073, 1.1075, 1.1080, 1.1088. It falls steeply out to about six, is flat to four decimal places from eight to twenty-four, and rises very slightly after. Every charge in the table picks a width inside or just short of that plateau, and a plateau is a place where being wrong costs nothing.

That is the whole explanation of the half per cent, and it is worth separating from the flatness of the objective, which is a different flatness with the same consequence.

The plateau an argmax is chosen out of. The average penalised objective against the band's width, each curve shifted so its own best is zero, on 400 samples of 120 rows. The band across the top is one log-likelihood unit — the depth inside which a criterion has no view. Under the convention's charge the objective is within a unit of its best at widths 6 and 8. Under the charge the weights derive it is within a unit at 6, 8, 12, 16, and under the measured charge at 8, 12, 16, 20, 24 — five widths spanning a factor of 3. A smaller charge does not only move the argmax; it flattens what the argmax is chosen out of. And the flatness is not the whole difficulty: one draw's version of this curve has a standard deviation of 8.17 units at the argmax, against a plateau one unit deep, so the width a criterion returns on a given sample is decided by the sample rather than by the shape.
Fig. 3 The average penalised objective, each curve shifted so its own best is zero. The shaded band is one log-likelihood unit — the depth inside which a criterion has no view at all.

Under Akaike’s charge the objective is within a unit of its best at widths 6 and 8. Under the weights charge it is within a unit at 6, 8, 12 and 16. Under the measured charge it is within a unit at 8, 12, 16, 20 and 24 — five widths spanning a factor of three. A smaller charge does not only move the argmax; it flattens what the argmax is chosen out of, because the charge is what supplies the downward slope that makes the maximum a maximum at all.

And the flatness is not the largest difficulty. One draw’s version of that curve has a standard deviation of 8.17 log-likelihood units at the argmax, against a plateau one unit deep. The average curve has a shape; a single sample’s curve is that shape plus eight units of noise. Which width comes back on a given sample is decided by the sample.

The rule of thumb, which is not in the argument and is in the table

One row of the table is not a charge at all and it is the row most applied work actually uses.

The automatic window, 4(n/100)^(2/9), is four lags here and four lags at two hundred rows and five at four hundred. It contains no likelihood, no penalty and no data — it is a function of the sample size, derived from an asymptotic mean-squared-error argument about estimating a long-run variance rather than about selecting anything. It picks 4.00 on every draw, with a standard deviation of exactly zero.

It delivers 1.11177, which is worse than Akaike’s by 0.00281 and better than Schwarz’s by 0.00113. On a table where the four principled answers span 0.00562, a rule with no principle in it lands in the middle of them.

That is not an argument for using it, and it is a sharp illustration of what the flatness means. The gap between the most and the least defensible thing anybody does here is smaller than the gap between any of them and the width that was best on the draw, and the ordering among them is stable but slight. A practitioner who read only this table would be right to conclude that the choice does not matter much, and wrong to conclude that the question does not matter — because the two conclusions are about different things, and the second is what the essay that measured the charge is for.

What a charge that could be trusted would have to be

The measured charge in this table is not a rule anybody can run, and it is worth saying exactly why, because the gap is instructive rather than technical.

It is 0.767 of the summed weights, and the 0.767 came from a Monte Carlo over draws from a known law. A practitioner has one sample and no law. So a usable version would have to estimate the share from the sample — and the share depends on the persistence, running from 0.738 under a first-order autoregression to 0.944 under a break, which is a spread of a quarter. Estimating it means estimating the dependence, which is the object whose dimension is being charged for.

That circle is not obviously vicious. The band’s own Ω̂ is available at every candidate width, and a parametric bootstrap from it would give a per-sample optimism directly: simulate from Ω̂, re-estimate, difference. It would cost a resample per width per draw, which is affordable, and it would be a genuine plug-in charge rather than a constant.

Whether it would be better is a separate question and the table above suggests the answer is no by a predictable margin. The whole span of feasible charges is 0.00562 of error and the best available width is 0.00536 further on; a plug-in charge could at most move within the first of those, and the noise it would add — a resampled optimism at four hundred draws has its own standard error — would come out of the same budget. The measurement that would settle it is not made here, and naming it as the next thing rather than an oversight is the honest form.

None of them is an estimate

Which is the second measurement, and it is the one that changes how the table should be read.

A table of four charges printed beside a column headed the best width on this draw invites a particular reading: these are four attempts at the same target, and the one closest to it is the best attempt. The measured charge picks 14.15 and the best width averages 13.90, which under that reading looks like a rule that has almost got there.

None of the rules is an estimate of the width it is read against. The correlation, across 400 draws of 120 rows, between the width each charge picks on a draw and the width that would actually have been best on that same draw. Every one is under 0.070: 0.037 for one unit a lag, 0.006 for half a log n a lag, 0.063 for the numbers the window leaves free, 0.065 for what the optimism actually is. The best width averages 13.90 with a standard deviation of 10.30 across draws, so there is a great deal to track and none of it is being tracked. A table of four charges beside a column headed the best width on this draw invites the reading that the rules are estimates of it and one of them is close. They are not estimates of it at all; they are four nearly deterministic functions of the sample size sitting in the neighbourhood of a very noisy target.
Fig. 4 The correlation, across draws, between the width each rule picks and the width that was best on that same draw.

The correlations are 0.037, 0.006, 0.063 and 0.065. Nothing is being tracked. The best width has a standard deviation of 10.30 across draws — it is 3 on some samples and past 30 on others — and every rule’s answer is very nearly independent of which sample it is looking at. The measured charge’s own spread is 5.52, about half the target’s, and none of that variation is aimed anywhere.

So 14.15 against 13.90 is not a near miss. It is a rule whose average happens to land near the average of a quantity it has no information about, in the same way that a rule of thumb lands near an average block length while being a constant. The right description of all four is that they are nearly deterministic functions of the sample size, sitting in the neighbourhood of a very noisy target, and the neighbourhood is wide enough that sitting anywhere in it delivers the same error.

The same shape, one field over

It is worth putting this beside the other tuning parameter this collection has priced, because the two readings are the same and were arrived at independently.

The block-length field compares two block windows and finds the ordering between them reversing depending on who chose the block length — and then measures the two quantities on one scale. What the window choice buys at the best available length is 2.12 points of error; what the best feasible rule gives up against that same length is 7.26, a factor of 3.4. The argument is a third of the size of the thing it is inside.

Here the arithmetic runs the same way with different numbers. What the charge choice buys is 0.00169; what the best feasible charge gives up against the best available width is 0.00536, a factor of 3.2. Two tuning parameters, two fields, two instruments sharing no arithmetic, and the same ratio to within a twentieth.

That is either a coincidence or a statement about what a plateau does to any rule that has to find an argmax on it, and there is a reason to think the second. A flat objective produces a rule whose answer is decided by noise and an error surface where being wrong is cheap; the ratio between what a rule recovers and what it leaves is then a ratio between two curvatures, and both curvatures are small for the same reason. Testing that properly means finding a tuning parameter whose error surface is not flat and checking whether the ratio moves. Nothing in this collection has one.

What the field concludes

Three sentences, and the order matters.

The convention is wrong. A band of lags is charged as though its lags were free parameters, and they are shrunk parameters; the measured optimism is 0.374 of a unit a lag against a charge of one, under every window and every law tested. That is a defect in a rule that has been in use unexamined, and it is worth knowing whether or not it costs anything.

Fixing it buys a fifth of what is available and half a per cent of the error. Both halves of that are true and quoting either alone misleads. The derived charge does beat the convention at six paired standard errors; the whole table of charges spans half a per cent; and what separates the best feasible rule from the best available width is a further half per cent that no charge reaches.

And the reason it buys so little is the reason the question was open. The penalised objective is flat, so its argmax is decided by noise; the error surface is flat, so an argmax decided by noise costs nothing. A criterion that is very nearly indifferent produces a noisy answer and a harmless one at the same time, and those two facts have the same cause. The right derivation does not repair a flat objective — it explains why the objective was flat and why nobody had noticed the charge was wrong.

Which is also why the finding is worth carrying rather than filing. There is nothing here that says a covariance band’s charge is harmless in general. It is harmless where the error surface is flat over the range the charge moves the width across, and that range is a factor of four. A setting where the error rises steeply past twenty lags would have the same three-fold error in the charge and a real cost attached to it, and nothing in the existing convention would have flagged it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A charge that reads the draw — both name bandwidth selection, estimation error, information criterion, mean squared error, model selection, optimism, out of sample, regret, tapering
  • A window for every candidate — both name bandwidth selection, information criterion, model selection, optimism, overfitting, regret, tapering, tuning parameter
  • A penalty is a trace — both name degrees of freedom, information criterion, mean squared error, model selection, optimism, out of sample, overfitting
  • A step that is not a ratio — both name bandwidth selection, information criterion, model selection, nested models, overfitting, regret, tuning parameter
  • A table and a list — both name bandwidth selection, information criterion, model selection, nested models, overfitting, regret, tuning parameter
  • The comparison that was not made — both name bandwidth, covariance matrix, degrees of freedom, information criterion, model selection, regret, tapering

Named objects

A flat tag is an object no other essay names yet.

BandwidthBandwidth selectionCovariance matrixDegrees of freedomEstimation errorInformation criterionMean squared errorModel selectionNested modelsOptimismOut of sampleOverfittingRegretTaperingTuning parameter