A charge that is not a straight line

The surface under every law

A draw-reading charge for a covariance band failed on one law, and the failure was deferred to a law with a steeper error surface. Run under all four laws in the collection, the surface is flat from six lags to thirty under every one, no draw-reading charge tracks anything, and long memory — the law expected to be steepest — is the flattest end to end. What the law moves is the level of the charge, not whether a draw can be read.

Worth reading first: A design is a number · The observations that repeat each other.

The essay that built three charges which read the sample measured them under one law and found them no better than the fixed rule they were calibrated to. It then asked the question a negative result always invites: was the failure a property of the problem, or of the law it happened to be measured on? A first-order autoregression has a short, summable memory, and its error surface across band widths was nearly level — so a rule that read the draw had very little to win, and any rule would have looked as good as any other. Under long memory, it suggested, where the autocorrelations do not sum, the band width might matter more, the surface might be steeper, and a rule that tracked the draw might at last have something to track.

That is a prediction, and it can be checked without building anything new. Four laws for the errors are already in use here: a first-order autoregression, a five-period moving average, long memory and a break in the persistence. The sweep that priced the three charges is a function of the law, so it can be run under all four with every other setting held — the same hundred and twenty rows, the same grid of band widths from one lag to thirty, the same four hundred draws and the same seeds, the same fitted pairs charge, the same three plug-ins. The answer is that the flatness belongs to the problem. It survives every law, and the law the prediction pointed at is the flattest of the four.

Four surfaces, one floor

Start with the thing the deferral was about: the error a band of each fixed width delivers, averaged over the draws, under each law. Each surface is drawn above its own law’s best single width, so that four laws with very different levels of error can be compared on one axis.

Four laws, one flat floor. The error a band of each fixed width delivers, averaged over 400 draws of 120 rows under each law, drawn above that law's own best single width. From one lag to the best width the error falls by 0.02763 under AR(1) at 0.8, 0.02517 under a five-period moving average, 0.01122 under long memory at d = 4/9, 0.03507 under a break in the persistence; from six lags to thirty the whole surface spans only 0.00269, 0.00445, 0.00255, 0.00521. The best single width is 12, 24, 24, 16 lags. The flattest surface end to end is the one under long memory at d = 4/9, which is the law a steeper surface was expected from.
Fig. 1 The error at every fixed band width under each of the four laws, drawn above that law’s best single width. Every surface falls steeply over the first few lags and is close to level from six lags on.

Every surface has the same shape: a steep fall over the first few lags, then a floor. From one lag to the best width the error falls by 0.02763 under the autoregression, 0.02517 under the moving average, 0.01122 under long memory and 0.03507 under the break. From six lags to thirty — the region every charge in this field actually chooses between — the whole surface spans 0.00269, 0.00445, 0.00255 and 0.00521. The fall is spent before six lags under all four, and what is left is a floor a few thousandths of a unit deep.

The best single width moves with the law, and that is the first thing the law does change: 12 lags under the autoregression, 24 under both the moving average and long memory, 16 under the break. A law whose autocorrelations persist wants a wider band on average, which is what anybody would expect. What it does not do is make the floor any steeper.

And long memory is the flattest of the four end to end. Its full fall, from one lag to the best, is less than half the autoregression’s and under a third of the break’s. The law the deferral expected to be steepest has the surface on which the choice of width matters least. One reading of why is available in the essay that counted a lag in the pairs it actually has: under long memory the autocorrelations that a wide band would need to estimate are small at every lag and badly estimated on a hundred and twenty rows, so widening the band adds noise at about the rate it adds signal. That is offered as a reading, not measured here; what is measured is the surface, and it is flat.

The same split, four times

The single-law essay turned its negative result into a measurement by splitting the gap the fixed charge leaves. The gap is the distance from the error the fixed pairs charge delivers to the error the best width on each draw would have delivered — an infeasible target, since nobody knows the best width on their own draw. It has two parts. One is the distance from the fixed charge to the best single width, which a better fixed rule could take. The other is the distance from the best single width to the best width on the draw, which only a rule that read the draw could take.

The same split under four laws. The error the fixed pairs charge delivers above the best band width on the same draw, cut in two, over 400 draws of 120 rows under each law. The first part is the distance from the fixed charge to the best single width — what a better fixed rule could take — and the second is the distance from the best single width to the best width on the draw. The first runs 0.00107, 0.00000, 0.00119, 0.00105 for AR(1) at 0.8, a five-period moving average, long memory at d = 4/9, a break in the persistence; the second runs 0.00493, 0.00327, 0.00704, 0.00721. The share a fixed rule could reach is 17.8%, 0.0%, 14.5%, 12.7%, and under the moving average the fixed charge already sits on the best single width.
Fig. 2 The gap the fixed pairs charge leaves, under each law, cut into the part a better fixed rule could reach and the part only the draw could.

The first part is 0.00107 under the autoregression, 0.00119 under long memory and 0.00105 under the break. Under the moving average it is nothing at all: the fixed charge picks bands of 27.73 lags on average and already delivers the best single width’s error, to within two hundred-thousandths. The second part is 0.00493, 0.00327, 0.00704 and 0.00721. So the share of the gap a better fixed rule could ever reach is 17.8%, 0.0%, 14.5% and 12.7%, and the rest belongs to the draw.

On the single law this was the result that mattered most. If five sixths of the gap is the distance from a fixed width to the per-draw optimum, then a rule that reads the draw is the only kind that could close most of it, and the question is whether the draw carries the information. Under every law that share is larger — between 85% and the whole of it — so on this split alone the case for reading the draw is stronger under the other three laws than under the one it was first measured on.

The split says nothing about whether the per-draw part can be read. That needs the tracking correlations.

Twelve correlations with nothing

Each of the three draw-reading charges is a rule for picking a width, and the question is whether the width it picks moves with the best width on the same draw. The three are the ones built for the first law: the pairs line scaled by the draw’s own summed squared autocorrelation, the pairs line reweighted lag by lag by how much this draw’s autocovariance at each lag moves, and the optimism estimated directly from a split of the sample. Each is recalibrated to the law it is run under, so that on the average draw it charges what the fixed rule charges; what it adds is the draw-to-draw movement.

Twelve correlations with nothing. The correlation between the band width each draw-reading charge picks and the best band width on the same draw, for each of the three charges under each of the four laws, over 400 draws apiece. The twelve readings run from −0.080 to 0.030, and the largest in absolute value is under a five-period moving average. Long memory, which was the law expected to give a draw-reading rule something to find, reads −0.049, −0.067, −0.062. Reading the draw fails under every law in the collection, not under one.
Fig. 3 The correlation between the width each draw-reading charge picks and the best width on the same draw, for three charges under four laws.

The twelve correlations run from −0.080 to 0.030. Long memory reads −0.049, −0.067 and −0.062, every one of them on the wrong side of zero. Nothing here tracks anything.

That is worth setting beside the one place in the collection where a plug-in of this kind does find something. The essay that chose a block length from the data builds rules that read a sample’s persistence and finds their correlations with the best block length small but real, because a more persistent sample genuinely needs a longer block to carry its dependence. A band width is not like that. Its best value is set by a trade-off between two small, noisy errors, and a sample that looks more persistent does not, on any of these four laws, want a band of a different width.

The statistic the first charge reads is not inert. Its coefficient of variation across draws is 0.312 under the autoregression, 0.225 under the moving average, 0.467 under long memory and 0.419 under the break — so long memory is the law where the draw’s own persistence moves most, by half again as much as on the law the failure was first found on. A rule reading it moves with it. What the rule moves with carries no information about the width the draw wanted, under the law where it moves most just as under the law where it moves least.

The cost side repeats as well. Against the fixed charge each was calibrated to match, the split-sample rule is worse by 0.00199 under long memory, at 3.67 paired standard errors — further behind than on the autoregression, where it was 2.58. The persistence rule is worse by 0.00020 at 0.54 standard errors. Under the break the three are worse by between 0.25 and 0.43 standard errors, and under the moving average the first two are better by one and two hundred-thousandths, which is the width of the line the figure draws them with.

Whether a draw wants the wide band

The tracking correlations say the three rules fail. The surface’s shape, read one draw at a time, says why any such rule would.

A rule that reads the draw can win only where the draw itself has a definite preference — where the error surface on this sample is steep enough that its slope has a predictable sign. The averaged surface hides that, because an average of noisy curves can be smooth while every individual curve is ragged. So read the same surface draw by draw, at the two widths that bracket the floor: six lags and thirty.

Whether a draw wants the wide band is nearly a coin. Two readings of how steep the error surface is from one draw's point of view, over 400 draws under each law. The first is the share of draws on which a band of thirty lags delivers less error than a band of six: 47.3% for AR(1) at 0.8, 68.5% for a five-period moving average, 52.3% for long memory at d = 4/9, 52.8% for a break in the persistence. The second is the share of draws whose best width is the narrowest or the widest on the grid: 25.5%, 39.0%, 44.0%, 32.3%. The standard deviation across draws of the difference between the errors at those two widths is 20.67, 2.39, 8.20, 6.88 times its mean, so the sign of the slope on one draw is mostly noise under every law, and least so under the moving average, where the fixed charge already picks the wide band.
Fig. 4 Under each law, the share of draws on which thirty lags beat six, and the share whose best width is the narrowest or the widest on the grid.

Thirty lags beat six on 47.3% of draws under the autoregression, 68.5% under the moving average, 52.3% under long memory and 52.8% under the break. Three of the four are a coin. The fourth has a preference — the moving average puts its weight at lags one to four and leaves the long band nothing to waste — and it is exactly the law where the fixed charge already sits on the best single width and the reachable share is zero. Where a draw’s preference is predictable, the average already knows it.

The noise in that slope can be stated directly. The standard deviation across draws of the difference between the errors at thirty lags and at six is 20.67 times its mean under the autoregression, 2.39 times under the moving average, 8.20 times under long memory and 6.88 times under the break. A per-draw rule would have to read the sign of a quantity whose standard deviation is two to twenty times its mean from a statistic that is itself noisy, and the correlations above are what that looks like.

The second reading in the figure is the share of draws whose best width is an end of the grid — one lag or thirty. It is 25.5%, 39.0%, 44.0% and 32.3%. On the single law a quarter of draws landing on the boundary was the sign that the best width on the draw was an argmin of a nearly flat surface rather than a property of the sample. Under long memory it is closer to half. The best width on the draw has a standard deviation of 11.44 lags there, against 10.19 on the autoregression, on a grid of thirty. The per-draw optimum is less a target under long memory than anywhere else.

What the law does move

So the law does not make the draw readable. It does change one thing, and it is worth isolating because it is reachable by a fixed rule and the fixed rule here is not reaching it.

The pairs charge is a line through the origin in the pairs the band uses, and its rate is fitted by describing the optimism profile under the law, without reference to error at all. The essay that priced four derived charges showed that under the autoregression the level of that line decides the width completely and is worth almost nothing in error. Run the same level grid — the fitted rate multiplied by a quarter, a half, one, two and four — under each law.

The level the law wants. The error the pairs charge delivers above each law's best single width when the law's own fitted rate is multiplied by a quarter through four, over 400 draws under each law. The fitted rate is 0.8133 for AR(1) at 0.8, 0.8185 for a five-period moving average, 0.9382 for long memory at d = 4/9, 1.0782 for a break in the persistence. It is the best level on the grid under AR(1) at 0.8 and a five-period moving average; under long memory at d = 4/9 and a break in the persistence half of it is better. The rings mark the autoregression's rate applied to each law, which is what a rule that did not know the law would charge, at 0.99×, 0.87×, 0.75× of the law's own.
Fig. 5 The error the pairs charge delivers above each law’s best single width as the law’s own fitted rate is multiplied by a quarter through four. The rings are the autoregression’s rate applied to each of the other laws.

The fitted rate is 0.8133 under the autoregression, 0.8185 under the moving average, 0.9382 under long memory and 1.0782 under the break. Under the first two it is the best level on the grid. Under long memory and the break, half of it is better: at half the rate the long-memory charge picks bands of 26.53 lags rather than 9.86 and delivers 1.63648 rather than 1.63768, slightly below even the best single width’s 1.63650; the break’s picks 27.57 lags rather than 13.06 and delivers 1.17990 rather than 1.18018.

The rate fitted to the optimism profile and the rate that delivers the least error are therefore the same thing on two laws and different things on two. That is a fact about the gap between what a criterion describes and what a forecast pays for, and it is the same gap the essay on what the correction assumes found from the side of the window rather than the law: a quantity calibrated to describe one thing is right about another only where the two happen to agree.

The rings in the figure make the practical point. A rule that did not know the law would charge the autoregression’s rate, which is 0.99, 0.87 and 0.75 of the other three laws’ own. Under the moving average that is the same charge to within a per cent. Under long memory and the break it is cheaper than the law’s own fitted rate, and so it is better: the long-memory error falls by 0.00048 and the break’s by 0.00066, with bands of 12.45 and 19.39 lags. A charge calibrated to the wrong law beats one calibrated to the right law under both of the laws where the right law’s rate is too dear.

None of this is large. Every gain in this section is under a fifth of the per-draw part of the gap under the same law, and most are under a tenth. But it is the one part of the problem the law moves, and it is a part a fixed rule can reach.

What the four laws agree on

Put the four together and the deferral is answered on every count.

The surface is flat from six lags to thirty under every law, within a few thousandths of a unit, and the law the deferral expected to be steepest has the flattest surface of the four, end to end.

The part of the gap a fixed rule could reach is a sixth or less under every law, so the case for reading the draw looks as strong under the other three laws as under the first — and it fails under all four, with twelve tracking correlations none of which reaches a tenth.

What a draw prefers is close to a coin under three laws, and on the fourth, where it is not, the fixed charge already gives the draw what it wants.

What the law moves is the level of the charge, and there the fitted rate is right under two laws and too dear under the other two, where half of it does better. A charge calibrated on the autoregression, applied blind, does better under long memory and the break than each law’s own.

That last reading has a consequence the sequence of charges in this field has not had to face before. Every essay since the charge nobody derived has calibrated the charge to the profile of optimism, on the reasoning that optimism is what a criterion is supposed to subtract. Under two of the four laws the optimism-fitted level is beaten by a cheaper one — which says the criterion that best describes optimism is not the criterion that best chooses a band, under those laws, at this sample size. On the first law the two coincide; and that coincidence was what made the fitted charge look like the end of the problem.

What this does not show

The comparison is honest only within its frame, and the frame is narrow in three ways worth naming.

Every reading is at a hundred and twenty rows. A band of thirty lags spends a quarter of the sample on its widest autocovariance, and the flatness of the floor may be as much about that ratio as about the problem. Nothing above says how the surface changes as the sample grows.

The grid stops at thirty lags. Under the moving average and long memory the best single width is twenty-four, and a quarter to two fifths of draws put their best width on an end of the grid. A wider grid would let some of those draws choose further out, and it would also change what the fitted charge sees.

Four laws are not every law. They are the four used throughout for the errors, chosen to differ in how their autocorrelations decay. A law with a sharp seasonal spike at one lag would give a band a single lag it must reach, and that is the kind of surface on which a per-draw rule might have a definite target. None of the four has one.

Still open: the surface as the sample grows

The next measurement is the first of those three. The prediction that a steeper surface would let a draw-reading charge earn its keep was wrong about the law; it may be right about the sample. At a hundred and twenty rows the floor from six lags to thirty is a few thousandths of a unit deep and the per-draw slope is noise many times its mean. At four hundred and eighty rows each autocovariance is estimated on four times the pairs, the cost of a wide band falls, and the floor could either steepen — giving a rule that reads the draw a definite slope to read — or flatten further, if the gain from each extra lag shrinks faster than its cost. The same sweep at two or three sample sizes, under the two laws whose fitted rate is too dear, would say which, and would say whether the level the law wants converges on the level the optimism profile fits as the sample grows.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A width that moves and an error that does not — both name bandwidth selection, estimation error, information criterion, mean squared error, out of sample, regret, tapering
  • How often it matters — both name bandwidth selection, estimation error, information criterion, long memory, monte carlo, regret, selection effect
  • The width a band is measured in — both name bandwidth selection, dependence, estimation error, information criterion, monte carlo, plug in estimate, tapering
  • A line that beats two curves — both name bandwidth selection, dependence, estimation error, information criterion, monte carlo, tapering
  • A rate times a size — both name bandwidth selection, estimation error, information criterion, monte carlo, regret, selection effect
  • A step that is not a ratio — both name bandwidth selection, dependence, information criterion, monte carlo, regret, selection effect

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionDependenceEstimation errorInformation criterionLong memoryMean squared errorMonte CarloOut of samplePlug in estimateRegretSelection effectStructural breakTaperingWorst case