A charge that is not a straight line

A charge that reads the draw

Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.

Worth reading first: A design is a number · The observations that repeat each other.

The essay that priced four derived charges ends on an error column that does not move, and names the reason rather than the remedy. Every charge on that table is a deterministic function of the band’s width, the window and the sample size. A deterministic function of the width cannot correlate with anything that varies from draw to draw at a fixed width, so no rule there is an estimate of the best width on its own sample, and the correlations are all within a few hundredths of nothing.

The remedy it names is a charge that reads the sample. This essay builds three, and measures what they buy.

What a charge would have to read

The optimism a plug-in covariance costs is the gap between the likelihood at the fitted Ω^\hat\Omega and the same objective on a world drawn afresh. What makes it large is how much γ^\hat\gamma moves, and there is a standard expression for that. Bartlett’s formula gives, for a lag kk,

nVarγ^(k)γ(0)2    j[r(j)2+r(j+k)r(jk)]\frac{n\,\mathrm{Var}\,\hat\gamma(k)}{\gamma(0)^2} \;\approx\; \sum_j \big[\, r(j)^2 + r(j+k)\,r(j-k) \,\big]

and every autocorrelation in it is estimable from the residuals of the model already fitted. That is the opening: the band’s charge ought to depend on how noisy this sample’s autocovariances are, and this sample can be asked.

Three rules follow, in order of how much of the sample each one uses.

The pairs line at the draw’s own persistence. Take S^=jr^(j)2\hat S = \sum_j \hat r(j)^2, which is one for white noise and grows with dependence, and charge a(S^/Sˉ)P(L)a\,(\hat S/\bar S)\,P(L). It keeps the denominator the field derived and moves only the level.

A charge at the draw’s own variance, lag by lag. Keep the shape open too: charge akLw(k)(1k/n)V^(k)/Vˉ(k)a \sum_{k \le L} w(k)(1 - k/n)\,\hat V(k)/\bar V(k), so a lag whose autocovariance moves more on this draw than it usually does is charged more. Averaged over draws the two denominators coincide, which is what makes this the pairs count with the lags reweighted rather than a different width.

The draw’s own split-sample optimism. The feasible version of the measurement the whole field rests on. Nobody outside a simulation has a second world, but anybody has two halves: fit the band on the first sixty rows, evaluate the concentrated likelihood on the half it was fitted to and on the half it was not, and difference. It estimates the thing directly, at half the sample size, on one draw, with no averaging at all.

Each is scaled so that on the average draw it charges exactly what the fixed pairs line charges. That is deliberate and it is generous: a practitioner does not have the average draw and would have to guess the constant, so every plug-in here is handed information the fixed rules were not given. What is left to compare is nothing but reading the draw.

What the three rules are handed, and what a reader would not have

The comparison is arranged to be as favourable to the plug-ins as it can honestly be made, and the arrangement is worth setting out because it decides what a negative result means.

Each plug-in is normalised by the sweep average of its own statistic — the mean persistence over four hundred draws, the mean lag-by-lag variance at each lag, the mean split-sample gap. Those three constants are what make the plug-in charge the same as the fixed charge on the average draw, and they are exactly the quantities nobody outside a simulation has. A practitioner running the persistence rule would have to guess Sˉ\bar S, and guessing it wrong shifts every charge by the same factor, which the level grid below prices.

The constants are also computed on a separate pass over the same seeds rather than taken from the scoring pass. That is a small thing and it is not a formality: a normaliser computed from the draws it normalises is fitted on its own sample, and the difference between a constant that is fitted and one that is not is what the field’s own four essays have been about.

So the plug-ins are given a constant each, at no charge, and the fixed rules are given nothing. Whatever the table shows is not a comparison the plug-ins were arranged to lose.

The statistics move

Before asking whether the charges help it is worth checking that they vary at all, because a plug-in whose statistic is constant is the fixed rule with extra arithmetic.

They vary. The draw’s own summed squared autocorrelation averages 4.3627 with a standard deviation of 1.3613 — a coefficient of variation of 0.3120 — and runs from 2.064 to 10.188 across four hundred draws. The lag-by-lag variance has a shape as well as a level: averaged over draws it falls from 8.725 at lag zero to 4.276 by the sixth lag and sits at 4.381 at the thirtieth, so the reweighting is a real reweighting rather than a rescaling. The split-sample estimate of what a thirty-lag band costs over the narrowest one averages 9.0153 with a standard deviation across draws of 7.0324.

So the three charges move, and the third moves a great deal.

And the answers do not

The correlation between the width a rule picks and the best width on the same draw is the quantity the deferral was about. Over four hundred draws the three fixed charges read −0.0692, −0.0661 and 0.0121. The three that read the sample read −0.0120, −0.0186 and −0.0411.

None of the six is distinguishable from nothing, and reading the draw did not move the reading in the direction it was supposed to.

Reading the draw changes what is charged, not what is tracked. The correlation between the band width each rule picks and the best band width on the same draw, over 400 draws. The three fixed charges read -0.069, -0.066, 0.012. The three that read the sample read -0.012, -0.019, -0.041. None of the six is distinguishable from nothing. The statistic the first plug-in reads does vary — the draw's own summed squared autocorrelation runs from 2.06 to 10.19 with a mean of 4.36 — so the failure is not that the charge stopped moving. It is that what it moves with carries no information about which width this draw wanted.
Fig. 1 The correlation between the width each rule picks and the best width on the same draw. The three that read the sample are no nearer to tracking it than the three that do not.

The reason is one line further back, and it is worth having as its own measurement rather than as an inference. The correlation between the statistic itself — the draw’s own persistence — and the best width on that draw is −0.0590.

A charge built on a statistic that does not know the answer cannot know the answer. Every step after that is machinery faithfully propagating no information, which is why all three plug-ins fail in the same way and to the same degree despite being three quite different constructions.

What it costs, which is not nothing

A charge that reads the data is a selection, and a selection has a variance the fixed rule does not. The right way to price it is to look at what moved rather than at what was gained, because on this table nothing was gained.

The fixed pairs line picks widths with a standard deviation of 6.55 lags — not zero, because the likelihood it is subtracted from varies even though the charge does not. The three plug-ins pick widths with standard deviations of 7.88, 7.89 and 10.07. The split-sample rule, which uses the most of the sample and averages the least, is half again as variable as the rule it replaces.

What reading the draw costs. How far each rule's chosen band width moves from draw to draw, over 400 draws. A charge that is a function of the width alone still picks a varying width, because the likelihood it is subtracted from varies — the pairs line's widths carry a standard deviation of 6.55 lags. Every charge that reads the sample carries more: 7.88, 7.89, 10.07 for the three here, the largest being the one that estimates the optimism from a split of the sample. That extra movement is the cost side of a data-dependent rule and it is paid on every draw, whether or not the rule is reading anything. The best width on the draw moves by 10.19, which is more than any rule and is what a rule would have to follow.
Fig. 2 How far each rule’s chosen width moves from draw to draw. Every charge that reads the sample is more variable than the fixed charge it is calibrated to match.

That extra movement shows up where it should. Against the best width on the same draw the fixed pairs line is short by 0.00600; the three plug-ins are short by 0.00627, 0.00637 and 0.00690. Paired against the fixed rule itself, they are worse by 0.000270, 0.000377 and 0.000906, the last at 2.58 paired standard errors.

Reading the draw made the answer worse, and the amount it made it worse by is ordered exactly as the extra variability is. That is the signature of a selection paying its own cost and collecting nothing: the noise arrives in full and the signal never does.

Three charges that read the draw, and none of them helps. How much error each charge delivers above the best band width on the same draw, paired, over 400 draws. The fixed pairs line is short by 0.00600. The three charges that read the sample — scaling that same line by the draw's own persistence, weighting each lag by how much this draw's autocovariance at that lag moves, and estimating the optimism from a split of the sample — are short by 0.00627, 0.00637, 0.00690. Every one is worse than the fixed rule it is calibrated to match on the average draw, the split-sample estimate by 2.58 paired standard errors. Reading the draw is a selection, and here it selects on nothing.
Fig. 3 The error each charge delivers above the best band width on the same draw, paired. Every rule that reads the sample sits behind the fixed rule it was calibrated to.

How much room was there in the first place

A negative result is worth much more with a ceiling beside it, because “the plug-in did not help” and “nothing could have helped” are different findings and only the second is about the problem.

The cheap version of the ceiling is a grid. Take the fixed pairs line and multiply its rate by a quarter, a half, one, two and four, and read the same error column. The width moves enormously — from 29.97 lags at the cheapest charge to 4.71 at the dearest, a factor of 6.366 — so the level of the charge decides the width completely, which is the same reading the essay that put the conventions beside the derived charges arrives at from the other side.

The error moves by 0.003287 across that whole factor of sixteen. And the rate the optimism sweep fitted, which was chosen by describing a profile and never saw an error at all, is the lowest point on the grid.

A factor of sixteen in the charge, and half a shortfall. The error the pairs line delivers when its fitted rate is multiplied by each of 5 factors spanning a factor of 16, over 400 draws. The band width picked runs from 4.71 lags to 29.97, a factor of 6.37, so the level of the charge decides the width completely. The error moves by 0.00329 across the whole grid, against a shortfall of 0.00600 from the fitted rate to the best width on the draw — and the rate the optimism sweep fitted, with no reference to error at all, is the best point on this grid. There is no room here for a better charge to work in, whether or not it reads the data.
Fig. 4 The error the pairs line delivers as its rate is multiplied by a quarter through four. The horizontal line is the best width on the draw.

The ceiling, measured properly

The grid prices one family. The honest ceiling is the error at every fixed width, because that is what any rule choosing one width per draw is aiming at.

Averaged over the same four hundred draws it reads 1.13242 at one lag, 1.11030 at four, 1.10543 at eight, 1.10479 at twelve, 1.10505 at sixteen, 1.10572 at twenty and 1.10748 at thirty. The best single width is twelve lags and the curve around it is nearly level: from six lags to thirty the whole thing spans 0.00269.

Now split the shortfall. The fixed pairs line delivers 1.10586; the best fixed width delivers 1.10479; the best width on the draw delivers 1.09986. So of the 0.00600 the pairs line is short by,

  • 0.00107 is the distance to the best single width — the part a better fixed rule could take, and it is a sixth of the gap;
  • 0.00493 is the distance from the best fixed width to the best width on the draw — five sixths, and it is the part a rule would have to read the draw to get.

The second number is where the essay ends, because of what it is made of. The best width on the draw has a standard deviation of 10.19 lags on a grid running from one to thirty, and 25.5% of draws have their best width at one end of the grid or the other. A quantity whose argmin lands on a boundary in a quarter of cases, over a risk surface that spans 0.00269 from six lags to thirty, is not a property of the draw that a statistic could estimate. It is the argmin of a nearly flat surface with noise on it.

So “the best width on the draw” is a lower envelope rather than a target. Five sixths of the gap every rule in this field leaves is the expected minimum of a noisy flat function, and the expected minimum of a noisy flat function is below every fixed value by construction, whether or not anybody has estimated anything.

The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004.
Fig. 5 The widths the four derived charges and the two conventions pick, in the essay that priced them, with the best width on the draw beside them.
Four derived charges, one error. How much error each charge delivers above the best band width on the same draw, paired, over 150 draws. The two conventions are short by 0.00673 and 0.01048; the four charges derived from the measured optimism are short by 0.00551, 0.00563, 0.00554 and 0.00557 — a span of 0.000115, which is 2.1% of the gap any of them leaves. Getting the scale right is worth something against a convention and getting the shape right beyond a straight line is worth nothing measurable, which is the same reading the earlier field arrived at one level up.
Fig. 6 And the errors those six deliver, where the whole span between the four derived charges is a fiftieth of the gap any of them leaves.

The two readings agree about which part of the gap is available. The width a charge picks is decided by the level of the charge and is worth a sixth of the shortfall; the rest belongs to a quantity that is not there to be estimated. A rule that read the draw perfectly and picked the best fixed width every time would collect the sixth, and the charge derived from the summed weights already collects most of it.

What the split costs before it estimates anything

The split-sample rule deserves a paragraph on its own, because it is the only one of the three that estimates the right quantity and it is the one that does worst.

Splitting a hundred and twenty rows leaves two halves of sixty, and the optimism of a plug-in covariance is not the same quantity at sixty rows as at a hundred and twenty. It is larger, roughly in the way the correction’s own arithmetic requires — the discount a lag carries is 1k/n1 - k/n, so halving nn doubles it. The scale that puts the estimate back on the right footing is a fourth constant, and it is estimated the same generous way as the others.

What the split cannot be given back is the sample it spent. A band of thirty lags fitted on sixty rows is spending half its rows on the covariance, which is the regime the field that measured what a window leaves free is careful to stay out of, and the estimate it returns carries a standard deviation across draws of 7.0324 on a mean of 9.0153. One draw’s reading is worth a little over one standard error of itself.

A quantity estimated at one standard error and then charged is noise with a mean in it, and that is what the table records: the rule that reads the most of the sample is the rule whose width moves most and whose error is worst.

Where a plug-in does work, and what is different there

It would be wrong to read this as a verdict on plug-in rules, and the collection contains the counter-example.

The field that chose a block length from the data builds rules of exactly this kind and finds their correlations with the best block length small but not zero. The difference is not the machinery; it is what the best value is a property of. A block length has to be long enough to carry the dependence in the sample in hand, and a sample that happens to be more persistent genuinely wants a longer block, so there is something in the draw for a statistic to find.

Here the best band width is set by a trade-off between two errors that are both small and both noisy, on a surface the sweep has just shown to be flat. The draw’s persistence does move the optimism — that is what the charge is approximating and the approximation is good — but the optimism is not what decides the width. It is a term in a criterion whose other term is a likelihood that moves further, on a risk surface that barely moves at all.

A plug-in can only help where the quantity it estimates is the quantity that decides, and three constructions failing identically is stronger evidence about that than one would have been.

What the three failures have in common

Each of the three rules is wrong in a way worth stating separately, because they fail for the same reason at three different distances from the data.

The persistence rule reads a statistic with a coefficient of variation of 0.3120 and passes it straight into the level of the charge, so its width moves by 20% more than the fixed rule’s. It is the mildest of the three and it is worse than the fixed rule by 1.21 paired standard errors — not significant, and not on the right side of zero either.

The lag-by-lag rule reads more of the sample and reweights the denominator rather than rescaling it. It moves the width almost identically, is worse by 1.58 standard errors, and the extra structure it reads buys nothing over the level alone.

The split-sample rule reads the most and estimates the right quantity, and it is the worst of the three. Its estimate of what a thirty-lag band costs carries a standard deviation across draws of 7.0324 on a mean of 9.0153, so a single draw’s reading is worth a little over one standard error. Charging that estimate is charging noise with a mean in it.

The ordering is the ordering of how much noise each rule admits, and there is no counterweight anywhere on the table.

What stays out

Three things this essay names and does not settle.

A plug-in for the shape rather than the level. All three rules here keep the pairs count’s structure and vary what multiplies it. A rule that estimated the whole optimism profile from the sample and fitted its own curve to it is a different object, and it would be a selection over a much larger space — so the expectation would be that it costs more, not less, but that expectation is untested.

Whether the flat risk surface is the law’s or the problem’s. Everything above is a first-order autoregression at 0.7 on a hundred and twenty rows. Under long memory the band width matters more, since the law the correction leaves worst is the one whose autocorrelations do not sum, and a steeper risk surface is exactly the setting where a rule that tracks the draw would have something to win. That sweep is not run here.

And the shrinkage a selection would need. The right response to a noisy plug-in is not to abandon it but to shrink it towards its own average, and the amount to shrink by is a quantity this sweep contains and does not extract: the ratio of the signal in the statistic to its total variance. Measured, it would say how much of each of these three rules should have been kept — and on these numbers the answer looks like none of it, which is a prediction rather than a measurement.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • The width a band is measured in — both name bandwidth selection, dependence, estimation error, information criterion, model selection, monte carlo, optimism, plug in estimate, sample autocovariance, tapering
  • A width that moves and an error that does not — both name bandwidth selection, estimation error, information criterion, mean squared error, model selection, optimism, out of sample, regret, tapering
  • A table and a list — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, regret, selection effect
  • The eighth that was not a constant — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, regret, selection effect
  • The quarrel that changes the winner — both name bandwidth selection, benchmark forecast, estimation error, information criterion, model selection, out of sample, regret, selection effect
  • The repair that was exact and made it worse — both name dependence, estimation error, information criterion, mean squared error, model selection, optimism, out of sample, selection effect

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionBenchmark forecastDependenceEstimation errorInformation criterionMean squared errorModel selectionMonte CarloOptimismOut of samplePlug in estimateRegretSample autocovarianceSelection effectTapering