A charge that is not a straight line

What a better charge buys

Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.

Worth reading first: A design is a number · The observations that repeat each other.

Three essays of this field have been about the shape of a charge. This one is about whether the shape does anything.

The construction is the earlier field’s width table, unchanged: for each draw, fit a tapered covariance band at every width on a grid, evaluate the concentrated likelihood, subtract a charge, take the argmax, and score the resulting coefficients against a fresh replicate. What changes is only which charges are on offer — the four derived in this field, and the two conventions beside them for scale.

The widths

Over a hundred and fifty draws of a hundred and twenty rows, under a first-order autoregression:

charge constants width picked
half a log n a lag 3.65
one unit a lag 6.08
a line in the summed weights 1 15.05
a fitted decaying rate 2 15.44
a fitted power law 2 15.47
a line in the pairs 1 15.60
the best width on the draw 14.13

The conventions and the derived charges are a different kind of answer. Half a log n a lag picks under four lags; a unit a lag picks six; every rule derived from the measured optimism picks between fifteen and fifteen and a half, against a best width on the draw of 14.13.

The scale a charge is levied on moves the width by a factor of 4.270. The shape of the charge moves it by 3.6%.

The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004.
Fig. 1 The width each charge picks, with the best width on the draw for comparison. The conventions are the two short bars.

The errors

The width is not what anybody cares about. The error the resulting fit delivers against a fresh replicate is, and the comparison is paired on the draw so that the world drops out.

Against the best width available on the same draw, the shortfalls are 0.00551 for the summed-weight line, 0.00554 for the power law, 0.00557 for the decaying rate and 0.00563 for the pairs line — every one of them about ten paired standard errors above zero, so each is genuinely short of the best and none of them is short by a different amount from the others. The conventions are short by 0.00673 and 0.01048.

The whole span between the four derived charges is 0.000115, which is 2.1% of the gap the smallest of them leaves. A reader choosing between them is choosing inside a fiftieth of a gap none of them closes.

Four derived charges, one error. How much error each charge delivers above the best band width on the same draw, paired, over 150 draws. The two conventions are short by 0.00673 and 0.01048; the four charges derived from the measured optimism are short by 0.00551, 0.00563, 0.00554 and 0.00557 — a span of 0.000115, which is 2.1% of the gap any of them leaves. Getting the scale right is worth something against a convention and getting the shape right beyond a straight line is worth nothing measurable, which is the same reading the earlier field arrived at one level up.
Fig. 2 What each charge costs above the best width on the same draw, paired, with the number of paired standard errors beside each bar.

What “paired” is doing here

The shortfalls above are differences taken on the same draw, and the pairing is what makes a span of 0.000115 readable at all.

The absolute errors these rules deliver are around 1.09 and their standard errors across the sweep are about 0.0075. So an unpaired comparison between two charges — 1.09356 against 1.09368 — would be two numbers a hundred times further apart in their own noise than they are from each other, and would say nothing whatever.

Paired on the draw, the difference between two charges is the difference between two widths chosen on the same world from the same likelihood profile, and on most draws that difference is exactly zero because the two charges pick the same width. What is left is a small number with a small standard error, and it is what the table above reports.

The same pairing is what makes the shortfall against the best width readable: 0.00551 with a standard error small enough to be ten of them away from zero, on a quantity whose unpaired spread is fifteen times larger. The comparison is only affordable because the world cancels, and that is a fact about the design rather than about the charges.

Which is not a disappointment

It would be easy to read that as the field having found nothing, and it is worth being careful about what has and has not been established, because the same shape of result recurs in this collection and is usually the interesting half.

What moved. Against a convention, the derived charges close about half the gap: 0.01048 for half a log n a lag, 0.00673 for a unit a lag, 0.00551 for the best derived rule. That is a real improvement and it is the earlier field’s finding rather than this one’s.

What did not. Getting the scale right — which is the whole of this field’s contribution — moves the width from 15.05 to 15.60 and the error by a ten-thousandth. Getting the shape right beyond a straight line moves it less.

So the honest summary is that this field found the right denominator for a quantity the answer is nearly insensitive to, and the reason that is still worth doing is in the next section.

What a rule would have to be to close the gap

“None of them tracks the answer” is the right reading of the correlation column, and the column also says how much tracking would be needed to matter.

The six correlations run from −0.006 to 0.017, and on a hundred and fifty draws a correlation’s own standard error is 1/√150 = 0.082. So every one of them is well inside a standard error of zero, and the sweep could not have detected a correlation below about 0.16 — meaning the honest claim is that the rules explain under three per cent of the best width’s variance rather than none of it.

What would be needed is far more than that. The best width’s standard deviation across draws is about 10.3 lags; a rule correlating at r leaves a width error with standard deviation 10.3√(1 − r²); and on a locally quadratic error curve, halving the shortfall takes the width error down by a factor of √2. That needs r = 0.71.

Six rules at under 0.02, against a bar at 0.71. The gap is not a matter of a better charge being thirty per cent short of a tracking one — nothing on the table is within a factor of thirty of the correlation that would be required, and the shortfall of 0.0055 is therefore almost entirely irreducible by anything of this kind.

The level bias is a twenty-sixth of the shortfall

The same arithmetic splits the shortfall into the part a better charge could fix and the part it could not.

All four derived rules pick between 15.05 and 15.60 against a per-draw best averaging 14.13 — every one of them about a lag too wide, in the same direction. That is the level bias, and it is the only thing a charge with the right shape and the wrong constant gets wrong.

On an error curve that is flat between eight and twenty lags, a lag and a quarter of level bias is worth roughly 0.0002 of error. The measured shortfall is 0.0055.

So of the gap the derived charges leave, about one part in twenty-six is the level and the rest is the failure to track a target that moves by ten lags between draws. Correcting the level exactly — which is the most a fifth essay in this field could do — would recover four per cent of what is left, against the two per cent that separates the four charges from each other.

That is the field’s own result stated as a bound rather than as a disappointment: the charge’s scale was worth a factor of four, its shape was worth two per cent, and its constant is worth four. There is nothing else of this kind left.

Because nothing here is an estimate of anything

There is a column in the table that says why the errors are flat, and it is the one nobody looks at.

For each rule, how much does the width it picks track the width that would have been best on that draw? Not on average — on the draw. The correlations are −0.006 for the summed-weight line, −0.003 for the pairs line, 0.017 for the power law, −0.001 for the decaying rate, 0.016 for a unit a lag and −0.004 for half a log n a lag.

None of them. Not one of the six rules carries any information about which width this particular sample wanted.

That is the same reading the earlier field arrives at — its four charges correlate 0.037, 0.006, 0.063 and 0.065 with the best width, against a target whose own standard deviation is 10.30 — and it explains the flat error column completely. Six rules that are all uncorrelated with the answer deliver the same error, whatever else is true of them, because the only thing left for a charge to get right is the average level, and the average level is the thing all four derived charges agree about.

The best width on the draw has a standard deviation of 10.44 across the sweep, against 5.86 for the summed-weight line’s picks and 6.50 for the pairs line’s. So the target moves twice as much as any rule that is meant to be estimating it, and moves independently of all of them.

None of the rules is an estimate of the width it is read against. The correlation, across 400 draws of 120 rows, between the width each charge picks on a draw and the width that would actually have been best on that same draw. Every one is under 0.070: 0.037 for one unit a lag, 0.006 for half a log n a lag, 0.063 for the numbers the window leaves free, 0.065 for what the optimism actually is. The best width averages 13.90 with a standard deviation of 10.30 across draws, so there is a great deal to track and none of it is being tracked. A table of four charges beside a column headed the best width on this draw invites the reading that the rules are estimates of it and one of them is close. They are not estimates of it at all; they are four nearly deterministic functions of the sample size sitting in the neighbourhood of a very noisy target.
Fig. 3 How much of the best width each rule’s answer tracks, in the field that first measured it — which is nearly none of it.

The grid, and what it coarsens

One economy is worth stating, because it puts a floor under how finely any of this can be read.

The widths are chosen from a grid of eleven — 1, 2, 3, 4, 6, 8, 12, 16, 20, 24 and 30 — which is the earlier field’s grid and is spaced to be dense where the optimism moves fastest and sparse where it does not. At the wide end the steps are four and six lags apart.

So an average width of 15.05 against 15.60 is not two rules differing by half a lag. It is two rules that agree on most draws and disagree by one grid step on a few, and the averages differ by half a lag because a step at that end of the grid is four lags wide. The same is true of the best width on the draw: its standard deviation of 10.44 is partly the grid’s coarseness at the top, since a draw whose true best width is 26 must report either 24 or 30.

That does not weaken the comparison — every rule picks from the same grid, so a difference between two of them is a difference in what they charge and not in what they are allowed to pick — but it does mean the resolution of the width column is a grid step rather than a lag. It is the error column that is measured finely, and the error column is flat.

What is being chosen from

The plateau figure is the mechanism and it is worth reading directly.

A criterion picks a width by maximising a penalised objective across the grid. If that objective is flat across many widths — if a dozen widths sit within a log-likelihood unit of the best — then the argmax is decided by whichever way the noise happens to point, and a charge that shifts the whole curve by a small amount shifts the argmax around inside the flat region without moving what the fit delivers.

That is what is happening. The objective’s plateau is wide, the four derived charges differ by a few per cent of the charge at the widest band, and the widths they pick differ by half a lag on a grid whose steps at that end are four and six lags wide.

A charge cannot be more precise than the objective it is subtracted from. This field spent three essays getting a denominator right and then measured that the answer barely moves, and both halves of that are the result.

There is a version of this that would have been worth acting on and it is worth saying why it did not happen. If the plateau were narrow — if the penalised objective had a sharp maximum — then a three per cent change in the charge would move the argmax by a controlled amount and the width would be an estimate of something. The plateau is not narrow, and the field that measured its flatness reports the widths a criterion cannot tell apart as a range rather than a curvature for exactly that reason.

Reading the two conventions honestly

The two conventions come out badly in this table and it is worth saying exactly how badly, because “a factor of four too narrow” invites more indignation than the numbers support.

Half a log n a lag picks 3.65 lags and delivers 1.09853. The best derived rule picks 15.05 and delivers 1.09356. The convention is therefore short of the best available width by 0.01048 where the derived rule is short by 0.00551 — so it closes none of the gap the derived rule closes, and it costs about twice as much.

But twice a small thing is a small thing. Both numbers sit on an error of about 1.09, so the whole span from the worst convention to the best derived rule is 0.5% of what a practitioner is actually paying. The dominant term is not the charge; it is that a hundred and twenty rows are not enough to estimate a covariance band well under any rule, which is what an error of 1.09 against a best-on-the-draw of 1.088 is saying.

That reading is the same one the earlier field arrives at, and it is the right frame for everything in this field: the choice between charges is a choice inside a shortfall several times larger than the choice, which is a shape this collection meets often enough to expect it.

What would have moved it

It is worth naming what a charge would have had to do to matter, because that is the deferral this field leaves.

It would have to be correlated with the draw. Every rule here is a deterministic function of the width, the window and the sample size, so every rule picks its width from the likelihood alone and the likelihood is flat. A charge that read something about this sample — how persistent it looks, how far the sample autocovariances fall off — would have a chance of tracking the best width, and none of the six does.

The earlier field’s plug-in rules do exactly that in a neighbouring problem, and their correlations with the best block length are also small but not zero. Whether the same trick helps here is not measured, and it is the one thing that would turn a flat error column into a moving one.

A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it.
Fig. 4 The fits this field’s four charges come from, in the previous essay.

What the field is left holding

Four essays, one denominator, and an error column that does not move. It is worth setting out what a reader should take from that, because a field can be right and small at the same time.

The curvature was real and it is explained. The share of its summed weights a Bartlett band spends falls at every step across the plateau, a constant share misses it by 0.6323 per width in units of the widths’ own errors, and counting the band’s width in pairs takes that to 0.0632 — on a correction that reads no data, checks against a closed form, and moves with the sample size in the direction its mechanism requires.

The explanation is one window’s. The other three windows in the table are improved by a factor of two or less where this one is improved by ten, and the field says so in every figure rather than reporting the one that works.

And the repair is worth about a ten-thousandth. Which is not a reason to have used the wrong denominator. A rule that is right on average and wrong at both ends will be wrong at both ends whether or not anybody has measured how much that costs today, and the cost is a property of the grid, the sample size and the law rather than of the rule.

The general form of that is worth carrying: a derivation that is exact for one quantity is not thereby the right denominator for another, and the way to notice is to read a ratio across the range rather than at one point.

What stays out

Three things this field could have measured and did not.

Whether the correction transports to a second sample size. Everything here is a hundred and twenty rows. The pairs width’s ratio to the summed weights is checked at sixty and four hundred and eighty rows, but the optimism profile is not re-measured there, so the claim that the corrected share is the same constant at another n is a prediction rather than a measurement.

What distinguishes the window the correction works on from the three it does not. All four are measured, the answer is reported, and no mechanism is offered. A window’s weight profile has a mean lag as well as a sum, and the three windows the correction fails on all concentrate their weight at short lags; whether that is the explanation is a guess.

And a charge that reads the sample. The flat error column above is entirely explained by no rule tracking the best width, and every rule here is a function of the width alone by construction. That is the shape of the next question rather than a gap in this one.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A penalty is a trace — both name degrees of freedom, dependence, information criterion, mean squared error, model selection, optimism, out of sample, overfitting
  • A table and a list — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, overfitting, regret
  • The displacement is a parameter count — both name benchmark forecast, degrees of freedom, information criterion, mean squared error, model selection, monte carlo, out of sample, overfitting
  • The eighth that was not a constant — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, overfitting, regret
  • The quarrel that changes the winner — both name bandwidth selection, benchmark forecast, estimation error, information criterion, model selection, out of sample, overfitting, regret
  • Two factors pointing opposite ways — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, overfitting, regret

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionBenchmark forecastDegrees of freedomDependenceEstimation errorInformation criterionMean squared errorModel selectionMonte CarloOptimismOut of sampleOverfittingPlug in estimateRegretTapering