A charge that is not a straight line

A line that beats two curves

A deferral asked for a curve. Fitted against the same measurements, a straight line in a variable nobody had to fit describes the plateau better than either curve does with a constant more — and for three windows out of four it does not.

Worth reading first: A design is a number · The observations that repeat each other.

The question this field was opened to answer was put plainly: the share of its summed weights a covariance band spends falls from 0.8447 at two lags to 0.7399 at thirty, so the right charge is a curve and every rule considered is a straight line through the origin. Fit the curve.

This essay fits two, puts them beside two lines, and reads off which describes the measurement at how many constants.

Four shapes

Each is fitted to the same eleven measured rises, on the same eight thousand draws, by weighted least squares.

A line in the summed weights. One constant: a rate per unit of Σ w(k). It is the earlier field’s own charge, and it is the thing to beat.

A line in the pairs. One constant: a rate per unit of Σ w(k)(1 − k/n). The previous essay is the argument for that denominator; here it is simply a second candidate scale, fitted the same way as the first.

A power law. Two constants: a times the summed weights to the b. It is the shape a reader would fit by eye on a log-log plot, and it is the natural answer to “the charge grows more slowly than the width” — the reading the deferral that opened this field proposed.

A rate that decays with the width. Two constants: a·w/(1 + w/c), a rate a at the narrow end falling towards half of it at a width of c. It is the natural answer to “the first lag is dearest”.

None of the four has an intercept, and that is a decision rather than an omission. A constant added to a criterion at every width cannot change which width it picks, so an intercept is a parameter that improves a fit and decides nothing — and a fit contest in which one entrant is handed a free parameter that does no work is not a contest.

A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it.
Fig. 1 How far each shape sits from the measurement across the plateau, in units of each width’s own standard error. Lower is a better fit.

How they are weighted, and why it matters here

Every width is weighted by the reciprocal of its own squared standard error.

That is the ordinary thing to do and it is load-bearing. The sweep resolves a two-lag band’s rise to ±0.0440 on a rise of 0.4224 and a thirty-lag band’s to ±0.1388 on a rise of 10.7279 — a tenth of the quantity at one end and a seventy-seventh at the other. An unweighted fit would let the two widths the sweep knows least about set the shape, and those two widths are also the two furthest from the plateau on every scale.

The consequence is worth stating because it cuts against this essay’s own conclusion. Weighted this way, the fit over all eleven widths favours the curves: the saturating rate sits at 0.0894 per width and the power law at 0.1047, against the pairs line’s 0.1005 and the summed-weight line’s 0.7460. A reader who stopped there would conclude the deferral was half right.

What that comparison is dominated by is the narrow end, where the pairs correction does almost nothing and both lines are wrong. So the fits are reported twice: over the whole sweep, and over the plateau where the field’s claim is actually made.

Across the plateau

From four lags to thirty, in units of each width’s own standard error:

shape constants χ² per width
a line in the summed weights 1 0.6357
a line in the pairs 1 0.0640
a fitted power law 2 0.1254
a fitted decaying rate 2 0.0368

The one-constant line in the pairs fits the plateau ten times better than the one-constant line in the summed weights, and twice as well as a fitted power law with a constant more. Its worst residual anywhere in the sweep is 0.51 of a standard error.

One shape fits better: the fitted decaying rate, at 0.0368, which is within a factor of two of the pairs line and has a constant more. So the cheapest shape that fits the plateau is the one with nothing fitted in its scale, and it is not the best-fitting shape on the table. That is the answer to the deferral, and it is not the answer the deferral expected.

What the curves’ constants come out as

The two curves are worth reading rather than dismissing, because both of them find the same thing in a more expensive way.

The power law fits an exponent of 0.9641 — sublinear, which is another way of saying the charge grows more slowly than the summed weights, which is the curvature. Its rate constant is 0.8200, and 0.8200 × w^0.9641 tracks 0.8133 × P(w) closely across the plateau because P(w)/w falls from about 0.98 to about 0.91 over exactly that range.

The decaying rate fits a scale of 181.04 — meaning the rate is meant to halve at a width of a hundred and eighty-one summed weights, which is twelve times the widest band in the sweep. A decay scale far outside the data is a polite way of saying the fit wants a gentle bend and has nowhere to put one, and its rate constant, 0.8003, is within a hundredth of the pairs line’s 0.8133.

Both curves are approximating the same discount, and both need a constant to do it that the pairs count does not.

The curvature is in the denominatorThe optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.0.8000.9001102030the band's width, in lagsthe measured charge, per unit of the band's widththe plateau startsin the pairs the band usesin the weights it spends2000 draws of 120 rows, Bartlettflat means a straight line is right
Fig. 2 The profile the four shapes are fitted to, on both denominators. The slider changes how many draws the sweep takes.

What a residual of a fifth of a standard error means

The pairs line’s plateau residuals are all under half a standard error, and a reader is entitled to be suspicious of a fit that good.

It is not too good. The eight plateau widths are not eight independent measurements: they come off the same two thousand draws, they share the same base row, and a draw whose optimism at twelve lags is high is very likely to be high at sixteen as well. So the residuals are correlated by construction, and a set of correlated residuals will routinely look smoother than the same number of independent ones.

What that means for the comparison is nothing, because every shape is fitted to the same correlated residuals. The χ² column is a relative reading — this shape against that shape on these measurements — and it is not a goodness-of-fit test against a distribution. A reader who wanted to know whether the pairs line is acceptable in an absolute sense would need independent sweeps at each width, which would cost eight times what this one does and would answer a question nobody here is asking.

The absolute reading that is available is the one in the first essay of the field: the summed-weight profile falls at every one of its seven steps across the plateau, so there is a real slope for a shape to have to follow, and the pairs line follows it by not having one.

Where the line loses, which is three windows of four

The result above is for the Bartlett window. Run the same four fits under the other three and the answer changes.

window line in weights line in pairs power law decaying rate cheapest
Bartlett 0.6357 0.0640 0.1254 0.0368 the pairs line
its square 1.6417 0.7822 0.0207 0.2396 the power law
its cube 0.9703 0.5143 0.1604 0.0732 the decaying rate
Parzen 5.9491 4.4315 0.7826 2.3904 the power law

Three of the four windows want a curve, and the two-constant shapes earn their constants on all three.

That is not a small qualification. The window the earlier field measured its charge on is the plain Bartlett one, and it is the one for which the correction works; the three it does not work for are the three that field carried specifically so that a claim about windows would not rest on one window.

So the answer to “is the charge a curve?” is not one answer. For the window the earlier field measured, it is a straight line in a variable nobody had to fit. For the other three, it is a curve in every variable tried here. What distinguishes them is not established, and saying so is the point of fitting all four rather than the one that works.

The repair is one window's. How far the measured charge per unit of width sits from a constant across the plateau, in units of each width's own standard error, on each of two scales, for each of four windows, over 2000 draws. A reading near zero means a straight line through the origin is the right shape in that scale. Counting the band's width in pairs rather than in summed weights improves the Bartlett window by a factor of 28.16 — from 0.2359 to 0.0084 — and does far less for the other three: 1.80, 1.61 and 1.31, on profiles that are ten to seventy times further from flat to begin with. So the correction, which has no fitted parameter and is stated for windows in general, repairs the one window the earlier field measured and leaves the rest wanting a curve.
Fig. 3 The spread of the charge across the plateau on both denominators, window by window, which is the same limit read as a flattening factor rather than as a fit.
Four windows, one line, and one that is off it. The optimism measured for each window at a band of 30 lags, against what that window's weights sum to, on 2000 pairs of independent samples of 120 rows. The diagonal is where a window that spent exactly its summed weights would sit. Three of the points are one shape at three levels — the Bartlett window, its square and its cube, whose sums stand in the ratio 6 : 4 : 3 — and they lie on a line through the origin at 0.767 of the diagonal, with 0.033 between the highest and the lowest. Scaling the weights scales the charge by the factor the weights predict, which is what makes the weights the mechanism. The Parzen window has a comparable sum and a different shape, and it sits at 0.871: its weights stay near one over the first few lags, and the first few lags are where the information is. A weight sum treats every lag as equally informative and no sample does.
Fig. 4 The one-shape family whose weight sums stand as six to four to three, in the field that assembled it.

What the two curves would do to a practitioner

A charge is not a description; it is something somebody levies. So it is worth asking what a reader who took each of the four shapes seriously would end up doing differently, before the next essay measures it.

The line in the summed weights charges 0.7557 per unit of Σ w(k). On a Bartlett band that is 0.3778 per lag, so it is about a third of what a unit-per-lag convention charges and about a seventh of what half a log n a lag charges on a hundred and twenty rows. A rule that charges a third as much picks a much wider band, and the earlier field measures how much wider.

The line in the pairs charges 0.8133 per unit of Σ w(k)(1 − k/n). Because the pairs width is below the summed weight width, and increasingly so, the two rules charge almost the same at the wide end and the pairs rule charges slightly more at the narrow end — which pushes it towards wider bands, not narrower ones.

The power law charges 0.8200 w^0.9641, which is above the summed-weight line at every width in the sweep and converges towards it slowly. The decaying rate charges 0.8003 w/(1 + w/181.04), which is below the pairs line everywhere by a per cent or two.

So all four are near neighbours in the units that matter, and the largest difference between any two of them at the widest band in the sweep is under three per cent of the charge. That is the arithmetic behind the next essay’s finding, and it is worth having in advance: a fit contest can be decisive about shapes and still be about a quantity that barely moves.

The rate the pairs line settles on

One number is worth carrying out of the fits: the pairs line’s rate is 0.8133 per unit of pairs width, and the summed-weight line’s is 0.7557 per unit of summed-weight width.

Neither is one, and neither is meant to be. A convention charging one unit a lag is charging thirty for a thirty-lag Bartlett band; the summed-weight line charges 0.7557 × 15 = 11.3; the pairs line charges 0.8133 × 13.6667 = 11.1. So the two derived rules land within two per cent of each other at the widest band in the sweep, which is exactly what a re-scaling should do — they were fitted to the same measurements and they agree where the measurements are densest.

Where they part is at the narrow end, and by construction: at four lags the summed-weight line charges 1.134 and the pairs line 1.196, a difference of five per cent in the same direction as the measured 1.1852.

The pairs rates for the other three windows come out at 0.8327, 0.8320 and 0.9512, which is the same ordering the earlier field reports for its own shares and is another way of saying that the correction has changed the shape without changing which windows are dear.

The shape nobody fitted, and why it is not on the table

There is a fifth candidate a reader will think of, and leaving it out is deliberate.

The pairs count discounts a lag by 1 − k/n. If the argument is that a long lag is estimated from less, then the discount could equally be (1 − k/n)² — because the variance of γ̂(k) grows roughly like 1/(n − k) — or (1 − k/n)^p for a p fitted from the profile. That family would certainly fit better than p = 1, because it contains p = 1.

It would also stop being the thing being claimed. The claim is that a correction with nothing fitted takes an order of magnitude off the misfit, and a correction with an exponent chosen to remove the misfit removes it by construction. A fitted power of the pairs count is a curve, and the curves are on the table already: the power law here fits an exponent of 0.9641 on the summed weights, which is arithmetically almost the same object.

So the table is four shapes rather than five, and the fifth is named here so that a reader who wants it knows it was declined rather than missed. The place it would belong is a field that had a second, independent reason to prefer one exponent — an argument from the sampling variance of a tapered autocovariance, say — and this one does not have it.

The correction helps everywhere and decides once

“Three of the four windows want a curve” is the right summary of which shape wins, and dividing the two one-constant columns says something else the table contains.

The pairs correction’s improvement over the summed-weight line is a factor of 9.93 on the Bartlett window, 2.10 on its square, 1.89 on its cube and 1.34 on Parzen. It is an improvement on every window — the correction is never the wrong direction — and it is an order of magnitude on exactly one of them.

That is a more useful statement than the win column. The derived denominator is right about the shape everywhere and decisive only where the profile is otherwise smooth, so a reader who wanted a single scale to charge on would take the pairs count on all four windows and lose a factor of two to a fitted curve on three of them, rather than take a different shape per window.

The other reading in that table is which window is hard, and it is not the one the win column suggests. The best fit available on each window is 0.0368, 0.0207, 0.0732 and 0.7826 — so Parzen’s best entrant is ten to forty times worse than any other window’s best. Three of the four windows are fitted well by something on the table and one is fitted well by nothing on it.

So “three of four want a curve” understates the qualification. One of the four wants a shape that was not offered, and no amount of choosing between these four settles what it is.

Where the four charges actually differ

The four shapes are near neighbours in the units a practitioner levies, and the two ends of the sweep say how near.

At the widest band the four charge 11.34, 11.11, 11.16 and 11.09 — a spread of 0.25 on about 11.2, which is 2.2%. At four lags they charge 1.134, 1.196, 1.212 and 1.191 against a measured 1.185, a spread of 6.6%.

And the spread at the narrow end is one entrant’s. Drop the summed-weight line and the other three sit within 2.3% of each other and within 2.3% of the measurement; keep it and it is 4.3% below the measurement on its own. The whole of the disagreement between the four charges, at the width where they disagree most, is the shape whose scale was not corrected.

That is the fit contest’s practical residue stated in the currency the next essay spends. Three of the four shapes are the same charge to within a fortieth at every width in the sweep; the fourth differs by a twentieth and only where the band is narrow. A field that set out to decide between four shapes has found that three of them are one shape, and that the one it can distinguish is the one it started from.

What a fit contest can and cannot settle

Two things about the method, because a table of χ² invites more than it can support.

It can say which shape describes the measurements. All four are fitted to the same numbers with the same weights, so the comparison is like for like, and counting constants is the only defensible way to break a tie between shapes that fit comparably.

It cannot say which shape is right. Four shapes were chosen and fitted; a fifth would have had a fifth answer. The reason the pairs line is reported as the answer is not that it won a contest — the power law wins two of the four windows, and a curve fits the Bartlett plateau better than it does — but that it is the only entrant whose scale was derived rather than chosen, and it is the cheapest shape that fits on the window the derivation was made for. A curve that fits better with a constant more has not explained anything; a line that fits with nothing fitted has proposed a mechanism, and the mechanism is checkable independently, which is what the previous essay does with the sample-size sweep. The same discipline is why the fields upstream of this one measure a plug-in’s shortfall against a maximiser rather than assuming one: a plug-in is not a maximiser, and the difference is only visible to somebody who computes both.

And none of it says whether the difference between the four is worth having. That is the last essay of the field, and the answer there is that the four charges pick widths within four per cent of each other and deliver errors within three per cent of the gap any of them leaves.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A penalty is a trace — both name closed form, degrees of freedom, dependence, information criterion, least squares, model selection, optimism, overfitting
  • The displacement is a parameter count — both name closed form, degrees of freedom, information criterion, least squares, model selection, monte carlo, nested models, overfitting
  • A step that is not a ratio — both name bandwidth selection, dependence, information criterion, model selection, monte carlo, nested models, overfitting
  • A table and a list — both name bandwidth selection, dependence, information criterion, model selection, monte carlo, nested models, overfitting
  • Nothing in the fit picks the width — both name closed form, degrees of freedom, information criterion, model selection, nested models, overfitting, tapering
  • One number for a table of candidates — both name closed form, degrees of freedom, dependence, information criterion, least squares, model selection, optimism

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionClosed formDegrees of freedomDependenceEstimation errorInformation criterionLeast squaresModel selectionMonte CarloNested modelsNormalisationOptimismOverfittingParameter uncertaintyTapering