How long the list is

The comparison that was not made

Choosing a whitening's window separately for every candidate costs 0.00401 of regret. The same question about an order was named and left, because the two lists are different lengths. The order's answer is 0.00360, and matching the lists changes almost nothing.

Worth reading first: A design is a number · The observations that repeat each other.

Choosing the whitening’s window separately for each candidate costs 0.00281 of delivered regret against choosing it once for the whole table, at two paired standard errors, and 0.01087 when the covariance is estimated per candidate as well. That measurement ends with a sentence naming what it could not do:

The window’s is measured and the order’s is not, because the order list runs to twelve where the window list runs to eight, and making the two the same length changes what each rule is.

The reason is a good one. A rule minimising a criterion over thirteen values and a rule minimising it over eight are not the same rule, so a difference between them would not be a difference between a window and an order. The direction was predicted — the longer list should cost more — and the size was not offered.

Both halves of that turn out to be wrong, and the second one interestingly.

The missing number

Choosing a sieve order per candidate rather than once, over its own list of thirteen values, costs 0.00360 of regret at 1.7 paired standard errors. On the same draws, in the same sweep, the window’s own figure is 0.00401 at 2.3.

They are the same size.

The number the comparison was missing. What it costs to choose the tuning parameter for every candidate separately rather than once for the table, under AR(1) at 0.8, paired on the draw. The window's figure is the one the earlier field reported; the order's is the one it named and did not make. They are the same size — 0.00401 against 0.00360, at 2.30 and 1.72 paired standard errors — and matching the lists at eight values leaves them the same size again. The prediction that the longer list would make the order's cost the larger of the two is not what happens; what happens is that the two rules cost the same once they are scored by the same criterion, which took a missing term to arrange.
Fig. 1 Both tuning parameters at their own list lengths and at a matched eight, with the paired standard error of each.

Matched at eight values apiece the window costs 0.00401 and the order 0.00415, which is closer still and in the opposite order. Matched at four they are 0.00223 and 0.00406. Nowhere in the sweep is the difference between the two parameters larger than the standard error of either.

So the answer to the question the earlier field asked is that there is nothing to see: a window and an order cost the same when they are chosen per candidate, whatever length of list either is chosen from.

What that costs to say honestly

The number above is not the number this measurement produced first, and the difference between the two versions is the whole of the next essay.

Scored the way this collection has scored a sieve since the estimated-covariance field, a per-candidate order costs −0.00011 at 0.2 standard errors — nothing, and slightly the wrong sign. That reading is what the machinery gave for a day, it is entirely consistent with the earlier field’s worry running the other way, and it is wrong because the two rules were not being scored by the same criterion. The term that was missing is the volume the whitening moves, the window’s rule has carried it since it was written, and the sieve’s never had it.

The point worth carrying from that here is narrower than the repair. A comparison between two procedures is a comparison between two implementations of two criteria, and the second pair is the one that has to be checked. Two rules that differ in a term neither prose describes will differ by whatever that term is worth, and it will look like a difference between the rules.

A difference between criteria, read as a difference between rules. The extra error from choosing a tuning parameter per candidate, under AR(1) at 0.8, three ways. Scored by the criterion this collection has used for a sieve since the estimated-covariance field — which carries no determinant — a per-candidate order costs -0.00011, at -0.16 standard errors, and the honest reading of that is nothing. Scored by the criterion the window's rule has always used, which carries the volume its whitening moves, the same rule costs 0.00415 at 2.02 — the window's own 0.00401. The rules were never different. The criteria were.
Fig. 2 The same rule under two criteria, beside the window’s, which is what the repair was.
The term that cancels, and the term that does not. The volume each candidate's whitening moves — log|Ω̂| — for a sieve of order 4 on one sample of 120 rows. Estimated once from the fullest candidate and used for the whole table, it is the same number for every candidate, so it drops out of every difference the criterion reads: that is why nothing in this collection has ever needed to carry it. Estimated from each candidate's own residuals it ranges over 23.87, which is more than a parameter is worth, and the criteria being compared are then fits made under different error models with no term saying so. The window's rule has carried this term since the estimated-covariance field and the sieve's never had it.
Fig. 3 The volume each candidate’s own whitening moves, which is the term the comparison needed.

One of the two rules cares about its list and the other does not

The sweep prints each rule at three list lengths, and reading down the columns rather than across them finds an asymmetry the matched comparison was designed to remove.

The window’s cost roughly doubles with its list. Four values cost 0.00223 and eight cost 0.00401.

The order’s does not move. Four values cost 0.00406, eight cost 0.00415 and thirteen cost 0.00360 — a range of fifteen per cent, in a quantity whose own standard error is about 0.0021, which is to say no movement at all.

So the earlier field’s prediction was not merely wrong about the order; it was wrong about the rule it applies to. A longer list does cost more, and it costs more for the window, whose list is the shorter of the two and whose ladder is geometric — the extra values it adds are in the middle, where they let two candidates land closer together and quarrel more cheaply. The order’s list is a dense run of integers over the same range at every length, so lengthening it adds resolution rather than reach, and resolution is not what a per-candidate rule is spending.

That also explains why matching the two at eight values makes them closer rather than clarifying anything: matching moves the window up its own curve and leaves the order where it was.

How close “the same size” is

The largest gap anywhere in the sweep is at four values on the list, where the window costs 0.00223 and the order 0.00406 — a difference of 0.00183.

Set against the rules’ own standard errors, which are about 0.0021 for the order and 0.0017 for the window, that gap is comparable to either. The standard error of the difference is larger than either of those, so the gap is inside it by a wider margin than the comparison of point estimates suggests.

At the two other list lengths the gaps are smaller still: 0.00041 at the rules’ own lists and 0.00014 matched at eight — the second being three ten-thousandths of a regret, or eight per cent of one standard error.

Which is worth saying plainly, because “the same size” can mean two rather different things. This is not two effects that happen to round to the same figure. It is a difference that the measurement cannot distinguish from zero at any list length, and which is smallest exactly where the two rules are made most comparable.

What choosing per candidate does

Both rules cost about four thousandths of regret, on a baseline where the best model available is worth about 0.02 more than the worst rule delivers. That is a fifth of the whole quantity the field is about, and it is worth being clear about which fifth.

A candidate’s criterion under a shared nuisance is a fit made under a common error model. Under a per-candidate nuisance it is a fit made under its own error model, and the table then ranks fits that were computed under different assumptions. The two things that go wrong are different sizes:

The comparison stops being like for like. Each candidate’s score now carries its own whitening, and the determinant that says how much volume that whitening moved does not cancel between candidates. That is a bias, it is computable, and it is what the missing term was for.

And each candidate’s criterion becomes a minimum over the list. A minimum over eight numbers is below a typical one, by an amount that differs between candidates, and the difference is what displaces the choice. The window field measures that displacement at a true null: a per-candidate window picks a larger candidate on 13.5% of draws against 3.0% smaller.

What a tuning list manufactures where there is nothing to find. At a true null under AR(1) at 0.8 every candidate contains the truth, so nothing distinguishes them and any systematic preference is arithmetic rather than discovery. Choosing the window from each candidate's own residuals rather than once for the table moves the choice to a larger candidate on 13.5% of 200 draws and to a smaller one on 3.0%. The mechanism is that each candidate's criterion becomes a minimum over 8 windows: that manufactures 19.5 units of criterion on average, and — the part that displaces — it manufactures 2.32 units more for one candidate than for another, where a parameter costs two.
Fig. 4 The displacement at a true null, where nothing can be discovered and any preference is manufactured.

The candidates disagree about what to use

The mechanism is visible without any criterion at all, in what the rules pick.

Over five candidates and a list of thirteen orders, the choices are not all equal on 39% of draws at a true null, and the spread between the largest and smallest chosen order averages 1.56. For the window at its own list of eight the disagreement rate is 42% and the spread 2.15.

Those are the same rates, from two lists of different lengths, for two parameters that are not alike in anything except being tuning parameters. Which is the first hint that the length of the list is not the acting quantity — a hint the next essay but one turns into a measurement.

What a longer list actually changes. How often the five candidates choose different tuning parameters, at a true null where every one of them contains the truth, so a disagreement is manufactured rather than discovered. The order's list is an interval of integers, and thinning it moves the rate smoothly from 0% at two values to 39% at thirteen. The window's is not an interval — it runs 0, 1, 2, 4, 8, 12, 20, 30 — so a thinned window list jumps depending on whether it happens to keep the width the criterion wants, between 0% and 42% with no order to it. So "the same length" was never quite the same thing for the two rules, and it is a smaller effect than the field it was invoked to explain.
Fig. 5 How often the candidates choose differently, against how many values they may choose from.

The two rules, written out

It is worth setting the two constructions side by side, because everything above depends on them being the same construction with one substitution.

The window rule. From a residual series, take the sample autocovariances out to LL lags, weight each by the Bartlett triangle so the result is positive definite, factor the resulting banded Toeplitz matrix, and whiten with the factor. The value of LL is chosen by a criterion on the residuals themselves: how white the whitened series is, charged two per lag spent.

The order rule. From the same residual series, fit an autoregression of order pp by the Yule–Walker recursion, form the covariance matrix that autoregression implies, factor it, and whiten with the factor. The value of pp is chosen by a criterion on the residuals: the fitted innovation variance, charged two per coefficient.

They differ in one place — what the sequence past the chosen lag looks like. A band says zero and a sieve says an exponential tail. That difference is what makes the sieve escape a ceiling the band sits under, and it is the reason the two are worth comparing at all rather than being two names for one thing.

What each construction carries, against what there was. The autocorrelation of a resampled error series at five lags, averaged over 60 samples of 40 resamples each. Three facts are in the picture. The residuals lie below the errors at every lag, which is the ceiling a multiplier cannot exceed. The blocked multiplier and the fixed-length block lie on top of each other below it — they attenuate identically, because the attenuation is the join — while the stationary bootstrap, whose runs are geometric rather than fixed, sits above them both. And the sieve is the exception in kind rather than in degree: at lag six it carries 0.0638 where the residuals have 0.0300 and the multiplier has -0.0011, because a fitted model extrapolates past the lags it was told about and a truncated sample sequence cannot.
Fig. 6 What each construction keeps of a dependence past its own cut-off, which is the one place the two rules differ.

Everything else — the candidate table, the sample, the criterion the candidate is scored by, the baseline the regret is measured against — is identical, and identical by construction rather than by description: both go through the same sweep with a scorer swapped.

Why the two costs being equal is a result

It would be easy to file a null result and move on, and it is worth saying why this one is worth the field.

The two tuning parameters are not alike. A window is a bandwidth: it says how many lags of the sample’s own autocovariance sequence to keep, it enters through a tapered band, and widening it moves a determinant a long way. An order is a model dimension: it says how many coefficients of an autoregression to fit, it enters through a whitening that has an exponential tail rather than a truncated one, and it is the one construction in this collection that escapes the ceiling a resampling of residuals runs into.

Two objects that different, priced on the same table by the same criterion, costing the same to within a standard error, is a statement that the cost is a property of choosing per candidate rather than of what is being chosen. That is a more useful thing to know than either number, because it transports: a practitioner who has a third tuning parameter now has an estimate of what per-candidate choice will cost them, and it is about four thousandths of regret whatever the parameter is.

What the shared rule is, and why it is the default

The shared rule chooses one value from the fullest candidate’s residuals, before any selection. That order is not incidental and it is the reason a shared nuisance is genuinely shared: an estimate made from whichever candidate wins is not shared at all, it is an estimate from the winner, and it carries the selection into the nuisance.

The fullest candidate is also the one whose residuals report the least dependence, which is a collision this collection has already recorded: a fit takes memory out and the fullest fit takes the most. So the shared rule is estimating its nuisance from the most attenuated series available, on purpose, because comparability matters more than accuracy — an error every candidate carries cancels out of every difference and an error only some of them carry does not.

That is the argument the per-candidate rule breaks, and the four thousandths above are what breaking it costs.

The list lengths, which is what the earlier field’s objection was

The objection deserves its own measurement rather than a dismissal, and it gets one two essays along. The short form of the answer is that both halves of it are true and neither is doing anything.

The lists are different lengths, and cutting one to match the other is not a neutral operation: an order list of thirteen values thinned to eight is eight consecutive-ish integers over the same range, and a window list of eight thinned to four is four values out of {0, 1, 2, 4, 8, 12, 20, 30}. The second is a coarser instrument in a way the first is not.

And the cost does not move with the length. Over every length either list can be thinned to, the cost sits between 0.0022 and 0.0042, and the standard errors are 0.0017 to 0.0021. There is no slope to find.

The list is not what separates them. What choosing the tuning parameter separately for every candidate costs, against how many values the rule may choose from, under AR(1) at 0.8. The earlier field's reason for not comparing a window against an order was that their lists are different lengths, and matching them changes very little: at four values the window costs 0.00223 and the order 0.00406; at eight, 0.00401 and 0.00415. Each rule's own list is marked. The two curves are inside each other's standard errors from four values on, and neither has a slope worth the name. At two values the order's cost is exactly zero, because a list of two leaves the five candidates nothing to disagree about — which is the shape of the mechanism and the whole of the list's contribution to it. What the list decides is how often the candidates disagree; what it does not decide is what the disagreement costs.
Fig. 7 The cost against the number of values on the list, for both parameters, with each rule’s own list marked.

What the length does move is the disagreement rate — the order’s runs from nothing at two values to 39% at thirteen, smoothly — and the disagreement rate is not the cost. That gap between a mechanism that responds and an outcome that does not is the whole content of the third essay here.

What a practitioner should do about four thousandths

Very little, and the reason is the same one that runs through this collection’s tuning-parameter results.

Four thousandths of regret sits inside an error the whole procedure is paying about 0.02 of, and inside a decision — which model to report — that is right or wrong far more often for other reasons. Choosing a tuning parameter once for the table is free, obvious, and slightly better, so it is what to do. What it is not is an important saving.

What is worth carrying is the diagnostic. If the candidates disagree about the tuning parameter on four draws in ten, the criteria being compared were computed under four different error models, and the fact that this costs little here is a measurement on this table under these laws. On a table where the candidates differ more — a wider range of dimensions, a stronger dependence, a shorter sample — the disagreement rate is the quantity to look at, and it is free to compute.

Where this sits in the collection

Three fields now price a tuning parameter chosen from the data, and they answer three different questions about it that are easy to run together.

The window field asks what choosing it per candidate costs against choosing it once, and answers four thousandths. This essay asks whether that answer is about the parameter, and answers no. The field after this one asks whether the fit can choose it at all, and answers that nested families buy about what a criterion charges, so the choice is noisy whoever makes it. And the block-length field asks what estimating it costs against the best available value, and answers that the shortfall is several times the quantity the choice was being argued about.

Four questions, four different quantities, one dial. It is worth keeping them apart, because a practitioner who reads “the tuning parameter costs four thousandths” and stops has read the smallest of the four.

A closing note on what the comparison could not have been made without. Both rules are scored on delivered regret against the same baseline — the best model the two rules that need no tuning parameter can produce — and holding that baseline fixed while the rules move is what makes four thousandths a number rather than an artefact.

What is claimed here, and what is not

This essay takes what choosing a sieve order per candidate costs, and how it compares with the same choice about a window. The claims are that the order’s cost at its own list of thirteen values is 0.00360 of delivered regret at 1.7 paired standard errors; that the window’s on the same draws is 0.00401 at 2.3; that matched at eight values apiece they are 0.00401 and 0.00415, and at four 0.00223 and 0.00406, so nowhere is the gap between them larger than either standard error; that the candidates disagree about the order on 39% of draws at a true null with a spread of 1.56, against 42% and 2.15 for the window; and that the same measurement scored without the whitening’s own determinant gives −0.00011 at 0.2 standard errors, which is the reading this collection’s sieve criterion would have produced.

What stays out, and is named as a decision: the covariance estimated per candidate as well. The window field measures that too — 0.01087 against 0.00281, a factor of four — and the order’s version of it would need a per-candidate sieve fit inside a per-candidate order choice, which is a third selection problem and would be priced on a table already carrying two. It is named here and measured nowhere.

Also out: a law with a break. Every number above is under a first-order autoregression, which is the law the earlier field’s figures are quoted at, and a non-stationary law would change what a sieve is estimating rather than what choosing it per candidate costs. Putting a fifth column in would be answering a different question in the same table.

The boundary against the window field is that it prices one tuning parameter and this one asks whether the price is about the parameter. It is not.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutoregressionBandwidthCovariance matrixDegrees of freedomDependenceInformation criterionModel selectionMonte CarloNuisance parameterPaired comparisonRegretSelection effectTaperingWhitening