The comparison that was not made
Worth reading first: A design is a number · The observations that repeat each other.
Choosing the whitening’s window separately for each candidate costs 0.00281 of delivered regret against choosing it once for the whole table, at two paired standard errors, and 0.01087 when the covariance is estimated per candidate as well. That measurement ends with a sentence naming what it could not do:
The window’s is measured and the order’s is not, because the order list runs to twelve where the window list runs to eight, and making the two the same length changes what each rule is.
The reason is a good one. A rule minimising a criterion over thirteen values and a rule minimising it over eight are not the same rule, so a difference between them would not be a difference between a window and an order. The direction was predicted — the longer list should cost more — and the size was not offered.
Both halves of that turn out to be wrong, and the second one interestingly.
The missing number
Choosing a sieve order per candidate rather than once, over its own list of thirteen values, costs 0.00360 of regret at 1.7 paired standard errors. On the same draws, in the same sweep, the window’s own figure is 0.00401 at 2.3.
They are the same size.
Matched at eight values apiece the window costs 0.00401 and the order 0.00415, which is closer still and in the opposite order. Matched at four they are 0.00223 and 0.00406. Nowhere in the sweep is the difference between the two parameters larger than the standard error of either.
So the answer to the question the earlier field asked is that there is nothing to see: a window and an order cost the same when they are chosen per candidate, whatever length of list either is chosen from.
What that costs to say honestly
The number above is not the number this measurement produced first, and the difference between the two versions is the whole of the next essay.
Scored the way this collection has scored a sieve since the estimated-covariance field, a per-candidate order costs −0.00011 at 0.2 standard errors — nothing, and slightly the wrong sign. That reading is what the machinery gave for a day, it is entirely consistent with the earlier field’s worry running the other way, and it is wrong because the two rules were not being scored by the same criterion. The term that was missing is the volume the whitening moves, the window’s rule has carried it since it was written, and the sieve’s never had it.
The point worth carrying from that here is narrower than the repair. A comparison between two procedures is a comparison between two implementations of two criteria, and the second pair is the one that has to be checked. Two rules that differ in a term neither prose describes will differ by whatever that term is worth, and it will look like a difference between the rules.
One of the two rules cares about its list and the other does not
The sweep prints each rule at three list lengths, and reading down the columns rather than across them finds an asymmetry the matched comparison was designed to remove.
The window’s cost roughly doubles with its list. Four values cost 0.00223 and eight cost 0.00401.
The order’s does not move. Four values cost 0.00406, eight cost 0.00415 and thirteen cost 0.00360 — a range of fifteen per cent, in a quantity whose own standard error is about 0.0021, which is to say no movement at all.
So the earlier field’s prediction was not merely wrong about the order; it was wrong about the rule it applies to. A longer list does cost more, and it costs more for the window, whose list is the shorter of the two and whose ladder is geometric — the extra values it adds are in the middle, where they let two candidates land closer together and quarrel more cheaply. The order’s list is a dense run of integers over the same range at every length, so lengthening it adds resolution rather than reach, and resolution is not what a per-candidate rule is spending.
That also explains why matching the two at eight values makes them closer rather than clarifying anything: matching moves the window up its own curve and leaves the order where it was.
How close “the same size” is
The largest gap anywhere in the sweep is at four values on the list, where the window costs 0.00223 and the order 0.00406 — a difference of 0.00183.
Set against the rules’ own standard errors, which are about 0.0021 for the order and 0.0017 for the window, that gap is comparable to either. The standard error of the difference is larger than either of those, so the gap is inside it by a wider margin than the comparison of point estimates suggests.
At the two other list lengths the gaps are smaller still: 0.00041 at the rules’ own lists and 0.00014 matched at eight — the second being three ten-thousandths of a regret, or eight per cent of one standard error.
Which is worth saying plainly, because “the same size” can mean two rather different things. This is not two effects that happen to round to the same figure. It is a difference that the measurement cannot distinguish from zero at any list length, and which is smallest exactly where the two rules are made most comparable.
What choosing per candidate does
Both rules cost about four thousandths of regret, on a baseline where the best model available is worth about 0.02 more than the worst rule delivers. That is a fifth of the whole quantity the field is about, and it is worth being clear about which fifth.
A candidate’s criterion under a shared nuisance is a fit made under a common error model. Under a per-candidate nuisance it is a fit made under its own error model, and the table then ranks fits that were computed under different assumptions. The two things that go wrong are different sizes:
The comparison stops being like for like. Each candidate’s score now carries its own whitening, and the determinant that says how much volume that whitening moved does not cancel between candidates. That is a bias, it is computable, and it is what the missing term was for.
And each candidate’s criterion becomes a minimum over the list. A minimum over eight numbers is below a typical one, by an amount that differs between candidates, and the difference is what displaces the choice. The window field measures that displacement at a true null: a per-candidate window picks a larger candidate on 13.5% of draws against 3.0% smaller.
The candidates disagree about what to use
The mechanism is visible without any criterion at all, in what the rules pick.
Over five candidates and a list of thirteen orders, the choices are not all equal on 39% of draws at a true null, and the spread between the largest and smallest chosen order averages 1.56. For the window at its own list of eight the disagreement rate is 42% and the spread 2.15.
Those are the same rates, from two lists of different lengths, for two parameters that are not alike in anything except being tuning parameters. Which is the first hint that the length of the list is not the acting quantity — a hint the next essay but one turns into a measurement.
The two rules, written out
It is worth setting the two constructions side by side, because everything above depends on them being the same construction with one substitution.
The window rule. From a residual series, take the sample autocovariances out to lags, weight each by the Bartlett triangle so the result is positive definite, factor the resulting banded Toeplitz matrix, and whiten with the factor. The value of is chosen by a criterion on the residuals themselves: how white the whitened series is, charged two per lag spent.
The order rule. From the same residual series, fit an autoregression of order by the Yule–Walker recursion, form the covariance matrix that autoregression implies, factor it, and whiten with the factor. The value of is chosen by a criterion on the residuals: the fitted innovation variance, charged two per coefficient.
They differ in one place — what the sequence past the chosen lag looks like. A band says zero and a sieve says an exponential tail. That difference is what makes the sieve escape a ceiling the band sits under, and it is the reason the two are worth comparing at all rather than being two names for one thing.
Everything else — the candidate table, the sample, the criterion the candidate is scored by, the baseline the regret is measured against — is identical, and identical by construction rather than by description: both go through the same sweep with a scorer swapped.
Why the two costs being equal is a result
It would be easy to file a null result and move on, and it is worth saying why this one is worth the field.
The two tuning parameters are not alike. A window is a bandwidth: it says how many lags of the sample’s own autocovariance sequence to keep, it enters through a tapered band, and widening it moves a determinant a long way. An order is a model dimension: it says how many coefficients of an autoregression to fit, it enters through a whitening that has an exponential tail rather than a truncated one, and it is the one construction in this collection that escapes the ceiling a resampling of residuals runs into.
Two objects that different, priced on the same table by the same criterion, costing the same to within a standard error, is a statement that the cost is a property of choosing per candidate rather than of what is being chosen. That is a more useful thing to know than either number, because it transports: a practitioner who has a third tuning parameter now has an estimate of what per-candidate choice will cost them, and it is about four thousandths of regret whatever the parameter is.
What the shared rule is, and why it is the default
The shared rule chooses one value from the fullest candidate’s residuals, before any selection. That order is not incidental and it is the reason a shared nuisance is genuinely shared: an estimate made from whichever candidate wins is not shared at all, it is an estimate from the winner, and it carries the selection into the nuisance.
The fullest candidate is also the one whose residuals report the least dependence, which is a collision this collection has already recorded: a fit takes memory out and the fullest fit takes the most. So the shared rule is estimating its nuisance from the most attenuated series available, on purpose, because comparability matters more than accuracy — an error every candidate carries cancels out of every difference and an error only some of them carry does not.
That is the argument the per-candidate rule breaks, and the four thousandths above are what breaking it costs.
The list lengths, which is what the earlier field’s objection was
The objection deserves its own measurement rather than a dismissal, and it gets one two essays along. The short form of the answer is that both halves of it are true and neither is doing anything.
The lists are different lengths, and cutting one to match the other is not a neutral operation: an order list of thirteen values thinned to eight is eight consecutive-ish integers over the same range, and a window list of eight thinned to four is four values out of {0, 1, 2, 4, 8, 12, 20, 30}. The second is a coarser instrument in a way the first is not.
And the cost does not move with the length. Over every length either list can be thinned to, the cost sits between 0.0022 and 0.0042, and the standard errors are 0.0017 to 0.0021. There is no slope to find.
What the length does move is the disagreement rate — the order’s runs from nothing at two values to 39% at thirteen, smoothly — and the disagreement rate is not the cost. That gap between a mechanism that responds and an outcome that does not is the whole content of the third essay here.
What a practitioner should do about four thousandths
Very little, and the reason is the same one that runs through this collection’s tuning-parameter results.
Four thousandths of regret sits inside an error the whole procedure is paying about 0.02 of, and inside a decision — which model to report — that is right or wrong far more often for other reasons. Choosing a tuning parameter once for the table is free, obvious, and slightly better, so it is what to do. What it is not is an important saving.
What is worth carrying is the diagnostic. If the candidates disagree about the tuning parameter on four draws in ten, the criteria being compared were computed under four different error models, and the fact that this costs little here is a measurement on this table under these laws. On a table where the candidates differ more — a wider range of dimensions, a stronger dependence, a shorter sample — the disagreement rate is the quantity to look at, and it is free to compute.
Where this sits in the collection
Three fields now price a tuning parameter chosen from the data, and they answer three different questions about it that are easy to run together.
The window field asks what choosing it per candidate costs against choosing it once, and answers four thousandths. This essay asks whether that answer is about the parameter, and answers no. The field after this one asks whether the fit can choose it at all, and answers that nested families buy about what a criterion charges, so the choice is noisy whoever makes it. And the block-length field asks what estimating it costs against the best available value, and answers that the shortfall is several times the quantity the choice was being argued about.
Four questions, four different quantities, one dial. It is worth keeping them apart, because a practitioner who reads “the tuning parameter costs four thousandths” and stops has read the smallest of the four.
A closing note on what the comparison could not have been made without. Both rules are scored on delivered regret against the same baseline — the best model the two rules that need no tuning parameter can produce — and holding that baseline fixed while the rules move is what makes four thousandths a number rather than an artefact.
What is claimed here, and what is not
This essay takes what choosing a sieve order per candidate costs, and how it compares with the same choice about a window. The claims are that the order’s cost at its own list of thirteen values is 0.00360 of delivered regret at 1.7 paired standard errors; that the window’s on the same draws is 0.00401 at 2.3; that matched at eight values apiece they are 0.00401 and 0.00415, and at four 0.00223 and 0.00406, so nowhere is the gap between them larger than either standard error; that the candidates disagree about the order on 39% of draws at a true null with a spread of 1.56, against 42% and 2.15 for the window; and that the same measurement scored without the whitening’s own determinant gives −0.00011 at 0.2 standard errors, which is the reading this collection’s sieve criterion would have produced.
What stays out, and is named as a decision: the covariance estimated per candidate as well. The window field measures that too — 0.01087 against 0.00281, a factor of four — and the order’s version of it would need a per-candidate sieve fit inside a per-candidate order choice, which is a third selection problem and would be priced on a table already carrying two. It is named here and measured nowhere.
Also out: a law with a break. Every number above is under a first-order autoregression, which is the law the earlier field’s figures are quoted at, and a non-stationary law would change what a sieve is estimating rather than what choosing it per candidate costs. Putting a fifth column in would be answering a different question in the same table.
The boundary against the window field is that it prices one tuning parameter and this one asks whether the price is about the parameter. It is not.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What fitting them together buys — both name covariance matrix, dependence, model selection, monte carlo, nuisance parameter, paired comparison, tapering, whitening
- A charge that reads the draw — both name dependence, information criterion, model selection, monte carlo, regret, selection effect, tapering
- A dependence with a shape — both name autoregression, covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
- A width that moves and an error that does not — both name bandwidth, covariance matrix, degrees of freedom, information criterion, model selection, regret, tapering
- The charge nobody derived — both name bandwidth, covariance matrix, degrees of freedom, information criterion, model selection, nuisance parameter, tapering
- The window that has to be chosen, and the term that was dropped — both name covariance matrix, degrees of freedom, dependence, information criterion, model selection, nuisance parameter, tapering
Named objects
A flat tag is an object no other essay names yet.
AutoregressionBandwidthCovariance matrixDegrees of freedomDependenceInformation criterionModel selectionMonte CarloNuisance parameterPaired comparisonRegretSelection effectTaperingWhitening