Fitted together, or fitted after

A window for every candidate

The window and the order a whitening needs are chosen once, from the fullest candidate, on an argument that was made about an estimated covariance. A tuning parameter is not a covariance, and the two cost different amounts.

Worth reading first: A design is a number · Choosing the order.

Two things in this collection are chosen from the sample and then held fixed across a whole table of candidates: the window a whitening uses and the order a fitted autoregression wants. Both are chosen once, from the fullest candidate’s residuals, and both are chosen that way on the strength of an argument the estimated-covariance field made about a covariance: a nuisance estimated per candidate makes the fifteen criteria incomparable, because each is then computed in a different metric.

The argument is right about a covariance. A tuning parameter is a different object, and the argument has never been run on one.

Three ways to attach a window to a table

The confusion is easy to have because a window and a covariance estimate arrive together — the window decides which lags go into Ω̂, so choosing the window per candidate and choosing Ω̂ per candidate look like the same decision. They separate cleanly:

Shared. One window, chosen from the fullest candidate’s residuals; one Ω̂ built from those residuals; the same whitening applied to all fifteen candidates. This is the status quo.

A window each, one Ω̂. The window is chosen from each candidate’s own residuals, and the covariance estimate is still built from the fullest candidate’s. Every criterion is then computed at a different window and from the same residual series.

A window each and an Ω̂ each. Both from each candidate’s own residuals. This is the rule the field already refuses.

The middle column is the one nothing has run, and it is the one the argument was actually about: the objection to the third is that the scale of the criterion moves between candidates, and a shared Ω̂ removes that objection while leaving the searching intact.

One window for the table, or one each. The regret of the same fifteen-candidate table under AR(1) at 0.8 over 150 draws, with the window attached three ways. Chosen once from the fullest candidate's residuals it gives up 0.02518. Chosen from each candidate's own residuals, with the covariance estimate still shared, it gives up 0.02799 — a paired cost of 0.00281 at 2.0 standard errors for the tuning parameter alone. Estimating the covariance per candidate as well costs 0.01087, so the objection already on record is about 3.9 times the size of the one that was not.
Fig. 1 The three ways of attaching a window, with the law on the slider.

Under a first-order autoregression, over a hundred and fifty draws, the three give up 0.02518, 0.02799 and 0.03605 of regret. The paired cost of the tuning parameter alone is 0.00281 at 2.0 standard errors; the paired cost of the covariance estimate as well is 0.01087 at 4.3.

So the objection the field already had is about four times the size of the one it did not — and the one it did not is not zero.

The same split, four times

The split holds in every world and the sizes move.

Under the moving average the three read 0.02421, 0.02693 and 0.03016: the tuning parameter costs 0.00273 at 1.9 standard errors and the covariance 0.00596 at 3.3. Under long memory 0.05135, 0.05219 and 0.05616 — the tuning parameter costs 0.00085 at 1.3, which is nothing, and the covariance 0.00482 at 3.2. Under the break 0.11063, 0.11469 and 0.11914, at 2.8 and 3.0.

The break is the world where the tuning parameter costs most relative to the covariance, and the reason is the reason that world exists: no stationary Ω̂ represents it, so estimating a bad covariance per candidate costs less than usual, and the extra freedom to search over windows costs the same as usual.

The factor of four is one world’s number

“About four times” is the autoregression’s ratio and it is the largest of the four. Written out with the break’s two costs recovered from its columns — 0.11469 − 0.11063 = 0.00406 for the tuning parameter and 0.11914 − 0.11063 = 0.00851 for the covariance as well — the four ratios are:

  • first-order autoregression: 0.01087 / 0.00281 = 3.9
  • five-period moving average: 0.00596 / 0.00273 = 2.2
  • long memory: 0.00482 / 0.00085 = 5.7
  • a break in the persistence: 0.00851 / 0.00406 = 2.1

So the objection the field already had is between two and six times the objection it did not, depending on the world, and the autoregression sits in the middle of that range rather than at the top of it.

Read as shares of the regret each world already carries, the grouping is different and cleaner. The tuning parameter costs 11.2% of the shared column’s regret under the autoregression and 11.3% under the moving average — the two worlds whose dependence has an edge for a window to find — and 1.7% and 3.7% under long memory and the break, where it does not. The covariance estimate’s share falls monotonically across the same four: 43.2%, 24.6%, 9.4%, 7.7%.

The field is measuring at its own resolution

Dividing each cost by its t statistic gives the paired standard error of the comparison, and the window comparison’s is remarkably stable: 0.0014 under the autoregression, 0.0014 under the moving average and 0.0015 under the break, on a hundred and fifty draws. Long memory’s is smaller at 0.00065.

Which puts the detection floor at about 0.0029 in three of the four worlds, against measured costs of 0.00281, 0.00273 and 0.00406. Two of the four readings are within a hair of the floor and the third is one and a half times it. The field is not measuring a comfortable effect; it is measuring one at the edge of what a hundred and fifty draws can see, which is why every t in the window column is between 1.3 and 2.8 and none is above three.

The one outright non-detection is cheap to settle. Long memory’s 0.00085 against a standard error of 0.00065 needs the error down by a factor of 1.5 to reach two, which is 2.4 times the draws — about 360 rather than 150. That is a small enough number to be worth saying, because “costs nothing” and “costs a third of what the autoregression’s costs, measured on too few draws to tell” are different claims and only the second is supported.

The stability of the standard error across worlds is itself informative. The regrets differ by a factor of four and a half across the four laws while the precision of the paired comparison barely moves, so the noise in the comparison is coming from the draw-to-draw variation in which candidate wins rather than from the size of the regret being compared.

Why the middle column is the interesting one

It is worth being careful about what the third column’s cost is made of, because two things are happening in it at once and the middle column separates them.

When each candidate estimates its own Ω̂, two candidates’ criteria are computed on two different whitened samples. The log-determinant term differs, the residual scale differs, and the number that is compared across the table is a number in a different unit for each row. That is the objection the estimated-covariance field states, and it is not about searching at all — it would apply to a rule that estimated Ω̂ per candidate with the window fixed in advance by decree.

When each candidate chooses its own window from a shared Ω̂-building series, the metric is the same for all fifteen — the same residuals, the same taper, the same arithmetic — and what differs is only how many lags each candidate was allowed to search over before reporting its best. That is not a units problem. It is a search, and it is priced the way this collection prices searches.

The two effects are of the same sign, which is why they were never separated: a rule doing both is worse than a rule doing neither, and nothing about that says which half is doing the damage. Holding one fixed while the other moves is the whole experimental design of this essay, and it is available only because the covariance estimate and the window choice can genuinely be decoupled — the window is used twice, once to pick a number and once to build a matrix, and the two uses can be given different series.

How far apart the windows are

How far apart the windows a table wants are. The gap between the largest and the smallest window fifteen candidates ask for, on the same sample, under AR(1) at 0.8 over 150 draws. One window for the table has a spread of zero by construction. Letting each candidate choose its own gives 1.73 lags of spread, which is the size of the thing being chosen: the list runs from 0 to 30. That spread is not a nuisance to be averaged away — it is the whole of the displacement, because a candidate that can move its window further can manufacture more criterion than one that cannot.
Fig. 2 The gap between the largest and the smallest window fifteen candidates ask for, on the same sample.

The mechanism is visible before any regret is computed. Fifteen candidates handed the same residual series and asked to choose from the same list of eight windows do not agree: the largest and the smallest choice differ by 1.7 lags on average under the geometric law, by 2.5 under the moving average, 1.9 under long memory and 1.9 under the break. The list runs from zero to thirty, so a spread of two is a real spread over a short list rather than a rounding.

That spread is not a nuisance to be averaged away. It is the displacement, because a candidate that can move its window further can manufacture more criterion than one that cannot.

What a minimum over a list manufactures

The cleanest measurement of a displacement is at a configuration where nothing can be discovered.

Set every coefficient to zero. Every candidate then contains the truth, all fifteen are equally accurate in population, and any systematic preference among them is arithmetic rather than a finding — the same configuration the displacement between candidates is derived at and the same one a specification search is calibrated at.

What a tuning list manufactures where there is nothing to find. At a true null under AR(1) at 0.8 every candidate contains the truth, so nothing distinguishes them and any systematic preference is arithmetic rather than discovery. Choosing the window from each candidate's own residuals rather than once for the table moves the choice to a larger candidate on 13.5% of 200 draws and to a smaller one on 3.0%. The mechanism is that each candidate's criterion becomes a minimum over 8 windows: that manufactures 19.5 units of criterion on average, and — the part that displaces — it manufactures 2.32 units more for one candidate than for another, where a parameter costs two.
Fig. 3 What choosing the window per candidate does where there is nothing to find.

Choosing the window per candidate moves the choice to a larger candidate on 13.5% of draws and to a smaller one on 3.0%. Under the moving average it is 8.0% against 1.5%, under long memory 10.5% against 5.0%, and under the break 13.0% against 2.0%. Four worlds, four times as often up as down.

The mechanism is a minimum. Each candidate’s criterion becomes the smallest of eight numbers rather than one number, and a minimum over eight is systematically below the one. Averaged over draws that manufactures 19.50 units of criterion under the geometric law — where a parameter costs two — so a minimum over a window list is worth about ten parameters.

That number is large and it is not the number that displaces. A constant amount of manufactured criterion would move all fifteen candidates alike and change nothing. What displaces is the spread: how much more one candidate manufactures than another, which is 2.32 units with a standard error of 0.14, against the 2 a parameter costs.

A per-candidate window is worth slightly more than a free parameter, handed unequally. That is exactly the shape of an unpriced search, and it is why the choice belongs where it is.

Why the effect is small, and why that is not a reason to stop

The costs above — 0.00281 of regret, one extra parameter’s worth of manufactured criterion — are small beside what the field’s other quantities are worth. Not knowing the dependence at all costs 0.063 under this law. Estimating the covariance per candidate costs four times as much as tuning per candidate. And getting the coefficient itself right is worth 0.0022.

The reason to have measured it anyway is that the alternative was an argument. A rule adopted on the strength of a claim about a different object is a rule that is right for a reason nobody checked, and the failure mode of that is not being wrong — it is being right until the object changes. When the joint fit arrives and the nuisance stops being estimated from residuals at all, the argument about metrics stops applying and the argument about searching does not. Only one of the two was ever a measurement.

What the dimension does

Regret is a summary and the mechanism is visible one level below it, in which model each rule picks.

Averaged over draws under the geometric law, the shared rule selects a candidate of 3.75 coefficients, the per-candidate-window rule 3.64, and the rule that also estimates Ω̂ per candidate 3.42. So on the world with real coefficients in it, the searching rules pick smaller models — the opposite of what the null configuration shows.

That is not a contradiction and it is worth a sentence, because a reader who has seen only one of the two numbers will draw the wrong conclusion from it. At a true null every candidate is equally accurate and the displacement is the whole of the signal, so it shows up cleanly as a preference for larger models. In a world with a true model in it there is also a real signal, and a per-candidate window degrades the whitening of the candidates that are ill-fitting — which are the large ones, since a large candidate’s residuals are the ones a window is most easily fooled by. The displacement pushes up and the degradation pushes down, and on this design the second is larger.

A displacement measured in a world with signal in it is a displacement plus something else, which is exactly why the null configuration is where it is measured. That is a discipline this collection applies everywhere — the corner a test is calibrated at is the same argument for a size — and it is what stops a measurement of one thing from being reported as a measurement of two.

A selection inside a selection

The general shape is worth naming because this collection meets it repeatedly and under different names.

A model-selection rule picks the smallest of fifteen criteria. A tuned model-selection rule picks the smallest of fifteen numbers, each of which is itself the smallest of eight. The second is a selection over one hundred and twenty things wearing the clothes of a selection over fifteen, and the correction for it is not a correction to a penalty — it is a change to what the penalty is a penalty for.

The displacement a search creates is about the number of comparisons; the optimism a criterion pays for is about the number of parameters; and a tuning list is a third thing that is neither, because the eight windows are not eight candidates and not eight parameters. What they are is eight chances, and the price of a chance is measured here rather than derived.

The same argument applies to the order, and it is why the order is chosen once too. It is not measured here: the order list runs to twelve rather than eight, so the manufacturing is larger, and the honest thing to say is that the direction is the same and the number is not this one.

One further asymmetry is worth recording, because it is the reason this cannot be repaired by simply enlarging the penalty. A displacement that were the same for every candidate would be absorbed by any constant added to the criterion, and this one is not: the spread across candidates is 2.32 units under the geometric law, 1.38 under the moving average, 2.62 under long memory and 2.45 under the break, so the size of the correction a table needs is a function of the world the table is fitted in. A correction that has to be estimated from the same sample is a third selection problem, and pricing it is what the essay on what a search costs does for a break point.

What is claimed here, and what is not

This essay takes what a tuning parameter chosen per candidate costs, separately from what a covariance estimated per candidate costs. The claims are that the two are separable, by holding the covariance estimate fixed while letting the window move; that on this table the tuning parameter costs 0.00281 of regret at 2.0 paired standard errors and the covariance estimate 0.01087 at 4.3, so the objection the field already had is four times the one it did not; that fifteen candidates handed the same residuals choose windows 1.7 lags apart; and that at a true null a per-candidate window picks a larger candidate four times as often as a smaller one, manufacturing 19.50 units of criterion of which 2.32 are handed unequally.

What stays out, and is named as a decision: the order. The same three columns exist for the sieve’s order and are not run here, because the honest comparison needs the order list and the window list to be the same length before the two numbers can be put beside each other, and making them the same length changes what each rule is. It is a measurement rather than a difficulty, and it is named rather than done.

The boundary against the field that estimates the covariance is that it establishes that the nuisance must be shared and this one asks which part of the nuisance. The answer is both, at four to one.

The checks, and the refusals that make them mean something

Three claims are gated. The shared column is required to have a window spread of exactly zero and the per-candidate columns a positive one, which is what says the three rules are the three rules. The displacement is required to move the chosen dimension upwards at a true null, where nothing can be discovered. And the manufactured criterion’s spread across candidates is required to exceed what a parameter costs, because a constant manufacture would displace nothing.

The refusal is the construction itself, offered as a rule. A criterion minimised over a tuning list and then compared across candidates is refused, with the manufactured spread printed beside the two units a parameter costs. What makes it inadmissible is not that it is worse — it is that the number it reports is a minimum over a set whose size differs from the number of things being compared.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A rate times a size — both name bandwidth selection, information criterion, long-run variance, model selection, regret, selection effect, tuning parameter, whitening
  • A width that moves and an error that does not — both name bandwidth selection, information criterion, model selection, optimism, overfitting, regret, tapering, tuning parameter
  • How often it matters — both name bandwidth selection, information criterion, long-run variance, model selection, regret, selection effect, tuning parameter, whitening
  • The eighth that was not a constant — both name bandwidth selection, data snooping, information criterion, model selection, overfitting, regret, selection effect, tuning parameter
  • A charge that reads the draw — both name bandwidth selection, information criterion, model selection, optimism, regret, selection effect, tapering
  • A step that is not a ratio — both name bandwidth selection, information criterion, model selection, overfitting, regret, selection effect, tuning parameter

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionData snoopingInformation criterionLong-run varianceModel selectionNuisance parameterOptimismOverfittingRegretSelection effectSpecification searchTaperingTuning parameterWhitening