What decides whether a tuning list decides

A step that is not a ratio

Run the separation sweep on a tuning list of integers rather than a geometric ladder and the two factors still point opposite ways. The invariant does not survive: along a row of integers the probability moves by 2.163 where along the geometric ladder it moves by 1.208.

Worth reading first: A design is a number · Choosing the order.

Everything this field has measured has been measured on one tuning parameter: the width of the band that whitens a candidate’s residuals, chosen from a list of eight values. That list is a geometric ladder — 0, 1, 2, 4, 8, 12, 20, 30 — and its steps run from one lag to ten.

The collection contains a second tuning parameter that does the same job by a different route, and its list is an interval of integers. An autoregression of order four and one of order five are different models and there is no coarser place to put a value, so the sieve order is chosen from 0 through 12 and every step is the same size.

Two questions follow, and only the first is the obvious one. Does the separation sweep behave the same way when the step is additive? And does the invariant the sweep was read against survive?

What a step is, on each list

The two lists differ in a way that can be stated before any data are drawn.

The window’s steps are 1, 1, 2, 4, 4, 8, 10 lags, a factor of 10.00 between the smallest and the largest. The order’s natural steps are all 1, a factor of 1.00. Fitting the size of a step against the value it is taken from, on logarithms, gives 0.7422 for the window and 0.0000 for the order — where one is a step proportional to where it is taken and nought is a step of the same size everywhere.

One list steps by a ratio, the other by one. How large each step is on the two tuning lists this collection chooses from, both cut to 8 values. The whitening window's list is a geometric ladder — 0, 1, 2, 4, 8, 12, 20, 30 — whose steps run from 1.0 to 10.0 lags, a factor of 10.00. The sieve order's list is an interval of integers — 0, 2, 3, 5, 7, 9, 10, 12 — whose steps are 1 or 2, a factor of 2.00. Fitting the step against the value it starts from on logarithms gives 0.742 for the window and 0.134 for the order, where one is a step proportional to where it is taken and nought is a step of the same size everywhere. The two rules therefore change a candidate's criterion by different amounts, and no choice of values makes them the same.
Fig. 1 The size of every step on the two tuning lists, both cut to eight values. The window’s steps grow with the value they start from; the order’s do not.

That is the whole difference, and it has a consequence the sweep has to be arranged around. A geometric ladder is coarse where its values are large and fine where they are small; an interval of integers is uniformly fine. Two candidates that would pick 8 and 12 lags on the ladder are two steps apart; two that would pick order 4 and order 5 are one step apart and are, in every sense that matters to a criterion, closer together.

What is held fixed, and what cannot be

The comparison is made at a matched list length of eight, so neither rule is minimising over more chances than the other. That confound is not hypothetical: it is the whole finding of the field that matched the two lists on a different quantity, which reports the list doing more of the work than the parameter.

The table, the law, the sample size and the seeds are also held. Cutting the order’s thirteen values to eight is done through the positions of the list rather than its values, so the endpoints are kept and the rule can still reach order twelve; what it gives up is resolution. The result is 0, 2, 3, 5, 7, 9, 10, 12 — steps of one or two, a factor of 2.00 rather than the window’s ten.

What cannot be held is the size of a step. An integer step and a geometric step are different amounts of change, and no choice of values makes them the same, because the two parameters do not measure in the same units and there is no exchange rate between a lag and an order. Any comparison stated in lags or orders is comparing quantities that are not comparable, and saying which is larger would be an artefact of the choice.

So the comparison is not made there.

The measure that is the same on both lists

This field already carries an instrument built for exactly this problem, and it was built for a different reason.

On each draw, how far a change of tuning parameter moves a candidate’s criterion is computed, and the count of candidates lying within one tuning change’s worth of criterion of the winner is recorded. It is a ratio of two quantities that both scale with the size of a step — what a step is worth, and how far apart the criterion puts the candidates — so it is the same statistic on a ladder and on an interval. That is what makes it the place the two parameters can be put beside each other.

Run it on both. At the standing coefficients the criterion spans 20.60 across the table on the order and 21.04 on the window — within a few per cent of each other, which is a coincidence rather than a construction and is what makes the next number readable. A change of order is worth 0.666 of criterion; a change of window is worth 1.857.

So an integer step, on tables the criterion spans equally, is worth about a third of a geometric step. The order’s list is finer in its own units and finer in the criterion’s, and the two facts do not have to agree — a coarse list of a parameter the criterion cares about would be the other way round.

That single number carries the rest of the essay. Fewer candidates are within reach on the order: 1.433 against the window’s 2.368 at the standing coefficients, and 1.900 against 3.635 at the near end of the sweep.

The measure that is the same on both lists. How many of the fifteen candidates sit within one tuning change's worth of criterion of the winner, at each separation, for each of the two tuning parameters. It is the one reading a geometric ladder and an interval of integers can both be scored on, because it is a ratio of two quantities that both grow with the size of a step: what a change of tuning parameter is worth on this draw, against how far apart the criterion puts the candidates. On the order it falls from 1.900 to 1.305; on the window, from 3.635 to 1.660. A change of order is worth 0.666 of criterion against the window's 1.857 at the standing coefficients, on spans of 20.60 and 21.04.
Fig. 2 How many candidates sit within one tuning change’s reach of the winner, on both parameters, across the separation. It is the one reading a ladder and an interval can both be scored on.

The two factors still point opposite ways

With the instrument settled, the sweep reproduces.

As the candidates are pulled apart, the order’s candidates disagree about the tuning parameter more often — 31.0%, 34.8%, 41.9%, 56.8% and 72.5% of draws across the five separations, a rise of 2.34. And the share of those disagreements that changes which candidate the table selects falls: 42.3%, 33.8%, 22.1%, 6.4% and 3.6%, a fall of 11.69.

On the window the same two columns rise by 2.80 and fall by 30.52. Both parameters, both directions, same ordering.

Opposite ways, on an additive step too. The two factors the cost of a per-candidate tuning parameter is a product of, as the candidates are pulled apart, with the tuning parameter the sieve order — a list of 8 integers rather than a geometric ladder — over 800 draws at each of 5 separations. How often the candidates disagree rises from 31.0% to 72.5%; the share of those disagreements that changes the winner falls from 42.3% to 3.6%. On the window the same two readings are 31.8% to 88.8% and 51.6% to 1.7%. The direction is a property of the separation and not of the list's spacing.
Fig. 3 The two factors across the separation, with the tuning parameter the sieve order. The rising line is how often the candidates disagree; the falling one is how often the disagreement decides.

So the direction is a property of the separation and not of the list’s spacing, which is the answer to the first question and is the one the essay that found the two factors would have predicted.

The magnitudes are not the same, and they fail to be in the direction the reach measurement already implied. The order’s fall is a factor of 11.69 against the window’s 30.52, and its rise is 2.34 against 2.80 — both movements smaller, the one that decides the product smaller by more. A tuning change that is worth a third as much cannot tip as many verdicts at either end of the sweep.

The product follows. The probability that the list changes the winner runs 13.1%, 11.8%, 9.3%, 3.6% and 2.6% on the order, a span of 5.00, against the window’s span of 11.75.

The same dial, on a list that steps by one. The probability that a per-candidate tuning list changes which candidate the table selects, at each of 5 separations, for both tuning parameters at a matched list length of 8. The sieve order runs from 13.1% to 2.6%, a factor of 5.00; the whitening window, from 16.4% to 1.5%, a factor of 11.75. What is held is the number of options, the table, the law, the sample size and the seeds; what cannot be held is the size of a step, since an integer step and a geometric step are different amounts of change. The dial moves both, and it moves them by 2.35 times as much on one as on the other.
Fig. 4 The probability that a per-candidate tuning list changes the winner, on both parameters at a matched list length, across five separations between the candidates.

The leverage arithmetic, which accounts for some of it

The field’s own explanation of why the decisive share falls is a ratio: the criterion’s span across the table against what a change of tuning parameter is worth. It is worth checking whether that ratio accounts for the difference between the two parameters, because if it does the essay is a re-run and if it does not there is something else here.

Across the separation the order’s criterion span grows from 5.99 to 162.14, a factor of 27.1, while what a change of order is worth grows from 0.460 to 1.220, a factor of 2.65. So the criterion’s grip on the table grows by about ten. On the window the same two readings are a factor of 24.5 against 2.03, a grip growing by about twelve.

Those are close, and the decisive shares are not: 11.69 against 30.52. The leverage ratio moves almost identically on the two parameters and the turnover falls two and a half times as far on one of them, so the ratio is not the whole account.

What the ratio leaves out is the level it starts from. At the near end the order already has only 1.900 candidates within a tuning change’s reach where the window has 3.635, so the order’s turnover starts at 42.3% against 51.6% and has less distance to fall through before it runs out of candidates to reach. By the far end both are near one — 1.305 and 1.660 — and both turnovers are small. The fall is bounded by where it began, and where it began is set by what a step is worth.

That is a smaller claim than the field’s mechanism and it is the one the measurement supports. The ratio says which direction the share moves; it does not price how far.

The near end, which is the control

The separation of zero deserves reading on its own, because it is where every candidate is correctly specified and whatever the criterion does there is what it does when there is nothing to find.

On the order the candidates disagree about the tuning parameter on 31.0% of draws — almost exactly the window’s 31.8% — and 42.3% of those disagreements change the winner against the window’s 51.6%. So on 13.1% of draws, one in eight, the model a reader is handed depends on whether the sieve order was chosen once for the table or once per candidate, in a world where the answer means nothing either way.

The agreement between the two disagreement rates at that cell is worth a sentence, because it is the one place the two lists behave identically. At a separation of zero every candidate’s residuals are the same error process plus estimation noise, so what the candidates are choosing between is noise, and a list of eight values samples noise about equally well whether its steps are ratios or integers. It is only once the residual series genuinely differ that the spacing of the list starts to matter — which is the same reading, arriving from the near end, as the section above.

The invariant does not survive

The second question is the one that goes the other way, and it is the reason this essay exists rather than confirming something.

The essay that read the probability against list length reports it flat: about an eighth, whatever the list length, and the grid that re-ran it on three tables finds the same thing on every row. The window’s probability moves along a row by factors of 1.158, 1.087 and 1.208 on the three tables, against 1.981 down a column. The list length is the dial that does nothing; the table is a dial that does.

On the order it is the other way round. Along a row the probability moves by 2.163, 1.921 and 1.960; down a column, by at most 1.898. The list length moves the answer more than the table does.

The invariant is the ladder's, not the problem's. The probability that a per-candidate tuning list changes the winner, at three list lengths on three candidate tables, over 800 draws in each of the nine cells — with the tuning parameter the sieve order rather than the whitening window. Along a row it moves by a factor of up to 2.16; down a column, by up to 1.90. On the whitening window the same two readings are 1.21 and 1.98, so there the list length is the dial that does nothing and the table is the dial that does. Here they change places. The earlier reading was taken on a geometric ladder, where thinning a list keeps its reach and loses resolution only at the wide end where the criterion is already flat; thinning an interval of integers loses resolution everywhere, and the candidates stop disagreeing because they are being handed the same value.
Fig. 5 Every table at every list length, with the tuning parameter the sieve order. Along a row the probability is not flat, which is the reverse of what the same grid shows on the window.

The numbers behind the worst row are worth setting out. On the fifteen-subset table the order’s probability reads 5.4% at four values, 9.3% at eight and 11.6% at thirteen. On the window’s grid the same three cells read 14.6%, 12.6% and 13.1%.

So the invariant is the ladder’s, not the problem’s. That is a real qualification of a result this field has been reading as a property of tuning lists in general, and it was invisible for the reason most such things are: it was measured once, on the parameter that happened to be in hand.

Flat along a row, apart between them. The probability that a per-candidate tuning list changes the winner, at three list lengths on three candidate tables, over 800 draws in each of the nine cells. Along a row — the reading the earlier field takes — it moves by a factor of at most 1.21, so that field's invariant survives on every table. Down a column it moves by up to 1.98. The list length is the dial that does not move this number and the table is one that does, and the earlier field varied only the first.
Fig. 6 The same grid on the whitening window, in the essay that established the invariant, where every row is flat.

Why thinning does different things to the two lists

The mechanism is in the first section, and it needs the decomposition to be visible.

Split the order’s movement along a row into its two factors. The disagreement rate reads 16.5% at four values, 41.9% at eight and 50.1% at thirteen — a factor of three. The share of disagreements that change the winner reads 32.6%, 22.1% and 23.2% — roughly flat, and not ordered. So the whole of the movement is in the rate, and the rate is a count of how often two candidates pick different values from the list.

Now ask what thinning does. Cutting a geometric ladder from eight values to four gives 0, 2, 12, 30: the reach is kept and the resolution is lost at the wide end, where a candidate that wanted 20 lags and one that wanted 30 both get 30 and where the criterion was nearly flat anyway. Few disagreements are removed. Cutting an interval of integers from thirteen values to four gives 0, 4, 8, 12: two candidates that wanted orders 5 and 6 now both get 4. Resolution is lost everywhere, including where the candidates actually differ, so the disagreements are removed wholesale.

That is why the order’s rate collapses by a factor of three across its own list and the window’s does not, and it follows from the step statistics rather than from anything about sieves. A list whose steps span a factor of ten has somewhere cheap to lose resolution; a list whose steps are all the same size does not.

The natural lists agree; the matched ones do not

There is one more reading on the table, and it repeats a lesson this line of fields has already paid for once.

At its own natural list length the order chooses from thirteen values, and the probability that its list changes the winner is 11.6%. The window at its own natural length of eight reads 13.1%. Those are close enough that a reader comparing the two parameters as they are actually run would conclude they behave the same.

At a matched list length of eight the same two numbers are 9.3% and 13.1% — the order forty per cent below the window.

The difference between the two comparisons is the list, not the parameter, which is exactly what the field that matched the two lists found for the cost of a per-candidate rule and what the field that put the two criteria side by side found when a difference between two tuning parameters turned out to be a difference between two criteria. Three fields, three quantities, one confound.

Which comparison a reader wants depends on the question. Matched answers “does the spacing of the list change how the sweep behaves”, which is this essay’s question. Natural answers “does it matter which of the two tuning parameters is being run”, which is a practitioner’s, and there the answer is that it does not matter much — and that the agreement is built out of two disagreements cancelling.

What the three tables say on the order

The table dial was the subject of the earlier grid and it is worth checking it has not changed character, because a finding about the list would be weaker if the table had stopped behaving.

On the order at eight values, the probability that the list changes the winner reads 9.3% on the fifteen subsets, 7.1% on the six pairs and 5.6% on the nested ladder — the same ordering the window gives, and a span of 1.64. The criterion’s spans are 20.60, 15.03 and 13.04, and a tuning change is worth 0.666, 0.628 and 0.421.

So the mechanism is intact: the nested ladder’s candidates are close together, a change of tuning parameter is worth less there than anywhere, and it is worth less by more than the span falls by. What has changed is only that the table is no longer the larger of the two dials.

How many draws each of these is measured on

The two ends of the sweep are not equally well resolved, and the imbalance runs the same way it does on the window.

The disagreement rate is a proportion over every draw, so at eight hundred draws it is known to about 1.8 percentage points anywhere on the sweep. The decisive share is a proportion over the disagreeing draws only, and there are 248 of those at a separation of zero rising to 580 at four. That looks like the far end being better resolved and it is the reverse: a share of 3.6% on 580 disagreements is twenty-one events, and twenty-one events fix a share to about eight tenths of a point in its own units.

So the far end is fine for saying the share is small and useless for saying whether it is 3.6% or 5.0%. Every claim above about the far end is a claim about an order of magnitude, and the claims that carry weight — the direction of both factors, the failure of the invariant — are made where the counts are large.

That ordering is the same one the field that priced the two averages measured directly, and it is why this field reports the probability rather than the conditional size: the probability is a product of two proportions and the size is an average of a heavy-tailed quantity over whichever draws happened to disagree.

How many candidates a tuning change can reach. How many of the fifteen candidates sit within one tuning change's worth of criterion of the winner, on the draw's own scale, at each separation. It is the measurement behind the word "indifferent": on each draw, how far a change of tuning parameter moves a candidate's criterion is computed and compared against how far apart the candidates are. It falls from 3.635 at the near end to 1.560 at twice the standing coefficients, and the criterion's own span across the table grows from 6.51 to 159.79 over the same sweep. One means the winner is out of reach of everything else and a tuning list cannot change it.
Fig. 7 The same reach measurement on the window alone, in the essay that introduced it, where the fall across the sweep is steeper.

What stays out

Three things this essay names and does not measure.

Whether the invariant fails for the ladder at a finer spacing. The diagnosis above says the window’s flatness comes from having somewhere cheap to lose resolution. A ladder with twice as many rungs over the same reach would have less of that, and the prediction is that its probability would start moving along a row. Nothing here builds that list, and the diagnosis stands or falls on it.

What the criterion’s spans agreeing to two per cent is. The order spans 20.60 across the fifteen-subset table and the window 21.04, which is what makes the two tuning changes directly comparable and which nothing here arranged. Whether that is a property of the table, of the law, or an accident of one setting is unmeasured, and the comparison of 0.666 against 1.857 is only as safe as it is.

And the second law. Every cell above is a first-order autoregression. The sieve order is the parameter whose natural list is longest and whose criterion carries a determinant term, and the field that showed what that term does works in one law as well. A dial that reverses which of two others is larger is exactly the kind of finding that should be re-run somewhere else before it is carried anywhere.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • How often it matters — both name bandwidth selection, granularity, information criterion, model selection, monte carlo, regret, selection effect, tuning parameter
  • The quarrel that changes the winner — both name bandwidth selection, granularity, information criterion, model selection, overfitting, regret, selection effect, tuning parameter
  • A charge that reads the draw — both name bandwidth selection, dependence, information criterion, model selection, monte carlo, regret, selection effect
  • A line that beats two curves — both name bandwidth selection, dependence, information criterion, model selection, monte carlo, nested models, overfitting
  • A width that moves and an error that does not — both name bandwidth selection, information criterion, model selection, nested models, overfitting, regret, tuning parameter
  • A window for every candidate — both name bandwidth selection, information criterion, model selection, overfitting, regret, selection effect, tuning parameter

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionConditional distributionDependenceEffect sizeGranularityInformation criterionModel selectionMonte CarloNested modelsOrder-selectionOverfittingRegretSelection effectSubset of parametersTuning parameter