Two factors pointing opposite ways
Worth reading first: A design is a number · Choosing the order.
The probability that a per-candidate tuning list changes which candidate a table selects is a product of two things: how often the candidates disagree about the tuning parameter, and how often a disagreement changes the winner.
The previous essay shows the product running from 17.6% to 1.5% as the candidates are pulled apart. This one is about the two factors, because they do not move together and one of them is the one a sweep can see.
The two columns
Across separations of 0, 0.5, 1, 2 and 4 times the standing coefficient vector, over eight hundred draws apiece:
| separation | they disagree | a disagreement decides | the product |
|---|---|---|---|
| nothing omitted | 31.8% | 51.6% | 16.4% |
| half | 35.9% | 49.1% | 17.6% |
| the standing vector | 45.1% | 29.1% | 13.1% |
| twice | 65.8% | 4.2% | 2.7% |
| four times | 88.8% | 1.7% | 1.5% |
The rate rises by a factor of 2.80. The share falls by a factor of 30.52.
The world where the candidates quarrel most about the tuning parameter is the world where the quarrel matters least. That is the sentence this essay exists for, and it is not a mild observation: it means a sweep that measures the disagreement rate and stops has measured the factor that points the wrong way.
The worst case is neither end of the sweep
The product does not fall the whole way. It reads 16.4% with nothing omitted, 17.6% at half the standing vector, and 13.1%, 2.7% and 1.5% after that.
So it rises before it falls, and the maximum is interior. Fitting a parabola through the first three points puts the peak at a separation of about 0.36, at a product a little above 17.7%.
That is a more useful statement than either endpoint. The world in which a per-candidate tuning list is most likely to change the answer is not the world where every candidate is equivalent, and it is not the world where one candidate is obviously right. It is the world where a small coefficient is at stake: large enough that the candidates’ residuals differ and they pick different windows, small enough that the criterion has not yet made up its mind. Both factors are near their middle and neither has taken over.
It is also the world an analyst is most often in. Nobody runs a fifteen-candidate table to discover that a coefficient four times the standing vector belongs in the model.
The share outruns the rate by a factor of eleven
The two movements are worth putting in the same units, because their ratio is the whole story and the product’s fall is nothing but that ratio.
Across the sweep the rate rises by 2.80 and the share falls by 30.52. Dividing gives 0.0917, and 16.4% × 0.0917 = 1.50%, which is the product at the far end to the last digit printed. The product is a ratio of two movements and nothing else.
In logarithms the share moves 3.3 times as far as the rate. That factor is what makes the sweep’s headline safe to state as a direction rather than as a balance: for the rate to win anywhere, its movement would have to be within a factor of one of the share’s, and it is not close at any separation past the peak.
What the decisive share is really tracking
The share falls because the verdict gets easier, and “easier” has a number: the criterion’s span from best to worst against what a change of tuning parameter is worth.
At the far end that is 159.79 against 2.984, the fifty-four to one the field quotes. At the near end the span is 159.79 / 24.5 = 6.5 units against a tuning move of 1.467 — 4.4 to one. So the leverage the criterion has over the tuning parameter grows by a factor of 12.2 across the sweep, which is the span’s 24.5 divided by the tuning move’s 2.0.
Against that, the decisive share falls by 30.5. On a log-log reading the share goes roughly as the leverage to the power −1.37: it falls somewhat faster than inversely, so doubling the criterion’s grip on the table more than halves the chance that a tuning quarrel decides anything.
That exponent is the quantity a reader would want if they had a different table and wanted to guess where on this curve they sit. It is measured on five points of one sweep and is offered as a slope rather than a law, which is the most this field’s design supports.
Why the rate rises
The disagreement rate is about the tuning parameter and not about the criterion’s verdict, and it rises for a reason that has nothing to do with which candidate is better.
Each candidate chooses its tuning value from its own residuals. As the coefficients grow, the residuals of a candidate that omits a large coefficient stop looking like the residuals of a candidate that keeps it: one is the error process, the other is the error process plus an omitted regressor. Different series, different persistence, different best whitening window.
So at a separation of four the fifteen candidates are looking at fifteen materially different residual series and pick different windows on 88.8% of draws. At a separation of zero every candidate’s residuals are the same error process plus estimation noise, and they agree on more than two thirds of draws.
The rate is therefore a measure of how different the candidates’ residuals are, which is a fact about the world and not about the table’s verdict at all.
It is worth checking that against the one place the two could be confounded. If growing the coefficients also made the tuning values themselves more spread out — if the candidates were picking windows further apart as well as more often different — then the rate and the size of a disagreement would be moving together and the rate would be a proxy for both. They are not. What a tuning change is worth in criterion units grows from 1.467 to 2.984 across the sweep, a factor of two, while the rate grows by 2.80 and the criterion’s span by 24.5. The three quantities move at three different speeds, and it is the span that runs away from the others.
Why the share falls
The share is about the verdict, and it falls because the verdict gets easier.
At a separation of four the full model beats every proper subset on almost every draw, by a margin no change of whitening window can close: the criterion’s span from best to worst across the table is 159.79 units and a change of tuning parameter moves a candidate’s criterion by 2.984. Fifty-four to one. Nothing the tuning list does can reach the second-placed candidate.
At a separation of zero the span is 6.51 and a tuning change is worth 1.467. Four and a half to one, with fifteen candidates spread inside it, so a tuning change routinely reaches two or three of them and the winner is whichever one the change happened to favour.
The previous essay reports that as a count: 3.635, 3.300, 2.368, 1.560 and 1.660 candidates within a tuning change’s reach of the winner. The span grows 24.5-fold across the sweep and what a tuning change is worth grows twofold, so the count falls even though both quantities rise.
The two ends are two different problems
It helps to think of the sweep as two regimes with a crossing in the middle, because the quantities behave differently in each and the language for them is different.
At the near end the table is a lottery. Every candidate is nearly as good as every other; the criterion is choosing on parameter counts and noise; a tuning change reaches three or four candidates and flips the answer half the time it happens. What a reader gets from such a table is not a model, it is a draw from the set of models that fit about equally, and the tuning list is one of several things deciding which draw.
At the far end the table is a formality. The full model wins by a margin fifty times what a tuning change is worth, so the criterion is confirming something the data would have told anybody. The tuning list decides nothing, the disagreement rate is enormous, and neither fact matters.
In between is the only regime where the question is interesting, and it is where the standing coefficient vector sits — a probability of 13.1%, a rate of 45.1%, a turnover of 29.1%, and two or three candidates within reach.
That is not an accident of this collection’s choices. A coefficient vector is chosen for a simulation so that the selection problem is non-trivial, and a non-trivial selection problem is exactly one in which several candidates are within reach of each other. So the setting anybody would build to study selection is the setting in which the tuning list matters most, and the number measured there is not transferable to a world where the answer is obvious.
What that does to a reading
The two factors point opposite ways, so any reading that has only one of them can be exactly backwards.
A disagreement rate is easy to measure and easy to report. It needs no notion of the winner, no risk, no replicate; it is a count over draws of whether two vectors of chosen tuning values are equal. It is the natural thing to compute and it is the natural thing to quote.
A turnover is not. It needs the table scored under both rules and compared on which candidate came out, which doubles the fitting on every draw and is the reason the earlier field’s own sweep computed it and summarised it away before anybody needed it.
So the failure mode is specific and predictable: a study that reports “the candidates disagreed about the tuning parameter on nine draws in ten” reads as an alarming number and is, in that setting, the reassuring one. The same study in a world where nothing is true would report three draws in ten and be describing a setting where the tuning list decides the answer on one draw in six.
What the tuning parameter is, and what it is not
A word about the object being disagreed over, because “the tuning parameter” is doing duty for two different things in this collection and only one of them is on this sweep.
The one here is the width of the whitening window each candidate’s criterion is computed through. A candidate’s criterion needs a covariance for the errors; the covariance is estimated by banding the candidate’s own sample autocovariances; and the band’s width is chosen from a list. The field that introduced the list is careful that a list of values is not a rule for choosing among them, and that is the distinction this whole line of fields turns on.
The other is the order of a sieve, which does the same job by a different route and comes with a determinant term the window’s version does not. The field that put the two beside each other finds that what looked like a difference between two tuning parameters was a difference between two criteria, and repairs it.
Everything in this field is the window. The order’s list is an interval of integers rather than a geometric ladder, so its steps are a different size and its disagreement rate moves differently, and reproducing this sweep on it is work this field does not do.
The mediator, in both directions
There is a third quantity that closes the argument, and it is the earlier field’s rather than this one’s: a disagreement costs something only when it changes the winner.
Split the disagreeing draws on that. At the standing coefficients the regret is 0.03163 when the winner changes and −0.00067 when it does not — two orders of magnitude apart, with the smaller on the wrong side of zero. At twice the standing coefficients, 0.01333 against −0.00015.
That is what makes the turnover the quantity worth chasing rather than one of several. It is not merely correlated with the cost; it is the switch. A draw on which the tuning list changed nobody’s mind carries a regret indistinguishable from zero, and a draw on which it changed the winner carries the entire cost.
How many draws each of these needs
The two factors also differ in how expensive they are to measure, and it runs the same way as everything else here.
The rate is a proportion over every draw. At eight hundred draws and a rate near a half its standard error is about 1.8 percentage points, so three of the five cells above are separated by ten standard errors or more.
The share is a proportion over the disagreeing draws only, and at the far end of the sweep there are seven hundred and ten of those but only twelve on which the winner changed. A share of 1.7% on twelve events has a standard error of about half a point in its own units, which is fine for saying it is small and useless for saying whether it is 1.7% or 2.3%.
And the conditional size is worse again: it is an average over the disagreeing draws, split further by whether the winner changed, so the far end of the sweep is averaging a heavy-tailed quantity over twelve draws.
That ordering is the reason this field reports the probability rather than the size. The earlier field measured it directly: a conditional average needs about five hundred and fifty draws where the sweep that first reported it used a hundred and twenty, and the unconditional average needs several times more again because most of its budget is spent counting exact zeros.
A rate that is not a rate of anything
One more reading of the rate is worth closing off, because it is the one that makes the far end of the sweep look alarming.
At a separation of four the candidates disagree about the tuning parameter on 88.8% of draws. A reader could take that as evidence that the tuning parameter is badly determined — that the criterion has no idea what window it wants, so the whole procedure is unstable.
It is the opposite. The candidates disagree because they are looking at genuinely different residual series and each one is picking the window its own series wants. The disagreement is the machinery working, and it is the setting where the candidates’ residuals are nearly identical that produces agreement.
What is unstable at the far end is nothing, because the verdict is decided by a margin the tuning parameter cannot reach. What is unstable at the near end is the verdict, and there the candidates mostly agree about the window.
So the rate is not a diagnostic. It measures how different the candidates’ residuals are, and how different the candidates’ residuals are is nearly the inverse of how much the tuning parameter can decide.
The zero cell, which is the control
The separation of zero deserves its own paragraph, because it is a control and not merely the end of a sweep.
At that setting every coefficient is exactly nothing. Every candidate on the table is correctly specified; the true model is the empty one; and the criterion is choosing between fifteen subsets that all fit the same law with different numbers of parameters. Whatever the criterion does there is what it does when there is nothing to find.
What it does is disagree about the tuning parameter on 31.8% of draws and change its mind on 51.6% of those. So on 16.4% of draws — one in six — the answer a reader gets depends on whether the tuning parameter was chosen once or fifteen times, in a world where the answer means nothing either way.
That is the strongest form of the field’s point. The tuning list decides most where the data decide least, and a setting in which nothing is true is precisely a setting in which somebody is likely to be reading the selected model seriously.
It is also the cell whose conditional size comes out negative — −0.01699, with the winner-changing draws at −0.02973 — meaning that on average, letting each candidate tune itself was slightly better there. A rule that makes almost no difference to the answer’s quality and a large difference to which answer it is is a rule worth naming rather than optimising.
What survives being read the other way
It is worth checking that this is a real reversal rather than an artefact of one measurement, so the same two columns are read on the other tuning parameter and on the other tables.
Across the three tables at one list length — fifteen subsets, six pairs, and a nested ladder — the rates are 45.1%, 41.1% and 24.4% and the shares are 29.1%, 30.4% and 27.2%. There the rate moves by a factor of nearly two and the share barely moves at all, which is the third essay’s subject and is a different shape from this one.
So the two factors do not have a fixed relationship. The separation dial moves them in opposite directions; the table dial moves one of them and leaves the other. That is what says the product is the quantity to report: it is the only one of the three columns that means the same thing whichever dial has been turned.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A charge that reads the draw — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, regret, selection effect
- What a better charge buys — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, overfitting, regret
- A line that beats two curves — both name bandwidth selection, dependence, information criterion, model selection, monte carlo, nested models, overfitting
- A width that moves and an error that does not — both name bandwidth selection, information criterion, model selection, nested models, overfitting, regret, tuning parameter
- A window for every candidate — both name bandwidth selection, information criterion, model selection, overfitting, regret, selection effect, tuning parameter
- A criterion is a prediction of the hold-out — both name information criterion, model selection, monte carlo, nested models, overfitting, selection effect
Named objects
A flat tag is an object no other essay names yet.
Bandwidth selectionBenchmark forecastConditional distributionDependenceEffect sizeInformation criterionModel selectionMonte CarloNested modelsOrder-selectionOverfittingRegretSelection effectTuning parameterVariance decomposition