The eighth that was not a constant
Worth reading first: A design is a number · Choosing the order.
The field that split what a per-candidate tuning parameter costs ends on an invariant. A disagreement about the tuning parameter costs something only when it changes which candidate the table selects; that happens on 14.2%, 11.9% and 12.3% of draws at four, six and eight values on the list; and the list length that moves the disagreement rate by half does not move it.
The essay that reports it is careful to call it an invariant of the list. It also says, in its last paragraph, what that leaves open: nothing in that sweep varies the table. The candidates, the criterion, the coefficients and the world are identical in every cell of it.
A number that does not move when one dial is turned is not a constant until a second dial has been tried.
The second dial
The obvious one is the table itself, and it is the third essay of this field. The one that turns out to matter is not.
Every candidate on the table is scored on the same data. How nearly indifferent the criterion is between them is therefore a property of how much the coefficients each candidate omits are worth: scale the true coefficient vector to nothing and every candidate is true, so the criterion is choosing between them on nothing but how many parameters they spend; scale it up and the full model wins on every draw.
That is one multiplier. It holds the table at fifteen subsets, the list at eight values, the law at a first-order autoregression, the sample at a hundred and twenty rows and the seeds where they were. Nothing else moves.
What it does
At multiples of 0, 0.5, 1, 2 and 4 of the standing coefficients, over eight hundred draws apiece, the probability that the tuning list changes which candidate wins is
16.4%, 17.6%, 13.1%, 2.7% and 1.5%.
A factor of 11.75 between the largest and the smallest, on the same table at the same list length. The 13.1% in the middle is the earlier field’s own setting, reproduced here rather than cited.
The eighth is not a constant. It is what the quantity happens to be in a world where the omitted coefficients are worth what this collection’s standing vector says they are worth, and it is a property of that world rather than of the problem.
What a reader was entitled to think
It is worth being fair to the reading being corrected, because it is not a careless one.
The earlier field does not claim the twelve per cent is universal. What it claims is that the probability is flat across list length while the disagreement rate rises by half — and that claim is exactly right, and this field re-runs it and confirms it in every row of its own grid.
But a number printed once, with an invariance beside it, reads as a property. Three values in a row that agree to within a fifth of themselves, in a field whose whole subject is which quantities move and which do not, invite the conclusion that this one does not move. And the setting it was measured in — one table, one coefficient vector, one law — is the setting every essay of that field works in, so there is nothing in the text to say what the number is conditional on.
That is the shape of the defect this collection meets most often. It is not a wrong number; it is a number whose conditioning is invisible, and the repair is not to correct it but to vary something else and watch it move.
Nothing is true and the list decides
The end of the sweep a reader will not expect is the zero end.
At a separation of zero every candidate on the table is true: the coefficients it omits are all exactly nothing, so every subset fits the same law and the criterion is choosing between them on how many parameters they spend and on noise. In that world the candidates disagree about the tuning parameter on 31.8% of draws — the lowest rate in the sweep — and 51.6% of those disagreements change the winner, which is the highest share by a factor of thirty.
So the probability that the list decides the outcome is 16.4% there, and it is higher than at any positive separation except the next one along.
The mechanism is not subtle once it is stated. A criterion that is nearly indifferent between two candidates can be tipped by anything, and a change of tuning parameter is anything. A criterion that has already decided cannot be tipped by a change of tuning parameter however often the change happens.
And at the far end nothing decides
At four times the standing coefficients the candidates disagree about the tuning parameter on 88.8% of draws — the highest rate in the sweep, nearly three times the rate at zero — and 1.7% of those disagreements change the winner.
The full model wins on almost every draw whatever tuning value it is scored at, because the coefficients it carries and the others omit are large enough that no plausible change of window can close the gap. Every candidate quarrels about the tuning parameter and the quarrel decides nothing.
The next essay is about those two columns, because they are the whole mechanism and they point in opposite directions.
The share moves three times as fast as the rate
The two factors are quoted at the ends of the sweep and dividing them says which one carries the product.
The disagreement rate runs from 31.8% to 88.8% — a factor of 2.79. The share of disagreements that decide runs from 51.6% to 1.7% — a factor of 30.4 the other way. Their product, 16.4% against 1.5%, is a factor of 10.9, and both ends multiply out exactly: 0.318 × 0.516 = 0.164 and 0.888 × 0.017 = 0.015.
On a logarithmic scale the rate moves by 1.03 and the share by 3.41, so the conditional share is three and a third times more responsive than the rate. The product is, to a good approximation, the share with a mild correction — which is the opposite of the picture the list-length sweep gives, where the rate moves and the share does not.
That is the sharpest form of what this field adds. The same two factors are the moving one and the fixed one depending on which dial is turned, so neither can be described as the mechanism. The list decides how often the criteria disagree; the separation decides whether a disagreement matters; and the product is invariant under the first dial and moves by a factor of eleven under the second.
The dangerous world is the nearly-null one
The sweep’s maximum is not at either end. It is 17.6% at half the standing coefficients, above the 16.4% at zero and well above everything after it.
That interior peak has a reading. At exactly zero the candidates are all true, so they disagree about the tuning parameter least often — the residual series are most alike — even though any disagreement is maximally likely to tip a criterion that has nothing to decide on. Raising the separation a little makes the residual series differ enough to quarrel more, while leaving the criterion still nearly indifferent. The two factors briefly both favour the list.
So the world in which a per-candidate tuning rule is most likely to decide the answer is one where the omitted coefficients are small but not zero — which is precisely the world a model selection problem is in whenever it is difficult.
Read across the whole sweep: the list decides about one draw in six while the separation is under one, one draw in thirty-seven at twice it, and one in sixty-seven at four times. The tuning choice matters only where the selection itself is uncertain, and it stops mattering exactly when the answer becomes obvious — which is the least useful place for a safeguard to be reliable.
What the rule being priced actually is
It is worth restating the rule once, plainly, because everything above is a probability about it.
A table of candidate regressions is scored by a criterion. The criterion needs a tuning parameter — here the width of the whitening window each candidate’s criterion is computed through — and there are two ways to supply it. Shared: choose one value from the fullest candidate’s residuals and score the whole table at it. Own: let every candidate choose its own value from its own residuals.
The second is what a practitioner does by default, because each candidate’s criterion is a function that takes residuals and every candidate has its own. The first is what a practitioner does if they have thought about it, because scoring a table at different tuning values is scoring it on different scales.
The regret is the paired difference in what the two rules deliver: the risk of the coefficients the winner under “own” produces, less the risk under “shared”, on the same draw. Everything in this field is about how often that difference is not zero and how large it is when it is not.
And it is not zero only when the two rules pick different winners. That is not an assumption; it is the earlier field’s measurement, and it is what makes the probability above the quantity worth reporting rather than one of several.
The measurement behind the word “indifferent”
“Nearly indifferent” is doing a lot of work above, and it is measured rather than asserted.
On each draw, two things are computed. How far a change of tuning parameter moves a candidate’s criterion — the average over candidates of the gap between the criterion at the shared tuning value and at that candidate’s own. And how far apart the candidates are — the criterion’s span from the best to the worst on the table.
The reading is then the count of candidates that lie within one tuning change’s worth of criterion of the winner. One means the winner is out of reach of everything else and a tuning list cannot change it; two means there is one rival within a tuning change’s worth.
Across the separation sweep it is 3.635, 3.300, 2.368, 1.560 and 1.660 of fifteen candidates, while the criterion’s own span across the table grows from 6.51 to 159.79 — a factor of 24.5 — and what a tuning change is worth grows only from 1.467 to 2.984.
Both quantities grow; the span grows eight times faster. That gap is the whole of the mechanism, and it is why the count of candidates within reach falls even though a tuning change is worth more at the far end than at the near one.
What the size does while the probability moves
There is a third column in the decomposition and it behaves differently again, which is worth reading because it is where the arithmetic stops being intuitive.
The conditional size — what a disagreement costs on the draws where the candidates disagree — runs −0.01699, −0.00419, 0.00873, 0.00041 and −0.00025 across the sweep. It changes sign twice, it is largest in magnitude at the end of the sweep where nothing is true, and its sign there is negative.
A negative regret means that choosing the tuning parameter per candidate beat choosing one for the whole table. In a world where every candidate is true that is not paradoxical: the shared rule tunes on the fullest candidate’s residuals and hands that value to everybody, and when no candidate is better than any other, letting each one pick its own value is a small genuine improvement rather than a small self-deception.
Split further, on the draws that disagree: when the winner changes, the regret is −0.02973 at zero separation and 0.03163 at the standing coefficients; when it does not, it is −0.00342 and −0.00067. The order-of-magnitude gap between those two columns is the earlier field’s own finding, and it survives the separation dial with its sign attached to which world is being worked in.
The probability is the readable quantity and the size is not. That is the practical reason this field reports the first: a probability of 16.4% against 1.5% is a statement anybody can act on, and a conditional average that changes sign between worlds is one nobody can.
The one place it is not monotone
The count of candidates within reach falls at every step to twice the standing coefficients and then ticks back up, from 1.560 to 1.660, while the probability the list decides keeps falling from 2.7% to 1.5%.
That is reported rather than smoothed. What produces it is that a tuning change is worth more in absolute criterion units at the far end — 2.984 against 2.385 — so it reaches slightly further into a table whose candidates are further apart. What it reaches is candidates that are worse by enough that reaching them does not make them win.
So “within a tuning change’s reach” is a necessary condition for the winner to change and not a sufficient one, and the last cell of the sweep is where the two come apart. A reader who wanted one number to predict the turnover would not get it from this one, and the honest statement is that the count explains the fall from 51.6% to 4.2% and does not explain the last step.
What is not being varied
Three things are held fixed throughout and each of them could be the next dial.
The law. Everything here is a first-order autoregression at the persistence this collection works at. The earlier field sweeps its decomposition across four laws and finds the probability moving by less than the list moves it; whether the separation dial’s factor of twelve survives a change of law is not measured here.
The sample size. A hundred and twenty rows throughout. The separation and the sample size are not independent levers on the criterion — doubling the rows and halving the coefficients move the same quantity in the criterion’s own units — so a sweep over n would very likely reproduce the sweep over β with a different scale on the axis, and would not be a second finding.
And what a tuning change is worth. The list is the same eight values in every cell, so the size of a tuning change is whatever those eight values make it. A finer list would move smaller distances and reach fewer candidates, which is exactly the dial the earlier field turned, and the grid in this field’s third essay is what puts the two dials on one picture.
Why a multiplier on the coefficients is the right dial
There are several ways to make a criterion more or less decided between its candidates, and the one used here is the crudest available. That is deliberate.
A multiplier on β changes nothing about the table, the list, the seeds, the residual variance or the design. Every candidate still omits the same predictors; every predictor still has the same persistence; the fifteen subsets are still the same fifteen subsets. What changes is how much the omissions cost, and therefore how far apart the criterion puts the candidates.
The alternatives all change two things at once. Adding predictors changes the table’s size as well as its spread. Correlating the columns changes what each candidate can recover as well as how nearly indifferent the criterion is. Lengthening the sample changes the criterion’s scale and the tuning parameter’s own precision together.
A dial that moves one quantity is worth a factor of twelve; a dial that moves two is worth an argument. The table sweep in this field’s third essay moves two on purpose, and reports both.
The identity, re-checked
Everything above rests on a decomposition that is exact, and the exactness is re-checked here rather than inherited.
When every candidate picks the value the shared rule picks, the two rules are the same rule: the same criteria, the same winner, the same coefficients, the same regret. So the paired difference on such a draw is not merely small, it is exactly zero, and
holds to machine precision. That is the earlier field’s own construction, and both the table and the coefficient vector are different here — so it is asserted again, at three tables and three separations, and the largest paired difference on a draw where nothing disagreed is zero at every one.
An identity that holds by construction is the only kind whose failure means the construction moved.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A charge that reads the draw — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, regret, selection effect
- A window for every candidate — both name bandwidth selection, data snooping, information criterion, model selection, overfitting, regret, selection effect, tuning parameter
- What a better charge buys — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, overfitting, regret
- A line that beats two curves — both name bandwidth selection, dependence, information criterion, model selection, monte carlo, nested models, overfitting
- A width that moves and an error that does not — both name bandwidth selection, information criterion, model selection, nested models, overfitting, regret, tuning parameter
- A criterion is a prediction of the hold-out — both name information criterion, model selection, monte carlo, nested models, overfitting, selection effect
Named objects
A flat tag is an object no other essay names yet.
Bandwidth selectionBenchmark forecastConditional distributionData snoopingDependenceEffect sizeInformation criterionModel selectionMonte CarloNested modelsOrder-selectionOverfittingRegretSelection effectTuning parameter