A table and a list
Worth reading first: A design is a number · Choosing the order.
The deferral that opened this field named the table as the dial to turn, and made a prediction about what turning it would do: a table of nested candidates differing by one coefficient would have a higher turnover at every list length.
The prediction is wrong in both of its readings, and the way it is wrong is the useful part.
Three tables
This collection carries three, and they differ in how many candidates they hold as well as in how close together those candidates are.
Fifteen subsets of four predictors — every non-empty one — which is the table every earlier field in this line works with. Six pairs, the two-predictor subsets, all of the same dimension so the criterion is choosing on content rather than on parameter count. And a nested ladder: one predictor, then two, then three, then four, each candidate differing from the next by exactly one coefficient. That last is the table the deferral asked for.
At the full list of eight tuning values, over eight hundred draws:
| table | candidates | they disagree | a disagreement decides | the product |
|---|---|---|---|---|
| fifteen subsets | 15 | 45.1% | 29.1% | 13.1% |
| six pairs | 6 | 41.1% | 30.4% | 12.5% |
| a nested ladder | 4 | 24.4% | 27.2% | 6.6% |
The nested ladder turns over less than the fifteen-subset table, not more. And its list changes the winner on 6.6% of draws against 13.1% — half as often.
Both readings of the prediction fail
The prediction can be read two ways and neither survives.
On the conditional probability — the share of tuning disagreements that change the winner, which is the quantity the word “turnover” names — the nested ladder reads 46.7%, 28.5% and 27.2% at four, six and eight values on the list, against the fifteen-subset table’s 51.1%, 29.8% and 29.1%. Lower at every one of the three.
On the unconditional probability — the chance a reader’s answer depends on how the tuning parameter was chosen, which is the quantity anybody would act on — it reads 8.0%, 6.6% and 6.6% against 14.6%, 12.6% and 13.1%. Roughly half, at every one.
So the table the deferral asked for is the table on which the tuning list matters least.
Which of the two factors carries the halving
The three tables differ on both shares, and only one of the differences is doing any work.
At the full list the turnovers are 29.1%, 30.4% and 27.2% — a range of 3.2 points, an eighth of the smallest of them. The disagreement rates are 45.1%, 41.1% and 24.4% — a range of 20.7 points, which is eighty-five per cent of the smallest.
Multiplied out, the ladder’s product is 0.541 of the fifteen-subset table’s through the rate and 0.935 through the turnover. The two together give 0.506, against the 6.6 / 13.1 = 0.504 measured.
So the halving is the disagreement rate, almost entirely. The turnover — the quantity the deferral actually made its prediction about — is nearly the same on all three tables, and is fractionally lower on the ladder in a way no reading of the prediction wanted.
That is a sharper statement of the failure than “both readings are wrong”. The prediction was that closeness would raise the share of quarrels that decide something. Closeness turns out to have almost no effect on that share, and a large effect on how often there is a quarrel at all — in the direction of fewer.
The candidate count does not explain it either
The obvious rival explanation is arithmetic rather than structural: the ladder holds four candidates and the fifteen-subset table holds fifteen, and more candidates have more chances to disagree.
It fits the two smaller tables and fails on the large one. Taking the chance that any one extra candidate agrees with the rest as a constant a, disagreement on 41.1% of draws with six candidates implies a = 0.899, and 24.4% with four implies a = 0.911 — two independent readings within a point and a half of each other.
Carry that to fifteen candidates and it predicts a disagreement rate of 75.5%. The measured rate is 45.1%.
Thirty points is not a discrepancy to be smoothed over. The fifteen-subset table’s candidates agree with each other far more than the smaller tables’ pairwise behaviour predicts, so the rate is not counting opportunities to disagree; it is measuring how alike the candidates’ residual series are, and the fifteen subsets of four predictors are more alike than four rungs of a nested ladder are different.
Which inverts the intuition the deferral was built on one more time. A nested ladder was proposed as the table whose candidates are closest together, and on the only measurement that separates the tables its candidates are the ones whose residuals differ most.
Why: close together is not nearly indifferent
The prediction’s premise is that candidates differing by one coefficient are close together, and that is true. What it assumes is that being close together makes the criterion nearly indifferent between them, and that is where it goes wrong.
The criterion’s span across the nested ladder is 11.44 units against the fifteen-subset table’s 21.04. Closer, as predicted. But what a change of tuning parameter is worth across the nested ladder is 0.797 against 1.857 — also smaller, and smaller by more.
The ratio is what decides the turnover, and it moves the wrong way: 0.797 against 11.44 is one part in fourteen, and 1.857 against 21.04 is one part in eleven.
Both quantities fall between the two tables, and the one that decides the turnover falls further. That is the whole mechanism, and it is measurable rather than arguable: the count of candidates within a tuning change’s reach of the winner is 1.298 of four on the nested ladder and 2.368 of fifteen on the subsets table.
And why the tuning change shrinks
The second half of that needs a reason, because it is not obvious that a nested ladder should make a tuning change worth less.
A candidate’s criterion is computed through a whitening window fitted to that candidate’s own residuals. On the nested ladder the four candidates are the four prefixes of one predictor list, so their residual series are nested too: each one is the next one plus an omitted regressor. The four series are similar, they want similar windows, and moving a candidate’s window from the shared value to its own therefore moves its criterion by less.
On the fifteen-subset table the candidates omit different predictors, so their residual series differ in what they contain rather than merely in how much. Moving a candidate’s window moves its criterion by more.
A table whose candidates share their residuals is a table on which the tuning parameter has less to do, in exactly the same proportion as it is a table whose candidates are hard to tell apart — except that the first effect is larger.
What the shortest list does to all three
The left column of the grid is worth reading on its own, because it is the one place all three tables behave the same way.
At four values on the list the turnover is 51.1%, 44.7% and 46.7% on the three tables — the highest reading in every row. At six values it is 29.8%, 31.0% and 28.5%, and at eight 29.1%, 30.4% and 27.2%.
So a coarse list produces disagreements that are much more likely to change the winner, and it produces them much less often: the rates are 28.6%, 25.8% and 17.1% at four values against 45.1%, 41.1% and 24.4% at eight.
That is the earlier field’s own mechanism and it is worth restating in these terms. A coarse list forces candidates that disagree to disagree by a large step, because the values on offer are far apart; a fine list lets them disagree by a small one. A large step moves a candidate’s criterion further and is more likely to reach past the winner.
The two effects very nearly cancel — the product moves by at most a fifth along any row — and the cancellation is exact enough, across three tables, that it is clearly not a coincidence of one setting. It is also not a law: nothing here says the two effects must cancel, only that they do across the range of lists this collection uses.
The list, which still does nothing
The reason to draw a grid rather than a row is that the earlier field’s invariant has to survive, and it does.
Along each row — three list lengths, one table — the probability that the list changes the winner moves by a factor of 1.158 on the fifteen subsets, 1.087 on the six pairs and 1.208 on the nested ladder. Down each column it moves by 1.828, 1.906 and 1.981.
So the list length is still the dial that does nothing, in every table, and the table is a dial that does something in every list length. The earlier field’s finding is not contradicted by any of this; it is placed.
Reading a prediction that failed
It is worth being precise about what kind of failure this is, because “the deferral was wrong” is the least informative thing to say about it.
The deferral was written at the end of the field that found the invariant, and it named the right dial for the right reason: the twelve per cent is a property of how nearly indifferent the criterion is between candidates, nothing in that sweep varies the table, and a nested table of near-identical candidates is the obvious way to vary it.
Every step of that is correct. What it then did was substitute one property for another — close together for nearly indifferent — and the substitution is exactly the kind that reads as a restatement and is not one.
Two candidates are close together if the criterion puts them near each other. They are nearly indifferent if the criterion puts them near each other relative to how far a change of procedure can move them. Those are the same quantity divided by different things, and the second denominator is not a constant: it shrinks along with the first when the candidates share their residuals.
That is the same shape as the finding in a neighbouring field of this round, where a curvature in a measured charge turned out to be a curvature in the denominator it was read against. A ratio whose denominator moves is a ratio whose numerator cannot be read alone, and the way to notice is to measure the denominator.
What the table dial is confounded with
The three tables differ in a second way, and it is stated rather than controlled because it cannot be controlled by choosing among these three.
They hold different numbers of candidates. A table of fifteen has more chances for one of its members to disagree with the shared value than a table of four, so part of the rate difference — 45.1% against 24.4% — is arithmetic rather than statistical. Fifteen draws from a distribution over eight values will disagree with a fixed value more often than four draws will, whatever the distribution.
That confound is real, and it is why the separation sweep in the first essay exists. That dial holds the table at fifteen subsets throughout, so nothing about the count changes, and it moves the probability by a factor of 11.75 where this dial moves it by 1.981.
The uncounfounded dial is worth six times what the confounded one is worth, which is the reason this essay is third in the field rather than first.
What the six pairs say
The middle row is worth reading because it separates two candidate explanations.
The six-pair table holds candidates that all have the same dimension. So the criterion cannot prefer one for spending fewer parameters; it is choosing entirely on fit. If the nested ladder’s low turnover were about parameter counts — about the criterion’s penalty term dominating — the six-pair table would look like the nested one.
It does not. It reads 12.5% against the fifteen-subset table’s 13.1% and the nested ladder’s 6.6%, so on the quantity that matters it sits with the larger table rather than the smaller one, despite holding six candidates rather than fifteen.
Its criterion span is 15.39 and its tuning change is worth 1.256 — one part in twelve, between the other two tables and nearer the fifteen-subset one.
So it is not the count of candidates and it is not the parameter counts. It is the ratio of what a tuning change is worth to how far apart the criterion puts the candidates, and the six pairs sit where that ratio puts them.
The conditional size, which is where the tables differ most
There is one column on which the nested ladder is not merely lower but different in kind, and it is the one the earlier fields spent the most effort on.
The conditional size — what a tuning disagreement costs on the draws where the candidates disagree — is 0.00873 on the fifteen subsets, 0.00918 on the six pairs and 0.00207 on the nested ladder. A quarter of the other two.
That is not a fourth independent finding; it follows from the turnover. A disagreement costs something only when it changes the winner, so a table on which fewer disagreements change the winner has a smaller average cost per disagreement — and 27.2% against 29.1% accounts for part of it. The rest is that when the nested ladder’s winner does change, it changes to the candidate one rung along, which differs by one coefficient and therefore delivers a nearly identical fit.
A table whose candidates are close together is a table on which changing the winner is cheap, which is the half of the deferral’s intuition that survives. What does not survive is the step before it: being close together made the change cheap and also made it rare, and the two effects reinforce rather than trade off.
At the shortest list the nested ladder’s conditional size is −0.00230 — negative, meaning that on those draws letting each candidate tune itself was slightly better. Twelve draws’ worth of a heavy-tailed quantity, so the sign is not to be leaned on, and it is reported because a table of positive numbers with one negative in it is more honest than a table with the negative rounded away.
What stays out of this field
Three things it could have measured and did not.
The other tuning parameter. Everything here is the whitening window. The sieve order has its own list — an interval of integers rather than a geometric ladder — and the field that put the two side by side finds them behaving differently in exactly the way a difference of list spacing would predict. Whether the separation dial’s factor of twelve reproduces on the order is a guess.
A table built to be indifferent. The three tables here are the three this collection already had, and none was constructed for this question. A table of candidates whose criteria were deliberately tuned to sit within a tuning change of each other would separate the ratio from everything it is confounded with here, and would be the clean version of this essay’s measurement.
And a diagnostic anybody could run. The count of candidates within a tuning change’s reach is computed here on simulated draws where the truth is known. On real data it needs no truth — a criterion’s span across a table and what a second tuning value does to it are both one extra pass — so the quantity is available, and nothing here checks whether it predicts the turnover on data it was not measured against.
What a reader should do with this
The practical form of the field is short.
The probability that a per-candidate tuning rule changes the answer is not a constant, and it is not a property of the list. It is a property of how nearly indifferent the criterion is between the candidates, measured in the units a change of tuning parameter moves that criterion by.
Both of those are computable on the data in hand, and neither needs a simulation. The criterion’s span across the table is one pass; what a tuning change is worth is one more pass at a second tuning value. Their ratio is the diagnostic, and the count of candidates within one tuning change’s reach of the winner is the readable form of it.
And the number to worry about is small when the answer is obvious. A table whose winner is out of reach of every rival is a table on which none of this matters. A table with three candidates inside a tuning change’s reach is one where the answer is partly a choice of procedure, and that is the setting a selection study is usually built in.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The quarrel that changes the winner — both name bandwidth selection, benchmark forecast, granularity, information criterion, model selection, overfitting, regret, selection effect, tuning parameter
- A charge that reads the draw — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, regret, selection effect
- What a better charge buys — both name bandwidth selection, benchmark forecast, dependence, information criterion, model selection, monte carlo, overfitting, regret
- A line that beats two curves — both name bandwidth selection, dependence, information criterion, model selection, monte carlo, nested models, overfitting
- A width that moves and an error that does not — both name bandwidth selection, information criterion, model selection, nested models, overfitting, regret, tuning parameter
- A window for every candidate — both name bandwidth selection, information criterion, model selection, overfitting, regret, selection effect, tuning parameter
Named objects
A flat tag is an object no other essay names yet.
Bandwidth selectionBenchmark forecastConditional distributionDependenceGranularityInformation criterionModel selectionMonte CarloNested modelsOrder-selectionOverfittingRegretSelection effectSubset of parametersTuning parameter