The quarrel that changes the winner
Worth reading first: A design is a number · The observations that repeat each other.
The decomposition one essay back leaves a question with no obvious answer. The list the candidates choose their tuning parameter from can be refined until they are always neighbours on it — they are one step apart on every list here — and refining it pulls the values disagreed about from 10.00 apart down to 4.14. What a disagreement costs does not move at all.
So the cost is not a function of how far apart the values are. Something else decides it, and there is only one thing it can be.
What a per-candidate tuning parameter can actually do
The rule under test compares five candidates by a criterion and selects the best. The tuning parameter enters that comparison in two places: it fixes the whitening each candidate is scored through, and it fixes the whitening the winner is finally fitted through.
A disagreement can therefore do two different things.
It can change the fitted coefficients while the table’s answer stays the same — the same candidate wins, fitted through a slightly different whitening. Or it can change which candidate wins, and then the coefficients change because the model changed.
Those are not small and large versions of one effect. They are different events, and the field that established the tuning parameter moves the winner is where the second one is measured as a displacement rather than as a cost.
Over twelve hundred draws at each list length, split that way:
When the winner changes, the regret is 0.02215, 0.03029 and 0.03145 at four, six and eight values on the list.
When it does not, it is −0.00243, −0.00069 and −0.00065 — negative at every length, and inside two standard errors of nothing at all.
Two orders of magnitude at the longer lists, and the smaller half is on the wrong side of zero.
The half that costs nothing, and why it is negative
The negative sign is worth taking seriously rather than rounding to zero, because it is consistent across every list length and every law and it says something about the rule.
When the table’s answer is unaffected by the tuning parameter, letting each candidate use its own window is a very slightly better rule than making them share one. That is not surprising once stated: the winner is being fitted through a whitening chosen from its own residuals rather than from the fullest candidate’s, and its own residuals are the right series to estimate its own dependence from. The field that priced shared against own estimation established the opposite ordering overall, and the reason it comes out the other way there is entirely the selection: a shared value is better because it keeps the criteria comparable, not because it fits better.
So the split separates two effects that were being averaged into one number of the wrong sign.
Sharing is worth −0.0007 in fit and +0.031 in selection, and the sum is positive only because the second is fifty times the first.
Which explains the flatness
Now the earlier essay’s puzzle answers itself.
The share of disagreements that change the winner is 49.6% at four values, 29.2% at six and 28.5% at eight. It falls as the list grows, by a factor of 0.57 over the range where the rate rises by 1.52.
That is the reciprocal movement the conditional size hides. The average over disagreements is a mixture of the two halves, weighted by how often each occurs:
with τ the turnover. That is an identity rather than a fit — the two halves partition the disagreeing draws — so it returns 0.00975 at four values and 0.00848 at eight exactly, which are the measured conditional sizes. What it adds is the reason they are so close: 0.496 × 0.02215 is 0.01098 and 0.285 × 0.03145 is 0.00895, and the small negative term takes 0.00123 off the first and 0.00047 off the second. The conditional size is nearly flat because the thing it averages is bimodal, and the mixing weight moves against the size of the larger mode.
And the mode itself is not flat: the regret given a changed winner rises from 0.02215 to 0.03145 across the same range. So there are now three things moving — the rate up, the turnover down, the decisive cost up — and the two-factor split that this field started from was one factor short.
How large a change of winner is, in the units the table is read in
The 0.031 is a regret in the coefficient error against a replicate, which is the scale the whole line of fields is measured on, and it is worth converting into something a reader can hold.
The best any rule achieves on these draws is the risk of the best candidate fitted through the true whitening, and the errors in the table run a little above one. So 0.031 is about three per cent of the error, incurred on the draws where the winner moves. For comparison, the entire difference between the best and worst charge for a covariance band’s width — a whole field’s worth of argument — is half a per cent, and the gap between the best feasible width and the best available one is another half.
A change of winner is six times the size of anything a tuning rule can buy by being cleverer about the same winner. That ordering is worth carrying, because it says where the attention belongs: the tuning parameter is worth arguing about only through the selection it perturbs, and the perturbation is rare.
It also says why the earlier field’s regret looked small. Averaged over every draw, three per cent of the error on an eighth of draws is under half a per cent — indistinguishable, on the face of it, from the charge arithmetic two fields away. The two numbers are the same size and they are not the same kind of thing at all: one is a small effect everywhere and the other is a large effect almost nowhere.
The four laws, which move both shares and not the product
Everything above is a first-order autoregression. The other three laws move the two shares a great deal, and they move them the same way.
The candidates disagree on 43.3% of draws under a first-order autoregression, 65.8% under a five-period moving average, 58.5% under long memory and 46.9% under a break in the persistence. The share of those disagreements that change the winner runs the other way: 28.5%, 20.5%, 13.7% and 25.6%.
A moving average is the extreme case and it is legible. Its autocovariance sequence stops dead after the fourth lag, so the criterion has almost nothing to choose between windows of eight and twenty and the five candidates land wherever their own residuals happen to fall — the highest disagreement rate of the four, and the lowest cost when they do: 0.00188 ± 0.00058 against the autoregression’s 0.00848. A great many quarrels, almost none of them about anything.
Long memory is the other extreme in the turnover: only 13.7% of its disagreements are decisive, because its sequence is still substantial at the twentieth lag and every candidate wants a wide window, so the values they pick are all at the top of the ladder and produce whitenings that differ very little.
In both cases the two shares move in opposite directions from the autoregression’s, which is the same pattern the list length produced and arrived at by a completely different route.
One draw in eight, whatever the law
The heading above claims the two shares move and the product does not, and the product is worth computing rather than asserted, because it is the only quantity here a reader can carry away.
Multiplying the disagreement rate by the turnover gives the share of all draws on which a per-candidate tuning parameter changes which candidate is selected:
- first-order autoregression: 0.433 × 0.285 = 12.3%
- five-period moving average: 0.658 × 0.205 = 13.5%
- long memory: 0.585 × 0.137 = 8.0%
- a break in the persistence: 0.469 × 0.256 = 12.0%
The two factors vary by 1.52 and 2.08 across the four laws and their product varies by 1.69, with three of the four inside a point and a half of 12.6%. Long memory is the one that is genuinely lower, and for the reason its own paragraph gives: every candidate wants the widest window available, so the disagreements it does have are between neighbouring rungs at the top of the ladder.
So the portable statement is that a decisive disagreement arrives on about one draw in eight, and the cost of one is about three per cent of the error. Multiplied out, an unconditional regret of about 0.0037 for the autoregression — which is the number the earlier field reported, arrived at from its two constituents rather than measured whole.
How nearly the two movements cancel
“The mixing weight moves against the size of the larger mode” is exact enough to be checked, and it is close rather than exact.
Across the list the turnover falls by a factor of 0.285 / 0.496 = 0.575 and the decisive cost rises by 0.03145 / 0.02215 = 1.420. Their product is 0.816, so the decisive half’s contribution to the conditional size falls by about a fifth rather than staying put.
What flat would have required is computable. Holding the conditional size at its four-value figure of 0.00975 once the turnover has fallen to 0.285 would take a decisive cost of 0.0358, against the 0.03145 measured. The rise delivered about eighty per cent of what exact cancellation would have needed, and the conditional size duly fell — from 0.00975 to 0.00848, thirteen per cent.
That is the right way to read the flatness. It is not a conservation law and nothing forces the two movements to match; it is two effects of similar size pulling opposite ways, leaving a residue small enough that a field measuring only the average would have called the whole thing constant and stopped.
Why refining the list makes disagreements less decisive
The mechanism is a fact about where the extra values go.
The window list is a geometric ladder: 0, 2, 12, 30 at four values, and 0, 1, 2, 4, 8, 12, 20, 30 at eight. The values it adds are in the middle. Two candidates that would have had to choose between 2 and 12 on the short list can now choose between 4 and 8, and the whitenings those two produce are far more alike than the ones the short list forced.
A criterion comparing five candidates through nearly identical whitenings gets nearly identical answers, so the ordering of the table survives. A criterion comparing them through a truncation at 2 and a truncation at 12 is comparing two quite different treatments of the dependence, and the ordering does not survive.
So a longer list buys disagreements that are cheap and does not reduce the number of expensive ones. Both halves of that matter and only the first is visible in the rate.
The half that is measured badly, and it is the interesting one
A last observation about the arithmetic, because it constrains how hard any of this can be pushed.
The decisive half is the smaller sample. At eight values on the list there are 520 disagreements and 148 of them change the winner, so the 0.03145 rests on 148 draws and carries a standard error of 0.00327 — a tenth of itself. The non-decisive half rests on 372 draws and carries 0.00022, which is a third of a very small number.
The half that carries all the signal is the half measured worst, and that is structural rather than a budgeting mistake: the decisive draws are the rare ones and their regrets are the heavy-tailed ones. Reaching a twentieth of relative precision on the decisive cost needs about four times the twelve hundred draws used here, which is why the rise from 0.02215 to 0.03145 across the list is reported as a movement rather than measured as one — it is 1.7 standard errors, and it is the third moving quantity the field found rather than a finding of its own.
What the split does not establish
Two limits are worth naming before the conclusion, because the split is a cross-tabulation rather than an experiment and a cross-tabulation can be read as more than it is.
The two halves are not randomly assigned. Whether a disagreement changes the winner is decided by the same sample that decides how large the regret is, so the draws in the decisive half are a selected set — they are the draws where the table’s two best candidates were close enough for a change of whitening to reorder them. Those draws would have carried more regret than average under any perturbation, so the 0.031 is not the causal effect of changing the winner; it is the regret on the draws where the winner was changeable, which is larger for a reason that has nothing to do with the tuning parameter.
That does not touch the field’s use of the split — the mixture identity is arithmetic, and the reciprocal movement of the two shares is the finding — but it does forbid the sentence changing the winner costs 0.031. What is established is that the regret lives entirely on those draws.
And the negative half is a fit effect measured in one direction. −0.0007 is what a per-candidate window buys on the draws where the table’s answer is unaffected, at this sample size, on this table of five candidates, under a Bartlett window. It is two to three standard errors from zero at the longer lists and it is not measured anywhere else. A reader tempted to conclude that per-candidate estimation is the better rule for fitting should note that the field that priced shared against own estimation directly does the comparison properly, without a selection step in it.
What a list would have to do to change the answer
The mechanism suggests its own test, and it is worth stating because it is the thing this field would do next.
If refining the list makes disagreements cheap by making neighbouring whitenings similar, then a list refined somewhere the whitenings are not similar should behave the other way. The window ladder’s neighbours are close in effect near the top, where a window of twenty and a window of thirty whiten almost identically, and far apart near the bottom, where zero and one are the difference between whitening and not.
So a list refined only at the bottom — 0, 1, 2, 3, 4, 30, say — should raise the disagreement rate and hold the turnover up, and its conditional size should rise rather than stay flat. A list refined only at the top — 0, 2, 12, 16, 20, 24, 30 — should raise the rate and collapse the turnover further.
Neither is measured here. The lists in this field are the ones the earlier field thinned from the natural ladder, taken unchanged so that the comparison is with its numbers rather than with a construction of this one’s. Where a list is refined is a free parameter nobody has swept, and the mechanism above says it should matter more than how long the list is.
What this makes of the original comparison
The field this began in asks whether a table’s tuning parameter should be shared or chosen per candidate, and answers with a single regret. That regret is now three quantities:
How often the candidates disagree — a property of the list alone, measurable with no regret in it at a thousandth of the cost, rising from nothing to 43.3% across the list.
How often a disagreement is decisive — a property of how far apart the list’s neighbouring values are in their effect on a whitening, falling from 49.6% to 28.5%.
And what a decisive disagreement costs — 0.031, which is the only one of the three that is about the selection problem rather than about the list.
The recommendation the earlier field arrives at is unchanged and its reason is not. Share the tuning parameter across the table is right, and it is right because sharing keeps the criteria comparable — not because a per-candidate window fits worse. It fits very slightly better, and pays for it fifty times over on the eighth of draws where it moves the winner.
That eighth is the next essay’s subject, and it is the one quantity in the field that does not depend on the list at all.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A table and a list — both name bandwidth selection, benchmark forecast, granularity, information criterion, model selection, overfitting, regret, selection effect, tuning parameter
- A charge that reads the draw — both name bandwidth selection, benchmark forecast, estimation error, information criterion, model selection, out of sample, regret, selection effect
- A step that is not a ratio — both name bandwidth selection, granularity, information criterion, model selection, overfitting, regret, selection effect, tuning parameter
- What a better charge buys — both name bandwidth selection, benchmark forecast, estimation error, information criterion, model selection, out of sample, overfitting, regret
- The comparison that was not made — both name information criterion, model selection, paired comparison, regret, selection effect, whitening
- The width a band is measured in — both name bandwidth selection, estimation error, information criterion, long-run variance, model selection, overfitting
Named objects
A flat tag is an object no other essay names yet.
Bandwidth selectionBenchmark forecastDiscretenessEstimation errorGranularityInformation criterionLong-run varianceModel selectionOut of sampleOverfittingPaired comparisonRegretSelection effectTuning parameterWhitening