How often it matters
Worth reading first: A design is a number · The observations that repeat each other.
Three quantities have now been separated out of one. How often five candidates disagree about a tuning parameter rises from nothing to 43.3% as the list they choose it from is lengthened. How often a disagreement changes which candidate the table selects falls from 49.6% to 28.5% over the same range. And what a decisive disagreement costs is about three per cent of the error and does not depend on the list in any way the measurement can see.
Two of those move. This essay is about their product, which does not.
The quantity
How often does the tuning list change which candidate wins? It is the disagreement rate times the turnover, and it is the only thing in the field a reader has to care about — everything downstream of the winner is a function of it, and everything upstream is bookkeeping about a rule.
It is 14.2%, 11.9% and 12.3% at four, six and eight values on the list. A spread of 2.3 percentage points, across a list length that moves the disagreement rate by a factor of 1.52 and the turnover by 0.57.
The two shares are reciprocal to within the noise, and the product they make is flat.
The two shares behind it are printed under each bar, and they are worth reading as a pair rather than as a product: at four values on the list a quarter of draws produce a quarrel and half of those quarrels decide the table; at eight, two fifths produce a quarrel and under a third decide anything. Twice as many arguments, half as many of them about something.
Which is a stranger result than it looks
It is easy to read that as an accounting identity and it is not one. Nothing forces two shares measured on different objects to move reciprocally.
The disagreement rate is a fact about five residual series and a list of values. It is computed with no regret in it, no replicate, and no reference to which candidate is best; the tuning values alone decide it.
The turnover is a fact about a table of criteria. It is computed from which candidate has the lowest score under two different whitenings, and it has nothing to do with how many values were on the list except through what those values do to a whitening.
Two independent measurements, on different objects, moving by factors of 1.52 and 0.57 in opposite directions across the same sweep. Their product is 12% throughout.
The mechanism is the one the previous essay traced: the values a longer window list adds are in the middle of the ladder, and the whitenings they produce are more alike than the ones the shorter list forced. So the list buys extra disagreements and buys them cheap, in the same act and by the same amount. A longer list changes how often the candidates quarrel and not how often the quarrel matters.
One more thing follows from the flatness and it is the practical half. A rule’s designer choosing how many values to put on a tuning list is usually choosing between a coarse ladder that is fast and a fine one that feels safer. The fine one produces more disagreements between candidates and no more changes of answer, so what it buys is a diagnostic that fires more often about nothing.
How flat, in standard errors
A flat quantity deserves the same scrutiny as a moving one, so it is worth saying what “flat” is worth on twelve hundred draws.
A share near 12% carries a standard error of √(0.12 × 0.88 / 1200) = 0.94 points, so a difference between two list lengths carries at most 1.33 unpaired. The three readings are 14.2%, 11.9% and 12.3%.
The largest gap, 14.2 against 11.9, is therefore about 1.7 standard errors — not a movement, and not a demonstrated zero either. The second gap, 11.9 against 12.3, is 0.3.
The shape of the three is the more persuasive part. If the product were declining with list length the middle reading would sit between the other two, and it is the lowest of the three. A flat value near 12.1% with the four-value point high by noise fits what is measured; a decline does not.
Settling it would take about three times the draws — the 2.3 points would need a standard error near 0.77 to read at three, against 1.33 now. That is worth naming rather than leaving implicit, because the field’s finding is a flatness and a flatness is exactly the claim a small count flatters.
The reciprocity is thirteen per cent from exact
The two shares are described as reciprocal, and the residual is worth having as a number.
Recovering the disagreement rates from the products and turnovers gives 28.6%, 40.8% and 43.3% at four, six and eight values. So the rate rises by a factor of 1.514 and the turnover falls by 0.575.
Exact reciprocity would need the turnover to fall by 1/1.514 = 0.661. It falls by 0.575, which is thirteen per cent further — and that thirteen per cent is the product’s own drop, from 14.2% to 12.3%.
So the honest summary is that two independently measured shares, each moving by half to three quarters of itself, are reciprocal to within thirteen per cent. That is a strong statement about a mechanism and it is not the exact cancellation the word reciprocal suggests.
What the longer list does to the diagnostic
The same numbers say what a designer is actually buying, read as a detector rather than as a pair of shares.
Every change of winner is preceded by a disagreement, so a quarrel between the candidates catches every decisive event there is — the diagnostic’s recall is one, at every list length, by construction.
What changes is its precision, which is the turnover: 49.6% of quarrels decide something at four values and 28.5% at eight. Read as false alarms, quarrels that change nothing run at 14.4% of all draws on the short list and 31.0% on the long one.
A finer ladder doubles the noise and leaves the signal alone. That is the same conclusion the section above reaches from the product, arrived at without dividing anything by anything, and it is the form a designer can act on: the extra values are not making the rule worse, they are making the one observable symptom of the rule’s instability less informative about it.
And it survives the law
The list is one dial. The law the errors follow is another, and it moves the two shares much harder.
The candidates disagree on 43.3% of draws under a first-order autoregression, 65.8% under a five-period moving average, 58.5% under long memory and 46.9% under a break in the persistence. The share of those that change the winner is 28.5%, 20.5%, 13.7% and 25.6%.
The products are 12.3%, 13.5%, 8.0% and 12.0%.
Three of the four are inside a point and a half of each other, on laws whose disagreement rates span twenty-two percentage points. Long memory is the exception at 8.0%, and it is the exception for a legible reason: its sequence is still substantially correlated at the twentieth lag, so every candidate wants a wide window, all five land near the top of the ladder, and the whitenings they produce differ least of any law here. It has the second-highest disagreement rate and by far the lowest turnover.
So the invariant is real and it is not exact. It holds across a list length that moves one factor by half and across three of four laws; the fourth breaks it by a third, in the direction its own dependence predicts.
What “does not move” is worth, given how it was measured
Before pushing the invariant further it is worth pricing it, because a flat line is the easiest thing in statistics to produce by accident.
Each of the three readings rests on twelve hundred draws. The product is a rate, so its standard error is the binomial one: at 12.3% on twelve hundred draws that is 0.9 percentage points. The three readings — 14.2%, 11.9%, 12.3% — sit inside a band about two standard errors wide, so what is established is that the product does not move by more than about two and a half points across a list length that moves its first factor by fourteen.
That is a flatness worth reporting and it is not a proof of exact constancy. A movement of a point and a half would be invisible here and would matter to nothing.
What makes it more than a null result is that both factors are measured to much better precision than their product needs, and both move by many standard errors. The disagreement rate moves from 28.6% to 43.3% — fifteen points, against a standard error under one and a half. The turnover moves from 49.6% to 28.5% — twenty-one points, against standard errors of 2.7 and 2.0. Two quantities each moving by five or more standard errors, whose product moves by one. That is the shape a genuine cancellation makes, and it is the shape the earlier field’s flat regret did not make, because there the factors were measured worse than the product.
Where it does not hold at all
The window is one of the two tuning parameters this collection chooses. The autoregressive order is the other, and it does not have the invariant.
Its product runs 5.5% at four values on the list, 9.3% at eight and 10.6% at thirteen — a spread of five points, nearly doubling. Its disagreement rate rises harder than the window’s, by a factor of 2.74, and its turnover falls less, from 30.8% to 21.6%. The two do not cancel.
The reason is the shape of the list, and it is the same observation the field that could not compare them makes for a different purpose. An order list is an interval of integers and a window list is a geometric ladder. Lengthening an order list from four values to thirteen adds orders 9 through 12 at the far end, where a candidate that wants a high order can now go; lengthening a window list adds values in the middle, where nobody wanted to go. So the order’s extra values create disagreements that are further apart — 2.99 steps at thirteen values against the window’s 1.04 — and further-apart values change more winners.
That is a real limit on the finding and it is worth stating as one. The invariant is a property of a geometric list, not of tuning parameters. A list that grows a tail behaves the other way.
The reading that is available to a practitioner
Everything above is measured over draws from a known law, which is the position nobody analysing data is in. It is worth asking what survives the translation, because two of the three quantities do.
The disagreement rate is observable on one sample. Fit the five candidates, take each one’s residuals, run the criterion on each, and see whether the five answers agree. No replicate, no law, no simulation — it is a property of the sample in hand, and it costs five criterion evaluations. A practitioner who finds all five picking the same window knows, with certainty rather than in expectation, that the convention did not matter on their data.
So is the turnover, one draw at a time. Run the table both ways and see whether the winner changes. That is one extra pass through a comparison already being made, and it converts the eighth from a probability into an observation.
What is not observable is the cost. Whether the change of winner made the answer better or worse needs the truth, which is what a replicate supplies here and what nothing supplies in practice. So the honest advice that comes out of this field is not a correction to apply — it is a diagnostic to run. Check whether the tuning convention changes which model gets reported, because on an eighth of samples it does and nothing else in the analysis says so.
That is a smaller recommendation than a regret figure and a more usable one, and it is available precisely because the field split the product. The two cheap factors are the two a practitioner can compute.
What an eighth is, and is not
Twelve per cent of draws, at 0.03145 of an error a little above one apiece, is under half a per cent of error on average — which is where the earlier field’s flat regret came from and is a number small enough to ignore.
It should not be ignored, and the reason is the shape rather than the size.
A small average made of a large effect on a few draws is a different object from a small effect everywhere. On seven draws in eight, sharing the tuning parameter across the table changes nothing a reader would notice; on the eighth, it changes which model is being reported. A practitioner running the analysis once has an eighth chance of the tuning convention deciding their answer, and no way to know which case they are in — the two look identical from inside a single sample, because the alternative rule is not run.
That is the same shape as a searched break’s charge and a forking path: a decision that is invisible on most datasets and decisive on some, whose average effect is small and whose conditional effect is the whole of the analysis. The average is the right thing to report and the wrong thing to reason from.
And an eighth is not a small probability for a convention nobody states. No paper that fits five models through an estimated whitening says which residuals the whitening came from. The choice is made silently in code, it is the kind of choice a second analyst would make differently, and it moves the reported model on one dataset in eight.
The three quantities, and which field each belongs to
It is worth setting the three side by side one last time, because they have been arrived at in the order they could be measured rather than the order they explain things.
The disagreement rate belongs to the list. It rises from nothing at two values to 43.3% at eight, and it would go on rising with a longer list. It is a property of how finely the tuning parameter may be recorded and it says nothing about how uncertain the tuning parameter is — the five residual series are as different at two values as at eight, and the criterion is as flat.
The turnover belongs to the criterion. It is the probability that a change of whitening reorders the two best candidates, and it falls as the list is refined because refined lists produce similar whitenings. It is also the quantity that moves most between laws — 13.7% under long memory against 28.5% under an autoregression — because how alike the candidates’ chosen whitenings are is a fact about the dependence.
And the decisive cost belongs to the table. Three per cent of the error, which is what changing the selected model is worth when it happens, and which the split cannot measure cleanly because the draws where the winner is changeable are the draws where any perturbation costs most.
The first is cheap and observable. The second is cheap and observable. The third is neither, and is the only one that needs a replicate. A sweep that reports their product spends the cost of the third to learn about the first, and the earlier field’s flat regret is exactly that trade made without anybody noticing it was a trade.
What would move it
Three things would, and none of them is the list.
A table whose candidates are closer together. The turnover is the probability that a change of whitening reorders the two best candidates, so it is a function of how close their criteria are. A table of five nested candidates differing by one coefficient would have a higher turnover than this one at every list length.
A sample size. Everything here is a hundred and twenty rows. A longer sample estimates the dependence better, so the five candidates’ chosen values would converge and the disagreement rate would fall — but it also separates the candidates’ criteria, so the turnover would fall too. Which effect wins is not measured, and it is the obvious sweep to run next.
And a criterion with a steeper penalty. The turnover is decided by how nearly indifferent the criterion is between candidates, which is a property of the charge it levies. The field that derived a charge for a covariance’s dimension found the penalised objective flat enough that its argmax was noise; the same flatness is what makes a change of whitening able to reorder a table.
None of the three is a dial anybody turns deliberately, which is the last thing worth saying about the invariant: it is stable against the parameter people do think about and not obviously stable against three they do not.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A step that is not a ratio — both name bandwidth selection, granularity, information criterion, model selection, monte carlo, regret, selection effect, tuning parameter
- A window for every candidate — both name bandwidth selection, information criterion, long-run variance, model selection, regret, selection effect, tuning parameter, whitening
- A charge that reads the draw — both name bandwidth selection, estimation error, information criterion, model selection, monte carlo, regret, selection effect
- A lag the sample has less of — both name bandwidth selection, estimation error, information criterion, long-run variance, model selection, monte carlo
- The reversal that was the instrument's — both name estimation error, long-run variance, monte carlo, persistence, robustness, tuning parameter
- The width a band is measured in — both name bandwidth selection, estimation error, information criterion, long-run variance, model selection, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Bandwidth selectionDiscretenessEstimation errorGranularityInformation criterionLong memoryLong-run varianceModel selectionMonte CarloPersistenceRegretRobustnessSelection effectTuning parameterWhitening