How long the list is

A list is not a rule

How often five candidates disagree about a tuning parameter runs from nothing at two values on the list to two draws in five at thirteen. What the disagreement costs does not move at all.

Worth reading first: A design is a number · The observations that repeat each other.

The window field’s reason for not comparing a window against an order was the list: eight values against thirteen, and a rule minimising a criterion over more options is a different rule. It is a good objection. It has a mechanism, it predicts a direction, and it is the kind of thing that is usually right.

It is testable, and testing it takes both halves — what the list does to the mechanism, and what it does to the outcome.

The mechanism responds

At a true null, where every candidate contains the truth and nothing can be discovered, the five candidates choose their own tuning parameter and the question is whether they choose the same one.

With two values on the order list they always do. With four, they disagree on 18% of draws; with eight, 37%; with thirteen, 39%. The sequence is 0.0, 6.7, 18.3, 21.7, 24.2, 25.0, 36.7, 35.8, 32.5, 38.3, 39.2 and 39.2 per cent as the list runs from two values to thirteen — smooth, monotone within noise, and exactly the shape the objection predicts.

What a longer list actually changes. How often the five candidates choose different tuning parameters, at a true null where every one of them contains the truth, so a disagreement is manufactured rather than discovered. The order's list is an interval of integers, and thinning it moves the rate smoothly from 0% at two values to 39% at thirteen. The window's is not an interval — it runs 0, 1, 2, 4, 8, 12, 20, 30 — so a thinned window list jumps depending on whether it happens to keep the width the criterion wants, between 0% and 42% with no order to it. So "the same length" was never quite the same thing for the two rules, and it is a smaller effect than the field it was invoked to explain.
Fig. 1 How often the candidates choose differently, against how many values they may choose from, for both tuning parameters.

That is the acting mechanism the objection is about, drawn. A longer list gives five criteria more places to land in, they land in more of them, and the criteria the table then compares were computed under more different error models.

The outcome does not

What the disagreement costs is another matter.

At two values on the order list the cost of choosing per candidate is exactly zero — not approximately, since with nothing to disagree about the per-candidate rule is the shared rule. At four values it is 0.00406, at eight 0.00415, at thirteen 0.00360, each with a standard error between 0.0020 and 0.0021.

Three readings inside a fifth of a standard error of each other, across a list that has more than tripled in length and a disagreement rate that has doubled.

The list is not what separates them. What choosing the tuning parameter separately for every candidate costs, against how many values the rule may choose from, under AR(1) at 0.8. The earlier field's reason for not comparing a window against an order was that their lists are different lengths, and matching them changes very little: at four values the window costs 0.00223 and the order 0.00406; at eight, 0.00401 and 0.00415. Each rule's own list is marked. The two curves are inside each other's standard errors from four values on, and neither has a slope worth the name. At two values the order's cost is exactly zero, because a list of two leaves the five candidates nothing to disagree about — which is the shape of the mechanism and the whole of the list's contribution to it. What the list decides is how often the candidates disagree; what it does not decide is what the disagreement costs.
Fig. 2 The cost against the number of values on the list, for both tuning parameters, with each rule’s own list marked.

The window’s version has the same shape with a coarser grid: 0.00016 at two values, 0.00223 at four, 0.00401 at eight. It rises out of zero as soon as disagreement becomes possible and then stops responding.

So the answer to the objection is that the list length switches the mechanism on and does not then scale it. Between “the candidates can disagree” and “the candidates disagree twice as often” there is nothing measurable to find.

The term that cancels, and the term that does not. The volume each candidate's whitening moves — log|Ω̂| — for a sieve of order 4 on one sample of 120 rows. Estimated once from the fullest candidate and used for the whole table, it is the same number for every candidate, so it drops out of every difference the criterion reads: that is why nothing in this collection has ever needed to carry it. Estimated from each candidate's own residuals it ranges over 23.87, which is more than a parameter is worth, and the criteria being compared are then fits made under different error models with no term saying so. The window's rule has carried this term since the estimated-covariance field and the sieve's never had it.
Fig. 3 What differs between candidates once the tuning parameter is theirs, which is what the disagreement rate is a rate of.

Why a mechanism can respond and an outcome not

The two quantities are separated by an average, and it is worth naming which one.

The cost of a per-candidate rule is the extra regret it delivers: the difference between the model it hands over and the model a shared rule hands over, in expected error. Most draws hand over the same model — the candidates disagree about the tuning parameter far more often than that disagreement changes which candidate wins — and on those draws the cost is exactly nothing.

So the cost is a rate times a size. A longer list raises the rate at which the criteria are incomparable, and lowers the size of each incomparability, because the extra values a longer list adds are the ones between the values already there. A list of {0, 30} and a list of {0, 1, 2, 4, 8, 12, 20, 30} disagree about very different amounts when they disagree.

That is the honest reading, and it is a reading rather than a measurement: the sweep here shows the product being flat and does not separate the two factors. Separating them would need the size of each disagreement’s effect measured conditional on there being one, and on these sweeps the conditional sample is too small to resolve it.

How far apart the windows a table wants are. The gap between the largest and the smallest window fifteen candidates ask for, on the same sample, under AR(1) at 0.8 over 150 draws. One window for the table has a spread of zero by construction. Letting each candidate choose its own gives 1.73 lags of spread, which is the size of the thing being chosen: the list runs from 0 to 30. That spread is not a nuisance to be averaged away — it is the whole of the displacement, because a candidate that can move its window further can manufacture more criterion than one that cannot.
Fig. 4 How far apart the candidates’ choices are when they disagree, which is the second factor of the product.

One list is an interval and one is not

The objection has a second half that turns out to be more substantial than its first, and it is visible in the same figure.

The order list is the integers from zero to twelve. Thinning it to mm values takes every 13/m13/m-th integer, which is a coarser grid over the same range and behaves like one: the disagreement rate rises smoothly and monotonically.

The window list is {0, 1, 2, 4, 8, 12, 20, 30} — geometric-ish, because a bandwidth’s useful range is multiplicative. Thinning that to mm values gives {0, 30}, then {0, 8, 30}, then {0, 2, 12, 30}, and the disagreement rate goes 0.8%, 0.0%, 20.0%, 8.3%, 40.8%, 8.3%, 41.7%. It jumps by a factor of five between adjacent lengths and back again, depending on whether the thinned list happens to keep the widths the criterion actually wants.

The window a whitening wants is not the memory of the errorsRegret under AR(1) at 0.8 as the tapered estimate is given more lags, over 120 draws at n = 120. The best window is L = 12; the automatic bandwidth is 4 and the error model's own likelihood chooses 5.2 on average. Both land in the same place and both are short, and the reason is the taper: a Bartlett weight at lag k is 1 − k/(L + 1), so a window of 8 keeps 0.556 of whatever the fourth lag carries and a window of 30 keeps 0.871. A window has to be several times the memory before it stops removing the memory. The dashed line is the rule told the errors are a first-order autoregression, which needs no window at all.0.0200.0300.0400.05000.3010.6020.9031.081.301.481.601.78log₁₀ of the window — how many lags the estimate carriesregret against the best model availabletold the form, no windowbest at L = 12the automatic bandwidth120 draws, geometricthe choices all land left of the optimum
Fig. 5 Which widths a criterion wants, which is what a thinned window list either keeps or loses.

So the same length was never quite the same thing for the two rules, and the earlier field was right to be uneasy about matching them. What it could not know without measuring is that the unease is about a quantity that does not reach the outcome.

The number the comparison was missing. What it costs to choose the tuning parameter for every candidate separately rather than once for the table, under AR(1) at 0.8, paired on the draw. The window's figure is the one the earlier field reported; the order's is the one it named and did not make. They are the same size — 0.00401 against 0.00360, at 2.30 and 1.72 paired standard errors — and matching the lists at eight values leaves them the same size again. The prediction that the longer list would make the order's cost the larger of the two is not what happens; what happens is that the two rules cost the same once they are scored by the same criterion, which took a missing term to arrange.
Fig. 6 The two rules at their own lists and matched, which is what the length was supposed to explain.

What a minimum over a list is doing

The mechanism deserves one paragraph of arithmetic, because the intuition behind the objection is a piece of order statistics and the arithmetic says how strong it should be.

A candidate’s tuning criterion evaluated at mm values is mm correlated draws, and the rule takes their minimum. The expected minimum of mm independent standard normals is 0.56 at two, 1.42 at eight and 1.67 at thirteen — the familiar 2logm\sqrt{2\log m} shape, which at these lengths overstates it by about forty per cent and is worth not using. So going from eight values to thirteen buys about 0.25 of a standard deviation more manufactured criterion, which is a sixth more, on a quantity that is already a small part of the difference between candidates.

Against a disagreement rate that doubles over the same range, that is the first sign the two quantities are not the same quantity. The rate is about where the minima land relative to each other; the manufactured amount is about how far below a typical value they land; and only the first of those is what makes two candidates’ criteria incomparable.

This collection has the same distinction elsewhere. The best of a set against the cost of having searched for it is about the size of a maximum, and what a search costs in parameters is about how that size is charged. Here a third quantity — how unevenly the size is distributed across the things being compared — turns out to be the one that acts.

The rate and the size, separated after all

The reading offered above — a longer list raises how often the criteria are incomparable and lowers what each incomparability is worth — is deferred as unmeasurable, and it is available from the numbers already on the page. The cost is exactly zero when the candidates agree, so E[cost] = P(disagree) × E[cost | disagree] holds as an identity, and dividing one column by the other gives the conditional size directly.

For the order: 0.00406/0.183 = 0.0222, 0.00415/0.367 = 0.0113, 0.00360/0.392 = 0.0092 at four, eight and thirteen values. For the window: 0.00016/0.008 = 0.0200, 0.00223/0.200 = 0.0112, 0.00401/0.417 = 0.0096.

Both fall by more than half as the rate roughly doubles, which is the mechanism the essay proposes. And the two sequences agree with each other at every matched rate — 0.0222 against 0.0200 near a fifth, 0.0113 against 0.0112 near two fifths, 0.0092 against 0.0096 at the top. Two different tuning parameters, two differently-shaped lists, one conditional size.

That is a stronger finding than the flat product it explains. The product is flat because two factors cancel; the factors are the same two factors for both rules; and the size that falls is a property of how finely the list is divided rather than of what the list is a list of. A longer list adds values between the values already there, so a disagreement about which of two adjacent entries to take is worth less the closer together they are — and it is worth the same amount whether the entries are integers or a geometric ladder.

A grid and a ladder, counted

The two lists’ disagreement sequences differ in a way that can be counted rather than described.

Thinning the order list gives 0.0, 6.7, 18.3, 21.7, 24.2, 25.0, 36.7, 35.8, 32.5, 38.3, 39.2 and 39.2 — eleven steps with two changes of sign. Thinning the window list gives 0.8, 0.0, 20.0, 8.3, 40.8, 8.3 and 41.7 — six steps with five.

Five reversals in six against two in eleven is the difference between a grid and a ladder, and it is the sharpest form the essay’s second objection takes. A thinned interval is a coarser instrument of the same kind; a thinned ladder is a different instrument each time, and which widths survive the thinning decides everything.

The order sequence also says where its mechanism lives. Going from four values to eight buys 18.4 points of disagreement and going from eight to thirteen buys 2.5. Almost the whole of the effect is switched on by the eighth value, which is the list the window rule actually uses.

What this says about matched comparisons generally

There is a general shape here worth extracting, because “match the two procedures on the obvious axis” is the standard move and it has two failure modes rather than one.

Matching can be impossible in kind. Two lists of eight values are not comparable when one is an interval of integers and the other is a geometric ladder: the second’s resolution changes with where in the range it is, so an eighth of its length means something different at each end.

And matching can be unnecessary. If the axis being matched does not reach the outcome, the matched comparison and the unmatched one give the same answer, and the effort spent constructing the matched version bought a reassurance rather than a correction.

Both apply here, which is why the field’s conclusion is a pair: the comparison is worth making, and the reason it had been avoided was not the reason it was hard. What made it hard was a missing term in one of the two criteria, and no amount of matching lists would have found that.

A difference between criteria, read as a difference between rules. The extra error from choosing a tuning parameter per candidate, under AR(1) at 0.8, three ways. Scored by the criterion this collection has used for a sieve since the estimated-covariance field — which carries no determinant — a per-candidate order costs -0.00011, at -0.16 standard errors, and the honest reading of that is nothing. Scored by the criterion the window's rule has always used, which carries the volume its whitening moves, the same rule costs 0.00415 at 2.02 — the window's own 0.00401. The rules were never different. The criteria were.
Fig. 7 What was actually separating the two rules, which the list length has nothing to do with.

The objection was about a real thing

It is worth being fair to the sentence this essay tests, because dismissing it would be the wrong reading.

Making the two lists the same length changes what each rule is is true. An order rule that may not fit past the eighth lag is a different rule from one that may fit to the twelfth, and on a law with a long tail — long memory at d=4/9d = 4/9, whose autocorrelation is still 0.637 at the eighth lag — the difference between those two rules is not small. The matching used here keeps each rule’s reach by thinning across the whole range rather than cutting the tail off, precisely because cutting the tail off would compare a rule that can reach the truth against one that cannot.

That is a decision, and it is the decision that makes the comparison mean anything. It is also the one place the objection’s worry bites: a thinned list is a coarser instrument, and a coarser instrument on a rule whose optimum sits between two of the remaining values is a rule that has been damaged in a way the length alone does not describe.

The list a practitioner should use

A different question from either of the above, and the sweep answers it incidentally.

The rules here are scored on delivered regret, so a list that reaches the values the criterion wants is better than one that does not, whatever it costs in per-candidate disagreement. On the window’s side the two-value list {0, 30} delivers 0.02221 against the eight-value list’s 0.02669 — the short list is better, because half of it is “no whitening at all” and the other half is a width so wide that the criterion rarely takes it, so the rule mostly falls back to least squares and least squares is not much worse than a badly-tuned whitening.

That is a warning about reading these sweeps too directly. The shared column is not a recommendation either: it moves with the list because the shared rule chooses from the same list.

What the field is measuring is the difference between the shared and own columns at a fixed list, and that difference is what does not move.

The one place the length does decide something

Two values on the list. There the per-candidate rule is the shared rule, and the cost is exactly zero rather than nearly zero — the order’s column reads 0.00000 with a standard error of 0.00000 across a hundred and twenty draws.

That is worth having in the field because it is the only entry in it that is exact, and it anchors the interpretation of everything else: the cost is caused by disagreement, since removing the possibility of disagreement removes the cost entirely. Without that row the flat curve would be consistent with the cost coming from something else the list happens not to move.

It is the same kind of anchor as a true null in a search field: a setting where the answer is known by construction, used to check that the machinery reports the known answer before any unknown one is believed.

Two readings a reader might take, one of which is wrong

Right: the difference between choosing a tuning parameter once and choosing it per candidate is about four thousandths of regret, it is caused by the candidates disagreeing, and how many values they may disagree over changes how often that happens without changing what it costs.

Wrong: longer lists are free. Nothing here says that. Every measurement is of the difference between a shared rule and a per-candidate rule at a fixed list, and both rules move when the list does. The window’s shared column runs 0.02221, 0.02356 and 0.02669 as the list goes from two values to eight — the shared rule gets worse with a longer list on this law, because the extra values give its own criterion more room to mis-tune. That is a separate effect, it is larger than anything in this essay, and it belongs to the field that prices the window itself.

What is claimed here, and what is not

This essay takes whether the length of the list is what separates two per-candidate tuning rules. The claims are that at a true null the five candidates disagree about the order on 0.0% of draws with two values on the list and 39.2% with thirteen, rising smoothly in between; that the cost of choosing per candidate is exactly zero at two values and 0.00406, 0.00415 and 0.00360 at four, eight and thirteen, all with standard errors near 0.0020, so the outcome does not respond to a length the mechanism responds to smoothly; that the window’s list is not an interval, so thinning it moves the disagreement rate between 0.0% and 41.7% with no order to the sequence; and that the two-value row’s exact zero is what ties the cost to the disagreement rather than to anything else.

What stays out, and is named as a decision: the rate and the size, separated. The reading offered above — that a longer list raises how often the criteria are incomparable and lowers how much each incomparability is worth — is consistent with the sweep and is not measured by it. Measuring it needs the regret conditional on the candidates having disagreed, and on a hundred and twenty draws with a 39% disagreement rate that conditional average has a standard error larger than the whole effect.

Also out: a list nobody would use. Every list here is a thinning of a list a practitioner might run. A list of twenty values, or one concentrated where the criterion wants it, would move the disagreement rate further and might move the cost — but a rule with a list chosen to make a point is not a rule anybody has, and this field is about a comparison between two rules that exist.

The boundary against the first essay of the field is that it reports the two costs and this one tests the reason they were never put beside each other. The reason was reasonable and does not survive being measured, which is the most ordinary outcome in this collection, and it is worth a field whenever it happens.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutoregressionBandwidthClosed formDependenceInformation criterionModel selectionMonte CarloNuisance parameterOrder statisticPaired comparisonRegretResolutionSelection effectWhitening