A rate times a size
Worth reading first: A design is a number · The observations that repeat each other.
The field that measured what a per-candidate tuning parameter costs closes on a puzzle it names and declines to solve. How often five candidates disagree about the parameter rises steeply with the length of the list they choose it from — nothing at two values, 39.2% at thirteen. What the disagreement costs does not move: 0.00406, 0.00415 and 0.00360 at four, eight and thirteen values, three readings inside a fifth of a standard error of each other.
Its own last paragraph says what is wrong with reading that as an answer. The sweep measures a product. A longer list could be raising the rate while lowering what each disagreement is worth, and the two movements would cancel into exactly the flat line the table shows. One number cannot tell the two stories apart.
This essay separates them, and the separation is exact rather than approximate — which is the part worth having before any of the numbers.
Why the decomposition is an identity
The two rules being compared differ in one thing only: which value of the tuning parameter each candidate is scored at. The shared rule takes one value from the fullest candidate’s residuals and gives it to the whole table. The per-candidate rule gives each candidate the value chosen from its own.
When every candidate happens to pick the value the shared rule picked, the two rules are the same rule. Same whitening, same criteria, same winner, same coefficients, same regret. The paired difference on that draw is not small; it is zero, and it is zero in the way that 3 − 3 is zero rather than in the way an average of noise is zero. So
holds to machine precision. It is a rearrangement rather than a model.
Checked over twelve hundred draws at each list length, the largest paired difference on a draw where nothing disagreed is exactly 0, and the product matches the average to the last bit at every length. That is asserted rather than assumed, because it is the only thing the rest of this rests on: if the two rules could differ while every candidate agreed, there would be a third term and the split would be a decomposition of something else.
What the sweep could have been seeing
There are three shapes a flat product can have, and the sweep is compatible with all of them.
The rate rises and the size falls by the reciprocal factor, so the product sits still — the reading the deferral proposed and the one that would make the list a genuine dial on the problem.
The rate rises and the size does not move, so the product rises and the sweep was too coarse to see it.
Or both are flat and the flat product is exactly what it looks like.
Distinguishing them is a question about precision rather than about statistics, and it is worth being blunt about how far short the original was.
Reaching a standard error a fifth of the quantity’s own size takes 569 draws for the average and 537 for the conditional size. The sweep used 120.
The near-equality of those two numbers is worth a sentence, because the natural hope is that conditioning is the cheaper measurement. It is not. Dropping the draws that are exactly zero removes most of the sample, and the quantity that is left is larger by exactly the factor it removed them by, so the relative precision comes out the same to within a ratio of 1.06. What conditioning buys is two readings where there was one — not either of them for less.
There is a fourth shape, and it is the one that turns out to be right, but it cannot be seen from the product at all: the rate rises, the size is flat, and a third quantity that neither of them is stays constant. Arriving at that needs the split to be made first, which is why this essay stops where it does.
The two factors, separated
Twelve hundred draws at each list length, on the whitening window, under a first-order autoregression.
The rate rises. It is 28.6% at four values on the list, 40.8% at six and 43.3% at eight — a factor of 1.52 over the range where the conditional average is measurable at all.
The size does not. It is 0.00975 ± 0.00224 at four values and 0.00848 ± 0.00113 at eight, which is 0.5 standard errors apart. The point estimate is lower at the longer list, so the deferral’s guess has the sign it expected and none of the size it needed.
Half a standard error is compatible with a real fall the measurement cannot see, so the comparison is made a second time with the world taken out of it. Restricting to the 151 draws that disagree under both lists and differencing within the draw, the difference is +0.00122 ± 0.00317 — the other way round, at 0.39 standard errors. Two comparisons, one paired and one not, that do not agree on the sign: which is what a quantity that is not moving looks like.
So the flatness has a different cause
Rate up by half and size flat means the product should be up by half, and it is up by a third: 0.00279, 0.00341, 0.00368 at four, six and eight values. That rise is not visible at 120 draws and is barely visible at 1,200 — its own standard error at the full list is about a tenth of it.
The sweep’s flat line was its precision, not its subject. Nothing was cancelling.
Which leaves the real question, and it is not the one the deferral asked. If a longer list makes disagreements half again as common, and each one costs the same, what is it that stays the same about a disagreement when the list changes underneath it? The values being disagreed about certainly do not.
At four values the list is 0, 2, 12, 30 and a disagreement spans 10.00. At eight values the list is 0, 1, 2, 4, 8, 12, 20, 30 and it spans 4.14. The candidates are essentially always neighbours on whatever list they are handed — 1.00 steps apart at four values and 1.04 at eight — so refining the list refines the quarrel, by a factor of 2.4 here.
And the cost does not follow it. A disagreement about whether the window is 8 or 12 is worth what a disagreement about whether it is 2 or 12 is worth. That is the finding this field hands on, and the essay that takes it apart shows why: what a disagreement costs is not a function of how far apart the values are, because the cost is not incurred by the values at all.
What a rate is, when the rule is the thing being measured
It is worth pausing on what the disagreement rate is a rate of, because the phrase makes it sound like a property of the data and it is not.
Five candidates are fitted, each leaves its own residuals, and each residual series is handed the same list and asked to pick a value from it. The rate is how often those five answers are not all identical. Nothing about the world has to change for that number to move: the residuals are what they were, the criterion is what it was, and lengthening the list changes only how finely the five answers may be recorded. At two values on the list the five can only disagree by disagreeing about a very coarse question, and they almost never do — 2.8% of draws, on thirty-three of twelve hundred, which is too few to average a conditional over and is why that row is not in the comparison. At three values the list is 0, 8, 30 and they disagree on none of the twelve hundred draws.
So the rate is an artefact of the recording rather than a fact about the sample, and that is not a criticism of it. It is exactly what makes it the right thing to sweep against the list, and it is why the field that measured it reads it as the mechanism. What it is not is a measure of how uncertain the tuning parameter is: the five residual series are as different at two values on the list as at eight, and the criterion is as flat.
That distinction is what makes the flat conditional size surprising rather than obvious. If the rate were tracking genuine uncertainty, a longer list would be surfacing more marginal disagreements and the average disagreement would get less serious. It is not, and it does not.
The other tuning parameter, which is not the same
Every number above is the whitening window. The autoregressive order is the collection’s other tuning parameter and it is the one the earlier field could not put beside it until the lists were matched.
Its rate rises further: 17.8% at four values to 48.9% at thirteen, a factor of 2.74. Its conditional size is 0.01100 ± 0.00218 at four and 0.00637 ± 0.00099 at thirteen — a fall of 0.00463 at 1.9 standard errors, which is not nothing and is not two.
So the two parameters do not answer the same way, and the difference is in the lists rather than in the parameters. An order list is an interval of integers, so lengthening it adds values at the far end; a window list is a geometric ladder, so lengthening it adds values in the middle. The order’s candidates disagree by 2.99 steps at thirteen values against 1.20 at four, where the window’s stay at one step throughout — the order’s list grows a tail the candidates can spread out along, and the window’s grows a middle they cannot.
Which is a warning about how far any of this transports. The clean statement is about one parameter and one list shape, and the second parameter agrees about the rate, disagrees about the size at two standard errors, and has a visibly different mechanism underneath.
Where a zero comes from, and why there are so many of them
One more thing follows from the identity and it changes how the whole sweep should be read.
At the full list the candidates agree on 56.7% of draws. On every one of those the two rules are the same rule and the regret is exactly zero — not approximately, not on average. So the average the sweep reports is already, before anything is measured, a number that is zero more than half the time and something else the rest of the time.
That has a consequence for every standard error in the earlier field. A mean over a sample that is mostly a point mass at zero has a spread driven by how many non-zeros happened to land in it, and the usual rule-of-thumb reading of a standard error — that the quantity is within two of them — is fine, but the reading of a difference between two such means is not, because two list lengths share their zeros and do not share their non-zeros. That is exactly why the paired comparison above is worth making separately: pairing on the draws that disagree under both lists removes the shared zeros and the shared world, and it moves the estimate from −0.00127 ± 0.00251 to +0.00122 ± 0.00317 without moving anything about the data.
A quantity that is zero on most draws is not well summarised by its mean, and this one is zero on most draws by construction rather than by luck. The field that measured it reports its regrets to five decimal places with standard errors attached and every one of them is a rate times a size wearing one number’s clothing.
The other parameter’s product, which is the case that was hypothesised
The order’s two factors are reported and its product is not, and multiplying them out says the two tuning parameters land on opposite sides of the question the deferral asked.
At four values the order’s rate is 17.8% and its conditional size 0.01100, so the product is 0.00196; at thirteen values, 48.9% and 0.00637 give 0.00312. The rate rose by a factor of 2.74 and the size fell by 0.579, and the product rose by 1.59.
So for the order the size is absorbing part of the rate’s rise — 42% of it — which is exactly the cancellation the deferral proposed and exactly what the window does not do. The window’s three sizes are 0.00975, 0.00836 and 0.00848 at four, six and eight values: a spread of 0.0014 against standard errors of 0.001 to 0.002, and not monotone, which is what noise looks like and what a trend does not.
That makes the field’s closing warning about transportability sharper than it states it. It is not merely that the two parameters have different mechanisms; it is that one of them exhibits the phenomenon the sweep was deferred to investigate and the other does not, at nearly two standard errors apart. A single sentence about what a longer list does to a disagreement’s cost would be right for the order, wrong for the window, and unfalsifiable from either sweep’s product.
Why 120 draws could not have seen it
The budget figure prices the original sweep, and running it backwards says how badly.
A standard error a fifth of the quantity’s own size takes 569 draws; the sweep used 120. Standard errors scale as 1/√n, so the sweep’s was 0.2 × √(569/120) = 0.44 of the number it was reporting. The rise it failed to see is 32% of that number, spread over three list lengths — so the signal was about seven tenths of one standard error at the endpoints and rather less between them.
A sweep whose standard error is nearly half its own quantity cannot report a third of a rise, and the flat line was not evidence of flatness in any direction. Stating it that way is more useful than saying the measurement was too coarse, because the ratio is computable before any draws are taken: a rate and a conditional size need about five hundred draws apiece, and the product needs the same.
The gap column can be given the same treatment and it sharpens the essay’s closing claim. The values being disagreed about get 2.42 times closer as the list grows — a span of 10.00 at four values against 4.14 at eight — while the cost of a disagreement moves by a factor of 0.87. Divide the two: the cost per unit of window disagreed about rises from 0.000975 to 0.002048, a factor of 2.1.
So the claim is not that the cost is loosely unrelated to the size of the quarrel. It is that halving the quarrel doubles what each unit of it is worth, exactly enough to leave the total where it was — which is a much stronger statement, and it is the one the essay that takes the disagreement apart has to explain.
What separating a product is worth in general
Two things, and the second is the reusable one.
A product of two quantities can be flat while neither is, and a sweep that reports only the product has no way to say so. That is a familiar warning and it is the smaller half here, because in this case nothing was cancelling: the flatness was a measurement four times too coarse for either factor.
The larger half is that the factors were not equally hard to get. The rate is a count and needs no regret at all: it is measured on the tuning values alone, at a thousandth of the cost of the sweep, and it is the factor that moves. Anything the field wanted to know about how a list behaves was available in the cheap half of the product, and it was invisible because the expensive half was multiplied into it.
That generalises past this field. Where a measured quantity is a rate times a size, the rate is usually a property of the rule and the size a property of the world, and they are usually not equally expensive. Reporting the product prices them as though they were.
The same shape is visible elsewhere in this collection once it is named. A searched break’s charge is how often the search finds something times how much it finds; a criterion’s regret against a hold-out is how often the two rules pick different candidates times what the difference is worth. Neither is reported split, and in both cases the rate is the cheap half. Splitting them is not done here and is the obvious next thing to do with the instrument.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- How often it matters — both name bandwidth selection, estimation error, granularity, information criterion, long-run variance, model selection, monte carlo, regret, selection effect, tuning parameter, whitening
- A window for every candidate — both name bandwidth selection, information criterion, long-run variance, model selection, regret, selection effect, tuning parameter, whitening
- A charge that reads the draw — both name bandwidth selection, estimation error, information criterion, model selection, monte carlo, regret, selection effect
- A lag the sample has less of — both name bandwidth selection, closed form, estimation error, information criterion, long-run variance, model selection, monte carlo
- The width a band is measured in — both name bandwidth selection, closed form, estimation error, information criterion, long-run variance, model selection, monte carlo
- What the correction assumes — both name bandwidth selection, closed form, estimation error, information criterion, long-run variance, model selection, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Bandwidth selectionClosed formEstimation errorGranularityInformation criterionLong-run varianceModel selectionMonte CarloPaired comparisonRegretSelection effectStandard errorTuning parameterVariance reductionWhitening