The rate and the size of a disagreement

A rate times a size

A sweep reported what it costs to let every candidate choose its own tuning parameter and found it flat across the list. It was reporting a product, and the two things multiplied together do not behave the same way at all.

Worth reading first: A design is a number · The observations that repeat each other.

The field that measured what a per-candidate tuning parameter costs closes on a puzzle it names and declines to solve. How often five candidates disagree about the parameter rises steeply with the length of the list they choose it from — nothing at two values, 39.2% at thirteen. What the disagreement costs does not move: 0.00406, 0.00415 and 0.00360 at four, eight and thirteen values, three readings inside a fifth of a standard error of each other.

Its own last paragraph says what is wrong with reading that as an answer. The sweep measures a product. A longer list could be raising the rate while lowering what each disagreement is worth, and the two movements would cancel into exactly the flat line the table shows. One number cannot tell the two stories apart.

This essay separates them, and the separation is exact rather than approximate — which is the part worth having before any of the numbers.

Why the decomposition is an identity

The two rules being compared differ in one thing only: which value of the tuning parameter each candidate is scored at. The shared rule takes one value from the fullest candidate’s residuals and gives it to the whole table. The per-candidate rule gives each candidate the value chosen from its own.

When every candidate happens to pick the value the shared rule picked, the two rules are the same rule. Same whitening, same criteria, same winner, same coefficients, same regret. The paired difference on that draw is not small; it is zero, and it is zero in the way that 3 − 3 is zero rather than in the way an average of noise is zero. So

E[regret]=P(disagree)×E[regretdisagree]\mathbb{E}[\text{regret}] = \mathbb{P}(\text{disagree}) \times \mathbb{E}[\text{regret} \mid \text{disagree}]

holds to machine precision. It is a rearrangement rather than a model.

The two routes to one number, which are the same arithmetic. The average regret from letting each candidate choose its own tuning parameter, computed twice: directly as a mean over all 1200 draws, and as the disagreement rate times the average regret on the draws that disagreed. The two agree to machine precision at every list length, and that is not an approximation — when every candidate picks the value the shared rule picks, the two rules are the same rule and the paired difference is exactly zero. So the average is a rate times a size by construction, and a sweep that reports only the average has multiplied two things together and printed the product. The product runs 0.00279, 0.00341, 0.00368 across the list lengths here, which is the flatness the field inherited and is the thing that has to be taken apart rather than explained.
Fig. 1 The average computed twice — as a mean over every draw, and as the rate times the conditional size. The two are identical rather than close.

Checked over twelve hundred draws at each list length, the largest paired difference on a draw where nothing disagreed is exactly 0, and the product matches the average to the last bit at every length. That is asserted rather than assumed, because it is the only thing the rest of this rests on: if the two rules could differ while every candidate agreed, there would be a third term and the split would be a decomposition of something else.

What the sweep could have been seeing

There are three shapes a flat product can have, and the sweep is compatible with all of them.

The rate rises and the size falls by the reciprocal factor, so the product sits still — the reading the deferral proposed and the one that would make the list a genuine dial on the problem.

The rate rises and the size does not move, so the product rises and the sweep was too coarse to see it.

Or both are flat and the flat product is exactly what it looks like.

Distinguishing them is a question about precision rather than about statistics, and it is worth being blunt about how far short the original was.

Neither quantity was measurable on the budget that reported them. How many draws each quantity needs for a standard error a fifth of its own size, computed from the spreads measured over 1200 draws at the full list. The average the sweep reports needs 569; the conditional size needs 537. The sweep used 120. The conditional is not the cheaper measurement and the near-equality is not a coincidence: it drops the draws that are exactly zero, which is 56.7% of them, and it is larger than the average by exactly the factor it drops them by, so the two effects cancel to a ratio of 1.06. What separating the two buys is a reading of two quantities where there was one, not a reading of either for less — and the flat product the sweep reported was not a finding about the world, it was the precision the budget bought.
Fig. 2 How many draws each quantity needs for a standard error a fifth of its own size, against what the sweep spent.

Reaching a standard error a fifth of the quantity’s own size takes 569 draws for the average and 537 for the conditional size. The sweep used 120.

The near-equality of those two numbers is worth a sentence, because the natural hope is that conditioning is the cheaper measurement. It is not. Dropping the draws that are exactly zero removes most of the sample, and the quantity that is left is larger by exactly the factor it removed them by, so the relative precision comes out the same to within a ratio of 1.06. What conditioning buys is two readings where there was one — not either of them for less.

There is a fourth shape, and it is the one that turns out to be right, but it cannot be seen from the product at all: the rate rises, the size is flat, and a third quantity that neither of them is stays constant. Arriving at that needs the split to be made first, which is why this essay stops where it does.

The two factors, separated

Twelve hundred draws at each list length, on the whitening window, under a first-order autoregression.

One factor moves and the other does notThe two factors of the same average, each drawn against its own largest value so that they share an axis. The rate at which the five candidates disagree about the tuning parameter rises from 28.6% at 4 values on the list to 43.3% at 8, a factor of 1.52. What a disagreement costs, given that there was one, is 0.00975 ± 0.00224 and 0.00848 ± 0.00113 at the same two points — 0.5 standard errors apart, and the paired comparison on the draws that disagree under both lists puts it the other way. The guess this field was written to test was that a longer list makes disagreements commoner and each one smaller. The first half is right and there is no second half.00.2500.5000.7501468values on the listeach factor, against its own largesthow often they disagree, of 43.3%what it costs when they do, of 0.014241200 draws a point, AR(1) at 0.8one rises, one sits still
Fig. 3 The two factors of one average, each against its own largest value. The slider changes which tuning parameter is being chosen.

The rate rises. It is 28.6% at four values on the list, 40.8% at six and 43.3% at eight — a factor of 1.52 over the range where the conditional average is measurable at all.

The size does not. It is 0.00975 ± 0.00224 at four values and 0.00848 ± 0.00113 at eight, which is 0.5 standard errors apart. The point estimate is lower at the longer list, so the deferral’s guess has the sign it expected and none of the size it needed.

Half a standard error is compatible with a real fall the measurement cannot see, so the comparison is made a second time with the world taken out of it. Restricting to the 151 draws that disagree under both lists and differencing within the draw, the difference is +0.00122 ± 0.00317 — the other way round, at 0.39 standard errors. Two comparisons, one paired and one not, that do not agree on the sign: which is what a quantity that is not moving looks like.

So the flatness has a different cause

Rate up by half and size flat means the product should be up by half, and it is up by a third: 0.00279, 0.00341, 0.00368 at four, six and eight values. That rise is not visible at 120 draws and is barely visible at 1,200 — its own standard error at the full list is about a tenth of it.

The sweep’s flat line was its precision, not its subject. Nothing was cancelling.

Which leaves the real question, and it is not the one the deferral asked. If a longer list makes disagreements half again as common, and each one costs the same, what is it that stays the same about a disagreement when the list changes underneath it? The values being disagreed about certainly do not.

A closer quarrel, at the same price. How far apart the values the candidates pick are, given that they picked different ones, against the length of the list. At 4 values the list is 0, 2, 12, 30 and a disagreement spans 10.00; at 8 values the list is 0, 1, 2, 4, 8, 12, 20, 30 and it spans 4.14. The mechanism is printed beside each point: the candidates disagree by 1.00 steps of the list at the short end and 1.04 at the long one, so they are essentially always neighbours on whatever list they are handed, and refining the list refines the quarrel. What does not follow, and what the next figure shows does not happen, is any fall in what the quarrel costs. A disagreement about whether the window is 8 or 12 is worth what a disagreement about whether it is 2 or 12 is worth.
Fig. 4 How far apart the values the candidates pick are, given that they picked different ones — and how many steps of the list that is.

At four values the list is 0, 2, 12, 30 and a disagreement spans 10.00. At eight values the list is 0, 1, 2, 4, 8, 12, 20, 30 and it spans 4.14. The candidates are essentially always neighbours on whatever list they are handed — 1.00 steps apart at four values and 1.04 at eight — so refining the list refines the quarrel, by a factor of 2.4 here.

And the cost does not follow it. A disagreement about whether the window is 8 or 12 is worth what a disagreement about whether it is 2 or 12 is worth. That is the finding this field hands on, and the essay that takes it apart shows why: what a disagreement costs is not a function of how far apart the values are, because the cost is not incurred by the values at all.

What a rate is, when the rule is the thing being measured

It is worth pausing on what the disagreement rate is a rate of, because the phrase makes it sound like a property of the data and it is not.

Five candidates are fitted, each leaves its own residuals, and each residual series is handed the same list and asked to pick a value from it. The rate is how often those five answers are not all identical. Nothing about the world has to change for that number to move: the residuals are what they were, the criterion is what it was, and lengthening the list changes only how finely the five answers may be recorded. At two values on the list the five can only disagree by disagreeing about a very coarse question, and they almost never do — 2.8% of draws, on thirty-three of twelve hundred, which is too few to average a conditional over and is why that row is not in the comparison. At three values the list is 0, 8, 30 and they disagree on none of the twelve hundred draws.

So the rate is an artefact of the recording rather than a fact about the sample, and that is not a criticism of it. It is exactly what makes it the right thing to sweep against the list, and it is why the field that measured it reads it as the mechanism. What it is not is a measure of how uncertain the tuning parameter is: the five residual series are as different at two values on the list as at eight, and the criterion is as flat.

That distinction is what makes the flat conditional size surprising rather than obvious. If the rate were tracking genuine uncertainty, a longer list would be surfacing more marginal disagreements and the average disagreement would get less serious. It is not, and it does not.

The other tuning parameter, which is not the same

Every number above is the whitening window. The autoregressive order is the collection’s other tuning parameter and it is the one the earlier field could not put beside it until the lists were matched.

Its rate rises further: 17.8% at four values to 48.9% at thirteen, a factor of 2.74. Its conditional size is 0.01100 ± 0.00218 at four and 0.00637 ± 0.00099 at thirteen — a fall of 0.00463 at 1.9 standard errors, which is not nothing and is not two.

So the two parameters do not answer the same way, and the difference is in the lists rather than in the parameters. An order list is an interval of integers, so lengthening it adds values at the far end; a window list is a geometric ladder, so lengthening it adds values in the middle. The order’s candidates disagree by 2.99 steps at thirteen values against 1.20 at four, where the window’s stay at one step throughout — the order’s list grows a tail the candidates can spread out along, and the window’s grows a middle they cannot.

Which is a warning about how far any of this transports. The clean statement is about one parameter and one list shape, and the second parameter agrees about the rate, disagrees about the size at two standard errors, and has a visibly different mechanism underneath.

Where a zero comes from, and why there are so many of them

One more thing follows from the identity and it changes how the whole sweep should be read.

At the full list the candidates agree on 56.7% of draws. On every one of those the two rules are the same rule and the regret is exactly zero — not approximately, not on average. So the average the sweep reports is already, before anything is measured, a number that is zero more than half the time and something else the rest of the time.

That has a consequence for every standard error in the earlier field. A mean over a sample that is mostly a point mass at zero has a spread driven by how many non-zeros happened to land in it, and the usual rule-of-thumb reading of a standard error — that the quantity is within two of them — is fine, but the reading of a difference between two such means is not, because two list lengths share their zeros and do not share their non-zeros. That is exactly why the paired comparison above is worth making separately: pairing on the draws that disagree under both lists removes the shared zeros and the shared world, and it moves the estimate from −0.00127 ± 0.00251 to +0.00122 ± 0.00317 without moving anything about the data.

A quantity that is zero on most draws is not well summarised by its mean, and this one is zero on most draws by construction rather than by luck. The field that measured it reports its regrets to five decimal places with standard errors attached and every one of them is a rate times a size wearing one number’s clothing.

The other parameter’s product, which is the case that was hypothesised

The order’s two factors are reported and its product is not, and multiplying them out says the two tuning parameters land on opposite sides of the question the deferral asked.

At four values the order’s rate is 17.8% and its conditional size 0.01100, so the product is 0.00196; at thirteen values, 48.9% and 0.00637 give 0.00312. The rate rose by a factor of 2.74 and the size fell by 0.579, and the product rose by 1.59.

So for the order the size is absorbing part of the rate’s rise — 42% of it — which is exactly the cancellation the deferral proposed and exactly what the window does not do. The window’s three sizes are 0.00975, 0.00836 and 0.00848 at four, six and eight values: a spread of 0.0014 against standard errors of 0.001 to 0.002, and not monotone, which is what noise looks like and what a trend does not.

That makes the field’s closing warning about transportability sharper than it states it. It is not merely that the two parameters have different mechanisms; it is that one of them exhibits the phenomenon the sweep was deferred to investigate and the other does not, at nearly two standard errors apart. A single sentence about what a longer list does to a disagreement’s cost would be right for the order, wrong for the window, and unfalsifiable from either sweep’s product.

Why 120 draws could not have seen it

The budget figure prices the original sweep, and running it backwards says how badly.

A standard error a fifth of the quantity’s own size takes 569 draws; the sweep used 120. Standard errors scale as 1/√n, so the sweep’s was 0.2 × √(569/120) = 0.44 of the number it was reporting. The rise it failed to see is 32% of that number, spread over three list lengths — so the signal was about seven tenths of one standard error at the endpoints and rather less between them.

A sweep whose standard error is nearly half its own quantity cannot report a third of a rise, and the flat line was not evidence of flatness in any direction. Stating it that way is more useful than saying the measurement was too coarse, because the ratio is computable before any draws are taken: a rate and a conditional size need about five hundred draws apiece, and the product needs the same.

The gap column can be given the same treatment and it sharpens the essay’s closing claim. The values being disagreed about get 2.42 times closer as the list grows — a span of 10.00 at four values against 4.14 at eight — while the cost of a disagreement moves by a factor of 0.87. Divide the two: the cost per unit of window disagreed about rises from 0.000975 to 0.002048, a factor of 2.1.

So the claim is not that the cost is loosely unrelated to the size of the quarrel. It is that halving the quarrel doubles what each unit of it is worth, exactly enough to leave the total where it was — which is a much stronger statement, and it is the one the essay that takes the disagreement apart has to explain.

What separating a product is worth in general

Two things, and the second is the reusable one.

A product of two quantities can be flat while neither is, and a sweep that reports only the product has no way to say so. That is a familiar warning and it is the smaller half here, because in this case nothing was cancelling: the flatness was a measurement four times too coarse for either factor.

The larger half is that the factors were not equally hard to get. The rate is a count and needs no regret at all: it is measured on the tuning values alone, at a thousandth of the cost of the sweep, and it is the factor that moves. Anything the field wanted to know about how a list behaves was available in the cheap half of the product, and it was invisible because the expensive half was multiplied into it.

That generalises past this field. Where a measured quantity is a rate times a size, the rate is usually a property of the rule and the size a property of the world, and they are usually not equally expensive. Reporting the product prices them as though they were.

The same shape is visible elsewhere in this collection once it is named. A searched break’s charge is how often the search finds something times how much it finds; a criterion’s regret against a hold-out is how often the two rules pick different candidates times what the difference is worth. Neither is reported split, and in both cases the rate is the cheap half. Splitting them is not done here and is the obvious next thing to do with the instrument.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • How often it matters — both name bandwidth selection, estimation error, granularity, information criterion, long-run variance, model selection, monte carlo, regret, selection effect, tuning parameter, whitening
  • A window for every candidate — both name bandwidth selection, information criterion, long-run variance, model selection, regret, selection effect, tuning parameter, whitening
  • A charge that reads the draw — both name bandwidth selection, estimation error, information criterion, model selection, monte carlo, regret, selection effect
  • A lag the sample has less of — both name bandwidth selection, closed form, estimation error, information criterion, long-run variance, model selection, monte carlo
  • The width a band is measured in — both name bandwidth selection, closed form, estimation error, information criterion, long-run variance, model selection, monte carlo
  • What the correction assumes — both name bandwidth selection, closed form, estimation error, information criterion, long-run variance, model selection, monte carlo

Named objects

A flat tag is an object no other essay names yet.

Bandwidth selectionClosed formEstimation errorGranularityInformation criterionLong-run varianceModel selectionMonte CarloPaired comparisonRegretSelection effectStandard errorTuning parameterVariance reductionWhitening