How long the list is

The volume a whitening moves

A sieve's whitening has a determinant and this collection's criterion for it never carried one. Shared across a table the term cancels exactly, which is why nothing ever noticed; used per candidate it is worth more than a parameter and the whole comparison turns on it.

Worth reading first: A design is a number · The observations that repeat each other.

A criterion compares fitted models by their residual sums, and a whitened criterion compares them by the residual sums of transformed samples. That is a substitution most of this collection makes without comment, and it is legitimate exactly when every candidate is transformed the same way.

When it is not, the comparison has to carry a term for it. The Gaussian likelihood has that term written into it — 12logΩ-\tfrac{1}{2}\log|\Omega|, the volume the transform moves — and every criterion in this collection is a Gaussian likelihood with a charge attached.

This essay is about the one place the term was missing, why nothing found it for five rounds, and what its absence was doing to the comparison the previous essay makes.

The term, and what it is for

Whitening multiplies the sample by Ω1/2\Omega^{-1/2}. A density transformed by a linear map picks up the Jacobian of that map, so the log-likelihood of the original sample and the log-likelihood of the whitened one differ by 12logΩ-\tfrac{1}{2}\log|\Omega|: not an adjustment, a change of variables.

The parametric version of this is one number. The Prais–Winsten transform for a first-order autoregression scales the first row by 1ρ2\sqrt{1-\rho^2} and differences the rest, so the whole determinant is 12log(1ρ2)\tfrac{1}{2}\log(1-\rho^2) — one term, not nn of them, and it is the entire difference between maximising and iterating.

The general version is the accumulated pivots of a Cholesky factor and there is nothing to say about it except that it exists.

The term that cancels, and the term that does not. The volume each candidate's whitening moves — log|Ω̂| — for a sieve of order 4 on one sample of 120 rows. Estimated once from the fullest candidate and used for the whole table, it is the same number for every candidate, so it drops out of every difference the criterion reads: that is why nothing in this collection has ever needed to carry it. Estimated from each candidate's own residuals it ranges over 23.87, which is more than a parameter is worth, and the criteria being compared are then fits made under different error models with no term saying so. The window's rule has carried this term since the estimated-covariance field and the sieve's never had it.
Fig. 1 The volume each candidate’s whitening moves, estimated once for the table and estimated per candidate.

Why it cancels, and why nobody noticed

Estimated once, from the fullest candidate’s residuals, before any selection, the term is the same number for every row of the table. On the sample drawn above it is −90.864 for all fifteen candidates, and it would be the same number for any sample: one estimate, one matrix, one determinant.

A criterion is read in differences. Adding a constant to every candidate’s score changes no ranking, no choice and no regret. So a criterion that omits a shared determinant is exactly as good as one that carries it, and this is not an approximation — it is an identity.

That is why the omission survived. Every whitened criterion in this collection until now has used a shared nuisance, on purpose and for a stated reason: an error every candidate carries cancels out of every difference, which is the argument that makes an estimated nuisance affordable at all. Under that regime the missing term is provably harmless, so nothing failed, nothing looked wrong, and the two rules’ implementations were free to differ in it.

Fifteen candidates, fifteen different dependences. The lag-one autocorrelation of each candidate's own residuals, on a design whose four predictors carry persistences 0.9, 0.6, 0.3 and 0 and whose errors are autocorrelated at 0.7. The truth is 0.700 and every candidate is below it, because a fit removes the part of the errors lying in a column space that is itself slow-moving. What matters here is not the bias but the spread: the estimates run 0.6287 to 0.6716 across the table, which is 8 times the standard error of any one of them. A criterion reads differences between candidates, so a nuisance every candidate shares cancels out of every difference and one that varies like this does not.
Fig. 2 The regime in which the term cancels: one estimate, made before any selection, used by every candidate.
The number the comparison was missing. What it costs to choose the tuning parameter for every candidate separately rather than once for the table, under AR(1) at 0.8, paired on the draw. The window's figure is the one the earlier field reported; the order's is the one it named and did not make. They are the same size — 0.00401 against 0.00360, at 2.30 and 1.72 paired standard errors — and matching the lists at eight values leaves them the same size again. The prediction that the longer list would make the order's cost the larger of the two is not what happens; what happens is that the two rules cost the same once they are scored by the same criterion, which took a missing term to arrange.
Fig. 3 The comparison the term makes possible, with both rules at their own lists and at a matched one.

And why it stops cancelling

Estimated from each candidate’s own residuals, the term is fifteen different numbers. On the same sample they run from −110.659 to −86.784, a spread of 23.875.

A criterion charges 2 for a coefficient. The volume term varies by twelve times that across the table, and it varies systematically: a candidate with more predictors leaves residuals with less dependence in them, so its estimated whitening is closer to the identity and moves less volume. The term is not noise added to each row; it is a slope across the table in the same direction as the thing the table is trying to measure.

Leaving it out therefore does not merely add variance to the comparison. It biases the ranking towards candidates whose whitening moved the least — which are the fullest ones, which is the direction a per-candidate rule is already displaced in for an entirely separate reason.

What it was doing

Scored without the term, choosing a sieve order per candidate costs −0.00011 of regret at 0.2 standard errors. Scored with it, on the same draws, it costs 0.00415 at 2.0 — and the window’s rule, which has carried the term since the day it was written, costs 0.00401 at 2.3.

A difference between criteria, read as a difference between rules. The extra error from choosing a tuning parameter per candidate, under AR(1) at 0.8, three ways. Scored by the criterion this collection has used for a sieve since the estimated-covariance field — which carries no determinant — a per-candidate order costs -0.00011, at -0.16 standard errors, and the honest reading of that is nothing. Scored by the criterion the window's rule has always used, which carries the volume its whitening moves, the same rule costs 0.00415 at 2.02 — the window's own 0.00401. The rules were never different. The criteria were.
Fig. 4 The same rule under two criteria, beside the window’s rule under its own.

So the whole apparent difference between a per-candidate window and a per-candidate order was the missing term. Two rules that cost the same, one of them scored by a criterion missing a twelve-parameter slope, read as one rule costing something and the other costing nothing.

It is worth naming what that would have looked like if it had shipped. A window chosen per candidate is a real cost and an order chosen per candidate is free is a clean, quotable, false result with a plausible mechanism available for it — a sieve’s whitening is smoother, an order is a coarser dial, pick one. The absence of the term is invisible in every figure and every gate.

Where the shared value sits inside the spread

The shared term is one of the fifteen — it is the fullest candidate’s own estimate, applied to everybody — so it can be located inside the range the per-candidate rule produces, and where it sits is a check on the mechanism.

It is −90.864 against a range of −110.659 to −86.784. That is 4.08 below the least negative reading and 19.80 above the most negative: the shared value sits in the upper fifth of the spread, which is what a whitening estimated from the residuals of the fullest fit has to do.

The 4.08 is worth noticing rather than rounding away. If the slope across the table were an ordering rather than a tendency, the fullest candidate would sit at the very top and the shared value would be the maximum. It is two coefficients’ worth of charge below it, so at least one candidate with fewer predictors leaves residuals whose whitening moves less volume than the fullest fit’s does. The direction of the slope is real; the monotonicity is not, and nothing in the argument needed it.

The repair leaves the two rules indistinguishable, not merely close

“The two rules read identically after it” is measured, and the margin is worth stating because it is much tighter than the sentence implies.

With the term, the order’s rule costs 0.00415 at 2.0 standard errors, so its own standard error is about 0.0021. The window’s rule costs 0.00401 at 2.3, so its standard error is about 0.0017. The two point estimates differ by 0.00014 — three and a half per cent of either, and under a tenth of the smaller standard error.

So the repair does not narrow a gap; it removes one. Two rules that differed by a factor of thirty-eight before the term was added differ by a twentieth of a standard error after it.

The omission was suppressing the noise as well as the number

One more thing falls out of the same two readings, and it is the reason the false result looked solid rather than merely small.

Scored without the term, the comparison read −0.00011 at 0.2 standard errors, so its standard error was about 0.00055. Scored with it, the standard error is 0.0021 — nearly four times larger.

That is not a coincidence: the volume term varies from draw to draw as well as from candidate to candidate, and adding it to every score adds its draw-to-draw variation to the comparison. A criterion missing the term is measuring a quieter quantity, and quietness is what a reader reads as precision.

It also settles what would have happened if the field had simply run more draws. The unrepaired comparison could have resolved an effect of 0.0011 at two standard errors, which is a quarter of the true 0.00415 — so more draws would not have found it. The omission moved the estimand rather than adding noise to it, and no amount of counting reaches a number that is not being computed.

What the term looks like as a function of the order

The volume a sieve’s whitening moves is not an arbitrary number, and the shape of it explains both halves of the finding.

An autoregression of order pp fitted to a series implies a covariance matrix whose determinant is, to a very good approximation, nn times the log of the fitted innovation variance plus a small end-correction. A higher order fits the residuals better, so the innovation variance falls, so the log-determinant falls. And a candidate whose own fit has already removed more of the dependence leaves residuals with a smaller innovation variance at every order.

So the term moves for two reasons at once — the order chosen and the candidate it was chosen on — and under a per-candidate rule both are varying. That is the worst arrangement for an omitted term: it is correlated with the choice being made and with the quantity being compared.

The window’s rule has the same structure and carries the term, which is why the two rules read differently before the repair and identically after it. Nothing about the sieve is worse. The criterion written for it was.

The check that was written

An omission has no failing test to point at, so what a repair like this needs is an assertion that would have caught it, and the one written here is a pair rather than a single claim.

A shared estimate gives every candidate the same volume, to nine decimal places, on one sample. That is arithmetic and it is asserted as arithmetic rather than averaged over draws: if a future change to the shared path ever makes the term differ between candidates, the shared rule has stopped being shared and the assertion says so before any regret is measured.

And a per-candidate estimate does not, by more than a coefficient is charged. That half is what makes the first half a statement rather than a tautology — a term that never varied would cancel everywhere and could be dropped everywhere.

The refusal beside them is the criterion this collection actually shipped: it asserts that the order’s cost with the term and without it agree, and they differ by a factor of six.

The shape of the defect

This collection has a name for this shape and it is worth using it: the symptom is absence.

Nothing was wrong with any number. The sieve’s criterion computed what it said it computed, the whitening was correct, the residual sum was right, the charge for the order was right. What was missing was a term nobody had written, and no check in this fleet asks whether a formula has all of its terms — a gate can ask whether what is drawn fits, contrasts and stays inside its frame, and none asks whether something that ought to be there is there.

The two other instances in this collection are exactly parallel. A shared ticks() that returned an empty array drew two figures with no gridlines at all, at full marks, because every check asks whether a label fits and none asks whether it exists. And a balancing rule that removes exactly none of an interaction is a guarantee of no protection that reads as a guarantee.

What found this one was not a gate either. It was making a comparison the two rules had never been in — a result too clean to believe, followed by reading both criteria’s arithmetic side by side rather than their descriptions.

Where else the term belongs

Having found one omission it is worth asking where else the same argument applies, and there are two places.

A width that is being varied. Choosing the band’s width means comparing covariances rather than candidates, and there the determinant is not optional for the same reason: two widths are two error models. That field carries it and reports what it does — it moves the answer outwards rather than bounding it, which is not the reason it is usually given.

A break that is being searched for. A two-regime whitening transforms each segment separately, so its determinant is the sum of two, and a rule comparing a one-regime fit against a two-regime one is comparing two transforms. The field that prices a searched break carries the per-segment Jacobian for exactly this reason, and states that the term is what stops the search putting a break at the trim and calling the extra scaled row a gain.

The general rule is easy to state and was apparently easy to forget: whenever two things being compared were computed under different transforms of the sample, the criterion owes the difference in volume. Shared, it is zero. Varying, it is not.

How much of this is about sieves

Almost none of it, and that is worth separating from the repair.

The finding is that a comparison between two procedures is a comparison between two implementations, and the term that differed between them was invisible in both descriptions. A reader could substitute any two whitenings, any two criteria, any two fields for the ones here and the lesson would be the same: the prose said the sieve’s whitening and the band’s whitening and both were accurate, and the criteria differed by a term neither sentence contained.

What makes it findable at all is having a quantity both rules should agree about. The two per-candidate costs are the same object measured through two implementations, and equality was the prediction — so a factor-of-six disagreement was a signal rather than a result. Without a shared prediction the wrong number would simply have been the answer.

That is the same discipline a residual autocovariance’s exact expectation supplies one field along, and the same one two routes to a coverage supplies across the whole collection. A number with only one route to it is a number nothing can contradict.

What the term is not

Two things the determinant does not do, both of which are natural guesses and both of which are measured elsewhere in this collection.

It is not a penalty for the covariance’s dimension. Those are different terms with different sizes: the dimension is charged per lag or per coefficient and the determinant is a property of the matrix. A criterion needs both once the covariance is being chosen, and carrying one is not carrying the other.

And it does not bound the fit. The natural argument for it — that a covariance free to become singular would explain any residual vector — is wrong when the covariance is normalised to a unit diagonal, because that fixes the trace. Both halves of the objective turn over inside the region, and the determinant’s job is to move the answer rather than to keep it finite.

The list is not what separates them. What choosing the tuning parameter separately for every candidate costs, against how many values the rule may choose from, under AR(1) at 0.8. The earlier field's reason for not comparing a window against an order was that their lists are different lengths, and matching them changes very little: at four values the window costs 0.00223 and the order 0.00406; at eight, 0.00401 and 0.00415. Each rule's own list is marked. The two curves are inside each other's standard errors from four values on, and neither has a slope worth the name. At two values the order's cost is exactly zero, because a list of two leaves the five candidates nothing to disagree about — which is the shape of the mechanism and the whole of the list's contribution to it. What the list decides is how often the candidates disagree; what it does not decide is what the disagreement costs.
Fig. 5 The repaired comparison, across every list length either rule can be run at.

What a reader should do with this

Two things, and the second is the one that generalises past this collection.

Where a nuisance is shared, keep it shared. The whole reason an estimated whitening is affordable is that its errors cancel between candidates, and the whole reason its determinant can be dropped is the same cancellation. Those are one argument, not two, and giving it up gives up both at once — which is why the per-candidate rule pays for the displacement and needs a term the shared rule never did.

And when comparing two published procedures, compare their objectives rather than their descriptions. The two rules here are described identically in every respect that matters and their criteria differ by a term worth twelve coefficients. That is not a rare failure: an objective is usually written once, in code, and described many times, in prose, and the description is what a comparison is built from.

What is claimed here, and what is not

This essay takes what an estimated whitening owes a criterion, and what happens when it is not paid. The claims are that whitening changes variables and the Gaussian likelihood’s 12logΩ-\tfrac{1}{2}\log|\Omega| is the Jacobian rather than an adjustment; that a nuisance estimated once gives every candidate the same term, −90.864 on the sample drawn here, so it cancels out of every difference exactly; that a nuisance estimated per candidate gives fifteen different terms spanning 23.875 on the same sample, which is twelve times what a coefficient is charged, and that the spread slopes with the candidate’s dimension; and that the per-candidate order’s regret reads −0.00011 at 0.2 standard errors without the term and 0.00415 at 2.0 with it, against the window’s 0.00401 at 2.3.

What stays out, and is named as a decision: an audit of every criterion in the collection. The argument here identifies two more places the term belongs and both already carry it, but the statement nothing else is missing it would need every whitened criterion in twelve fields read line by line against its own likelihood, and that has not been done. What has been done is stated: the sieve’s rule was missing it, the band’s was not, and the two places named above were checked.

Also out: the same question about a resampling. A block or multiplier resample also transforms the sample, and whether a criterion comparing resampling constructions owes a volume term is a question with a different answer, because a resample is a distribution rather than a transform of one sample. What each construction keeps is measured in its own field and the criterion question there is not asked.

The boundary against the previous essay is that it reports a comparison and this one reports what the comparison needed to be made at all.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutoregressionCholeskyClosed formCovariance matrixDegrees of freedomDependenceDeterminantInformation criterionJacobianLikelihoodModel selectionNuisance parameterRegretWhitening