The volume a whitening moves
Worth reading first: A design is a number · The observations that repeat each other.
A criterion compares fitted models by their residual sums, and a whitened criterion compares them by the residual sums of transformed samples. That is a substitution most of this collection makes without comment, and it is legitimate exactly when every candidate is transformed the same way.
When it is not, the comparison has to carry a term for it. The Gaussian likelihood has that term written into it — , the volume the transform moves — and every criterion in this collection is a Gaussian likelihood with a charge attached.
This essay is about the one place the term was missing, why nothing found it for five rounds, and what its absence was doing to the comparison the previous essay makes.
The term, and what it is for
Whitening multiplies the sample by . A density transformed by a linear map picks up the Jacobian of that map, so the log-likelihood of the original sample and the log-likelihood of the whitened one differ by : not an adjustment, a change of variables.
The parametric version of this is one number. The Prais–Winsten transform for a first-order autoregression scales the first row by and differences the rest, so the whole determinant is — one term, not of them, and it is the entire difference between maximising and iterating.
The general version is the accumulated pivots of a Cholesky factor and there is nothing to say about it except that it exists.
Why it cancels, and why nobody noticed
Estimated once, from the fullest candidate’s residuals, before any selection, the term is the same number for every row of the table. On the sample drawn above it is −90.864 for all fifteen candidates, and it would be the same number for any sample: one estimate, one matrix, one determinant.
A criterion is read in differences. Adding a constant to every candidate’s score changes no ranking, no choice and no regret. So a criterion that omits a shared determinant is exactly as good as one that carries it, and this is not an approximation — it is an identity.
That is why the omission survived. Every whitened criterion in this collection until now has used a shared nuisance, on purpose and for a stated reason: an error every candidate carries cancels out of every difference, which is the argument that makes an estimated nuisance affordable at all. Under that regime the missing term is provably harmless, so nothing failed, nothing looked wrong, and the two rules’ implementations were free to differ in it.
And why it stops cancelling
Estimated from each candidate’s own residuals, the term is fifteen different numbers. On the same sample they run from −110.659 to −86.784, a spread of 23.875.
A criterion charges 2 for a coefficient. The volume term varies by twelve times that across the table, and it varies systematically: a candidate with more predictors leaves residuals with less dependence in them, so its estimated whitening is closer to the identity and moves less volume. The term is not noise added to each row; it is a slope across the table in the same direction as the thing the table is trying to measure.
Leaving it out therefore does not merely add variance to the comparison. It biases the ranking towards candidates whose whitening moved the least — which are the fullest ones, which is the direction a per-candidate rule is already displaced in for an entirely separate reason.
What it was doing
Scored without the term, choosing a sieve order per candidate costs −0.00011 of regret at 0.2 standard errors. Scored with it, on the same draws, it costs 0.00415 at 2.0 — and the window’s rule, which has carried the term since the day it was written, costs 0.00401 at 2.3.
So the whole apparent difference between a per-candidate window and a per-candidate order was the missing term. Two rules that cost the same, one of them scored by a criterion missing a twelve-parameter slope, read as one rule costing something and the other costing nothing.
It is worth naming what that would have looked like if it had shipped. A window chosen per candidate is a real cost and an order chosen per candidate is free is a clean, quotable, false result with a plausible mechanism available for it — a sieve’s whitening is smoother, an order is a coarser dial, pick one. The absence of the term is invisible in every figure and every gate.
Where the shared value sits inside the spread
The shared term is one of the fifteen — it is the fullest candidate’s own estimate, applied to everybody — so it can be located inside the range the per-candidate rule produces, and where it sits is a check on the mechanism.
It is −90.864 against a range of −110.659 to −86.784. That is 4.08 below the least negative reading and 19.80 above the most negative: the shared value sits in the upper fifth of the spread, which is what a whitening estimated from the residuals of the fullest fit has to do.
The 4.08 is worth noticing rather than rounding away. If the slope across the table were an ordering rather than a tendency, the fullest candidate would sit at the very top and the shared value would be the maximum. It is two coefficients’ worth of charge below it, so at least one candidate with fewer predictors leaves residuals whose whitening moves less volume than the fullest fit’s does. The direction of the slope is real; the monotonicity is not, and nothing in the argument needed it.
The repair leaves the two rules indistinguishable, not merely close
“The two rules read identically after it” is measured, and the margin is worth stating because it is much tighter than the sentence implies.
With the term, the order’s rule costs 0.00415 at 2.0 standard errors, so its own standard error is about 0.0021. The window’s rule costs 0.00401 at 2.3, so its standard error is about 0.0017. The two point estimates differ by 0.00014 — three and a half per cent of either, and under a tenth of the smaller standard error.
So the repair does not narrow a gap; it removes one. Two rules that differed by a factor of thirty-eight before the term was added differ by a twentieth of a standard error after it.
The omission was suppressing the noise as well as the number
One more thing falls out of the same two readings, and it is the reason the false result looked solid rather than merely small.
Scored without the term, the comparison read −0.00011 at 0.2 standard errors, so its standard error was about 0.00055. Scored with it, the standard error is 0.0021 — nearly four times larger.
That is not a coincidence: the volume term varies from draw to draw as well as from candidate to candidate, and adding it to every score adds its draw-to-draw variation to the comparison. A criterion missing the term is measuring a quieter quantity, and quietness is what a reader reads as precision.
It also settles what would have happened if the field had simply run more draws. The unrepaired comparison could have resolved an effect of 0.0011 at two standard errors, which is a quarter of the true 0.00415 — so more draws would not have found it. The omission moved the estimand rather than adding noise to it, and no amount of counting reaches a number that is not being computed.
What the term looks like as a function of the order
The volume a sieve’s whitening moves is not an arbitrary number, and the shape of it explains both halves of the finding.
An autoregression of order fitted to a series implies a covariance matrix whose determinant is, to a very good approximation, times the log of the fitted innovation variance plus a small end-correction. A higher order fits the residuals better, so the innovation variance falls, so the log-determinant falls. And a candidate whose own fit has already removed more of the dependence leaves residuals with a smaller innovation variance at every order.
So the term moves for two reasons at once — the order chosen and the candidate it was chosen on — and under a per-candidate rule both are varying. That is the worst arrangement for an omitted term: it is correlated with the choice being made and with the quantity being compared.
The window’s rule has the same structure and carries the term, which is why the two rules read differently before the repair and identically after it. Nothing about the sieve is worse. The criterion written for it was.
The check that was written
An omission has no failing test to point at, so what a repair like this needs is an assertion that would have caught it, and the one written here is a pair rather than a single claim.
A shared estimate gives every candidate the same volume, to nine decimal places, on one sample. That is arithmetic and it is asserted as arithmetic rather than averaged over draws: if a future change to the shared path ever makes the term differ between candidates, the shared rule has stopped being shared and the assertion says so before any regret is measured.
And a per-candidate estimate does not, by more than a coefficient is charged. That half is what makes the first half a statement rather than a tautology — a term that never varied would cancel everywhere and could be dropped everywhere.
The refusal beside them is the criterion this collection actually shipped: it asserts that the order’s cost with the term and without it agree, and they differ by a factor of six.
The shape of the defect
This collection has a name for this shape and it is worth using it: the symptom is absence.
Nothing was wrong with any number. The sieve’s criterion computed what it said it computed, the whitening was correct, the residual sum was right, the charge for the order was right. What was missing was a term nobody had written, and no check in this fleet asks whether a formula has all of its terms — a gate can ask whether what is drawn fits, contrasts and stays inside its frame, and none asks whether something that ought to be there is there.
The two other instances in this collection are exactly parallel.
A shared ticks() that returned an empty array drew two
figures with no gridlines at all, at full marks, because every check asks whether a label fits and
none asks whether it exists. And
a balancing rule that removes exactly none of an interaction
is a guarantee of no protection that reads as a guarantee.
What found this one was not a gate either. It was making a comparison the two rules had never been in — a result too clean to believe, followed by reading both criteria’s arithmetic side by side rather than their descriptions.
Where else the term belongs
Having found one omission it is worth asking where else the same argument applies, and there are two places.
A width that is being varied. Choosing the band’s width means comparing covariances rather than candidates, and there the determinant is not optional for the same reason: two widths are two error models. That field carries it and reports what it does — it moves the answer outwards rather than bounding it, which is not the reason it is usually given.
A break that is being searched for. A two-regime whitening transforms each segment separately, so its determinant is the sum of two, and a rule comparing a one-regime fit against a two-regime one is comparing two transforms. The field that prices a searched break carries the per-segment Jacobian for exactly this reason, and states that the term is what stops the search putting a break at the trim and calling the extra scaled row a gain.
The general rule is easy to state and was apparently easy to forget: whenever two things being compared were computed under different transforms of the sample, the criterion owes the difference in volume. Shared, it is zero. Varying, it is not.
How much of this is about sieves
Almost none of it, and that is worth separating from the repair.
The finding is that a comparison between two procedures is a comparison between two implementations, and the term that differed between them was invisible in both descriptions. A reader could substitute any two whitenings, any two criteria, any two fields for the ones here and the lesson would be the same: the prose said the sieve’s whitening and the band’s whitening and both were accurate, and the criteria differed by a term neither sentence contained.
What makes it findable at all is having a quantity both rules should agree about. The two per-candidate costs are the same object measured through two implementations, and equality was the prediction — so a factor-of-six disagreement was a signal rather than a result. Without a shared prediction the wrong number would simply have been the answer.
That is the same discipline a residual autocovariance’s exact expectation supplies one field along, and the same one two routes to a coverage supplies across the whole collection. A number with only one route to it is a number nothing can contradict.
What the term is not
Two things the determinant does not do, both of which are natural guesses and both of which are measured elsewhere in this collection.
It is not a penalty for the covariance’s dimension. Those are different terms with different sizes: the dimension is charged per lag or per coefficient and the determinant is a property of the matrix. A criterion needs both once the covariance is being chosen, and carrying one is not carrying the other.
And it does not bound the fit. The natural argument for it — that a covariance free to become singular would explain any residual vector — is wrong when the covariance is normalised to a unit diagonal, because that fixes the trace. Both halves of the objective turn over inside the region, and the determinant’s job is to move the answer rather than to keep it finite.
What a reader should do with this
Two things, and the second is the one that generalises past this collection.
Where a nuisance is shared, keep it shared. The whole reason an estimated whitening is affordable is that its errors cancel between candidates, and the whole reason its determinant can be dropped is the same cancellation. Those are one argument, not two, and giving it up gives up both at once — which is why the per-candidate rule pays for the displacement and needs a term the shared rule never did.
And when comparing two published procedures, compare their objectives rather than their descriptions. The two rules here are described identically in every respect that matters and their criteria differ by a term worth twelve coefficients. That is not a rare failure: an objective is usually written once, in code, and described many times, in prose, and the description is what a comparison is built from.
What is claimed here, and what is not
This essay takes what an estimated whitening owes a criterion, and what happens when it is not paid. The claims are that whitening changes variables and the Gaussian likelihood’s is the Jacobian rather than an adjustment; that a nuisance estimated once gives every candidate the same term, −90.864 on the sample drawn here, so it cancels out of every difference exactly; that a nuisance estimated per candidate gives fifteen different terms spanning 23.875 on the same sample, which is twelve times what a coefficient is charged, and that the spread slopes with the candidate’s dimension; and that the per-candidate order’s regret reads −0.00011 at 0.2 standard errors without the term and 0.00415 at 2.0 with it, against the window’s 0.00401 at 2.3.
What stays out, and is named as a decision: an audit of every criterion in the collection. The argument here identifies two more places the term belongs and both already carry it, but the statement nothing else is missing it would need every whitened criterion in twelve fields read line by line against its own likelihood, and that has not been done. What has been done is stated: the sieve’s rule was missing it, the band’s was not, and the two places named above were checked.
Also out: the same question about a resampling. A block or multiplier resample also transforms the sample, and whether a criterion comparing resampling constructions owes a volume term is a question with a different answer, because a resample is a distribution rather than a transform of one sample. What each construction keeps is measured in its own field and the criterion question there is not asked.
The boundary against the previous essay is that it reports a comparison and this one reports what the comparison needed to be made at all.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A dependence with a shape — both name autoregression, covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
- The window that has to be chosen, and the term that was dropped — both name covariance matrix, degrees of freedom, dependence, information criterion, likelihood, model selection, nuisance parameter
- The order the tail is drawn at — both name autoregression, information criterion, model selection, nuisance parameter, regret, whitening
- The window a whitening wants — both name covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
- Where the generality runs out — both name covariance matrix, information criterion, model selection, nuisance parameter, regret, whitening
- A lag the sample has less of — both name closed form, degrees of freedom, dependence, information criterion, model selection
Named objects
A flat tag is an object no other essay names yet.
AutoregressionCholeskyClosed formCovariance matrixDegrees of freedomDependenceDeterminantInformation criterionJacobianLikelihoodModel selectionNuisance parameterRegretWhitening