Counting what is independent

The repair that was exact and made it worse

A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.

Worth reading first: What the model says next · Choosing the order.

The correction is right and the rule is worse. That sentence is the whole of this essay, and it is worth dwelling on how uncomfortable it is before it gets explained.

An information criterion’s penalty estimates the optimism of an in-sample fit. On rows that repeat each other, 2q is not the optimism and 2·tr(HΩ) is — checked from both sides, at every candidate, to within six per cent. Substitute the correct quantity for the incorrect one and the selection gets worse: at a persistence of 0.85 the uncorrected rule gives up 0.08589 against the best model available and the corrected one gives up 0.10433, which is 21.5% more.

There are three things that could be happening. Two of them can be ruled out by measurement rather than by argument, which is the only reason the third is worth believing.

It is not the logarithm

The additive penalty is a linearisation. The optimism theorem is stated in the scale of a mean square — an estimate of a candidate’s risk is RSS/n plus 2·tr(HΩ)σ̂²/n, which is Mallows’ form — and the familiar n·log(RSS/n) + 2q comes from taking logarithms and using σ̂² ≈ RSS/n. That approximation is harmless when the penalty is 2q and a good deal less obviously harmless when it is two and a half times larger, so it is the first suspect.

It is innocent. The same substitution made on Mallows’ scale, where nothing is linearised and σ̂² is estimated properly from the fullest candidate with its own trace correction, behaves the same way: 0.08619 for the parameter count against 0.10711 for the trace, where the logarithmic forms give 0.08589 and 0.10433. The two scales agree to within three thousandths about both rules, and both scales say the correction hurts.

The row count entered twice, and a penalty is one placeRegret against the best available model as the errors are made persistent. Counting rows more than quadruples; both penalty repairs — the trace, and the scalar effective sample size — are worse than it at every persistence measured; and whitening the sample and keeping the ordinary penalty *falls*, recovering 86.9% of what counting rows gives up at ρ = 0.85. Doing it at an estimated ρ recovers 80.6%, so having to estimate the dependence from the rows being selected on costs 7.3% of what knowing it is worth. Mallows' forms are drawn beside the logarithmic ones and behave the same, which is what rules the linearisation out.00.0500.10000.2000.4000.6000.800how much the errors repeatregret against the best model availablepenalty repairednothing repairedfit whitened250 draws at each of 5 persistences, n = 120whitening recovers 86.9%
Fig. 1 Seven rules as the errors are made persistent. The dashed pair are Mallows’ forms of the same two penalties, drawn beside the logarithmic ones so that the linearisation can be ruled out rather than assumed innocent. The number of draws each point averages is on a slider.

Three numbers that describe the same change

The correction can be summarised in the three quantities it moves, and putting them side by side is what makes the discomfort precise rather than merely stated.

The price of a coefficient roughly triples. The per-parameter correction the trace requires runs from about 2.55 to 4.74 across the table, against a count of 1.

The model selected roughly halves. At a persistence of 0.85 the corrected rule chooses about two coefficients on average, against a best-available model that has grown to four.

And the regret rises by a fifth, from 0.08589 to 0.10433.

Read together, those say something about the shape of the problem rather than about the correction. A threefold change in the penalty and a halving of the model selected move the delivered regret by 21.5% — so the criterion’s surface is flat, and being far from the best model costs much less than being far from the best penalty suggests.

That flatness is what makes the paradox possible. If the regret rose steeply with the wrong model choice, a correction that halved the model would be catastrophic rather than mildly bad, and there would be no room for a rule to be right about the optimism and wrong about the selection by a fifth.

It is not a mistake about the optimism

The second suspect is the trace itself. It is a closed form derived under a Gaussian design and a stated error covariance, and there are several places a factor could go missing.

It does not. Measured against the optimism directly — a fit held, the design lived again with a fresh error path, the difference between what the fit scores in sample and what it scores on the replicate — the trace reproduces it at every candidate in the table, with ratios running from 1.006 to 1.059. The departures are not noise and are worth naming, because they are the one genuine imperfection in the correction: they are largest for the candidates that omit the most persistent predictors, whose effective error contains those predictors and is therefore more persistent than the Ω the trace was computed from.

That is a real limitation and it is not large enough to be the explanation. A six per cent error in a penalty does not turn a repair into a regression.

The row count entered twice

What is left is that a criterion has two halves and the row count is in both of them.

The penalty is the second half. The first is the fit. Ordinary least squares is the efficient estimator when the errors are independent, and it stops being the efficient estimator the moment they are not — the efficient one is generalised least squares, which weights by the inverse of the error covariance. A criterion built on ordinary least squares is therefore choosing between candidates that were all estimated by a rule which is throwing information away, and no penalty can put that back, because a penalty is a number added after the fit.

A penalty can change which candidate is named. It cannot make the named candidate any better.

And under strong dependence the difference between a least-squares fit and an efficient one is larger than the difference between any two candidates in the table. That is the sense in which correcting only the penalty is correcting the smaller of the two things.

Whitening, and the penalty that needs no repair

The repair is one line and it is not new. Scale the first row by √(1 − ρ²) and replace every later row by itself minus ρ times its predecessor — the Prais–Winsten transform — and the transformed errors are independent with the same variance. Least squares on the transformed sample is generalised least squares on the original.

What matters here is what it does to the criterion. In the transformed sample the rows really are independent, so the hat matrix’s trace really is q, so the optimism really is 2qσ²/n, so the ordinary penalty is correct again. There is nothing left to repair. The whole of the fix is in the fit, and the penalty that looked like the problem turns out to be right as long as it is applied to a fit that deserves it.

Repairing the penalty, and repairing the fit. What each rule gives up against the best model available from any of them, at ρ = 0.85. The first three fit by least squares and differ only in the penalty; the next two are the same substitutions on Mallows' scale, where there is no linearisation to blame; the last two whiten the sample and keep the ordinary penalty 2q. A penalty computed from the trace is exact about the optimism and selects worse than the parameter count it corrects. Whitening is better than every penalty, and doing it at an estimated ρ rather than a known one costs 0.00544.
Fig. 2 The seven rules at a persistence of 0.85. The two that whiten the sample are an order below the five that do not, and the two penalty repairs are behind the rule they were built to improve.

At ρ = 0.85 that rule gives up 0.01123 against the best available model, where counting rows gives up 0.08589. It recovers 86.9% of the gap. The penalty repairs recover −21.5% and −19.6% of it.

Selection quality and delivered quality

Two regrets are worth separating, because a rule can be good at one and poor at the other and the penalty repairs are.

Selection quality is how far a rule’s chosen candidate is from the best candidate under that rule’s own estimator. It asks whether the rule picked well, with the question of how well anything was fitted divided out. Delivered quality is how far the model that actually gets handed over is from the best model available from any of the rules. It is what an analyst experiences, and it contains both.

Counting rows has a selection regret of 0.0363 at ρ = 0.85 and a delivered regret of 0.0859, so more than half of what it gives up is the estimator underneath it rather than the choice it made. The whitened rule has 0.0071 and 0.0112 — it selects better and delivers better, and the second gap is the smaller of the two.

The trace penalty’s numbers are the awkward ones: 0.0547 and 0.1043. It selects worse than the rule it corrects, on the same estimator, which is the part that no story about efficiency explains and which the section before this one is about.

Rows that repeat each other are not less information

The most surprising number in the sweep is the direction the whitened rule moves in.

Counting rows gets steadily worse as the persistence rises: 0.01936 at ρ = 0, 0.04449 at 0.6, 0.08589 at 0.85. That is the expected shape and it is what the rows are worth less would predict.

The whitened rule goes the other way. 0.01936 at ρ = 0 — identical, as it must be, since whitening at ρ = 0 is the identity — then 0.02807 at 0.6, then 0.02134 at 0.75, then 0.01123 at 0.85. Its regret is lower on the most persistent design than on the independent one.

This needs stating carefully, because it is easy to over-read. The problem does get harder in absolute terms: the best model available is worth 1.0274 at ρ = 0 and 1.0887 at ρ = 0.85, so whitening does not make a persistent sample as good as an independent one. Differencing a persistent design at ρ takes the signal out along with the noise — a predictor moving at 0.9 differenced at 0.85 has a fifth of its variance left — so there is genuinely less information about the coefficients.

What falls is the gap between what the rule delivers and the best it could deliver. Under dependence the candidates separate more sharply once they are estimated efficiently, so the selection problem becomes easier even as the estimation problem becomes harder, and the whitened rule collects almost all of what is there.

So rows that repeat each other carry less information is not quite the right sentence. They carry information a least-squares fit does not know how to read. Some of it is genuinely gone and the rest is recoverable, and the split between the two is what the sweep measures.

There is a companion statement about the design that is worth having in the same breath. The best available model grows as the persistence rises — from four coefficients at ρ = 0 to a table whose efficient fits separate more sharply at 0.85 — and the whitened rules follow it, selecting an average of 3.27 coefficients at ρ = 0 and 3.97 at 0.85. The penalty repairs move the other way, to 2.73 and 2.02. A rule that buys fewer coefficients exactly as coefficients become better value is failing in a direction, and the direction is the one a heavier charge always produces.

What the obstacle actually costs

There is one thing wrong with everything above: Ω is not known. The whitened rule at the true ρ is infeasible, and the essay that named this repair named the reason — estimating a dependence from the same rows the selection is running on is a second selection problem with the same shape as the first, and nothing said how large it is.

It is 7.3% of what the repair is worth.

The feasible rule estimates ρ from each candidate’s own residuals, before any selection, and whitens with that. At ρ = 0.85 it gives up 0.01667 against the known-ρ rule’s 0.01123 and the uncorrected 0.08589 — so it recovers 80.6% of the gap where knowing ρ recovers 86.9%, and the difference between them is a small fraction of either.

What each repair buys. The average number of coefficients each rule selects as the errors are made more persistent. Counting rows drifts slightly; scaling the penalty by n/n_eff collapses to the smallest candidate in the table; whitening the fit selects larger models as the persistence rises, which is the direction the best available model moves in this design. A rule that selects the right size for the wrong reason is not distinguished from one that selects it for the right reason by this figure — the regret is what does that.
Fig. 3 What each rule buys. The whitened rules select larger models as the persistence rises, which is the direction the best available model moves; the penalty repairs go the other way.

The reason it is cheap is worth having, because it is the same reason a similar worry turns out to be misplaced in a field about two-arm trials. ρ̂ is estimated from a hundred and twenty residuals; the thing it feeds is a comparison between fifteen candidates whose scores differ by a few hundredths. An error in ρ̂ moves every candidate’s score in nearly the same direction, and a criterion only reads differences.

Two biases in ρ̂, pulling opposite ways

The estimate is biased, and the shape of the bias is the interesting part.

A fitted model’s residuals are less persistent than its errors, because the projection removes the part of the errors lying in a column space that is itself slow-moving. That pulls ρ̂ down, and it pulls it down hardest for the candidates that fit the most persistent columns.

Against it, a candidate that omits a persistent predictor has that predictor in its residuals, and it is persistent, so its own effective error looks more dependent than the true error is. That pulls ρ̂ up, and it pulls hardest for the candidates that omit most.

The two do not cancel and they do run in opposite directions across the table. Measured at ρ = 0.7, the smallest candidates give ρ̂ = 0.664 and the largest gives 0.630 — the bias grows from −0.036 to −0.070 as the candidate grows. The estimate looks best for the candidates that leave the most out, and it looks best there for the wrong reason: two errors partly cancelling rather than one estimate being good.

That is what makes the refusal at the end of this field’s library the refusal it is. Estimating ρ from the winner’s residuals after selecting looks like a tidy economy and reads a quantity the selection produced. The winner is disproportionately the candidate whose residuals happened to look least dependent, so ρ̂ from the winner comes out below the average across the table — and a rule that re-whitens with it whitens by too little, on the one draw where it has already been shown that whitening is what matters.

Why the second selection problem is small here and not everywhere

It would be wrong to take 7.3% as a general figure, and the reason it is small is specific enough to say.

A criterion reads differences between candidate scores. An error in ρ̂ that is shared across the table shifts every score by nearly the same amount and cancels; only the part of the error that differs between candidates survives into the comparison. And ρ̂ is estimated per candidate from a hundred and twenty residuals, so its sampling error is a few hundredths while the systematic part — the two biases below — is of the same order. The comparison is protected by cancellation rather than by precision.

Where that protection is absent, the same construction would be far more expensive. A rule that used one ρ̂ for a level rather than for a comparison — a standard error, an interval, a stopping decision — would carry the whole error rather than its differences, which is exactly the situation the essays about intervals under dependence measure. The cheapness here is a property of what a criterion is, not of the estimate.

The general form of the finding

Three things were true at once and only one of them was being repaired.

The optimism was wrong, and the trace fixes it exactly. The fit was inefficient, and nothing in a criterion fixes it. And the target the criterion is derived for — the same design lived again — is not the target a forecaster wants, because the rows that come next have errors correlated with the sample’s and the theorem’s fresh row does not.

A repair aimed at the first alone is exactly right about the first alone. The measurement says that is not enough, and that a repair aimed at the fit gets the first for free.

What is claimed here, and what is not

This essay takes why a correct penalty is not a correct rule, and the claims are that the linearisation is not the explanation because Mallows’ scale behaves identically, that the optimism is not the explanation because the trace reproduces it to within six per cent, that whitening the sample recovers 86.9% of what counting rows gives up, that estimating the dependence costs 7.3% of that, and that the whitened rule’s regret falls as the persistence rises while the problem itself gets harder.

What stays out and is named as a decision: any dependence structure other than a first-order autoregression, since the transform is written for that one and a general Ω would have to be estimated rather than parameterised; and the target question — the same design lived again against the rows that come next — which is measured and reported and not resolved, because which of the two an analyst wants depends on what the model is for.

The boundary against the essay that derived the trace is that it is about what the penalty should be and this one is about whether the penalty was the problem.

The checks, and the refusals that make them mean something

Three claims are gated in this field’s library. The trace penalty is required to be worse than the parameter count on both scales, which is the finding and would be a bug in either direction. The whitened rule is required to beat both. And the cost of estimating the dependence is required to be a small share of what knowing it is worth — a claim that would fail if the second selection problem were the obstacle it was feared to be.

The refusal for this essay is the dependence estimated from the selected candidate’s residuals. The winner is the candidate whose residuals looked least dependent, so ρ̂ from the winner is below the average across the table, and a rule that re-whitens with it under-whitens exactly where whitening is the thing that mattered.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationDependenceEfficiencyEstimation errorGeneralised least squaresInformation criterionLeast squaresMean squared errorModel selectionOptimismOut of samplePersistenceRegressionSelection effectSpecification search