A penalty is a trace
Worth reading first: Choosing the order · What the model says next.
The essay that priced a specification search against a rolling hold-out ends by naming a repair and declining to build it. A criterion beats a hold-out at every split there is on independent rows, loses to it once the rows repeat each other, and loses because its penalty is counting rows. The honest fix, that essay says, is a penalty computed from an effective sample size rather than from n.
Before anything can be computed from an effective sample size it is worth being precise about what the penalty is for, because the usual statement of it is a coincidence dressed as a definition.
What 2q is the answer to
Akaike’s criterion is n·log(RSS/n) + 2q. The 2q is universally described as a charge of two units per fitted parameter, and that description is what makes the repair look impossible: if the penalty counts parameters, and the defect is about rows, there is nowhere for a sample size to go.
The penalty does not count parameters. It estimates the optimism of the fit — the amount by which an in-sample residual sum understates what the same fit scores on data it has not seen — and that quantity has a closed form which mentions the parameter count nowhere:
where H is the projection onto the candidate’s columns and Σ is the covariance of the errors. Under independent errors Σ = σ²I, so tr(HΣ) = σ² tr(H) = qσ², and the optimism is 2qσ²/n. The parameter count is the value a trace happens to take, not the thing being charged for.
So the repair writes itself. With Σ = σ²Ω the same derivation gives a penalty of 2·tr(HΩ), and that is not an approximation to anything — it is the same theorem with the same premises, minus the one that was dropped.
Measuring it takes two things that must not be allowed to share code. The trace comes from the design matrix and the error correlation, and never sees an outcome. The optimism comes from fitting on a sample and then scoring that same fit on the same design lived again — the identical rows, with an independent draw of the error process. Averaging over the fresh errors leaves σ² plus the mean squared distance between the fit and the truth, so that side needs no simulation at all once the fit exists.
The two agree at every candidate in the table, with the worst departure at 5.9% and the smallest at 0.6%. The departures are not noise and are worth naming: they are largest for the candidates that leave out the most persistent predictors, whose effective error is not the error the trace was computed for. That is a real limitation of the correction and it is measured rather than assumed away.
The count is the average over orientations
There is a sense in which 2q is exactly right, and finding it says precisely what goes wrong for real candidates.
Ω is a correlation matrix, so its trace is n. If H were a projection onto a randomly oriented q-dimensional subspace, the expected value of tr(HΩ) would be — the count, exactly, at every persistence and every n.
So the parameter count is not merely the independent-errors special case. It is the average of the trace over all orientations a candidate might have had.
Which is why it fails in one direction rather than at random. A real candidate’s columns are not randomly oriented: they are predictors somebody recorded, they are persistent for the same reasons the errors are, and persistence means alignment with Ω’s large eigenvalues. Every candidate in a table of this kind sits on the same side of the average, and the measured corrections run from 2.550 to 4.743 per parameter against a count of 1.
How far the trace can be from the count
The spectrum bounds it, and the bound is wide.
For a projection of rank q, tr(HΩ) lies between the sum of Ω’s q smallest eigenvalues and the sum of its q largest. A first-order autoregression’s correlation matrix has eigenvalues running roughly from to , which at ρ = 0.7 is 0.18 to 5.67.
So a q-parameter candidate’s true penalty is somewhere between 0.18q and 5.67q — a factor of thirty-two — against a criterion that charges q regardless.
The measured range of 2.55 to 4.74 sits comfortably inside that and comfortably above one, which is the useful statement. Under positive dependence the count is not a bound and it is not conservative; it is a number the trace passes on its way up, and how far past it goes is a property of the candidate’s columns rather than of the errors alone.
That is also why nothing about the defect can be repaired by charging more per parameter. A heavier constant would move every candidate by the same factor, and what the trace says is that the candidates need moving by different amounts — 2.55 for one and 4.74 for another, in the same table, at the same persistence.
Where the count and the trace come apart
At ρ = 0 the trace is q, exactly, at every candidate and on every draw — which is the check that this generalises the rule rather than replacing it. The interesting question is what it becomes when it stops being the count.
Two things have to be true of the design for the answer to be interesting, and only one of them is usually arranged. The errors must repeat each other, which is what Ω being non-diagonal means. And the predictors must differ in how much they repeat each other. A table of candidates in which every regressor has the same autocorrelation makes tr(HΩ)/q the same number for every candidate of the same size, and the whole distinction between a per-candidate correction and a scalar is invisible.
So the design here carries four predictors at four persistences — 0.9, 0.6, 0.3 and 0 — which is not an exotic world. A table of candidates in which every regressor moves at the same speed is the exotic one.
At ρ = 0.7 the corrections run from 2.550 to 4.743 per parameter. Both ends are candidates with the same number of coefficients: the one that fits only the near-unit-root predictor needs 4.743 and the one that fits only the white one needs 3.258, a gap of 1.485 in a quantity that a single scalar would have to answer for both.
The pattern is exactly what the mechanism predicts. A candidate’s correction rises with how persistent its own columns are, because the optimism is a trace of a projection against the errors’ covariance and a slow-moving column is aligned with a slow-moving covariance. A fit spends its coefficients tracking the errors’ low-frequency movement, and it can only do that with columns that move at the same speed.
Why it takes both premises
It is worth being clear about which two things have to be true, because the field this one continues found the same pair and the pairing is not obvious.
Persistent errors alone do almost nothing. The optimism is a trace of a projection against the error covariance, and if the design is a fresh draw of independent rows then the projection is not aligned with the errors’ slow movement — the trace comes out near qσ² whatever Ω is. There is nothing for the fit to spend its coefficients on.
Persistent predictors alone do nothing either, for the mirror reason: the columns are slow and the errors are not, so there is nothing slow in the errors for the columns to track.
It takes both, and the size of the effect is a product rather than a sum. That is exactly why the correction has to be per-candidate: the alignment between a candidate’s columns and the errors’ shape is a property of those columns, and two candidates fitting the same number of coefficients can be aligned to entirely different degrees. Averaged over the fifteen candidates here the correction is 3.318 per parameter; the best-aligned candidate needs 4.743 and the worst-aligned 2.550, and the ordering follows the average persistence of the columns each one fits with no exceptions.
The classical effective sample size for a first-order autoregression is n(1 − ρ)/(1 + ρ), which at n = 120 and ρ = 0.7 is 21.18. Its finite-sample form — the one that actually governs the variance of a mean, Σ(1 − |k|/n)ρ^|k| — gives 21.67. Neither of those numbers is in the table above, and the reason they are not is the subject of the next essay.
Two ways to say the same thing, and only one of them is right
There is a second reason the correction has to be per-candidate, and it is arithmetic rather than mechanism.
A criterion compares candidates. Only differences between scores matter, so what a repair changes is the relative charge for going from one candidate to another. Under the parameter count that charge is 2 per extra coefficient, the same everywhere. Under the trace it is 2 times the difference in tr(HΩ), which at ρ = 0.7 is about 2.5 times as large on average and is not constant across the table: adding the persistent predictor to a candidate costs more than adding the white one, which is a statement no scalar multiplier of 2q can make.
A scalar can change the exchange rate. It cannot change which coefficient is expensive.
That is the sentence the whole first half of this field turns on, and the next essay is what happens when it is ignored.
What the trace is not
Two things this correction does not do, both worth stating because both are easy to assume.
It does not repair the fit. Least squares stops being the efficient estimator the moment the errors are correlated, and the trace penalty leaves the fit alone. Every candidate is still estimated by a rule that is throwing information away; the criterion is simply charging more accurately for what that rule spent. Whether that is enough to improve the selection is a measurement rather than a corollary, and it turns out not to be.
It does not know what the fit will be used for. The optimism theorem is a statement about a fresh replicate of the same sample: the same design rows, a new draw of the errors. A forecaster wants something else — the rows that come next, whose errors are correlated with the sample’s, so that a fit on a persistent design partly tracks the error it is about to be scored on. The theorem has no term for that, because its fresh row is independent of everything that produced the fit. Both targets are measured in this field and they order the candidates differently.
What it is worth, which is less than nothing
The trace is the right correction for the quantity it corrects, checked from both sides. Selecting with it is worse than selecting with the parameter count.
At ρ = 0.85 the rule that counts rows gives up 0.08589 against the best model available; the rule that computes the trace gives up 0.10433. That is 21.5% more regret for a repair that is exactly right about what it repairs, and the sign is not a rounding matter — it holds at every persistence measured, and the gap widens as the persistence rises.
Two explanations offer themselves and both can be ruled out rather than argued about.
It is not the logarithm. The n·log(RSS/n) form is a linearisation of the optimism argument, which is stated in the scale of a residual mean square: an estimate of the risk is RSS/n plus 2·tr(HΩ)σ̂²/n, and taking logarithms of that and dropping a term is where the additive penalty comes from. So the same substitution can be made on Mallows’ scale, where there is no linearisation to blame. It behaves identically: 0.08619 for the parameter count and 0.10711 for the trace, against 0.08589 and 0.10433 in the logarithmic form. The two scales agree to within a thousandth about both rules.
It is not a mistake about the optimism. That was checked first, from both sides, and the trace is the optimism to within six per cent at every candidate.
What is left is that a correct estimate of each candidate’s risk is not the same thing as a good selection rule. The trace charges the persistent columns two and a half times what the count charges them, which is correct on average and is applied to a residual sum that is itself much noisier under dependence — and the small residual misspecification, the six per cent, falls hardest on exactly the candidates whose charge moved most. The selection that comes out prefers models that are too small: an average of 2.73 coefficients at ρ = 0.85, against 3.18 under the parameter count and against a best available model that grows rather than shrinks.
The repair the trace was pointing at
None of this makes the diagnosis wrong. The row count really is the defect, and the correction really is a trace. What the measurement says is that a penalty is the second place the row count enters and not the first, and repairing the second alone leaves an object whose two halves are answering different questions.
Whiten the sample instead — scale the first row by √(1 − ρ²) and difference every later one at ρ — and both halves are repaired at once. In the transformed sample the rows really are independent, the trace of the hat matrix really is q, and the ordinary penalty is correct again. That rule recovers 86.9% of what counting rows gives up at ρ = 0.85, where the trace penalty recovers −21.5% of it. It is the essay after next.
The trace does not become useless. It is what says how large the correction should be, and it is what proves that the correction cannot be a scalar. What it is not is a rule to select with.
And there is one more thing it is, which is a way of reading the failure it fails at. A rule that charges each candidate 2·tr(HΩ) is charging correctly and selecting badly; a rule that charges 2q is charging incorrectly and selecting better. That combination only makes sense if the criterion is being asked to do two things at once — to estimate each candidate’s risk, and to order the candidates — and the two are not the same job. An unbiased estimate of every candidate’s risk is exactly what a selection rule wants when the estimates are precise, and is not what it wants when the estimates are noisy and their noise is bigger for the candidates whose charge is largest. Under dependence the residual sum is much noisier than the independent-row derivation supposes, so the second condition is the one that holds.
That is not an argument for leaving the penalty wrong. It is an argument that a penalty was never going to be enough on its own, which is what the measurement says and what the whitened rule confirms.
What is claimed here, and what is not
This essay takes what an information criterion’s penalty is actually charging for, and the claims are that 2q is the value of a trace rather than a count of parameters, that the trace is the optimism to within six per cent on a persistent design, that the correction it implies differs by 1.485 between two candidates of the same size, and that selecting with it is 21.5% worse than not correcting at all.
What stays out and is named as a decision: whether a scalar effective sample size can stand in for the trace, which is the next essay; what the repair that works actually is, which is the one after; and anything about reference distributions or rejection rates, since nothing here is a test.
The boundary against the field this one continues is that it measures the failure and this one measures the thing that fails. Its crossing at ρ = 0.81 is what makes the question worth asking; nothing here changes that number.
The checks, and the refusals that make them mean something
Three claims are gated in this field’s library. The trace is required to equal the parameter count to machine precision when the errors are independent, which is the identity that makes this a generalisation. The measured optimism is required to reproduce 2·tr(HΩ)/n at every candidate, on the replicate target and not on any other. And the corrections at a fixed number of coefficients are required to spread, since a table where they did not would make a scalar correct.
The refusal for this essay is a scalar effective sample size applied to a table of candidates. It is the mean’s variance inflation, the table’s candidates are not means, and it is above every one of them — 5.537 against corrections that run from 2.550 to 4.743. A number that is outside the range of what it stands in for is not an approximation to it.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A line that beats two curves — both name closed form, degrees of freedom, dependence, information criterion, least squares, model selection, optimism, overfitting
- The displacement is a parameter count — both name closed form, degrees of freedom, information criterion, least squares, mean squared error, model selection, out of sample, overfitting
- What a better charge buys — both name degrees of freedom, dependence, information criterion, mean squared error, model selection, optimism, out of sample, overfitting
- A width that moves and an error that does not — both name degrees of freedom, information criterion, mean squared error, model selection, optimism, out of sample, overfitting
- The width a band is measured in — both name closed form, degrees of freedom, dependence, information criterion, model selection, optimism, overfitting
- The window that has to be chosen, and the term that was dropped — both name autocorrelation, degrees of freedom, dependence, information criterion, least squares, model selection, persistence
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationClosed formDegrees of freedomDependenceEffective sample sizeHat matrixInformation criterionLeast squaresMean squared errorModel selectionOptimismOut of sampleOverfittingPersistenceRegression