One number for a table of candidates
Worth reading first: The observations that repeat each other · Choosing the order.
The effective sample size is one of the few pieces of vocabulary that crosses freely between fields. A hundred and twenty observations from a series that repeats itself are worth fewer than a hundred and twenty independent ones; the standard expression says how many fewer; and it is quoted in survey statistics, in Monte Carlo, in meteorology and in clinical trials as though it were a property of the sample.
It is a property of the sample and of one estimator, and the estimator is the mean. That is not a quibble about generality. It decides whether an entire class of repairs can work.
Where the number comes from
For a first-order autoregression at ρ, the variance of a sample mean of n observations is not σ²/n but σ²/n times an inflation factor, and the factor is the sum of the autocorrelations weighted by how many pairs are that far apart:
Take n to infinity and it becomes (1 + ρ)/(1 − ρ), so the effective sample size is n(1 − ρ)/(1 + ρ) — the form everybody quotes. At n = 120 and ρ = 0.7 the limit gives 21.18 and the finite sum gives 21.67, a difference of about two per cent that is worth carrying because it is free.
Every part of that derivation is about a mean. The weights are 1/n each; the inflation is a weighted sum of autocorrelations with those weights; there is no design matrix anywhere in it.
There is one more thing to notice about the derivation, and it is the part that travels. The inflation is a weighted sum of autocorrelations, with the weights coming from the estimator. For a mean the weights are the pair counts, which is what (1 − |k|/n) is. Change the estimator and the weights change, and there is no reason the sum should come out the same. That is the whole of the argument below, written before the counter-example.
The two per cent is a persistence, not a constant
The finite sum has a closed form, and having it turns “a difference of about two per cent” into a statement about where the quoted limit can be trusted.
Summing the weighted geometric series gives
so the limit overstates the inflation by very nearly . At n = 120 and ρ = 0.7 that is 1.4 / 10.8 = 0.1296, which takes 5.66667 to 5.53704 — the number measured, to five figures.
What the expression says is that the gap closes like 1/n and opens like , and the second is much the stronger. Holding n at 120:
- ρ = 0.7: the limit is 5.667 against a truth of 5.537, 2.3% high
- ρ = 0.9: 19.00 against 17.50, 8.6% high
- ρ = 0.95: 39.00 against 32.67, 19.4% high
So the correction that is free to make and worth two per cent at moderate persistence is worth a fifth of the answer at the persistence where somebody is most likely to reach for an effective sample size in the first place. A series persistent enough that the phrase feels necessary is a series where the quoted form is worst.
The mechanism is the one the derivation makes obvious in hindsight. The limit assumes the autocorrelation has decayed before the sample runs out; at ρ = 0.95 the correlation at the hundredth lag is still 0.006 and the weights have been cutting the sum down the whole way. The limit is not the large-sample form of the finite sum so much as the form for a sample long compared with , which at ρ = 0.95 is twenty observations of memory against a hundred and twenty of data.
The identity, and what it is an identity about
Now put the same quantity beside the trace an information criterion’s penalty should be charging.
For a fit with a single constant column, the projection is 1/n in every entry, so
which is the finite inflation factor exactly. So n/n_eff and tr(HΩ) are the same number for the intercept-only fit — not close to it, the same, to machine precision. That identity is what makes an effective sample size look like a general answer: it is a correct special case of the right object, which is the most persuasive kind of wrong generalisation there is.
Two things follow immediately and they point in opposite directions.
The finite form is exact for the mean, and the limit form is not: 5.6667 against a truth of 5.5370 at n = 120 and ρ = 0.7. Quoting the limit as though it were the answer is a two per cent error on the one candidate the scalar is right about, and it is free to avoid.
And the average correction the table’s candidates need is 3.318 per parameter, against the scalar’s 5.537. The scalar is not merely wrong for them; it is above every one of them. The largest correction any candidate in the table needs is 4.743 and the smallest is 2.550.
Why the scalar is above all of them
The mechanism is worth a paragraph, because the mean is the worst case is a useful thing to know and is not obvious.
A constant column is maximally aligned with the low-frequency part of a persistent error process — it is the zero-frequency direction. Any other column has variation the errors do not share, so its projection picks up less of the errors’ slow movement, and its contribution to the trace is smaller. A candidate with an intercept and one predictor has a trace of the intercept’s 5.537 plus something under it, so the per-parameter average is dragged down. Adding more predictors drags it down further.
So the ordering is structural: the mean’s inflation is an upper bound on the per-parameter inflation of any candidate containing it, and a rule that applies the mean’s inflation to every candidate over-charges every one of them. That is exactly what a penalty repair built on n_eff does, and it is why the selection it produces collapses onto the smallest candidate in the table — an average of 2.02 coefficients at ρ = 0.85, against a best-available model that has grown to four.
One consequence is worth stating in the direction nobody states it. If a table contained only means — a set of candidates each of which is an intercept and nothing else — the scalar would be exactly right for every one of them, and a repair built on it would be exactly the trace repair. The scalar is not an approximation that degrades as the model grows. It is an identity that holds at one point of the table and has nothing to say anywhere else, and the reason it looks like an approximation is that nobody checks it away from the point.
How far above them it is, in the units of their own spread
“Above every one of them” is the finding, and the margin is worth putting on the scale of the variation the scalar is failing to represent.
The scalar charges 5.537. The fifteen candidates need between 2.550 and 4.743, averaging 3.318, and at a fixed number of coefficients they spread by 1.485.
Three readings of the same four numbers:
- The scalar sits 2.219 above the average candidate, which is 1.5 times the spread at fixed q. The systematic error of using one number is half again as large as the whole variation the number cannot express.
- The candidates’ full range is 4.743 − 2.550 = 2.193. The scalar’s distance above their mean is larger than their entire range.
- Even the most inflated candidate — the one with the near-unit-root predictor, which is as close to a mean as anything else in the table gets — is over-charged by 0.794, seventeen per cent of the correction it actually needs.
That last line is what rules out the obvious salvage. If the scalar were merely a poor average one could imagine rescaling it — dividing by 5.537/3.318 — and getting something usable on the table as a whole. It would then be right on average and still wrong candidate by candidate by up to 1.485 in a quantity averaging 3.318, which is a relative error of nearly a half on the coefficient price. And it would need the trace computed to find the rescaling, at which point the trace is available and the scalar is not needed.
The substitution that does nothing at all
There is a blunter problem, and it comes first in practice.
The natural way to use an effective sample size is to substitute it for n. That is what the phrase means and it is what software does when it offers the option. Applied to Akaike’s criterion,
it changes nothing whatsoever, because the penalty is 2q and 2q contains no sample size. Not approximately nothing: the selection is identical on every draw, which is checkable as an equality rather than as an agreement. Two hundred draws, two hundred identical selections, with n_eff = 21.7 standing in for n = 120.
The substitution does move the finite-sample form, whose penalty is 2q + 2q(q + 1)/(n − q − 1), and it moves Schwarz’s, whose penalty is q·log n. Both of those are answering different questions from the one the defect is in. The one criterion whose penalty is exactly the optimism is the one an effective sample size cannot touch.
To make the scalar do anything to Akaike at all it has to be used as a scale — a penalty of 2q·n/n_eff rather than 2q — which is a different construction with a different justification, and it is the construction measured above.
The two objects that share the name
There is a second construction called an effective sample size, it appears later in this same round, and confusing the two is easy enough to be worth heading off.
The one above is about dependence between observations: n rows that repeat each other are worth n/(1 + 2Σρ_k) independent ones for estimating a mean. The other is about dependence between draws: B steps of a Markov chain are worth B/(1 + 2Σγ_k) independent draws for estimating whatever the chain is being used to estimate. The arithmetic is the same expression and the objects are not related — one is a fact about a sample, the other a fact about a sampler — and the second is entirely correct, since the thing being estimated by a chain really is a mean over its draws.
That is the pattern worth carrying: the inflation factor is right whenever the estimator is an average with equal weights, and the question is always whether it is. A sample mean is. A Monte Carlo average is. A least-squares coefficient on a design matrix is not, and neither is a criterion’s penalty, and both of those are cases where the phrase gets used anyway.
A scalar changes an exchange rate
The general statement is short, and it is what separates this repair from the one that works.
A criterion compares candidates, so only differences of scores matter, so what a penalty determines is the price of an extra coefficient. A scalar multiplier on 2q makes every coefficient dearer by the same factor: the exchange rate moves and the ordering of what to buy first does not.
The trace does something else. At ρ = 0.7 it prices the near-unit-root predictor at 4.743 per parameter and the white one at 3.258, so it says not only that coefficients are dear but which ones. That is a statement about the table, and no scalar can make it.
The measurement bears it out and does so in the most awkward way available: the scalar and the trace are the two worst rules in the table, and they are worst for opposite reasons. The scalar over-charges uniformly, so it buys almost nothing — 0.10269 of regret at ρ = 0.85, on an average of two coefficients. The trace charges correctly on average and gets the relative prices right, and gives up 0.10433. Neither beats the uncorrected count’s 0.08589.
What a repair would have to be able to say
It is worth writing down what the failed repairs have in common, because the successful one is defined by not having it.
Both of the penalty repairs leave the fit alone. Every candidate is still estimated by ordinary least squares, which stopped being the efficient estimator the moment the errors became correlated, so the model that gets handed over is worse than the same data could produce whichever candidate is named. A penalty can change which candidate is named. It cannot make the named candidate’s coefficients any better, and under a persistence of 0.85 the difference between a least-squares fit and an efficient one is larger than the difference between any two candidates in the table.
That is why the measurement below has two regrets in it rather than one. Selection quality is how far a rule’s chosen candidate is from the best candidate under that rule’s own estimator, and on that measure the penalty repairs are merely mediocre. Delivered quality is how far the model that actually gets handed over is from the best model available from anything, and on that measure they are far behind, because the estimator underneath them is the wrong one and no penalty knows it.
An effective sample size cannot make that distinction either, and neither can a trace. Both are statements about how to score a fit. The thing that was lost when the rows started repeating each other was not only the scoring.
What the number is still good for
None of this makes the effective sample size a bad quantity. It is exactly right about the thing it was derived for, and the things it is derived for are not rare: a mean and its standard error, a difference of two means, the width of an interval about a level. A time series of a hundred and twenty persistent observations really does carry about twenty-two independent observations’ worth of information about its own mean.
What it is not is a property of the sample. The information a sample carries depends on what is being estimated, and a design matrix is exactly a statement of what is being estimated. Two candidates on the same hundred and twenty rows are entitled to different corrections and this field measures them differing by a factor of nearly two.
The honest form of the sentence is: this sample is worth 21.7 independent observations for estimating its mean. The clause at the end is not pedantry; it is the whole content.
And there is a diagnostic that follows from it and costs nothing. Whenever an effective sample size is about to be used for something other than an average, the question to ask is what the weights are. For a mean they are 1/n. For a least-squares coefficient they are a row of (X′X)⁻¹X′, which depends on the design and is different for every candidate — so the answer is different for every candidate, and any procedure that reports one number has already decided not to look. The reason nobody notices is that almost every table anybody fits has regressors that move at roughly the same speed, and on such a table the spread this field measures collapses to nothing. The scalar is not usually wrong by much. It is wrong by a factor of two on the first table built to find out.
What is claimed here, and what is not
This essay takes what an effective sample size is a property of, and the claims are that the finite inflation factor is exactly the trace for an intercept-only fit, that the quoted limit form is not even right about that one, that the scalar is above the correction every other candidate in the table needs, and that substituting it into Akaike’s criterion changes the selection on zero draws out of two hundred.
What stays out and is named as a decision: whether any penalty repair can work, which is settled in the next essay and settled against; anything about resampling, which is the other half of this field; and the effective sample size of a Markov chain, which is a different object with the same name and belongs to a field about assignments.
The boundary against the field about dependence in series is that it uses the inflation factor for the estimator it is correct for — an interval about a mean — and this one asks what happens when it is carried to an estimator it is not.
The checks, and the refusals that make them mean something
Two claims are gated in this field’s library. The finite effective sample size is required to equal the intercept-only trace to within a billionth, which is an identity rather than an agreement and would break under any error in either. And substituting an effective sample size into Akaike’s criterion is required to select the identical candidate on every draw — an equality, because there is nothing in 2q for a sample size to move.
The refusal for this essay is one effective sample size applied to every candidate. It is not rejected for being approximate: it is rejected for being outside the range of what it stands in for, at 5.537 against corrections running from 2.550 to 4.743. An approximation that is nearer to none of its targets than they are to each other is not a summary of them.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The window that has to be chosen, and the term that was dropped — both name autocorrelation, degrees of freedom, dependence, generalised least squares, information criterion, least squares, model selection, persistence
- A covariance with no parameter in it — both name autocorrelation, dependence, generalised least squares, information criterion, least squares, model selection, persistence
- A line that beats two curves — both name closed form, degrees of freedom, dependence, information criterion, least squares, model selection, optimism
- A lag the sample has less of — both name closed form, degrees of freedom, dependence, information criterion, model selection, optimism
- The width a band is measured in — both name closed form, degrees of freedom, dependence, information criterion, model selection, optimism
- What the correction assumes — both name closed form, degrees of freedom, dependence, information criterion, model selection, optimism
Named objects
A flat tag is an object no other essay names yet.
AutocorrelationClosed formDegrees of freedomDependenceEffective sample sizeGeneralised least squaresHat matrixInformation criterionLeast squaresModel selectionOptimismPersistenceSample sizeStandard errorVariance