Out of sample — where it appears
Named by 19 essays across 9 fields — each of them below, with the objects they name alongside it.
A criterion is a prediction of the hold-out
A rolling hold-out spends half the sample measuring what a criterion computes from all of it. Against an oracle that is arithmetic rather than an estimate, the criterion gives up 0.01701 and the hold-out 0.03200 — and the number the hold-out reports for its own winner is optimistic by more than either.
A penalty is a trace
Akaike's 2q is not a count of coefficients. It is the answer a trace collapses to when the rows are independent — and once they are not, the trace is still the right object and is no longer the count.
A table of nested models
A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.
The charge nobody derived
A band of lags is charged one log-likelihood unit apiece, because that is what a regression coefficient costs. A band's numbers are not regression coefficients, and measuring what they actually cost puts the convention out by a factor of nearly three.
When the benchmark is a candidate
A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.
A null with a model in it
The distribution to read the winner of a table against cannot be resampled from the data, because the data does not contain the null. It has to be generated from a model — which is the assumption the resampling was chosen to avoid.
The displacement is a parameter count
A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.
The quarrel that changes the winner
A disagreement about the tuning parameter costs 0.031 when it changes which candidate the table selects and −0.0007 when it does not. The distance between the values disagreed about has nothing to do with it.
What a search costs in parameters
An information criterion's penalty is an estimate of the optimism a fit carries. For a break point the optimism can be measured and cannot be counted, and it comes to about two and a half parameters.
Where the two searches cross
The obvious dial between a criterion and a hold-out is how much of the sample to hold out, and moving it never changes the answer. The dial that does is one nobody chooses — how much each row repeats the one before it — and the two rules change places at about 0.81.
A width that moves and an error that does not
Four charges give four widths a factor of four apart and four errors half a per cent apart. The derived charge wins, significantly, by a quarter of what was on offer — and none of the four is an estimate of anything.
Residuals that keep their own variance
A reference distribution for a search has to be generated from a fitted model, and the generator draws residuals. Four ways of drawing them keep four different things — and the one this site has reached for three times repairs nothing at all here.
The repair that was exact and made it worse
A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.
The weight that is a vector
Two forecasts have a best combination and one number describes it. Eight have a best combination too, and the vector describing it puts nothing at all on the forecast with the smallest mean squared error.
Two defects and one resampling
Four resamplings, each the repair for one defect and wrong about the other. Put both defects in the same world and the statistic's 5% point is 3.8028, where the best of the four reaches 2.8326 — until a multiplier that stays on its own row and shares a sign with its neighbours reaches 2.9988.
Calibrated and useless
Six forecasters that are calibrated to 2·10⁻³³ run from resolution exactly 0 to 0.092758, and three forecasters with reliabilities from 0 to 0.013025 have areas under the ROC curve identical to every bit a double carries. Each measure is exactly blind to what the other one sees.
The corner the test is calibrated at
"No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.
What a better charge buys
Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.
A charge that reads the draw
Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.
Named alongside it
The objects these essays reach for when they reach for this one.
Model selectionInformation criterionMean squared errorSpecification searchMonte CarloNested modelsBenchmark forecastOverfittingDegrees of freedomEstimation errorOptimismReference distribution