Concept

Specification search — where it appears

Comparing a benchmark model with a table of variants of it, where every variant contains the benchmark and the winner is reported. Take the benchmark out of the table — let it be one candidate among many, chosen by the same data — and the reading rather than the arithmetic decides the error rate.

Named by 26 essays across 10 fields — each of them below, with the objects they name alongside it.

The profile a break point is chosen from. One sample of 120 rows under a break in the persistence, fitted as two first-order regimes at every admissible break point. The maximum is at row 78, where the true break is at 60. The shaded band is every break point within two log-likelihood units of the best one — 8 of the 73 positions searched, which is 11% of the range. The horizontal line is the one-regime fit the search is compared against; the whole profile is above it, at every position, which is the point: a maximum over 73 candidates is above the null by construction and not by evidence.

A break that was looked for

A two-regime whitening finds its change point by maximising a profile, and then reads a criterion that counts parameters. Under no break there is no parameter to count, because every position describes the same model.

charged · Break point
What each rule gives up against an oracle that is arithmetic. Expected squared error of the candidate each rule selects, minus the expected squared error of the best candidate in the table, over 500 draws of 120 rows. Both quantities are closed forms — σ_S²(1 + q/(n − q − 1)) — so the only Monte Carlo here is over which candidate got picked. The hold-out spends half its rows measuring what the criterion computes, and pays 1.8 times as much for it. Schwarz's criterion is worst because it is answering a different question: which candidate contains the truth, rather than which one forecasts best.

A criterion is a prediction of the hold-out

A rolling hold-out spends half the sample measuring what a criterion computes from all of it. Against an oracle that is arithmetic rather than an estimate, the criterion gives up 0.01701 and the hold-out 0.03200 — and the number the hold-out reports for its own winner is optimistic by more than either.

proxy · Order-selection
What each variant loses before anything has been searched for. The mean loss differential of each of the eight variants against the benchmark, over 600 tables of 60 origins, with every fit given 71 rows. The series is an AR(1) and every variant adds a lag whose coefficient is zero, so in population the two forecasts are the same forecast and the difference drawn here is estimation noise and nothing else. The marked line is σ²(q₁ − q₀)/n = -0.01408, which is an expression in how many coefficients each model has and how many rows it was fitted on — it knows nothing about the series, the persistence or which lag the variant added, and every bar is within a fifth of it. This is the amount a reference distribution recentred at each column's own sample mean believes the candidates are already behind by.

A table of nested models

A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.

search · Forecast
What each rung is made of. Each pair of searches, over 300 draws, split into the two effects its excess is the difference of. The overlap is what the second search loses by having the first already run at its own answer; the interaction is what the joint search finds by moving the first off it. They subtract to the excess exactly, on every draw, because the pinned supremum cancels. Two disjoint dictionaries of independent columns read an excess of 0.000011 and are made of 0.000514 and 0.000503. A break paired with a dictionary of step columns has an interaction of exactly 0 and is all overlap. And a break paired with an independent column has an overlap of -0.004395 against an interaction of 0.002364, which is what puts its excess below zero.

Two effects in one number

How much two searches over one sample share is measured as the net of two things — ground both of them find, and configurations only the joint search reaches. One extra supremum per draw separates them exactly.

separate · Break point
Two searches find some of the same luck. What each search reports on a sample with no break in it, and what the two report together, on four dependences. The dashed line is the sum of the two — what a rule charging each search separately would levy — and the two together always come in below it: 19.30, 15.44, 14.24, 26.64 short, on 100%, 99%, 99%, 100% of draws. The shortfall is not a rounding. Under AR(1) at 0.8 it is 19.30 of the 34.70 the break search manufactures on its own, which is more than half of it. Two searches over one sample are looking at the same noise, and the second one has less left to find.

Two searches, one sample

A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.

twice · Break point
How much of one search the other has already found. Five pairs of searches on one sample, on a scale whose zero and one are both fixed by construction. Zero is two searches over disjoint sets of independent columns: they remove shares of the residual sum that add, at 0.8 standard errors from exactly additive, and they read 0.004. One is a break search paired with a step column it contains, which reads exactly one on every draw because the step adds nothing at all. Between them: two dictionaries of step columns cut a few rows apart read 0.125, and the pair the earlier field measured — a break and a whitening window, both reading the same residual series — reads 0.762, three quarters of the way to one search containing the other. And below zero, a break paired with a search over independent columns reads -0.306: the joint search finds configurations neither half of it contains, so charging the two separately under-charges.

Two searches that share nothing

Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.

apart · Criterion
One true null, one table, five readings. every subset of four, fifteen models, at a null where nothing any candidate holds is worth anything, over 500 draws. Each bar is the share of draws on which that reading declares a difference at a nominal 5%. The reading is the whole of the difference between the bars: the data is identical. An open search over all 210 ordered pairs rejects 76.2%; the table's own 5% point is 3.163 against the 1.671 a single comparison uses. Bonferroni takes the open reading to 0.6% — and on the nested ladder the same correction does not reach the nominal level at all, because there the excess is a shift in the mean rather than a maximum over many.

When the benchmark is a candidate

A specification search with a benchmark nailed down is the case with a closed form. Take the nail out — let the model that would have been reported be one of sixteen, chosen by the same data as its rivals — and the same true null is read three ways, at 2.0%, 7.8% and 76.2%.

select · Forecast
The distribution of the largest statistic in the table. Fit the benchmark to the whole series, resample its residuals, simulate 199 series in which the null is true by construction, re-run the entire eight-variant search on each, and keep the largest statistic. That is the distribution drawn here, and it is the distribution of the thing a specification search actually reports. It is centred at 1.045 — the maximum of eight statistics is not centred at zero however well each of them behaves — and its 5% point is 2.536. A table read against 1.671 is reading the distribution of one statistic; a Bonferroni correction reads it against 2.577 and is nearly right here, because eight variants that each add a different lag are nearly eight separate chances.

A null with a model in it

The distribution to read the winner of a table against cannot be resampled from the data, because the data does not contain the null. It has to be generated from a model — which is the assumption the resampling was chosen to avoid.

search · Bootstrap
What the second search finds, alone and afterwards. For four of the pairs, what the second search removes on its own and what it removes once the first has already run. The gap between the two is the overlap in absolute terms. Where the searches share nothing the two readings are the same: an independent column removes 0.0261 alone and 0.0260 afterwards. Where one contains the other they are 0.1387 and exactly zero. The pair the earlier field measured sits between: a whitening window removes 0.5033 alone and 0.3120 after a break search has run. This is the earlier field's own reading of its pair, on the share scale rather than in log-likelihood units, and it is the number a rule that runs both searches actually has to charge for.

A search that is already the other

A break search shifts every coefficient after a row, so a step column is one of the directions it can move in. Paired with a dictionary of them it reads exactly one, on every draw, and that fixes the top of the scale.

apart · Criterion
Three charges, and only one of them is a test. What each of three thresholds does to the same decision, under AR(1) at 0.8, against the size of a genuine break in the mean at row 60. A chi-square on the 5 coefficients a split adds — 11.07 — declares a break on 73.6% of samples that have none: it is not a test at all. The break search's own 95% point, 69.6, carried into a rule that also chooses its window, fires on 0.0% of null samples and on 0.0% of samples with the largest break measured — the natural way of combining two published corrections does not lose a little power, it switches the test off. The calibrated charge, 27.2, holds 5.6% at no break and reaches 29.2% at the largest.

The charge that is not a sum

Charging two searches what each costs on its own is conservative, and conservative here means the test never fires. At the largest break measured it declares nothing, on every draw, while a calibrated threshold reaches 29%.

twice · Break point
Every candidate is behind by what its parameter count says. Each dot is one of the fifteen subsets of four predictors, fitted on a rolling window of 80 rows and scored against the benchmark out of sample over 60 origins, at a null where every one of them contains the truth. The line is σ²(q₀/(R − q₀ − 1) − q/(R − q − 1)), which is arithmetic on two integers and a window length. Most of these pairs are not nested — a subset of two predictors and a different subset of two share neither model — and the closed form does not care: the displacement is a statement about how many coefficients each side estimates. The candidates of the benchmark's own dimension sit at zero.

The displacement is a parameter count

A nested variant is behind its benchmark out of sample before anything is searched for. The closed form for how far turns out to have nothing about nesting in it — only two integers and a window length — and it prices a table where no candidate contains any other.

select · Multiplicity
A window with two wrong ends. The regret of a rule whitened by a Bartlett-tapered Ω̂, as the window widens, against three rules that need no window at all. At L = 0 the estimate is the identity and the rule is exactly least squares — 0.08556, the same number to five places. It falls to 0.01754 at L = 20 and rises again by L = 30, because a quarter of the sample's lags are then being estimated from it. The automatic bandwidth a practitioner would reach for, 4(n/100) to the power 2/9, which at this sample size is 4, gives 0.02963 — 69% above the best available window. The rule told the dependence is an AR(1) sits at 0.01440 throughout, which is the price of not knowing the form.

The window that has to be chosen, and the term that was dropped

An estimated covariance has a bandwidth in it, and both ends of the dial are wrong for different reasons. The rule a practitioner would reach for is two thirds worse than the best window there is.

banded · Order-selection
Both halves grow; the difference does not. The control pair's two components and their difference, against how much each of its two searches can find, over 1200 draws at each dictionary size. Two disjoint sets of independent columns are additive at every size — the excess stays inside a standard error or two of zero throughout — and it is not because there is nothing there. The overlap grows from 0.000112 at two columns to 0.000870 at ten, a factor of 7.76, and the interaction grows with it, staying within a factor of two of the overlap at every size. Two searches competing for one residual sum share ground and find configurations neither has alone, in almost equal measure, and their difference is what the earlier field's scale calls zero.

What a zero is made of

Two disjoint dictionaries of independent columns read an excess of 0.000116 and are made of an overlap of 0.000583 and an interaction of 0.000467. The control the whole scale is anchored on reads zero because two effects cancel.

separate · Break point
The crossing is in the dependence, not in the split. Regret of each rule as the design and the errors are made persistent at the same coefficient, scored on fresh rows because the closed form assumes exactly what is being taken away. An optimism theorem counts rows; when the rows repeat each other there are fewer of them than there are rows, the penalty is too small for the fit it is correcting, and the criterion starts buying coefficients it should not — its average winner grows from 3.31 coefficients to 3.90. The hold-out never used the theorem and overtakes at ρ ≈ 0.81. Schwarz's criterion, worst of the three on independent rows, is best on repeating ones — its heavier penalty is right for the wrong reason.

Where the two searches cross

The obvious dial between a criterion and a hold-out is how much of the sample to hold out, and moving it never changes the answer. The dial that does is one nobody chooses — how much each row repeats the one before it — and the two rules change places at about 0.81.

proxy · Forecast
What a search costs is not a property of that search. The likelihood ratio a searched break in the regression reports, two ways, on every law. On its own — the whole rule being a split of the sample, at no whitening — it averages 34.70 under AR(1) at 0.8, against the 11.07 a chi-square on the five coefficients a split adds would use as a threshold. Inside a rule that also chooses a window from a list of eight, the same search adds only 15.40 — less than half. Most of what a break search finds under correlated errors is the correlation, and a whitening chosen from the same sample has taken it already. A charge measured for one search, carried into a rule that makes two, is not conservative in some harmless direction: it is measuring a different quantity.

A charge that depends on the rule

The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.

twice · Break point
The split depends on the order. How much two searches share, measured both ways round, over 300 draws. Pinning the first search at its own answer and searching the second gives one overlap; pinning the second and searching the first gives another. A break paired with an independent column reads -0.004395 one way and 0.002163 the other, at 8.70 paired standard errors and on opposite sides of zero. The excess the two components subtract to is the same in both orders by construction, so what changes is only how it is attributed. There is no order-free way to say which of two searches found ground both can reach, and the two orders bracket it.

A split that depends on the order

Run the second search first and pin that instead, and the same draw gives a different overlap and a different interaction — with the same difference. And one pair has no second order at all.

separate · Break point
How often the split is taken, and by which rule. Over 400 draws on each of five laws. The first two rows have no break in them at all, the last two have one at row 60, and the middle one is a moving average. A criterion that counts a fitted two-regime model's parameters and nothing else takes the split on 99% of draws where there is no break. Counting the break point as one more parameter brings that to 67%. Charging what the search actually manufactures — 5.16 units, measured on a law with no break — brings it to 16%, and still takes the split on 61% of draws where there is one.

Choosing whether to break

Charging what the search manufactures takes a rule from splitting a stationary sample on 99% of draws to 16%. It also costs regret, because the two mistakes a rule can make are not the same size.

charged · Break point
Which repair goes with which defect. The share of true nulls rejected at a nominal 5% by a reference distribution generated from the fitted benchmark, over 120 draws with 59 resamples each. Where the errors are well behaved every resampling is fine and all four are conservative. Where the variance is a function of the design, the two that detach a residual from its own row reject 5.8% and 8.3% — and the block bootstrap, which is the resampling three earlier fields on this site reach for, repairs nothing at all, because the dependence it is built for is between origins and the rolling scheme reproduces that on its own. Where the errors are skewed the symmetric multiplier is the one that is wrong, and Mammen's two-point version is the only one of the four that is right in both columns.

Residuals that keep their own variance

A reference distribution for a search has to be generated from a fitted model, and the generator draws residuals. Four ways of drawing them keep four different things — and the one this site has reached for three times repairs nothing at all here.

select · Bootstrap
The row count entered twice, and a penalty is one place. Regret against the best available model as the errors are made persistent. Counting rows more than quadruples; both penalty repairs — the trace, and the scalar effective sample size — are worse than it at every persistence measured; and whitening the sample and keeping the ordinary penalty falls, recovering 86.9% of what counting rows gives up at ρ = 0.85. Doing it at an estimated ρ recovers 80.6%, so having to estimate the dependence from the rows being selected on costs 7.3% of what knowing it is worth. Mallows' forms are drawn beside the logarithmic ones and behave the same, which is what rules the linearisation out.

The repair that was exact and made it worse

A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.

effective · Forecast
Two constructions on one triangle, and a third that is not. Three resamplings that all keep runs of neighbours, on the same residuals at a block length of 5, with the lags running past ℓ so that the tapers separate. A blocked multiplier never moves a residual; a fixed-length moving block moves every one; and they attenuate identically, worst gap 1.4 standard errors, both sitting on γ_resid(k)(1 − k/ℓ)⁺ and both exactly zero past ℓ — so the attenuation is the block boundary rather than the multiplier. The third is the stationary bootstrap, whose runs are geometric rather than fixed: its taper is γ_resid(k)(1 − 1/ℓ)^k, it agrees with the other two at the first lag and at no other, and at lag 6 it still carries 0.0081 where they carry -0.0005.

The triangle that was not the multiplier's

A resampling that leaves each residual on its own row can keep only what the residuals have, times a triangle. A construction that moves every one of them has the same triangle — and the one in this collection's own table has a different taper entirely.

banded · Bootstrap
The ladder is the same ladder under every law. Each pair's overlap under each of the four laws, over 400 draws apiece. The scale is fixed at both ends by construction: two searches over disjoint sets of independent columns read -0.001, 0.013, -0.002, 0.004, and a break search paired with a step column it already contains reads exactly one under every law. Between them the pair that reads one residual series twice runs 0.763, 0.805, 0.577, 0.752 — lowest under long memory, where a whitening has most to do and the break search has least left to find that the whitening has not taken. The rung that moves most is the pair of step dictionaries, from -0.148 under a moving average to 0.419 under a break; and the pair that is negative is negative under all four.

Three quarters of the way to one search

The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.

apart · Criterion
Each repair is for its own defect, and one is for both. The 95% point of the statistic's own distribution in each world, against the mean 95% point of five reference distributions built from one sample. Where the error variance is a function of the design, the two resamplings that detach a residual from its row fall short and the two multipliers that keep it there do not; where the rows repeat each other it is the other way round. With both defects at once the blocked multiplier — drawn once per run of 5 rows, so the residual never moves and its neighbours share a sign — is the closest of the five, at 2.999 against a truth of 3.803. It is still short by 0.804, and that shortfall is the next figure.

Two defects and one resampling

Four resamplings, each the repair for one defect and wrong about the other. Put both defects in the same world and the statistic's 5% point is 3.8028, where the best of the four reaches 2.8326 — until a multiplier that stays on its own row and shares a sign with its neighbours reaches 2.9988.

proxy · Bootstrap
What a second break adds. Over 200 draws, the likelihood ratio a search over one break point reports, and how much more a search over an ordered pair adds on top of it. Under AR(1) at 0.8, which has no break at all, the first search manufactures 5.697 and the second adds 4.278. Under a law with exactly one break — where a second one is as absent as the first was in the row above — the first search reports 9.442 and the second still adds 5.800. Searching for something that is not there costs the same whether or not something else was there to find.

A second break on a flat profile

Searching a hundred and twenty rows for one change point manufactures five units of likelihood. Searching for a second manufactures four more, on a series that has at most one — and on a profile whose whole range is under seven.

charged · Break point
One window for the table, or one each. The regret of the same fifteen-candidate table under AR(1) at 0.8 over 150 draws, with the window attached three ways. Chosen once from the fullest candidate's residuals it gives up 0.02518. Chosen from each candidate's own residuals, with the covariance estimate still shared, it gives up 0.02799 — a paired cost of 0.00281 at 2.0 standard errors for the tuning parameter alone. Estimating the covariance per candidate as well costs 0.01087, so the objection already on record is about 3.9 times the size of the one that was not.

A window for every candidate

The window and the order a whitening needs are chosen once, from the fullest candidate, on an argument that was made about an estimated covariance. A tuning parameter is not a covariance, and the two cost different amounts.

together · Order-selection
A bias against a variance, with the answer in between. How wrong one sample's reference distribution is, split into the two things it is wrong by. Sharing the multiplier over more rows keeps more of the dependence and closes the bias from 1.688 to 0.835; every row it is shared over also removes an independent sign from the 101 the sample started with, and the spread of the resulting quantile rises from 1.307 to 2.172. The distance a practitioner with one sample is actually exposed to is the two together, and it is smallest at ℓ = 5.

How long a block a multiplier shares

Sharing a sign over more rows keeps more of the dependence and leaves fewer independent signs to build a distribution from. The bias falls from 1.6885 to 0.8479 and the spread rises from 1.3073 to 2.1716, and the rejection rate walks straight through its nominal level on the way from 11.3% to 1.3%.

proxy · Reference
What the corner costs when the table is full of hopeless candidates. The benchmark holds two predictors, one of which is worth 1; a third predictor, worth the amount on the horizontal axis, is held only by candidates the benchmark does not contain. At the left the null is true and both procedures hold their level. To the right there is a genuinely better candidate, and the uncorrected reality check finds it 2.7% of the time while the same test with the clearly bad columns recentred finds it 51.0% of the time. The columns doing the damage are the ones nobody would have looked at twice: they are so far behind that they cannot win, and calibrating as though they might is what makes the test blind.

The corner the test is calibrated at

"No candidate is better than the benchmark" is not a null but a face of a region, and a reality check is calibrated at one corner of it. Fill the table with candidates that are hopeless rather than equal and the test finds a genuine improvement 0.0% of the time.

select · Reference

Named alongside it

The objects these essays reach for when they reach for this one.

Monte CarloSelection effectModel selectionCritical valueLikelihood ratioStructural breakSupremum statisticInformation criterionOut of sampleDegrees of freedomDependenceOverfitting

All concepts