What a search costs in parameters
Worth reading first: A design is a number · The observations that repeat each other.
An information criterion has a penalty in it, and the penalty is not a convention. It is an estimate of the optimism a fit carries: how much better a model looks on the sample it was fitted to than it will be on a fresh one. The essay that identifies the penalty with the trace makes that exact for a least-squares fit, where the optimism is tr(HΩ) and equals the parameter count when the rows are independent.
The two-regime whitening has a parameter the count cannot see, and the optimism it carries is therefore not 2 per parameter and not anything a count produces. It can, however, be measured.
The charge is a mean
Run the search on a law with no break in it and the whole of the likelihood ratio it reports is manufactured. Averaged over four hundred draws that is 5.080 on the geometric law, 4.839 on the moving average and 4.748 under long memory — three different shapes of dependence, three answers within a quarter of a unit of each other.
A criterion charging two per parameter, with the break point counted as free, charges 2. Counting the break as one more parameter charges 4. What the search actually costs is about 2.5 parameters, and the number is a measurement rather than a count.
The agreement across the three laws is the part that makes it usable. If the charge depended strongly on the shape of the dependence it would have to be estimated from the same sample the criterion is reading, which is a third selection problem inside the second one; it does not, so a single number computed once on a convenient null is a defensible charge for all three.
What the search itself costs, separated from the coefficient
The two conventional charges differ by one parameter, so subtracting says what the search costs as distinct from the fit.
Averaging the three laws gives 4.889 units, or 2.44 parameters. One of those is the extra coefficient a two-regime fit carries. So the break point itself — the fact that its position was chosen by looking — is worth 1.44 parameters.
Forty-four per cent more than counting it as one parameter, and infinitely more than counting it free.
Which says how many break points the search effectively tried
That figure can be read backwards into the quantity a designer would want, which is how much of a search the sample actually supports.
If the hundred or so admissible break positions gave independent tries, the expected maximum of that many chi-squares on one degree of freedom would be around 6.6 units, or 3.3 parameters. The measured cost is 1.44.
Inverting the same approximation, a manufactured excess of 2.9 units corresponds to a maximum over about eleven independent tries.
So a search over a hundred candidate break points behaves like a search over eleven. The reason is immediate once stated — a break at row 60 and a break at row 61 produce nearly the same two-segment fit, so the hundred candidates are a heavily overlapping set rather than a hundred separate chances — and it is what keeps the charge as small as it is.
It also predicts the stability across laws. The effective count is a property of how much two neighbouring break positions share, which is a fact about a hundred and twenty rows rather than about the dependence between them, and the three laws agree to within 0.17 of a parameter.
What the charge is a charge for
The penalty in an information criterion is doing one job and it is easy to forget which. It is not there to express a preference for small models, and it is not a prior. It is there because the fitted value of the likelihood is a biased estimate of the out-of-sample value, biased upward by exactly the amount the fit was able to chase the sample’s own noise, and subtracting the expected size of that chase is what makes two models’ scores comparable.
That reading is what makes the measurement above the right one rather than a plausible one. Under a law with no break, the two-regime model and the one-regime model are the same model, so their out-of-sample performances are equal by construction and the whole of the in-sample gap is optimism. The trace identity makes the same statement for a least-squares fit, where the optimism is computable in closed form and equals the parameter count when the rows are independent; here it is not computable and it is countable to four decimal places by drawing.
A quantity that can be measured does not need to be countable. The count is a shortcut that works when the model is smooth in its parameters, and it is the shortcut rather than the principle.
Why the mean and not the tail
There are two quantities a search’s cost could be measured by and they are not interchangeable.
A test needs a tail. A hypothesis test for a break asks how often a searched ratio exceeds a threshold when there is nothing there, and the answer is a quantile: the 95% points of the three stationary laws are 8.708, 8.150 and 9.126 against a chi-square-on-two point of 5.991. A test using the chi-square table would reject a true null far more often than it claims.
A criterion needs a mean. An information criterion is not a test. It compares two numbers and takes the smaller, and the penalty that makes that comparison unbiased is the expected optimism — an average, not a quantile. So the number this field charges is 5.156, measured on the geometric law, and not 8.708.
The two are quoted together throughout because a reader who has one of them will otherwise assume the other, and they differ by 70%. A break point that is cheap on one scale is not automatically cheap on the other.
What is being charged for
The manufactured likelihood is made of two things and it is worth separating them, because only one is about the search.
Fitting two coefficients instead of one is a genuine extra parameter and would cost about one unit of mean ratio even if the break point were named in advance by somebody who had never seen the data. Searching over seventy-three positions for the place that makes the two coefficients look most different is the rest.
The split is visible in the arithmetic: a chi-square on one degree of freedom has a mean of 1, so the extra coefficient accounts for roughly one of the 5.080 and the search for roughly four. Four units of ratio is what looking costs, and it is twice what the thing being looked for is worth.
That ratio is the field’s headline and it is a fact about this sample size rather than about change points. At a thousand rows the same search would manufacture about the same amount — the number of admissible positions grows with n, but so does the information in each segment, and the two nearly cancel — while a real break of the same size would be worth many times more. The artefact does not shrink and the signal does, which is why a search for a break is a small-sample problem and not an asymptotic one.
An arithmetic that is not a coincidence
One further reading of the same three numbers is worth having, because it turns a measurement into something that can be predicted for a search nobody has run.
A chi-square on one degree of freedom has mean 1. If the seventy-three candidate positions were uninformative about each other, the maximum of seventy-three of them would average about 8.9. If they were perfectly informative — one effective position — it would average 1. The measured 5.080 corresponds to an effective number of independent positions of about eleven, which is the searched range of seventy-three rows divided by roughly seven.
Seven rows is about what it takes for two neighbouring splits to disagree meaningfully about a lag-one coefficient, and that is a quantity a practitioner can reason about before running anything: it is set by how many rows a coefficient needs, which is a property of the model rather than of the sample. So the charge for a wider trim, or a longer sample, or a search over the order as well, is predictable to within a unit from the same arithmetic — and predictable is not the same as measured, which is why the number quoted throughout is the measured one.
The three charges
Three rules present themselves and only one of them is a measurement.
The break point is free. A criterion reading a fitted two-regime model’s parameter list sees two coefficients where the one-regime model has one, charges 2, and stops. On a law with no break at all it takes the split on 98.8% of draws.
The break point is one parameter. The common patch, and it charges 4. It takes the split on 66.5%.
The break point costs what it manufactures. Charging 2 + 5.156 takes it on 15.8%, and still takes it on 61.0% of draws where there really is a break.
Fifteen per cent is not five per cent, and it is not meant to be: this is a criterion rather than a test, and a criterion comparing two models has no size to control. What the number says is that the charge has moved the rule from always split to split when there is something to find, which is what a penalty is for.
What it costs to leave it uncharged
A charge that is too small does not merely make a rule slightly wrong. It makes a rule that always does one thing.
At a charge of 2 the two-regime fit is taken on 98.8%, 98.0% and 95.3% of draws on the three laws that have no break in them — and on 99.3% of draws on the law that does. A rule that says the same thing whatever the data is, is not a rule; it is a decision that was taken before the sample arrived and is being reported as though the sample had said it.
That failure mode is the one this collection meets most often and under the most names. It is the interval that is always short, it is the resampling that always exists at one taper and never at another, and it is the search that always finds something. The symptom is a number close to one where a number close to the truth was expected, and the diagnosis is always that something free was being spent.
What the charge is not
Two things this measurement does not deliver, and both are named rather than implied.
It is not a critical value. The distribution above was measured on four laws at one sample length with one trim, and every one of those is part of the answer. A break search’s asymptotic theory replaces the maximum over positions with a functional of a Brownian bridge, and the resulting critical values depend on the trim in a way a table can carry; that is not what this collection does, and a number read off four hundred draws of one design is a number for that design.
It is not transportable to a different search. Widen the trim and there are more positions to maximise over, so the charge rises. Search over pairs of positions instead of one and it rises again, by about as much as the first search cost. Search over the order of each regime as well and it rises further. The charge is a property of the search, which is why it has to be measured for the search actually run.
Two routes to the same floor
There is a second way to see that four of the five units are the search, and it needs no simulation.
Every candidate position gives a likelihood ratio, and under the null those seventy-three ratios are each approximately a chi-square on one degree of freedom — the one restriction ρ₀ = ρ₁, at a fixed b. They are strongly correlated, because adjacent positions share nearly all their rows, so the maximum of seventy-three of them is not the maximum of seventy-three independent chi-squares. But it is bounded below by the maximum of a much smaller effectively independent set, and above by the maximum of seventy-three independent ones.
The second of those is computable — the expected maximum of seventy-three independent chi-squares on one degree of freedom — and it is 8.9; the first, taking the effective number to be the range divided by a segment length that resolves a coefficient, is about 4. The measured mean of 5.080 sits between them and nearer the lower bound, which is what strong correlation between neighbouring positions predicts. That is not a derivation of the number, and it is a check that the number is the right order of magnitude for the reason claimed rather than for some other reason.
The same shape, three fields along
The charge measured here has two siblings in this collection and putting the three side by side says what kind of quantity it is.
A criterion’s penalty charges for parameters that were fitted, and it is a trace: computable, exact, and equal to the count only when the rows are independent. A specification search’s displacement charges for candidates that were compared, and it is about dimensions rather than about containment. A tuning list’s manufacture charges for settings that were tried, and it is measured at a true null because that is where nothing can be discovered.
All three are the same principle — a rule pays for every degree of freedom it used, whether or not the degree of freedom was a parameter — and all three are different arithmetic. The break point’s charge is the one where the count is not merely inaccurate but undefined, which is why it is measured rather than derived and why the essay before this one spends its length establishing that.
What is claimed here, and what is not
This essay takes what a searched break point should be charged. The claims are that the honest charge for an information criterion is the expected likelihood a search manufactures on a law with no break; that the three stationary laws agree on that quantity to within a quarter of a unit, at 5.080, 4.839 and 4.748; that this is about two and a half parameters where a count would charge one or two; that a criterion needs that mean and a test would need the 95% point, which is 8.708 and 70% larger; and that roughly one of the five units is the extra coefficient and four are the search.
What stays out, and is named as a decision: the charge as a function of the trim and the sample length. Both change it, and neither is swept here. A charge measured for a trim of 0.2 at a hundred and twenty rows is stated as being for a trim of 0.2 at a hundred and twenty rows, which is the only honest way to quote a number that is a property of a search rather than of a model.
The boundary against the essay that describes the search is that it establishes what the statistic is and this one asks what to do with it. The boundary against the essay that spends the charge is that this one produces a number and that one measures what the number buys.
The checks, and the refusals that make them mean something
Three claims are gated. The searched ratio is required to be non-negative on every draw, on every law, which is what makes it a supremum. Its mean is required to exceed what two parameters would cost, on every stationary law rather than on one. And its 95% point is required to exceed the chi-square point for two restrictions, which is the separate statement a test would need and is checked separately for that reason.
The refusal is what the field is named for. A searched break point counted as one parameter is refused, with the manufactured mean and the manufactured tail printed beside the counts, and with the reason a better count cannot fix it: under the null the break point indexes a family of identical distributions, so there is no parameter for a count to be a count of.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two searches, one sample — both name chi squared, critical value, degrees of freedom, identification, information criterion, likelihood ratio, supremum statistic
- A width that moves and an error that does not — both name degrees of freedom, information criterion, model selection, nested models, optimism, out of sample
- The charge nobody derived — both name degrees of freedom, information criterion, model selection, nested models, optimism, out of sample
- A charge that depends on the rule — both name critical value, identification, likelihood ratio, model selection, supremum statistic
- A line that beats two curves — both name degrees of freedom, information criterion, model selection, nested models, optimism
- A search that is already the other — both name change point, degrees of freedom, likelihood ratio, nested models, supremum statistic
Named objects
A flat tag is an object no other essay names yet.
Bias correctionChange pointChi squaredCritical valueDegrees of freedomIdentificationInformation criterionLikelihood ratioModel selectionNested modelsOptimismOut of sampleSupremum statistic