Choosing whether to break
Worth reading first: A design is a number · The observations that repeat each other.
The charge a searched break point deserves is about five units of likelihood ratio, measured as what a search manufactures on a law with no break in it. This essay spends it.
The rule is the one the estimated-covariance field leaves open: fit two regimes if the criterion says so and one otherwise, then select a model from fifteen candidates using whichever whitening was chosen. What changes between the columns below is only what the break point is charged.
What the charge does to the rate
Three rules, five laws, four hundred draws each. The first two rows of the picture have no break in them at all and the last two do.
Counted as free, the two-regime fit is taken on 98.8%, 98.0% and 95.3% of draws where there is no break — and on 99.3% and 100% where there is. That is not a rule. It is a decision taken before the data arrived, reported as though the data had made it.
Counted as one parameter takes it on 66.5%, 61.5% and 59.0% against 93.8% and 95.5%. Better, and still splitting a stationary sample on two draws in three.
Charged what it manufactures takes it on 15.8%, 13.0% and 15.0% against 61.0% and 59.8%. That is a rule: it says different things about different data, and what it says is mostly right.
The 15.8% is not five per cent and is not meant to be. A criterion comparing two models has no size to control, and calibrating one to reject at a rate would be turning it into a test — which is a different instrument with a different requirement, needing a tail where this needs a mean.
What the charge does to the decision
The rate is not the point. What the rule is for is choosing a model, and what it should be judged on is what the model it chooses gives up.
Under the geometric law, over two hundred and fifty draws, the free rule gives up 0.03147 of regret, the one-parameter rule 0.02813, and the measured charge 0.02338 — a paired gain of 0.00810 at 2.3 standard errors over the free rule. Under the moving average the same three read 0.03577, 0.03288 and 0.02939, a gain of 0.00638 at 3.3. On both of the laws that have no break, the charge earns its keep.
The two fixed rules bracket the whole comparison and are worth reading. Never splitting gives up 0.01773 under the geometric law; always splitting gives up 0.03145. So on a stationary sample, splitting unnecessarily costs 0.0137 — real, and small.
Now the other side. Under the break never splitting gives up 0.09853 and always splitting 0.06364, so missing a real break costs 0.0349 — two and a half times as much. The measured charge, which declines to split on 40% of those draws, gives up 0.07069 against the free rule’s 0.06359: it loses 0.00710. On the same law turned over it loses 0.02079 at 4.1 standard errors.
The three rules as one number each
Five laws is five rates a column, and the columns are easier to compare as two averages: the rate on the three laws with no break and the rate on the two with one.
- Counted as free: 97.4% where there is no break, 99.7% where there is. A separation of 2.3 points.
- Counted as one parameter: 62.3% and 94.7%. A separation of 32.4.
- Charged what it manufactures: 14.6% and 60.4%. A separation of 45.8.
Read as balanced accuracy — the average of getting a break right and getting no break right — the three rules score 51.1%, 66.2% and 72.9%.
The first of those is the useful one to say out loud. A rule that charges nothing for a searched break point performs at 51.1%, which is a coin, and it does so while taking the two-regime fit on 98% of draws and therefore looking decisive rather than random.
What the last unit of charge bought, against the first
The two steps are not equally good bargains, and the exchange rate between the two errors says by how much.
Going from free to one parameter costs 5.0 points of correct breaks and saves 35.1 points of spurious ones — seven points saved per point spent.
Going from one parameter to the measured charge costs a further 34.3 points of correct breaks and saves 47.7 — 1.4 points saved per point spent.
So the exchange rate deteriorates by a factor of five between the two steps. The first correction is nearly free; the second is a genuine trade, and it is being made at close to one for one.
That places the measured charge rather than merely endorsing it. It is well past the knee — the separation was rising at about thirty points per unit of charge over the first unit and about three per unit over the last four — so there is very little left to gain by charging more, and what remains would be bought at worse than one for one.
It also names the cost the field cannot argue away. Even charged correctly, the rule misses two breaks in five. A 60.4% detection rate is not a consequence of the charge being too heavy; it is what a criterion comparing a one-regime fit with a two-regime fit can do on a hundred and twenty rows, and no setting of the charge produces a rule that both finds breaks reliably and leaves stationary samples alone.
The two errors, priced
Before the comparison can be read, the two mistakes have to be priced, and the two fixed rules do it. A rule that never splits and a rule that always splits are both available to anyone and neither reads the data at all, so the gap between them on each law is exactly what the decision is worth there.
On the three stationary laws, always splitting costs 0.01084 more than never splitting, averaged. On the two that have a break, never splitting costs 0.03883 more than always splitting. Those two numbers are the whole of the asymmetry, and they are measured on rules with no charge in them — so nothing about how the charge is calibrated can move them.
The reason for the asymmetry is not subtle. A spurious two-regime whitening is a slightly wrong whitening: it fits two coefficients where one would do, both estimated on half the rows, and the sample is whitened a little worse than it might have been. A missing break is a structurally wrong whitening: no single ρ describes a covariance that changes with position, and the laws field measures what that costs at several times what any tuning does. One error is a matter of degree and the other is a matter of kind.
The two charges disagree
Sweeping the charge instead of picking three of them makes the shape plain. Charging more always helps on a stationary law and always hurts on a broken one, and the two curves do not balance.
Averaged over the laws, missing a real break costs 0.03883 where splitting a stationary sample costs 0.01084 — a ratio of 3.6 to one. So the charge that minimises regret over an equal mixture of the five worlds is 2: the free rule, the one that splits nearly always.
That is a real disagreement rather than a rounding, and it has a mechanism. An information criterion’s penalty is calibrated so that two models’ scores are unbiased estimates of the same thing — a statement about estimation. A decision rule is judged on what it loses, and its two errors cost different amounts — a statement about loss. Nothing in the optimism argument ever promised the two would agree, and here they do not.
What the sweep says about the shape
The sweep is worth reading for more than its minimum, because the two curves have different shapes and the difference is the whole argument.
The stationary curve falls from 0.04168 at a charge of nothing to 0.03084 at a charge of sixteen and then stops moving, because past sixteen the rule never splits and there is nothing left to improve. The broken curve rises from 0.06334 to 0.10230 over the same range, and it is still rising at the end. So the cost of over-charging is unbounded in the range that matters and the benefit is capped — which is the asymmetry again, in a second form.
The two-parameter point sits almost exactly at the bottom of the mixture curve and the mixture curve is nearly flat from a charge of nothing to a charge of four: 0.05035, 0.05029, 0.05061. Over that range the decision is insensitive, and the difference between “the break point is free” and “the break point is one parameter” is worth 0.0003 on an equal mixture. Both of the rules a practitioner would reach for without measuring anything are about equally good, and the measured charge is worse. That is not the result this field expected to produce, and it is the one it produced.
The honest form of the answer
The equal weighting above is a statement about what is expected rather than a measurement, and it is doing a great deal of work. So the answer has to be stated as a share rather than as a verdict.
With R(c, p) = p·(regret on broken laws) + (1 − p)·(regret on stationary ones), the free rule and a charge of six change places at
Below three worlds in ten carrying a break, charge for the search; above it, do not. That is a number a practitioner can compare against their own subject, and it is the form the answer has to take because the measurement cannot supply the prevalence and the practitioner can.
At a charge of four the crossing is at 30.6% and at a charge of eight it is at 24.4%, so the answer is not sensitive to which charge is being compared against — the crossing sits between a quarter and a third whatever the alternative.
The rung this sits on
It is worth placing the whole comparison beside the prices the estimated-covariance field measures, because the numbers here are small and their smallness is informative.
Least squares on the break, told nothing about the dependence, gives up 0.18187. A whitening told the whole covariance gives up 0.01605. Between those two the entire question of whether and where to split lives, and the best any rule here manages is 0.06359. So the split is worth about two thirds of what knowing the covariance is worth, and choosing whether to split is worth a twentieth of the split.
That ordering is the same one this collection keeps finding. Knowing the form of the dependence is worth far more than estimating its parameter well; estimating the parameter jointly is worth far more than choosing the tuning around it. Each rung down the ladder is worth an order of magnitude less than the one above, and every one of them takes a field to measure.
Why the split rate and the regret disagree so much
The gap between the two readings is large enough to be worth a mechanism rather than a shrug.
At the measured charge the rule declines to split on 40% of draws where there is a real break. But the draws it declines on are not a random 40%: they are the ones where the search found the least, which are disproportionately the ones where the break is hardest to locate — and a break that is hard to locate is one whose two regimes are hard to tell apart, which is one where fitting a single regime costs less. So the rule errs where the error is cheapest, and its regret cost is much smaller than its rate would suggest.
The same logic runs the other way on the stationary laws. The 15.8% of draws the charged rule still splits on are the ones where the search found the most, which are the ones where the spurious two-regime fit is most different from the one-regime fit — and therefore most expensive. So its saving is smaller than its rate suggests too.
A rate is a count of decisions and a regret is a sum of their costs, and the two are only the same when every decision costs the same. They never do.
What a criterion is being asked to do
There is a general point underneath the specific one, and it is worth separating from the arithmetic because it applies wherever a criterion is used as a rule.
An information criterion is an estimator. Its penalty makes the in-sample score an unbiased estimate of the out-of-sample score, and calibrating it well means the two models’ scores are comparable. That is a statement with no losses in it anywhere.
A rule that reads the criterion and acts is a decision procedure, and a decision procedure has losses in it by construction. Making the estimates unbiased does not make the decisions optimal, and it does not even tend to: the decision that minimises expected loss is the one that picks the larger model whenever the expected gain exceeds the expected cost, and those two expectations are weighted by how often each kind of world occurs and by what each error costs in it. Neither quantity appears in the optimism argument.
This collection meets the same separation in what the correction corrects, where a familywise rate and a false discovery rate are two different promises about the same procedure, and in the fixed-width interval that keeps one promise and not the other. The pattern is always that a quantity calibrated for one purpose is read as though it had been calibrated for another, and the repair is always to say which of the two is wanted before choosing the number.
Here the answer happens to be that the estimation-calibrated charge is worse for the decision on this mixture, and better on a mixture where breaks are rarer than three in ten. That is a statement about the subject rather than about the arithmetic, and it is the kind of statement a measurement can set up and cannot settle.
What is claimed here, and what is not
This essay takes what charging for a searched break point buys. The claims are that a rule counting the break point as free splits a stationary sample on 95% to 99% of draws and is therefore not a rule; that charging what the search manufactures takes that to 13–16% while still splitting on 60% of draws that have a break; that this earns 0.00810 and 0.00638 of regret at 2.3 and 3.3 paired standard errors on the two stationary laws and loses 0.00710 and 0.02079 on the two that have breaks; that missing a real break costs 3.6 times what an unnecessary split does; and that the two charges change places at a prevalence of 28.9%.
What stays out, and is named as a decision: the prevalence. Nothing here can supply how often a real series has a break in it, and the essay does not pretend to. What it supplies is the crossing point, which is the quantity that turns an unanswerable question into an answerable one.
Also out: a loss function other than this one. Regret against the best model available is the loss this whole collection uses, and a practitioner who cares about forecasting rather than about selecting would weight the two errors differently again. The 3.6 is a property of this loss.
The boundary against the essay that measures the charge is that it produces a number from an argument about optimism, and this one finds that the number is right about optimism and wrong about the decision. Both are true and they are about different quantities.
The checks, and the refusals that make them mean something
Three claims are gated. The measured charge is required to cut the split rate on every law that has no break, and the laws that do have one are required still to be split on a majority of draws — because a charge large enough to stop every split is available to anyone and is not a correction. Missing a break is required to cost more than an unnecessary split, which is what makes the two charges disagree. And the crossing is required to be a share strictly inside the unit interval, because a crossing at zero or one would mean one rule dominates and the essay would have nothing in it.
The refusal is the one carried from the field’s first essay and still standing here: a searched break point counted as one parameter is refused, on the grounds that the parameter does not exist under the null. What this essay adds is that a correctly counted charge is not automatically the right charge either, which is a different objection to a different rule.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A window for every candidate — both name information criterion, model selection, optimism, regret, selection effect, specification search, whitening
- A charge that depends on the rule — both name identification, model selection, selection effect, specification search, structural break, whitening
- Two searches, one sample — both name identification, information criterion, selection effect, specification search, structural break, whitening
- A charge that reads the draw — both name information criterion, model selection, optimism, regret, selection effect
- A list is not a rule — both name information criterion, model selection, regret, selection effect, whitening
- A rate times a size — both name information criterion, model selection, regret, selection effect, whitening
Named objects
A flat tag is an object no other essay names yet.
Change pointDecision theoryError rateIdentificationInformation criterionLoss differentialModel selectionOptimismRegretRisk setSelection effectSpecification searchStructural breakWhitening