The charge that is not a sum
Worth reading first: The observations that repeat each other · A design is a number.
The pair of searches manufactures 99.40 where the two apart manufacture 118.70. That is an arithmetic fact about two suprema and it does not yet say what anybody should do.
This essay attaches it to a decision — is there a break in this regression? — and reports what three defensible thresholds do to the answer. Two of the three are not tests. The third is, and it is the one nobody would have arrived at by combining published advice.
Three thresholds
A count of coefficients. A split adds five, so a chi-square on five degrees of freedom, whose 95% point is 11.07. This is what a table of critical values gives and what a great deal of applied work uses.
The break search’s own charge. The searched break’s likelihood ratio has a distribution that can be simulated, and its 95% point under a first-order autoregression is 69.55. That is the honest threshold for the search this essay’s field is about — as long as it is the only search being run.
The charge the pair needs. The statistic a rule doing both searches actually reads is what the break search adds after the window search, and its 95% point is 27.18.
What each one does
A break of size zero is a null and a break of size one is an alternative, and the same sweep gives both.
The count of coefficients declares a break on 73.6% of samples that have none. Under long memory, 87.6%. That is not a test with an inflated size; there is no sense in which it is testing anything. Its power at the largest break measured is 90.0%, which is only slightly above its size and is the signature of a threshold that fires on almost everything.
The break search’s own charge declares a break on 0.0% of null samples — and on 0.0% of samples with the largest break measured. Not a low rate: none, at every shift in the sweep, across two hundred and fifty draws each.
The calibrated charge holds 5.6% at no break and reaches 9.6%, 16.8% and 29.2% at shifts of 0.3, 0.6 and 1.
Conservative is the wrong word
The middle row is the interesting one, because it is what a careful practitioner would do.
Two literatures exist. One says a searched break point must not be charged as a parameter and gives a simulated critical value. The other says a whitening’s bandwidth should be chosen by a criterion and gives a rule for it. Neither mentions the other. Combining them means taking the break’s published threshold and applying it inside a procedure that also chooses its window — which is exactly the middle row.
The word that gets used for the result is conservative, and it carries a suggestion that the cost is a little power in exchange for safety. Here the cost is all of the power, at every alternative measured, and the safety was already there: the calibrated charge holds its size at 5.6%.
The arithmetic of why is the previous essay’s. The break search’s own charge is calibrated against a statistic whose average is 34.70; the statistic a joint rule reads averages 15.40. Threshold 69.55 against a statistic that rarely exceeds forty is a threshold nothing crosses.
The sum would be worse
It is worth pricing the arrangement the field’s title is about, since it is the one a reader will imagine.
Charging the two searches separately means a threshold of 63.5 for the break plus 126.1 for the window, or 189.6 for the pair, against a joint statistic whose own 95% point is 146.4. A rule using that threshold rejects less often than one calibrated at 146.4, which already rejects at 5% by construction — so the sum is a threshold at roughly the 99.7th percentile of the statistic it is applied to, and its power against everything in this sweep is nothing.
Every one of the three over-charging arrangements ends in the same place: a test that cannot fire. The count of coefficients ends at the other extreme, a test that cannot help firing. Only the statistic-specific calibration lands anywhere useful, and it is not obtainable from either literature alone.
A threshold should be about twice its statistic’s null mean
The three thresholds are 11.07, 69.55 and 27.18, and the three statistics they belong beside have null means of 34.70, 34.70 and 15.40. Dividing gives a diagnostic sharper than any of the rates.
The count of coefficients sits at 0.32 of its statistic’s mean. The break search’s own charge sits at 2.00 of it. The calibrated charge sits at 1.76.
A 95% point for a supremum of this kind lands at about twice the null mean, and the two thresholds that were derived for a statistic do. The count of coefficients is at a third of it, which is not a slightly loose threshold — it is below the middle of the distribution it is being applied to, so a statistic exceeds it on most draws by construction. That is why 73.6% and 87.6% are the rates rather than something in the twenties.
The ratio is worth having because it is checkable without a simulation. Any threshold whose value is under the average of the statistic it is applied to is a threshold that fires more often than not, and that comparison costs one mean. A count of parameters applied to a searched statistic fails that test immediately, and it would have failed it before any of this field’s sweeps were run.
What the sum is, in percentiles
The sum of the two separate charges is 189.6 against a joint statistic whose 95% point is 146.4, and converting says how far past the end it lands.
The joint statistic’s null mean is 99.40, so a 95% point at 146.4 puts its standard deviation near 28.6. A threshold of 189.6 is then 3.15 standard deviations above the mean, which for a roughly normal statistic is a size under one tenth of one per cent.
So the arrangement the field’s title is about is not a conservative test. It is a test at a nominal level three orders of magnitude below the one it claims, and its power against everything in this sweep is what a test at a level of 0.001 has against a one-standard-deviation step detected across about ten independent observations: nothing.
Read that way the three arrangements are one scale rather than three procedures. A count of coefficients sets a level of roughly three quarters, a separate charge sets one of roughly a thousandth, and the calibrated charge sets 5% — and the two failures are the same failure with the sign reversed, both produced by applying a threshold derived for one statistic to another.
A break the sweep can find, and one it cannot
The alternative in this sweep is a shift added to the response after row sixty — an intercept break, which is the simplest structural break there is. The shifts measured are 0.3, 0.6 and 1, against errors of unit variance, so the largest is a one-standard-deviation step in the middle of a hundred and twenty correlated rows.
The calibrated threshold reaches 29.2% there. That is low, and it is worth saying why rather than treating it as a disappointment.
A one-standard-deviation step in sixty rows of an autoregression at 0.8 is not a large signal. The effective number of independent observations in those sixty rows is about a sixth of their count, so the step is being detected against roughly ten independent observations on each side — and it has to clear a threshold set by a search over a hundred break points and eight window widths.
So the low power is a statement about the sample rather than about the threshold, and the comparison that matters is between the three thresholds at the same alternative rather than against any absolute standard. At the largest shift the three read 90.0%, 0.0% and 29.2%, and only one of them was holding its size while it did.
Where the calibrated charge does not transport
The calibrated threshold has its own limitation and it is the same one the earlier field’s charge has: it is calibrated on one law and it does not carry to the others.
At 27.18 the size is 5.6% under the autoregression it was calibrated on, 0.8% under the moving average, and 18.0% under long memory and under a break in the persistence. Three and a half times the nominal rate on two of the four laws.
That is not a defect of this calibration; it is the standing situation in this collection whenever a reference distribution depends on a nuisance. The estimated-covariance field’s charge does not transport either, and the reason is the same: the distribution of a searched statistic depends on the dependence, and the dependence is the thing being estimated.
The practical form is that the calibration has to be run on the series in hand, by simulating from a fitted model of its errors, rather than looked up. That is one more construction on top of a procedure that is already two searches deep, and the honest reading of this field is that a rule doing two searches is expensive to use correctly.
What the sizes say about the four laws
Reading down the size column at the calibrated threshold — 5.6%, 0.8%, 18.0% and 18.0% — the pattern is not random.
The two laws where the threshold is too small are the two whose dependence a band cannot represent: long memory has an autocorrelation still at 0.637 at the eighth lag, and the break law is not stationary at all, so the family the window search chooses from does not contain either. The window search leaves more of the dependence in the residuals, the break search finds more of it, and the statistic runs larger than the calibration expects.
The law where the threshold is too large is the moving average, whose sequence stops dead at the fourth lag and which a band represents exactly. There the window search removes almost all of the dependence, and the break search afterwards has almost nothing to find.
So the size distortion is predictable from a quantity computed before any test is run — how much of the law the band family can hold — and that is worth more than the sizes themselves.
What the size column is measuring
One clarification, because “size” is doing two jobs in this essay and the two are worth separating.
The null here is no break in the mean, and all four laws satisfy it: the coefficients are constant under every one of them. What differs between the four is the error structure, which is a nuisance under this null rather than part of it. So the size column is asking whether the test holds its rate across a nuisance it was not calibrated on, and the answer — 5.6%, 0.8%, 18.0%, 18.0% — is that it does not.
That is the same question a randomisation test answers by construction and a likelihood-ratio test answers by simulation, and it is the reason this collection reaches for the first wherever it can. Here it cannot: a break in a regression’s coefficients is not a hypothesis about an assignment, so there is no permutation to condition on and the reference distribution has to be generated from a fitted model of the errors.
Generating from a fitted model means the reference distribution inherits whatever the fit got wrong, which is exactly what the size column shows. It is the standing cost of a test whose null is composite in the nuisance.
What a practitioner should take
Three sentences.
Do not add published charges. Two corrections derived in two places are calibrated against two statistics, and a rule running both reads a third. The sum is not conservative in a useful sense; on this construction it switches the test off entirely.
Calibrate against the statistic the rule actually reads. That means simulating the whole procedure — window search and break search together — from a fitted null, and taking the quantile of what comes out. It is the only one of the four arrangements here that holds its size.
And expect the calibration not to transport. Between four dependences all standardised to the same first-lag correlation, the same threshold gives sizes from 0.8% to 18.0%. A calibration is a calibration for the series it was run on.
The decision this is not
One thing this essay does not do is compare the errors rather than the rates, and the omission is deliberate rather than an oversight.
The searched-break field ends on exactly that comparison and finds that the two errors are not the same size — missing a real break costs 0.03883 where splitting a stationary sample costs 0.01084, a ratio of 3.6 to one — so the charge that minimises regret is not the charge that holds a 5% size. The same question here would need a loss for a wrongly split regression, and every candidate for it is a modelling choice rather than a measurement.
What is reported instead is size and power, which is what a test’s own literature reports, and which is enough to say that two of the three thresholds are not tests.
Why the middle row is the one to remember
The count of coefficients failing is not news. Every field in this collection that searches for anything reports the same thing, and the searched-break field reports it about a different break in the same series.
What is new here is the middle row, and it is new because it is the row a reader who has already learned the lesson produces. Somebody who knows that searched break points must not be charged as parameters, who looks up the right correction and applies it, and who separately knows that a bandwidth should be chosen by a criterion, ends up with a procedure that never rejects anything.
That is worth a sentence of its own: the failure mode of following two pieces of correct advice is different from the failure mode of following neither, and it is not obviously better. The naive rule declares a break three times in four and a practitioner will notice; the careful rule declares one never, and a practitioner will conclude there was nothing there.
The symptom is absence again, one level up from where this collection usually finds it. A test that never fires produces no wrong numbers, no failing check, and a clean run of nulls.
What is claimed here, and what is not
This essay takes what three thresholds do to a decision made after two searches. The claims are that a chi-square on the five coefficients a split adds declares a break on 73.6% of samples with none under a first-order autoregression and 87.6% under long memory; that the break search’s own 95% point of 69.55, applied inside a rule that also chooses its window, declares a break on 0.0% of null samples and 0.0% of samples with the largest break measured; that the charge calibrated on the statistic the joint rule reads, 27.18, holds 5.6% at no break and reaches 9.6%, 16.8% and 29.2% at shifts of 0.3, 0.6 and 1; and that the same threshold gives sizes of 0.8%, 18.0% and 18.0% on the other three laws, with the direction predicted by whether the band family can represent the law.
What stays out, and is named as a decision: a loss function. Comparing charges by regret rather than by size and power needs a price for a wrongly split regression against a missed break, and unlike the persistence-break field — where both errors are errors in a whitening and can be scored on one prediction — a structural break changes what is being reported, not just how well. Any number put there would be a modelling choice wearing a measurement’s clothes.
Also out: a threshold calibrated per sample. The recommendation above is to simulate from a fitted null on the series in hand, and that is not what is measured here: every rate above uses one threshold calibrated once on one law. A per-sample calibration has its own error, and pricing it would be a third search on top of two.
The boundary against the previous essay is that it measures what the searches manufacture and this one measures what happens when the manufacture is charged for. The boundary against the next is that the calibrated threshold used here is taken as given, and where it comes from is that essay’s subject.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A split that depends on the order — both name dependence, likelihood ratio, monte carlo, selection effect, specification search, structural break, supremum statistic
- Two effects in one number — both name dependence, likelihood ratio, monte carlo, selection effect, specification search, structural break, supremum statistic
- A second break on a flat profile — both name critical value, likelihood ratio, selection effect, specification search, structural break, supremum statistic
- What a zero is made of — both name dependence, likelihood ratio, monte carlo, selection effect, specification search, supremum statistic
- A criterion is a prediction of the hold-out — both name monte carlo, reference distribution, selection effect, specification search
- A detector built for the ordering — both name critical value, monte carlo, statistical power, supremum statistic
Named objects
A flat tag is an object no other essay names yet.
BandwidthCalibrationChi squaredCritical valueDependenceLikelihood ratioMonte CarloReference distributionSelection effectSpecification searchStatistical powerStructural breakSupremum statisticType i error