A charge that depends on the rule
Worth reading first: The observations that repeat each other · A design is a number.
There are three numbers a rule that searches for a break might be charged, and the field has been treating two of them as the interesting pair. The third is the one a joint rule actually reads, and it is the finding.
On its own, at a fixed window, a searched structural break manufactures a likelihood ratio averaging 34.70 under a first-order autoregression with no break in it.
After a window search, the same break search on the same draws adds 15.40.
Less than half. The charge for a search is not a property of the search.
Why the second number is the right one
The quantity a decision reads has to be the quantity the rule computes, and a rule doing both searches computes the joint supremum. Its base is not the no-break likelihood at a fixed window — that arrangement is not something the rule ever visits.
What it visits is: choose the width, then look for the break. The likelihood the break search adds to that is the joint supremum minus the window-only supremum, and it is what the decision hangs on. Everything else is bookkeeping about a rule nobody ran.
That reframing is the whole of the essay, and it is worth noticing that the previous essay’s calibrated threshold is the 95% point of exactly this statistic — 27.18 — rather than of the break search’s own, which is 69.55.
Why it halves
The mechanism is short and it explains the sizes in the previous essay too.
A break search on unwhitened residuals is mostly finding the dependence. A run of correlated errors above the line, followed by a run below it, is what a level shift looks like, and a search over a hundred break points on a hundred and twenty rows of an autoregression at 0.8 finds one almost every time.
A window search removes that. Whitening at a band of four to eight lags leaves residuals with far less run structure, so the break search afterwards has far less to find.
The two searches overlap because they are looking at the same feature of the sample from two directions, and that overlap is the sub-additivity the field opens with seen from inside.
On every law
Under the moving average the break search’s charge goes from 27.47 to 12.04; under long memory from 33.44 to 19.20; under a break in the persistence from 46.38 to 19.73. Every one is at least a 40% reduction, and the two ends of the range are the two extremes of how much of the law the window’s family can hold.
Under the moving average — the one law a band represents exactly — the reduction is 56%, because the whitening removes nearly all of the dependence. Under long memory it is 43%, the smallest of the four, because a band cannot reach a sequence still at 0.637 at the eighth lag and the residuals keep some of their run structure.
So the reduction is predictable from a quantity computed before any search runs, which is the same pattern the previous essay’s size distortions follow.
The four reductions, and what they do and do not order
Worked out from the eight numbers, the reductions are 55.6% under the autoregression, 56.2% under the moving average, 42.6% under long memory and 57.5% under the break in the persistence.
Long memory stands apart and the other three are inside two points of each other, with the break law fractionally the largest rather than the moving average. So the quantity computed before any search runs — how much of each law a band can hold — separates one law from three and does not order the three. That is what a mechanism of this shape should be expected to do: a band represents a moving average exactly and an autoregression at 0.8 nearly so, and the difference between exactly and nearly is smaller than the noise on three hundred draws.
In absolute points the ordering is different again and is not close. The break law’s charge falls by 26.65, against 19.30, 15.43 and 14.24 for the others — because it starts at 46.38, which is a third above the next highest. A law that manufactures more spurious likelihood has more for a whitening to take away, and the relative reduction hides that entirely.
Where the residual charge sits against its own threshold
The four charges after the window search — 15.40, 12.04, 19.20 and 19.73 — are all comfortably below the calibrated 95% point of 27.18, as they must be, since that point is the 95th percentile of the same statistic under the autoregression.
The ratio is worth one line because it says something about the shape of the two distributions. The joint statistic averages 15.40 against a 95% point of 27.18, a ratio of 0.57. The break search’s own statistic averages 34.70 against a 95% point of 69.55, a ratio of 0.50.
So the statistic the joint rule reads is not merely smaller; it is relatively tighter, with its upper tail closer to its centre. That is the reason the two thresholds are not in the same proportion as the two means — 27.18 is 39% of 69.55 while 15.40 is 44% of 34.70 — and it is why a threshold cannot be rescaled from one arrangement to the other even approximately.
Each search moves the other’s answer
If the two searches were separable the charges would be a bookkeeping question. They are not, and the evidence is in where each one lands.
The break point found with the window searched differs from the one found at a fixed window on 75.6% of draws, and by more than ten rows on 28.4%. The standard deviation of the displacement is 20 rows on a sample of a hundred and twenty — which is to say the two rules are not disagreeing about a detail, they are choosing different break points.
And the set of rows the profile cannot separate from its best grows from 8.9 to 18.1 when the window is searched too. A joint search returns a break point with twice as wide a plateau around it.
The other direction is smaller and is still there. Without a break in the model the rule takes four lags on 37.6% of draws and eight on 61.2%; with a break searched too the mass moves to the shorter widths — four on 63.2% and eight on 29.6% — because a split of the sample has already absorbed some of what the band was there for.
A worked reading of one sample
The averages hide something a single sample shows plainly, and it is the picture the field opens with for a reason.
On one draw of a hundred and twenty rows, the searched break profile drawn at four widths of the whitening peaks at rows 49, 80, 92 and 96. Four widths, four break points, spread over half the sample. Nothing about the sample changed between the four curves; the only thing that changed is how much of the errors’ dependence was removed before the search ran.
A practitioner who fixes the window at eight lags and reports a break at row 92 has reported a quantity that would have been row 49 at no whitening. That is not a confidence-interval-sized disagreement, and no output of either rule says the other exists.
That draw is one of the 75.6% where the two rules disagree, chosen for that. What makes it worth drawing rather than tabulating is that the four curves are all perfectly reasonable-looking profiles with clear maxima. There is nothing in any of them to suggest the answer depends on a nuisance.
The displacement is larger than either plateau
Two numbers in the section above are about the same rows and they are usually read separately. Put together they say the disagreement between the two rules is not a flatness effect.
The plateau is 8.9 rows at a fixed window and 18.1 with the window searched, so the half-width of the wider one is about nine rows. The standard deviation of the displacement between the two rules’ break points is 20 rows.
The two rules land outside each other’s plateaus. If the disagreement were the profile being flat — two positions the likelihood cannot separate, with the rule picking whichever happens to be higher on a given draw — the displacement would be contained by the plateau and its spread would be smaller than nine rows, not more than twice it.
The shape of the displacement says the same thing more sharply. A quarter of draws move not at all; 28.4% move by more than ten rows; and the standard deviation is 20. Those three cannot be reconciled by a smooth distribution. Taking the draws inside ten rows as contributing a few tens to the mean square, the far group has to carry nearly all of the 400 — which puts its root-mean-square displacement near 37 rows, close to a third of the sample.
So the picture is a spike and a long tail rather than a wobble. On most draws the window search moves the break a little or not at all, and on something under a third of them it moves it to a different part of the series entirely — which is exactly what the four-width profile of one sample shows, with peaks at rows 49 and 96.
What the plateau means
The doubling of the plateau is worth more than a sentence, because it is a statement about what a joint rule can report.
A break point’s interval, in the convention this collection uses, is the set of positions the profile cannot separate from its best at two log-likelihood units. Nine rows wide is already a poor interval on a hundred and twenty-row sample — the earlier field refuses a two-unit interval read as a 95% one — and eighteen rows is a sixth of the sample.
So a rule that searches its window too does not merely have a smaller statistic to test with. It has a less identified break point: the extra freedom in the nuisance flattens the profile, because a break at a slightly different place can be partly compensated by a slightly different width.
That is a general property of profiles over jointly searched nuisances and it is the reason two searches cost less than the sum: the extra freedom that flattens the profile is the same freedom that lets the second search find what the first already took.
The direction the window moves
The smaller half of the interaction has its own reading and it is the one a practitioner can act on.
With no break in the model, the rule takes eight lags on 61.2% of draws and four on 37.6%. With a break searched too, four on 63.2% and eight on 29.6% — the mass moves down.
The reason is that the two constructions are partly interchangeable. A run of correlated errors can be absorbed by a wider band or by a split of the sample, and a rule allowed both takes some of it each way. So a narrower window is what a joint rule wants, and a practitioner who fixes the window using a rule calibrated for a no-break model is using a width chosen for a different problem.
The size of that effect is small beside the break point’s displacement — a shift of about one step on the window list against twenty rows on the break — and it points the same way: neither search has a well-defined answer until the other is settled.
The reading, stated as a rule
A charge is a property of the procedure, not of the step. The published correction for a searched break point is correct for a rule whose only search is the break. Dropped into a rule that also chooses a bandwidth it is nearly a factor of two too large, and what that does to a test is to switch it off.
The general form is uncomfortable, because it means corrections do not compose. A practitioner who assembles a procedure from three published steps, each with its own correction, has a procedure whose correction is none of the three and is not their sum either. The only reliable route is to simulate the whole assembled procedure under a fitted null and take the quantile — which is what the reference distribution a randomisation test supplies for free does in a setting where it is available, and this is not one of them.
What this does not say
It does not say the break search’s own charge is wrong. It is right for the rule it was measured on, and that rule is a real one — a practitioner who fixes the window in the protocol and searches only for the break is running exactly it.
Nor does it say the reduction is always downwards. Two searches over one sample are always sub-additive, so the joint charge is always below the sum; whether the second search’s own contribution falls depends on whether the two look at the same feature. Here they do, dramatically. Two searches over genuinely different features would overlap little and the second’s contribution would be close to its solo value.
What is measured is this pair, and the size of the overlap is this pair’s.
Where the field ends
Three essays and one sentence: what a search costs depends on what else the rule is doing.
The first essay measures that two searches manufacture less together than apart — 99.40 against 118.70 under the autoregression, on 99.7% of draws. The second attaches it to a decision and finds that following two correct pieces of advice produces a test that never fires. The third locates the quantity a joint rule actually reads and finds it is less than half the published charge.
None of the three is a criticism of the corrections. All three are consequences of a procedure being assembled from parts that were each measured alone, which is how nearly every applied procedure is built.
Two things this shares with its neighbours
The finding has two relatives elsewhere in this collection, and naming them is worth more than either on its own.
A quantity that is a property of the procedure rather than of the step. The order’s shared-against-own comparison turned out to be identical to the window’s once both were scored by the same criterion — the difference had been a property of two implementations rather than of two rules. Here a charge turns out to be a property of the assembled procedure rather than of the search it is a charge for. Both are the same shape: a number attributed to a component belongs to the assembly.
And a correction that does not compose. Every field in this collection prices something a rule does — a search, a tuning choice, an estimated nuisance — and prices it with everything else held fixed. That is the only way any of them can be measured, and it is exactly the condition a real procedure violates. What this field measures is how badly, on one pair, and the answer is a factor of two on the charge and a switched-off test on the decision.
One more thing about the plateau is worth recording, because it is the part a practitioner meets first. A break point reported with a wider plateau around it is not a break point that is harder to find; it is one that is less well located, and the two are easy to confuse in a summary. The joint rule finds a break at least as often as the fixed-window rule does — its statistic is larger by construction — and it says less about where. So a report that gives a break point and a likelihood ratio, with no interval, hides exactly the quantity the joint rule degraded. Nothing about the output of either rule reveals which one produced it, and the difference between eight rows of plateau and eighteen is a sixth of the sample.
What is claimed here, and what is not
This essay takes whether the charge for a search is a property of the search. The claims are that a searched structural break manufactures 34.70 of likelihood ratio on its own under a first-order autoregression with no break in it, and 15.40 once a window has been chosen from the same sample; that the same reduction holds on every law — 27.47 to 12.04, 33.44 to 19.20 and 46.38 to 19.73 — with the largest reduction on the law a band represents exactly; that the break point found with the window searched differs from the one found at a fixed window on 75.6% of draws and by more than ten rows on 28.4%, with a displacement standard deviation of 20 rows; and that the set of rows the profile cannot separate from its best grows from 8.9 to 18.1.
What stays out, and is named as a decision: the same measurement with the searches in the other order. Nothing here runs a break search first and then asks what a window search adds. The joint supremum is symmetric, so C is the same either way, but C − A — what the window search adds after the break search — is a different number and is not reported. It would answer a question nobody is asking: a practitioner chooses a whitening before testing, not after.
Also out: an interval for the break point. The plateau widths above are the profile’s two-unit sets, which the earlier field refuses to read as 95% intervals and which are reported here only as a comparison between two rules. A calibrated interval for a break point found by a joint search would need its own simulation, and it would inherit the transport problem the previous essay measures.
The boundary against the first essay of the field is that it measures the pair and this one measures the part, and the part is what a decision reads.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A split that depends on the order — both name dependence, identification, likelihood ratio, monte carlo, profile likelihood, selection effect, specification search, structural break, supremum statistic
- Two effects in one number — both name dependence, likelihood ratio, monte carlo, profile likelihood, selection effect, specification search, structural break, supremum statistic
- A list is not a rule — both name bandwidth, dependence, model selection, monte carlo, nuisance parameter, selection effect, whitening
- Three quarters of the way to one search — both name critical value, likelihood ratio, selection effect, specification search, structural break, supremum statistic, whitening
- What fitting them together buys — both name dependence, model selection, monte carlo, nuisance parameter, profile likelihood, structural break, whitening
- A family before a fit — both name dependence, identification, model selection, nuisance parameter, profile likelihood, whitening
Named objects
A flat tag is an object no other essay names yet.
BandwidthCritical valueDependenceIdentificationLikelihood ratioModel selectionMonte CarloNuisance parameterProfile likelihoodSelection effectSpecification searchStructural breakSupremum statisticWhitening