Two searches over one sample

A charge that depends on the rule

The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.

Worth reading first: The observations that repeat each other · A design is a number.

There are three numbers a rule that searches for a break might be charged, and the field has been treating two of them as the interesting pair. The third is the one a joint rule actually reads, and it is the finding.

On its own, at a fixed window, a searched structural break manufactures a likelihood ratio averaging 34.70 under a first-order autoregression with no break in it.

After a window search, the same break search on the same draws adds 15.40.

Less than half. The charge for a search is not a property of the search.

What a search costs is not a property of that search. The likelihood ratio a searched break in the regression reports, two ways, on every law. On its own — the whole rule being a split of the sample, at no whitening — it averages 34.70 under AR(1) at 0.8, against the 11.07 a chi-square on the five coefficients a split adds would use as a threshold. Inside a rule that also chooses a window from a list of eight, the same search adds only 15.40 — less than half. Most of what a break search finds under correlated errors is the correlation, and a whitening chosen from the same sample has taken it already. A charge measured for one search, carried into a rule that makes two, is not conservative in some harmless direction: it is measuring a different quantity.
Fig. 1 The same search, charged two ways, on four dependences.

Why the second number is the right one

The quantity a decision reads has to be the quantity the rule computes, and a rule doing both searches computes the joint supremum. Its base is not the no-break likelihood at a fixed window — that arrangement is not something the rule ever visits.

What it visits is: choose the width, then look for the break. The likelihood the break search adds to that is the joint supremum minus the window-only supremum, and it is what the decision hangs on. Everything else is bookkeeping about a rule nobody ran.

That reframing is the whole of the essay, and it is worth noticing that the previous essay’s calibrated threshold is the 95% point of exactly this statistic — 27.18 — rather than of the break search’s own, which is 69.55.

Two searches find some of the same luck. What each search reports on a sample with no break in it, and what the two report together, on four dependences. The dashed line is the sum of the two — what a rule charging each search separately would levy — and the two together always come in below it: 19.30, 15.44, 14.24, 26.64 short, on 100%, 99%, 99%, 100% of draws. The shortfall is not a rounding. Under AR(1) at 0.8 it is 19.30 of the 34.70 the break search manufactures on its own, which is more than half of it. Two searches over one sample are looking at the same noise, and the second one has less left to find.
Fig. 2 The three suprema, where the halving is the difference between two of them.

Why it halves

The mechanism is short and it explains the sizes in the previous essay too.

A break search on unwhitened residuals is mostly finding the dependence. A run of correlated errors above the line, followed by a run below it, is what a level shift looks like, and a search over a hundred break points on a hundred and twenty rows of an autoregression at 0.8 finds one almost every time.

A window search removes that. Whitening at a band of four to eight lags leaves residuals with far less run structure, so the break search afterwards has far less to find.

The two searches overlap because they are looking at the same feature of the sample from two directions, and that overlap is the sub-additivity the field opens with seen from inside.

Three charges, and only one of them is a test. What each of three thresholds does to the same decision, under AR(1) at 0.8, against the size of a genuine break in the mean at row 60. A chi-square on the 5 coefficients a split adds — 11.07 — declares a break on 73.6% of samples that have none: it is not a test at all. The break search's own 95% point, 69.6, carried into a rule that also chooses its window, fires on 0.0% of null samples and on 0.0% of samples with the largest break measured — the natural way of combining two published corrections does not lose a little power, it switches the test off. The calibrated charge, 27.2, holds 5.6% at no break and reaches 29.2% at the largest.
Fig. 3 What the two charges do to a decision, which is what the halving is worth.

On every law

Under the moving average the break search’s charge goes from 27.47 to 12.04; under long memory from 33.44 to 19.20; under a break in the persistence from 46.38 to 19.73. Every one is at least a 40% reduction, and the two ends of the range are the two extremes of how much of the law the window’s family can hold.

Under the moving average — the one law a band represents exactly — the reduction is 56%, because the whitening removes nearly all of the dependence. Under long memory it is 43%, the smallest of the four, because a band cannot reach a sequence still at 0.637 at the eighth lag and the residuals keep some of their run structure.

What the band family does not contain. A band at eight lags holds every autocorrelation past the eighth at exactly zero. Three of these four laws are still correlated there — 0.168 for AR(1) at 0.8, 0.637 for long memory at d = 4/9, 0.663 for a break in the persistence — and cutting a sequence off does not merely approximate it: the truncated sequence stops being a covariance matrix, so there is no point in the family to call the truth. Only the moving average, whose autocorrelations are zero past the fourth lag by construction, is inside. That is what makes it the law the joint fit is worth anything under.
Fig. 4 How much of each law a band can hold, which is what decides how much of the break search’s charge the window search takes away.

So the reduction is predictable from a quantity computed before any search runs, which is the same pattern the previous essay’s size distortions follow.

The four reductions, and what they do and do not order

Worked out from the eight numbers, the reductions are 55.6% under the autoregression, 56.2% under the moving average, 42.6% under long memory and 57.5% under the break in the persistence.

Long memory stands apart and the other three are inside two points of each other, with the break law fractionally the largest rather than the moving average. So the quantity computed before any search runs — how much of each law a band can hold — separates one law from three and does not order the three. That is what a mechanism of this shape should be expected to do: a band represents a moving average exactly and an autoregression at 0.8 nearly so, and the difference between exactly and nearly is smaller than the noise on three hundred draws.

In absolute points the ordering is different again and is not close. The break law’s charge falls by 26.65, against 19.30, 15.43 and 14.24 for the others — because it starts at 46.38, which is a third above the next highest. A law that manufactures more spurious likelihood has more for a whitening to take away, and the relative reduction hides that entirely.

Where the residual charge sits against its own threshold

The four charges after the window search — 15.40, 12.04, 19.20 and 19.73 — are all comfortably below the calibrated 95% point of 27.18, as they must be, since that point is the 95th percentile of the same statistic under the autoregression.

The ratio is worth one line because it says something about the shape of the two distributions. The joint statistic averages 15.40 against a 95% point of 27.18, a ratio of 0.57. The break search’s own statistic averages 34.70 against a 95% point of 69.55, a ratio of 0.50.

So the statistic the joint rule reads is not merely smaller; it is relatively tighter, with its upper tail closer to its centre. That is the reason the two thresholds are not in the same proportion as the two means — 27.18 is 39% of 69.55 while 15.40 is 44% of 34.70 — and it is why a threshold cannot be rescaled from one arrangement to the other even approximately.

Each search moves the other’s answer

If the two searches were separable the charges would be a bookkeeping question. They are not, and the evidence is in where each one lands.

The break point found with the window searched differs from the one found at a fixed window on 75.6% of draws, and by more than ten rows on 28.4%. The standard deviation of the displacement is 20 rows on a sample of a hundred and twenty — which is to say the two rules are not disagreeing about a detail, they are choosing different break points.

And the set of rows the profile cannot separate from its best grows from 8.9 to 18.1 when the window is searched too. A joint search returns a break point with twice as wide a plateau around it.

Each search moves the other's answer. Which window a rule chooses, with and without a break in the model, under AR(1) at 0.8. Without one it takes 4 or 8 lags on 99% of draws; with a break searched too the mass moves to the shorter widths, because a split of the sample has already absorbed some of what the band was there for. The other direction is larger: the break point found with the window searched differs from the one found at a fixed window on 76% of draws and by more than ten rows on 28%, and the set of rows the profile cannot separate from its best grows from 8.9 to 18.1.
Fig. 5 Which window a rule picks with and without a break in the model, and how far the break point moves when the window is searched.

The other direction is smaller and is still there. Without a break in the model the rule takes four lags on 37.6% of draws and eight on 61.2%; with a break searched too the mass moves to the shorter widths — four on 63.2% and eight on 29.6% — because a split of the sample has already absorbed some of what the band was there for.

A worked reading of one sample

The averages hide something a single sample shows plainly, and it is the picture the field opens with for a reason.

On one draw of a hundred and twenty rows, the searched break profile drawn at four widths of the whitening peaks at rows 49, 80, 92 and 96. Four widths, four break points, spread over half the sample. Nothing about the sample changed between the four curves; the only thing that changed is how much of the errors’ dependence was removed before the search ran.

A practitioner who fixes the window at eight lags and reports a break at row 92 has reported a quantity that would have been row 49 at no whitening. That is not a confidence-interval-sized disagreement, and no output of either rule says the other exists.

Four windows, four profiles, four break points. One sample of a hundred and twenty rows under AR(1) at 0.8, with the searched break profile drawn at four widths of the error covariance. Each curve is measured against its own no-break likelihood, so what is compared is the shape rather than the level. This is one sample of the 76% on which the two rules disagree, chosen for that; the four peaks are at rows 49, 80, 92, 96, and a rule that chooses the window and then searches for a break is not choosing between four readings of one profile, it is choosing between four profiles. The unwhitened one is the tallest, which is most of what this field is about — a break search under correlated errors reports the correlation, and a whitening chosen from the same sample has already taken it.
Fig. 6 One sample’s break profile at four widths of the error covariance, with the peak marked on each.

That draw is one of the 75.6% where the two rules disagree, chosen for that. What makes it worth drawing rather than tabulating is that the four curves are all perfectly reasonable-looking profiles with clear maxima. There is nothing in any of them to suggest the answer depends on a nuisance.

The displacement is larger than either plateau

Two numbers in the section above are about the same rows and they are usually read separately. Put together they say the disagreement between the two rules is not a flatness effect.

The plateau is 8.9 rows at a fixed window and 18.1 with the window searched, so the half-width of the wider one is about nine rows. The standard deviation of the displacement between the two rules’ break points is 20 rows.

The two rules land outside each other’s plateaus. If the disagreement were the profile being flat — two positions the likelihood cannot separate, with the rule picking whichever happens to be higher on a given draw — the displacement would be contained by the plateau and its spread would be smaller than nine rows, not more than twice it.

The shape of the displacement says the same thing more sharply. A quarter of draws move not at all; 28.4% move by more than ten rows; and the standard deviation is 20. Those three cannot be reconciled by a smooth distribution. Taking the draws inside ten rows as contributing a few tens to the mean square, the far group has to carry nearly all of the 400 — which puts its root-mean-square displacement near 37 rows, close to a third of the sample.

So the picture is a spike and a long tail rather than a wobble. On most draws the window search moves the break a little or not at all, and on something under a third of them it moves it to a different part of the series entirely — which is exactly what the four-width profile of one sample shows, with peaks at rows 49 and 96.

What the plateau means

The doubling of the plateau is worth more than a sentence, because it is a statement about what a joint rule can report.

A break point’s interval, in the convention this collection uses, is the set of positions the profile cannot separate from its best at two log-likelihood units. Nine rows wide is already a poor interval on a hundred and twenty-row sample — the earlier field refuses a two-unit interval read as a 95% one — and eighteen rows is a sixth of the sample.

So a rule that searches its window too does not merely have a smaller statistic to test with. It has a less identified break point: the extra freedom in the nuisance flattens the profile, because a break at a slightly different place can be partly compensated by a slightly different width.

That is a general property of profiles over jointly searched nuisances and it is the reason two searches cost less than the sum: the extra freedom that flattens the profile is the same freedom that lets the second search find what the first already took.

The direction the window moves

The smaller half of the interaction has its own reading and it is the one a practitioner can act on.

With no break in the model, the rule takes eight lags on 61.2% of draws and four on 37.6%. With a break searched too, four on 63.2% and eight on 29.6% — the mass moves down.

The reason is that the two constructions are partly interchangeable. A run of correlated errors can be absorbed by a wider band or by a split of the sample, and a rule allowed both takes some of it each way. So a narrower window is what a joint rule wants, and a practitioner who fixes the window using a rule calibrated for a no-break model is using a width chosen for a different problem.

The size of that effect is small beside the break point’s displacement — a shift of about one step on the window list against twenty rows on the break — and it points the same way: neither search has a well-defined answer until the other is settled.

The reading, stated as a rule

A charge is a property of the procedure, not of the step. The published correction for a searched break point is correct for a rule whose only search is the break. Dropped into a rule that also chooses a bandwidth it is nearly a factor of two too large, and what that does to a test is to switch it off.

The general form is uncomfortable, because it means corrections do not compose. A practitioner who assembles a procedure from three published steps, each with its own correction, has a procedure whose correction is none of the three and is not their sum either. The only reliable route is to simulate the whole assembled procedure under a fitted null and take the quantile — which is what the reference distribution a randomisation test supplies for free does in a setting where it is available, and this is not one of them.

What this does not say

It does not say the break search’s own charge is wrong. It is right for the rule it was measured on, and that rule is a real one — a practitioner who fixes the window in the protocol and searches only for the break is running exactly it.

Nor does it say the reduction is always downwards. Two searches over one sample are always sub-additive, so the joint charge is always below the sum; whether the second search’s own contribution falls depends on whether the two look at the same feature. Here they do, dramatically. Two searches over genuinely different features would overlap little and the second’s contribution would be close to its solo value.

What is measured is this pair, and the size of the overlap is this pair’s.

Where the field ends

Three essays and one sentence: what a search costs depends on what else the rule is doing.

The first essay measures that two searches manufacture less together than apart — 99.40 against 118.70 under the autoregression, on 99.7% of draws. The second attaches it to a decision and finds that following two correct pieces of advice produces a test that never fires. The third locates the quantity a joint rule actually reads and finds it is less than half the published charge.

None of the three is a criticism of the corrections. All three are consequences of a procedure being assembled from parts that were each measured alone, which is how nearly every applied procedure is built.

The charge a criterion wants, and the charge a decision wants. The regret of a rule that fits two regimes when its criterion says so, swept over what the break point is charged, on 5 laws over 250 draws each. The upper line is the average over the two laws that really have a break and the lower is the average over the three that do not; the middle is an equal-weighted average of all five. Charging more always helps on a stationary law and always hurts on a broken one, and the two do not balance: missing a real break costs 0.03883 where splitting a stationary sample costs 0.01084, a ratio of 3.6 to one. So the charge that minimises regret over an equal mixture is 2, below the 7.16 the optimism argument gives — and the two change places at a share of 29% of worlds carrying a break.
Fig. 7 The sweep over charges for one search, which is the object this field’s three numbers had to be fitted onto.

Two things this shares with its neighbours

The finding has two relatives elsewhere in this collection, and naming them is worth more than either on its own.

A quantity that is a property of the procedure rather than of the step. The order’s shared-against-own comparison turned out to be identical to the window’s once both were scored by the same criterion — the difference had been a property of two implementations rather than of two rules. Here a charge turns out to be a property of the assembled procedure rather than of the search it is a charge for. Both are the same shape: a number attributed to a component belongs to the assembly.

And a correction that does not compose. Every field in this collection prices something a rule does — a search, a tuning choice, an estimated nuisance — and prices it with everything else held fixed. That is the only way any of them can be measured, and it is exactly the condition a real procedure violates. What this field measures is how badly, on one pair, and the answer is a factor of two on the charge and a switched-off test on the decision.

One more thing about the plateau is worth recording, because it is the part a practitioner meets first. A break point reported with a wider plateau around it is not a break point that is harder to find; it is one that is less well located, and the two are easy to confuse in a summary. The joint rule finds a break at least as often as the fixed-window rule does — its statistic is larger by construction — and it says less about where. So a report that gives a break point and a likelihood ratio, with no interval, hides exactly the quantity the joint rule degraded. Nothing about the output of either rule reveals which one produced it, and the difference between eight rows of plateau and eighteen is a sixth of the sample.

What is claimed here, and what is not

This essay takes whether the charge for a search is a property of the search. The claims are that a searched structural break manufactures 34.70 of likelihood ratio on its own under a first-order autoregression with no break in it, and 15.40 once a window has been chosen from the same sample; that the same reduction holds on every law — 27.47 to 12.04, 33.44 to 19.20 and 46.38 to 19.73 — with the largest reduction on the law a band represents exactly; that the break point found with the window searched differs from the one found at a fixed window on 75.6% of draws and by more than ten rows on 28.4%, with a displacement standard deviation of 20 rows; and that the set of rows the profile cannot separate from its best grows from 8.9 to 18.1.

What stays out, and is named as a decision: the same measurement with the searches in the other order. Nothing here runs a break search first and then asks what a window search adds. The joint supremum is symmetric, so C is the same either way, but C − A — what the window search adds after the break search — is a different number and is not reported. It would answer a question nobody is asking: a practitioner chooses a whitening before testing, not after.

Also out: an interval for the break point. The plateau widths above are the profile’s two-unit sets, which the earlier field refuses to read as 95% intervals and which are reported here only as a comparison between two rules. A calibrated interval for a break point found by a joint search would need its own simulation, and it would inherit the transport problem the previous essay measures.

The boundary against the first essay of the field is that it measures the pair and this one measures the part, and the part is what a decision reads.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A split that depends on the order — both name dependence, identification, likelihood ratio, monte carlo, profile likelihood, selection effect, specification search, structural break, supremum statistic
  • Two effects in one number — both name dependence, likelihood ratio, monte carlo, profile likelihood, selection effect, specification search, structural break, supremum statistic
  • A list is not a rule — both name bandwidth, dependence, model selection, monte carlo, nuisance parameter, selection effect, whitening
  • Three quarters of the way to one search — both name critical value, likelihood ratio, selection effect, specification search, structural break, supremum statistic, whitening
  • What fitting them together buys — both name dependence, model selection, monte carlo, nuisance parameter, profile likelihood, structural break, whitening
  • A family before a fit — both name dependence, identification, model selection, nuisance parameter, profile likelihood, whitening

Named objects

A flat tag is an object no other essay names yet.

BandwidthCritical valueDependenceIdentificationLikelihood ratioModel selectionMonte CarloNuisance parameterProfile likelihoodSelection effectSpecification searchStructural breakSupremum statisticWhitening