Two searches over one sample

Two searches, one sample

A searched break in a regression manufactures 34.7 of likelihood ratio where a count of coefficients says 11.1. A searched window manufactures 84.0. The two together manufacture 99.4, not 118.7.

Worth reading first: The observations that repeat each other · A design is a number.

A break point that has been looked for costs more than a count of parameters says, because under the null the break point does not exist — every position describes the same model — and no count of restrictions describes the distribution of a supremum over a parameter that is not there. That field measures the charge for one search, on one profile, and says so.

A rule that also chooses a window is running two.

The construction

The break here is a break in the regression, which is what the word usually means outside this collection and is not what the earlier field searches for: that one breaks the persistence of the errors. A structural break splits the sample at some row and fits separate coefficients on each side, so a split adds five coefficients and a count of parameters says a chi-square on five, whose 95% point is 11.07.

The window is the width of a tapered band estimated from the residuals and used to whiten the sample before any of that happens. It is chosen from a list of eight, it carries the volume it moves — which is not optional once widths are being compared — and it is charged its own dimension, without which the likelihood simply takes the widest band on every draw.

Three suprema, all measured against the same base of no break at no whitening:

A, the break alone at a fixed window. B, the window alone with no break in the model. C, both together.

Two searches find some of the same luck. What each search reports on a sample with no break in it, and what the two report together, on four dependences. The dashed line is the sum of the two — what a rule charging each search separately would levy — and the two together always come in below it: 19.30, 15.44, 14.24, 26.64 short, on 100%, 99%, 99%, 100% of draws. The shortfall is not a rounding. Under AR(1) at 0.8 it is 19.30 of the 34.70 the break search manufactures on its own, which is more than half of it. Two searches over one sample are looking at the same noise, and the second one has less left to find.
Fig. 1 What each search reports on a sample with no break in it, and what the two report together, on four dependences.

What each search manufactures

Under a first-order autoregression with no break anywhere, the break search alone reports a likelihood ratio of 34.70 on average. The chi-square point a count of coefficients would compare it against is 11.07, and its own 95% point is 63.5.

Two things are in that number and they are worth separating. Part of it is what the earlier field measures: a supremum over an unidentified location. The rest is the dependence — a Chow test on correlated errors is inflated even without any search, because the residuals on each side of any split are not independent of each other, and least squares has not been told.

The window search alone reports 84.00. That is not a manufactured quantity in the same sense: the errors really are correlated, so a whitening really does buy likelihood, and B mixes genuine gain with search luck. What matters here is not what B is made of but that it is large — nothing in this collection has ever charged for having chosen a window, and eighty-four is not a rounding.

The two together report 99.40.

The charges do not add

34.70 plus 84.00 is 118.70. The pair reports 99.40, so a rule charging the two searches separately over-charges by 19.30, at 28 standard errors, on 99.7% of draws.

That is not a small correction to a bookkeeping question. The excess is 56% of what the break search manufactures on its own.

Three charges, and only one of them is a test. What each of three thresholds does to the same decision, under AR(1) at 0.8, against the size of a genuine break in the mean at row 60. A chi-square on the 5 coefficients a split adds — 11.07 — declares a break on 73.6% of samples that have none: it is not a test at all. The break search's own 95% point, 69.6, carried into a rule that also chooses its window, fires on 0.0% of null samples and on 0.0% of samples with the largest break measured — the natural way of combining two published corrections does not lose a little power, it switches the test off. The calibrated charge, 27.2, holds 5.6% at no break and reaches 29.2% at the largest.
Fig. 2 What three thresholds do to the same decision, which is where the arithmetic above becomes a test that works or does not.

The same pattern holds on every law: 15.44 under the moving average, 14.24 under long memory and 26.64 under a break in the persistence, on 99.0%, 99.3% and 99.7% of draws. The direction is guaranteed — a supremum over a product of two sets is at least the supremum over either slice through it, so C is at least the larger of A and B, and A + B ≥ C is the claim that the two searches find some of the same luck. What is measured is the size.

Each search moves the other's answer. Which window a rule chooses, with and without a break in the model, under AR(1) at 0.8. Without one it takes 4 or 8 lags on 99% of draws; with a break searched too the mass moves to the shorter widths, because a split of the sample has already absorbed some of what the band was there for. The other direction is larger: the break point found with the window searched differs from the one found at a fixed window on 76% of draws and by more than ten rows on 28%, and the set of rows the profile cannot separate from its best grows from 8.9 to 18.1.
Fig. 3 Where each search lands once the other is running, which is the overlap seen as two distributions.

Why they overlap

The mechanism is the reason the next essay’s finding exists, so it is worth stating carefully.

A break search on unwhitened residuals is largely reporting the dependence. A run of correlated errors looks like a level shift, and the best split of a hundred and twenty correlated rows finds one almost every time — which is exactly why the earlier field’s charge is 5.08 against the 2 a parameter costs, and why this field’s is 34.70 against 11.07.

A window search removes that. Whitening at a band of eight lags leaves residuals with much less of the run structure in them, so the break search run afterwards has much less to find.

The two searches are therefore looking at the same feature of the sample from two directions, and whichever runs second finds what the first left. A + B − C is the size of the overlap, and it is about a fifth of the pair.

Four windows, four profiles, four break points. One sample of a hundred and twenty rows under AR(1) at 0.8, with the searched break profile drawn at four widths of the error covariance. Each curve is measured against its own no-break likelihood, so what is compared is the shape rather than the level. This is one sample of the 76% on which the two rules disagree, chosen for that; the four peaks are at rows 49, 80, 92, 96, and a rule that chooses the window and then searches for a break is not choosing between four readings of one profile, it is choosing between four profiles. The unwhitened one is the tallest, which is most of what this field is about — a break search under correlated errors reports the correlation, and a whitening chosen from the same sample has already taken it.
Fig. 4 One sample’s break profile drawn at four widths of the error covariance, which is the overlap made visible: a rule choosing the window is choosing between four profiles rather than reading one.

The window search is not free, and nothing had charged for it

The measurement that has no precedent in this collection is B.

Every field before this one chooses a window from a list and then reads a criterion as though the window had been handed over. The window a whitening wants chooses it by a criterion on the residuals; the estimated-covariance field chooses it once from the fullest candidate; the per-candidate version chooses one for each row. None of them charges anything for the choosing, and every downstream statistic is read against a distribution computed as though a width had been given.

Eighty-four units of likelihood ratio is what that omission is worth on this construction. It is larger than the break search’s own charge, and the break search’s own charge is the thing an entire earlier field exists to measure.

Three ways to choose a window, and the one that would have been bestRegret under AR(1) at 0.8, over 120 draws. The automatic bandwidth is the plug-in rule a practitioner reaches for, and it is derived to estimate a long-run variance — a sum over the band — where a whitening inverts the matrix that band builds; at n = 120 it is 4, and it costs 50% more than the best window available. Choosing the window by the error model's own likelihood, two per lag, lands at 5.2 on average and costs 35% more. The rule that reads what a whitening actually promises — how flat it leaves the residuals — is not on this chart because it is not a rule: it is monotone in the window and picks the widest one it is offered, on 100% of draws under three of the four laws.regret against the best model available — smaller is betterAR(1) at 0.8 · n = 120 · 120 drawsthe best fixed window, L = 120.01555chosen on the sample, L̄ = 5.20.02097the automatic bandwidth, L = 40.02336told the form: AR(1) at ρ̂0.01174told the whole covariance0.01416120 drawsevery feasible choice lands short
Fig. 5 The width a rule picks and how much its choice moves, which is the search nothing had been charging for.

The honest qualification is the one made above: B is not all manufactured, because the whitening buys real accuracy on a law whose errors really are correlated. Separating those two parts would need a world with no dependence in it at all, where the window search has nothing to find and the whole of B is luck — and that world is not in this field’s four laws, because a field about whitening under dependence has nothing to say at independence.

What the profile looks like at four widths

The overlap has a picture, and the picture says something the three numbers do not.

Take one sample and draw the searched break profile at four widths of the whitening. Four curves, each measured against its own no-break likelihood so what is compared is shape rather than level. On the sample drawn above the four peaks are at rows 49, 80, 92 and 96 — four different break points from four widths of a nuisance nobody is interested in.

So a rule that chooses the window and then searches for a break is not choosing between four readings of one profile. It is choosing between four profiles, and the one it takes is the one whose peak is highest after the volume term and the width’s charge are paid. That is the sense in which the two searches are not separable: neither one has a well-defined answer until the other has been settled.

It also explains the sub-additivity without any order statistics. The unwhitened profile is the tallest of the four, because it is the one still carrying the dependence, and the joint search’s answer is a compromise between a tall profile and a well-charged width.

Under a law with a break

The fourth law breaks its own persistence half way through, and its column is the largest in the table: A of 46.38 and an excess of 26.64.

There is no break in the mean under that law — the coefficients are constant, as they are under all four — so the searched structural break is still finding nothing that is there. What it is finding is a change in the error structure, read through a test that has no way to express one, and the answer is to report a break in the regression instead.

That is a specific and useful warning. A searched Chow test on a series whose variance structure changes will find a break in the mean, at 46 units against a threshold of 11, and nothing about the output says which of the two it found.

The base the three suprema are measured from

One choice in the construction is worth defending because it decides what the numbers mean.

All three suprema are measured against the same base: no break, no whitening, least squares on the whole sample. That is why A and B are directly comparable and why C can be compared with their sum. An alternative would be to measure B from no-break-at-the-chosen-window and A from break-at-no-window, which is what a practitioner reading two separate literatures would implicitly do — and that arrangement has no common origin, so the sum of the two is a sum of quantities measured from two different places.

Having one base is also what makes C − B meaningful, which is the next-but-one essay’s whole subject. It is the ordinary discipline this collection applies to every regret measurement: hold the benchmark fixed while the rules move, or a sweep over a rule becomes a sweep over its own benchmark.

What a search costs is not a property of that search. The likelihood ratio a searched break in the regression reports, two ways, on every law. On its own — the whole rule being a split of the sample, at no whitening — it averages 34.70 under AR(1) at 0.8, against the 11.07 a chi-square on the five coefficients a split adds would use as a threshold. Inside a rule that also chooses a window from a list of eight, the same search adds only 15.40 — less than half. Most of what a break search finds under correlated errors is the correlation, and a whitening chosen from the same sample has taken it already. A charge measured for one search, carried into a rule that makes two, is not conservative in some harmless direction: it is measuring a different quantity.
Fig. 6 The statistic a joint rule actually reads, beside the one a count of coefficients would be compared against.

What a count of parameters says

It is worth stating what would happen to somebody who charged in the ordinary way.

A split adds five coefficients, so a chi-square on five, whose 95% point is 11.07. On a sample with no break in the mean under a first-order autoregression, that threshold declares a break on 73.6% of draws. Under long memory, 87.6%.

Where the search says the break is, when there is none. 400 draws under a break in the persistence, which has no break anywhere in it. The search still returns one every time, and what it returns is spread across the whole of the range the trim allows — rows 24 to 96 — with 89.5% of the mass in the interior bins. That shape is the diagnosis. A parameter that is identified under the null has a true value the estimate concentrates on; this one has none, because when the two regimes share a coefficient every position describes the same model. The dashed line is what a flat spread would look like.
Fig. 7 Where a searched break lands when there is nothing to find, which is why a count of restrictions does not describe it.

Three in four is not a test with a poor size; it is not a test. The two fields’ conclusions coincide here and the sizes differ: the earlier field’s persistence break at a naive threshold rejects far too often, and this one’s regression break at a naive threshold rejects almost always.

What comes next

Two essays follow, and the second one is the reason the field exists rather than being a paragraph.

The charge that is not a sum takes the arithmetic above into a decision: what happens to size and power when a practitioner charges each search what its own literature says, and the answer is that the test stops firing at all.

A charge that depends on the rule takes the sharper half. The quantity a rule doing both searches actually reads is not A and not C; it is C − B, the likelihood the break search adds once the window search has run — and it is 15.40 where A is 34.70. What a search costs is not a property of that search.

The log scale understates the overlap it is measuring

The excess of 19.30 is 56% of what the break search manufactures on its own, and the same pair read on the calibrated share scale of the field built to give overlaps an origin comes out at 0.762. Those are two readings of one relation and the gap between them is not noise.

The share scale’s zero is two searches that share nothing; the log scale’s is not, because a log of a ratio of residual sums is super-additive before any overlap — two searches each removing a tenth of the residual sum remove a fifth together, and the logarithm makes the second one’s gain look larger for having gone second. That artefact pushes A + B − C down, so it partially cancels a genuine overlap.

So 56% and 76% are the same finding with one of them net of an artefact, and the log-scale figure is the conservative one. That is worth knowing before the number is carried anywhere: a rule charging the two searches separately over-charges by 19.30 units of likelihood, and the share of the smaller search that the two genuinely have in common is three quarters rather than a half.

Six point nine by its mean, four by its tail

The naive threshold’s failure can be read two ways and they do not agree, which is the same signature the single-search field reports.

By its mean, the searched break statistic is 34.70 against a chi-square-on-five’s 5 — so the search is worth 6.9 chi-squares. By its rejection rate, a chi-square on five scaled by k rejects at 11.07 with probability 0.736 when 11.07/k sits at the 26th percentile of χ²₅, which is about 2.75 — so k is 4.0.

Mean-equivalent 6.9, tail-equivalent 4.0, and no single scaled chi-square matches both. That is the same disagreement, in the same direction and at nearly the same ratio, that the persistence-break search shows between five degrees of freedom by its mean and three and a half by its 95% point. Two different searches, two different laws, one shape of failure: the manufactured statistic is less dispersed than any chi-square with its mean and more dispersed than any with its tail, so choosing a bigger number of degrees of freedom repairs one end and breaks the other.

The four profile peaks say the same thing about the window in the other direction. Rows 49, 80, 92 and 96 span two thirds of the searched range, and the steps between them are 31, 12 and 4 — so almost all of the movement is between no whitening and some, and the choice among widths past that moves the answer by a few rows. A rule that has decided to whiten at all has made most of the decision the break search is sensitive to, which is why C − B at 15.40 is so much smaller than A at 34.70.

Where this generalises

Two searches over one sample is not a special situation; it is the ordinary one, and the pattern of findings here should transfer wherever two of them overlap.

Any pair of nuisance searches that look at the same feature of the data will be sub-additive, and the overlap will be large exactly when the two are looking at the same thing from different directions. Here the shared feature is the run structure of the errors: a break search reads it as a level shift and a window search reads it as autocorrelation, and each is diminished by the other having run.

Where the two searches look at genuinely different features the overlap should be small, and this field cannot show that because both of its searches read the same residual series. It is worth naming as the limit: the size of the overlap is a property of the pair, and 19.30 is this pair’s.

The one thing that does transport without qualification is the sign. Any two searches over one sample are sub-additive, always, on every draw, because a supremum over a product set is at most the sum of what the two slices manufacture. A rule that charges separately is therefore conservative — and what conservative costs is not what the word suggests.

Everything here is measured on stationary designs with a constant mean, so the null is a null in the strong sense and every unit of likelihood a search reports is manufactured.

What is claimed here, and what is not

This essay takes what two searches over one sample manufacture together. The claims are that a searched break in the regression reports an average likelihood ratio of 34.70 under a first-order autoregression with no break in it, against a chi-square-on-five point of 11.07; that a searched window reports 84.00 and that nothing in this collection has charged for it; that the two together report 99.40 rather than 118.70, so a rule charging them separately over-charges by 19.30 at 28 standard errors and on 99.7% of draws; that the same sub-additivity holds on all four laws, by 15.44, 14.24 and 26.64; and that a count of coefficients declares a break on 73.6% of samples that have none under the autoregression and 87.6% under long memory.

What stays out, and is named as a decision: what B is made of. The window search’s eighty-four units mix genuine gain from whitening a genuinely correlated series with luck from having chosen the width. Separating them needs a world with no dependence, where the whole of B is luck, and a field about whitening has no such world in it. The consequence is that B should be read as what the window search adds, never as what the window search manufactures.

Also out: a break in the errors and a break in the mean at once. The earlier field searches for the first and this one for the second, and a rule that searched for both would be running three searches. That is a real construction and it is not measured here; what is measured is that the second search picks up the first when the first is present, at 46.38 against 34.70.

The boundary against the searched break field is that it prices one search exactly and this one asks what happens when a second is running beside it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A second break on a flat profile — both name critical value, degrees of freedom, identification, information criterion, likelihood ratio, selection effect, specification search, structural break, supremum statistic
  • The comparison that was not made — both name bandwidth, degrees of freedom, dependence, information criterion, monte carlo, selection effect, whitening
  • What a search costs in parameters — both name chi squared, critical value, degrees of freedom, identification, information criterion, likelihood ratio, supremum statistic
  • A list is not a rule — both name bandwidth, dependence, information criterion, monte carlo, selection effect, whitening
  • A search that is already the other — both name degrees of freedom, likelihood ratio, selection effect, specification search, structural break, supremum statistic
  • Choosing whether to break — both name identification, information criterion, selection effect, specification search, structural break, whitening

Named objects

A flat tag is an object no other essay names yet.

BandwidthChi squaredCritical valueDegrees of freedomDependenceIdentificationInformation criterionLikelihood ratioMonte CarloSelection effectSpecification searchStructural breakSupremum statisticWhitening