Paying for a search

A break that was looked for

A two-regime whitening finds its change point by maximising a profile, and then reads a criterion that counts parameters. Under no break there is no parameter to count, because every position describes the same model.

Worth reading first: The observations that repeat each other · A design is a number.

The law no stationary estimate can represent has a persistence of 0.95 for its first sixty rows and 0.65 for the rest. The repair for it is a whitening in two regimes, and the place the regimes meet has to be found — by fitting every admissible split and taking the one whose profile likelihood is highest.

That is a search, and nothing in this collection has charged for it.

What the profile looks like

The profile a break point is chosen from. One sample of 120 rows under a break in the persistence, fitted as two first-order regimes at every admissible break point. The maximum is at row 78, where the true break is at 60. The shaded band is every break point within two log-likelihood units of the best one — 8 of the 73 positions searched, which is 11% of the range. The horizontal line is the one-regime fit the search is compared against; the whole profile is above it, at every position, which is the point: a maximum over 73 candidates is above the null by construction and not by evidence.
Fig. 1 One sample of a hundred and twenty rows, fitted as two first-order regimes at every admissible break point.

The construction is one line. For each candidate position b the two segments are fitted separately, the concentrated Gaussian log-likelihood of the whole series is read off, and the maximiser is b̂. The trim keeps the search away from the ends, where a segment of six rows reports a lag-one autocorrelation of almost anything.

On the sample drawn here the true break is at row 60 and the maximum is at row 78. The two-unit set — every position within two log-likelihood units of the best one, which is the convention a practitioner reads off a profile — holds 8 of the 73 positions searched, and the true break is not among them.

The horizontal line is the one-regime fit the search is compared against, and the whole profile is above it. Every one of the seventy-three positions gives a higher likelihood than not breaking at all. That is not evidence about this sample; it is arithmetic, because a two-regime model contains the one-regime model as the case ρ₀ = ρ₁ and a maximised likelihood over a superset cannot be smaller.

Where the search says the break is when there is none

Where the search says the break is, when there is none. 400 draws under AR(1) at 0.8, which has no break anywhere in it. The search still returns one every time, and what it returns is spread across the whole of the range the trim allows — rows 24 to 96 — with 77.5% of the mass in the interior bins. That shape is the diagnosis. A parameter that is identified under the null has a true value the estimate concentrates on; this one has none, because when the two regimes share a coefficient every position describes the same model. The dashed line is what a flat spread would look like.
Fig. 2 Four hundred draws from a law with no break anywhere in it. The search returns one every time.

Run the same search on a first-order autoregression, which is stationary and has no break anywhere in it. The estimate spreads across the whole of the searched range — rows 24 to 96 — with 77.5% of the mass in the interior bins and no bin holding more than a seventh of the draws. Its average is 60.4 with a standard deviation of 22.3, which is a way of saying it is nearly uniform on the trim rather than a way of saying it prefers the middle.

That shape is the diagnosis, and it is a shape rather than a number.

A parameter that is identified under the null has a true value the estimate concentrates on. This one does not. When the two regimes share a coefficient, every value of b describes the same model — the same likelihood, the same fit, the same residuals — so there is nothing for b̂ to be near and nothing for a standard error to be a standard error of. What the search returns is the argument of a maximum over seventy-three correlated random numbers, and where that lands is a fact about the maximum rather than about the series.

The statistic is a maximum

What a search finds where there is nothing. The likelihood ratio between a two-regime fit at the best break point and a one-regime fit, over 400 draws under AR(1) at 0.8. It is never negative — a maximum over candidate breaks cannot be worse than not breaking — and it averages 5.080 with a 95% point at 8.708. The curve is the chi-square density on two degrees of freedom, which is what a count of restrictions would predict: its mean is 2 and its 95% point is 5.99. The statistic is a supremum over a parameter that does not exist under the null, so it is the distribution of a maximum rather than of a quadratic form, and no count of parameters describes it.
Fig. 3 What the search reports on a law with no break, against the chi-square density a count of restrictions predicts.

The likelihood ratio a two-regime fit reports against a one-regime fit — twice the gain — has three properties worth separating.

It is never negative, on any draw, in any world. That is what makes it a supremum and not a comparison: the search would have to choose a break that made the fit worse, and it never will.

Its mean is 5.080 on the geometric law, with a standard error of 0.094. A chi-square on two degrees of freedom, which is what a count of restrictions would give — one extra coefficient and one break point — has a mean of 2.

Its 95% point is 8.708 where the chi-square’s is 5.991. So a test that read a searched ratio against a chi-square table would reject a true null far more often than it claims, and an information criterion that charged two per parameter is charging for a search that costs about two and a half of them.

The picture is not a chi-square that has been shifted. It is the distribution of a maximum over a whole profile, and it has the shape a maximum has: a floor at zero, a mode above the chi-square’s, and a tail that runs further out.

Why no count of parameters can be right

It is tempting to treat the gap between 5.080 and 2 as a calibration problem — charge two and a half parameters instead of one and move on. The reason that is the wrong repair is worth stating because it decides what the next essay can and cannot do.

A likelihood ratio between nested models is distributed as a chi-square with one degree of freedom per restriction, and the argument behind that needs two things. The restricted model has to sit in the interior of the parameter space, and the likelihood has to be smooth in a neighbourhood of it. Here the restricted model is ρ₀ = ρ₁, which is interior, and the likelihood is smooth in the two coefficients — so far so good. What fails is that b is a third parameter that exists only under the alternative: it indexes the alternative and disappears from the null, so the null is not a point in a three-dimensional space, it is a whole line of points all describing the same distribution.

The consequence is that there is no Fisher information for b and no quadratic expansion around a true value. The statistic is a supremum of a family of chi-squares indexed by b, and its distribution depends on how correlated the family’s members are — which depends on the trim, on the sample length, and on the process the segments are fitted to. A number can be measured; a count cannot be corrected.

The same thing in every world

The three stationary laws report means of 5.080, 4.839 and 4.748, with 95% points of 8.708, 8.150 and 9.126. Different shapes of dependence, the same manufacture.

The law that really does have a break reports 8.757. So the true break at row 60, with a persistence changing from 0.95 to 0.65, is worth 3.68 units of ratio over what the search would have found in a series with nothing in it at all — less than the search manufactures on its own.

What a search manufactures, law by law. The average likelihood ratio a search over 120 rows reports, on each of the four laws, over 400 draws. Three of them have no break at all and report 5.080, 4.839 and 4.748; the fourth has one and reports 8.757, so the real break is worth only 3.677 beyond what the search would have found anyway. The dashed line is 2, which is what an information criterion charges for one extra parameter. A search costs about two and a half of them, and the number is a measurement rather than a count.
Fig. 4 What a search manufactures, law by law, against what a count of parameters charges.

That is the sentence this field exists for. A real break in the persistence of a hundred and twenty rows is smaller than the artefact of looking for one, and any rule that adds the two together without separating them is reading mostly the artefact.

Where the estimate lands when there is a break

The estimated break leans towards the side with more memory. Where the profile likelihood puts the break, over 200 draws, on a sample that changes at row 60. When the persistent half comes first the estimate averages 74.6 — 14.6 rows late, at 17 standard errors. Turning the law over, so that the persistent half comes second, moves the average to 46.1, which is 13.9 rows early. A bias with the same sign both ways would be the trim or the estimator's arithmetic; one that reverses is the likelihood being steeper where the dependence is stronger, so the search hands that side rows it does not own.
Fig. 5 The estimated break against the true one, on a law and on the same law turned over.

The estimated break leans towards the persistent side — by about fifteen rows on a break from 0.95 to 0.65, and by about fourteen rows the other way when the law is turned over. Here that reads as an average estimate of 74.9 against a truth of 60, with a standard deviation of 11.8.

Half the spread of the null case, and a mean fifteen rows out. Both halves matter: the search is informative, in that it points at something rather than nowhere, and it is not accurate, in that where it points is systematically past the truth on the side the likelihood is steeper.

The mechanism is worth one line because it is what makes the lean a mechanism rather than a trim artefact. A segment’s contribution to the concentrated likelihood is dominated by how well its own coefficient whitens it, and a persistent segment has more to whiten — so handing the persistent regime a few rows it does not own costs less than taking a few from it, and the maximiser drifts that way. Turning the law over turns the drift over, which is the only comparison that could tell this apart from a preference for the middle of the trim. The laws field runs exactly that comparison, and the answer reverses.

What the interval convention delivers

A practitioner reading a profile takes the positions within two log-likelihood units of the best one and calls them an interval. It is the right shape — it is what a quadratic likelihood in an identified parameter gives — and neither condition holds here.

How much of the profile the likelihood cannot separate. The number of break points within two log-likelihood units of the best one, over 250 draws on each law, out of 73 positions searched. Under the law that has a break at row 60 the set averages 13.0 positions and contains the truth on 41.2% of draws; turning the law over takes that to 20.8%. The two-unit convention is read off a quadratic likelihood in an identified parameter and neither condition holds here, so what it delivers is not 95% and is not any fixed number: it is whatever the flatness of the profile happens to give.
Fig. 6 How much of the profile the likelihood cannot separate, and how often the set contains the truth.

Over two hundred and fifty draws, the two-unit set holds 13.0 of the 73 searched positions on average — about a sixth of the range — and it contains the true break on 41.2% of draws. Turning the law over takes that to 20.8%, because the lean reverses and the set moves with it. On the stationary law it holds 22.9 positions, nearly a third of the range, and there is no truth for it to contain.

Nothing here is 95%, and nothing here is any fixed number. The two-unit convention delivers whatever the flatness of the profile happens to give, and the flatness is a property of the sample.

The total range of the profile is worth reporting beside the width, because the two together say what kind of object it is. Over the same draws the highest and lowest points of the profile differ by 6.78 log-likelihood units under the break and 5.39 under the stationary law. A profile whose whole range is under seven units and whose two-unit set is a sixth of its domain is not a peaked function with a well-determined maximum; it is a shallow one, and every statement about where the break is is a statement about a maximum of a shallow function of a noisy series.

That is the same shape a flat profile takes in the design field, where an optimum sitting on a tie rather than on a slope is what makes a maximin design robust and what makes its argument hard to locate. Here the flatness is a defect rather than a feature, because the thing being located is a row of the data rather than a value of a parameter.

Nearly uniform, and slightly more than uniform

“Nearly uniform on the trim” is the right description of where the search lands under the null, and it can be made exact, because a uniform distribution over the seventy-three searched positions has a mean and a standard deviation that need no simulation.

Rows 24 to 96 give a uniform mean of 60.0 and a standard deviation of √((73² − 1)/12) = 21.07. The counted values are 60.4 and 22.3.

The mean is inside half a row of the centre, which rules out any preference for the middle of the trim. The standard deviation is 5.8% above the uniform’s, which is the more interesting half: the distribution is not merely flat, it is very slightly anti-concentrated, with a little more mass at the ends of the searched range than a coin over positions would put there.

That is the right direction for a maximum over a correlated family. A candidate near the edge of the trim splits the series very unevenly, so one of its two segments is short and its fitted coefficient is noisy — which makes the profile at that position more variable, and a more variable member of a family wins a maximum more often than a flat comparison of means suggests. The edges are not attractive because they fit better on average; they win because they miss further in both directions.

The practical form of that is a warning about the trim itself. Tightening it does not remove the effect, it only relocates it, since whatever the trim’s new edge is becomes the position with the shortest admissible segment. The pile-up is at the boundary wherever the boundary is put, which is why the shape here is reported rather than repaired.

Five by its mean, three and a half by its tail

The manufactured ratio is said not to be a chi-square that has been shifted, and the two numbers already quoted say so without any further computation.

Its mean is 5.080, and a chi-square’s mean is its degrees of freedom, so on the mean the search is worth five restrictions. Its 95% point is 8.708, which sits between the chi-square points for three (7.815) and four (9.488) degrees of freedom — so on the tail it is worth about three and a half.

The two readings disagree by a factor of one and a half, and they disagree in the direction that matters. A chi-square on five degrees of freedom, which matches the mean exactly, has a 95% point of 11.07 — a quarter further out than the measured one. The searched statistic is less dispersed than any chi-square with its mean, and more dispersed than any chi-square with its 95% point.

So the failure is not a missing constant, and it is not even a missing constant in one direction. A criterion charging 5.08 units per search would be right about the average penalty and would over-charge the extreme cases; a critical value taken from a chi-square on three and a half degrees of freedom would be right about the tail and would under-charge the average. The one repair that works for both is the one this field takes: measure the distribution and use it, rather than choosing which of its moments to match.

What this means for every rule that reads a break

Three constructions in this collection place a break and none of them charges for having looked.

The two-regime whitening is the one this essay is about, and it is selected against a one-regime whitening by a criterion that counts coefficients. The sweep over where to place the break reads the profile’s shape as evidence about the best place to split a sample, on a profile whose whole range is under seven units. And the fitted model that extrapolates is fitted to a residual series whose regime structure was decided by the same search.

None of the three is wrong in what it computes. What changes is that the number each of them treats as information — a likelihood gain from breaking — now has a measured floor under it, and the floor is larger than the signal the design was built to carry.

What is claimed here, and what is not

This essay takes what a searched break point is, before anything is charged for it. The claims are that the profile of a two-regime fit is above the one-regime fit at every admissible position, by construction; that on a law with no break the estimate spreads over the whole searched range with 77.5% of the mass in the interior; that the likelihood ratio the search reports is never negative and averages 5.080 on such a law against the 2 a count of two restrictions would predict, with a 95% point of 8.708 against 5.991; that a real break in this design is worth 3.68 units over the artefact; and that the two-unit interval convention contains the true break on 41.2% of draws under one law and 20.8% under the same law turned over.

What stays out, and is named as a decision: a test. Nothing here is offered as a hypothesis test for a break. The distribution measured is the distribution under four particular laws at one sample size with one trim, and a critical value read off it would be a critical value for those; the literature’s answer to that problem is a functional limit and it is not what this collection does. What the next essay takes instead is the criterion’s question, which needs a mean rather than a tail and has an answer that can be measured.

The boundary against the essay that found the break is that it takes what a two-regime whitening is worth once the break is placed, and this one takes what placing it costs. Neither is a claim about whether the repair is worth making.

The checks, and the refusals that make them mean something

Three claims are gated. The fast profile — the one this field computes through prefix sums, so that a search over pairs is possible at all — is required to find the same break and report the same log-likelihood at every candidate position as the direct loop the laws field uses, to ten decimal places, because everything here rests on the two being one function. The searched ratio is required to be non-negative on every draw, which is what says it is a supremum. And the estimate is required not to concentrate under the null, stated as the interior bins carrying a substantial share rather than as any particular shape, because a shape would be a claim about the trim.

The refusal is the field’s own subject. A searched break point counted as one parameter is refused, with the manufactured mean and 95% point printed beside the chi-square points a count would use — and with the reason the count cannot be made right by choosing a better number: under the null the break point is not identified, so there is no parameter for a count to be a count of.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A second break on a flat profile — both name change point, critical value, degrees of freedom, identification, information criterion, likelihood ratio, model selection, profile likelihood, selection effect, specification search, structural break, supremum statistic
  • Three quarters of the way to one search — both name chi squared, critical value, degrees of freedom, likelihood ratio, selection effect, specification search, structural break, supremum statistic, whitening
  • A split that depends on the order — both name change point, identification, likelihood ratio, profile likelihood, selection effect, specification search, structural break, supremum statistic
  • A search that is already the other — both name change point, degrees of freedom, likelihood ratio, selection effect, specification search, structural break, supremum statistic
  • Two effects in one number — both name change point, likelihood ratio, profile likelihood, selection effect, specification search, structural break, supremum statistic
  • Two searches that share nothing — both name chi squared, degrees of freedom, likelihood ratio, selection effect, specification search, structural break, supremum statistic

Named objects

A flat tag is an object no other essay names yet.

Change pointChi squaredCritical valueDegrees of freedomIdentificationInformation criterionLikelihood ratioModel selectionProfile likelihoodSelection effectSpecification searchStructural breakSupremum statisticWhitening