A search that is already the other
Worth reading first: A design is a number · A break that was looked for.
A scale with a zero and no one is half a scale. The essay that fixed the zero put two searches over independent columns at 0.004 ± 0.005 and showed the reading stable across dictionary sizes. What is missing is the other end: a pair where the second search adds nothing at all, so that a number near one has a meaning rather than merely being large.
The temptation is to take whichever pair measures largest and call it one. That would make the scale a ranking with a normalisation on it, and every reading would depend on which pairs happened to be in the table. The top has to be fixed the same way the bottom was — by a construction that says in advance what the answer must be.
A break contains a step
A break in the regression at row fits one set of coefficients to the rows before it and another to the rows after. Written as an augmented design, that is the original columns plus the same columns multiplied by the indicator — and the intercept’s copy of that indicator is the centred step column itself.
So a search over step columns cut at any of six points inside the admissible range is searching a subset of the directions a break search at those same points already spans. Every configuration the step search can reach, the break search can reach. It follows that the joint supremum equals the break search’s supremum, not approximately but exactly, and that the step search adds nothing.
Over twelve hundred draws, a step search alone removes 0.1387 of the residual sum. After a break search has run it removes exactly zero — the joint supremum and the break supremum agree to the last bit on every draw. The overlap reads 1.000.
That is not a measurement in the usual sense; it is a check that the instrument agrees with an argument made before it. Had it come back at 0.98 the instrument would have been wrong, and the value of running it is that 0.98 is exactly what a subtly mis-indexed joint search would return.
What the same feature at different places is worth
With both ends fixed, the interesting rung is the one in between: two searches over the same kind of feature at places that do not coincide.
Two dictionaries of six step columns each, with cut points interleaved along the admissible range so that every column in one has a near-twin a couple of rows away in the other. Neither dictionary contains the other. Both search over exactly the same feature — where the level shifts — and they do it at different places.
They read 0.125 ± 0.027.
That number is smaller than it looks like it should be, and the reason is worth having.
Two adjacent step columns are highly correlated — an indicator for and one for differ on six rows out of a hundred and twenty — so a great deal of what one finds, the other has already found. But “already found” and “cannot add” are different things. The second column’s residual against the first is a short indicator on those six rows, and searching six such residuals still turns up whichever one happens to align with the noise. The overlap is high in correlation and low in what the second search adds, because what a search adds is a maximum over what is left rather than a share of what was there.
The step dictionaries remove 0.1387 and 0.1390 of the residual sum, add to 0.2777, and jointly remove 0.2604. The shortfall is 0.01729 ± 0.00376, at 4.6 standard errors — real, and an eighth of the smaller of the two.
The arithmetic of what is left
The point is worth making with numbers rather than only in words, because it is the whole of why the step rung sits where it does.
Write for the share the first search removes and for the share the two remove together. What the second search adds is , and the overlap is one minus that divided by . For the step dictionaries: 0.2604 − 0.1387 = 0.1218 added, against 0.1390 alone, so seven eighths of what the second search would have found on its own is still there after the first has run.
Now the correlation. The two dictionaries’ chosen columns are typically neighbours, and neighbouring step columns at a hundred and twenty rows share all but a handful of rows. A correlation of that size between two fixed regressors would leave the second one almost nothing: its residual against the first is a short indicator, its own explanatory power is tiny, and adding it to a fit would move very little.
But neither column is fixed. Each is the argmax of a search over six, and the second search re-chooses after seeing nothing — it is choosing among six residual indicators, and the best of six short indicators against noise is not small. The search recovers most of what the correlation took, which is a general fact about searches and not about steps: correlation between the candidates reduces what a search finds much less than it reduces what any one candidate finds.
The field that measured the best of many correlated candidates makes the same point from the other side, and the arithmetic underneath both is the same order statistic.
Which means correlation is the wrong intuition
It is worth pressing on that, because the intuition it displaces is the natural one.
A reader told that two searches overlap reaches for how similar the things being searched over are. On that reading the step dictionaries — interleaved, highly correlated, searching one feature — ought to be near the top of the scale, and the pair of a break and a whitening window ought to be near the bottom, since a break location and a covariance width have nothing obvious in common.
The measurements are the other way round: 0.125 for the step dictionaries and 0.762 for the break and the window. What decides the reading is not similarity but whether the second search’s directions have already been used up, and a search over a large set of nearly-parallel directions leaves a great deal unused.
That is also why containment is the right thing to put at one. A contained search has nothing left by definition, whatever the correlations are.
Why containment is exact and not merely large
One detail of the containment result is worth dwelling on, because it is what makes the top of the scale a construction rather than a very good approximation.
The joint search maximises over break rows and step columns together. It could, in principle, prefer a configuration with a step column in it — a break at row 40 plus a step at row 70, say — which the break search alone cannot reach. It never does, on any of twelve hundred draws.
The reason is that the step column at row 70, added to a design already containing a break at row 40, is not orthogonal to anything the break search could not have reached instead: a break at row 70 dominates it, because a break shifts every coefficient there and the step shifts only the intercept. So for every configuration the joint search might prefer there is a break-only configuration at least as good, and the supremum over the larger set equals the supremum over the smaller one.
That is a statement about the two search sets and not about the data, which is why it holds exactly and on every draw rather than in expectation. An instrument that returned 0.997 here would be reporting an implementation error — an off-by-one in the admissible range, a step column centred against the wrong mean — and the check is worth running for that reason rather than for the number.
Where 0.125 sits on a scale with both ends pinned
With the ends fixed at 0.004 ± 0.005 and exactly 1.000, the middle rung can be read as a position rather than as a number.
0.125 ± 0.027 is 4.6 standard errors above the zero end and 32 below the one end. Two searches over the same kind of feature, at cut points a couple of rows apart, are seven times closer to sharing nothing than to sharing everything.
The same reading in the units the searches work in. A step dictionary alone removes 0.1387 of the residual sum; after a rival dictionary of near-twin columns has run, it still removes 0.1387 × 0.875 = 0.1214. Two searches over interleaved copies of one feature are 87.5% additive.
And the scale has resolution as well as ends. At two standard errors the middle rung distinguishes about eighteen levels between zero and one, so this is a measurement scale rather than a three-point ordering — a pair reading 0.30 and a pair reading 0.20 are different pairs, not the same pair measured twice.
Why the top end had to be exact rather than merely large
The containment check returns 1.000 with the two suprema agreeing on every one of twelve hundred draws, and the exactness is not a stylistic preference. It is what gives the check any power at all.
Read as a measurement, zero failures in twelve hundred draws bounds the failure rate at 0.25% by the rule of three. Read as what it is — an argument that one search’s feasible set contains the other’s, followed by a check that the code agrees — it either holds on every draw or the code is wrong.
Now compare that with the alternative the essay names: a subtly mis-indexed joint search returning 0.98. The middle rung’s standard error is 0.027, so a rung predicted to be 0.98 could not be distinguished from one at 1.00 by any number of draws this field could afford — the difference is 0.7 of a standard error and would need about two hundred times the draws to resolve.
A prediction of exactly one is testable and a prediction of about one is not. That is the whole reason the top of the scale was fixed by a containment argument rather than by taking the largest pair in the table: the largest pair would have come with a standard error, and a standard error at the top of a scale is a scale with no top.
What the step dictionaries cannot settle
The middle rung is one construction and it is worth saying what a different one would have given, because 0.125 reads as a small number and the choice that produced it was arbitrary.
The cut points are interleaved along the whole admissible range at twelve positions, so adjacent columns are about six rows apart. Interleave them at a hundred and twenty positions and adjacent columns differ on one row: the correlation goes up and the residual each leaves goes down, so the overlap should rise. Space them across two disjoint halves instead and they differ by fifty rows: the overlap should fall.
Neither is measured here. The rung is a point on a curve nobody has swept, and the curve’s parameter is how far apart two searches over the same feature are placed. That parameter has no natural value, which is why the rung is reported as an illustration of the middle of the scale rather than as a measurement of what two step searches cost.
The two ends do not have this problem, and that is the difference between them and everything between them. Zero is fixed by independence and one by containment; both are relations rather than distances, and neither has a dial on it.
The law moves the middle and not the ends
The step rung is the one that moves, and it moves a long way.
Under a first-order autoregression the step dictionaries read 0.120. Under a five-period moving average they read −0.148 — below zero, meaning the two searches together find more than the sum of what each finds alone. Under long memory they read 0.385 and under a break in the persistence 0.419.
The mechanism is what the errors themselves look like. Under long memory the error series wanders, so a step column has a great deal of genuine drift to align with and the two dictionaries align with the same drift; under a moving average the series has no drift at all past the fourth lag, so what each dictionary finds is local noise at its own cut points and the two find different noise.
The ends do not move. The control reads 0.013, −0.002 and 0.004 under the other three laws, and containment reads exactly one under all four. A scale whose ends were empirical would have moved with them.
What a rule doing both searches has to charge
The reason any of this is more than bookkeeping is that a charge is what a rule levies before it declares anything, and two searches charged wrongly produce a test with the wrong size.
The field that measured the window-and-break pair shows what that costs in decisions: a chi-square point on the coefficients a split adds declares a break on 73.6% of samples that have none, the break search’s own 95% point applied inside a rule that also chooses its window fires on none at all, and only a charge calibrated on the statistic the joint rule actually reads holds its size. Three thresholds, two of them not tests.
The ladder here says which of those failures to expect from a given pair. A pair near one is a pair where the second search’s own charge is nearly all wasted — levying it makes the rule far too conservative, which is the middle threshold’s failure. A pair near zero is a pair where the two charges genuinely add, and separate charging is correct. The size of the mistake is the overlap, and it was previously a single number measured on one pair.
What the ladder does not supply is the critical value. Overlap is a statement about means, and a threshold is a statement about a tail; the two move together but not by the same amount, which the earlier field’s own numbers show — its means differ by 19.30 and its 95% points by more.
Where else a search contains another
The relation is commoner than it looks once it is named, and three places in this collection have it.
A second break inside a first. The field that priced a second break searches for another split on a profile that is already flat, and a two-break model contains every one-break model. Its excess is the same quantity measured here, one level down and without the scale.
A wider band inside a narrower one. Band widths are nested — a band at lags is a band at with the last entry held at zero — which is exactly why the maximised likelihood cannot fall as the width grows. A search over the narrower family would read one against a search over the wider.
And a balancing rule’s basis inside another’s span. The field that is about the geometry establishes that a rule cannot tell one basis from another with the same span, which is containment in both directions at once and would read one either way round.
What is uncommon is the pair rather than the relation: two searches over one sample where neither contains the other and both read the same thing. That is where the interesting readings are, and the two constructed ends exist so that they can be placed.
What this makes available
Two things, and the second is what the next essay uses.
An overlap is now a number with a meaning rather than a ranking. 0.125 says the second step search adds seven eighths of what it would have found alone; 1.000 says it adds nothing; 0.004 says it adds all of it. Those are statements about a rule that runs both searches and has to charge for them, and they can be compared across pairs that share no construction.
And the scale has room below zero, which the construction did not anticipate and which the control’s stability rules out as an artefact. A pair reading below zero is a pair whose joint search reaches configurations neither half of it contains — the two searches are not merely independent, they are complementary, and a rule charging them separately under-charges. That is what a break paired with an independent column does, at −0.306, and it is the finding that says the earlier field’s sign does not transport.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two searches, one sample — both name degrees of freedom, likelihood ratio, selection effect, specification search, structural break, supremum statistic
- A charge that depends on the rule — both name likelihood ratio, selection effect, specification search, structural break, supremum statistic
- What a search costs in parameters — both name change point, degrees of freedom, likelihood ratio, nested models, supremum statistic
- A criterion is a prediction of the hold-out — both name nested models, residual, selection effect, specification search
- Choosing whether to break — both name change point, selection effect, specification search, structural break
- The arcsine that closes it, and the error that was overstated — both name basis functions, discreteness, orthogonality, variance explained
Named objects
A flat tag is an object no other essay names yet.
Basis functionsChange pointDegrees of freedomDiscretenessLeast squaresLikelihood ratioNested modelsOrthogonalityResidualSelection effectSpecification searchStructural breakSupremum statisticVariance explained