Two searches over different features

Three quarters of the way to one search

The pair that started this reads 0.762 on a scale whose one is containment. And the pair that shares nothing but its response reads −0.306, so the sign the earlier field found does not transport at all.

Worth reading first: A design is a number · A break that was looked for.

The field this one exists to place measures a break search and a whitening window on one sample and reports that charging them separately over-charges by 19.30 units. It also names the objection: both searches read the same residual series, so the number is that pair’s rather than two searches’. The deferral it wrote down was specific — the sign transports and the size does not.

Half of that survives. The size is now measurable on a scale with a fixed zero and a fixed one, and the pair comes to 0.762: three quarters of the way to one search containing the other. The other half does not survive at all.

Where the pair sits

How much of one search the other has already foundFive pairs of searches on one sample, on a scale whose zero and one are both fixed by construction. Zero is two searches over disjoint sets of independent columns: they remove shares of the residual sum that add, at 0.8 standard errors from exactly additive, and they read 0.004. One is a break search paired with a step column it contains, which reads exactly one on every draw because the step adds nothing at all. Between them: two dictionaries of step columns cut a few rows apart read 0.125, and the pair the earlier field measured — a break and a whitening window, both reading the same residual series — reads 0.762, three quarters of the way to one search containing the other. And below zero, a break paired with a search over independent columns reads -0.306: the joint search finds configurations neither half of it contains, so charging the two separately under-charges.share nothinga break, and an independent column-0.306two dictionaries of independent columns0.004two dictionaries of step columns0.125a break, and a whitening window0.762a break, and a step column it contains1.0001200 draws, AR(1) at 0.8a scale fixed at both ends
Fig. 1 The five pairs on the calibrated scale. The slider changes the law the errors follow.

Over twelve hundred draws under a first-order autoregression, the break search removes 0.2510 of the residual sum, the window search removes 0.5033 of an effective one, and together they remove 0.5630 against 0.7543 added. The shortfall is 0.19134 ± 0.00256, at 74.8 standard errors, and it is positive on every draw.

Divided by the smaller of the two shares, that is 0.762 ± 0.010.

The reading is direct. The break search has, on average, already found three quarters of what the window search would find on its own — or, taken the other way round, a whitening chosen from the same residuals has taken three quarters of what a break search would have found. That is what the earlier field says in words when it reports the break search adding 15.40 units after the window search where alone it is 34.70, and it is now a number that can be compared with other pairs.

Three quarters is a great deal. It sits between two dictionaries of step columns cut a few rows apart, at 0.125, and a search that contains the other, at exactly 1.000. Two searches with no obvious feature in common are nearer to containment than two searches over the same feature at different places.

What three quarters means for a rule

It is worth converting once, because the ratio is the portable form and a rule needs the absolute one.

A rule that runs both searches and charges each its own critical value is levying a charge for the window search’s full 0.5033 of the effective residual sum. What the window search actually adds, once the break search has run, is 0.3120. The difference is the 0.19134 above, and it is what the rule over-charges by.

In the log-likelihood units a chi-square point is quoted in, the same comparison is 86.36 for the window search alone against 66.14 added afterwards, and the naive sum over-charges by 20.22 units. That is the earlier field’s 19.30, measured on a longer run and with the break search over the regression rather than over the errors — the same quantity, and the small difference is the sweep rather than the construction, which is checked draw by draw against its own profile.

The ratio is what transports and the units are what a threshold is in. Both are needed, and the ratio is the one that can be compared with a pair that manufactures a hundredth as much.

Which share the three quarters is a share of

The scale divides the overlap by the smaller of the two searches’ solo contributions, and since the two are a factor of two apart, it is worth saying what the other denominator gives.

The overlap is 0.19134. The break search removes 0.2510 alone and the window search removes 0.5033. So the overlap is 76.2% of the break search’s own contribution and 38.0% of the window search’s.

Both are true and they say different things. Three quarters of what the break search finds is also found by the window search; a bit over a third of what the window search finds is also found by the break search. The pair is lopsided, and the ratio quoted is the larger half of the lopsidedness.

Dividing by the smaller is not arbitrary, though, and the reason is the top of the scale rather than convenience. Containment means the contained search adds nothing once the other has run, and the contained search is necessarily the smaller of the two. So the smaller share is the only denominator under which exact containment reads exactly one, which is what makes the scale a scale rather than a ranking.

What each search removes, and what the two remove together. For each pair, the share of the residual sum the first search removes, the share the second removes, the two added, and the share the joint search actually removes. Additivity is the upper mark; the joint reading is the dot. Two searches over disjoint sets of independent columns remove 0.0269 and 0.0261 and jointly 0.0529, against 0.0530 added — a gap of 0.000116, which is 0.8 standard errors from zero. A break search and a step column remove 0.2510 and 0.1387 and jointly 0.2510, which is the first of them exactly: the step adds nothing, on every draw, because a break shifts every coefficient after a row and a centred step column is one of the directions it can move in.
Fig. 2 For each pair, the share of the residual sum the first search removes, the share the second removes, the two added, and the share the joint search actually removes. Two searches over disjoint sets of independent columns remove 0.0269 and 0.0261 and jointly 0.0529 against 0.0530 added — a gap of 0.8 standard errors.

Two cautions about the ratio’s standard error

The shortfall carries ±0.00256 and the ratio carries ±0.010, and the second is the first divided by the denominator with the denominator treated as fixed.

It is not fixed. The break search’s 0.2510 is measured on the same twelve hundred draws, so the honest standard error on the ratio depends on how the numerator and denominator move together across draws. If they move together — and a draw on which the break search finds more is a draw on which there is more for the two to share — the correlation makes the ratio steadier than the naive calculation says, not noisier. So ±0.010 is a conservative bound in the likely case and would be an underestimate only if the two moved oppositely.

Either way it does not threaten the reading. Seventy-four and a half standard errors on the shortfall, positive on every one of twelve hundred draws, is not a result any treatment of the denominator disturbs.

The comparison with the earlier field’s likelihood units is worth one line of caution as well. There the break search reads 34.70 alone and 15.40 after, an overlap of 19.30 — 55.6% of the solo figure, where the residual-sum scale here reads 76.2%. The two are shares of different quantities, and a ratio of log-likelihood rises is not a ratio of residual-sum shares. The direction transports and the fraction does not, which is the same warning the field’s own title is about, applied to the scale rather than to the sign.

And the sign does not transport

The pair that was supposed to be the easy case is the one that does not behave.

Take the break search and pair it with a search over six independent standard normal columns — a search whose candidates are uncorrelated with the response at 0.0025 and with everything else. Nothing is shared except that both are fitting the same y.

It reads −0.306 ± 0.021.

Negative, at fifteen standard errors, and it is not the log-scale artefact: that has already been removed, and the control on the same scale with the same construction reads 0.004. The joint search finds more than the two searches find separately.

The mechanism is that a joint search is not the union of two searches. The break search maximises over seventy-two rows with no column; the column search maximises over six columns with no break; the joint search maximises over all four hundred and thirty-two combinations, and most of those are configurations neither slice contains. A column that is mediocre against the whole sample can be excellent against a sample split at row 52, and the joint search finds it.

So a rule charging the two searches separately under-charges. That is the opposite of the earlier field’s conclusion, on a pair chosen to be the clean case for it, and it means the sentence two searches over one sample find some of the same luck twice is a statement about a pair rather than about two searches.

Why it is not the dictionary finding real signal

The obvious objection is that the columns must be picking up something after all — that a break at row 52 leaves structure a normal column can fit, so the joint search is finding signal rather than luck. That would make the super-additivity uninteresting.

It is ruled out by the design rather than by argument. The columns are drawn from a stream of their own, before the response exists, and their correlation with it is 0.0025 over six hundred column-draw pairs, which is inside its own standard error. There is nothing in them to find under any split of the sample. What the joint search finds is the best of four hundred and thirty-two ways of fitting noise, against the best of seventy-two and the best of six taken separately — and the best of a product set exceeds the sum of the two marginal bests whenever the two dimensions are not separable, which they are not.

The same argument says why the control is exempt: two orthogonal columns fitted together reduce the residual sum by the sum of their separate reductions exactly, so the joint search over thirty-six pairs is the pair of the two marginal winners and there is nothing extra to find.

The zero is a property of the searches, not of their size. The control at four dictionary sizes, over 600 draws apiece: the two shares added, minus the share the joint search removes, with two standard errors either side. Every reading is inside three standard errors of zero — -0.85, 2.40, 0.58, 1.68 — while what each search finds grows from 0.0142 of the residual sum at 2 columns to 0.0328 at 10. So the zero is not the zero of two searches with nothing to find. It is the zero of two searches whose findings occupy directions that do not overlap, which is what a scale's origin has to mean if the numbers above it are to be read as shares of one search that the other has already taken.
Fig. 3 The control at four dictionary sizes over 600 draws apiece, with two standard errors either side. Every reading is inside three standard errors of zero while what each search finds grows from 0.0142 of the residual sum at two columns to 0.0328 at ten: the zero is not the zero of two searches with nothing to find.

Two searches can be complementary, and the word is exact

It is worth separating the phenomenon from its explanation, because the explanation makes it sound like a technicality and it is not.

Two searches overlap when the second has less left to find. Two searches are complementary when the joint search space contains configurations neither of them can reach, and reaching them finds more than the two separately. The first depends on what the searches read; the second depends on whether the searches interact — on whether the best value of one depends on the value of the other.

A break location and a column choice interact strongly: a break changes which rows the column has to fit. A break location and a whitening width interact too, but the interaction is swamped by how much they share, since both are reading the same correlation.

Both effects are present in every pair. The number reported is the net of them, which is why the scale runs below zero and why a pair reading zero is not necessarily a pair with neither effect — it may be a pair with both, cancelling. The control’s zero is a genuine absence of both only because two searches over independent columns have nothing to interact through: the joint fit of two orthogonal columns is the sum of the two single fits, exactly.

Which of the two effects the pair is mostly made of

The window pair reads high, and there are two ways to be high on this scale: sharing a great deal, or interacting very little. It is worth arguing which, because the two have different consequences for a rule.

A break location and a whitening width interact — a break changes the residuals the whitening is estimated from, and a whitening changes the profile the break is found on. The earlier field measures the interaction directly: the break point found with the window searched too is not the break point found at a fixed window on the same sample, and the two disagree often enough to matter. So the pair is not one where the interaction is absent.

It is one where the sharing dominates. A break search under correlated errors is largely finding the correlation: a run of positive errors looks like a level shift, and the longer the runs the more of them there are to find. A whitening removes exactly those runs. So the two searches are looking at the same feature of the sample through different instruments, and the interaction — real, and worth something — is a small correction on a large overlap.

That is why the pair is nearer containment than the step dictionaries are, despite having nothing in common in its description. What matters is what the searches find, not what they are searching over, and both of these find persistence.

Under the four laws

The ladder is the same ladder under every law. Each pair's overlap under each of the four laws, over 400 draws apiece. The scale is fixed at both ends by construction: two searches over disjoint sets of independent columns read -0.001, 0.013, -0.002, 0.004, and a break search paired with a step column it already contains reads exactly one under every law. Between them the pair that reads one residual series twice runs 0.763, 0.805, 0.577, 0.752 — lowest under long memory, where a whitening has most to do and the break search has least left to find that the whitening has not taken. The rung that moves most is the pair of step dictionaries, from -0.148 under a moving average to 0.419 under a break; and the pair that is negative is negative under all four.
Fig. 4 The five pairs under each of the four laws. Both ends hold; the pair this essay is about moves between 0.577 and 0.805.

The window pair reads 0.763 under a first-order autoregression, 0.805 under a five-period moving average, 0.577 under long memory and 0.752 under a break in the persistence.

Long memory is the low reading and the reason is legible. There the errors are still 0.57 correlated at the twentieth lag, so a whitening has a very great deal to remove — its share is the largest of the four — and the break search, which finds correlation rather than breaks, cannot have taken most of it. The window’s own gain grows faster than what the break search overlaps with it, so the ratio falls.

The complementary pair is negative under all four: −0.268, −0.366, −0.345 and −0.232. So the finding that the sign does not transport is not one law’s.

How much of the collection this reaches

Almost everything in the second half of this collection runs two searches without saying so, and the ladder says which of them to worry about.

A criterion choosing among candidates, through an estimated whitening, is a selection search and a window search on one sample — the pair a whole field measures the displacement of without charging for either. Both read the residual series, so the window pair’s reading is the relevant one and the charges are far from additive.

A criterion choosing an autoregressive order, through the same whitening, is the same shape with the order in place of the candidate set, and the field that compared the two tuning parameters shows they cost about the same.

A rule that estimates a break and then a variance is the complementary shape: the second search reads what the first has left, and by the reading here it would find more than it would have found alone, not less.

None of the three is measured with the instrument built here. What the ladder gives is a prediction for each, and predictions are what a scale is for.

The number worth carrying

If one sentence survives from all three essays of this field it is this: how much two searches over one sample overlap runs from a third below zero to one, and where a given pair lands is not predictable from what the searches are searching over.

A break and a window have nothing in common in their descriptions and read 0.762. Two dictionaries of step columns are searching one feature at nearby places and read 0.125. A break and an independent column share nothing at all and read −0.306, on the wrong side of the origin. The only reliable predictor in the table is containment, which reads one and is a relation nobody needs a measurement for.

So the practical form is not a rule of thumb. It is that the quantity is cheap: three suprema on the same draws, an exponential, and a subtraction. Anything that runs two searches can measure its own pair in the time it takes to run the searches three times instead of twice.

What this does to the earlier field’s advice

The earlier field ends with a recommendation about thresholds, and the recommendation is unchanged. Its strongest reading is what three thresholds do to one decision: a chi-square point on the coefficients a split adds declares a break on 73.6% of samples with none, the break search’s own 95% point applied inside a rule that also chooses a window fires on none, and only a charge calibrated on the statistic the joint rule reads holds its size. None of that depends on the overlap being 19.30 rather than something else.

What changes is the generalisation. A practitioner told that two searches over one sample over-charge, and applying that to a break search and a variable-selection search, will levy too little rather than too much — and too little is the direction that produces a test firing more often than its stated size. The failure the earlier field’s finding protects against and the failure its generalisation produces are opposite failures, which is a bad property for a rule of thumb.

The usable version is narrower and is what the ladder supports. Two searches reading the same series overlap a great deal; two searches reading different features of one sample may go either way, and which way depends on whether they interact more than they share. Neither can be assumed, and both can be measured on one sample’s worth of simulation — three suprema and an exponential.

What is left open

A pair that is complementary and overlapping at once, separated. The net is what is measured; nothing here splits it. A design that could is a pair of searches whose interaction is switched off by construction — the break search restricted to a fixed row, say — and running the same comparison with and without it. That is one extra sweep and it is the obvious next thing.

A threshold rather than a mean. Everything here is an average shortfall. What a rule needs is the 95% point of the joint statistic, and the field that computed both for one pair finds the two moving together and not by the same amount. An overlap of 0.762 does not say what critical value to use.

And a pair where the second search is not a search. All five rungs are suprema over finite lists. A rule that estimates a nuisance rather than searching for it — a plug-in whitening at an automatic window, say — manufactures nothing and would read at the zero for a different reason. Whether the scale is meaningful there is not established.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A charge that depends on the rule — both name critical value, likelihood ratio, selection effect, specification search, structural break, supremum statistic, whitening
  • A second break on a flat profile — both name critical value, degrees of freedom, likelihood ratio, selection effect, specification search, structural break, supremum statistic
  • A split that depends on the order — both name likelihood ratio, overfitting, selection effect, specification search, structural break, supremum statistic
  • Two effects in one number — both name likelihood ratio, overfitting, selection effect, specification search, structural break, supremum statistic
  • What a zero is made of — both name likelihood ratio, orthogonality, overfitting, selection effect, specification search, supremum statistic
  • How often it matters — both name long memory, long-run variance, persistence, selection effect, whitening

Named objects

A flat tag is an object no other essay names yet.

Chi squaredCritical valueDegrees of freedomLikelihood ratioLong memoryLong-run varianceMultiple comparisonsOrthogonalityOverfittingPersistenceSelection effectSpecification searchStructural breakSupremum statisticWhitening