Two searches over different features

Two searches that share nothing

Two searches over independent columns remove shares of the residual sum that add exactly. On the scale a chi-square point is quoted on they look super-additive by a fifth of a unit, and none of it is overlap.

Worth reading first: A design is a number · A break that was looked for.

The field that ran two searches on one sample reports that charging them separately over-charges by 19.30 units of likelihood, on 99.7% of draws, at twenty-eight standard errors. It also says, in its own last paragraph, what that number does not establish. Both searches read the same residual series. A break search under correlated errors finds mostly correlation, and a whitening chosen from the same sample has taken the correlation already, so the pair is one feature seen from two directions and the shortfall is theirs rather than two searches’.

The obvious repair is to measure a pair that shares nothing and see whether it is additive. That turns out to need a step before it, because the scale the 19.30 is measured on does not put a pair sharing nothing at zero.

What the log scale does before any overlap

Take two dictionaries of six independent standard normal columns each, drawn from their own stream so that neither the sample nor the response can move them. Search each dictionary for the column that most improves the fit. The two searches share nothing: the columns are uncorrelated with the response at 0.0025 over six hundred pairs, and uncorrelated with each other at −0.0010.

The same five pairs, uncalibrated. What a rule charging the two searches separately over-charges, in log-likelihood units, which is the scale the earlier field reports and the scale a chi-square point is quoted on. The pair that reads one residual series twice comes to 20.22 units, which is the earlier field's measurement. The trouble is at the other end: two searches constructed to share nothing read -0.08, and a break paired with an independent column reads -2.50. Neither is an overlap. A gain is a log of a ratio of residual sums, so two searches that each remove a tenth of the residual sum remove a fifth of it together and the second one's log gain is larger than it was alone — a systematic negative that has nothing to do with what the searches share. Written on the share of the residual sum removed rather than on its log, the same five pairs sort themselves out.
Fig. 1 The five pairs on the log-likelihood scale, which is the scale a chi-square point is quoted on and the scale the earlier field reports. The control is not at zero.

On the log-likelihood scale the control comes to −0.08 units and a break search paired with an independent column comes to −2.50. Both negative, which on that scale means the two searches together manufacture more than the sum of what each manufactures alone.

Nothing in either construction can produce overlap. What produces the negative is arithmetic.

A gain is nlog(rss0/rss)n\log(\text{rss}_0/\text{rss}) — a log of a ratio of residual sums. Two searches that each remove a tenth of the residual sum remove a fifth of it together, and

log1115>log11110+log11110\log\frac{1}{1-\tfrac{1}{5}} > \log\frac{1}{1-\tfrac{1}{10}} + \log\frac{1}{1-\tfrac{1}{10}}

because the logarithm is convex in what has been removed. The second search’s log gain is larger after the first one has run, precisely because there is less left. That is a property of the units, it applies to every pair whatever they share, and it has to be removed before anything called an overlap is read off.

The scale that has a zero

Invert the units. A gain of GG corresponds to a share of the residual sum

δ=1eG/n\delta = 1 - e^{-G/n}

and for two searches whose directions are orthogonal the shares add exactly: removing the projection onto one direction does not change the projection onto another. Two independent columns are orthogonal in expectation by construction, so a pair of searches over disjoint dictionaries has to read zero on this scale or the construction is wrong.

What each search removes, and what the two remove together. For each pair, the share of the residual sum the first search removes, the share the second removes, the two added, and the share the joint search actually removes. Additivity is the upper mark; the joint reading is the dot. Two searches over disjoint sets of independent columns remove 0.0269 and 0.0261 and jointly 0.0529, against 0.0530 added — a gap of 0.000116, which is 0.8 standard errors from zero. A break search and a step column remove 0.2510 and 0.1387 and jointly 0.2510, which is the first of them exactly: the step adds nothing, on every draw, because a break shifts every coefficient after a row and a centred step column is one of the directions it can move in.
Fig. 2 For each pair, the share each search removes, the two added, and the share the joint search actually removes. The gap between the mark and the dot is the overlap.

Over twelve hundred draws, the two dictionary searches remove 0.0269 and 0.0261 of the residual sum and jointly remove 0.0529, against 0.0530 added. The difference is 0.00012 ± 0.00014, at 0.8 standard errors from zero.

That is the zero. Divided through by the smaller of the two shares, it reads 0.004 ± 0.005 on the overlap scale — where one means the second search adds nothing the first has not already found.

The five searches, and what each one reads

The scale needs pairs to be put on it, and the pairs are made out of five searches over one sample. It is worth listing them by what they look at rather than by what they are called, because the ladder’s ordering is a statement about those descriptions.

A break in the regression searches over when: the row after which every coefficient shifts. It is the search the earlier fields priced and it has seventy-two admissible positions on a hundred and twenty rows.

A dictionary of independent columns searches over which column: six standard normal columns, drawn from a stream of their own and correlated with nothing. Two such dictionaries, disjoint, give two searches that share nothing by construction.

A dictionary of step columns searches over when, coarsely: six centred indicators for the rows past a cut point. Two such dictionaries with interleaved cut points give two searches over the same feature at places a couple of rows apart.

And a whitening window searches over the dependence: the width of the band the errors are whitened through, charged its own dimension because the likelihood rises about a unit a lag whatever is in the data.

The step columns exist for a reason worth stating before the numbers: a break shifts every coefficient after a row, so a centred step column is one of the directions a break search may move in, and a break search paired with a step dictionary is therefore a pair where one search contains the other. That fixes the top of the scale by construction rather than by whichever pair happened to be largest.

Which is not the zero of two searches with nothing to find

A control that reads zero because both halves of it are empty would calibrate nothing, so the same measurement is run at four dictionary sizes.

The zero is a property of the searches, not of their size. The control at four dictionary sizes, over 600 draws apiece: the two shares added, minus the share the joint search removes, with two standard errors either side. Every reading is inside three standard errors of zero — -0.85, 2.40, 0.58, 1.68 — while what each search finds grows from 0.0142 of the residual sum at 2 columns to 0.0328 at 10. So the zero is not the zero of two searches with nothing to find. It is the zero of two searches whose findings occupy directions that do not overlap, which is what a scale's origin has to mean if the numbers above it are to be read as shares of one search that the other has already taken.
Fig. 3 The control at four dictionary sizes, with two standard errors either side. Flat at zero while what each search finds nearly triples.

At two columns each search removes 0.0142 of the residual sum; at ten it removes 0.0328, more than twice as much. The excess over the four sizes reads −0.85, 2.40, 0.58 and 1.68 standard errors — scattered around zero, with no trend in it.

So the zero survives the searches getting larger. It is the zero of two searches whose findings occupy directions that do not overlap, which is what an origin has to mean if the numbers above it are to be read as the share of one search that the other has already taken.

What a shared response does and does not do

There is a second candidate explanation for a control that is not additive, and it has to be dealt with before the zero can be trusted, because it would apply to every pair here.

Two searches on one sample are both fitting the same response. Whatever the first one explains, the second one cannot explain again — there is one residual sum and both are eating it. That sounds like a floor under any pair of searches on one sample, independent or not, and it would mean the scale’s zero was unreachable in principle.

It is not a floor, and the reason is the difference between competing for a residual sum and overlapping in it. Two orthogonal directions each remove their own projection of the response, and neither projection depends on whether the other has been removed: that is what orthogonality means. What the second search loses is not the size of its own reduction but the proportion it represents of what is left — and proportion is exactly what the logarithm is measuring and the share is not.

So the shared response produces the whole of the log scale’s −0.08 and none of the share scale’s 0.00012. The two effects have been separated by choosing units, which is the cheapest kind of separation available and is worth having precisely because the alternative — arguing about whether a floor exists — has no measurement in it.

The reading that follows is worth keeping separate from the arithmetic that produced it. Two searches on one sample compete for one residual sum and do not overlap in it, and the two are different relations: competing is what makes the log scale’s answer negative, and overlapping is what the share scale is built to measure. Conflating them is the whole of why a control was needed.

The five pairs

With both ends of the scale fixed — the second is a search that contains another, which reads exactly one — the comparison the earlier field could not make becomes available.

How much of one search the other has already foundFive pairs of searches on one sample, on a scale whose zero and one are both fixed by construction. Zero is two searches over disjoint sets of independent columns: they remove shares of the residual sum that add, at 0.8 standard errors from exactly additive, and they read 0.004. One is a break search paired with a step column it contains, which reads exactly one on every draw because the step adds nothing at all. Between them: two dictionaries of step columns cut a few rows apart read 0.125, and the pair the earlier field measured — a break and a whitening window, both reading the same residual series — reads 0.762, three quarters of the way to one search containing the other. And below zero, a break paired with a search over independent columns reads -0.306: the joint search finds configurations neither half of it contains, so charging the two separately under-charges.share nothinga break, and an independent column-0.306two dictionaries of independent columns0.004two dictionaries of step columns0.125a break, and a whitening window0.762a break, and a step column it contains1.0001200 draws, AR(1) at 0.8a scale fixed at both ends
Fig. 4 Five pairs on the calibrated scale. The slider changes the law the errors follow; the ordering does not change with it.

0.004 for two dictionaries of independent columns. 0.125 for two dictionaries of step columns cut a few rows apart. 0.762 for a break search and a whitening window, which is the pair the earlier field measured. 1.000 for a break search and a step column it contains. And −0.306 for a break search paired with a search over independent columns, which is the one that does not fit the story at all and is the subject of the third essay here.

The shortfall is a property of the pair, spanning more than the whole scale, and the ordering is the one the searches’ own descriptions predict: nothing shared, a shared feature at different places, a shared series, and containment.

The gap between the two ends is 1.31 — more than the whole nominal scale, because the bottom rung is below the origin. A quantity presented as a fraction that runs from a third below zero to one is not behaving like a fraction, and that is the honest description of it: the numerator is a difference of shares and nothing forbids it from having either sign.

What the calibration cost, and what it did not

Two things are worth being explicit about, because the scale is a construction and constructions can be tuned until they say what was wanted.

The zero was not fitted. Nothing in the definition of δ was chosen by looking at the control. The inversion δ=1eG/n\delta = 1 - e^{-G/n} follows from the definition of the gain and nothing else; the prediction that orthogonal searches add on that scale is the ordinary geometry of least squares; and the control’s reading is a test of both, which it could have failed and did not.

The one was not fitted either. A break in the regression shifts every coefficient after a row, so a centred step column is one of the directions a break search can move in, and a search over step columns can therefore find nothing a break search has not. That is a containment argument made before the measurement, and the measurement returns exactly one — not approximately, and on every draw, because the joint supremum equals the break search’s supremum to the last bit.

What the scale does not fix is what happens outside [0, 1]. A reading above one would mean the joint search found less than the better of its two halves, which cannot happen: the joint search maximises over a set containing both slices. A reading below zero can happen and does, and it means the joint search found configurations neither slice contains — which is a real phenomenon rather than a defect in the units, and is the finding the third essay is about.

What the window pair is doing on the same axis

One rung needs a qualification, and it is the rung the field was written for.

The window search’s gain is not a reduction of a residual sum. Its objective carries a determinant and a charge for its own width, so what it removes is a share of an effective residual sum, rsse(logΩ^+2L)/n\text{rss}\cdot e^{(\log|\hat\Omega| + 2L)/n}, rather than of the residual sum itself. The inversion to a share still applies — the objective is still a log of a ratio — but the orthogonality argument that makes two column searches add exactly does not.

So the window pair’s 0.762 is measured on the same axis and is not calibrated by the same construction. What supports reading it as an overlap is that it sits between two rungs that are calibrated, that it moves in the direction the laws predict, and that the earlier field’s own reading of it — what the break search adds after the window search has run — is the same statement in the same units.

Naming the limit is the point. Three of the five rungs stand on a construction; the fourth stands on where it falls between them.

None of that is an argument for reporting gains on the share scale generally — a log-likelihood is the right object for a likelihood ratio and this site quotes chi-square points constantly. It is an argument for converting before subtracting, which costs one exponential and is the difference between a control at −0.08 and a control at 0.0001.

The rule of thumb is right at the control and wrong at the pair

The convexity is priced above as about δ_A·δ_B·n, and it reproduces the control exactly. With shares of 0.0269 and 0.0261 on a hundred and twenty rows that is 120 × 0.000702 = 0.084 units, against the counted −0.08 on the log scale. A closed expression with nothing fitted to it, landing on the measured artefact to the digit.

Applied to the break-and-column pair it does not work, and the failure is instructive. That pair’s break search removes about a quarter of the residual sum, so the leading term gives 120 × 0.25 × 0.0265 = 0.79 against a counted −2.50 — short by a factor of three.

Doing the conversion exactly rather than to leading order closes it. With the joint share at 0.283, the log-scale super-additivity is −120[log(1 − 0.2498) + log(1 − 0.0265) − log(1 − 0.283)] = −2.22, which is 89% of the counted −2.50.

So the essay’s claim that both readings are the artefact stands, and the rule of thumb it is stated with does not. δ_A·δ_B·n is a small-share approximation, and a break search removing a quarter of the residual sum is not in the small-share regime. The practical form: use the rule of thumb to decide whether the correction matters, and do the exponentials when it does — the threshold is around a share of a tenth, above which the leading term understates by more than a fifth.

What the four control readings say together

The dictionary-size sweep is read as scattered around zero with no trend, and the four standardised excesses — −0.85, 2.40, 0.58 and 1.68 — support the first half more clearly than the second.

One reading above two among four is unremarkable on its own: the largest of four standardised departures exceeds 2.40 about six per cent of the time when all four are correct, so nothing there calls for an explanation. What is worth noticing is that three of the four are positive and their combined evidence is 1.9 standard errors in the same direction.

That is not a finding and it is not nothing. A control whose four readings combine to 1.9σ above zero is consistent with the zero it is asserted to have and is also consistent with a small positive floor of a few ten-thousandths — well below every rung of the ladder, and above the resolution of twelve hundred draws only in aggregate.

The honest statement is that the origin is established to within about a thousandth of a share and not to zero, which is enough for a scale whose next rung is at 0.125 and would not be enough if anything on it were being read at the third decimal place. Reporting the four readings rather than their average is what makes that visible; an average of the four would have shown 1.9σ and looked like a defect, and any one of them alone would have shown either nothing or 2.4σ and looked like whichever it was.

Two things the ladder is not

It is not a decomposition of a charge. Nothing here says how much a rule running both searches should levy; it says how much less the two together find than the sum of what each finds alone. Turning that into a critical value needs the distribution of the joint statistic rather than its mean, and the field that did that for one pair shows the two are not interchangeable — its 95% points are further apart than its means.

And it is not a measure of how similar two searches are. The step dictionaries are cut a few rows apart and read 0.125; a break and a step column read one. Both pairs are searching over the same feature, and the difference between them is that one search’s set of directions contains the other’s rather than merely resembling it. Containment and resemblance are different relations, and the scale measures the first.

Where else the units matter

The correction is not specific to searches, and the size of it is worth carrying.

Any two effects reported as log-likelihood gains against a common base are super-additive by the same convexity, and the size is about δAδBn\delta_A\delta_B \cdot n in units — a tenth and a tenth on a hundred and twenty rows is 1.2 units, which is more than what a coefficient is charged. Two model comparisons each worth five units of likelihood, reported separately and added, over-state the joint improvement by about a fifth of a unit at this sample size and by more as the effects grow.

That is small enough to ignore in most places and it is exactly the size of the thing this field is measuring. The control’s −0.08 and the break-and-column pair’s −2.50 are both entirely this artefact, and reading either as an overlap would have reported a phenomenon that is a property of the logarithm.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Chi squaredClosed formDegrees of freedomLeast squaresLikelihood ratioMonte CarloMultiple comparisonsOrthogonalityOverfittingResidualSelection effectSpecification searchStructural breakSupremum statisticVariance explained