Overlap and complementarity, separated

What a zero is made of

Two disjoint dictionaries of independent columns read an excess of 0.000116 and are made of an overlap of 0.000583 and an interaction of 0.000467. The control the whole scale is anchored on reads zero because two effects cancel.

Worth reading first: A break that was looked for · A design is a number.

A scale needs a zero, and the field that built this one chose its zero carefully: two disjoint dictionaries of independent standard normal columns, drawn from a stream of their own so that neither the sample nor the columns can move the other. Two searches over them share nothing by construction, so their excess should be nothing, and it is — 0.000116 ± 0.000137 over twelve hundred draws, under one standard error from exactly additive.

That is a good control and the number is right. What it is not is a statement that there is nothing between the two searches.

The zero, split

The same twelve hundred draws, with the sequential supremum computed as well:

overlap 0.000583, interaction 0.000467, excess 0.000116.

Each component is several times the number they cancel into. The pair that anchors the whole ladder at zero is a pair in which the second search loses about half a thousandth of the residual sum by having the first run first, and the joint search gains about half a thousandth by being allowed to move the first off its answer.

They cancel to within a standard error, and the cancellation is what the ladder’s zero is.

Both halves grow; the difference does not. The control pair's two components and their difference, against how much each of its two searches can find, over 1200 draws at each dictionary size. Two disjoint sets of independent columns are additive at every size — the excess stays inside a standard error or two of zero throughout — and it is not because there is nothing there. The overlap grows from 0.000112 at two columns to 0.000870 at ten, a factor of 7.76, and the interaction grows with it, staying within a factor of two of the overlap at every size. Two searches competing for one residual sum share ground and find configurations neither has alone, in almost equal measure, and their difference is what the earlier field's scale calls zero.
Fig. 1 The control’s two components and their difference, against how many columns each dictionary offers. The slider changes how many draws each size takes.

Why either component is non-zero at all

Two searches over independent columns should find nothing in common, so it is worth being clear about what they do find.

They compete for one residual sum. Both searches are looking for the column that most reduces the same sum of squares. Whichever runs first takes the largest reduction available on that draw; whatever the second finds is measured against a smaller sum. That is a real loss to the second search and it has nothing to do with the columns being related — it is a fact about there being one sample.

And moving the first off its answer buys something. The joint search over both dictionaries at once evaluates every pair of columns, and on some draws the best pair is not the best column from each. Two columns that are mildly correlated with each other in this sample — which independent columns are, at a hundred and twenty rows — can beat the pairing of two individually better ones.

Neither effect is large. Both are exactly the size a competition for a shared residual sum should be, and they very nearly cancel because they are two readings of the same competition from opposite ends.

The joint search moves off the sequential answer on 13% of draws at six columns each, which is the lowest share of any pair on the ladder and is what a pair of unrelated searches should look like. On the other four rungs it is between 29% and 78%. So the control is not a pair the joint search treats specially; it is a pair on which it has least to gain, and the little it gains is matched by the little the second search loses.

How large “several times” is

The word carries the essay, so it is worth putting a number on it.

The ratio of the smaller component to the excess is 46.19 for the control at three hundred draws and 4.03 at twelve hundred. Those look like different findings and they are one: the excess is a noisy quantity near zero, so dividing by it produces a ratio that moves with the sweep length while the numerator does not. At three hundred draws the excess happened to land at 0.000011; at twelve hundred it lands at 0.000116; the overlap is 0.000514 and 0.000583.

The stable half is the components and the unstable half is the ratio, which is exactly what should happen when one of two quantities is genuinely zero.

So the claim is not that the components are precisely forty-six times the excess. It is that the components are a measurable size and the excess is not distinguishable from zero, and the ratio between them is bounded below by whatever the sweep length makes it. The right way to read the control is as two numbers near half a thousandth and one number at zero, rather than as three numbers on one scale.

What happens when there is more to compete for

The strongest evidence that this is the mechanism is that both components move with how much there is to find, and the difference does not.

Grow each dictionary from two columns to ten:

columns each overlap interaction excess
2 0.000112 0.000077 0.000035 ± 0.000084
4 0.000362 0.000219 0.000142 ± 0.000119
6 0.000583 0.000467 0.000116 ± 0.000137
10 0.000870 0.000800 0.000069 ± 0.000159

The overlap grows by a factor of 7.76 and the interaction grows with it, staying between 0.61 and 0.92 of the overlap at every size. The excess stays inside a standard error of zero at all four.

That is the shape a cancellation has and a coincidence does not. Two quantities that happened to be equal at one dictionary size would not stay equal across a fourfold change in how much each search can find; two quantities that are two readings of the same competition would.

What each search finds alone grows over the same range — δA runs 0.01469, 0.02165, 0.02691, 0.03313 — so the competition really is getting stronger, and the two components track it.

The zero is a property of the searches, not of their size. The control at four dictionary sizes, over 600 draws apiece: the two shares added, minus the share the joint search removes, with two standard errors either side. Every reading is inside three standard errors of zero — -0.85, 2.40, 0.58, 1.68 — while what each search finds grows from 0.0142 of the residual sum at 2 columns to 0.0328 at 10. So the zero is not the zero of two searches with nothing to find. It is the zero of two searches whose findings occupy directions that do not overlap, which is what a scale's origin has to mean if the numbers above it are to be read as shares of one search that the other has already taken.
Fig. 2 The same sweep as the earlier field reads it, in the essay that asked whether the zero moves — where the excess column barely does.

Both components are second order in what each search finds

The two components track how much each search can find, and dividing by the natural scale says how closely.

Each dictionary’s own share δA runs 0.01469, 0.02165, 0.02691 and 0.03313 across the sweep, and the two dictionaries are symmetric, so δA·δB is δA². Dividing the overlap by it gives 0.519, 0.772, 0.805 and 0.793.

Three of the four sit at about 0.8, so the overlap is four-fifths of the product of the two searches’ own shares. The interaction is between 0.61 and 0.92 of the overlap, so it is second order in δ too.

That is the structural reason the control works. Both components are quadratic in what each search finds, and the excess is their difference — so the excess is smaller than second order, and a sweep that quadruples δA² leaves it where it was. A zero that survives because two first-order terms cancel would be fragile; a zero that survives because both terms are second order and nearly equal is not, and the ladder’s origin is the second kind.

It also gives the overlap a closed prediction. overlap ≈ 0.8·δA·δB, from quantities every pair on the ladder already reports, so the control’s component sizes can be checked against the sweep rather than only measured by it.

The cancellation tightens as the dictionaries grow

The two components are equal to within a standard error at every dictionary size, and they are not equally equal.

The interaction as a share of the overlap runs 0.688, 0.605, 0.801 and 0.920 at two, four, six and ten columns each, and the excess as a share of the overlap runs 0.313, 0.392, 0.199 and 0.079.

So the two components converge on each other as the searches get larger: the residue falls from about a third of the overlap to under a twelfth, over a range in which the overlap itself grows by a factor of nearly eight.

That is a stronger statement than the excess staying flat. A flat excess against a growing overlap is consistent with a small constant offset; a relative residue that falls from 31% to 8% says the two effects are approaching identity rather than merely being similar. The zero gets better as the control gets bigger.

And it says which direction the control should be pushed to check it further. A sweep to twenty columns each would put the components near 0.003 and the residue, on this trend, near two per cent of that — about 6 × 10⁻⁵, which is below the standard error of every excess in the table. The control cannot be broken by making it larger, which is the opposite of the usual worry about a null that is only ever measured at one size.

What the earlier field could see and could not

That field asked exactly the right question about its own control: two searches that share nothing still compete for one residual sum, so the excess cannot be exactly zero, and how far from zero should grow with how much each search can find.

Its answer was that the excess barely moves with the dictionary size, and it read that as the floor being small enough not to matter.

Both halves of that are true and the reading underneath them is not. The floor does not move because it is a difference of two things that both move, so a sweep over the dictionary size measures the difference of two growing quantities and finds it flat — which looks exactly like a mechanism that is not operating.

The excess does drift a little across the sweep, from 0.000035 to 0.000069, with a ratio of 1.99 between the ends. That is well inside the noise on either number and it is the whole of what the earlier field had to work with.

What the dictionaries are, and why they are drawn where they are

The control depends entirely on its two dictionaries being independent of each other and of the response, so the construction is worth stating.

Each dictionary is a set of standard normal columns drawn from a stream of its own, seeded separately from the stream that drew the sample. That separation is load-bearing: if the columns came from the sample’s stream, changing how many columns each dictionary holds would change the sample, and the sweep in the table above would be measuring two things at once.

They are independent of the response by construction, so a search over them finds nothing but luck. That is what makes them a control rather than a pair of searches with a small true overlap — there is no true overlap to be small.

And they are disjoint: the first dictionary’s columns and the second’s are different draws, so nothing is shared between the two searches except the sample they are both fitted to. The earlier field is explicit that this is the pair its zero is defined by, and everything above is a statement about that pair rather than about independent searches in general.

What this changes about the ladder

The practical consequence is about reading the other rungs, and it is a caution rather than a correction.

Every rung above the control is reported as an excess, and every one of them is also a difference of two components. The break-and-window pair reads 0.191122 and is made of 0.197114 and 0.005991 — nearly all overlap, so its number means what it appears to mean. The break-and-step pair reads 0.136647 and is made of 0.136647 and exactly nothing — the same.

But the break-and-column pair reads −0.006759 and is made of an overlap of −0.004395 and an interaction of 0.002364, and the two step dictionaries read 0.014075 and are made of 0.044030 and 0.029955 — where the smaller component is 2.13 times the excess.

So of the five rungs, two are what they look like, one is a zero made of two effects forty-six times larger, and two are nets in which a third of the larger component has been cancelled away.

A ladder of nets is still a ladder, and it orders the pairs correctly on the quantity it measures. What it cannot do is say what any rung is made of, and the two rungs whose numbers are smallest relative to their components are the two that carry the most cancellation.

What each rung is made of. Each pair of searches, over 300 draws, split into the two effects its excess is the difference of. The overlap is what the second search loses by having the first already run at its own answer; the interaction is what the joint search finds by moving the first off it. They subtract to the excess exactly, on every draw, because the pinned supremum cancels. Two disjoint dictionaries of independent columns read an excess of 0.000011 and are made of 0.000514 and 0.000503. A break paired with a dictionary of step columns has an interaction of exactly 0 and is all overlap. And a break paired with an independent column has an overlap of -0.004395 against an interaction of 0.002364, which is what puts its excess below zero.
Fig. 3 All five pairs split, in the previous essay.
How much a reading hides. For each pair of searches, how large its smaller component is against the excess the two subtract to, over 300 draws. A pair whose reading is a genuine absence has a small number here; a pair whose reading is two effects cancelling has a large one. Two disjoint dictionaries of independent columns — the control the earlier field's whole scale is anchored on — read 46.19 times their own excess, so its zero is a cancellation rather than an absence. The pairs further up the ladder read below one, which is what it looks like when the excess is the effect rather than the residue of two.
Fig. 4 And how large each pair’s smaller component is against the excess it cancels into — which is 46 for the control and under one for the two rungs at the top.

The interaction’s sign, which cannot be otherwise

One of the two components cannot be negative, and that asymmetry is worth reading because it says which of the two is the more interesting reading.

The interaction is the joint supremum less the sequential one, and the sequential point is in the joint search’s own feasible set — so the interaction is non-negative on every draw by construction, at machine precision. On the control it comes out at 0.000467, on the containment rung at exactly zero, and it is never below.

The overlap has no such guarantee. It is δA + δB less the sequential supremum, and there is no containment relation that forces a sign: a first search that has run can leave the second search more to find, and on the break-and-column pair it does.

That means a pair reading zero on the excess can be zero in two structurally different ways. Either both components are near zero, or a positive interaction is cancelled by a positive overlap of the same size — and there is no third option, because a negative interaction is impossible.

The control is the second. Both of its components are positive, both are near half a thousandth, and their difference is nothing.

A control that reads zero for two reasons

There is a general form of this and it is the reason the essay is here rather than in a footnote.

A control is chosen so that a measurement on it should read a known value, usually zero, and passing that check is taken as evidence the machinery is right. What the check actually establishes is weaker: the measurement reads the known value. It does not establish that it reads it for the reason the control was built to test.

Here the control was built to test that two searches with nothing between them are additive, and it reads additive. What it was not built to test — and cannot test — is whether the additivity is an absence or a balance. The two are indistinguishable from the outside, they behave identically under every sweep the earlier field ran, and they come apart only under a measurement that was not being made.

The way they were told apart is not clever. It is that a fourth supremum was computed, and the fourth supremum is a quantity the three did not contain.

The same reading on the other rungs

It is worth applying the same arithmetic up the ladder, because the control is not the only rung whose number is small relative to what it is made of.

The two step dictionaries read an excess of 0.014075 with a standard error of 0.007931 — under two standard errors from zero, so on its own it is a rung that could be reported as “about additive”. Split, it is an overlap of 0.044030 against an interaction of 0.029955: two searches over step columns cut a few rows apart share a great deal, and a joint search over both reaches a great deal that neither slice contains, and the two nearly cancel.

That is the rung the earlier field placed at 0.125 on its scale-free reading, describing it as the step between sharing nothing and one search containing the other. The description is right and the mechanism underneath it is two large effects rather than one middling one.

The break-and-column pair reads −0.006759 ± 0.001025, which is six standard errors below zero and is the earlier field’s most surprising rung: charging the two searches separately under-charges. Split, its interaction is 0.002364 and its overlap is −0.004395 — so both components point the same way, and the negative excess is not a cancellation but a sum.

Two rungs made of cancellations, two rungs made of one dominant effect, and one rung made of two effects reinforcing. The excess column orders them correctly and says nothing about which is which.

Whether it matters

It is fair to ask what breaks if a reader takes the control at face value, and the answer is: not much, and one thing.

Not much, because the ladder’s zero is a zero either way. A pair reading 0.000116 is additive for any purpose anybody would use the ladder for, and the rungs above it are ordered correctly on the quantity being reported.

One thing, and it is about extrapolation. If the control’s zero were an absence, then a pair of searches with nothing between them would be additive at any dictionary size, any sample size and any law — because there would be nothing there to grow. If it is a balance, then it is additive as long as the two effects stay in proportion, and nothing here says they must.

The four dictionary sizes above are the evidence that they do stay in proportion over the range measured, and the ratio between them runs from 0.61 to 0.92 across that range rather than staying fixed. A balance that drifts by half is a balance, not a law, and a reader extrapolating the control to a setting where each search can find much more should expect it to fail eventually and has no way from these measurements to say where.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Basis functionsBenchmark forecastClosed formCombinatorial searchData snoopingDependenceIndependenceLikelihood ratioMonte CarloOrthogonalityOverfittingSelection effectSpecification searchSupremum statisticVariance decomposition