What a dictionary buys and what it costs

What the extra function buys

A rule balancing the mean of each covariate has a worst case of exactly zero. Adding the median split — the other thing every trial balances — leaves it at exactly zero, and one square moves it.

Worth reading first: A design is a number · The variance removed before the data.

A balancing rule is chosen before the outcome exists, so what it is worth is what it removes of the shape it is worst against. That is a maximin reading, it is the one the basis field takes, and here it has an answer that a designer can act on in one sentence.

Seven shapes an outcome might have: linear in one covariate, its square, a median split of it, a threshold one standard deviation out, the product of the covariates, the product of two median splits, and one covariate times the other’s square. Five dictionaries, in the order a trial would build them up. Every entry closed form at a correlation of a half.

The worst case of the dictionary everybody holds is exactly zero

What the worst case is worth, one function at a timeThe smallest share each dictionary removes, over seven outcome shapes, at a correlation of 0.5. A rule balancing the mean of each covariate has a worst case of exactly zero — against the square, and against both products. Adding a median split to it, which is the second thing every trial balances, leaves the worst case at exactly zero, because a median split is odd and so is a mean. Adding the square instead moves it to 6.8%, and the extra functions after that move it to 7.4%. The worst case is decided by which parities the dictionary contains rather than by how many functions are in it.the smallest share removed, over seven outcome shapes — larger is bettera correlation of 0.5, both covariates constrained the same waythe mean — 2 functionsexactly 0the mean and a median split — 4 functionsexactly 0the mean and the square — 4 functions6.8%the mean, the square and a median split — 6 functions6.8%and a cut at one as well — 8 functions7.4%closed form at ρ = 0.5, seven shapesparity, not count
Fig. 1 The smallest share each dictionary removes, over the seven shapes, with the correlation on a slider.

A rule balancing the mean of each covariate removes all of a linear outcome, 63.7% of a median split and 43.9% of a threshold at one — and exactly nothing of the square, of the product of the covariates, or of the product of the two median splits. Its worst case is zero.

A rule balancing the mean and the median split of each covariate removes all of a linear outcome, all of a median split, 46.3% of a threshold, and exactly nothing of the same three shapes. Four constraints instead of two, and the worst case is still exactly zero.

A rule balancing the mean and the square removes 6.8% of the shape it is worst against, which is the product of the two median splits. That is a small number and it is the first non-zero one in the table.

The reason is parity: a mean and a median split are both odd, an interaction between two odd functions is even, and an odd function is orthogonal to an even one at every correlation. Adding a second odd function to a dictionary that is already odd cannot touch the shapes the first one could not touch, however many of them are added.

Which is not the same as saying it is useless

The zero is a statement about the worst case and the rest of the table is not zero.

Every dictionary against every shape. What each of five dictionaries removes of each of seven outcome shapes, at a correlation of 0.5. Reading across, the shapes are: linear in one covariate, its square, a median split of it, a threshold at one standard deviation, the product of the covariates, the product of two median splits, and one covariate times the other's square. A dictionary removes a shape entirely when the shape is in its span — the flat runs at the top — and removes exactly none of a shape of the opposite parity, which is the floor two of these curves sit on. The mean alone removes 63.7% of a median split, because a cut and a covariate are correlated even though neither is in the other's span.
Fig. 2 Every dictionary against every shape. The flat runs at the top are shapes inside the span; the floors are parity.

Adding the median split to the mean takes the median-split shape from 63.7% removed to 100%, and the mixed interaction — one covariate times the other’s square — from 25.0% to 38.3%. Both are real gains and neither touches the worst case. A designer who believes the outcome is smooth in the covariates has bought something; a designer who is protecting against the shape they have not thought of has bought nothing.

That is the distinction the maximin reading exists to make, and it cuts against the usual instinct. The usual instinct is that more constraints are more protection, and the table says more constraints are more protection against the shapes they span and their relatives, with a floor that does not move until the parity of the dictionary changes.

The cheapest repair, and what it costs

One square of one covariate is enough to move the floor, and the square is available at no cost in data: it is a function of a number the trial already recorded.

What it costs is admissible assignments. A constraint is a constraint, and the set the rule has left shrinks with every function added — at two hundred units and a tolerance of one, going from two constraints to four takes the admitted share from 45.7% to 19.8%, and from four to six takes it to 7.8%. So the honest comparison is not does this extra function help but does it help more than another function of the same cost.

On that comparison parity decides it, and the arithmetic is worth doing rather than asserting: two extra constraints per covariate cost the same admissible fraction whichever functions they are, so the comparison is between a floor of 6.8% and a floor of 0.0% at an identical price. Adding the two squares moves the worst case from 0 to 6.8% and costs the same two constraints per covariate as adding the two median splits, which moves it from 0 to 0. There is no reading of the table on which the median splits win, and there is a reading on which they are the natural choice: they are what a trial’s covariate is usually reported as.

The cost is not only in assignments admitted. Past a certain thinness the sampling method has to change — a walk over the admissible set replaces a hunt for admissible assignments — and where that happens is a fact about the share admitted rather than about which functions produced it. So the price of an extra constraint is paid in the same currency whatever its parity, which is what makes the comparison above a fair one rather than an accounting trick.

Three of the seven shapes are out of reach before anything is chosen

Counting the seven shapes by parity says how much of the maximin problem is decided before a single dictionary is compared.

Odd, and therefore reachable by an odd dictionary: the linear term, the median split, and one covariate times the other’s square — which is odd in the first factor and even in the second, so odd under the joint flip.

Even, and therefore unreachable by any odd dictionary however large: the square, the product of the covariates, and the product of the two median splits.

Neither: the threshold at one standard deviation, which is a mixture of the two parities.

So three of the seven — 43% of the set the worst case is taken over — carry a guarantee of exactly zero from any dictionary made of odd functions, and adding odd functions cannot change which three. The maximin is being taken over a set nearly half of which is inaccessible in principle.

Adding functions of the parity you already have. Four dictionaries against the interaction each is aimed at. Two covariates remove exactly nothing of their own product, at every correlation — the result usually stated for the independent case, and it holds at all of them. Two median splits remove exactly nothing of theirs, which is the same statement. A covariate and its own median split removes exactly nothing either, which is what makes the pair of results one result: all three dictionaries are odd, the interactions are even, and an odd function is orthogonal to an even one whatever ρ is. Two squares remove 64.00% at a correlation of a half, and the squares are the first even functions in the list.
Fig. 3 Four dictionaries against the interaction each is aimed at. Two covariates remove exactly nothing of their own product at every correlation, two median splits remove exactly nothing of theirs, and a covariate with its own median split removes exactly nothing either — all three dictionaries are odd and the interactions are even.

What the second odd function buys on the one shape it can touch

The mean rule removes 43.9% of the threshold and the mean-and-split rule removes 46.3%. Two extra constraints, 2.4 points.

The reason that is so small is the same parity accounting, applied to the one mixed shape. The threshold at one standard deviation is 59.4% odd, so an odd dictionary can remove at most 59.4% of it, and that ceiling holds however many odd functions are added.

Against that ceiling, the mean alone reaches 43.9/59.4 = 74% of what is available, and the mean and split together reach 78%.

The second function is not buying little because it is a poor function. It is buying little because the first one already took three quarters of everything an odd dictionary is ever allowed to have on that shape, and the remaining quarter is all that four constraints can compete for.

Which puts the 6.8% in the mean-and-square row in its proper light. That is not a small improvement on zero; it is the first entry in the table that is on the other side of a wall, bought by adding a function of the opposite parity rather than by adding another function.

The floor rises slowly after that

The last two rows are worth reading together. A dictionary holding the mean, the square and the median split of each covariate — six constraints — has a worst case of 6.8%, the same as the mean and square alone. Adding a cut at one standard deviation to make eight takes it to 7.4%.

So the floor moves from 0 to 6.8% with the first even function and then barely at all. The shape that holds it down is the product of the two median splits, which is even, and the only even functions available against it are the squares and the even parts of off-median cuts — and a square is a poor proxy for a product of two indicators.

That is a statement about this shape set rather than about dictionaries in general, and it is the one place this essay’s answer depends on the list of shapes chosen. Drop the product of the median splits from the list and the worst case becomes the ordinary product, which the squares handle at 64.0%; keep it and the floor stays under a tenth however much is added.

Reading the table down a column

The rows say what a dictionary is worth. The columns say when a shape first becomes protected, and they are more useful to somebody who has a guess about the outcome.

A linear outcome is removed entirely by every dictionary here, because every one of them contains the mean.

A square is removed entirely by every dictionary that contains a square and exactly none by every one that does not. There is no partial credit: the mean and the median split of both covariates, four constraints, remove 0.0% of it.

A median split is removed entirely once the dictionary holds one, and 63.7% by a rule holding only the means — which is the most interesting entry in the table, because a mean and a median split are different functions and neither is in the other’s span. What connects them is that they are correlated: a covariate and its own sign share most of their variation, and balancing one balances most of the other by accident.

A threshold one standard deviation out is the shape that separates the last two dictionaries: 43.9%, 46.3%, 65.8%, 68.2%, and then 100% once a cut at one is actually held. Nothing short of holding it reaches it, and holding the median split instead adds two and a half points.

The product of the covariates is 0.0% for the odd dictionaries and 64.0% for every dictionary containing the squares, with no intermediate value anywhere.

The product of two median splits is the shape that holds the floor down: 0.0%, 0.0%, 6.8%, 6.8%, 7.4%. Even the widest dictionary here removes under a tenth of it.

And one covariate times the other’s square — the one odd interaction — runs 25.0%, 38.3%, 25.0%, 38.3%, 39.2%. It rises with the median splits and not with the squares, which is parity again read from the other side: an odd interaction is reached by odd functions.

The column pattern is the row pattern seen edge-on. Every entry in this table is either a hundred per cent, an exact zero, or a number that depends on how much two functions of the same parity happen to share.

How the floor moves with the correlation

The zeros are exact at every correlation, because parity does not depend on ρ. What moves is everything else, and it moves in the direction that makes the correlated case the interesting one.

At ρ = 0.2 the best dictionary here has a worst case of 1.57%. At ρ = 0.5 it is 7.35%, and at ρ = 0.8 it is 10.54%. So the protection a dictionary supplies against the shape it is worst against grows with the correlation, which is the opposite of the usual reading of a correlated design as a harder one.

The mechanism is the same one that destroyed the interaction guarantee in the first place: at a correlation an interaction acquires a component along the main effects, and a rule holding those removes it. A design whose covariates are strongly correlated is a design in which balancing a few things balances a great many others by accident.

None of which helps the dictionaries whose worst case is zero. At every one of the three correlations the mean alone and the mean-with-median-split have a worst case of exactly 0.00%, because parity is not a matter of degree.

What a designer should take from it

Three things, in the order they matter.

Check the parity of the dictionary before counting its entries. It is the one property of a balancing rule that produces an exact zero rather than a small number, and an exact zero is the only kind of gap that no sample size closes. It takes one line — average a function against its own reflection — and it decides an exact zero rather than a small number. A dictionary that is entirely odd has a worst case of exactly zero against interactions, and no number of additional odd functions changes that.

Hold at least one even function of each covariate. The square is the obvious one and it is free.

And do not read the worst case as the expected case, which is the same warning the essay on what a rule removes makes about a guarantee. The mean-and-median-split rule removes all of two of these seven shapes and a substantial share of two more; its worst case is zero because one shape in seven is untouchable by it, not because it does nothing. A trial with a reason to believe the outcome is monotone in the covariates is buying the protection it needs and is entitled to say so — provided the reason is stated, which is what turns a worst case into a stated assumption rather than an unstated one.

The rule that is not in the table

There is a dictionary this table does not contain and a trial might: the interaction itself. Nothing stops a rerandomisation from balancing the product of two covariates directly — it is a number computed from the same recorded columns as everything else.

It is left out because it changes the question rather than answering it. A rule handed the interaction removes all of that interaction and nothing of the next one, and the table would then be about which interactions were guessed rather than about what a dictionary of main effects is worth. The maximin over a list of shapes is interesting precisely when the list is longer than the dictionary, and a dictionary that can be extended to contain any named shape has a worst case decided by whichever shape was not named.

What the parity rule says about that choice is still worth having. A product of two odd functions is even, so holding it adds an even function to the dictionary — and by the argument above, the first even function is where the floor moves. A trial that balances one interaction has, incidentally, bought some protection against every other even shape, in a way that balancing another main effect would not.

That is the useful form of the whole essay: the parities in a dictionary matter more than its length, and a designer choosing what to add should be choosing a parity first and a function second.

What is claimed here, and what is not

This essay takes what a balancing dictionary is worth against the shape it is worst against. The claims are that a rule balancing the mean of each covariate has a worst case of exactly zero over these seven shapes; that adding the median split leaves it at exactly zero while taking two other shapes to 100% and 38.3%; that adding the square instead moves it to 6.8%; that six and eight constraints reach 6.8% and 7.4%; and that the ordering is decided by which parities the dictionary contains rather than by how many functions it holds.

What stays out and is named as a decision: the shape set. A maximin is a minimum over a list, and the list here is seven shapes chosen to span the cases this collection has closed forms for. It contains two even interactions and one odd one, and a list with a different balance of parities would give a different floor. That is not a defect of the method — it is the thing choosing which shapes are worth protecting is about — but a floor quoted without its list is not a number.

Also out: the correlation. Everything above is at a correlation of a half, and the slider on the first figure moves it. The exact zeros are exact at every correlation, because parity does not depend on ρ; the non-zero entries all shrink towards zero as the covariates become independent, which is the case where a rule protects against interactions least and needs to least.

The checks, and the refusal

Two claims are gated. The worst case of an all-odd dictionary is required to be exactly zero rather than small, at machine precision, because the whole argument is that it is a zero and not a rounding. And adding one even function is required to move it, which is the assertion that the first claim is about parity rather than about these particular functions.

The refusal is the reading this essay exists to prevent. A dictionary credited with protection because it holds more functions is rejected: four constraints of one parity have the same worst case as two, exactly, and a design calculation that counts constraints rather than reading them is counting the wrong thing.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Basis functionsClosed formCorrelationCovariate adjustmentCovariate balanceDesign criterionExperimental designInteractionMaximin designOrthogonalityProjectionRerandomisationRobustnessVariance explained