Choosing what the rule reads

Which shapes are worth protecting

Choosing a basis by its worst case is a finite problem with an exact answer. The answer has no tie in it, which a maximin optimum is supposed to have — and the tie comes back, along with twice the guarantee, when the basis is drawn rather than chosen.

Worth reading first: The variance removed before the data · A design is a number.

An experimenter has to hand a balancing rule some functions of the covariate before the trial starts, and every choice has a shape it is helpless against. The obvious way to choose is by the worst case, which is what a design field two subjects away does with parameter values, and it is worth doing here because the arithmetic is exact and the answer is not what anybody would guess.

The problem is finite. There are eight functions in the dictionary and six shapes in the list; a basis of k functions is a k-subset; every cell of the table is a closed form. Scoring all of them and keeping the row with the best worst cell is an exact answer to an exactly stated question, and it takes no simulation and no search.

The guarantee, as the basis is allowed more functions. The lower line is the best worst case over the six named shapes for a basis of each size, found by scoring every subset of the dictionary — an exact answer, since the problem is finite. One function guarantees 2.3%, which is nearly nothing; three guarantee 59.0% and the basis that does it is the covariate, its square and its cube, with no indicator in it. The upper line is the same problem with the basis drawn rather than fixed, which is worth 2.09 times as much at two functions and 1.32 at three. The two lines converge because a basis large enough to protect everything has nothing left to randomise over.
Fig. 1 The guarantee as the rule is allowed more functions, with the basis that attains each point named on it. The upper line is the same problem with the basis drawn rather than fixed.

The answers, and the first surprise

One function guarantees 2.3%. The best single function to hand the rule is an indicator at 2 — not the covariate. Every polynomial in the dictionary is exactly orthogonal to some shape in the list and therefore has a worst case of zero; an indicator has a component on every Hermite order and so has none. The best of them still guarantees almost nothing, at 0.0233.

Two guarantee 26.8%, and the basis is two indicators — at the median and at 2 — with no polynomial in it at all.

Three guarantee 59.0%, and the basis flips entirely: the covariate, its square and its cube. Once three polynomial orders are available they cover the low-order content of every shape in the list, including the two thresholds, better than any mixture of cuts does.

Four guarantee 69.8%, and the basis is those three polynomials plus an indicator at 2 — the shape in the list that polynomials reach last.

The composition of the winning basis is not monotone: the best pair contains no polynomial and the best triple contains nothing else. That is a straightforward consequence of a maximin criterion — a row is judged by its weakest cell, and the weakest cell moves as the basis grows — and it means there is no such thing as adding the next-best function.

Why every polynomial scores zero, and what is left

The first answer has a structural reason behind it that is worth separating from the measurement, because it cuts the search down before any subset is scored.

The dictionary’s polynomials are Hermite functions, which are orthogonal to one another. So a polynomial of one order is exactly orthogonal to a shape that is a polynomial of another order, and removes precisely none of it. Every polynomial in the dictionary therefore has a worst case of exactly zero, whichever it is and however good it is against the other five shapes.

That leaves the indicators, which are the only entries with a component on every Hermite order at once and therefore the only ones that can have a non-zero worst case at all. The maximin search at one function is not a search over eight candidates; it is a search over the indicators, and the polynomials are eliminated by an argument rather than by a score.

What the best single function is, read as a rule

The winner is an indicator at 2, which is an indicator for the top 1Φ(2)=2.3%1 - \Phi(2) = 2.3\% of the covariate — roughly one unit in forty-four.

That is a strange thing for a maximin to choose and it follows from the same argument. A cut far out in the tail has the flattest spread across Hermite orders, so it is the function least orthogonal to anything, and being least orthogonal to everything is exactly what a worst case rewards.

It also guarantees almost nothing: 0.0233, a fortieth of the shape it is worst against. And the second function takes that to 0.268 — a factor of 11.5 for one more constraint, which is the steepest step anywhere on the curve.

So the honest reading of the first two points is that one function is not a balancing rule at all, and that the maximin’s advice at k = 1 is best understood as a statement about what is impossible rather than about what to do.

The tie that is not there

The design field’s maximin has a property that makes an answer feel like an answer: the worst case is attained at more than one point. A maximin design over a rectangle of parameter values sits on a tie at three of the four corners, and that equalisation is what a least favourable prior would be supported on.

Nothing like it happens here. At every basis size the optimum’s worst cell is attained at exactly one shape, and the second-worst is some way above it. The best pair guarantees 26.8% against a quadratic and 27.2% against a threshold at 1 — close, but not equal — and 100% against two of the six.

The reason is that the choice is discrete. Equalisation is a first-order condition: at an interior optimum of a continuous problem, if the worst case were attained at one point only, the design could be moved a little towards that point and the worst case would improve. There is no “a little” here. The bases are a finite list and the optimum is simply the best row.

That is worth recording because the tie has been used as a check elsewhere on this site — the design field verifies its maximin answer by confirming the equalisation rather than by a second search, since its second route would not converge. Here the exhaustive enumeration is exact, so no second route is needed; and the property that would have been the check is absent, for a reason that has nothing to do with the answer being wrong.

What the insurance costs, shape by shape. The bars are what the maximin basis of 2 functions removes of each shape; the marks beside them are what the best basis for that shape alone removes. The gap is the premium, and it is paid in every column: the maximin basis is nobody's best basis. Against quadratic it gives up the most — 73.2% of the covariate's imbalance — and in exchange the row has no cell below 26.8%, where the best basis for any single shape has one at 0.0%.
Fig. 2 What the maximin basis of two functions removes of each shape, against what the best basis for that shape alone removes. The premium is paid in every column, and it is not paid evenly.

Where the tie comes back

The equalisation is not gone, it is in a different problem, and the different problem turns out to be a design an experimenter can actually run.

Suppose the basis is drawn rather than fixed: a stated randomisation, decided before the trial, that picks one of several bases. The guarantee is then the worst shape’s average removal over the draw, and the problem is a finite two-player game — the experimenter picking a distribution over bases, the world picking a shape.

Solved by fictitious play, at two functions the value is 0.5617 against the best fixed pair’s 0.2685. Drawing the basis is worth 2.09 times the best thing anybody can commit to.

And the equalisation comes with it. The randomised basis removes 0.5674, 0.5617, 0.5623, 0.5618, 0.5619 and 0.5617 of the six shapes: five of them are equal to four decimal places and the sixth is a whisker above. That is the tie, arriving exactly when the problem becomes convex.

The basis that is drawn rather than chosen. Choosing one basis of 2 functions and standing by it guarantees 26.8% against the worst of the six shapes. Drawing the basis from these weights before the trial guarantees 56.2% — 2.09 times as much — because a shape that defeats one basis is no use against a basis nobody has committed to. The randomised design also has the property the fixed one lacks: it protects every shape by nearly the same amount (0.567, 0.562, 0.562, 0.562, 0.562, 0.562), which is the equalisation a maximin optimum is supposed to have and which the discrete choice destroys.
Fig. 3 The weights the randomised basis of two functions is drawn from, and what it guarantees. Four bases carry nearly all the weight and none of them is the fixed optimum by much.

What a randomised basis actually is

It is worth being concrete, because “randomise the design” is a phrase with several meanings and only one of them is this.

Before any unit arrives, the experimenter runs a stated randomisation — a published table of probabilities — and it returns a basis: the median cut and the cut at 2 with probability 0.290, the cube and the cut at 1 with 0.281, the covariate and its square with 0.219, the square and the cube with 0.172, and two more with the remaining weight. That basis is then fixed for the whole trial and the balancing rule reads it.

Nothing about the analysis changes. Nothing about the trial’s conduct changes. The unit-level randomisation is exactly what it was. The only thing that has been randomised is a decision the experimenter would otherwise have made by judgement, and the judgement was the thing with a worst case of 26.8% in it.

Why it works is the ordinary minimax reason: a shape that defeats one basis is no use against a basis nobody has committed to. The adversary in the fixed problem gets to see the choice; in the randomised problem it does not, and the value of the game rises accordingly.

The weights the world puts on the shapes are the other half of the solution, and they say which shapes are actually binding: quadratic 0.219, threshold at 2 0.200, threshold at 1 0.194, cubic 0.193, median 0.158 and linear 0.036. The linear shape carries almost nothing — every basis in the support handles it well enough that the world gains nothing by choosing it — which is a precise version of the observation that the covariate itself is the shape nobody needs protecting against.

The basis that is drawn rather than chosen. Choosing one basis of 1 functions and standing by it guarantees 2.3% against the worst of the six shapes. Drawing the basis from these weights before the trial guarantees 28.1% — 12.09 times as much — because a shape that defeats one basis is no use against a basis nobody has committed to. The randomised design also has the property the fixed one lacks: it protects every shape by nearly the same amount (0.285, 0.283, 0.282, 0.283, 0.282, 0.281), which is the equalisation a maximin optimum is supposed to have and which the discrete choice destroys.
Fig. 4 At one function the same arithmetic is dramatic: drawing a single function guarantees 28.1% where the best fixed single function guarantees 2.3%.

The result that ought to be uncomfortable

At one function the randomised design guarantees 0.2813 and the gain over the fixed choice is 12.09 times. That is a large number and it produces a comparison worth stating on its own:

Drawing one function at random protects more than any fixed pair. 0.2813 against 0.2685. An experimenter who cannot afford to balance two functions, and randomises which single one to balance, ends up with a better guarantee than an experimenter who balances two and picks them as well as anybody can.

At three functions the gain falls to 1.32 times — 0.7767 against 0.5900 — because a large basis covers most of the list on its own and there is less for the randomisation to hide. The two lines converge, and the place where randomising is worth the most is exactly the place where an experimenter has the least to spend.

The basis that is drawn rather than chosenChoosing one basis of 3 functions and standing by it guarantees 59.0% against the worst of the six shapes. Drawing the basis from these weights before the trial guarantees 77.7% — 1.32 times as much — because a shape that defeats one basis is no use against a basis nobody has committed to. The randomised design also has the property the fixed one lacks: it protects every shape by nearly the same amount (0.861, 0.777, 0.777, 0.777, 0.777, 0.777), which is the equalisation a maximin optimum is supposed to have and which the discrete choice destroys.x + x² + x³0.5441{x>0} + 1{x>1} + 1{x>2}0.272x³ + 1{x>1} + 1{x>2}0.103x² + 1{x>0} + 1{x>2}0.062x² + 1{x>1} + 1{x>2}0.018how often each basis is drawnguarantee 77.7% against 59.0% fixeda finite game on 56 bases and 6 shapesfictitious play, and the value is attained at 5 shapes at once
Fig. 5 At three functions the fixed polynomial basis takes 54% of the weight and the gain has fallen to a third. Drag the basis size: randomisation buys most where the budget is smallest.

Why fictitious play, and how far to trust it

The randomised design’s value is the value of a finite two-player zero-sum game, and there are several ways to compute one. The one used here is the oldest: repeatedly let each side best-respond to the history of the other, and average.

It is worth saying why that is trustworthy here when the design field’s second route was not. There, the outer problem was a search over continuous designs with an inner optimisation at every evaluation, and two implementations converged to worst-case values well above the direct search’s — which meant the check was reporting a disagreement between two routes when what it had found was its own optimiser failing. Here the game is a matrix: twenty-eight rows by six columns at two functions, every entry a closed form, no inner search anywhere. Fictitious play on a finite zero-sum game converges to the value, and the arithmetic per step is a pair of dot products.

The evidence that it has converged is not the iteration count but the equalisation: the value is attained at five of the six shapes to four decimal places, which is a property of the solution rather than of the solver. A run that had not converged would show one shape well below the others, because the world’s best response would still be finding somewhere to go.

And the answer has an upper bound that agrees. The maximum of the mixture-weighted rows, evaluated at the converged shape weights, is the same number as the minimum of the shape values at the converged basis weights — the two sides of a saddle point, computed from opposite ends of the same table.

The basis that is drawn rather than chosen. Choosing one basis of 2 functions and standing by it guarantees 26.8% against the worst of the six shapes. Drawing the basis from these weights before the trial guarantees 56.2% — 2.09 times as much — because a shape that defeats one basis is no use against a basis nobody has committed to. The randomised design also has the property the fixed one lacks: it protects every shape by nearly the same amount (0.567, 0.562, 0.562, 0.562, 0.562, 0.562), which is the equalisation a maximin optimum is supposed to have and which the discrete choice destroys.
Fig. 6 The support again, with the six shape weights underneath the argument: the linear shape carries almost nothing, because every basis in the support handles it.

What the insurance costs, column by column

A guarantee is only interesting beside what it gives up, and the maximin basis is nobody’s best basis by construction.

At three functions the fixed maximin basis — the covariate, its square and its cube — removes 1.0000 of a linear outcome, 1.0000 of a quadratic and 1.0000 of a cubic, because it spans all three exactly. Against the threshold at 1 it removes 0.6579, against the median split 0.7427 and against the threshold at 2 0.5900. The best basis for any single one of those shapes removes all of it, since the dictionary contains each shape itself.

So the premium is between a quarter and four tenths of the covariate’s imbalance, paid on the three shapes the polynomial basis reaches least well, in exchange for a floor at 0.5900 instead of a floor at zero.

That is a much better trade than the corresponding one in the design field, where protecting a fourfold rectangle of parameter values costs about thirty points of efficiency at the guess. The reason is the currency, and it is the same reason the earlier balancing field found its insurance nearly free: protection bought with runs is expensive, and protection bought with assignments costs only the degree to which the constraints compete for the same units.

What the insurance costs, shape by shape. The bars are what the maximin basis of 3 functions removes of each shape; the marks beside them are what the best basis for that shape alone removes. The gap is the premium, and it is paid in every column: the maximin basis is nobody's best basis. Against cut at 2 it gives up the most — 41.0% of the covariate's imbalance — and in exchange the row has no cell below 59.0%, where the best basis for any single shape has one at 59.0%.
Fig. 7 The premium at three functions. Three of the six shapes are removed entirely and the other three carry the whole cost.

What an experimenter would put in a protocol

The four things that would have to be written down before the first unit arrives are short, and none of them is a statistical technique.

The shapes. A list, with a sentence each saying why that shape is plausible for this outcome. Six is a workable number; the arithmetic does not care and the argument does.

The dictionary. What the rule is allowed to be handed. This is not the same list — the dictionary here contains four polynomial orders and four cut points, and the shape list contains six things, and neither is a subset of the other.

The budget. How many functions the rule reads, which is the axis of the first figure and is set by how many constraints the assignment can carry rather than by anything about shapes.

And whether the basis is drawn. If it is, the table of probabilities goes in the protocol along with the seed, exactly as the unit-level randomisation does.

That last line is the one that would cause an argument, and it is worth noting that it causes the argument for a good reason: a drawn basis means two experimenters running the identical protocol on the identical population balance different functions, and their designs are then not comparable in the way two fixed-basis trials are. What they share is the guarantee, which is the thing that was being bought.

What none of this decides

The list. Every number in this essay is conditional on the six shapes, and moving one of them moves every answer.

That is not a limitation to be apologised for; it is the substance. A maximin guarantee is a statement of the form “against anything in this set, at least this much”, and the set is the part that a statistician cannot supply. What the arithmetic can do is make the consequences of a set visible — the best basis for it, what that basis costs, and how much randomising the choice is worth — so that somebody who knows the subject matter can argue about the right thing.

And the alternative to naming a set is not neutrality. It is a guarantee of zero, which is the next essay entire.

What is claimed here, and what is not

This essay takes the basis chosen by its worst case, and the claims are the maximin values and bases at four sizes, the absence of a tie at the discrete optimum, the value and support of the randomised basis, and the premium the fixed maximin pays column by column.

What stays out and is named as a decision: any weighting of the shapes by how plausible they are, which turns a maximin into a Bayes problem with a different answer and a different argument; the choice of dictionary, which bounds every number here from above; and whether a referee would accept a randomised basis, which is a question about practice rather than about arithmetic and on which this site has no evidence.

The boundary against the design field is the currency and the geometry. That a worst case is the right thing to optimise, that its optimum sits on a tie in a continuous problem, and that a least favourable prior supports it, are established there. What is new is that a discrete choice loses the tie, that randomising the choice recovers it, and that here the randomisation is something an experimenter can actually perform.

The checks, and the refusals that make them mean something

One claim is gated in this field’s library and it has three parts, because the argument fails if any of them is wrong. The best fixed basis of two functions is required to have its worst case at exactly one shape — no tie. The randomised basis is required to equalise, with every shape it protects within a fixed distance of every other. And the randomised guarantee is required to be more than twice the fixed one, which is what makes the equalisation worth reporting rather than a curiosity.

The refusal is a basis reported at the shape it is best against. Every row in the table has a cell at 1.0000 and a much smaller worst cell, so a design described by its best column is a design described by the assumption it was built on — and the whole of this essay is about the other end of the row.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Allocation ruleBasis functionsCovariate balanceDesign criterionEfficiencyEqualisationLeast favourable priorMaximin designModel misspecificationProjectionRandomisationRobustnessThresholdVariance reduction